mod funnel_bo¶
- module funnel_bo¶
First-derivative interface for HMC-style samplers. Choosing which morphology to search next, by expected improvement.
The searches here fail by funnel, not by local optimization. At 75 points every failing run reaches the icosahedral plateau and stays; the relaxation is fine, the descent is fine, and the answer is in a region the chain never visits. Eight mechanisms built to make the chain leave a funnel were measured and none helped, because leaving is not the problem: the chain has nowhere better to be told to go.
That is a decision problem over a small space, and it is the shape Bayesian optimization is for. Not over the coordinates, where three hundred dimensions puts a Gaussian process out of reach, but over a structural descriptor: the share of points in each local environment, five numbers from
crate::structure::ptm_fractions. A model of “how low does this morphology go” over those five numbers is cheap to fit and is exactly the surface the search is blind to.What is modelled
For each distinct morphology the search has visited, the lowest energy found there. A Gaussian process over that gives a mean and a variance everywhere, including at morphologies never visited, and expected improvement turns the pair into a single number: how much better than the incumbent this morphology is likely to be, integrated over the model’s own uncertainty.
The point is what expected improvement does with a region that has never been sampled. Its mean reverts to the prior and its variance is large, so it scores highly; a region sampled repeatedly and found mediocre scores low however close it is to the incumbent. That is the opposite of what a bias does, and it is the missing half: a bias says where not to go, this says where to go instead.
Why a Gaussian process and not something larger
The observation count is the number of distinct morphologies, which is hundreds rather than millions, and the dimension is five. Exact inference is a Cholesky factorisation of a few-hundred-square matrix, done once per refit. Nothing here needs approximating.
Structs and Unions
- struct FunnelCompression¶
Auditable covariance loss from a pivoted-Cholesky kernel compression.
- input_count: usize¶
Number of observations offered to this compression step.
- retained_rank: usize¶
Number of kernel pivots retained by the GP.
- residual_fraction: f64¶
Remaining prior-covariance trace divided by the original trace.
- rank_limited: bool¶
Whether the rank ceiling was reached above the numerical covariance floor.
- struct FunnelModel¶
A Gaussian process over a structural descriptor.
Squared-exponential kernel with a single length scale, which is right when the inputs are fractions on a common scale, and a noise term that also keeps the factorisation well conditioned when two morphologies are nearly equal.
- length_scale: f64¶
Kernel length scale, in units of the descriptor.
- amplitude: f64¶
Signal standard deviation.
- noise: f64¶
Observation noise standard deviation.
Implementations
- impl FunnelModel¶
Functions
- fn assign_by_mean(&mut self, candidates: &[(usize, Vec<f64>)], q: usize) -> Vec<usize>¶
qcopies of the lowest posterior mean. The greedy lie q-EI replaces.
- fn assign_gibbon(&mut self, candidates: &[(usize, Vec<f64>)], q: usize, max_family_size: usize, minimum_samples: usize) -> Vec<usize>¶
Greedily fill a batch by GIBBON under a hard per-source family cap.
The candidate table is a finite search domain. Posterior minimum samples are joint Gaussian draws over that domain, so correlated descriptors are not treated as independent opportunities. Greedy filling is the standard submodular approximation to the corresponding determinantal MAP problem.
- fn assign_q_ei(&mut self, candidates: &[(usize, Vec<f64>)], q: usize) -> Vec<usize>¶
Rank-and-cycle q-EI on a discrete family table.
Independent argmax EI repeats the same family
qtimes. Rank the table by Jones EI and cycle without replacement so a WAVE spreads across families that still have remaining improvement.
- fn compress(&mut self, maximum_rank: usize) -> FunnelCompression¶
Bound the retained kernel rank by pivoted Cholesky.
The lowest observed response is always a pivot. Remaining pivots greedily maximize conditional prior variance until either the covariance floor or
maximum_rankis reached. Exact caller-owned observations can therefore remain outside this numerical surrogate.
- fn expected_improvement(&mut self, x: ArrayView1<f64>) -> f64¶
Expected improvement at a morphology, for a minimisation.
Zero where the model is confident nothing better lives, and large where the mean is low or the uncertainty is high. The second half is what makes this different from following the model’s mean: a morphology never visited scores on its variance alone, which is how a search reaches a funnel it has no evidence about.
- fn gibbon_batch(&mut self, batch: &[ArrayView1<f64>], minima: &[f64]) -> f64¶
Batch GIBBON acquisition for prospective descriptor-space evaluations.
The singleton information terms target the unknown minimum value. The half log-determinant of the predictive correlation matrix discounts redundant launches without an independently tuned diversity weight.
- fn gibbon_information(&mut self, x: ArrayView1<f64>, minima: &[f64]) -> f64¶
GIBBON information about the minimum value from one noisy evaluation.
This is the minimization form of the closed-form lower bound in Moss et al. The signal-to-observation variance ratio accounts for the model’s observation noise;
minimacontains posterior samples of the minimum function value over the candidate set.
- fn incumbent(&self) -> Option<f64>¶
The best value seen, which is what improvement is measured against.
- fn is_empty(&self) -> bool¶
Whether nothing has been observed.
- fn len(&self) -> usize¶
Morphologies observed.
- fn max_expected_improvement_at_data(&mut self) -> f64¶
Largest expected improvement at the morphologies already observed.
Occupancy retire uses this as the remaining-improvement bound on the seen packing codebook. Unseen families stay leftover-dwell’s job. Empty and unfitted models return infinity so they cannot look exhausted.
- fn max_value_entropy(&mut self, x: ArrayView1<f64>, minima: &[f64]) -> f64¶
Wang and Jegelka MES for a minimizer, given samples of the min value.
(I(E^star; y) = H[y] - mathbb{E}_{E^star}[H[y mid E^star]]). Jones EI is kept for retire (
Self::max_expected_improvement_at_data). Occupancy Leave ranks holes with this, not with EI. Independent marginal draws, not the GIBBON determinant bound.
- fn new(length_scale: f64, amplitude: f64, noise: f64) -> Self¶
A model over descriptors, with a length scale in descriptor units.
- fn new_euclidean(length_scale: f64, amplitude: f64, noise: f64) -> Self¶
A model that always uses Euclidean distance between descriptor vectors.
Universal multiblock descriptors are not probability simplices even when every coordinate happens to be nonnegative. This constructor keeps their block amplitudes and concatenated geometry intact.
- fn observe(&mut self, x: ArrayView1<f64>, y: f64)¶
Records the lowest energy found at a morphology, bumping
Self::versionwhen that changes the data a prediction is built from.A morphology already present is updated to the lower of the two rather than added again: the quantity modelled is how low a region goes, not how often it was sampled.
Non-finite coordinates or scores are dropped rather than stored: a NaN in the design matrix poisons every later prediction, and the caller that produced it is an exhausted budget answering with an infinity, not a landing worth keeping.
- fn predict(&mut self, x: ArrayView1<f64>) -> (f64, f64)¶
Posterior mean and standard deviation at a morphology.
- fn predictive_observation_correlation(&mut self, left: ArrayView1<f64>, right: ArrayView1<f64>) -> f64¶
Posterior correlation between two prospective noisy observations.
Independent observation noise contributes to each marginal variance, but not to the covariance between two prospective evaluations.
- fn sample_minima(&mut self, extras: &[ArrayView1<f64>], n_samples: usize) -> Vec<f64>¶
Independent-marginal samples of the posterior minimum at the observed sites plus
extras. Seeded from the book version so a ranking is reproducible without a caller rng.
- fn set_prior_mean(&mut self, mean: f64)¶
Fix the GP prior mean independently of the retained observation subset.
- fn similarity(&self, a: ArrayView1<f64>, b: ArrayView1<f64>) -> f64¶
Hellinger (or Euclidean) kernel between two descriptors.
- fn version(&self) -> u64¶
Changes to the observed data since this model was created.
Self::max_expected_improvement_at_datapredicts at every observed site and each prediction is quadratic in their number, so the sweep is cubic: at 497 observations that is of order 1e8 operations, and the occupancy coordinator asked for it once per policy request per replica per slice. Measured with perf on a live 48-replica run, FunnelModel::predict was 74 percent of the coordinator’s cycles. The version lets a caller keep the verdict while the data has not moved.