Asset sets matrix
PortfolioOptimisers.UniverseSets — Type
struct UniverseSets{__T_xkey, __T_uxkey, __T_fkey, __T_ufkey, __T_zkey, __T_dict} <: AbstractEstimatorDeclares the universes a portfolio problem is written against, and any groupings or partitions of them.
Constraint generation and the estimator routines read it to expand group references, to map a group name to its member list, and to validate membership.
It declares every axis it carries: xkey/uxkey for assets, fkey/ufkey for factors, and zkey for features. Assets are the primary axis — haskey(dict, xkey) is required, and it is the axis a view slices. The factor and feature axes are optional: requiring either would invalidate every sets object built for a problem with no factor model or no feature program, so a consumer that needs one and does not find it throws at the point of need rather than at construction.
If a key in dict starts with the same value as xkey, it means that the corresponding group must have the same length as the asset universe, dict[xkey]. This is useful for defining partitions of the asset universe, for example when using asset_sets_matrix with NestedClustered.
If a key in dict starts with the same value as uxkey, it identifies a unique-entry group variant. The corresponding xkey-prefixed group must exist in dict with the same length as the asset universe, and is used to match each asset to a unique entry from the uxkey-prefixed group. This enables constraint generation using unique entries even in NestedClustered optimisations.
The fkey/ufkey prefixes mean the same thing on the factor axis, but they buy something different. On the asset side the conventions serve views; factors are never sliced by an asset index, so on the factor side they buy length validation at construction and one shared mental model.
zkey has no prefix convention at all, and that asymmetry is the point. xkey and fkey each have a unique-entry sibling because each names an axis that partitions are written over; nothing is written over the feature axis. A graded feature program's taxonomy keys are xkey-prefixed and asset-length, and its column nodes are named directly out of the flat list dict[zkey]. So zkey carries exactly one rule — allunique(dict[zkey]), so ReturnsResult's own uniqueness check cannot be reached with a duplicate — and no length rule whatever.
A key matching none of the four prefixes is a plain group: expanded by name and axis-blind, which is why a factor group needs no machinery of its own.
Fields
xkey: Key indictidentifying the primary asset list. Required, and the axis a view slices.
uxkey: Key prefix for unique-entry asset group variants indict.
fkey: Key indictidentifying the factor list. Optional — a consumer that needs it and does not find it throws at the point of need.
ufkey: Key prefix for unique-entry factor group variants indict. Validated at construction, never recomputed by a view.
zkey: Key indictidentifying the declared feature axis — the node list a graded feature program writes its columns against. Optional, likefkey, and it carries no prefix convention: nothing is partitioned over the feature axis, so it has no unique-entry sibling and no length rule beyondallunique.
dict: Dictionary mapping group identifiers to member labels.
Constructors
UniverseSets(; xkey::AbstractString = "nx", uxkey::AbstractString = "ux", fkey::AbstractString = "nf", ufkey::AbstractString = "uf", zkey::AbstractString = "nz", dict::AbstractDict{<:AbstractString, <:Any}) -> UniverseSetsKeywords correspond to the struct's fields.
Validation
!isempty(dict).haskey(dict, xkey).- No two of
xkey,uxkey,fkey,ufkey,zkeymay be a prefix of one another (20 ordered checks, which also rules out any two being equal). - If
haskey(dict, zkey),allunique(dict[zkey]). - If a key in
dictstarts with the same value asxkey,length(dict[k]) == length(dict[xkey]). - If a key in
dictstarts with the same value asuxkey, there must be a corresponding key indictwhere theuxkeyprefix is replaced by thexkeyprefix, and its length must equallength(dict[xkey]). - If a key in
dictstarts with the same value asfkey,haskey(dict, fkey)andlength(dict[k]) == length(dict[fkey]). - If a key in
dictstarts with the same value asufkey, there must be a corresponding key indictwhere theufkeyprefix is replaced by thefkeyprefix, and its length must equallength(dict[fkey]).
Examples
julia> UniverseSets(; xkey = "nx", dict = Dict("nx" => ["A", "B", "C"], "group1" => ["A", "B"]))UniverseSets xkey ┼ String: "nx" uxkey ┼ String: "ux" fkey ┼ String: "nf" ufkey ┼ String: "uf" zkey ┼ String: "nz" dict ┴ Dict{String, Vector{String}}: Dict("nx" => ["A", "B", "C"], "group1" => ["A", "B"])Related
PortfolioOptimisers.AssetSetsMatrixEstimator — Type
struct AssetSetsMatrixEstimator{__T_val} <: AbstractConstraintEstimatorNames the group name key a binary asset-group membership matrix is built from.
The key is read out of a UniverseSets by asset_sets_matrix, which returns one row per distinct group value and one column per asset. A row of that matrix is the set indicator a group weight constraint sums the weights over.
Fields
val: Group name key for asset set membership matrix extraction.
Constructors
AssetSetsMatrixEstimator(; val::AbstractString) -> AssetSetsMatrixEstimatorKeywords correspond to the struct's fields.
Validation
!isempty(val).
Examples
julia> sets = UniverseSets(; xkey = "nx", dict = Dict("nx" => ["A", "B", "C"], "nx_sector" => ["Tech", "Tech", "Finance"]));julia> est = AssetSetsMatrixEstimator(; val = "nx_sector")AssetSetsMatrixEstimator val ┴ String: "nx_sector"julia> asset_sets_matrix(est, sets)2×3 transpose(::BitMatrix) with eltype Bool: 1 1 0 0 0 1Related
References
- [4] D. Cajas. Advanced Portfolio Optimization: A Cutting-edge Quantitative Approach (Springer Nature Switzerland, 2025). Section 9.1, Equations 9.2-9.4.
PortfolioOptimisers.MatNum_ASetMatE — Type
const MatNum_ASetMatE = Union{<:AssetSetsMatrixEstimator, <:MatNum}Alias for an asset sets matrix estimator or a numeric matrix.
Matches either an AssetSetsMatrixEstimator or a plain numeric matrix. Used internally in constraint generation that accepts a pre-computed membership matrix or an estimator.
Related
PortfolioOptimisers.MatNum_ASetMatE_VecMatNum_ASetMatE — Type
const MatNum_ASetMatE_VecMatNum_ASetMatE = Union{<:MatNum_ASetMatE, <:VecMatNum_ASetMatE}Alias for a single or vector of asset sets matrix estimators or numeric matrices.
Matches either a single MatNum_ASetMatE or a vector of them. Used for dispatch in asset set matrix operations that accept one or many estimators or matrices.
Related
PortfolioOptimisers.VecMatNum_ASetMatE — Type
const VecMatNum_ASetMatE = AbstractVector{<:MatNum_ASetMatE}Alias for a vector of asset sets matrix estimators or numeric matrices.
Represents a collection of MatNum_ASetMatE elements, enabling batch processing.
Related
PortfolioOptimisers.taxonomy_column — Function
taxonomy_column(
sets::UniverseSets,
key::AbstractString,
need::AbstractString
) -> Any
Read the taxonomy column sets.dict[key], checking that the key exists.
The sibling of factor_universe and feature_universe for a group name key, and written for the same reason: one shared helper whose message names the key and says what to do about it, so every consumer of a caller-supplied taxonomy key fails the same way. A bare sets.dict[key] raises a KeyError carrying the key alone, which says nothing about which of the two producers asked for it, and offers no help with a typo.
need names the consumer. The suggestion comes from suggest_declared_key, the looser configuration shared by every declaration-key suggestion: the candidates here are sets.dict keys the caller authored, not asset names, so the info-leak boundary of ADR 0026 does not apply.
The exception stays a KeyError, as factor_universe and feature_universe do, because a missing key is what it is — only the message improves.
Arguments
sets: AUniverseSetsobject specifying the asset universe and groupings.key: The group name key to read.need: Names the consumer in the error message.
Returns
col: The value ofsets.dict[key], one group per asset.
Related
PortfolioOptimisers.asset_sets_matrix — Function
asset_sets_matrix(
smtx::AbstractString,
sets::UniverseSets
) -> LinearAlgebra.Transpose{Bool, BitMatrix}
Construct a binary asset-group membership matrix from asset set groupings.
asset_sets_matrix generates a binary (0/1) matrix indicating asset membership in groups or categories, based on the key or group name smtx in the provided UniverseSets. Each row corresponds to a unique group value, and each column to an asset in the universe. This is used in constraint generation and portfolio construction workflows that require mapping assets to groups or categories.
Arguments
smtx: The key or group name to extract from the asset sets.sets: AUniverseSetsobject specifying the asset universe and groupings.
Validation
haskey(sets.dict, smtx), viataxonomy_column.- Throws an
AssertionErrorif the length ofsets.dict[smtx]does not match the asset universe.
Returns
A: Thetransposeof aBitMatrix, of size (number of groups) × (number of assets), whereA[i, j] == 1if assetjbelongs to groupi.
Details
- The function checks that
smtxexists insets.dictand that its length matches the asset universe. - Each unique value in
sets.dict[smtx]defines a group. - The output matrix is transposed so that rows correspond to groups and columns to assets.
Examples
julia> sets = UniverseSets(; xkey = "nx", dict = Dict("nx" => ["A", "B", "C"], "nx_sector" => ["Tech", "Tech", "Finance"]));julia> asset_sets_matrix("nx_sector", sets)2×3 transpose(::BitMatrix) with eltype Bool: 1 1 0 0 0 1Related
asset_sets_matrix(smtx::Option{<:MatNum}, args...)No-op fallback for asset set membership matrix construction.
This method returns the input matrix smtx unchanged. It is used as a fallback when the asset set membership matrix is already provided as an MatNum or is nothing, enabling composability and uniform interface handling in constraint generation workflows.
Arguments
smtx: An existing asset set membership matrix (MatNum) ornothing.args...: Additional positional arguments (ignored).
Returns
smtx::Option{<:MatNum}: The input matrix ornothing, unchanged.
Related
asset_sets_matrix(smtx::AssetSetsMatrixEstimator, sets::UniverseSets)This method is a wrapper calling:
asset_sets_matrix(smtx.val, sets)It is used for type stability and to provide a uniform interface for processing constraint estimators, as well as simplifying the use of multiple estimators simulatneously.
Related
asset_sets_matrix(smtx::VecMatNum_ASetMatE,
sets::UniverseSets)Broadcasts asset_sets_matrix over the vector.
Provides a uniform interface for processing multiple constraint estimators simulatneously.
PortfolioOptimisers.assert_feature_keys — Function
assert_feature_keys(vals::AbstractVector{<:AbstractString})
Assert that a list of taxonomy keys can produce a graded feature matrix.
Shared by asset_sets_features and AssetSetsFeatures, so the estimator rejects a bad key list at construction and the bare entry point rejects it at the call, from one encoding of the rule.
Two keys is a floor, not a style preference
Every Distances.jl semimetric is invariant to permuting coordinates, so a one-hot feature matrix — which is what a single partition gives — can only distinguish "same group" from "different group": the distance takes at most two values for every metric, and clustering it recovers the partition it was built from. Grading relatedness needs at least two partitions to agree or disagree about.
Duplicate keys are refused rather than deduplicated: repeating a key doubles its block, silently reweighting that partition against the others.
Arguments
vals: Group name keys to stack into the feature axis.
Validation
length(vals) >= 2.allunique(vals).
Returns
nothing.
Related
PortfolioOptimisers.AbstractFeatureValue — Type
abstract type AbstractFeatureValue <: AbstractAlgorithmAbstract supertype for markers that reinterpret a value in a graded feature program.
A bare Number in asset_sets_features' pair grammar sets a cell, absolutely. A marker says the number means something else, resolved against the cell's natural value by resolve_feature_value. Scale is the sole member; the family is open so a second reading can be added without touching the resolver.
Why the marker, and not nesting depth
The alternative was to let position decide — a top-level number scales, a nested one sets. That is invisible on a one-hot taxonomy, where the natural value is always 1.0 and the two readings coincide, and it diverges silently the moment a taxonomy carries numbers, which is exactly what the graded grammar adds. A marker makes the reading something the caller writes down.
Related
PortfolioOptimisers.Scale — Type
struct Scale{__T_val} <: AbstractFeatureValueMultiply a cell's natural value by val, rather than setting the cell to val.
Fields
val:val: Multiplier applied to the cell's natural value.
Constructors
Scale(; val::Number = 1.0) -> ScaleKeywords correspond to the struct's fields.
Validation
isfinite(val).
The natural value, and the two properties it forces
Scale scales the key's own datum, never the value already accumulated in the cell. At the top level of a program there is no accumulated value — the matrix starts at zero, and scaling zero is useless — so the only coherent referent is the underlying datum:
- a numeric taxonomy key: the asset's own number (
nx_esgat0.30,0.80,0.50); - a one-hot key, an asset node or a group member:
1.0when the row belongs,0.0when it does not.
Two consequences follow, and neither is a defect:
- The program stays a pure overwrite. Every write is
resolve_feature_value(v, natural)and never a read-modify-write, so last-wins ordering survives the marker. - Scaling a cross edge gives zero.
"C" => ["nx_country" => "US" => Scale(0.3)]scales C's US membership, and C is UK — so the cell is0.0, not0.3. Use a bare number to set a cross edge.
Examples
julia> resolve_feature_value(Scale(1.3), 0.5)0.65julia> resolve_feature_value(0.7, 0.5)0.7Related
PortfolioOptimisers.resolve_feature_value — Function
resolve_feature_value(v::Number, natural::Number) -> Number
resolve_feature_value(v::Scale, natural::Number) -> NumberResolve a graded-feature-program value v against the cell's natural value.
A bare number is absolute and ignores natural; a Scale multiplies it. Extending AbstractFeatureValue means adding one method here and nothing else.
Related
PortfolioOptimisers.Num_AFeatVal — Type
const Num_AFeatVal = Union{<:Number, <:AbstractFeatureValue}Alias for a resolved value in a graded feature program: a bare number, or a marker that reinterprets one.
This is the grammar's value production, and it is what distinguishes a value from a nested target list when a program entry is parsed.
Related
PortfolioOptimisers.asset_sets_features — Function
asset_sets_features(
vals::AbstractVector{<:AbstractString},
sets::UniverseSets;
strict
) -> Matrix{Float64}
Stack taxonomy memberships into an assets × features feature matrix.
asset_sets_features concatenates one asset_sets_matrix block per key in vals, transposed to assets-major, into the exogenous feature matrix FeatureDistance consumes. Feature k reads "belongs to group k", so two assets are close when they are classified together across many taxonomies.
This is the exogenous feature source: a sector, industry or country classification is structure that return correlations do not contain, which is exactly what a feature distance exists to bring in. Every other producer in the library derives Z from the returns.
Arguments
vals: Group name keys insets.dict, at least two (seeassert_feature_keys).sets: AUniverseSetsobject specifying the asset universe and groupings.strict: Accepted for a uniform interface with the graded method and ignored.strictgoverns name resolution, and on this path every name is asets.dictkey whose absence is an unconditionalKeyErrorfromtaxonomy_column— there is no soft failure for it to govern.
Validation
length(vals) >= 2andallunique(vals)(seeassert_feature_keys).- Each key exists in
sets.dictand has the length of the asset universe (enforced byasset_sets_matrix).
Returns
Z::Matrix{Float64}: Anassets × featuresmatrix,Z[i, k] == 1when assetibelongs to groupk. The feature count is the total number of distinct group values across every key.
Nested versus crossed taxonomies
The reading depends on how the keys relate, and both are useful:
- Nested (sector ⊃ industry ⊃ sub-industry): two assets share a level only if they share every coarser one, so the cosine similarity
shared / Lcounts the classification levels they agree on — a depth-graded relatedness. - Crossed (sector + country): the keys are independent, so the same count reads as how many independent attributes two assets happen to share.
The row norms are equal, exactly
asset_sets_matrix builds its groups from unique(all_sets), so every key is a partition: each asset lands in exactly one group per key, giving every row exactly L = length(vals) ones. All rows therefore have norm sqrt(L) and
cos(i, j) = shared(i, j) / L ∈ [0, 1]exactly, with no standardisation needed — the only producer with that property.
Float64, not BitMatrix
The result is dense Float64 rather than the BitMatrix asset_sets_matrix returns, so AngularDist keeps its BLAS gemm path.
Views
An asset view of a UniverseSets slices the groups prefixed by sets.xkey and leaves the rest alone, so a key named for a view to reach must carry that prefix — "nx_sector", not "sector". An unprefixed key does not fail silently: asset_sets_matrix's length check throws on the next call, because the sliced universe no longer matches the unsliced group.
Examples
julia> sets = UniverseSets(; xkey = "nx", dict = Dict("nx" => ["A", "B", "C"], "nx_sector" => ["Tech", "Tech", "Finance"], "nx_country" => ["US", "UK", "UK"]));julia> Z = asset_sets_features(["nx_sector", "nx_country"], sets)3×4 Matrix{Float64}: 1.0 0.0 1.0 0.0 1.0 0.0 0.0 1.0 0.0 1.0 0.0 1.0Feeding the user-supplied carrier
ReturnsResult requires nz whenever Z is set. Take it from asset_sets_feature_names rather than rebuilding the column order by hand:
ReturnsResult(; nx = nx, X = X, nz = asset_sets_feature_names(vals, sets), Z = asset_sets_features(vals, sets))Related
asset_sets_features(vals::AbstractVector{<:Pair}, sets::UniverseSets;
strict::Bool = false) -> Matrix{Float64}Resolve an ordered edge-authoring program into an assets × features matrix over the declared feature axis sets.dict[sets.zkey].
This is the graded contract. The group-name-key method above is the degenerate case of it — a partition stack with every written cell at 1.0 — and the two are separated by dispatch on vals' element type, so today's callers, today's matrix and today's "<key>=<group>" names are untouched.
The grammar
entry := rowsel => targets # row scope, then explicit columns | taxkey [=> group] => value # diagonal: those rows, their own membershiprowsel := asset | group | taxkey [=> group]target := taxkey [=> group] => value | asset => value | group => valuevalue := Number # sets, absolutely | <:AbstractFeatureValue # Scale(x): x × the key's natural valuetargets is one target or a vector of them. Entries are applied in order and each write is a pure overwrite, so last wins — repeating a key is the point, not a mistake, which is why allunique does not carry into this path.
Every target names its column in full. There is no ambient scope and no fallback chain: "UK" inside a nx_country entry would otherwise be resolved as a country by proximity rather than by what the caller wrote, and UK is also a real ticker.
Two things the declared axis buys
- Column order. A
Dicthas none, so without a declared list the feature axis would be whatever order the taxonomy happened to iterate in. - Fold invariance.
size(Z, 2)does not change under an asset view, becauseport_opt_view(::UniverseSets, i, args...)passeszkeythrough — the axis is authored, not summarised. This is the exact opposite of the group-name-key path, and both are documented because both are true. The cost is accepted: an asset node whose asset a view dropped survives as an all-zero column, benign for every blessed metric exceptDistances.CorrDist, which centres each row.
Names are bare, and what that forbids
Nodes are named plainly — "Tech", "US", "esg" (a numeric key with sets.xkey * "_" stripped), "A" for an asset node — because the caller authored the axis and qualifying it would make them write the prefix twice.
The accepted cost: a nested taxonomy with a repeated value is inexpressible in graded mode. With nx_industry and nx_subindustry both containing IntegratedOil, both land on the one bare node and the later entry overwrites the earlier — harmless under one-hot, where both wrote 1.0, silently lossy under grading. The group-name-key path qualifies its names and stays the tool for that case.
What is refused, and what is not
strict governs names only — an unknown node, asset or group warns with a did_you_mean suggestion, and throws under strict. Nothing structural is refused:
- An all-zero row is legal, and is the one genuine silent-wrongness case grading opens: an asset no entry touches has a zero row, and
FeatureDistance's zero-norm convention declares zero rows mutually identical, so forgotten assets cluster together at distance0. - A one-column matrix is legal:
assert_feature_keys' two-key floor is a property of stacking partitions and does not carry here.
Only non-emptiness is unconditional. A malformed entry — one that matches no production — throws regardless of strict, because there is no reading of it to fall back to.
Arguments
vals: The ordered program.sets: AUniverseSetswhosedictdeclares the feature axis undersets.zkey.strict: Whether an unresolvable name throws instead of warning.
Validation
!isempty(vals).haskey(sets.dict, sets.zkey)(seefeature_universe).
Returns
Z::Matrix{Float64}: Anassets × length(sets.dict[sets.zkey])matrix, zero-initialised.
Examples
julia> sets = UniverseSets(; xkey = "nx", zkey = "nz", dict = Dict{String, Any}("nx" => ["A", "B", "C"], "nz" => ["Tech", "Finance", "esg"], "nx_sector" => ["Tech", "Tech", "Finance"], "nx_esg" => [0.30, 0.80, 0.50]));julia> asset_sets_features(["nx_sector" => 2.0, "nx_esg" => Scale(1.3), "B" => ["nx_sector" => "Finance" => 0.2]], sets)3×3 Matrix{Float64}: 2.0 0.0 0.39 2.0 0.2 1.04 0.0 2.0 0.65Related
PortfolioOptimisers.asset_sets_feature_names — Function
asset_sets_feature_names(
vals::AbstractVector{<:AbstractString},
sets::UniverseSets
) -> Vector{String}
Name the columns asset_sets_features produces, in their own order.
ReturnsResult requires nz whenever Z is set, so the user-supplied carrier needs these names — and reproducing the column order by hand is exactly the kind of restatement that drifts. The pair is built from one traversal of vals, so the two can only agree.
Names are qualified by their key
Each name is "<key>=<group>" rather than the bare group value, because a nested taxonomy reuses its values across levels: an integrated-oil producer is in industry IntegratedOil and sub-industry IntegratedOil, so bare values would collide and ReturnsResult's uniqueness check would reject them.
Arguments
vals: The same group name keys, in the same order, passed toasset_sets_features.sets: AUniverseSetsobject specifying the asset universe and groupings.
Returns
nz::Vector{String}: One name per feature column,length(nz) == size(Z, 2).
Examples
julia> sets = UniverseSets(; xkey = "nx", dict = Dict("nx" => ["A", "B", "C"], "nx_sector" => ["Tech", "Tech", "Finance"], "nx_country" => ["US", "UK", "UK"]));julia> asset_sets_feature_names(["nx_sector", "nx_country"], sets)4-element Vector{String}: "nx_sector=Tech" "nx_sector=Finance" "nx_country=US" "nx_country=UK"Related
asset_sets_feature_names(vals::AbstractVector{<:Pair}, sets::UniverseSets) -> Vector{String}Name the columns a graded asset_sets_features program produces: the declared axis itself, sets.dict[sets.zkey].
The pairing with the matrix is trivial here, and deliberately so — the caller authored the axis, so the recipe below keeps the shape it has on the group-name-key path while the program itself carries no naming rule at all.
Names are bare, not qualified
The exact opposite of the group-name-key method above, and for a stated reason. That path derives its axis by stacking partitions, so it must qualify ("nx_industry=IntegratedOil") or a nested taxonomy would collide with itself. This path's axis is authored, so the names are whatever the caller wrote — "Tech", "US", "esg", "A" — and qualifying them would mean writing the prefix twice, once in dict[zkey] and again in every target.
The cost of bareness is a graded-mode limitation, documented on asset_sets_features: a nested taxonomy with a repeated value cannot be expressed, because both levels land on the one node.
Arguments
vals: The program. Read only for dispatch — the axis does not depend on it.sets: AUniverseSetswhosedictdeclares the feature axis undersets.zkey.
Returns
nz::Vector{String}: The declared feature axis,length(nz) == size(Z, 2).
Examples
julia> sets = UniverseSets(; xkey = "nx", zkey = "nz", dict = Dict{String, Any}("nx" => ["A", "B", "C"], "nz" => ["Tech", "Finance", "esg"], "nx_sector" => ["Tech", "Tech", "Finance"]));julia> asset_sets_feature_names(["nx_sector" => 2.0], sets)3-element Vector{String}: "Tech" "Finance" "esg"Related
PortfolioOptimisers.feature_program_candidates — Function
feature_program_candidates(sets::UniverseSets, nz) -> Vector{String}Build the did_you_mean pool for a graded feature program: asset names, every sets.dict key, every distinct value of every taxonomy key, and every declared feature node.
A graded program's names live in four namespaces at once, so a pool narrower than their union would answer "did you mean" with silence on the commonest typo.
Related
PortfolioOptimisers.is_feature_taxonomy_key — Function
is_feature_taxonomy_key(k, sets::UniverseSets) -> BoolWhether a name in a graded feature program is a taxonomy key: a sets.xkey-prefixed key of sets.dict.
UniverseSets guarantees every sets.xkey-prefixed dict key is asset-parallel, so the prefix rule alone decides the question — no new convention is needed, and row-selector precedence reduces to prefix, then asset, then group, exactly as estimator_to_val resolves.
Related
PortfolioOptimisers.is_feature_factor_key — Function
is_feature_factor_key(k, sets::UniverseSets) -> BoolWhether a name in a graded feature program declares the factor axis, by carrying the sets.fkey or sets.ufkey prefix.
Such a name is factor-length, so it can index neither the rows (assets) nor the declared nodes. It is refused by name rather than falling through to the plain-group branch, where it would fail later on a length mismatch that names neither the axis nor the cause. feature_rows puts the test between the asset branch and the group branch, and feature_factor_key_msg writes the diagnostic.
Related
PortfolioOptimisers.feature_grammar_msg — Function
feature_grammar_msg(term) -> StringBuild the error text for a malformed term of a graded feature program — an entry or a target — and print the grammar itself.
A malformed term is a syntax error, so the fastest fix is seeing the production it missed. Unlike the name diagnostics, this one never routes through strict_diagnostic: its callers throw an ArgumentError whatever strict says, because there is no reading of the term to fall back to.
Related
PortfolioOptimisers.feature_factor_key_msg — Function
feature_factor_key_msg(k, sets::UniverseSets) -> StringBuild the warning/error text for a name k that a graded feature program wrote in a row-selector or column-target position, but that declares the factor axis (see is_feature_factor_key).
The message names the two prefixes and the reason, and names sizes rather than universes — the same info-leak-safe discipline as unknown_variable_msg. It carries no did_you_mean suggestion: the name resolved perfectly well, on the wrong axis, so there is no typo to propose.
Related
PortfolioOptimisers.feature_missing_group_value_msg — Function
feature_missing_group_value_msg(key, group, col) -> StringBuild the warning/error text for a group value group of the taxonomy key key that matches no asset, where col is that key's column of sets.dict.
The message names the count of distinct values under the key, never the values themselves — the info-leak-safe discipline of unknown_variable_msg — and appends a did_you_mean suggestion drawn from those values. It drops group from its own candidate pool first, because the pool of a graded program is deliberately wide enough to contain a name that is nonetheless invalid in the position it was written.
Related
PortfolioOptimisers.feature_unknown_name_msg — Function
feature_unknown_name_msg(name, nx, key, pool; axis::AbstractString = "asset") -> StringBuild the warning/error text for a name of a graded feature program that resolves in no namespace of the axis it was written on.
The function wraps unknown_variable_msg, so it inherits that message's shape and names the size of nx rather than its members. axis is "asset" for a row selector and "feature" for a column target, and key is sets.xkey or sets.zkey to match. pool is the wide candidate pool of feature_program_candidates; the function drops name from it before suggesting, because that pool is deliberately wide enough — taxonomy values and declared nodes included — to contain a name that is nonetheless invalid in the position it was written.
Related
PortfolioOptimisers.feature_numeric_column — Function
feature_numeric_column(col) -> BoolWhether a taxonomy key's data are numbers, which is what decides the diagonal form's column and natural value.
A categorical key writes each asset into the node its own group value names, at a natural value of 1.0. A numeric key has one node — the key with sets.xkey * "_" stripped — and the asset's own number is the natural value. The two are the same production in the grammar and differ only here.
Related
PortfolioOptimisers.feature_write! — Function
feature_write!(Z::Matrix{Float64}, rows, node, natural, v, sets::UniverseSets, nz, zidx,
pool, strict::Bool) -> NothingWrite one column of a graded feature program: resolve node against the declared axis, then set Z[i, col] = resolve_feature_value(v, natural(i)) for every i in rows.
Every write in the program funnels through here, which is what makes last-wins a property of the traversal rather than a rule each production has to honour: the assignment is a pure overwrite, never a read-modify-write, so a later entry simply replaces an earlier one.
The column is resolved before natural is ever called, so an unknown node costs one diagnostic and no work — and natural never has to be defined for a node that does not exist.
Related
PortfolioOptimisers.feature_diagonal! — Function
feature_diagonal!(Z::Matrix{Float64}, key, group, v, sets::UniverseSets, nx, nz, zidx,
pool, strict::Bool) -> NothingApply the grammar's diagonal production, taxkey [=> group] => value: write each selected asset's own membership, rather than a column named from outside.
This is the production that makes a bare number on the right of a taxonomy key unambiguous, and it is why the uniform "targets are always fully qualified" rule costs nothing — "nx_country" => "UK" => 0.5 says what the two-level nested form used to say, one bracket shorter.
Related
PortfolioOptimisers.feature_rows — Function
feature_rows(sel, sets::UniverseSets, nx, pool, strict::Bool) -> Option{Vector{Int}}Resolve a non-taxonomy row selector — an asset or an asset group — to row indices, or nothing when the name does not resolve.
Taxonomy keys are handled by the caller, because the sets.xkey prefix rule settles them before any lookup. What is left is exactly estimator_to_val's precedence, asset first and then group, resolved by the shared resolve_axis_name, with the factor axis refused between them so a factor-length list is diagnosed by name rather than by an eventual length mismatch. The factor test therefore runs before the resolution, not after: a factor key that is also a sets.dict key would otherwise expand to factor names and be reported as a group of missing assets.
Related
PortfolioOptimisers.feature_target! — Function
feature_target!(Z::Matrix{Float64}, rows, target::Union{<:Pair, <:AbstractVector{<:Pair}}, sets::UniverseSets, nx, nz, zidx, pool,
strict::Bool) -> NothingApply the grammar's target production inside an already-resolved row scope, singly or over a vector.
Every target names its column in full, so this reads left to right with no ambient state: a taxonomy key with a group value names that value's node, a numeric taxonomy key names its own node, an asset names its own node, and a group expands to one node per member. The natural value is the row's membership of the node the target named — which is why scaling a cross edge gives zero.
Related
PortfolioOptimisers.feature_entry! — Function
feature_entry!(Z::Matrix{Float64}, entry::Pair, sets::UniverseSets, nx, nz, zidx, pool,
strict::Bool) -> NothingApply one entry of a graded feature program.
How the two productions are told apart
entry := rowsel => targets | taxkey [=> group] => value, and Julia's => is right-associative, so both arrive as a Pair whose right side may nest. Three tests separate them, in this order:
- The left side is a taxonomy key and the right side bottoms out in a value —
"nx_sector" => 2.0,"nx_country" => "UK" => 0.5. That is the diagonal, by decision: a bare value on the right of a taxonomy key always means "these rows, their own membership". - The left side is a taxonomy key, the right side is
g => tailwithtaila target list, andgis itself a taxonomy key. Thengstarts a target, so the left side is a bare row selector over the whole universe. This is the one genuine ambiguity in the grammar and it is resolved by the same prefix rule everything else uses. - Otherwise the left side is a row selector — restricted by
gwhen the left side is a taxonomy key, an asset or a group when it is not.
Related
PortfolioOptimisers.port_opt_view — Method
port_opt_view(smtx, i; kwargs...)Get a column view or subset of an asset sets membership matrix for asset index i.
Returns a column view for matrix inputs, the estimator unchanged for estimator inputs, or processes vectors element-wise.
Arguments
smtx: Asset sets matrix, estimator, or vector thereof.i: Asset index or range to slice.kwargs...: Additional keyword arguments.
Returns
- Column view of the matrix, or the estimator unchanged.
Related
References
- [4]
- D. Cajas. Advanced Portfolio Optimization: A Cutting-edge Quantitative Approach (Springer Nature Switzerland, 2025).