Super-large scale product structure configuration and index matching method based on composite rule graph
By constructing a composite rule graph and a Markov decision process, combined with reinforcement learning methods, the problem of insufficient expression of complex constraints in ultra-large-scale product structure configuration is solved, achieving efficient product structure configuration and indicator matching, and generating excellent configuration schemes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies are insufficient in expressing complex constraints and have low solution efficiency in ultra-large-scale product structure configurations. They are unable to fully characterize complex semantics and multi-dimensional constraints in a unified model, resulting in a surge in computational load and low efficiency of traditional methods, making it difficult to efficiently obtain structural configuration schemes with excellent overall performance.
We construct a Markov decision process based on a composite rule graph, optimize product structure configuration through reinforcement learning, and use the composite rule graph to uniformly represent the multi-layer semantic information and combination constraints of the product. We model the product configuration process as a Markov decision process and train the agent through a policy network to achieve dynamic pruning and adaptive configuration.
It achieves unified modeling of multi-layer semantic information and complex constraints of products, significantly improves the solution efficiency of ultra-large-scale product structure configuration, and generates configuration schemes that meet the constraints of composite rule graphs and have excellent overall performance.
Smart Images

Figure CN121921084A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of product structure configuration technology, and specifically to a method for ultra-large-scale product structure configuration and index matching based on composite rule graphs. Background Technology
[0002] With the serialization and platformization of complex products such as home appliances and transportation equipment, the number of selectable modules and components within the same product family has increased dramatically, resulting in an extremely large combination space for configuration schemes. In practical engineering, the product structure configuration process not only needs to ensure that the product's functions and performance indicators meet user needs, but also must simultaneously meet multi-dimensional constraints such as cost, reliability, noise, and vibration. Complex cross-level mutual exclusion, dependency, and coordination relationships are commonly found between product structural units, making the configuration problem typically characterized by a large candidate set, complex coupling relationships, and an exponentially growing solution space.
[0003] Existing product structure configuration methods largely rely on human experience, rule bases, or configuration trees, and often employ heuristic search strategies. These methods typically only express local structural constraints, lacking a unified representation of multi-layered semantic information ("requirements-functions-behaviors-structure") and their global combination relationships. For ultra-large-scale product structure configuration and indicator matching problems, traditional exhaustive search or fixed heuristic search strategies result in a surge in computational cost and low efficiency, making it difficult to efficiently obtain structural configuration solutions with superior overall performance. Therefore, there is an urgent need to research product structure configuration and indicator matching methods that can fully characterize complex semantics and multi-dimensional constraints in a unified model and achieve efficient and intelligent optimization in ultra-large-scale configuration spaces. Summary of the Invention
[0004] To address the common problems of insufficient expression of complex constraints and low solution efficiency in existing technologies for ultra-large-scale product structure configuration, this invention proposes a method for ultra-large-scale product structure configuration and index matching based on composite rule graphs. The specific technical solution is as follows:
[0005] A method for configuring and matching metrics for ultra-large-scale product structures based on composite rule graphs includes the following steps:
[0006] 1) Based on the target product and its application scenarios, the product's requirements, system functions, physical behavior and structural units are decomposed, and a composite rule graph integrating four layers of nodes (requirement layer, function layer, behavior layer, and structure layer) and layer constraints is constructed to uniformly represent the product's multi-layer semantic information, parameter attributes and combination constraints.
[0007] 2) Based on the composite rule diagram, the product configuration process is modeled as a Markov decision process: the structural layer is divided into several structural groups according to function or structural characteristics. Each structural group contains a set of structural units to be configured. Starting from empty configuration, each structural group is configured in a predetermined order.
[0008] 3) Using the Markov decision process as a reinforcement learning environment, a policy network is constructed as an agent, with the state vector as input and the selection probability of all possible actions in that state as output. In this environment, the configuration process from the initial state to the final state is repeatedly executed, and the policy network parameters are iteratively trained using a reinforcement learning method based on policy gradient. The trained policy network is used to configure structural units for each structure group starting from the initial state, thereby obtaining the product structure configuration scheme.
[0009] Furthermore, step 1) specifically includes:
[0010] 1.1) Decompose the product's requirements, system functions, physical behavior, and structural units to determine the node sets of the requirement layer, functional layer, behavioral layer, and structural layer, which are used to describe the product's multi-layered semantic information;
[0011] 1.2) Based on the semantic mapping relationship between each layer and nodes in the same layer in step 1.1), construct layer constraints including directed edges of rules, super edges of rules, and parameter sets of rule nodes to obtain an RFBS structure that reflects the multi-layer semantics and constraint relationship of the product, and formally define the composite rule graph of the product structure configuration.
[0012] Furthermore, in step 1.1):
[0013] The demand layer is used to describe product performance requirements, including at least one target value / range or normalized preference weight for a performance indicator.
[0014] The functional layer is used to describe the key functions of the product;
[0015] The behavior layer is used to describe the product's operational behavior;
[0016] The structural layer is divided into several structural groups, and each structural group contains a set of structural units to be configured.
[0017] Furthermore, in step 1.2):
[0018] The directed edge according to the rule is specifically represented as: e k =(v i ,v j ,ω(v i ,v j )), v i ,v j∈V represents the start and end points of the directed edges of the rules, respectively, and V is the set of nodes in the requirement layer, functional layer, behavioral layer, and structural layer. ω(v i ,v j Let ω be the edge weight function ω on the node pair (v) i ,v j The value at ) is used to quantize node v. i For node v j The weights of semantic dependencies or constraints are used as the path for decomposing and allocating high-level objectives to lower-level elements in the RFBS structure.
[0019] The rule-defined hyperedge is specifically represented as: each rule-defined hyperedge e h e h =(N(e),τ(e),θ e ), Let τ(e) be the set of nodes associated with the rule hyperedge, representing a combination constraint that associates at least two structural unit nodes or related cross-layer nodes; τ(e) is the type of rule hyperedge; θ e This is the vector of regular hyperedge parameters, used to describe the constraint weights or influence coefficients of the combined constraints on each performance dimension;
[0020] The set of rule node parameters is specifically represented as: Π=Y∪Q, where Y={y u |u∈S str} represents the set of performance attribute vectors for structural units, used to reflect performance-related physical parameters; Q = {q u |u∈S str} represents the set of cost attribute vectors for structural units, used to reflect the values of structural units in terms of cost attributes.
[0021] Furthermore, the rule-based superedge specifically includes:
[0022] Mutual exclusion rule superedges are used to describe combinations of components that are not allowed to be used simultaneously.
[0023] Dependency coexistence class rule hyperedges are used to describe combinations of components that must be used simultaneously;
[0024] Collaborative classification rule hyperedges are used to describe the performance gain of a combination of components when they appear simultaneously.
[0025] Conflict reduction rule superedges are used to describe the adverse effects on system performance when component combinations occur.
[0026] Furthermore, step 2) specifically includes:
[0027] 2.1) Structural layer S str Based on functional or structural characteristics, the structures are divided into T groups, denoted as:
[0028]
[0029] Each structural group Corresponding to a set of structural units to be configured Grouped by the first structure Up to the Tth structural group For the predetermined configuration order;
[0030] 2.2) Define the state space X: Denote any state vector in the state space as:
[0031] x t =(U t ,Y(U t ),L(U t )), x t ∈X
[0032] Among them, U t For the currently selected set of structural units, Y(U) t ) represents the performance attribute vector calculated based on the composite rule graph, L(U) t The cost attribute vector is calculated based on the composite rule graph; the system-level performance vector is calculated through the inter-layer mapping matrix of the composite rule graph, and the structural layer aggregation vector is normalized to obtain the relative importance weight of each structural unit.
[0033] 2.3) Define the action space A: for any non-terminal state x t The following action a t ∈A(x t This represents the grouping of candidate actions within the current structure after performing validity checks based on a composite rule graph and pruning the candidate actions. The corresponding set of structural units Select one structural unit from the set and add it to the selected set U. t ;
[0034] 2.4) Define the state transition function P: for any non-terminating state x t Execute action a t Then, the next state x is obtained. t+1 :
[0035]
[0036] Among them, U t+1 =U t ∪{u} represents the updated set of selected structural units; Y(U t+1 ) and L(U t+1 () are sets U based on composite rule graphs. t+1 The recalculated performance and cost attribute aggregation vector; when t = T, the x obtained after the transitionT+1 This indicates that the configuration has been terminated.
[0037] 2.5) Define the reward function R: Define the reward function as follows:
[0038]
[0039] For all non-terminating states, the reward function value is 0; when the configuration process terminates, the state is x. T+1 And obtain the complete configuration scheme U T+1 When calculating the comprehensive evaluation score J(U) of the scheme, T+1 As a final reward.
[0040] Furthermore, in step 2.2):
[0041] From the composite rule graph, the directed edges of rules from the requirement layer to the functional layer, from the functional layer to the behavior layer, and from the behavior layer to the structure layer are extracted and organized into an inter-layer mapping matrix W. RF W FB W BS The aggregated representations of the requirement layer, functional layer, behavioral layer, and structural layer are denoted as vectors r, f, b, and s, respectively, such that the inter-layer mapping relationship satisfies:
[0042] f = W RF r, b = W FB f, s = W BS b
[0043] Normalizing the aggregation vector s of the structural layer yields the relative importance weights λ of the structural units. u ,
[0044] In any state x t Next, calculate the system-level performance aggregation vector for the current configuration:
[0045]
[0046] The cost attribute of structural unit u in the nth dimension is denoted as q. u,n (n = 1, ..., m), where m is the number of cost attributes; the cost aggregation function f is selected. n (·) Perform aggregation to obtain the current set of selected structural units U. t The aggregation result of the nth cost attribute l n (U t )=f n ({q u,n |u∈U t}), thus obtaining the cost attribute aggregate vector:
[0047] L(U t )=(l1(U t),...,l m (U t ))
[0048] Furthermore, in step 2.3):
[0049] The candidate action pruning specifically involves: in any state x t Next, for each candidate structural unit u in the current structural group, construct a temporary configuration set U. tmp =U t ∪{u};
[0050] Calculate L(U) tmp ) and compare with the cost attribute constraints. If the current action exceeds the allowable range in any cost attribute dimension, remove the current action from the set of actionable actions A(x). t Delete it;
[0051] Retrieve mutually exclusive rule superedges in the composite rule graph, when U tmp When all structural units associated with a certain mutually exclusive rule superedge are already included, the current action is removed from the action set A(x). t Delete it;
[0052] In the composite rule graph, retrieve dependency coexistence rule hyperedges and remove candidate structural units from subsequent structural groups that are incompatible with the dependency relationship. If the set of candidate structural units in a subsequent structural group is empty after constraint search, remove the current action from the set of available actions A(x). t ) was deleted.
[0053] Furthermore, in step 2.5):
[0054] The comprehensive score calculation involves: after completing all structural grouping configurations, checking the hyperedge parameters corresponding to the hyperedges of collaborative addition and conflict subtraction rules in the composite rule graph, and calculating the performance correction amount ΔY(U). T+1 ), and correct the final system-level performance aggregation vector, Y'(U T+1 )=Y(U T+1 )+ΔY(U T+1 Construct a comprehensive scoring function:
[0055] J(U T+1 )=λ cost J cost (U T+1 )+λ match J match (U T+1 )+λ mag J mag (U T+1 )
[0056] Among them, J costFor L(U) T+1 The cost attribute utilization score obtained from the cost constraint; J match Based on Y'(U T+1 ) and demand R req The score for the degree of matching between the indicators; J mag Based on Y'(U T+1 Overall performance strength score; λ cost +λ match +λ mag =1 is the weighting coefficient.
[0057] Furthermore, step 3) specifically includes:
[0058] 3.1) Constructing a reinforcement learning environment: The Markov decision process described above is used as the reinforcement learning environment. This environment receives the current state x... t Then, based on the composite rule graph, the legality of the candidate action set is judged and pruned, and the set of available actions A(x) in the current state is returned. t Upon receiving action a t Then, based on the state transition function P(x) t+1 |x t ,a t Return to the next state x t+1 And based on the reward function R(x) t ,a t ,x t+1 Return the reward;
[0059] 3.2) Constructing the intelligent agent: Constructing an agent based on the state vector x t Given the input, the set of actions A(x) that can be performed in this state. t The policy network in which the probability of choosing each action is the output is the policy network. As a reinforcement learning agent, in which This is the parameter vector of the policy network;
[0060] 3.3) Training the policy network: The agent is run in the reinforcement learning environment for multiple rounds of iterative training. Each round of training starts from the initial state x1 and makes T decisions to obtain the final state x. T+1 With reward R T =R(x) T ,a T ,x T+1 );
[0061] In step t, the policy network receives x returned by the environment. t With A(x) t After that, output A(x) t The probability of choosing each action in ) And select action a according to probability.t Execution, calculated by the environment, returns the next state x. t+1 With the action set A(x) t+1 );
[0062] To maximize reward R T The expected value is the objective, and a policy gradient-based reinforcement learning method is used to optimize the policy network parameters. Update;
[0063] When the improvement in the overall score during training is less than a preset threshold ε for M consecutive evaluation periods, the policy network is considered to have converged, and the optimal policy network is obtained.
[0064] 3.4) Configuration Inference: This involves configuring the trained, converged policy network. As a decision-making strategy; for a new configuration task, starting from the initial state x1, at each step, select the action with the highest probability, i.e. Structural units are selected sequentially for each structural group until a complete configuration scheme U is formed. T+1 .
[0065] The beneficial effects of this invention include:
[0066] 1. Achieve unified knowledge modeling for multi-layered semantics and complex constraints: By constructing a composite rule graph that integrates four layers of nodes (requirements, functions, behaviors, and structures) along with directed edges and hyperedges, a unified modeling of multi-layered semantic information and complex engineering constraints of the product is achieved. This graph structure integrates performance, cost, and various constraint information, overcoming the limitations of fragmented knowledge and semantic separation in traditional rule bases and configuration trees, and providing a complete, consistent, and computable knowledge representation foundation for product structure configuration.
[0067] 2. A Markov decision-making framework based on dynamic pruning mechanism is proposed: the product configuration process is formalized as a Markov decision process. System-level performance and cost attributes calculated in real time by composite rule graph are dynamically integrated in the state space. Candidate structural units are intelligently pruned in the action space according to rule hyperedges and cost constraints. This significantly compresses the action space of each decision step, effectively avoids the combinatorial explosion problem of traditional exhaustive search and heuristic methods in ultra-large-scale, strongly coupled constraint scenarios, and significantly improves the solution efficiency of multi-constraint, multi-index configuration tasks.
[0068] 3. Implementing Reinforcement Learning-Based Adaptive Configuration Strategy Optimization: A reinforcement learning environment and policy network are constructed based on a formal decision model. Through multiple rounds of interactive iteration, configuration strategies for ultra-large-scale candidate sets are automatically learned, achieving a comprehensive trade-off between system-level performance, cost, and constraint satisfaction without relying on human experience or complex rule design. After training, a convergent strategy can quickly generate configuration schemes that satisfy composite rule graph constraints and exhibit excellent overall performance, providing reliable and efficient intelligent decision support for ultra-large-scale product structure configuration. Attached Figure Description
[0069] Figure 1 This is a framework diagram of the ultra-large-scale product structure configuration and index matching method based on composite rule graphs in this embodiment;
[0070] Figure 2 This is the knowledge graph of a refrigerator as an example in this embodiment;
[0071] Figure 3 This is a schematic diagram of the reinforcement learning agent in this embodiment. Detailed Implementation
[0072] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments described herein are intended to aid in understanding the technical process and calculation methods of the present invention; the data, parameters, tables, and specific configurations are merely examples and can be adjusted according to actual product types and design requirements. The scope of protection of the present invention is defined by the claims.
[0073] like Figure 1 As shown in this embodiment, a method for ultra-large-scale product structure configuration and index matching based on composite rule graphs is proposed. A composite rule graph is constructed with household refrigerators as the target product. Based on the composite rule graph, the refrigerator configuration process is modeled as a Markov decision process, and a reinforcement learning method is used to solve it. The method includes the following steps:
[0074] 1) Construction of Composite Rule Graph: The requirements, system functions, physical behaviors, and structural units of the refrigerator are decomposed to determine the node sets of the Requirement, Function, Behavior, and Structure layers, which are used to describe the multi-layer semantic information of the refrigerator. Based on this, the semantic mapping relationship between each layer and the nodes in the same layer is established, the constraint relationship of the combination of structural units and cross-layer combination is identified, and the RFBS model reflecting the multi-layer semantics and constraint relationship of the product is obtained. The Composite Rule Graph (CRG) of the product structure configuration is formally defined.
[0075] like Figure 2 As shown, the details are as follows:
[0076] 1.1) RFBS Model Construction:
[0077] 1.1.1) Demand Layer R req Modeling: In this embodiment, user requirements are abstracted into a three-dimensional requirement vector:
[0078]
[0079] Where: r cool This represents the preference weight for cooling capacity; r energy This represents the preference weight for low energy consumption and energy efficiency; r noise This indicates the preference weight for quietness performance.
[0080] In different application scenarios, different demand weight vectors r can be set according to user groups and market positioning. In this embodiment, a set of user preferences representing "cooling priority, while taking into account energy efficiency and quietness" is selected as typical demand weights for subsequent calculation.
[0081]
[0082] 1.1.2) Functional Layer F fun The structure: The functional layer describes the key functions of the refrigerator in terms of refrigeration, heat exchange, airflow organization, control, insulation, door, vibration reduction, and defrosting. This embodiment defines 9 functional units, as shown in Table 1.
[0083] Table 1 List of Functional Layer Nodes
[0084]
[0085] 1.1.3) Behavioral Layer B beh Modeling: The behavior layer describes behaviors such as refrigerant circulation, heat exchange, airflow organization, temperature field, dynamic response, heat loss, noise, and defrosting process. The node list is shown in Table 2.
[0086] Table 2 List of Behavioral Layer Nodes
[0087]
[0088]
[0089] 1.1.4) Structural layer S str Modeling: In this embodiment, structural layer S str Subdivided into 9 structural groups Each structural group Contains a set of structural units to be configured In this embodiment, the various subsystems of the refrigerator are shown in Table 3.
[0090] Table 3 List of nodes in the structural layer subsystem
[0091]
[0092] 1.1.5) Based on the modeling of the above layers, the node set representation of each layer can be obtained: V = R req ∪F fun ∪B beh ∪S str .
[0093] 1.2) Construction of Directed Edges Based on Rules: In this embodiment, the mapping relationship between the requirement layer, functional layer, behavioral layer and structural layer is constructed based on a multi-layer node set, and the directed edge set E is used as the basis. k Representation is performed. The rule is a directed edge e. k =(v i ,v j ,ω(v i ,v j )), v i ,v j ∈V represent the start and end points of the regular directed edge, respectively, ω(v i ,v j Let ω be the edge weight function ω on the node pair (v) i ,v j The value at ) is used to quantize node v. i For node v j The weights of semantic dependencies or constraints are used as the path for decomposing and allocating high-level objectives to lower-level elements in the RFBS structure.
[0094] Some exemplary mapping relationships are shown in Table 4.
[0095] Table 4 shows examples of directed edges with rules.
[0096]
[0097] The aforementioned directed edges are used to illustrate typical semantic dependencies between the functional layer and the behavioral layer, and between the behavioral layer and the structural layer. In practice, this embodiment uses the complete RFBS mapping matrix to generate all directed edges and their weights.
[0098] 1.3) Construction of rule-based superedges:
[0099] For the regular hyperedge set E h Each rule in the super-edge e h e h =(N(e),τ(e),θ e ) , Let τ be the set of nodes associated with the rule's hyperedge, representing a combination constraint that associates at least two structural unit nodes or related cross-layer nodes; τ(e) is the type of rule hyperedge, including mutually exclusive, dependent coexistence, cooperative addition, and conflict reduction; θ e This is the vector of hyperedge parameters for the rules, used to describe the constraint weights or influence coefficients of the combined constraints on each performance dimension.
[0100] Examples of its node set N(e) and type τ(e) are shown in Table 5.
[0101] Table 5. Examples of Mutually Exclusive / Dependency Coexistence Hyperedges
[0102]
[0103] 1.3.1) Mutual exclusion and dependency coexistence rule hyperedges: In this embodiment, mutual exclusion rule hyperedges are used to describe component combinations that are not allowed to be selected at the same time; dependency coexistence rule hyperedges are used to describe component combinations that must be selected at the same time.
[0104] 1.3.2) Cooperative and Conflict-Related Rule Hyperedges: In this embodiment, cooperative and classification rule hyperedges are used to describe the performance gain of system when component combinations occur simultaneously; conflict reduction rule hyperedges are used to describe the adverse effects of component combinations on system performance. An example is shown below:
[0105] Synergistic addition classification rule superedge HPY_Bonus_001: When a high-end energy-saving and quiet combination occurs simultaneously, it is assumed that there is a synergistic gain between the overall energy efficiency and quiet performance.
[0106]
[0107] The corresponding parameter vector is
[0108] Conflict reduction classification rule overedge HPY_Penalty_001: When a high-power compressor is combined with minimum insulation, foundation door, and foundation vibration reduction, energy consumption and noise are considered to be significantly worsened.
[0109]
[0110] The corresponding parameter vector is
[0111] When configuration scheme U t When it includes all nodes of a certain rule's superedge, let's denote it as e. h ∈E h (U t And in the system-level performance aggregation calculation, it is based on its parameter vector θ e Adjust the corresponding performance dimensions by gaining or reducing them.
[0112] 1.4) Construction of the rule node parameter set Π: Π=Y∪Q, where Y={y u |u∈S str} represents the set of performance attribute vectors for structural units, used to reflect performance-related physical parameters; Q = {q u |u∈S str} represents the set of cost attribute vectors for structural units, used to reflect the values of structural units in cost and other cost attributes.
[0113] Specifically, for each structural unit node u∈S unit The three-dimensional performance vector is obtained by normalizing the performance parameters related to cooling, energy efficiency, and noise reduction.
[0114]
[0115] in, This indicates the performance of the structural unit in terms of cooling performance; This indicates the energy efficiency performance of the structural unit; This indicates the performance of the structural unit in terms of noise reduction;
[0116] The cost attribute only includes one item: manufacturing cost, denoted as q. u , which represents the manufacturing cost when the structural unit is selected. A detailed example of the structural unit is shown in Table 6.
[0117] Table 6. Detailed List of Structural Layer Components
[0118]
[0119]
[0120] In summary, the metamodel of the composite rule graph is represented as follows:
[0121] CRG = <V,E k E h ,Π>
[0122] 2) Modeling the product configuration process as a Markov decision process: In this embodiment, the nine structural subsystems of the structural layer are grouped into nine structural groups S1, S2…S9, and the configuration order from the refrigeration system to the accessory system is preset. Based on the composite rule diagram, the product structural configuration process is modeled as a Markov decision process: MDP = (X, A, P, R), as follows:
[0123] 2.1) Structural layer S str Based on functional or structural characteristics, the structures are divided into T groups, denoted as:
[0124]
[0125] Each structural group Corresponding to a set of structural units to be configured Grouped by the first structure Up to the Tth structural group This is the pre-defined configuration order.
[0126] 2.2) Define the state space X: In this embodiment, the state vector is:
[0127] x t =(U t ,Y(U t ),L(U t ))
[0128] Among them, U t For the currently selected set of structural units, Y(U) t L(U) is a three-dimensional performance attribute vector representing cooling, energy efficiency, and noise. t The manufacturing cost attribute is used. The system-level performance vector is calculated using the inter-layer mapping matrix of the composite rule graph. The structural layer aggregation vector is then normalized to obtain the relative importance weight λ of each structural unit. u The weights of the structural subsystem in terms of cooling capacity, energy efficiency, and noise were calculated, and the specific calculation method is as follows:
[0129] From the composite rule graph, the directed edges of rules from the requirement layer to the functional layer, from the functional layer to the behavior layer, and from the behavior layer to the structure layer are extracted and organized into an inter-layer mapping matrix W. RF W FB W BS The aggregated representations of the requirement layer, functional layer, behavioral layer, and structural layer are denoted as vectors r, f, b, and s, respectively, such that the inter-layer mapping relationship satisfies:
[0130] f = W RF r, b = W FB f, s = W BS b
[0131] Normalizing the aggregation vector s of the structural layer yields the relative importance weights λ of the structural units. u ,
[0132] In any state x t Next, calculate the system-level performance aggregation vector for the current configuration:
[0133]
[0134] The cost attribute of structural unit u in the nth dimension is denoted as q. u,n (n = 1, ..., m), where m is the number of dimensions of the cost attribute. Choose an appropriate cost aggregation function f. n(·) Perform aggregation to obtain the current set of selected structural units U. t The aggregation result of the nth cost attribute l n (U t )=f n ({q u,n |u∈U t}), thus obtaining the cost attribute aggregate vector:
[0135] L(U t )=(l1(U t ),...,l m (U t ))
[0136] Table 7 shows the weights of the subsystem layers:
[0137] Table 7 Weights of Structural Subsystem Layers
[0138]
[0139] 2.3) Define the action space A: In this embodiment, any state x t The set of actions A(x) t It consists of candidate structural units from the current structural group. The candidate action set is pruned; if adding a candidate structural unit results in a manufacturing cost q... u If the action exceeds the allowed range, causes a mutually exclusive superedge relationship with the selected structural unit, or causes a prerequisite relationship conflict with subsequent structural groupings, the action will be removed from A(x). t Delete it. Details are as follows:
[0140] For any non-termination state x t (i.e., possible actions a under t < T) t ∈A(x t This represents the grouping of candidate actions within the current structure after performing validity checks based on a composite rule graph and pruning the candidate actions. Corresponding candidate set of structural units Select one structural unit from the set and add it to the selected set U. t ;
[0141] In any state x t Next, for each candidate structural unit x in the current structural grouping t Construct a temporary configuration set U tmp =U t ∪{u};
[0142] Calculate L(U) tmp ) and compare with the cost attribute constraints. If the current action exceeds the allowable range in any cost attribute dimension, remove the current action from the set of actionable actions A(x). t Delete it;
[0143] Retrieve mutually exclusive rule superedges in the composite rule graph, when U tmp When all structural units associated with a certain mutually exclusive rule superedge are already included, the current action is removed from the action set A(x). t Delete it;
[0144] In the composite rule graph, retrieve dependency coexistence rule hyperedges and remove candidate structural units from subsequent structural groups that are incompatible with the dependency relationship. If the set of candidate structural units in a subsequent structural group is empty after constraint search, remove the current action from the set of available actions A(x). t Delete it;
[0145] The remaining candidate structural units that simultaneously satisfy the cost constraint and will not cause violations in subsequent steps retain the corresponding actions in the set of possible actions A(x) of state x. t )middle.
[0146] 2.4) Define the state transition function P: for any non-terminating state x t Execute action a t Then, the next state x is obtained. t+1 :
[0147]
[0148] Among them, U t+1 =U t ∪{u} represents the updated set of selected structural units; Y(U t+1 ) and L(U t+1 () are sets U based on composite rule graphs. t+1 The recalculated performance and cost attribute aggregation vector; when t = T, the x obtained after the transition T+1 This indicates that the configuration has been terminated.
[0149] In this embodiment, the state transition follows the aforementioned configuration process, adding the structural unit corresponding to the current action to the selected set and updating the parameters of the state vector. When the structural grouping index t = 9, it indicates that the configuration of the 9 subsystems is complete, and the corresponding state is the termination state. The final configuration scheme is denoted as U. 10 .
[0150] 2.5) Define the reward function R: Define the reward function as follows:
[0151]
[0152] For all non-terminating states, the reward function value is 0; when the configuration process terminates, the state is x. T+1 And obtain the complete configuration scheme U T+1 When calculating the overall score J(U) of the scheme,T+1 This will be awarded as the final reward. The overall score will be calculated as follows:
[0153] After completing all structural grouping configurations, check the hyperedge parameters corresponding to the hyperedges of collaborative addition and conflict subtraction rules in the composite rule graph, and calculate the performance correction amount ΔY(U). T+1 ), and correct the final system-level performance aggregation vector, Y'(U T+1 )=Y(U T+1 )+ΔY(U T+1 Construct a comprehensive scoring function:
[0154] J(U T+1 )=λ cost J cost (U T+1 )+λ match J match (U T+1 )+λ mag J mag (U T+1 )
[0155] Among them, J cost For L(U) T+1 The cost attribute utilization score obtained from the cost constraint; J match Based on Y'(U T+1 ) and demand R req The score for the degree of matching between the indicators; J mag Based on Y'(U T+1 Overall performance strength score; λ cost +λ match +λ mag =1 is the weighting coefficient.
[0156] In this embodiment, the reward is 0 for intermediate steps where not all structural grouping configurations have been completed; the reward is only applied to the final configuration scheme U at the termination state, i.e., t=9. 10 The reward J(U) is calculated based on cost utilization, performance matching degree, and overall performance strength. 10 ):
[0157] J(U 10 ) = 0.60J cost (U 10 )+0.35J match (U 10 +0.05S mag (U 10 )
[0158] Among them, the cost utilization weight is the highest at 0.6, followed by the indicator matching weight at 0.35, and the overall performance strength weight at 0.05, reflecting that this embodiment pays more attention to the indicator matching degree under the premise of meeting the cost constraint.
[0159] 3) such as Figure 3 As shown, the Markov decision process is solved using reinforcement learning methods:
[0160] 3.1) Constructing a reinforcement learning environment: The Markov decision process is used as the reinforcement learning environment. The environment receives the current state x... t Then, based on the composite rule graph, the legality of the candidate action set is judged and pruned, and the set of available actions A(x) in the current state is returned. t Upon receiving action a t Then, based on the state transition function P(x) t+1 |x t ,a t Return to the next state x t+1 And based on the reward function R(x) t ,a t ,x t+1 Return the reward;
[0161] 3.2) Constructing a policy network Construct a state vector x t Given the input, the set of actions A(x) that can be performed in this state. t The policy network in which the probability of choosing each action is the output is the policy network. As a reinforcement learning agent, in which This is the parameter vector of the policy network;
[0162] 3.3) Training the policy network: The agent is run in the reinforcement learning environment for multiple rounds of iterative training. Each round of training starts from the initial state x1 and makes T decisions to obtain the final state x. T+1 With reward R T =R(x) T ,a T ,x T+1 ).
[0163] In step t, the policy network receives x returned by the environment. t With A(x) t After that, output A(x) t The probability of choosing each action in ) And select action a according to probability. t Execution, calculated by the environment, returns the next state x. t+1 With the action set A(x) t+1 ).
[0164] To maximize reward RT The expected value is the objective, and a policy gradient-based reinforcement learning method is used to optimize the policy network parameters. The update will be performed, and the update rules are as follows:
[0165]
[0166] Where α is the learning rate.
[0167] When the improvement in the overall score during training is less than a preset threshold ε for M consecutive evaluation periods, the policy network is considered to have converged, and the optimal policy network is obtained.
[0168] 3.4) Configuration Inference: This involves configuring the converged policy network. As a fixed decision-making strategy; for a new configuration task, starting from the initial state x1, at each step t, the action with the highest probability is selected, i.e. Structural units are selected sequentially for each structural group until a complete configuration scheme U is formed. T+1 .
[0169] In this embodiment, a two-layer fully connected feedforward neural network is used as the policy network, and a neural network that takes the state vector as input and outputs the probability of each action selection based on the pruned action set in the current state is used as the reinforcement learning agent. During the training phase, the REINFORCE algorithm based on policy gradients is used, with an upper limit of 20,000 training epochs and a learning rate α = 1 × 10⁻⁶. -3 During training, if the absolute value of the improvement in the overall score for 20 consecutive times is less than the preset threshold ε = 10... -4 When the training process has converged, the update of the policy network parameters is stopped.
[0170] After the policy network training converges, the policy network is run for configuration reasoning: starting from the initial state (t=1), at each decision step, based on the selection probability output by the policy network in the current state, the action with the highest probability is selected from the set of possible actions, and the structural units of each structural group are determined sequentially until a complete set of structural configuration schemes U is obtained. 10 And calculate its comprehensive evaluation score J(U) 10 While keeping the other modeling parameters unchanged, the cost constraints were set to 700, 800, 1000, 1200, 1600 and 2000 respectively. The policy network training and configuration solution were repeated to obtain the optimal structural configuration schemes and their performance evaluation results under different cost constraints, as shown in Tables 8 and 9.
[0171] Table 8. Performance indicators of the optimal configuration schemes obtained under different cost constraints.
[0172]
[0173] Table 9. Detailed Structure of Optimal Configuration Schemes Obtained Under Different Cost Constraints
[0174]
[0175] In summary, this invention first constructs a composite rule graph integrating four layers of nodes (demand, function, behavior, and structure) as well as directed edges and hyperedges, achieving a unified representation of multi-layered semantic information, parameter attributes, and complex combination constraints of a product. Then, based on this composite rule graph, the product configuration process is formalized as a Markov decision process: system-level performance and cost attribute information calculated in real-time by the composite rule graph is integrated in the state space; candidate structural units are dynamically pruned using rule hyperedges and cost constraints in the action space; and the performance index matching degree of the final configuration scheme is quantitatively evaluated through a comprehensive scoring function. Finally, the decision process is solved using reinforcement learning, training a policy network with state vectors as input and action selection strategies as output. This allows the agent to iteratively learn configuration strategies for a massive candidate set through multiple rounds of interaction with the environment. Using the resulting policy network, a product structure configuration scheme with superior overall performance and high index matching degree is automatically generated while satisfying complex engineering constraints.
[0176] As can be seen, the composite rule graph constructed in this invention can simultaneously express requirements, performance, costs, and multiple types of engineering constraints within a unified semantic space, enabling the construction of configuration states and constraint determination to be completed within the same graph structure. Based on the Markov decision process formed by this graph model, system-level performance and cost information is integrated into the states, and candidate structural units are dynamically pruned step-by-step at the action end, effectively compressing the ultra-large-scale configuration space. Even under multiple constraints, the feasibility and efficiency of the solution are maintained, providing a unified computational framework for product structure configuration and indicator matching.
[0177] Building upon this foundation, the policy network trained through reinforcement learning can automatically learn the trade-off mechanism between performance and cost, generating structural configuration schemes that satisfy the constraints of the composite rule graph and exhibit excellent overall performance under different cost constraints. Test results show that various performance indicators exhibit a monotonic and interpretable improvement trend as cost constraints are relaxed, and structural differences can be traced back to their source through the mapping links and constraints of the composite rule graph. This achieves unified knowledge modeling, dynamic constraint pruning, and adaptive policy optimization, providing accurate, efficient, and interpretable intelligent decision support for ultra-large-scale product structural configuration and indicator matching based on composite rule graphs, demonstrating significant engineering application value.
Claims
1. A method for configuring and matching indicators of ultra-large-scale product structures based on composite rule graphs, characterized in that, Includes the following steps: 1) Based on the target product and its application scenarios, the product's requirements, system functions, physical behavior and structural units are decomposed, and a composite rule graph integrating four layers of nodes (requirement layer, function layer, behavior layer, and structure layer) and layer constraints is constructed to uniformly represent the product's multi-layer semantic information, parameter attributes and combination constraints. 2) Based on the composite rule diagram, the product configuration process is modeled as a Markov decision process: the structural layer is divided into several structural groups according to function or structural characteristics. Each structural group contains a set of structural units to be configured. Starting from empty configuration, each structural group is configured and quantitatively evaluated in a predetermined order. 3) Using the Markov decision process as a reinforcement learning environment, a policy network is constructed as an agent, with the state vector as input and the selection probability of all possible actions in that state as output. In this environment, the configuration process from the initial state to the final state is repeatedly executed, and the policy network parameters are iteratively trained using a reinforcement learning method based on policy gradient. The trained policy network is used to configure structural units for each structure group starting from the initial state, thereby obtaining the product structure configuration scheme.
2. The method according to claim 1, characterized in that, Step 1) specifically includes: 1.1) Decompose the product's requirements, system functions, physical behavior, and structural units to determine the node sets of the requirement layer, functional layer, behavioral layer, and structural layer, which are used to describe the product's multi-layered semantic information; 1.2) Based on the semantic mapping relationship between each layer and nodes in the same layer in step 1.1), construct layer constraints including directed edges of rules, super edges of rules, and parameter sets of rule nodes to obtain an RFBS structure that reflects the multi-layer semantics and constraint relationship of the product, and formally define the composite rule graph of the product structure configuration.
3. The method according to claim 2, characterized in that, In step 1.1): The demand layer is used to describe product performance requirements, including at least one target value / range or normalized preference weight for a performance indicator. The functional layer is used to describe the key functions of the product; The behavior layer is used to describe the product's operational behavior; The structural layer is divided into several structural groups, and each structural group contains a set of structural units to be configured.
4. The method according to claim 2, characterized in that, In step 1.2): The directed edge according to the rule is specifically represented as: e k =(v i ,v j ,ω(v i ,v j )), v i ,v j ∈V represents the start and end points of the directed edges of the rules, respectively, and V is the set of nodes in the requirement layer, functional layer, behavioral layer, and structural layer. ω(v i ,v j Let ω be the edge weight function ω on the node pair (v) i ,v j The value at ) is used to quantize node v. i For node v j The weights of semantic dependencies or constraints are used as the path for decomposing and allocating high-level objectives to lower-level elements in the RFBS structure. The rule-defined hyperedge is specifically represented as: each rule-defined hyperedge e h e h =(N(e),τ(e),θ e ) , Let τ(e) be the set of nodes associated with the rule hyperedge, representing a combination constraint that associates at least two structural unit nodes or related cross-layer nodes; τ(e) is the type of rule hyperedge; θ e This is the vector of regular hyperedge parameters, used to describe the constraint weights or influence coefficients of the combined constraints on each performance dimension; The set of rule node parameters is specifically represented as: Π=Y∪Q, where Y={y u |u∈S str } represents the set of performance attribute vectors for structural units, used to reflect performance-related physical parameters; Q = {q u |u∈S str } represents the set of cost attribute vectors for structural units, used to reflect the values of structural units in terms of cost attributes.
5. The method according to claim 4, characterized in that, The rule-based superedge specifically includes: Mutual exclusion rule superedges are used to describe combinations of components that are not allowed to be used simultaneously. Dependency coexistence class rule hyperedges are used to describe combinations of components that must be used simultaneously; Collaborative classification rule hyperedges are used to describe the performance gain of a combination of components when they appear simultaneously. Conflict reduction rule superedges are used to describe the adverse effects on system performance when component combinations occur.
6. The method according to claim 5, characterized in that, Step 2) specifically includes: 2.1) Structural layer S str Based on functional or structural characteristics, the structures are divided into T groups, denoted as: Each structural group Corresponding to a set of structural units to be configured Grouped by the first structure Up to the Tth structural group For the predetermined configuration order; 2.2) Define the state space X: Denote any state vector in the state space as: x t =(U t ,Y(U t ),L(U t )), x t ∈X Among them, U t For the currently selected set of structural units, Y(U) t ) represents the performance attribute vector calculated based on the composite rule graph, L(U) t The cost attribute vector is calculated based on the composite rule graph; the system-level performance vector is calculated through the inter-layer mapping matrix of the composite rule graph, and the structural layer aggregation vector is normalized to obtain the relative importance weight of each structural unit. 2.3) Define the action space A: for any non-terminal state x t The following action a t ∈A(x t This represents the grouping of candidate actions within the current structure after performing validity checks based on a composite rule graph and pruning the candidate actions. The corresponding set of structural units Select one structural unit from the set and add it to the selected set U. t ; 2.4) Define the state transition function P: for any non-terminating state x t Execute action a t Then, the next state x is obtained. t+1 : Among them, U t+1 =U t ∪{u} represents the updated set of selected structural units; Y(U t+1 ) and L(U t+1 () are sets U based on composite rule graphs. t+1 The recalculated performance and cost attribute aggregation vector; when t = T, the x obtained after the transition T+1 This indicates that the configuration has been terminated. 2.5) Define the reward function R: Define the reward function as follows: For all non-terminating states, the reward function value is 0; when the configuration process terminates, the state is x. T+1 And obtain the complete configuration scheme U T+1 When calculating the comprehensive evaluation score J(U) of the scheme, T+1 As a final reward.
7. The method according to claim 6, characterized in that, In step 2.2): From the composite rule graph, the directed edges of rules from the requirement layer to the functional layer, from the functional layer to the behavior layer, and from the behavior layer to the structure layer are extracted and organized into an inter-layer mapping matrix W. RF W FB W BS The aggregated representations of the requirement layer, functional layer, behavioral layer, and structural layer are denoted as vectors r, f, b, and s, respectively, such that the inter-layer mapping relationship satisfies: f=W RF r,b=W FB f,s=W BS b Normalizing the aggregation vector s of the structural layer yields the relative importance weights λ of the structural units. u , In any state x t Next, calculate the system-level performance aggregation vector for the current configuration: The cost attribute of structural unit u in the nth dimension is denoted as q. u,n (n = 1, ..., m), where m is the number of cost attributes; Choose the cost aggregation function f n (·) Perform aggregation to obtain the current set of selected structural units U. t The aggregation result of the nth cost attribute l n (U t )=f n ({q u,n |u∈U t }), thus obtaining the cost attribute aggregate vector: L(U t )=(l1(U t ),...,l m (And t ))。 8. The method according to claim 6, characterized in that, In step 2.3): The candidate action pruning specifically involves: in any state x t Next, for each candidate structural unit u in the current structural group, construct a temporary configuration set U. tmp =U t ∪{u}; Calculate L(U) tmp ) and compare with the cost attribute constraints. If the current action exceeds the allowable range in any cost attribute dimension, remove the current action from the set of actionable actions A(x). t Delete it; Retrieve mutually exclusive rule superedges in the composite rule graph, when U tmp When all structural units associated with a certain mutually exclusive rule superedge are already included, the current action is removed from the action set A(x). t Delete it; In the composite rule graph, retrieve dependency coexistence rule hyperedges and remove candidate structural units from subsequent structural groups that are incompatible with the dependency relationship. If the set of candidate structural units in a subsequent structural group is empty after constraint search, remove the current action from the set of available actions A(x). t ) was deleted.
9. The method according to claim 6, characterized in that, In step 2.5): The comprehensive score calculation involves: after completing all structural grouping configurations, checking the hyperedge parameters corresponding to the hyperedges of collaborative addition and conflict subtraction rules in the composite rule graph, and calculating the performance correction amount ΔY(U). T+1 ), and correct the final system-level performance aggregation vector, Y'(U T+1 )=Y(U T+1 )+ΔY(U T+1 Construct a comprehensive scoring function: J(U T+1 )=λ cost J cost (IN T+1 )+λ match J match (IN T+1 )+λ mag J mag (IN T+1 ) Among them, J cost For L(U) T+1 The cost attribute utilization score obtained from the cost constraint; J match Based on Y'(U T+1 ) and demand R req The score for the degree of matching between the indicators; J mag Based on Y'(U T+1 Overall performance strength score; λ cost +λ match +λ mag =1 is the weighting coefficient.
10. The method according to claim 6, characterized in that, Step 3) specifically includes: 3.1) Constructing a reinforcement learning environment: The Markov decision process described above is used as the reinforcement learning environment. This environment receives the current state x... t Then, based on the composite rule graph, the legality of the candidate action set is judged and pruned, and the set of available actions A(x) in the current state is returned. t Upon receiving action a t Then, based on the state transition function P(x) t+1 |x t ,a t Return to the next state x t+1 And based on the reward function R(x) t ,a t ,x t+1 Return the reward; 3.2) Constructing the intelligent agent: Constructing an agent based on the state vector x t Given the input, the set of actions A(x) that can be performed in this state. t The policy network in which the probability of choosing each action is the output is the policy network. As a reinforcement learning agent, in which This is the parameter vector of the policy network; 3.3) Training the policy network: The agent is run in the reinforcement learning environment for multiple rounds of iterative training. Each round of training starts from the initial state x1 and makes T decisions to obtain the final state x. T+1 With reward R T =R(x) T ,a T ,x T+1 ); In step t, the policy network receives x returned by the environment. t With A(x) t After that, output A(x) t The probability of choosing each action in ) And select action a according to probability. t Execution, calculated by the environment, returns the next state x. t+1 With the action set A(x) t+1 ); To maximize reward R T The expected value is the objective, and a policy gradient-based reinforcement learning method is used to optimize the policy network parameters. Update; When the improvement in the overall score during training is less than a preset threshold ε for M consecutive evaluation periods, the policy network is considered to have converged, and the optimal policy network is obtained. 3.4) Configuration Inference: This involves configuring the trained, converged policy network. As a decision-making strategy; for a new configuration task, starting from the initial state x1, at each step, select the action with the highest probability, i.e. Structural units are selected sequentially for each structural group until a complete configuration scheme U is formed. T+1 .