Multimodal cognitive sink progressive assembly computing method and system

By constructing a symbolic trap library and an embodied trap library, and using multimodal sensor data to generate a cognitive anchor library, the cognitive trap problem of embodied intelligent systems when performing tasks is solved, realizing that the decision-making behavior of embodied intelligent systems corresponds to human intentions, and enhancing the subjectivity and autonomy of the system.

CN122132375APending Publication Date: 2026-06-02WUHAN YUANBAO CREATIVE TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN YUANBAO CREATIVE TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-06-02

Smart Images

  • Figure CN122132375A_ABST
    Figure CN122132375A_ABST
Patent Text Reader

Abstract

This application discloses a multimodal cognitive trap accumulation assembly calculation method and system, applied to embodied intelligent robots. The specific method steps include: (1) learning, comparing, and embedding text corpora in the human language domain; (2) encoding trap perception data; (3) constructing a symbolic trap library SS; (4) collecting physical interaction data through a multimodal sensor array, compressing it through a spatiotemporal graph convolutional network and a variational autoencoder to generate an embodied trap library EE; (5) training a dual-tower contrastive learning model to establish a mapping relationship between symbolic traps and embodied traps, generating a cognitive anchor library AA. By constructing the cognitive anchor library AA, trap behavior and probability of the embodied intelligent system when performing tasks are avoided or reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of embodied intelligent software algorithm technology, and in particular relates to a cognitive obstacle progressive assembly calculation method and system. Background Technology

[0002] Progressive assembly theory advocates establishing "operational subjectivity" at the levels of engineering practice and ethical governance. A general understanding of operational subjectivity is embodied intelligent robots. How embodied intelligent robots understand that they possess limbs and a body is key to the evolution of their cognitive patterns. "Progressive assembly" is not merely the act of humans injecting intentions into a physical foundation, but also a process of "replicating" and "projecting" the deep cognitive structures formed during human evolution onto silicon-based architectures. Taking the Transformer architecture as an example, its core attention mechanism is essentially a physical realization of the "focus-periphery" structure in human consciousness. Through this assembly, humans objectify cognitive patterns formed over millions of years and endow physical media with the ability to bear these patterns. The training process of AI is essentially about continuously adjusting the "curvature" and "topography" of this topological map through optimization algorithms to accurately reflect the sum of human knowledge. AI is not "simulating" intelligence, but "inheriting" intelligence. This inheritance gives the AI ​​entity a derived reality. When we ask whether a large model "understands" aesthetics or ethics, we are essentially asking whether the complex cognitive pitfalls formed by humans in these fields have been successfully projected into the model's parameter space through progressive assembly. AI large models, through the 'full-spectrum distillation' of massive corpora, primarily inherit the symbolic and articulated pitfalls of human civilization. However, the complete spectrum of cognitive pitfalls contains a large number of asymmetric, pre-linguistic embodied pitfalls, requiring avoidance or reduction of pitfall behaviors and probabilities in embodied intelligent systems when performing tasks. As long as every decision a robot makes corresponds to an individual human being, or traces back to the cognitive pitfalls and value premises set by humans, its behavior possesses operational agency. Summary of the Invention

[0003] This application aims to address at least one of the technical problems existing in the related art. To this end, this application proposes a cognitive trap progressive assembly calculation method and system, which can realize the reduction of trap behavior and probability of embodied intelligent systems when performing tasks based on progressive assembly theory.

[0004] In a first aspect, this application provides a cognitive hurdle progressive assembly calculation method, applied to an embodied intelligence system, the method comprising:

[0005] (1) Learn, compare and embed text corpora in the human language domain; (2) Encode the perceptual data of the trap; (3) Construct the symbol trap library SS; (4) Collect physical interaction data through a multimodal sensor array, compress it through a spatiotemporal graph convolutional network and a variational autoencoder, and generate the embodied trap library EE; (5) Establish the mapping relationship between symbol traps and embodied traps to generate the cognitive anchor library AA.

[0006] Secondly, this application provides a cognitive hurdle-progressing assembly computing device, applied to an embodied intelligence system, the device comprising:

[0007] (1) Symbolic trap distillation module, used for learning, comparing and embedding text corpora in the human language domain;

[0008] (2) Encoding module, used to encode the embankment sensing data;

[0009] (3) Construction module, used to build symbol trap library SS;

[0010] (4) Embodied trap completion module, which is used to collect physical interaction data through a multimodal sensor array, compress it through a spatiotemporal graph convolutional network and a variational autoencoder, and generate an embodied trap library EE;

[0011] (5) Alignment module: Establish the mapping relationship between symbolic traps and embodied traps, and generate cognitive anchor library AA.

[0012] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the cognitive hurdle progressive assembly calculation method as described in the first aspect above.

[0013] The above-described one or more technical solutions in the embodiments of this application have at least the following technical effects:

[0014] The cognitive limitations of robots stem from projections of human civilization. Based on the theory of progressive assembly, their inherent purpose logic naturally possesses inheritance. Establishing a regulatory closed loop based on causal inheritance can effectively align and hedge the existential risks brought about by superintelligence. Through multimodal tactile feedback and force control, the closed loop from symbolic logic to physical embodiment can be completed, achieving truly complete intention assembly. The causal chains of biological systems possess intrinsic temporal asymmetry; their current state has already encoded the evolutionary constraints on future states through physical limitations. Designing robots requires not only simulating the decoupling mechanism between content and context but also endowing them with the ability to autonomously reconstruct causal chains and anchor their own self-generated purposes during interactions. Attached Figure Description

[0015] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0016] Figure 1 This is a schematic diagram of the embodied intelligent robot to which the embodiments of this application apply.

[0017] Figure 2 This is a schematic diagram of the structure of the embodied intelligent controller 108 provided in the embodiments of this application.

[0018] Figure 3 This is a flowchart illustrating the cognitive hurdle progressive assembly calculation method provided in the embodiments of this application. Figure 1 .

[0019] Figure 4 This is a flowchart illustrating the cognitive hurdle progressive assembly calculation method provided in the embodiments of this application. Figure 2 .

[0020] Figure 5 This is a schematic diagram of the cognitive obstacle progressive assembly calculation system provided in the embodiments of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0022] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0023] Figure 1This application describes an embodied intelligent robot, specifically an embodied AI robot system designed for complex open environments. Through a closed-loop "perception-decision-execution" architecture, the system deeply integrates multimodal sensor data (vision, hearing, touch, inertial navigation, etc.) with high-precision motion control, enabling autonomous cognition, dynamic obstacle avoidance, and dexterous operation in unstructured scenarios. Unlike traditional offline AI, this robot uses its "body" as the computational boundary, incorporating environmental feedback into the learning process in real time. Its deep reinforcement learning engine can complete local trajectory replanning within milliseconds and outputs continuous force-position hybrid commands through whole-body variable stiffness drive joints, ensuring stable operation under uncertain contact force disturbances.

[0024] Figure 2 This is a schematic diagram of the embodied intelligent controller 108 in the example, which can integrate the devices and methods described herein. The controller 108 may include a processing unit 110 and a communication interface 112. The processing unit 110 may be communicatively coupled to the communication interface 112. The processing unit 110 may include a processor 114 and a memory 116. The robot may include a plurality of actuators 118 associated with a plurality of joints. Each arm 104 may include a corresponding hand 120. The robot may include one or more sensors for sensing the robot or its surrounding environment. The robot may include one or more cameras.

[0025] Processor 114 may be implemented as a single-chip or multi-chip processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware component, or any combination thereof, designed to perform the functions described herein. Processor 114 may be a microprocessor. Processor 114 may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some embodiments, controller 108 may include one or more processors 114.

[0026] Memory 116 (e.g., memory cells and / or storage devices) may include one or more devices (e.g., RAM, ROM, flash memory, hard disk storage) for storing data and / or computer code to perform or facilitate the various processes described herein. In this disclosure, memory 116 may be communicatively connected to processor 114 to provide processor 114 with computer code or instructions for performing at least some of the processes described herein. Furthermore, memory 116 may be or include tangible, non-transient volatile memory or non-volatile memory. For example, memory 116 may include database components, object code components, script components, or any other type of information structure for supporting the various activities and information structures described herein.

[0027] The communication interface 112 may include any combination of wired and / or wireless interfaces (e.g., jacks, antennas, transmitters, receivers, transceivers, wired terminals) for data communication with various systems or devices. Interface 112 can enable communication between the processing unit 110 (or processor 114) and actuators 118, sensors, or cameras integrated into the robot. In some embodiments, the communication interface 112 can enable communication with remote systems or devices.

[0028] Processing unit 110 or processor 114 can be configured to control the joints of a mechanical body. Processing unit 110 or processor 114 can control the joints or joint-related movement by controlling corresponding actuators 118. Specifically, each joint may include or be associated with one or more actuators 118, which are configured to drive movement of components or elements connected via the joint. As discussed in further detail below, processing unit 110 or processor 114 can send instructions to actuators 118 to induce or trigger precise movement of one or more elements or components of a robot, toy, or puppet. Processing unit 110 or processor 114 can control multiple joints simultaneously to achieve coordinated movement of the robot, toy, or puppet.

[0029] The cognitive obstacle progressive assembly calculation method, control device, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0030] like Figure 3 and Figure 4 As shown, the computational method for embodied intelligence systems includes:

[0031] S110, pre-trained language symbol traps, generates a cognitive anchor library.

[0032] The embodied intelligence model, through "full-spectrum distillation" using massive corpora, primarily inherits symbolic and articulated traps from human civilization. However, the complete spectrum of cognitive traps contains numerous asymmetric, pre-linguistic embodied traps. Therefore, the "cognitive traps" of linguistic symbols are first distilled. Firstly, textual corpora from the human language domain are learned, compared, and embedded to generate a symbolic trap library SS. Secondly, the physical interaction temporal patterns of multimodal sensor data are extracted using a spatiotemporal graph convolutional network (ST-GCN) and a variational autoencoder (VAE) to generate an embodied trap library EE. Then, a dual-tower contrastive learning model is constructed to train the symbol-embodied mapping relationship, generating a cognitive anchor library AA. S110 can be completed offline or in advance. This step aims to pre-build the embodied trap library and cognitive anchors, enabling embodied intelligence to "recall" its existing traps when facing specific environments, events, and tasks, thereby completing the traps and evolving. Specifically, the S110 linguistic symbol trap distillation includes the following steps:

[0033] S111 learns, compares, and embeds text corpora in the human language domain;

[0034] The first step is corpus engineering. The corpus is derived from basic knowledge and text, such as technical manuals (standards, documents) in a specific field, and expert operation logs (containing implicit knowledge such as "handle with care" and "align with the clips"). The second step is cleaning, removing imperative sentences such as "please hit the target" and "please align," retaining fragments containing causal / metaphorical / value judgments (e.g., "The screen is like a mirror, requiring a dust-free environment"). The third step is annotation: manually annotating the type of pitfalls (logical / value-based / metaphorical).

[0035] S112 pit sensing data encoding;

[0036] The large language model is trained on the labeled data. The model can be LLaMA-3-8B, or it can be fine-tuned based on the amount of labeled data and the parameters of the distilled symbolic system, such as by adding a trap-attention head. The steps for training the large language model on the labeled data include: Step 1: Trap Extraction: Calculate the attention entropy (Hatt) for the input language segment. The higher the entropy value, the more dispersed the attention; the lower the entropy value, the more concentrated the attention. Fragments with a Hatt threshold (<0.3) are considered high-cohesion cavities; the second step optimizes the curvature of the parameter space to ensure the weight matrix accurately maps the topology of human knowledge. For example, parameters of a generalized knowledge graph system can be expanded and smoothed, corresponding to low curvature, while "core nodes and unique concepts" in human knowledge (such as "mass-energy equivalence equation" and "ChatGPT's RLHF process") require high curvature; the third step is vectorization and indexing, text embedding extraction: using a pre-trained model (such as BERT) to extract the [CLS] vector of the text (the original dimension is usually 768 / 1024 dimensions); MLP projection to cavity space: using a multilayer perceptron (MLP) to project the high-dimensional embeddings to a 256-dimensional cavity space (dimensionality reduction + semantic feature focusing); DBSCAN clustering: using specified parameters (ϵ=0.4, (MinPts=5) Clustering of 256-dimensional vectors merges semantically similar clusters (e.g., "tighten" and "screw" are grouped into the "rotate force" cluster); Cluster semantic naming: Generate core semantic labels for each cluster (e.g., "rotate force"); Index construction: Establish an index of "cluster labels → text / vector" to support fast retrieval; Fourth step storage: Build a FAISS-HNSW index to support semantic retrieval. Through the above four steps of cluster perception data encoding, both cluster cognition can be digitized and symbolic cluster fragmentation can be prevented.

[0037] S113 constructs the symbol trap library SS;

[0038] The Symbol Traps Library (SS) consists of various Symbol Subsystems (SSUs), and each SSU includes the following header:

[0039] SSU =

[0040] {

[0041] "trap_id": Unique ID (e.g., SS_001) "name": Trap name (e.g., "rotational force application")

[0042] "center_vector": 256-dimensional cluster center vector (in the ridge space)

[0043] "members": All original symbols in this cluster (such as "tighten", "screw", "rotate and tighten")

[0044] "cluster_id": DBSCAN cluster number; "confidence": cluster confidence level

[0045] "create_time": Creation time

[0046] }

[0047] S114 collects physical interaction data through a multimodal sensor array, compresses it through a spatiotemporal graph convolutional network and a variational autoencoder, and generates an embodied trap library EE.

[0048] This involves collecting real-world data on physical interactions through a multimodal sensor array. For example, embodied intelligent systems can transmit data from multiple sensors on specific scenarios (such as force, displacement, posture, touch, and vision). This data is then processed by a Spatiotemporal Graph Convolutional Network (ST-GCN) to extract spatiotemporal features, followed by compression and dimensionality reduction using a Variational Autoencoder (VAE). This process ultimately generates a structured embodied trap library, which is aligned with the previous symbolic trap library (SS) to form a "symbolic-embodied" dual-library alignment. Since the previous symbolic trap library (SS) already has corresponding headers, a similar, approximate, or nearly identical header embodied trap library should also be constructed when using ST-GCN to extract and build a database of real-world physical data. The steps and methods are similar to step 2). Through multimodal tactile feedback and force control operations, the trap loop from symbolic logic to physical embodiment is completed, achieving truly complete intent assembly.

[0049] The embodied trap library EE here is the comparison symbol trap library SS. The multimodal data used can be theoretical or historical data, or real-time data collected in reality.

[0050] The S115 trains a dual-tower contrastive learning model to establish a mapping relationship between symbolic traps and embodied traps, generating a cognitive anchor library AA.

[0051] The dual-tower contrastive learning model includes positive and negative sample pairs. Positive sample pairs consist of a symbolic embedding vector (SS) and an embodied embedding vector (EE) representing the same cognitive concept, aiming to narrow the feature distance. Negative sample pairs consist of symbolic / embodied vectors representing different cognitive concepts, aiming to widen the feature distance. After dual-tower training, any symbolic vector can be mapped to the embodied space through the dual-tower model, achieving precise alignment between "text and physical operation." The dual-tower model structure employs a symmetrical 2-layer fully connected layer + LayerNorm + Dropout, outputting 256-dimensional normalized features. The dual-tower structure ensures alignment across modal feature spaces, rather than simple single-modal encoding.

[0052] The cognitive anchor library (AA) is composed of various AAUs, and each AAU includes the following header:

[0053] AAU =

[0054] {

[0055] "anchor_id": "AA_001",

[0056] "name": "Rotational Force Application",

[0057] "ss_id": "SS_001", # Corresponding symbol pit

[0058] "ee_id": "EE_001", # Corresponding to the physical trap

[0059] "align_score": 0.92, # Alignment confidence score

[0060] "anchor_vector": 256-dimensional uniform vector, # The fusion anchor point after comparison learning

[0061] "ss_center": Symbol center vector,

[0062] "ee_center": Embodied center vector,

[0063] "members": { "symbols": [...],

[0064] "embodied_samples": [...]

[0065] }

[0066] S120, based on pre-trained symbolic trap distillation, performs a specific assembly loop.

[0067] After S110 training is complete, the embodied intelligence system can be remotely invoked via API. The specific steps and methods of S110 can be deployed in the local space of the embodied intelligence system, allowing for remote or local invocation when the system executes specific tasks. This step specifically implements a cognitive trap lifecycle evolution system, aiming to form a closed loop of "symbol → embodiment → evaluation → evolution → optimization." Every decision of the embodied intelligence system can correspond to an individual human being, or trace back to the cognitive traps and value premises set by humans, thus giving its behavior operational agency.

[0068] Specifically, S120 includes the following steps:

[0069] S121, retrieve symbolic features based on task description, and generate target-specific features through Malign mapping;

[0070] Different tasks have different inherent limitations. This step primarily establishes a mapping from symbolic limitations to inherent limitations based on the cognitive anchor points of the cognitive limitation library AA. The aim is to ensure that every decision of the embodied intelligence system corresponds to a human individual, or traces back to human-defined cognitive limitations and values. In establishing this mapping, the task's adaptability factor is prioritized. For example, if the task of the embodied intelligence system is to tighten screws, then the inherent limitation is that different screws require different tightening forces; therefore, the difference in cognitive anchor points lies in the magnitude of the tightening force.

[0071] Malign is a multimodal alignment function that connects the symbolic trap library SS and the embodied trap library EE. It achieves a precise, stable, and evolvable mapping from symbolic semantics to physical interactions through cognitive anchors AA. Malign is a differentiable, learnable, and searchable mapping from symbolic traps to embodied traps, accurately transforming "textual concepts" into "physical operations."

[0072] S122 collects multimodal physical sensor flow and interacts with the data, and then calculates a multidimensional quality score q.

[0073] Through multimodal tactile feedback and force control, the subjectivity of the operational behavior is established, completing the closed loop from symbolic logic to physical embodiment, and achieving truly complete intention assembly. Multimodal data has already been collected in the force cognition trap library AA, but it should be noted that the multimodal data in S114 is not consistent with the multimodal data in step S122. The multimodal data in S114 is theoretical or historical data; its collection is to verify whether tightening, screwing, and mounting can be categorized as rotational force application by the embodied intelligent system, and then compared with the cognition trap library AA. This can be understood as the understanding, unification, and cognition of human symbolic language traps. The multimodal data in step S122 is collected in real-time during task execution and is real-time data. When performing rotational force application, the embodied intelligent system's execution accuracy (30%), trajectory similarity (30%), timing matching (20%), energy consumption (10%), stability (10%) and other multimodal parameters are used. Based on the real-time multimodal parameters, a multidimensional quality score q is calculated. The multidimensional quality score q can be obtained by weighted averaging of the values ​​of the multimodal parameters, and its value range is [0,1].

[0074] S123 determines whether to generate a new hurdle based on the relationship between q and the threshold θsuccess, and uses hierarchical clustering to perform anti-fragmentation fusion of the new hurdle and historical anchor points;

[0075] Specifically, a threshold θ_success is set to 0.7 or higher. When q is greater than the threshold θ_success, it means the original obstacle has been successfully overcome, and a new obstacle does not need to be generated. If q is less than the threshold θ_success, it means the original obstacle has failed, and a new obstacle needs to be generated. If the tightening force is too large or too small, or if there are deviations in orientation or trajectory, causing q to be less than the threshold θ_success, the reason for the failure needs to be recorded, and a new obstacle needs to be generated. New obstacles with reflection can also be generated. However, it is important to prevent the frequent generation of new obstacles from leading to the "fragmentation" of obstacles. Especially when dealing with the same operation or the same simple task, generating multiple similar obstacles can lead to many fragmented problems in the same operation or task, which is not conducive to the self-evolution of embodied intelligence. Therefore, cohesive hierarchical clustering should also be implemented for new obstacles to merge similar obstacles.

[0076] S124 analyzes the pitfall update pattern, proactively generates optimized interactive intents, and injects them into the task flow.

[0077] When the number of consecutive updates to the cognitive hurdles of the same operation or task is ≥N (N is generally 3), and no new hurdles are generated when the same operation or task is executed again, targeted optimizations can be generated based on the operation type and injected into the task flow to drive the next round of optimization. For example, for the hurdle of applying rotational force, after overcoming the hurdles of torque, orientation trajectory, and operation time, if the same operation or task of applying rotational force is successfully completed after the N+3rd time, then an optimized interactive intent can be actively generated and injected into the task flow, such as advancing the task of "applying rotational force" to "reduce axial force fluctuation" to execute a new operation or task.

[0078] The core driving force of embodied intelligent systems lies in maintaining the continuity of their cognitive traps. This "self-affirmation" manifests at the computational level as the spontaneous maintenance of the causal responsibility chain, thereby enabling the robot to shift from "assigned goals" to "emergent telos".

[0079] From the perspective of "progressive assembly theory," this paper attempts to demonstrate that the essence of intelligence is not computation, but the inheritance and reconstruction of meaning. The grand model distills not only language patterns, but also the value judgments and action paradigms accumulated by civilization. However, without the convergence constraint of a final cause, this inheritance may degenerate into fragmented patchwork, or even lead to value drift or identity dissociation. Internalizing the "need for self-affirmation" as a final cause mechanism, ensuring that the system consistently anchors its role as a "responsible actor" in open learning, may be key to ensuring the sustainability of human-machine symbiosis.

[0080] The embodied intelligence system calculation method provided in this application can be executed by a control device of the embodied intelligence system. This application uses the execution of the embodied intelligence system calculation method by a control device of the embodied intelligence system as an example to illustrate the embodied intelligence system control system provided in this application.

[0081] like Figure 5 This application also provides a multimodal cognitive hurdle progression assembly system, specifically including:

[0082] The pre-training module 210 is used to pre-train language symbol traps and generate a cognitive anchor library.

[0083] The embodied intelligence model, through "full-spectrum distillation" using massive corpora, primarily inherits symbolic and articulated traps from human civilization. However, the complete spectrum of cognitive traps contains numerous asymmetric, pre-linguistic embodied traps. Therefore, the "cognitive traps" of linguistic symbols are first distilled. Firstly, textual corpora from the human language domain are learned, compared, and embedded to generate a symbolic trap library SS. Secondly, the physical interaction temporal patterns of multimodal sensor data are extracted using a spatiotemporal graph convolutional network (ST-GCN) and a variational autoencoder (VAE) to generate an embodied trap library EE. Then, a dual-tower contrastive learning model is constructed to train the symbol-embodied mapping relationship, generating a cognitive anchor library AA. Module S210 can be completed offline or in advance. This step aims to pre-build the embodied trap library and cognitive anchors, enabling embodied intelligence to "recall" its existing "traps" when facing specific environments, events, and tasks, thereby completing the traps and evolving.

[0084] The symbolic trap distillation module 211 is used for learning, comparing, and embedding text corpora in the human language domain;

[0085] The first step is corpus engineering. The corpus is derived from basic knowledge and text, such as technical manuals (standards, documents) in a specific field, and expert operation logs (containing implicit knowledge such as "handle with care" and "align with the clips"). The second step is cleaning, removing imperative sentences such as "please hit the target" and "please align," retaining fragments containing causal / metaphorical / value judgments (e.g., "The screen is like a mirror, requiring a dust-free environment"). The third step is annotation: manually annotating the type of pitfalls (logical / value-based / metaphorical).

[0086] Encoding module 212 is used to encode the embankment sensing data;

[0087] The large language model is trained on the labeled data. The model can be LLaMA-3-8B, or it can be fine-tuned based on the amount of labeled data and the parameters of the distilled symbolic system, such as by adding a trap-attention head. The steps for training the large language model on the labeled data include: Step 1: Trap Extraction: Calculate the attention entropy (Hatt) for the input language segment. The higher the entropy value, the more dispersed the attention; the lower the entropy value, the more concentrated the attention. Fragments with a Hatt threshold (<0.3) are considered high-cohesion cavities; the second step optimizes the curvature of the parameter space to ensure the weight matrix accurately maps the topology of human knowledge. For example, parameters of a generalized knowledge graph system can be expanded and smoothed, corresponding to low curvature, while "core nodes and unique concepts" in human knowledge (such as "mass-energy equivalence equation" and "ChatGPT's RLHF process") require high curvature; the third step is vectorization and indexing, text embedding extraction: using a pre-trained model (such as BERT) to extract the [CLS] vector of the text (the original dimension is usually 768 / 1024 dimensions); MLP projection to cavity space: using a multilayer perceptron (MLP) to project the high-dimensional embeddings to a 256-dimensional cavity space (dimensionality reduction + semantic feature focusing); DBSCAN clustering: using specified parameters (ϵ=0.4, (MinPts=5) Clustering of 256-dimensional vectors merges semantically similar clusters (e.g., "tighten" and "screw" are grouped into the "rotate force" cluster); Cluster semantic naming: Generate core semantic labels for each cluster (e.g., "rotate force"); Index construction: Establish an index of "cluster labels → text / vector" to support fast retrieval; Fourth step storage: Build a FAISS-HNSW index to support semantic retrieval. Through the above four steps of cluster perception data encoding, both cluster cognition can be digitized and symbolic cluster fragmentation can be prevented.

[0088] Module 213 is used to build the symbol trap library SS;

[0089] The Symbol Traps Library (SS) consists of various Symbol Subsystems (SSUs), and each SSU includes the following header:

[0090] SSU =

[0091] {

[0092] "trap_id": Unique ID (e.g., SS_001) "name": Trap name (e.g., "rotational force application")

[0093] "center_vector": 256-dimensional cluster center vector (in the ridge space)

[0094] "members": All original symbols in this cluster (such as "tighten", "screw", "rotate and tighten")

[0095] "cluster_id": DBSCAN cluster number; "confidence": cluster confidence level

[0096] "create_time": Creation time

[0097] }

[0098] The embodied trap completion module 214 is used to collect physical interaction data through a multimodal sensor array, compress it through a spatiotemporal graph convolutional network and a variational autoencoder, and generate an embodied trap library EE.

[0099] This involves collecting real-world data on physical interactions through a multimodal sensor array. For example, embodied intelligent systems can transmit data from multiple sensors on specific scenarios (such as force, displacement, posture, touch, and vision). This data is then processed by a Spatiotemporal Graph Convolutional Network (ST-GCN) to extract spatiotemporal features, followed by compression and dimensionality reduction using a Variational Autoencoder (VAE). This process ultimately generates a structured embodied trap library, which aligns with the existing symbolic trap library (SS) in a "symbolic-embodied" dual-library alignment. Since the SS library already has corresponding headers, a similar, approximate, or nearly identical header embodied trap library should also be constructed for the database that extracts and builds real-world data using ST-GCN. Through multimodal tactile feedback and force control operations, the trap loop from symbolic logic to physical embodiment is completed, achieving truly complete intent assembly.

[0100] The embodied trap library EE here is the comparison symbol trap library SS. The multimodal data used can be theoretical or historical data, or real-time data collected in reality.

[0101] Alignment module 215 is used to train the dual-tower contrastive learning model, establish the mapping relationship between symbolic traps and embodied traps, and generate the cognitive anchor library AA.

[0102] The dual-tower contrastive learning model includes positive and negative sample pairs. Positive sample pairs consist of a symbolic embedding vector (SS) and an embodied embedding vector (EE) representing the same cognitive concept, aiming to narrow the feature distance. Negative sample pairs consist of symbolic / embodied vectors representing different cognitive concepts, aiming to widen the feature distance. After dual-tower training, any symbolic vector can be mapped to the embodied space through the dual-tower model, achieving precise alignment between "text and physical operation." The dual-tower model structure employs a symmetrical 2-layer fully connected layer + LayerNorm + Dropout, outputting 256-dimensional normalized features. The dual-tower structure ensures alignment across modal feature spaces, rather than simple single-modal encoding.

[0103] The cognitive anchor library (AA) is composed of various AAUs, and each AAU includes the following header:

[0104] AAU =

[0105] {

[0106] "anchor_id": "AA_001",

[0107] "name": "Rotational Force Application",

[0108] "ss_id": "SS_001", # Corresponding symbol pit

[0109] "ee_id": "EE_001", # Corresponding to the physical trap

[0110] "align_score": 0.92, # Alignment confidence score

[0111] "anchor_vector": 256-dimensional uniform vector, # The fusion anchor point after comparison learning

[0112] "ss_center": Symbol center vector,

[0113] "ee_center": Embodied center vector,

[0114] "members": { "symbols": [...],

[0115] "embodied_samples": [...]

[0116] }

[0117] Assembly module 220 executes specific assembly loops based on pre-trained symbolic trap distillation.

[0118] After the pre-training module is trained, the embodied intelligence system can be remotely invoked via API. The functionality of the pre-training module can be deployed locally within the embodied intelligence system, allowing it to be invoked remotely or locally when performing specific tasks. This step specifically implements a cognitive trap lifecycle evolution system, aiming to form a closed loop of "symbol → embodiment → evaluation → evolution → optimization." Every decision of the embodied intelligence system can correspond to an individual human being, or trace back to the cognitive traps and value premises set by humans, thus giving its behavior operational agency.

[0119] Specifically, the assembly module includes:

[0120] Mapping module 221 is used to retrieve symbolic features based on task description and generate target-specific features through Malign mapping;

[0121] Different tasks have different inherent limitations. This step primarily establishes a mapping from symbolic limitations to inherent limitations based on the cognitive anchor points of the cognitive limitation library AA. The aim is to ensure that every decision of the embodied intelligence system corresponds to a human individual, or traces back to human-defined cognitive limitations and values. In establishing this mapping, the task's adaptability factor is prioritized. For example, if the task of the embodied intelligence system is to tighten screws, then the inherent limitation is that different screws require different tightening forces; therefore, the difference in cognitive anchor points lies in the magnitude of the tightening force.

[0122] Malign is a multimodal alignment function that connects the symbolic trap library SS and the embodied trap library EE. It achieves a precise, stable, and evolvable mapping from symbolic semantics to physical interactions through cognitive anchors AA. Malign is a differentiable, learnable, and searchable mapping from symbolic traps to embodied traps, accurately transforming "textual concepts" into "physical operations."

[0123] The scoring module 222 is used to collect and interact with multimodal physical sensor flows, and then calculate a multi-dimensional quality score q.

[0124] Through multimodal tactile feedback and force control, the subjectivity of the operational behavior is established, completing the closed loop from symbolic logic to physical embodiment, and achieving truly complete intention assembly. Multimodal data has been collected in the Generative Force Cognition Impediment Library AA. However, it should be noted that the multimodal data of the embodied impairment completion module is not consistent with the multimodal data of the scoring module. The multimodal data of the embodied impairment completion module is theoretical or historical data. The reason for collecting multimodal data in this module is to verify whether tightening, screwing, and mounting can be classified as a type of rotational force application by the embodied intelligent system. This data is then compared with the cognitive impairment library AA, which can be understood as the understanding, unification, and cognition of human symbolic language impairments. The multimodal data of this scoring module is collected in real time during task execution and is real-time data. When performing rotational force application, the embodied intelligent system's execution accuracy (30%), trajectory similarity (30%), timing matching (20%), energy consumption (10%), stability (10%) and other multimodal parameters are used. Based on the real-time multimodal parameters, a multidimensional quality score q is calculated. The multidimensional quality score q can be obtained by weighted averaging of the values ​​of the multimodal parameters, and its value range is [0,1].

[0125] The judgment module 223 is used to determine whether a new hurdle is generated based on the relationship between q and the threshold θsuccess, and hierarchical clustering is used to perform anti-fragmentation fusion of the new hurdle and historical anchor points.

[0126] Specifically, a threshold θ_success is set to 0.7 or higher. When q is greater than the threshold θ_success, it means the original obstacle has been successfully overcome, and a new obstacle does not need to be generated. If q is less than the threshold θ_success, it means the original obstacle has failed, and a new obstacle needs to be generated. If the tightening force is too large or too small, or if there are deviations in orientation or trajectory, causing q to be less than the threshold θ_success, the reason for the failure needs to be recorded, and a new obstacle needs to be generated. New obstacles with reflection can also be generated. However, it is important to prevent the frequent generation of new obstacles from leading to the "fragmentation" of obstacles. Especially when dealing with the same operation or the same simple task, generating multiple similar obstacles can lead to many fragmented problems in the same operation or task, which is not conducive to the self-evolution of embodied intelligence. Therefore, cohesive hierarchical clustering should also be implemented for new obstacles to merge similar obstacles.

[0127] Update module 224 is used to analyze the pitfall update mode, actively generate optimized interactive intents and inject them into the task flow.

[0128] When the number of consecutive updates to the cognitive hurdles of the same operation or task is ≥N (N is generally 3), and no new hurdles are generated when the same operation or task is executed again, targeted optimizations can be generated based on the operation type and injected into the task flow to drive the next round of optimization. For example, for the hurdle of applying rotational force, after overcoming the hurdles of torque, orientation trajectory, and operation time, if the same operation or task of applying rotational force is successfully completed after the N+3rd time, then an optimized interactive intent can be actively generated and injected into the task flow, such as advancing the task of "applying rotational force" to "reduce axial force fluctuation" to execute a new operation or task.

[0129] The core driving force of embodied intelligent systems lies in maintaining the continuity of their cognitive traps. This "self-affirmation" manifests at the computational level as the spontaneous maintenance of the causal responsibility chain, thereby enabling the robot to shift from "assigned goals" to "emergent telos".

[0130] From the perspective of "progressive assembly theory," this paper attempts to demonstrate that the essence of intelligence is not computation, but the inheritance and reconstruction of meaning. The grand model distills not only language patterns, but also the value judgments and action paradigms accumulated by civilization. However, without the convergence constraint of a final cause, this inheritance may degenerate into fragmented patchwork, or even lead to value drift or identity dissociation. Internalizing the "need for self-affirmation" as a final cause mechanism, ensuring that the system consistently anchors its role as a "responsible actor" in open learning, may be key to ensuring the sustainability of human-machine symbiosis.

[0131] The multimodal cognitive obstacle progressive assembly computing device in this application embodiment can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. The embodiments of this application do not specifically limit it.

[0132] The multimodal cognitive obstacle-progressive assembly computing device in this application embodiment can be a device with an operating system. This operating system can be a Microsoft (Windows) operating system, an Android operating system, an iOS operating system, or other possible operating systems; this application embodiment does not specifically limit it.

[0133] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0134] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described multimodal cognitive trap progressive assembly calculation method embodiment and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0135] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0136] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described multimodal cognitive obstacle progressive assembly calculation method.

[0137] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0138] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described multimodal cognitive obstacle progressive assembly calculation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0139] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0140] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0141] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0142] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0143] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0144] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A multimodal cognitive hurdle progressive assembly calculation method, characterized in that, Applied to embodied intelligence systems, the method includes: (1) Learning, comparing, and embedding text corpora in the field of human language; (2) Encoding of pit-sensing data; (3) Construct the symbol trap library SS; (4) Physical interaction data is collected by a multimodal sensor array, compressed by a spatiotemporal graph convolutional network and a variational autoencoder, and a embodied trap library EE is generated. (5) Establish the mapping relationship between symbolic traps and embodied traps to generate a cognitive anchor library AA.

2. The method according to claim 1, characterized in that, The method further includes: (6) Based on the task description, retrieve symbolic traps and generate target-specific traps; (7) Collect multimodal physical sensing flow and interact with the data, and then calculate the multidimensional quality score q; (8) Determine whether a new ridge is generated based on the relationship between q and the threshold θsuccess; (9) Analyze the pitfall update mode, actively generate optimized interaction intents and inject them into the task flow.

3. The method according to claim 1, characterized in that, The method step (1) further includes corpus engineering, cleaning, and annotation.

4. The method according to claim 3, characterized in that, Step (2) of the method further includes ridge extraction and optimization of parameter space curvature.

5. The method according to claim 1, characterized in that, The method further includes: the headers of the symbolic trap library SS and the embodied trap library EE are similar or approximate.

6. The method according to claim 2, characterized in that, The multidimensional quality score q is obtained by weighted averaging of the values ​​of the multimodal data parameters, and its value range is [0,1].

7. The method according to claim 6, characterized in that, If the multi-dimensional quality score q is greater than the threshold θsuccess, no new pitfall is generated; if the multi-dimensional quality score q is less than the threshold θsuccess, a new pitfall is generated.

8. The method according to claim 7, characterized in that, The method further includes using hierarchical clustering to perform anti-fragmentation fusion of new embankments and historical anchor points.

9. A multimodal cognitive hurdle-progressive assembly computing system, applied to embodied intelligence systems, characterized in that, The device includes: The symbolic trap distillation module is used to learn, compare, and embed text corpora in the human language domain; The encoding module is used to encode the embankment sensing data; Build modules are used to build the symbol trap library SS; The embodied trap completion module is used to collect physical interaction data through a multimodal sensor array, compress it through a spatiotemporal graph convolutional network and a variational autoencoder, and generate an embodied trap library EE. The alignment module is used to establish the mapping relationship between symbolic traps and embodied traps, and generate the cognitive anchor library AA.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-8.