Robot multi-mode sensing fusion method and system for precise assembly
By using feature mapping and consistency scoring of multimodal perception data, the problems of insufficient perception conflict handling mechanism and single fusion strategy in robot precision assembly are solved, and highly robust assembly operation in complex environments is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZAOZHUANG UNIV
- Filing Date
- 2026-04-14
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, the lack of a perception conflict handling mechanism and the simplistic fusion strategy during the precision assembly process of robots result in insufficient robustness in complex environments.
By collecting multimodal perception data, extracting the temporal variation features of force perception data, mapping them to the same semantic feature space, calculating feature distance and generating cross-modal semantic consistency scores, selecting predefined fusion modes to fuse perception results, and triggering active perception correction when unreliable.
It achieves quantitative assessment of cross-modal perception conflicts and adaptive switching fusion strategies, which improves the reliability and robustness of assembly operations and reduces the risk of assembly failure due to perception errors.
Smart Images

Figure CN122058375A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control and intelligent sensing technology, and in particular to a multimodal sensing fusion method and system for robots oriented towards precision assembly. Background Technology
[0002] In the field of precision robotic assembly, single sensors (such as vision) are easily affected by factors such as lighting, occlusion, and high reflectivity, leading to perception failure. Therefore, fusing multimodal information such as vision and force perception has become the mainstream technological direction.
[0003] Current technologies have disclosed a two-stage strategy that uses visual guidance for coarse localization and force information for fine assembly, a dynamic weight adaptive fusion method, and a consistency score concept for weight allocation in multimodal post-fusion.
[0004] However, these technologies currently have the following shortcomings: Lack of perceptual conflict handling mechanism: When visual and force information conflict (for example, the visual display is aligned, but the force feedback shows abnormal contact), the existing technology lacks an effective quantification and decision-making mechanism, making it difficult to determine which modality should be used as the standard.
[0005] The current fusion methods are mostly based on continuous weight adjustment and fail to perform discrete strategy switching based on the reliability of the perceived information, resulting in insufficient robustness in complex assembly environments. Summary of the Invention
[0006] The purpose of this invention is to provide a multimodal perception fusion method and system for robots oriented towards precision assembly, which solves the problems of insufficient perception conflict handling mechanism and single fusion strategy in the current technology.
[0007] To achieve the above objectives, this invention provides a multimodal perception fusion method for robots in precision assembly, comprising the following steps: Collect multimodal perception data in the robot assembly environment; the multimodal perception data includes visual modal data and force modal data; Extract the temporal variation characteristics of force sensory modal data, and identify the current assembly stage based on the temporal variation characteristics; Feature extraction and preliminary perception processing are performed on the multimodal perception data to obtain the initial perception results of the visual modality and the initial perception results of the force modality. The initial perception results of the visual modality and the initial perception results of the force modality are mapped to the same semantic feature space, and the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space is calculated. A cross-modal semantic consistency score is generated based on the feature distance; the cross-modal semantic consistency score is negatively correlated with the feature distance. Based on the cross-modal semantic consistency score and the current assembly stage, a fusion mode is selected from the predefined fusion modes to fuse the initial perception results of the visual modality and the initial perception results of the force modality, generating a fused perception result; the predefined fusion modes include: collaborative fusion mode, hybrid fusion mode and single-modal dominant mode; Based on cross-modal semantic consistency scores and fusion perception results, robot assembly control instructions are generated.
[0008] Preferably, mapping the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, and calculating the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space, specifically includes: The pairing of the initial visual modal perception results and the initial force modal perception results acquired under the assembly alignment state is used as a positive sample pair, and the pairing of the initial visual modal perception results and the initial force modal perception results acquired under the assembly deviation state is used as a negative sample pair. The mapping function is determined with the optimization objective of minimizing the feature distance between positive sample pairs in the semantic feature space and maximizing the feature distance between negative sample pairs in the semantic feature space. A mapping function is used to map the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, resulting in visual projection vectors and force projection vectors; The feature distance is obtained by calculating the vector distance between the visual projection vector and the force projection vector.
[0009] Preferably, the formula for calculating the feature distance is: ; in, For visual projection vectors, For force projection vector, For feature distance, i For the dimension of the semantic feature space, i =1,2,…, d , d It is a positive integer. For the visual projection vector at the th i Dimensional components, For the force projection vector at the th i Components in a dimension.
[0010] Preferably, the expression for the cross-modal semantic consistency score is: ; in, s For cross-modal semantic consistency scoring, For visual projection vectors, For force projection vector, For feature distance, This is the preset maximum distance threshold.
[0011] Preferably, based on the cross-modal semantic consistency score and the current assembly stage, a fusion mode is selected from predefined fusion modes to fuse the initial perception results of the visual modality and the initial perception results of the force modality. The specific content of the generated fused perception result includes: The cross-modal semantic consistency score is compared with a first preset threshold and a second preset threshold to obtain the comparison result; the second preset threshold is less than or equal to the first preset threshold. Based on the comparison results and the current assembly stage, select a fusion mode from the predefined fusion modes: Based on the selected fusion mode, the initial perception results of the visual modality and the initial perception results of the force modality are fused to obtain the fused perception result; the fused perception result includes: pose information and force information.
[0012] Preferably, the specific details of selecting a fusion mode from predefined fusion modes based on the comparison results and the current assembly stage include: When the cross-modal semantic consistency score is higher than the first preset threshold, the collaborative fusion mode is selected; When the cross-modal semantic consistency score is lower than the second preset threshold, the single-modal dominant mode is selected; When the cross-modal semantic consistency score is between the second preset threshold and the first preset threshold, the hybrid fusion mode is selected.
[0013] Preferably, the collaborative fusion mode includes: calculating the cross-modal attention weights of the initial visual modality perception result and the initial force modality perception result through an attention mechanism, and performing weighted fusion of the initial visual modality perception result and the initial force modality perception result according to the cross-modal attention weights; The single-modal dominant mode includes: identifying modalities with low cross-modal semantic consistency scores as untrusted modalities, isolating untrusted modalities from the fusion path, and using only the perception results of trusted modalities as the fusion perception results; The hybrid fusion mode includes: maintaining parallel processing of the initial perception results of the visual modality and the initial perception results of the force modality, using a weighted average method for fusion, wherein the weight coefficients in the weighted average method are positively correlated with the cross-modal semantic consistency score, and an uncertainty label is added to the fused perception results.
[0014] Preferably, the specific content of the robot assembly control instructions generated based on the cross-modal semantic consistency score and fusion perception results includes: Obtain the trend of cross-modal semantic consistency scores over time; State identification based on changing trends: When the cross-modal semantic consistency score is detected to drop sharply from above the first preset threshold to below the second preset threshold, the assembly operation is determined to be in an abnormal state; when the cross-modal semantic consistency score is detected to remain stable within the range above the first preset threshold, the assembly operation is determined to be in a normal state; thus, the state identification result is obtained. Based on the state recognition results, combined with pose information and force information, corresponding robot assembly control instructions are generated: if the assembly operation is determined to be in an abnormal state, control instructions containing an abnormality handling subroutine are generated; if the assembly operation is determined to be in a normal state, force-position hybrid control instructions are generated based on pose information and force information; the abnormality handling subroutine includes pausing the action, returning to a safe position, and triggering active perception correction.
[0015] Preferably, the multimodal perception fusion method for precision assembly robots further includes: verifying the fused perception results and performing active perception correction; The fused perception results are matched and verified with a preset assembly knowledge base; the assembly knowledge base contains prior information on the expected perception features under standard assembly conditions. When the deviation between the fused perception result and the prior information exceeds a preset deviation threshold, the current fused perception result is determined to be unreliable. In response to the determination that the current fusion perception result is unreliable, an active perception behavior is triggered, which controls the robot's end effector to perform a trial action to obtain updated force modal data and controls the vision sensor to change its pose to obtain updated visual modal data from a new perspective. Based on the updated force modality data and the updated visual modality data, the corrected fusion perception results are regenerated.
[0016] This invention also provides a multimodal perception fusion system for robots in precision assembly, used to implement the aforementioned multimodal perception fusion method for robots in precision assembly, comprising: The data acquisition module is used to collect multimodal perception data in the robot assembly environment; the multimodal perception data includes visual modal data and force modal data. The feature recognition module is used to extract the temporal variation features of force sensory modal data and identify the current assembly stage based on the temporal variation features; The data processing module is used to perform feature extraction and preliminary perception processing on multimodal perception data to obtain initial perception results for visual modality and initial perception results for force modality. The mapping calculation module is used to map the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, and to calculate the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space. The scoring generation module is used to generate a cross-modal semantic consistency score based on feature distance; the cross-modal semantic consistency score is negatively correlated with feature distance. The perception fusion module is used to select a fusion mode from predefined fusion modes based on the cross-modal semantic consistency score and the current assembly stage, and fuse the initial perception results of the visual modality and the initial perception results of the force modality to generate a fused perception result; the predefined fusion modes include: collaborative fusion mode, hybrid fusion mode and single-modal dominant mode; The instruction generation module is used to generate robot assembly control instructions based on cross-modal semantic consistency scores and fusion perception results.
[0017] In summary, the multimodal perception fusion method and system for precision assembly robots provided by this invention have the following advantages compared to traditional technologies: the generated cross-modal semantic consistency score intuitively quantifies the degree of semantic consistency between visual perception and force perception, effectively handling the problem of perception conflict, and realizing the quantitative evaluation of cross-modal perception conflict in assembly scenarios; by selecting a predefined fusion mode and combining it with assembly stage recognition, it can decisively switch to a reliable single modality when perception conflict is severe, solving the problem of a single fusion strategy and avoiding the impact of low-quality data on the fusion results.
[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0019] Figure 1 This is a flowchart of a multimodal perception fusion method for robots aimed at precision assembly, as described in this invention. Figure 2 This is a block diagram of a multimodal perception fusion system for robots designed for precision assembly, as described in this invention. Detailed Implementation
[0020] The technical method of the present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of this application.
[0021] The following description of at least one exemplary embodiment is merely illustrative and is not intended to limit the scope of this application or its application or use.
[0022] Techniques, systems, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the instruction manual.
[0023] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0024] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0025] This invention provides a multimodal perception fusion method for robots in precision assembly, such as... Figure 1 As shown, it includes the following steps: S1. Collect multimodal perception data in the robot assembly environment. The multimodal perception data includes visual modal data and force modal data.
[0026] S2. Extract the temporal variation characteristics of the force sensory modal data, and identify the current assembly stage based on the temporal variation characteristics. The assembly stage includes the approach stage, contact stage, insertion stage, and positioning stage.
[0027] S3. Perform feature extraction and preliminary perception processing on the multimodal perception data to obtain the initial perception results of the visual modality and the initial perception results of the force modality.
[0028] S4. Map the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, and calculate the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space.
[0029] Furthermore, step S4 specifically includes the following steps: S401. The pairing of the initial visual modal perception results and the initial force modal perception results acquired under the assembly alignment state is taken as a positive sample pair, and the pairing of the initial visual modal perception results and the initial force modal perception results acquired under the assembly deviation state is taken as a negative sample pair.
[0030] S402. Determine the mapping function with the optimization objective of minimizing the feature distance between positive sample pairs in the semantic feature space and maximizing the feature distance between negative sample pairs in the semantic feature space.
[0031] S403. Using a mapping function, the initial perception results of the visual modality and the initial perception results of the force modality are mapped to the same semantic feature space to obtain the visual projection vector and the force projection vector.
[0032] S404. The feature distance is obtained by calculating the vector distance between the visual projection vector and the force projection vector. The formula for calculating the feature distance is: ; in, For visual projection vectors, For force projection vector, For feature distance, i For the dimension of the semantic feature space, i =1,2,…, d , d It is a positive integer. For the visual projection vector at the th i Dimensional components, For the force projection vector at the th i Components in a dimension.
[0033] S5. Generate a cross-modal semantic consistency score based on feature distance. The cross-modal semantic consistency score is negatively correlated with feature distance. The expression for the cross-modal semantic consistency score is: ; in, s For cross-modal semantic consistency scoring, This is the preset maximum distance threshold.
[0034] This invention maps the initial visual and force-sensing results to the same semantic feature space, calculates the feature distance between the two in the same semantic feature space, and generates a cross-modal semantic consistency score that can intuitively quantify the degree of semantic consistency between visual and force perception. When there is a conflict between the two modal perception results (such as visual display showing alignment but force feedback showing abnormal contact), the consistency score will be significantly reduced, providing a clear quantitative basis for subsequent fusion decisions. Moreover, compared with the current technology, which cannot effectively handle the problem of perception conflict, this invention achieves quantitative evaluation of cross-modal perception conflict in assembly scenarios.
[0035] S6. Based on the cross-modal semantic consistency score and the current assembly stage, select a fusion mode from the predefined fusion modes to fuse the initial perception results of the visual modality and the initial perception results of the force modality, generating a fused perception result. The predefined fusion modes include: collaborative fusion mode, hybrid fusion mode, and single-modal dominant mode.
[0036] Furthermore, step S6 can be replaced by the following steps: S601. Compare the cross-modal semantic consistency score with the first preset threshold and the second preset threshold respectively to obtain the comparison result. Wherein, the second preset threshold is less than or equal to the first preset threshold.
[0037] S602. Based on the comparison results and the current assembly stage, select a fusion mode from the predefined fusion modes. Specifically, step S602 includes: When the cross-modal semantic consistency score is higher than a first preset threshold, a collaborative fusion mode is selected. The collaborative fusion mode includes: calculating the cross-modal attention weights of the initial visual modality perception results and the initial force modality perception results through an attention mechanism, and performing weighted fusion of the initial visual modality perception results and the initial force modality perception results according to the cross-modal attention weights.
[0038] When the cross-modal semantic consistency score is lower than a second preset threshold, a single-modal dominant mode is selected. The single-modal dominant mode includes: identifying the modality with a low cross-modal semantic consistency score as an untrusted modality, isolating the untrusted modality from the fusion path, and using only the perception result of the trustworthy modality as the fusion perception result.
[0039] When the cross-modal semantic consistency score is between the second preset threshold and the first preset threshold, a hybrid fusion mode is selected. The hybrid fusion mode includes: maintaining the parallel processing of the initial visual modality perception results and the initial force modality perception results, using a weighted average method for fusion, wherein the weight coefficients in the weighted average method are positively correlated with the cross-modal semantic consistency score, and adding an uncertainty label to the fused perception result.
[0040] S603. Based on the selected fusion mode, the initial perception results of the visual modality and the initial perception results of the force modality are fused to obtain the fused perception result. The fused perception result includes pose information and force information.
[0041] This invention identifies the current assembly stage (approach stage, contact stage, insertion stage, and placement stage) based on the temporal variation characteristics of force modal data. Then, based on two dimensions—cross-modal semantic consistency score and the current assembly stage—it selects the optimal fusion strategy from three modes: collaborative fusion, hybrid fusion, and single-modal dominant fusion. Under different assembly stages (e.g., the insertion stage where vision is easily occluded) and different consistency scores, the system can adaptively adjust the fusion method: when the consistency score is high, collaborative fusion is used to fully exploit multimodal information; when the consistency score is low, a single-modal dominant fusion mode is used to avoid low-quality data affecting the fusion result. Furthermore, compared to the current single-weighted fusion method, the three-mode discrete switching strategy of this invention has stronger robustness and scene adaptability.
[0042] S7. Generate robot assembly control instructions based on cross-modal semantic consistency scores and fusion perception results.
[0043] Furthermore, step S7 specifically includes the following steps: S701. Obtain the trend of cross-modal semantic consistency scores over time.
[0044] S702. State recognition of change trend: When the cross-modal semantic consistency score is detected to drop sharply from a state higher than the first preset threshold to a state lower than the second preset threshold, the assembly operation is determined to be an abnormal state; when the cross-modal semantic consistency score is detected to remain stable in the range higher than the first preset threshold, the assembly operation is determined to be a normal state; thus, the state recognition result is obtained.
[0045] S703. Based on the state recognition result, combined with pose information and force information, generate corresponding robot assembly control commands: if the assembly operation is determined to be in an abnormal state, generate control commands containing an abnormality handling subroutine; if the assembly operation is determined to be in a normal state, generate force-position hybrid control commands based on pose information and force information. The abnormality handling subroutine includes pausing the action, returning to a safe position, and triggering active perception correction.
[0046] In an exemplary embodiment of the present invention, the multimodal perception fusion method for robots in precision assembly further includes: verifying the fused perception results and performing active perception correction. Further, the specific content of verifying the fused perception results and performing active perception correction includes: The fused perception results are matched and verified against a pre-set assembly knowledge base. This assembly knowledge base contains prior information about the expected perceived features under standard assembly conditions.
[0047] When the deviation between the fused perception result and the prior information exceeds a preset deviation threshold, the current fused perception result is determined to be unreliable.
[0048] In response to the determination that the current fusion perception result is unreliable, an active perception behavior is triggered, which controls the robot's end effector to perform exploratory actions to obtain updated force modal data and controls the vision sensor to change its pose to obtain updated visual modal data from a new perspective.
[0049] Based on the updated force modality data and the updated visual modality data, the corrected fusion perception result is regenerated. Specifically, the updated force modality data and the updated visual modality data are used as new inputs, and the multimodal perception fusion process of steps S1-S6 is re-executed to generate the corrected fusion perception result.
[0050] This invention verifies the fused perception results by matching them with a pre-set assembly knowledge base. When the fused results are unreliable, it triggers proactive sensing actions (performing exploratory actions to obtain updated force data and adjusting the pose of the visual sensors to obtain updated visual data), and regenerates a corrected fused perception result based on the updated data. This complete mechanism realizes a paradigm shift from passive fusion to proactive verification and correction, enabling the system to proactively acquire better data for correction when perception is unreliable, significantly improving the reliability and fault tolerance of assembly operations.
[0051] Another exemplary embodiment of the present invention takes a typical shaft-hole assembly task as an example, as follows: Assembly stage identification and cross-modal consistency score generation: The robot collects data in real time through a vision camera and a six-dimensional force sensor at its end. First, the assembly stage recognition module identifies the assembly stage based on the changing trends of the force data: approach stage (zero contact force), contact stage (force data shows significant changes for the first time), insertion stage (force data shows stable insertion characteristics), and positioning stage (force data reaches a preset threshold).
[0052] During the insertion phase (when vision is easily occluded), a cross-modal consistency assessment is initiated. It is assumed that visual perception indicates the axis is aligned with the hole center, but force perception indicates a lateral contact force at the axis-hole edge. In this case, the mapping function projects both onto a physical semantic space with positional deviation and contact force amplitude as dimensions. The large feature distance between the two results in a low consistency score, indicating a perceptual conflict.
[0053] Adaptive fusion strategy based on consistency score: Upon receiving a low consistency score, the adaptive fusion controller, considering the current insertion phase, determines that the score is below a second preset threshold and therefore selects the single-modal dominant mode. Since force perception is typically more reliable at this stage, the system sets the force perception modality as the reliable modality, isolating interference information from the visual modality and using only the force perception result for subsequent control. Conversely, if in the proximity phase, visual perception is clear and highly consistent with force perception (non-contact), the score is high, and the system selects the collaborative fusion mode, employing attention-based weighted fusion. If the score is moderate (e.g., in low-light environments), the system selects the hybrid fusion mode, allocating weights according to the score ratio and adding an uncertainty marker to the result.
[0054] Active perception and correction: After executing the single-modal dominant mode, the control command generation module generated a downward insertion command. However, the fusion result verification unit matched the current result with the assembly knowledge base and found abnormal fluctuations in the actual force characteristics, determining that the fusion result was unreliable.
[0055] At this time, the active perception correction module is triggered and generates active perception instructions: (1) control the robot to perform a small lifting-lowering probing action to obtain a more accurate contact state; (2) control the camera to move to another preset viewpoint to obtain new image data to verify whether there are burrs.
[0056] This new data was re-injected into the perception fusion process, the consistency score was recalculated, and the strategy was corrected (e.g., a rotation-insertion action was performed), successfully completing the assembly.
[0057] This demonstrates that the present invention, through its complete technical approach of quantifying conflicts via consistency scoring, addressing conflicts through discrete switching between three modes, and actively verifying and correcting results via closed-loop verification, effectively solves the perception failure problem of current technologies in complex assembly environments such as occlusion, high reflectivity, and changing lighting. Compared to current technologies, the present invention significantly improves the intelligence level, success rate, and system robustness of precision assembly tasks such as shaft-hole assembly and threaded connections, while reducing the risk of assembly failure due to perception errors.
[0058] This invention provides a multimodal perception fusion system for robots designed for precision assembly, used to implement the aforementioned multimodal perception fusion method for robots designed for precision assembly, such as... Figure 2 As shown, it includes: The data acquisition module is used to collect multimodal perception data from the robot assembly environment. This multimodal perception data includes visual modal data and force modal data.
[0059] The feature recognition module is used to extract the temporal variation features of the force sensory modal data and identify the current assembly stage based on the temporal variation features.
[0060] The data processing module is used to perform feature extraction and preliminary perception processing on multimodal perception data to obtain initial perception results for visual modality and initial perception results for force modality.
[0061] The mapping calculation module is used to map the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, and to calculate the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space.
[0062] The scoring generation module is used to generate a cross-modal semantic consistency score based on the feature distance; the cross-modal semantic consistency score is negatively correlated with the feature distance.
[0063] The perception fusion module is used to select a fusion mode from predefined fusion modes based on the cross-modal semantic consistency score and the current assembly stage, and to fuse the initial perception results of the visual modality and the initial perception results of the force modality to generate a fused perception result. The predefined fusion modes include: collaborative fusion mode, hybrid fusion mode, and single-modal dominant mode.
[0064] The instruction generation module is used to generate robot assembly control instructions based on cross-modal semantic consistency scores and fusion perception results.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multimodal perception fusion method for robots in precision assembly, characterized in that, Includes the following steps: Collect multimodal perception data in the robot assembly environment; the multimodal perception data includes visual modal data and force modal data; Extract the temporal variation characteristics of force sensory modal data, and identify the current assembly stage based on the temporal variation characteristics; Feature extraction and preliminary perception processing are performed on the multimodal perception data to obtain the initial perception results of the visual modality and the initial perception results of the force modality. The initial perception results of the visual modality and the initial perception results of the force modality are mapped to the same semantic feature space, and the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space is calculated. A cross-modal semantic consistency score is generated based on the feature distance; the cross-modal semantic consistency score is negatively correlated with the feature distance. Based on the cross-modal semantic consistency score and the current assembly stage, a fusion mode is selected from the predefined fusion modes to fuse the initial perception results of the visual modality and the initial perception results of the force modality, thereby generating a fused perception result; The predefined fusion modes include: collaborative fusion mode, hybrid fusion mode, and single-modal dominant mode; Based on cross-modal semantic consistency scores and fusion perception results, robot assembly control instructions are generated.
2. The multimodal perception fusion method for robots in precision assembly according to claim 1, characterized in that, Mapping the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, and calculating the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space, includes the following: The pairing of the initial visual modal perception results and the initial force modal perception results acquired under the assembly alignment state is used as a positive sample pair, and the pairing of the initial visual modal perception results and the initial force modal perception results acquired under the assembly deviation state is used as a negative sample pair. The mapping function is determined with the optimization objective of minimizing the feature distance between positive sample pairs in the semantic feature space and maximizing the feature distance between negative sample pairs in the semantic feature space. A mapping function is used to map the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, resulting in visual projection vectors and force projection vectors; The feature distance is obtained by calculating the vector distance between the visual projection vector and the force projection vector.
3. The multimodal perception fusion method for robots oriented towards precision assembly according to claim 2, characterized in that, The formula for calculating the feature distance is: ; in, For visual projection vectors, For force projection vector, For feature distance, i For the dimension of the semantic feature space, i =1,2,…, d , d It is a positive integer. For the visual projection vector at the th i Dimensional components, For the force projection vector at the th i Components in a dimension.
4. The multimodal perception fusion method for robots in precision assembly according to claim 3, characterized in that, The expression for the cross-modal semantic consistency score is: ; in, s For cross-modal semantic consistency scoring, For visual projection vectors, For force projection vector, For feature distance, This is the preset maximum distance threshold.
5. The multimodal perception fusion method for robots oriented towards precision assembly according to claim 3, characterized in that, Based on the cross-modal semantic consistency score and the current assembly stage, a fusion mode is selected from predefined fusion modes to fuse the initial perception results of the visual modality and the initial perception results of the force modality. The specific content of the generated fused perception result includes: The cross-modal semantic consistency score is compared with a first preset threshold and a second preset threshold to obtain the comparison result; the second preset threshold is less than or equal to the first preset threshold. Based on the comparison results and the current assembly stage, select a fusion mode from the predefined fusion modes: Based on the selected fusion mode, the initial perception results of the visual modality and the initial perception results of the force modality are fused to obtain the fused perception result; the fused perception result includes: pose information and force information.
6. The multimodal perception fusion method for robots in precision assembly according to claim 5, characterized in that, Based on the comparison results and the current assembly stage, the specific details of selecting a fusion mode from the predefined fusion modes include: When the cross-modal semantic consistency score is higher than the first preset threshold, the collaborative fusion mode is selected; When the cross-modal semantic consistency score is lower than the second preset threshold, the single-modal dominant mode is selected; When the cross-modal semantic consistency score is between the second preset threshold and the first preset threshold, the hybrid fusion mode is selected.
7. The multimodal perception fusion method for robots in precision assembly according to claim 6, characterized in that, The collaborative fusion mode includes: calculating the cross-modal attention weights of the initial visual modality perception results and the initial force modality perception results through an attention mechanism, and performing weighted fusion of the initial visual modality perception results and the initial force modality perception results based on the cross-modal attention weights; The single-modal dominant mode includes: identifying modalities with low cross-modal semantic consistency scores as untrusted modalities, isolating untrusted modalities from the fusion path, and using only the perception results of trusted modalities as the fusion perception results; The hybrid fusion mode includes: maintaining parallel processing of the initial perception results of the visual modality and the initial perception results of the force modality, using a weighted average method for fusion, wherein the weight coefficients in the weighted average method are positively correlated with the cross-modal semantic consistency score, and an uncertainty label is added to the fused perception results.
8. The multimodal perception fusion method for robots in precision assembly according to claim 6, characterized in that, Based on the cross-modal semantic consistency score and fusion perception results, the specific content of the generated robot assembly control instructions includes: Obtain the trend of cross-modal semantic consistency scores over time; State identification based on changing trends: When the cross-modal semantic consistency score is detected to drop sharply from above the first preset threshold to below the second preset threshold, the assembly operation is determined to be in an abnormal state; when the cross-modal semantic consistency score is detected to remain stable within the range above the first preset threshold, the assembly operation is determined to be in a normal state; thus, the state identification result is obtained. Based on the state recognition results, combined with pose information and force information, corresponding robot assembly control instructions are generated: if the assembly operation is determined to be in an abnormal state, control instructions containing an abnormality handling subroutine are generated; if the assembly operation is determined to be in a normal state, force-position hybrid control instructions are generated based on pose information and force information; the abnormality handling subroutine includes pausing the action, returning to a safe position, and triggering active perception correction.
9. The multimodal perception fusion method for robots in precision assembly according to claim 1, characterized in that, The robot multimodal perception fusion method for precision assembly further includes: verifying the fused perception results and performing active perception correction; The fused perception results are matched and verified with a preset assembly knowledge base; the assembly knowledge base contains prior information on the expected perception features under standard assembly conditions. When the deviation between the fused perception result and the prior information exceeds a preset deviation threshold, the current fused perception result is determined to be unreliable. In response to the determination that the current fusion perception result is unreliable, an active perception behavior is triggered, which controls the robot's end effector to perform a trial action to obtain updated force modal data and controls the vision sensor to change its pose to obtain updated visual modal data from a new perspective. Based on the updated force modality data and the updated visual modality data, the corrected fusion perception results are regenerated.
10. A multimodal perception fusion system for robots designed for precision assembly, characterized in that, The method for implementing the robot multimodal perception fusion for precision assembly as described in any one of claims 1-9 includes: The data acquisition module is used to collect multimodal perception data in the robot assembly environment; the multimodal perception data includes visual modal data and force modal data. The feature recognition module is used to extract the temporal variation features of force sensory modal data and identify the current assembly stage based on the temporal variation features; The data processing module is used to perform feature extraction and preliminary perception processing on multimodal perception data to obtain initial perception results for visual modality and initial perception results for force modality. The mapping calculation module is used to map the initial perception results of the visual modality and the initial perception results of the force modality to the same semantic feature space, and to calculate the feature distance between the initial perception results of the visual modality and the initial perception results of the force modality in the same semantic feature space. The scoring generation module is used to generate a cross-modal semantic consistency score based on feature distance; the cross-modal semantic consistency score is negatively correlated with feature distance. The perception fusion module is used to select a fusion mode from predefined fusion modes based on the cross-modal semantic consistency score and the current assembly stage, and fuse the initial perception results of the visual modality and the initial perception results of the force modality to generate a fused perception result; the predefined fusion modes include: collaborative fusion mode, hybrid fusion mode and single-modal dominant mode; The instruction generation module is used to generate robot assembly control instructions based on cross-modal semantic consistency scores and fusion perception results.