A 3c-oriented complex operation skill knowledge base construction method

CN118966328BActive Publication Date: 2026-08-21TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410916334.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-08-21
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

机器人无法把语言和视觉观察联系起来,并实现细粒的理解学习

Benefits of technology

[0041] The method for constructing a complex operation skill knowledge base for 3C proposed in this invention realizes the hierarchical decoupling of different types of knowledge. It comprehensively, effectively and accurately represents robot assembly operation skill knowledge from eight levels: scene layer, intelligent agent layer, entity layer, task layer, skill layer, action layer, cognition layer and perception layer, based on two dimensions: static knowledge and dynamic knowledge. This method is helpful for assembly knowledge reasoning and robot skill learning and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118966328B_ABST
    Figure CN118966328B_ABST
Patent Text Reader

Abstract

The application provides a 3C-oriented complex operation skill knowledge base construction method, comprising the following steps: constructing an intelligent assembly line; constructing a multi-modal virtual environment based on the intelligent assembly line; in the multi-modal virtual environment, a perception model containing vision, touch and depth is established, and multi-modal teaching data is collected; the multi-modal teaching data is analyzed in multiple levels and stages; and a robot assembly skill knowledge base is constructed according to the analysis result. Through the method, comprehensive and accurate expression of the robot assembly skill is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing. Background Technology

[0002] With the rapid development of information technology, people's demand for computer, communication, and consumer electronics (3C products) is constantly increasing. In the existing 3C assembly, the intelligent assembly of rigid parts has been basically realized, but the assembly of flexible deformable parts, especially variable linear objects (DLOs), still relies on manual assembly and cannot achieve switching and migration between different production lines.

[0003] In the 3C assembly industry, it is still necessary to build intelligent flexible assembly production lines to solve problems such as material heterogeneity and easy deformation, and complex assembly processes in flexible assembly, improve assembly accuracy and efficiency, meet the needs of intelligent production, break through the limitations of traditional manual assembly, and adapt to the rapid development of intelligent assembly.

[0004] Skill learning and development are crucial for robots to achieve precise control and fine manipulation, and knowledge representation is a core element of skill learning. Currently, knowledge representation for assembly scenarios lacks precision and effectiveness, and robots still face significant challenges in intelligent assembly tasks, including task understanding, motion control, memory, and the acquisition of new skills and tasks.

[0005] First, a wealth of technical knowledge (geometry, physics, function, technology, data, etc.) is required to support the modeling, planning, simulation, control, and optimization stages of intelligent assembly. Existing knowledge representation systems are mainly based on static knowledge bases, which lack sufficient detail and flexibility in describing the objects and behaviors manipulated by robots. Furthermore, the system's structural design is flawed, leading to low query efficiency. In short, there is currently no professionally available operational skills knowledge base, and data collection and analysis are difficult.

[0006] Directly implementing intelligent assembly strategies is often impractical due to data scarcity, environmental complexity, and security risks. Therefore, a learning and training environment and a training and testing platform for intelligent algorithms based on this digital twin environment are needed. This framework platform must possess physical properties consistent with real-world scenarios to transfer trained deep reinforcement learning strategies to the physical robot. However, current methods for transferring operational skills from simulation environments to real assembly environments face challenges such as weak task comprehension, poor environmental adaptability, and low execution efficiency. It is necessary to improve the robot's ability to understand tasks and learn skills, enhance its adaptability to various complex environments, improve execution efficiency, reduce learning costs, and improve its generalization ability.

[0007] Secondly, unlike humans who can naturally learn from videos and natural language to improve their language skills and complete tasks, robots cannot connect language with visual observation and achieve fine-grained understanding and learning. Precise and effective knowledge representation methods not only help improve the robot's operational accuracy in complex assembly tasks and its ability to remember previous experiences, but also facilitate automated programming of the underlying hardware.

[0008] Therefore, how to break down different knowledge representation layers to facilitate robot understanding of tasks and skill learning, and how to enable robots to understand language instructions and make correct learning and imitations, are still considerable challenges for robots. Summary of the Invention

[0009] The present invention aims to at least partially solve one of the technical problems in the related art.

[0010] Therefore, the purpose of this invention is to propose a method for constructing a complex operation skill knowledge base for 3C, which is used to construct a robot assembly skill knowledge base.

[0011] To achieve the above objectives, a first aspect of the present invention proposes a method for constructing a knowledge base for complex operation skills oriented towards 3C (computers, communications, and consumer electronics), comprising:

[0012] Build intelligent assembly production lines;

[0013] Based on the intelligent assembly line, a multimodal virtual environment is constructed; in the multimodal virtual environment, a perception model including vision, touch, and depth is established to realize the collection of multimodal teaching data;

[0014] The multimodal teaching data is subjected to multi-level and multi-stage parsing.

[0015] A robot assembly skills knowledge base is constructed based on the results of the analysis.

[0016] In addition, the method for constructing a complex operation skills knowledge base for 3C products according to the above embodiments of the present invention may also have the following additional technical features:

[0017] Furthermore, in one embodiment of the present invention, the intelligent assembly line includes:

[0018] Intelligent assembly units for cyclic production lines, including: a cyclic guide rail control unit, a 3C flexible mobile phone carrier platform loading and unloading assembly unit, a 3C flexible mobile phone flexible flat cable intelligent assembly unit, a 3C flexible mobile phone front camera intelligent assembly unit, a 3C flexible mobile phone SIM card slot intelligent assembly unit, and a 3C flexible mobile phone coaxial cable intelligent assembly unit.

[0019] Furthermore, in one embodiment of the present invention, the parsing of the multimodal teaching data includes:

[0020] The multimodal teaching data is divided into eight layers: region layer, agent layer, entity layer, task layer, skill layer, action layer, cognition layer, and perception layer, to obtain the multi-layer knowledge representation structure of the multimodal teaching data.

[0021] The system is divided into several layers: the region layer stores environmental and area knowledge during robot operation; the entity layer stores the objects operated on during robot operation and the relationships between different objects; the agent layer stores the robot, sensors, end effectors, and related configurations during robot operation; the task layer stores the goals and tasks during robot operation; the skill layer stores the motion planning during robot operation, showing how sequential skill combinations can complete a task; the action layer stores the smallest action unit that the robot can perform during robot operation; the cognition layer stores the spatial, temporal, and attribute relationships between teaching data at various levels during robot operation; and the perception layer stores the state and attribute information of the robot, environment, and operated entities perceived by the robot during robot operation in real time.

[0022] Based on the multi-level knowledge representation structure, the multi-level knowledge representation of robot operation skills is optimized and divided into three levels: ontology, parent template, and child template.

[0023] The ontology level is a set of semantic knowledge of the static environment layer and the dynamic operation layer, which does not reflect the association and temporal relationship between hierarchical knowledge. The parent template level is oriented towards a certain category of operation tasks, and adds necessary logical relationships on the basis of relevant semantic knowledge to represent an abstract action sequence. The sub-template is instantiated on the basis of the parent template, oriented towards specific operation tasks, and represents a specific action sequence. Each action sequence has specific execution parameters.

[0024] Furthermore, in one embodiment of the present invention, the method further includes parameterizing the action layer and encoding the task layer:

[0025] The data of the action layer is classified according to contact type, force direction and motion dimension, and corresponding parameters are assigned to the action according to the type. Among them, position and displacement attributes are represented by three-dimensional coordinate parameters [x,y,z], rotation angle attributes are represented by four-dimensional coordinate parameters [x,y,z,w], force, time and velocity attributes are represented by floating-point data, and direction attributes are represented by vector.

[0026] The data of the task layer is represented by a 2-bit binary code to represent task nodes, and a 6-bit binary number to represent specific actions. In the encoding of the node, 00 indicates the start of the task, 01 indicates the next action, and 11 indicates the end of the task. In the encoding of the action, the first two bits are used to identify the action type, and the last four bits are used to encode the trajectory type and spatial position of the motion, corresponding to the x-axis, y-axis, z-axis, and rotational motion in the spatial coordinate system.

[0027] The data of the cognitive layer is represented by 8-bit binary encoding to represent knowledge nodes at each level, and 16-bit binary numbers to represent specific node relationships. The first 8 bits of the node encoding are used to identify the node type, and the last 8 bits are used to encode the relationship between nodes, corresponding to spatial relationships, temporal relationships, and attribute relationships.

[0028] Furthermore, in one embodiment of the present invention, it further includes:

[0029] Based on a multi-level knowledge representation structure, the part assembly logic and task timing logic are analyzed. The task timing logic includes timing logic, synchronization logic, or combination logic between different tasks. The timing logic specifically refers to the sequential relationship between adjacent tasks in terms of operation time. The synchronization logic refers to two or more tasks being performed simultaneously. The combination logic refers to a task being completed by combining several other tasks.

[0030] Furthermore, in one embodiment of the present invention, it further includes:

[0031] Based on a multi-level knowledge representation structure, the collaborative relationship between multiple agents is analyzed. The collaborative relationship between multiple agents includes the temporal relationship, the synchronization relationship, and the master-slave relationship between each node. The temporal relationship indicates that there is a sequential execution relationship between actions. The synchronization relationship indicates that multiple agents can start certain actions at the same time. The master-slave relationship indicates that multiple agents cooperate to complete the assembly of the same part.

[0032] Furthermore, in one embodiment of the present invention, it further includes:

[0033] Based on a multi-level knowledge representation structure, the operation state switching rules are analyzed. These rules include six states, with different actions resulting in corresponding state transitions. State 1 is the "standby state," such as when the robot is powered on or waiting for subsequent instructions. State 2 determines whether an operable entity exists; if so, its name, location, and orientation are obtained; otherwise, it returns to State 1. State 3 determines whether the operation's prerequisites are met, such as whether the end effector or sensor is in the correct position. If the prerequisites are met, it enters State 4; otherwise, the action is re-executed. State 4 determines whether the operation has reached a stage state, such as whether parts are correctly assembled, correctly positioned, or whether force conditions are met. If the stage goal is reached, it enters State 5; otherwise, the operation is re-executed until the goal is achieved. State 5 determines whether state recovery should be initiated, such as returning to the initial execution state if the operation has not achieved the specified goal. State 6 determines whether the task goal state has been reached; if so, it returns to the standby state. By judging the states achievable by actions, the robot can achieve error correction and autonomous target state detection functions during operation.

[0034] To achieve the above objectives, a second aspect of the present invention provides an apparatus for constructing a complex operation skills knowledge base for 3C products, comprising the following modules:

[0035] The first building module is used to build intelligent assembly production lines;

[0036] The conversion module is used to construct a multimodal virtual environment based on the intelligent assembly production line; in the multimodal virtual environment, a perception model including vision, touch, and depth is established to realize the collection of multimodal teaching data;

[0037] The parsing module is used to perform multi-level and multi-stage parsing on the multimodal teaching data;

[0038] The second construction module is used to construct a robot assembly skill knowledge base based on the results of the analysis.

[0039] To achieve the above objectives, a third aspect of the present invention provides a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a method for constructing a complex operation skills knowledge base for 3C as described above.

[0040] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements a method for constructing a complex operation skills knowledge base for 3C as described above.

[0041] The method for constructing a complex operation skill knowledge base for 3C proposed in this invention realizes the hierarchical decoupling of different types of knowledge. It comprehensively, effectively and accurately represents robot assembly operation skill knowledge from eight levels: scene layer, intelligent agent layer, entity layer, task layer, skill layer, action layer, cognition layer and perception layer, based on two dimensions: static knowledge and dynamic knowledge. This method is helpful for assembly knowledge reasoning and robot skill learning and development. Attached Figure Description

[0042] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0043] Figure 1 This is a flowchart illustrating a method for constructing a complex operation skills knowledge base for 3C products, as provided in an embodiment of the present invention.

[0044] Figure 2 This is a diagram of a multi-level, multi-layered knowledge representation system for constructing a skill base for complex 3C operations, provided in an embodiment of the present invention.

[0045] Figure 3 This is a schematic diagram illustrating static and dynamic knowledge representation, action parameterization, and task encoding for the construction of a 3C complex operation skill base, provided by an embodiment of the present invention.

[0046] Figure 4 This is a schematic diagram of a long-sequence task knowledge representation structure for constructing a skill base for complex 3C operations, provided by an embodiment of the present invention.

[0047] Figure 5 This is a schematic diagram of a long-sequence multi-agent collaborative method for constructing a skill base for complex 3C operations, provided in an embodiment of the present invention.

[0048] Figure 6 This is a schematic diagram of a long sequence of mobile phone assembly tasks for building a skill library for complex 3C operations, provided by an embodiment of the present invention.

[0049] Figure 7 This is a schematic diagram of an operation state switching rule for constructing a complex 3C operation skill library, provided by an embodiment of the present invention.

[0050] Figure 8 This is a schematic diagram of a complex operation skills knowledge base construction device for 3C provided in an embodiment of the present invention. Detailed Implementation

[0051] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0052] The following describes a method for constructing a complex operation skills knowledge base for 3C applications, based on an embodiment of the present invention, with reference to the accompanying drawings.

[0053] Figure 1 This is a flowchart illustrating a method for constructing a complex operation skills knowledge base for 3C products, as provided in an embodiment of the present invention.

[0054] like Figure 1 As shown, this method for constructing a complex operation skills knowledge base for 3C includes the following steps:

[0055] S101: Building an intelligent assembly production line;

[0056] Furthermore, in one embodiment of the present invention, the intelligent assembly line includes:

[0057] Intelligent assembly units for cyclic production lines, including: a cyclic guide rail control unit, a 3C flexible mobile phone carrier platform loading and unloading assembly unit, a 3C flexible mobile phone flexible flat cable intelligent assembly unit, a 3C flexible mobile phone front camera intelligent assembly unit, a 3C flexible mobile phone SIM card slot intelligent assembly unit, and a 3C flexible mobile phone coaxial cable intelligent assembly unit.

[0058] S102: Based on the intelligent assembly line, construct a multimodal virtual environment; in the multimodal virtual environment, establish a perception model including vision, touch, and depth to realize the collection of multimodal teaching data;

[0059] S103: Perform multi-level and multi-stage parsing on multimodal teaching data;

[0060] like Figure 2 As shown, this is a multi-level knowledge representation system built for complex 3C operation skills. The process from the parent template to the child template is an abstract-to-concrete process, and also a process from simple to complex.

[0061] Furthermore, in one embodiment of the present invention, parsing the multimodal teaching data includes:

[0062] Multimodal teaching data is divided into eight layers: region layer, agent layer, entity layer, task layer, skill layer, action layer, cognition layer, and perception layer, resulting in a multi-layered knowledge representation structure for the multimodal teaching data. Specifically, the region layer stores environmental and regional knowledge during the robot's actual operation; the entity layer stores the objects manipulated during the robot's actual operation and the relationships between different objects; the agent layer stores the robot, sensors, end effectors, and related configurations during the robot's actual operation; the task layer stores the goals and tasks during the robot's actual operation; the skill layer stores the motion planning during the robot's actual operation, showing how sequential skill combinations can complete a task; the action layer stores the smallest action unit that the robot can execute during the actual operation; the cognition layer stores the spatial, temporal, and attribute relationships between the teaching data from each layer during the robot's actual operation; and the perception layer stores the state and attribute information of the robot, environment, and manipulated entities perceived by the robot during the actual operation.

[0063] Based on the multi-level knowledge representation structure, the multi-level knowledge representation of robot operation skills is optimized and divided into three levels: ontology, parent template, and child template.

[0064] The ontology level is a collection of semantic knowledge from the static environment layer and the dynamic operation layer. It does not reflect the association and temporal relationship between hierarchical knowledge. The parent template level is oriented towards a certain category of operation tasks. It adds necessary logical relationships on the basis of relevant semantic knowledge and represents an abstract sequence of actions. The child template is instantiated from the parent template and is oriented towards specific operation tasks. It represents a specific sequence of actions, and each sequence of actions has specific execution parameters.

[0065] like Figure 3 As shown, this is a static and dynamic knowledge representation and action parameterization and task encoding for building a skill base for complex 3C operations.

[0066] Furthermore, in one embodiment of the present invention, the method further includes parameterizing the action layer and encoding the task layer:

[0067] The data of the action layer is divided into categories according to contact type, force direction and motion dimension, and corresponding parameters are assigned to the action according to the type. Among them, the position and displacement attributes are represented by three-dimensional coordinate parameters [x,y,z], the rotation angle attribute is represented by four-dimensional coordinate parameters [x,y,z,w], the force, time and velocity attributes are all represented by floating-point data, and the direction attribute is represented by vector.

[0068] The task layer data uses 2-bit binary encoding to represent task nodes and 6-bit binary numbers to represent specific actions. In the node encoding, 00 indicates the start of the task, 01 indicates the next action, and 11 indicates the end of the task. In the action encoding, the first two bits are used to identify the action type, and the last four bits encode the trajectory type and spatial position of the motion, corresponding to the x-axis, y-axis, z-axis, and rotational motion in the spatial coordinate system.

[0069] The task encoding starts with 00 and ends with 11, and 01 indicates the next relationship of the action.

[0070] The data of the cognitive layer is represented by 8-bit binary encoding to represent knowledge nodes at each level, and 16-bit binary numbers to represent specific node relationships. The first 8 bits of the node encoding are used to identify the node type, and the last 8 bits are used to encode the relationship between nodes, corresponding to spatial relationships, temporal relationships, and attribute relationships.

[0071] like Figure 4 As shown, this is a knowledge representation structure for long-sequence tasks built for a 3C complex operation skill base. A complete mobile phone assembly task can be viewed as a long sequence of complex tasks composed of different part assembly tasks. Therefore, it is necessary to analyze the knowledge representation structure of long-sequence tasks based on the existing multi-level knowledge representation structure. The static layer mainly describes entities within the region and their corresponding physical characteristics, such as size and material. The dynamic layer mainly describes tasks, subtasks, and states. Subtasks can be understood as skills, and a sequence of subtasks can be combined into a task. The switching between adjacent states of a task during the assembly process can be understood as actions.

[0072] like Figure 5 As shown, this is a long-sequence multi-agent cooperative method for building a skill library for complex 3C operations.

[0073] Furthermore, in one embodiment of the present invention, it further includes:

[0074] Based on a multi-level knowledge representation structure, the part assembly logic and task timing logic are analyzed. The task timing logic includes timing logic, synchronous logic, or combinational logic between different tasks. Timing logic specifically refers to the sequential relationship between adjacent tasks in terms of operation time. Synchronous logic refers to two or more tasks being performed simultaneously. Combinatorial logic refers to a task being completed by combining several other tasks.

[0075] like Figure 6 As shown, this is a long-sequence multi-agent collaborative method for building a skill library for complex 3C operations. It includes three types of relationships between nodes: temporal relationships, synchronization relationships, and master-slave relationships. The synchronization point is used to align actions when different agents collaborate, in order to solve the time difference problem caused by the speed of task execution.

[0076] Furthermore, in one embodiment of the present invention, it further includes:

[0077] Based on a multi-level knowledge representation structure, we analyze the multi-agent collaborative relationship, which includes the temporal relationship, synchronization relationship and master-slave relationship between nodes. The temporal relationship indicates that there is a sequential execution relationship between actions, the synchronization relationship indicates that multiple agents can start some actions at the same time, and the master-slave relationship indicates that multiple agents cooperate to complete the assembly of the same part.

[0078] like Figure 7 The diagram illustrates the operation state switching rules for building a complex 3C operation skill library. The operation state switching rules include six states; different actions result in corresponding state transitions.

[0079] Furthermore, in one embodiment of the present invention, it further includes:

[0080] Based on a multi-level knowledge representation structure, the operation state switching rules are analyzed. These rules include six states, with different actions resulting in corresponding state transitions. State 1 is the "standby state," such as when the robot is powered on or waiting for subsequent instructions. State 2 determines if an operable entity exists; if so, its name, location, and orientation are retrieved; otherwise, the robot returns to State 1. State 3 determines if the operation's prerequisites are met, such as whether the end effector or sensor is in the correct position. If the prerequisites are met, the robot enters State 4; otherwise, the action is re-executed. State 4 determines if the operation has reached a stage state, such as whether parts are correctly assembled, correctly positioned, or if force conditions are met. If the stage goal is achieved, the robot enters State 5; otherwise, the operation is re-executed until the goal is reached. State 5 determines whether state recovery should be initiated, such as returning to the initial execution state if the operation has not achieved the specified goal. State 6 determines if the task goal state has been reached; if so, the robot returns to the standby state. By judging the states achievable by actions, the robot can achieve error correction and autonomous target state detection functions during operation.

[0081] S104: Construct a robot assembly skills knowledge base based on the analysis results.

[0082] The method for constructing a complex operation skill knowledge base for 3C proposed in this invention realizes the hierarchical decoupling of different types of knowledge. It comprehensively, effectively and accurately represents robot assembly operation skill knowledge from eight levels: scene layer, intelligent agent layer, entity layer, task layer, skill layer, action layer, cognition layer and perception layer, based on two dimensions: static knowledge and dynamic knowledge. This method is helpful for assembly knowledge reasoning and robot skill learning and development.

[0083] To achieve the above embodiments, the present invention also proposes a device for constructing a knowledge base of complex operation skills for 3C products.

[0084] Figure 8 This is a schematic diagram of a device for constructing a complex operation skills knowledge base for 3C products, provided in an embodiment of the present invention.

[0085] like Figure 8 As shown, the device for building a complex operation skills knowledge base for 3C devices includes: a first construction module 100, a conversion module 200, a parsing module 300, and a second construction module 400, wherein...

[0086] The first building module is used to build intelligent assembly production lines;

[0087] The conversion module is used to build a multimodal virtual environment based on the intelligent assembly production line; in the multimodal virtual environment, a perception model including vision, touch and depth is established to realize the collection of multimodal teaching data;

[0088] The parsing module is used to perform multi-level and multi-stage parsing of multimodal teaching data;

[0089] The second construction module is used to build a robot assembly skill knowledge base based on the parsing results.

[0090] To achieve the above objectives, a third aspect of the present invention provides a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method for constructing a complex operation skills knowledge base for 3C as described above.

[0091] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the method for constructing a complex operation skills knowledge base for 3C as described above.

[0092] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0093] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0094] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for constructing a knowledge base of complex operation skills for 3C products, characterized in that, Includes the following steps: Build intelligent assembly production lines; Based on the intelligent assembly line, a multimodal virtual environment is constructed; in the multimodal virtual environment, a perception model including vision, touch, and depth is established to realize the collection of multimodal teaching data; The multimodal teaching data is subjected to multi-level and multi-layered parsing, including: The multimodal teaching data is divided into eight layers: region layer, agent layer, entity layer, task layer, skill layer, action layer, cognition layer, and perception layer, resulting in a multi-layered knowledge representation structure for the multimodal teaching data. Specifically, the region layer stores environmental and regional knowledge during the robot's actual operation; the entity layer stores the objects operated on during the robot's actual operation and the relationships between different objects; the agent layer stores the robot, sensors, end effectors, and related configurations during the robot's actual operation; the task layer stores the goals and tasks during the robot's actual operation; the skill layer stores the motion planning during the robot's actual operation, showing how sequential skill combinations can complete a task; the action layer stores the smallest action unit that the robot can execute during the robot's actual operation; the cognition layer stores the spatial, temporal, and attribute relationships between the teaching data at each layer during the robot's actual operation; and the perception layer stores the state and attribute information of the robot, environment, and operated entities perceived by the robot during the actual operation in real time. Based on the multi-level knowledge representation structure, the multi-level knowledge representation of robot operation skills is optimized and divided into three levels: ontology, parent template, and child template. The ontology level is a set of semantic knowledge of the static environment layer and the dynamic operation layer, which does not reflect the association and temporal relationship between hierarchical knowledge. The parent template level is oriented towards a certain category of operation tasks, and adds necessary logical relationships on the basis of relevant semantic knowledge to represent an abstract action sequence. The sub-template is instantiated on the basis of the parent template, oriented towards specific operation tasks, and represents a specific action sequence. Each action sequence has specific execution parameters. A robot assembly skills knowledge base is constructed based on the results of the analysis.

2. The method according to claim 1, characterized in that, The intelligent assembly line includes: Intelligent assembly units for cyclic production lines, including: a cyclic guide rail control unit, a 3C flexible mobile phone carrier platform loading and unloading assembly unit, a 3C flexible mobile phone flexible flat cable intelligent assembly unit, a 3C flexible mobile phone front camera intelligent assembly unit, a 3C flexible mobile phone SIM card slot intelligent assembly unit, and a 3C flexible mobile phone coaxial cable intelligent assembly unit.

3. The method according to claim 1, characterized in that, This also includes parameterizing the action layer and encoding the task layer and the cognitive layer: The data of the action layer is classified according to contact type, force direction and motion dimension, and corresponding parameters are assigned to the action according to the type. Among them, position and displacement attributes are represented by three-dimensional coordinate parameters [x,y,z], rotation angle attributes are represented by four-dimensional coordinate parameters [x,y,z,w], force, time and velocity attributes are represented by floating-point data, and direction attributes are represented by vector. The data of the task layer is represented by a 2-bit binary code to represent task nodes, and a 6-bit binary number to represent specific actions. In the node code, 00 indicates the start of the task, 01 indicates the next action, and 11 indicates the end of the task. In the action code, the first two bits are used to identify the action type, and the last four bits encode the trajectory type and spatial position of the motion, corresponding to the x-axis, y-axis, z-axis, and rotational motion in the spatial coordinate system. The data of the cognitive layer is represented by 8-bit binary encoding to represent knowledge nodes at each level, and 16-bit binary numbers to represent specific node relationships. The first 8 bits of the node encoding are used to identify the node type, and the last 8 bits are used to encode the relationship between nodes, corresponding to spatial relationships, temporal relationships, and attribute relationships.

4. The method according to claim 1, characterized in that, Also includes: Based on a multi-level knowledge representation structure, the part assembly logic and task timing logic are analyzed. The task timing logic includes timing logic, synchronization logic, or combination logic between different tasks. The timing logic specifically refers to the sequential relationship between adjacent tasks in terms of operation time. The synchronization logic refers to two or more tasks being performed simultaneously. The combination logic refers to a task being completed by combining several other tasks.

5. The method according to claim 1, characterized in that, Also includes: Based on a multi-level knowledge representation structure, the collaborative relationship between multiple agents is analyzed. The collaborative relationship between multiple agents includes the temporal relationship, the synchronization relationship, and the master-slave relationship between each node. The temporal relationship indicates that there is a sequential execution relationship between actions. The synchronization relationship indicates that multiple agents can start certain actions at the same time. The master-slave relationship indicates that multiple agents cooperate to complete the assembly of the same part.

6. The method according to claim 1, characterized in that, Also includes: Based on a multi-level knowledge representation structure, the operation state switching rules are analyzed. These rules include six states, with different actions resulting in corresponding state transitions. State 1 is the "standby state," such as when the robot is powered on or waiting for subsequent instructions. State 2 determines if an operable entity exists; if so, its name, location, and orientation are obtained; otherwise, it returns to State 1. State 3 determines if the operation's prerequisites are met, such as whether the end effector or sensor is in the correct position. If the prerequisites are met, it enters State 4; otherwise, the action is re-executed. State 4 determines if the operation has reached a stage state, such as whether parts are correctly assembled, correctly positioned, or whether force conditions are met. If the stage goal is reached, it enters State 5; otherwise, the operation is re-executed until the goal is achieved. State 5 determines whether state recovery should be initiated, such as returning to the initial execution state if the operation has not achieved the specified goal. State 6 determines if the task goal state has been reached; if so, it returns to the standby state. By judging the states achievable by actions, the robot can achieve error correction and autonomous target state detection functions during operation.

7. A device for constructing a knowledge base of complex operation skills for 3C products, characterized in that, Includes the following modules: The first building module is used to build intelligent assembly production lines; The conversion module is used to construct a multimodal virtual environment based on the intelligent assembly production line; in the multimodal virtual environment, a perception model including vision, touch, and depth is established to realize the collection of multimodal teaching data; The parsing module is used to perform multi-level and multi-layered parsing of the multimodal teaching data, including: The multimodal teaching data is divided into eight layers: region layer, agent layer, entity layer, task layer, skill layer, action layer, cognition layer, and perception layer, resulting in a multi-layered knowledge representation structure for the multimodal teaching data. Specifically, the region layer stores environmental and regional knowledge during the robot's actual operation; the entity layer stores the objects operated on during the robot's actual operation and the relationships between different objects; the agent layer stores the robot, sensors, end effectors, and related configurations during the robot's actual operation; the task layer stores the goals and tasks during the robot's actual operation; the skill layer stores the motion planning during the robot's actual operation, showing how sequential skill combinations can complete a task; the action layer stores the smallest action unit that the robot can execute during the robot's actual operation; the cognition layer stores the spatial, temporal, and attribute relationships between the teaching data at each layer during the robot's actual operation; and the perception layer stores the state and attribute information of the robot, environment, and operated entities perceived by the robot during the actual operation in real time. Based on the multi-level knowledge representation structure, the multi-level knowledge representation of robot operation skills is optimized and divided into three levels: ontology, parent template, and child template. The ontology level is a set of semantic knowledge of the static environment layer and the dynamic operation layer, which does not reflect the association and temporal relationship between hierarchical knowledge. The parent template level is oriented towards a certain category of operation tasks, and adds necessary logical relationships on the basis of relevant semantic knowledge to represent an abstract action sequence. The sub-template is instantiated on the basis of the parent template, oriented towards specific operation tasks, and represents a specific action sequence. Each action sequence has specific execution parameters. The second construction module is used to construct a robot assembly skill knowledge base based on the results of the analysis.

8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for constructing a complex operation skills knowledge base for 3C as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for constructing a complex operation skills knowledge base for 3C as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Construction method and system of agent multi-mode cognitive map and medium

    CN114925176A

  • Method for extracting operation skill information from human demonstration and constructing knowledge base

    CN118070890A