Digital twin system based on multiple modes and simulation method
By designing a multimodal identification digital twin system with integrated devices and digital twins, the existing system has solved the problems of high complexity, limited data fusion analysis capabilities and insufficient intelligent control capabilities, and efficient data processing and decision-making analysis are achieved, enhancing the performance and availability of the system.
Patent Information
- Application Number
- CN202411784849.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-02
AI Technical Summary
The existing multimodal recognition system has high complexity, limited data fusion analysis capabilities, and insufficient intelligent control and learning capabilities, making it impossible to achieve high-precision scenario deduction and decision-making.
A digital twin system based on multimodal recognition is designed to achieve efficient data acquisition, processing, transmission, analysis and feedback control through integrated devices, digital twins and information transmission media. The system includes a multimodal signal acquisition device, an intelligent controller, a digital twin and an information transmission medium, which can adaptively improve accuracy and efficiency.
The system improves the performance and availability of multimodal recognition system, realizes efficient data processing and decision analysis, enhances intelligent control and learning capabilities, and can adaptively improve accuracy and efficiency.
Smart Images

Figure CN119917951A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal digital twin system and simulation method, especially to the fields of smart agriculture, smart manufacturing, virtual humans, humanoid robots, etc. controlled by digital twins. Background Art
[0002] At present, fields such as smart agriculture, intelligent manufacturing, virtual humans and humanoid robots are developing rapidly, among which multimodal recognition technology plays an important role in these fields. Multimodal recognition refers to the technology of identifying and understanding objects or scenes based on multiple perceptual modalities (such as video, audio, image, text, etc.). By integrating information from multiple perceptual modalities, a more comprehensive and accurate analysis and decision-making basis can be provided. However, there are some problems with the existing multimodal recognition system. First, the data acquisition and processing process usually requires multiple independent devices and algorithms, resulting in increased system complexity and cost. Secondly, the existing system has certain limitations in data fusion and analysis, and cannot achieve high-precision scene deduction and decision-making. In addition, the intelligent control and learning capabilities of the existing system are limited, and it is impossible to adaptively improve accuracy and efficiency. Therefore, a new type of digital twin system based on multimodal recognition is needed to solve the problems of the existing system and improve the performance and availability of the system. . Summary of the invention
[0003] The present invention proposes a digital twin system based on multimodal recognition, which aims to solve the problems of high complexity, limited data fusion and analysis capabilities, and insufficient intelligent control and learning capabilities of multimodal recognition systems in the fields of smart agriculture, intelligent manufacturing, virtual humans, humanoid robots, etc. The system realizes efficient data collection, processing, transmission, analysis, and feedback control by integrating equipment, digital twins, and information transmission media, thus improving the performance and availability of the system.
[0004] Specifically, the system includes equipment, digital twins, and information transmission media. The equipment part includes an execution body 1, an intelligent controller 2, and a multimodal signal acquisition device 3. The multimodal signal acquisition device 3 can simultaneously collect information in multiple modes such as video, audio, pictures, text, and signals, clean and compress the data by calling local small models and knowledge graphs, and then send it to the digital twin 5 through a wired or wireless information transmission medium 4. The digital twin 5 includes a simulator 6 and a multimodal large model 7, which can receive information from the device, perform data fusion analysis, and perform scenario deduction. According to the results of the scenario deduction, the multimodal large model 7 can drive the simulator 6 to perform simulation operations, and send feedback information to the device through the information transmission medium 4 to guide the execution body 1 of the device to perform precise operations.
[0005] In addition, the digital twin system of the present invention also has intelligent learning and adaptability. When the intelligent controller 2, the multimodal signal acquisition device 3 or the multimodal large model 7 fails or is abnormal, the execution subject 1 can be taught by the manual control device, thereby updating the professional knowledge base of the intelligent controller 2, the multimodal signal acquisition device 3 and the multimodal large model 7, so that the system has self-learning ability and gradually improves accuracy and efficiency. At the same time, the intelligent controller 2 can also feed back the environmental information and status information around the execution subject 1 to the digital twin 5 in real time, so that the user can monitor it in real time through the simulator 6 to ensure the safety and reliability of the system.
[0006] In terms of data processing and analysis, the present invention uses professional knowledge graphs to filter the uncertainties generated by the generative models built into the intelligent controller 2, the multimodal signal acquisition device 3, and the multimodal large model 7, ensuring the high precision requirements of the execution of the main work. Through the collaborative work of the multimodal large model 7, the multimodal signal acquisition device 3, and the intelligent controller 2, accurate and efficient data processing and decision analysis are achieved under the constraints of the knowledge graph.
[0007] To sum up, the digital twin system based on multimodal recognition proposed in the present invention has efficient data acquisition, processing, transmission, analysis and feedback control capabilities, can adaptively improve accuracy and efficiency, and provides a new solution for smart agriculture, intelligent manufacturing, virtual humans, humanoid robots and other fields.
[0008] Attached Figure 1 It is a system structure diagram of a multimodal digital twin modeling and simulation method.
[0009] 1-Executing entity, 2-intelligent controller, 3-multimodal signal acquisition device, 4-information transmission medium, 5-digital twin, 6-simulator, 7-multimodal large model. DETAILED DESCRIPTION
[0010] Embodiment 1: Combining with the attached Figure 1 The present invention is described in detail by taking an intelligent agricultural robot as an example.
[0011] A multimodal intelligent agricultural robot digital twin system includes a robot device, a digital twin, and an information transmission medium, wherein the robot device includes: an execution body 1, an intelligent controller 2, and a multimodal signal acquisition device 3; the digital twin 5 includes: a simulator 6 and a multimodal large model 7; the information transmission medium 4 can be a 5G-A or remote operation information transmission method; in the robot device, the multimodal signal acquisition device 3 collects the posture, movement, shape, color of the fruit, temperature, humidity and other signal information of the agricultural robot through visual recognition, calls the local small model and knowledge graph for data cleaning and data compression, and then sends it to the digital twin 5 through the information transmission medium 4 to drive the simulator 6. The simulator 6 processes the data according to the trained local model combined with the knowledge graph. If the problem cannot be solved, it will be connected to the multimodal large model 7 through the information transmission medium 4 for data fusion analysis, and scenario deduction is performed based on the analysis results. The multimodal large model 7 drives the simulator 6 according to the result of the scenario deduction, the simulator 6 feeds back to the multimodal signal acquisition device 3 through the information transmission medium 4, the multimodal signal acquisition device 3 feeds back to the intelligent controller 2 through the transmission medium 4, and the intelligent controller 2 drives the execution subject 1 to pick the fruit after data processing.
[0012] The multimodal signal acquisition device 3, the small model and the local knowledge base on the intelligent controller 2 are updated by the multimodal large model 7 through knowledge distillation and the cloud-based professional knowledge base, and the update can be done in an OTA (Over-The-Air) manner.
[0013] When the intelligent controller 2, the multimodal signal acquisition device 3, and the multimodal large model 7 cannot work normally, the manual control device installed on the execution body 1 can be used to issue instructions to the intelligent controller 2 to manually control the execution body 1. The teaching information of this process can update the professional knowledge base of the intelligent controller 2, the multimodal signal acquisition device 3, and the multimodal large model 7, so that the digital twin system can perform self-learning and gradually improve its accuracy.
[0014] The intelligent controller 2 can transmit information such as rain, fog, temperature, humidity, pests and diseases, and fruit maturity status of the surrounding environment of the execution subject 1 to the digital twin through the information transmission medium 4, so that the user can monitor it in real time through the simulator 6.
[0015] Embodiment 2: Combining with Figure 1 The present invention is described in detail by taking the production of virtual humans as an example.
[0016] A virtual human digital twin system based on multimodality includes a human or robot device, a digital twin, and an information transmission medium, wherein the human or robot device includes: an execution subject 1, an intelligent controller 2, and a multimodal signal acquisition device 3; the digital twin 5 includes: a simulator 6, a multimodal large model 7; the information transmission medium 4 can be a 5G-A or remote operation information transmission method; in the human or robot device, the multimodal signal acquisition device 3 collects signal information such as the head shape, facial expression, posture, movement, sound, text, etc. of the human or robot device through a matrix camera and a matrix microphone, calls a local small model and a knowledge graph for data cleaning and data compression, and then sends it to the digital twin 5 through the information transmission medium 4 to drive the simulator 6. The simulator 6 performs three-dimensional modeling and control driving based on the trained local model combined with the knowledge graph. If the problem cannot be solved, it will be connected to the multimodal large model 7 through the information transmission medium 4 for data fusion analysis, and model construction and scene deduction are performed according to the analysis results. The multimodal large model 7 drives the simulator 6 according to the result of the scenario deduction, the simulator 6 feeds back to the multimodal signal acquisition device 3 through the information transmission medium 4, the multimodal signal acquisition device 3 feeds back to the intelligent controller 2 through the transmission medium 4, and the intelligent controller 2 drives the execution subject 1 to act after data processing.
[0017] The multimodal signal acquisition device 3, the small model and the local knowledge base on the intelligent controller 2 are updated by the multimodal large model 7 through knowledge distillation and the cloud-based professional knowledge base, and the update can be done in an OTA (Over-The-Air) manner.
[0018] When the intelligent controller 2, the multimodal signal acquisition device 3, and the multimodal large model 7 cannot work normally, the manual control device installed on the execution body 1 can be used to issue instructions to the intelligent controller 2 to manually control the execution body 1. The teaching information of this process can update the professional knowledge base of the intelligent controller 2, the multimodal signal acquisition device 3, and the multimodal large model 7, so that the digital twin system can perform self-learning and gradually improve its accuracy.
[0019] The intelligent controller 2 can transmit information such as rain, fog, temperature, humidity, green screen, and device status information of the surrounding environment of the execution subject 1 to the digital twin through the information transmission medium 4, so that the user can monitor it in real time through the simulator 6.
Claims
1. The digital twin system based on multimodal recognition is characterized by: The digital twin system based on multimodal recognition includes equipment, digital twins, and information transmission media, wherein the equipment includes: an execution subject 1, an intelligent controller 2, and a multimodal signal acquisition device 3; the digital twin 5 includes: a simulator 6 and a multimodal large model 7; the information transmission medium 4 can be wired or wireless; in the equipment, the multimodal signal acquisition device 3 collects video, audio, pictures, text and signal information of the equipment, and after data processing, sends it to the digital twin 5 through the information transmission medium 4 to drive the simulator 6, and the simulator 6 connects to the multimodal large model 7 through the information transmission medium 4 to perform data fusion analysis and scene deduction. The multimodal large model 7 drives the simulator 6 according to the scene deduction result, and the simulator 6 feeds back to the multimodal signal acquisition device 3 through the information transmission medium 4, and the multimodal signal acquisition device 3 feeds back to the intelligent controller 2 through the transmission medium 4, and the intelligent controller 2 drives the execution subject 1 after data processing.
2. According to the digital twin system of claim 1, it is characterized in that The multimodal signal acquisition device 3 can collect video, audio, picture, text and signal information at the same time, call the local small model and knowledge graph to perform data cleaning and data compression, and send it to the digital twin 5 through the information transmission medium 4. The local small model and knowledge graph can be updated by the multimodal large model 7 through the information transmission medium 4 through knowledge distillation and professional knowledge base, and the update can be done in an OTA (Over-The-Air) manner.
3. According to the digital twin system of claim 1, it is characterized in that When the intelligent controller 2, the multimodal signal acquisition device 3, and the multimodal large model 7 cannot work normally, the manual control device installed on the execution body 1 can be used to issue instructions to the intelligent controller 2 to control the execution end body 1. Through this process teaching information, the professional knowledge base of the intelligent controller 2, the multimodal signal acquisition device 3, and the multimodal large model 7 can be updated, so that the digital twin system can perform self-learning and gradually improve its accuracy.
4. The digital twin system according to claim 1 is characterized in that ,The intelligent controller 2 can transmit the surrounding environment information and ,status information feedback of the executing subject 1 to the digital twin through the ,information transmission medium 4, so that the user can monitor it in real ,time through the simulator 6.
Citation Information
Cited By
Digital human configuration method and system based on industrial element universe digital twinning
CN120763197A
Digital human configuration method and system based on industrial metaverse digital twinning
CN120763197B
Textile design optimization method based on artificial intelligence
CN120893330A
A multi-modal three-dimensional modeling and virtual environment simulation system based on an industrial scene
CN122632648A