A multimodal embodied intelligent device control method, system and terminal

By deploying and training a multimodal model on a smart card driver, and combining this with voice information alignment processing, the problem of data misalignment during multimodal model training was solved, thereby improving model training efficiency and device control accuracy.

CN120012045BActive Publication Date: 2025-10-28SHENZHEN MAITEXIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411893029.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-10-28
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing technologies for multimodal models suffer from data misalignment during training, leading to inefficiency and inaccurate results.

Method used

By acquiring the programmed smart card driver, deploying a multimodal model, training it, aligning the user's input voice information, outputting the alignment result, and controlling the operation of the embodied smart device to obtain the running result.

Benefits of technology

This improves the efficiency and performance of multimodal model training, ensuring the accuracy of the trained model and the accuracy of device control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012045B_ABST
    Figure CN120012045B_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, and terminal for controlling embodied intelligent devices based on multimodality. The method includes: acquiring a programmed smart card driver; controlling the smart card driver; deploying a multimodal model and training the multimodal model to obtain a target multimodal model; acquiring user-inputted voice information and inputting the voice information into the target multimodal model, outputting prompt words; transmitting the prompt words to the embodied intelligent device; controlling the operation of the embodied intelligent device; and acquiring the operation results. In training the multimodal model, this invention preprocesses the data, transforming different modal data into data of the same scale or distribution, and aligns the data through a graph structure, thereby improving the efficiency of model training and the performance of the trained model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, system, terminal, and computer-readable storage medium for controlling embodied intelligent devices based on multimodality. Background Technology

[0002] Embodied intelligence is a field of development in artificial intelligence, referring to the ability of an intelligent system or machine to interact with its environment in real time through perception and interaction. The generalization ability of large-scale multimodal models also lays the foundation for the autonomous learning ability of robots, helping them to adapt to varied tasks.

[0003] Currently, large-scale multimodal technology is developing rapidly, combining technologies from multiple fields such as computer vision, natural language processing, and speech recognition.

[0004] However, data from different modalities are often misaligned in time, space, and semantics, and have different statistical properties and distributions. This requires the model to be able to process heterogeneous data and extract useful information from it. In the training process of multimodal models, multimodal data usually requires complex annotation, which is not only costly, but the quality of the annotation directly affects the performance of the model.

[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0006] The main objective of this invention is to provide a control method, system, terminal, and computer-readable storage medium for embodied intelligent devices based on multimodality, aiming to solve the problems of data misalignment during the training process of multimodal models in the prior art, which leads to low efficiency and inaccurate results in the training process of multimodal models.

[0007] To achieve the above objectives, the present invention provides a multimodal embodied intelligent device control method, which includes the following steps:

[0008] Obtain the programmed smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain the target multimodal model;

[0009] Acquire user-inputted voice information, input the voice information into the target multimodal model for alignment processing, and output the alignment result;

[0010] The alignment result is transmitted to the embodied intelligent device, the embodied intelligent device is controlled to operate according to the alignment result, and the operation result of the embodied intelligent device is obtained.

[0011] Optionally, the method for controlling a multimodal embodied intelligent device, wherein acquiring the programmed smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model, specifically includes:

[0012] Obtain the initial smart card inserted into the PC and select AI bitstream data, wherein the AI ​​bitstream data is used to deploy a multimodal model;

[0013] The AI ​​bitstream data is burned into the initial smart card to obtain the smart card driver;

[0014] Construct a multimodal model and control the smart card driver to deploy the multimodal model to the PC based on the AI ​​bitstream data;

[0015] Historical recognition data is extracted from the knowledge base, a training set is constructed based on the historical recognition data, and the multimodal model is trained using the training set to obtain the target multimodal model.

[0016] Optionally, the method for controlling embodied intelligent devices based on multimodal characteristics, wherein acquiring user-inputted voice information, inputting the voice information into the target multimodal model for alignment processing, and outputting the alignment result specifically includes:

[0017] The system acquires user-input voice information, analyzes the voice information using a prompt word writing framework, and obtains the target voice information.

[0018] The target speech information is input into the target multimodal model, which performs alignment processing on the target speech information and outputs the alignment result.

[0019] Optionally, the method for controlling embodied intelligent devices based on multimodality, wherein the step of inputting the target speech information into the target multimodal model, and the target multimodal model performing alignment processing on the target speech information and outputting the alignment result, specifically includes:

[0020] The target speech information is input into the target multimodal model, which extracts the visual and linguistic features of the target speech information and generates positive and negative sample pairs.

[0021] The target multimodal model aligns the visual features and the linguistic features based on the image-text relationship in the target speech information to obtain an initial alignment result;

[0022] The target multimodal model minimizes the distance between the positive sample pairs to obtain the minimum positive sample pair distance, and maximizes the distance between the negative sample pairs to obtain the maximum negative sample pair distance;

[0023] The target multimodal model optimizes the initial alignment result based on the minimum positive sample pair distance and the maximum negative sample pair distance, and outputs the alignment result of the target speech information.

[0024] Optionally, the control method for embodied intelligent devices based on multimodality includes any one of a humanoid robot, a robotic arm, a robot, and a robotic dog.

[0025] The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device specifically includes:

[0026] Obtain the address of the humanoid robot, the robotic arm, the robot, or the robotic dog, and transmit the alignment result to the humanoid robot, the robotic arm, the robot, or the robotic dog according to the address;

[0027] Based on the alignment result, control the humanoid robot, the robotic arm, the robot, or the robotic dog to perform tasks and receive the task results.

[0028] Optionally, the multimodal-based embodied intelligent device control method further includes, before transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device:

[0029] Extract the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot from the historical recognition data;

[0030] The results of the first historical task are input into the target multimodal model for training to obtain a loop generalized linear model, and the smart card driver is controlled to deploy the text-to-speech conversion tool into the humanoid robot.

[0031] The results of the second historical task are input into the target multimodal model for training to obtain an environmental object recognition model;

[0032] The results of the third historical task are input into the target multimodal model for training to obtain the edge language model.

[0033] Optionally, the multimodal-based embodied intelligent device control method, wherein controlling the humanoid robot, the robotic arm, the robot, or the robotic dog to perform tasks based on the alignment result, and receiving the task results, specifically includes:

[0034] If the control target is the humanoid robot, the alignment result is input into the loop generalized linear model, and a first control instruction is generated. According to the first control instruction, the text-to-speech conversion tool is controlled to convert the speech information into text information and output a semantic response until the dialogue with the user is completed. The semantic response is generated by the humanoid robot based on the text information.

[0035] If the control target is the robotic arm, the alignment result is input into the environmental object recognition model, and a second control command is generated. The robotic arm is controlled to perform a second target task according to the second control command, and the execution result of the second target task is obtained.

[0036] If the control target is the robot, the alignment result is input into the edge language model, and a third control instruction is generated. The robot is then controlled to perform a third target task according to the third control instruction, and the execution result of the third target task is obtained.

[0037] If the control target is the robotic dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. The robotic dog is then controlled to perform a fourth target task according to the fourth control instruction, and the execution result of the fourth target task is obtained.

[0038] Furthermore, to achieve the above objectives, the present invention also provides a multimodal embodied intelligent device control system, wherein the multimodal embodied intelligent device control system includes:

[0039] The model training module is used to acquire the programmed smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain a target multimodal model.

[0040] The information processing module is used to acquire the voice information input by the user, input the voice information into the target multimodal model for alignment processing, and output the alignment result;

[0041] The device control module is used to transmit the alignment result to the embodied intelligent device, control the embodied intelligent device to operate according to the alignment result, and obtain the operating result of the embodied intelligent device.

[0042] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a multimodal embodied intelligent device control program stored in the memory and executable on the processor, wherein when the multimodal embodied intelligent device control program is executed by the processor, it implements the steps of the multimodal embodied intelligent device control method described above.

[0043] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal embodied intelligent device control program, which, when executed by a processor, implements the steps of the multimodal embodied intelligent device control method described above.

[0044] In this invention, a programmed smart card driver is acquired, and the smart card driver is used to deploy a multimodal model and train the multimodal model to obtain a target multimodal model. User-inputted voice information is acquired and input into the target multimodal model for alignment processing, and the alignment result is output. The alignment result is transmitted to an embodied intelligent device, which is then controlled to operate according to the alignment result, and the operating results of the embodied intelligent device are obtained. This invention improves the efficiency of model training and the performance of the trained model by preprocessing the data during multimodal model training, converting different modal data into data of the same scale or distribution, and aligning the data through a graph structure. Attached Figure Description

[0045] Figure 1 This is a flowchart of a preferred embodiment of the multimodal embodied intelligent device control method of the present invention;

[0046] Figure 2 This is a schematic diagram of AI model-driven bitstream programming, which is a preferred embodiment of the multimodal embodied intelligent device control method of the present invention.

[0047] Figure 3 This is a flowchart of AI model-driven bitstream programming, which is a preferred embodiment of the multimodal embodied intelligent device control method of the present invention.

[0048] Figure 4 This is an AI model interaction diagram of a preferred embodiment of the embodied intelligent device control method based on multimodality of the present invention;

[0049] Figure 5 This is a flowchart of a localized smart card driver device, which is a preferred embodiment of the multimodal embodied smart device control method of the present invention.

[0050] Figure 6This is a flowchart of a preferred embodiment of the custom prompt word voice of the embodied intelligent device control method based on multimodality of the present invention;

[0051] Figure 7 This is a field diagram of the custom prompt words in a preferred embodiment of the embodied intelligent device control method based on multimodality of the present invention;

[0052] Figure 8 This is a structural diagram of a preferred embodiment of the multimodal embodied intelligent device control system of the present invention;

[0053] Figure 9 This is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0055] The preferred embodiment of the present invention describes a multimodal-based embodied intelligent device control method, such as... Figure 1 As shown, the multimodal-based embodied intelligent device control method includes the following steps:

[0056] Step S10: Obtain the programmed smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain the target multimodal model.

[0057] Among them, such as Figure 2 and Figure 3 As shown, insert the smart card into the PC and use the programming software to program the preset bitstream data into the smart card to obtain the smart card driver.

[0058] Specifically, the process involves obtaining an initial smart card inserted into a PC and selecting AI bitstream data, wherein the AI ​​bitstream data is used to deploy a multimodal model; burning the AI ​​bitstream data into the initial smart card to obtain the smart card driver; constructing the multimodal model and controlling the smart card driver to deploy the multimodal model to the PC based on the AI ​​bitstream data; extracting historical recognition data from a knowledge base, constructing a training set based on the historical recognition data, and using the training set to train the multimodal model to obtain a target multimodal model.

[0059] Among them, such as Figure 4As shown, the AI ​​model is trained using historical recognition data from the knowledge base, thereby enabling the construction of the smart card driver (i.e., burning bitstream data onto the smart card). By enriching the knowledge base of the large model, the efficiency of training the intelligent capabilities of the large model is improved. At the same time, the ability to recognize language semantics is enhanced, enabling precise voice control.

[0060] Step S20: Obtain the voice information input by the user, input the voice information into the target multimodal model for alignment processing, and output the alignment result.

[0061] The alignment process here can be implemented using the ALBEF (Align before Fuse, a multimodal representation learning method). This method can efficiently learn the joint representation of vision and language through image-text alignment and momentum distillation techniques, thereby improving the accuracy of speech information recognition and ultimately enhancing the accuracy of subsequent control of smart devices.

[0062] Specifically, the system acquires user-input voice information, analyzes the voice information using a prompt word writing framework to obtain target voice information, inputs the target voice information into the target multimodal model, performs alignment processing on the target voice information, and outputs the alignment result.

[0063] Among them, such as Figure 5 As shown, by analyzing the user's input voice information (i.e., preprocessing to reduce invalid information and improve processing efficiency), target voice information (which can be converted into text) is obtained. Then, a multimodal model is controlled to compare each target voice information with the input image (the image can be captured and input by a smart device). This comparison process can be performed by calculating the similarity between the voice and the image.

[0064]

[0065] in, Let I represent the similarity between the m-th image and the text, and let T represent the label of the text. m Let represent the label of the m-th image, exp() represent the exponential function, and τ represent the learnable temperature parameter. Let I represent the similarity between the m-th text and the image, T represent the image label, and I represent the image label. m This represents the label of the m-th text, where M represents the number of texts.

[0066] Further, the target speech information is input into the target multimodal model, which extracts visual and linguistic features from the target speech information and generates positive and negative sample pairs. Based on the image-text relationship in the target speech information, the target multimodal model aligns the visual and linguistic features to obtain an initial alignment result. The target multimodal model minimizes the distance between the positive sample pairs to obtain a minimum positive sample pair distance and maximizes the distance between the negative sample pairs to obtain a maximum negative sample pair distance. Based on the minimum positive sample pair distance and the maximum negative sample pair distance, the target multimodal model optimizes the initial alignment result and outputs the alignment result of the target speech information.

[0067] In multimodal models, the calculated similarity can be used to compare and learn by minimizing the distance between positive samples (positive sample pairs) and maximizing the distance between negative samples (negative sample pairs).

[0068]

[0069] Among them, L itc The minimum distance between positive samples and the maximum distance between negative samples are represented by D, which is the learnable distance between samples and takes values ​​in the range (I, T). (I,T)~D H(y) represents the probability of a positive or negative sample pair. i2t (I), P i2t (I)) represents the distance between positive sample pairs, H(y) i2t (I), P i2t (I)) represents y i2t (I) and P i2t The cross-entropy of (I), y i2t (I) represents the formal similarity of the location labels corresponding to the currently selected text, P i2t (I) represents the probability set corresponding to all texts, H(y) t2i (T), P t2i (T) represents y t2i (T) and P t2i The cross-entropy of (T), y t2i (T) represents the formal similarity of the location labels corresponding to the currently selected image, P t2i (T) represents the probability set corresponding to all images. The smaller the cross-entropy, the closer the model's predicted distribution is to the true distribution, and the better the model's performance.

[0070] Furthermore, from the historical recognition data, the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot are extracted; the first historical task result is input into the target multimodal model for training to obtain a cyclic generalized linear model, and the smart card driver is controlled to deploy the text-to-speech conversion tool into the humanoid robot; the second historical task result is input into the target multimodal model for training to obtain an environmental object recognition model; the third historical task result is input into the target multimodal model for training to obtain an edge-side language model.

[0071] After the alignment operation is completed, different models are trained for different smart devices. Through training with aligned data and alternating optimization, these models have demonstrated strong performance in a variety of application scenarios, including content understanding, intelligent recommendation, and intelligent customer service, which significantly improves the efficiency of device training and the accuracy of training results.

[0072] Step S30: Transmit the alignment result to the embodied intelligent device, control the embodied intelligent device to run according to the alignment result, and obtain the running result of the embodied intelligent device.

[0073] The embodied intelligent device includes any one of the following: humanoid robot, robotic arm, robot, and robotic dog.

[0074] Specifically, the addresses of the humanoid robot, the robotic arm, the robot, or the robotic dog are obtained, and the alignment result is transmitted to the humanoid robot, the robotic arm, the robot, or the robotic dog according to the addresses; the humanoid robot, the robotic arm, the robot, or the robotic dog is controlled to perform tasks according to the alignment result, and the task results are received.

[0075] Specifically, based on the addresses of different embodied intelligent devices, the prompts input by the terminal are rewritten and sent to the corresponding devices, thereby reviewing the driver instructions and controlling each device; for example... Figure 6 As shown, after rewriting the prompt words, the prompt words also need to be evaluated to determine whether the rewritten prompt words are suitable as the final prompt words to generate driving instructions. The evaluation criteria include words of the nature such as nouns, adverbs, and verbs in the voice information. In the specific embodiment, these include "capabilities and roles," "insights," "statements," and "personality." If unqualified prompt words are found, they need to be supplemented according to the evaluation criteria until the generated prompt words are qualified. Then, driving instructions are issued according to the address to drive different embodied intelligent devices to operate, further improving the accuracy of device operation.

[0076] Further, if the control target is the humanoid robot, the alignment result is input into the loop generalized linear model, and a first control instruction is generated. Based on the first control instruction, the text-to-speech conversion tool is controlled to convert the speech information into text information and output a semantic response until a dialogue with the user is completed. The semantic response is generated by the humanoid robot based on the text information. If the control target is the robotic arm, the alignment result is input into the environmental object recognition model, and a second control instruction is generated. Based on the second control instruction, the robotic arm is controlled to perform a second target task, and the execution result of the second target task is obtained. If the control target is the robot, the alignment result is input into the edge-side language model, and a third control instruction is generated. Based on the third control instruction, the robot is controlled to perform a third target task, and the execution result of the third target task is obtained. If the control target is the robotic dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. Based on the fourth control instruction, the robotic dog is controlled to perform a fourth target task, and the execution result of the fourth target task is obtained.

[0077] Among them, such as Figure 7 As shown, after completing the above steps, you can directly input prompts on the terminal to control the operation of the target embodied smart device. You can also name this task. When the target embodied smart device receives the driving command, the terminal will display the server that received the command and provide real-time prompts to the user on the operation status of the target embodied smart device, making the task execution process visible and improving the efficiency of device management.

[0078] In training a multimodal model, this invention preprocesses the data and transforms different modal data into the same scale or distribution. By aligning the data through a graph structure, it improves the efficiency of model training and the performance of the trained model.

[0079] Furthermore, such as Figure 8 As shown, based on the above-described multimodal embodied intelligent device control method, the present invention also provides a multimodal embodied intelligent device control system, wherein the multimodal embodied intelligent device control system includes:

[0080] The model training module 51 is used to acquire the programmed smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain a target multimodal model.

[0081] Information processing module 52 is used to acquire voice information input by the user, input the voice information into the target multimodal model for alignment processing, and output the alignment result;

[0082] The device control module 53 is used to transmit the alignment result to the embodied intelligent device, control the embodied intelligent device to run according to the alignment result, and obtain the running result of the embodied intelligent device.

[0083] Furthermore, such as Figure 9 As shown, based on the above-described multimodal embodied intelligent device control method and system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 9 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0084] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a multimodal embodied intelligent device control program 40, which can be executed by the processor 10 to implement the multimodal embodied intelligent device control method of this application.

[0085] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the multimodal embodied intelligent device control method.

[0086] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.

[0087] In one embodiment, when the processor 10 executes the multimodal-based embodied smart device control program 40 in the memory 20, the following steps are performed:

[0088] Obtain the programmed smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain the target multimodal model;

[0089] Acquire user-inputted voice information, input the voice information into the target multimodal model for alignment processing, and output the alignment result;

[0090] The alignment result is transmitted to the embodied intelligent device, the embodied intelligent device is controlled to operate according to the alignment result, and the operation result of the embodied intelligent device is obtained.

[0091] The steps of acquiring the programmed smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model specifically include:

[0092] Obtain the initial smart card inserted into the PC and select AI bitstream data, wherein the AI ​​bitstream data is used to deploy a multimodal model;

[0093] The AI ​​bitstream data is burned into the initial smart card to obtain the smart card driver;

[0094] Construct a multimodal model and control the smart card driver to deploy the multimodal model to the PC based on the AI ​​bitstream data;

[0095] Historical recognition data is extracted from the knowledge base, a training set is constructed based on the historical recognition data, and the multimodal model is trained using the training set to obtain the target multimodal model.

[0096] The step of acquiring user-inputted voice information, inputting the voice information into the target multimodal model for alignment processing, and outputting the alignment result specifically includes:

[0097] The system acquires user-input voice information, analyzes the voice information using a prompt word writing framework, and obtains the target voice information.

[0098] The target speech information is input into the target multimodal model, which performs alignment processing on the target speech information and outputs the alignment result.

[0099] Specifically, the step of inputting the target speech information into the target multimodal model, and the target multimodal model performing alignment processing on the target speech information and outputting the alignment result includes:

[0100] The target speech information is input into the target multimodal model, which extracts the visual and linguistic features of the target speech information and generates positive and negative sample pairs.

[0101] The target multimodal model aligns the visual features and the linguistic features based on the image-text relationship in the target speech information to obtain an initial alignment result;

[0102] The target multimodal model minimizes the distance between the positive sample pairs to obtain the minimum positive sample pair distance, and maximizes the distance between the negative sample pairs to obtain the maximum negative sample pair distance;

[0103] The target multimodal model optimizes the initial alignment result based on the minimum positive sample pair distance and the maximum negative sample pair distance, and outputs the alignment result of the target speech information.

[0104] The embodied intelligent device includes any one of the following: humanoid robot, robotic arm, robot, and robotic dog;

[0105] The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device specifically includes:

[0106] Obtain the address of the humanoid robot, the robotic arm, the robot, or the robotic dog, and transmit the alignment result to the humanoid robot, the robotic arm, the robot, or the robotic dog according to the address;

[0107] Based on the alignment result, control the humanoid robot, the robotic arm, the robot, or the robotic dog to perform tasks and receive the task results.

[0108] The process of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device, further includes the following steps:

[0109] Extract the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot from the historical recognition data;

[0110] The results of the first historical task are input into the target multimodal model for training to obtain a loop generalized linear model, and the smart card driver is controlled to deploy the text-to-speech conversion tool into the humanoid robot.

[0111] The results of the second historical task are input into the target multimodal model for training to obtain an environmental object recognition model;

[0112] The results of the third historical task are input into the target multimodal model for training to obtain the edge language model.

[0113] Specifically, controlling the humanoid robot, the robotic arm, the robot, or the robotic dog to perform tasks based on the alignment result, and receiving the task results, includes:

[0114] If the control target is the humanoid robot, the alignment result is input into the loop generalized linear model, and a first control instruction is generated. According to the first control instruction, the text-to-speech conversion tool is controlled to convert the speech information into text information and output a semantic response until the dialogue with the user is completed. The semantic response is generated by the humanoid robot based on the text information.

[0115] If the control target is the robotic arm, the alignment result is input into the environmental object recognition model, and a second control command is generated. The robotic arm is controlled to perform a second target task according to the second control command, and the execution result of the second target task is obtained.

[0116] If the control target is the robot, the alignment result is input into the edge language model, and a third control instruction is generated. The robot is then controlled to perform a third target task according to the third control instruction, and the execution result of the third target task is obtained.

[0117] If the control target is the robotic dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. The robotic dog is then controlled to perform a fourth target task according to the fourth control instruction, and the execution result of the fourth target task is obtained.

[0118] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal embodied intelligent device control program, which, when executed by a processor, implements the steps of the multimodal embodied intelligent device control method described above.

[0119] In summary, this invention provides a method and related equipment for controlling an embodied intelligent device based on multimodality. The method includes: acquiring a programmed smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model; acquiring user-inputted voice information, inputting the voice information into the target multimodal model for alignment processing, and outputting the alignment result; transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and acquiring the operating result of the embodied intelligent device. This invention, during multimodal model training, preprocesses the data, converts different modal data into the same scale or distribution, and aligns the data using a graph structure, thereby improving the efficiency of model training and the performance of the trained model.

[0120] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.

[0121] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0122] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for controlling embodied intelligent devices based on multimodal modes, characterized in that, The multimodal-based embodied intelligent device control method includes: Obtain the programmed smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain the target multimodal model; Acquire user-inputted voice information, input the voice information into the target multimodal model for alignment processing, and output the alignment result; The process of acquiring user-inputted voice information, inputting the voice information into the target multimodal model for alignment processing, and outputting the alignment result specifically includes: The system acquires user-input voice information, analyzes the voice information using a prompt word writing framework, and obtains the target voice information. The target speech information is input into the target multimodal model, the target multimodal model performs alignment processing on the target speech information, and outputs the alignment result; The step of inputting the target speech information into the target multimodal model, the target multimodal model performing alignment processing on the target speech information, and outputting the alignment result specifically includes: The target speech information is input into the target multimodal model, which extracts the visual and linguistic features of the target speech information and generates positive and negative sample pairs. The target multimodal model aligns the visual features and the linguistic features based on the image-text relationship in the target speech information to obtain an initial alignment result; The target multimodal model minimizes the distance between the positive sample pairs to obtain the minimum positive sample pair distance, and maximizes the distance between the negative sample pairs to obtain the maximum negative sample pair distance; The target multimodal model optimizes the initial alignment result based on the minimum positive sample pair distance and the maximum negative sample pair distance, and outputs the alignment result of the target speech information; The alignment result is transmitted to the embodied intelligent device, the embodied intelligent device is controlled to operate according to the alignment result, and the operation result of the embodied intelligent device is obtained.

2. The control method for embodied intelligent devices based on multimodality according to claim 1, characterized in that, The steps of acquiring the programmed smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model specifically include: Obtain the initial smart card inserted into the PC and select AI bitstream data, wherein the AI ​​bitstream data is used to deploy a multimodal model; The AI ​​bitstream data is burned into the initial smart card to obtain the smart card driver; Construct a multimodal model and control the smart card driver to deploy the multimodal model to the PC based on the AI ​​bitstream data; Historical recognition data is extracted from the knowledge base, a training set is constructed based on the historical recognition data, and the multimodal model is trained using the training set to obtain the target multimodal model.

3. The control method for embodied intelligent devices based on multimodal operation according to claim 1, characterized in that, The embodied intelligent device includes any one of the following: humanoid robot, robotic arm, robot, and robotic dog; The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device specifically includes: Obtain the address of the humanoid robot, the robotic arm, the robot, or the robotic dog, and transmit the alignment result to the humanoid robot, the robotic arm, the robot, or the robotic dog according to the address; Based on the alignment result, control the humanoid robot, the robotic arm, the robot, or the robotic dog to perform tasks and receive the task results.

4. The control method for embodied intelligent devices based on multimodality according to claim 3, characterized in that, The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device, further includes: Extract the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot from the historical recognition data; The results of the first historical task are input into the target multimodal model for training to obtain a loop generalized linear model, and the smart card driver is controlled to deploy the text-to-speech conversion tool into the humanoid robot. The results of the second historical task are input into the target multimodal model for training to obtain an environmental object recognition model; The results of the third historical task are input into the target multimodal model for training to obtain the edge language model.

5. The control method for embodied intelligent devices based on multimodality according to claim 4, characterized in that, The step of controlling the humanoid robot, the robotic arm, the robot, or the robotic dog to perform tasks based on the alignment result, and receiving the task results, specifically includes: If the control target is the humanoid robot, the alignment result is input into the loop generalized linear model, and a first control instruction is generated. According to the first control instruction, the text-to-speech conversion tool is controlled to convert the speech information into text information and output a semantic response until the dialogue with the user is completed. The semantic response is generated by the humanoid robot based on the text information. If the control target is the robotic arm, the alignment result is input into the environmental object recognition model, and a second control command is generated. The robotic arm is controlled to perform a second target task according to the second control command, and the execution result of the second target task is obtained. If the control target is the robot, the alignment result is input into the edge language model, and a third control instruction is generated. The robot is then controlled to perform a third target task according to the third control instruction, and the execution result of the third target task is obtained. If the control target is the robotic dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. The robotic dog is then controlled to perform a fourth target task according to the fourth control instruction, and the execution result of the fourth target task is obtained.

6. A control system for embodied intelligent devices based on multimodal modes, characterized in that, The multimodal-based embodied intelligent device control system is applied to the multimodal-based embodied intelligent device control method as described in any one of claims 1-5, wherein the multimodal-based embodied intelligent device control system comprises: The model training module is used to acquire the programmed smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain a target multimodal model. The information processing module is used to acquire the voice information input by the user, input the voice information into the target multimodal model for alignment processing, and output the alignment result; The device control module is used to transmit the alignment result to the embodied intelligent device, control the embodied intelligent device to operate according to the alignment result, and obtain the operating result of the embodied intelligent device.

7. A terminal, characterized in that, The terminal includes: a memory, a processor, and a multimodal embodied intelligent device control program stored in the memory and executable on the processor. When the multimodal embodied intelligent device control program is executed by the processor, it implements the steps of the multimodal embodied intelligent device control method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multimodal embodied intelligent device control program, which, when executed by a processor, implements the steps of the multimodal embodied intelligent device control method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Intelligent multi-modal evaluation implementation method and universe intelligent evaluation system

    CN117875918A