Multi-mode-based intelligent device control method, system and terminal

By deploying and training multimodal models in smart card drivers and aligning the voice information input by users, the problem of data misalignment in multimodal model training is solved, and training efficiency and model performance are improved.

CN120012045AActive Publication Date: 2025-05-16SHENZHEN MAITEXIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411893029.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

In the prior art, the data of multimodal models are not aligned during training, resulting in low training efficiency and inaccurate results.

Method used

By obtaining the burned smart card driver, deploying the multimodal model, and training the model, the target multimodal model is obtained. Then, the voice information input by the user is obtained, the alignment process is performed, the alignment result is output, and the result is transmitted to the embodied intelligent device to control the operation of the device.

Benefits of technology

By preprocessing the data and aligning the graph structure, the efficiency and performance of multimodal model training are improved, and the training efficiency and inaccurate results are solved due to data misalignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012045A_ABST
    Figure CN120012045A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal-based control method, system and terminal for a smart device with a body, and the method comprises the steps: obtaining a burnt smart card drive, controlling the smart card drive, deploying a multi-modal model, and training the multi-modal model to obtain a target multi-modal model; acquiring voice information input by a user, inputting the voice information into the target multi-modal model, and outputting prompt words; and transmitting the prompt word to the own intelligent equipment, controlling the own intelligent equipment to operate, and obtaining an operation result. When the multi-modal model is trained, after the data is preprocessed, the data of different modals are converted into the data of the same scale or distribution, and the data are aligned through a graph structure, so that the efficiency of model training and the performance of the trained model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal embodied intelligent device control method, system, terminal and computer-readable storage medium. Background Art

[0002] Embodied intelligence is a developing field of artificial intelligence, which refers to the ability of an intelligent system or machine to interact with the environment in real time through perception and interaction. The multimodal generalization capability of large models also lays the foundation for the robot's autonomous learning ability, helping the robot to adapt to changing tasks.

[0003] Currently, large-model multimodal technology is developing rapidly, combining technologies from multiple fields such as computer vision, natural language processing, and speech recognition.

[0004] However, data of different modalities are often misaligned in time, space, and semantics, and have different statistical properties and distributions. This requires the model to be able to process heterogeneous data and extract useful information from it. In the process of multimodal model training, multimodal data usually requires complex annotation, which is not only costly, but the quality of the annotation directly affects the performance of the model.

[0005] Therefore, the prior art still needs to be improved and developed. Summary of the invention

[0006] The main purpose of the present invention is to provide a multimodal-based embodied intelligent device control method, system, terminal and computer-readable storage medium, aiming to solve the problem in the prior art that the multimodal model training process is inefficient and the results are inaccurate due to data misalignment during the training process.

[0007] To achieve the above object, the present invention provides a multi-modal embodied intelligent device control method, the multi-modal embodied intelligent device control method comprising the following steps:

[0008] Obtaining a burned smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model;

[0009] Acquire voice information input by a user, input the voice information into the target multimodal model for alignment processing, and output the alignment result;

[0010] The alignment result is transmitted to the embodied intelligent device, the embodied intelligent device is controlled to operate according to the alignment result, and the operation result of the embodied intelligent device is obtained.

[0011] Optionally, the multimodal-based embodied smart device control method, wherein the obtaining of a burned smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model, specifically includes:

[0012] Obtain an initial smart card inserted into a PC and select AI bitstream data, wherein the AI ​​bitstream data is used to deploy a multimodal model;

[0013] Burning the AI ​​bit stream data into the initial smart card to obtain the smart card driver;

[0014] Constructing a multimodal model, and controlling the smart card driver to deploy the multimodal model to the PC according to the AI ​​bitstream data;

[0015] Historical recognition data is extracted from a knowledge base, a training set is constructed according to the historical recognition data, and the multimodal model is trained using the training set to obtain a target multimodal model.

[0016] Optionally, the multimodal-based embodied intelligent device control method, wherein the acquiring of voice information input by the user, inputting the voice information into the target multimodal model for alignment processing, and outputting the alignment result, specifically includes:

[0017] Acquire voice information input by the user, analyze the voice information through the prompt word writing framework, and obtain target voice information;

[0018] The target speech information is input into the target multimodal model, and the target multimodal model performs alignment processing on the target speech information and outputs an alignment result.

[0019] Optionally, the multimodal-based embodied intelligent device control method, wherein the step of inputting the target voice information into the target multimodal model, the target multimodal model performing alignment processing on the target voice information, and outputting an alignment result, specifically includes:

[0020] Inputting the target speech information into the target multimodal model, the target multimodal model extracting visual features and language features of the target speech information and generating positive sample pairs and negative sample pairs;

[0021] The target multimodal model performs alignment processing on the visual features and the language features according to the image-text association relationship in the target speech information to obtain an initial alignment result;

[0022] The target multimodal model minimizes the distance between the positive sample pairs to obtain the minimum positive sample pair distance, and maximizes the distance between the negative sample pairs to obtain the maximum negative sample pair distance;

[0023] The target multimodal model optimizes the initial alignment result according to the minimum positive sample pair distance and the maximum negative sample pair distance, and outputs the alignment result of the target speech information.

[0024] Optionally, in the multimodal embodied intelligent device control method, the embodied intelligent device comprises: any one of a humanoid robot, a mechanical arm, a robot and a mechanical dog;

[0025] The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device specifically includes:

[0026] Acquire the address of the humanoid robot, the robotic arm, the robot or the robotic dog, and transmit the alignment result to the humanoid robot, the robotic arm, the robot or the robotic dog according to the address;

[0027] The humanoid robot, the robotic arm, the robot or the robotic dog is controlled to perform a task according to the alignment result, and a task result is received.

[0028] Optionally, the multimodal embodied intelligent device control method, wherein the step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device, further comprises:

[0029] Extracting the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot from the historical recognition data;

[0030] Inputting the first historical task result into the target multimodal model for training to obtain a ring generalized linear model, and controlling the smart card driver to deploy a text-to-speech conversion tool into the humanoid robot;

[0031] Inputting the second historical task result into the target multimodal model for training to obtain an environmental object recognition model;

[0032] The third historical task result is input into the target multimodal model for training to obtain a terminal-side language model.

[0033] Optionally, the multimodal embodied intelligent device control method, wherein the step of controlling the humanoid robot, the robotic arm, the robot or the robotic dog to perform a task according to the alignment result and receiving the task result specifically includes:

[0034] If the control target is the humanoid robot, the alignment result is input into the ring generalized linear model, and a first control instruction is generated. According to the first control instruction, the text-to-speech conversion tool is controlled to convert the voice information into text information and output a semantic response until the conversation with the user is completed, wherein the semantic response is generated by the humanoid robot according to the text information;

[0035] If the control target is the robot arm, the alignment result is input into the environmental object recognition model, and a second control instruction is generated, the robot arm is controlled to perform a second target task according to the second control instruction, and the execution result of the second target task is obtained;

[0036] If the control target is the robot, input the alignment result into the terminal-side language model, generate a third control instruction, control the robot to perform a third target task according to the third control instruction, and obtain the execution result of the third target task;

[0037] If the control target is the robot dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. According to the fourth control instruction, the robot dog is controlled to perform a fourth target task, and the execution result of the fourth target task is obtained.

[0038] In addition, to achieve the above-mentioned purpose, the present invention further provides a multi-modal embodied intelligent device control system, wherein the multi-modal embodied intelligent device control system comprises:

[0039] A model training module is used to obtain a burned smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain a target multimodal model;

[0040] An information processing module, used for acquiring voice information input by a user, inputting the voice information into the target multimodal model for alignment processing, and outputting an alignment result;

[0041] The device control module is used to transmit the alignment result to the embodied intelligent device, control the embodied intelligent device to operate according to the alignment result, and obtain the operation result of the embodied intelligent device.

[0042] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a multimodal embodied intelligent device control program stored in the memory and executable on the processor, and the multimodal embodied intelligent device control program, when executed by the processor, implements the steps of the multimodal embodied intelligent device control method as described above.

[0043] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal-based embodied intelligent device control program, and when the multimodal-based embodied intelligent device control program is executed by a processor, the steps of the multimodal-based embodied intelligent device control method as described above are implemented.

[0044] In the present invention, a burned smart card driver is obtained, the smart card driver is controlled to deploy a multimodal model, and the multimodal model is trained to obtain a target multimodal model; the voice information input by the user is obtained, and the voice information is input into the target multimodal model for alignment processing, and the alignment result is output; the alignment result is transmitted to the embodied smart device, the embodied smart device is controlled to run according to the alignment result, and the running result of the embodied smart device is obtained. When training a multimodal model, the present invention converts different modal data into the same scale or distribution after preprocessing the data, and aligns the data through a graph structure, thereby improving the efficiency of model training and the performance of the model after training. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a flow chart of a preferred embodiment of the multi-modal embodied intelligent device control method of the present invention;

[0046] Figure 2 It is a schematic diagram of AI model driven bitstream burning of a preferred embodiment of the multi-modal embodied intelligent device control method of the present invention;

[0047] Figure 3 It is a flowchart of AI model driven bitstream burning of a preferred embodiment of the multi-modal embodied intelligent device control method of the present invention;

[0048] Figure 4 It is an AI model interaction diagram of a preferred embodiment of the multi-modal embodied intelligent device control method of the present invention;

[0049] Figure 5 It is a flow chart of a smart card localized driving device of a preferred embodiment of the multi-modal embodied smart device control method of the present invention;

[0050] Figure 6It is a flow chart of a custom prompt word voice of a preferred embodiment of the multi-modal embodied intelligent device control method of the present invention;

[0051] Figure 7 It is a field diagram of a custom prompt word of a preferred embodiment of the multi-modal embodied intelligent device control method of the present invention;

[0052] Figure 8 It is a structural diagram of a preferred embodiment of the multi-modal embodied intelligent device control system of the present invention;

[0053] Fig. 9 Schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0055] The multi-modal embodied intelligent device control method described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the multi-modal embodied intelligent device control method includes the following steps:

[0056] Step S10: obtain a burned smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain a target multimodal model.

[0057] Among them, Figure 2 and Figure 3 As shown, insert the smart card into the PC, and burn the preset bit stream data into the smart card through the burning software to obtain the smart card driver.

[0058] Specifically, an initial smart card inserted into a PC is obtained, and AI bit stream data is selected, wherein the AI ​​bit stream data is used to deploy a multimodal model; the AI ​​bit stream data is burned into the initial smart card to obtain the smart card driver; a multimodal model is constructed, and the smart card driver is controlled to deploy the multimodal model to the PC according to the AI ​​bit stream data; historical recognition data is extracted from a knowledge base, a training set is constructed according to the historical recognition data, and the multimodal model is trained using the training set to obtain a target multimodal model.

[0059] Among them, Figure 4As shown, the AI ​​model is trained through the historical recognition data in the knowledge base, so as to realize the construction of the smart card driver (that is, burning the bitstream data onto the smart card). By enriching the knowledge base of the large model, the efficiency of training the intelligent ability of the large model is improved. At the same time, the recognition ability of language semantics is improved to achieve precise voice control.

[0060] Step S20: Acquire the voice information input by the user, input the voice information into the target multimodal model for alignment processing, and output the alignment result.

[0061] Among them, the method for implementing alignment processing here can be the ALBEF (Align before Fuse, a multimodal representation learning method) learning method, which can efficiently learn the joint representation of vision and language through image-text alignment and momentum distillation technology, improve the accuracy of recognizing voice information, and thus improve the accuracy of subsequent control of smart devices.

[0062] Specifically, voice information input by the user is obtained, and the voice information is analyzed through a prompt word writing framework to obtain target voice information; the target voice information is input into the target multimodal model, and the target multimodal model performs alignment processing on the target voice information and outputs an alignment result.

[0063] Among them, Figure 5 As shown, by analyzing the voice information input by the user (i.e. preprocessing, reducing invalid information in the voice information, and improving processing efficiency), the target voice information (which can be converted into text) is obtained, and then the multimodal model is controlled to compare each target voice information with the input image (the image can be taken and input by a smart device). This comparison process can be performed by calculating the similarity between the voice and the image:

[0064]

[0065] in, represents the similarity between the mth image and the text, I represents the label of the text, T m represents the label of the mth image, exp() represents the exponential function, τ represents the learnable temperature parameter, represents the similarity between the mth text and the image, T represents the label of the image, and I m Represents the label of the mth text, and M represents the number of texts.

[0066] Furthermore, the target speech information is input into the target multimodal model, and the target multimodal model extracts the visual features and language features of the target speech information and generates positive sample pairs and negative sample pairs; the target multimodal model aligns the visual features and the language features according to the image-text association relationship in the target speech information to obtain an initial alignment result; the target multimodal model minimizes the distance between the positive sample pairs to obtain the minimum positive sample pair distance, and maximizes the distance between the negative sample pairs to obtain the maximum negative sample pair distance; the target multimodal model optimizes the initial alignment result according to the minimum positive sample pair distance and the maximum negative sample pair distance, and outputs the alignment result of the target speech information.

[0067] Among them, in the multimodal model, the distance between positive samples (positive sample pairs) can be minimized and the distance between negative samples (negative sample pairs) can be maximized for comparative learning based on the calculated similarity:

[0068]

[0069] Among them, L itc It represents the minimum distance between positive samples and the maximum distance between negative samples, which is determined by D. D represents the distance between learnable samples, and its value range is (I, T). (I,T)~D represents the probability of a positive sample pair or a negative sample pair, H(y i2t (I), P i2t (I)) represents the distance between positive sample pairs, H(y i2t (I), P i2t (I)) represents y i2t (I) and P i2t (I) The cross entropy, y i2t (I) represents the formal similarity of the position tag corresponding to the currently selected text, P i2t (I) represents the probability set corresponding to all texts, H(y t2i (T), P t2i (T)) represents y t2i (T) and P t2i The cross entropy of (T), y t2i (T) represents the form similarity of the position label corresponding to the currently selected image, P t2i (T) represents the probability set corresponding to all images, where the smaller the cross entropy, the closer the predicted distribution of the model is to the true distribution, and the better the performance of the model.

[0070] Furthermore, in the historical recognition data, the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot are extracted; the first historical task result is input into the target multimodal model for training to obtain a ring generalized linear model, and the smart card driver is controlled to deploy a text-to-speech conversion tool to the humanoid robot; the second historical task result is input into the target multimodal model for training to obtain an environmental object recognition model; the third historical task result is input into the target multimodal model for training to obtain an end-side language model.

[0071] After completing the alignment operation, different models are trained according to different smart devices. Through aligned data training and alternating optimization, these models have demonstrated strong performance in a variety of application scenarios, including content understanding, intelligent recommendation, and intelligent customer service, significantly improving the efficiency of device training and the accuracy of training results.

[0072] Step S30: transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device.

[0073] The embodied intelligent device includes any one of a humanoid robot, a robotic arm, a robot and a robotic dog.

[0074] Specifically, the address of the humanoid robot, the robotic arm, the robot or the robotic dog is obtained, and the alignment result is transmitted to the humanoid robot, the robotic arm, the robot or the robotic dog according to the address; the humanoid robot, the robotic arm, the robot or the robotic dog is controlled to perform a task according to the alignment result, and the task result is received.

[0075] Among them, according to the addresses of different embodied intelligent devices, the prompt words input by the terminal are rewritten and sent to the corresponding devices, so as to review the driving instructions and control each device; Figure 6 As shown, after rewriting the prompt word, it is necessary to evaluate the prompt word to determine whether the rewritten prompt word is suitable for generating a driving instruction as the final prompt word. The evaluation criteria include words of the nature of nouns, adverbs and verbs in the voice information. Specifically in the embodiment, they include "ability and role", "insight", "statement" and "personality". If unqualified prompt words are found, they need to be supplemented according to the evaluation criteria until the generated prompt word is qualified. Then, a driving instruction is issued according to the address to drive different embodied intelligent devices to run, thereby further improving the accuracy of device operation.

[0076] Furthermore, if the control target is the humanoid robot, the alignment result is input into the ring generalized linear model, and a first control instruction is generated. According to the first control instruction, the text-to-speech conversion tool is controlled to convert the voice information into text information, and a semantic response is output until the conversation with the user is completed, wherein the semantic response is generated by the humanoid robot according to the text information; if the control target is the robotic arm, the alignment result is input into the environmental object recognition model, and a second control instruction is generated. According to the second control instruction, the robotic arm is controlled to perform a second target task, and the execution result of the second target task is obtained; if the control target is the robot, the alignment result is input into the terminal language model, and a third control instruction is generated. According to the third control instruction, the robot is controlled to perform a third target task, and the execution result of the third target task is obtained; if the control target is the robotic dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. According to the fourth control instruction, the robotic dog is controlled to perform a fourth target task, and the execution result of the fourth target task is obtained.

[0077] Among them, Figure 7 As shown, after implementing the above steps, you can directly enter prompt words in the terminal to control the operation of the target embodied intelligent device, and you can also name this task. When the target embodied intelligent device receives the driving instruction, the terminal will display the server that receives the command, and prompt the user in real time about the operation status of the target embodied intelligent device, so that the task execution process is visualized and the management efficiency of the device is improved.

[0078] When training a multimodal model, the present invention converts different modal data into the same scale or distribution after preprocessing the data, and aligns the data through a graph structure, thereby improving the efficiency of model training and the performance of the model after training.

[0079] Furthermore, if Figure 8 As shown, based on the above-mentioned multi-modal embodied intelligent device control method, the present invention also provides a multi-modal embodied intelligent device control system, wherein the multi-modal embodied intelligent device control system includes:

[0080] The model training module 51 is used to obtain a burned smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain a target multimodal model;

[0081] The information processing module 52 is used to obtain the voice information input by the user, input the voice information into the target multimodal model for alignment processing, and output the alignment result;

[0082] The device control module 53 is used to transmit the alignment result to the embodied intelligent device, control the embodied intelligent device to operate according to the alignment result, and obtain the operation result of the embodied intelligent device.

[0083] Furthermore, if Fig. 9 As shown, based on the above-mentioned multi-modal embodied intelligent device control method and system, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Fig. 9 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0084] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed in the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a multimodal embodied intelligent device control program 40 is stored on the memory 20, and the multimodal embodied intelligent device control program 40 can be executed by the processor 10, thereby realizing the multimodal embodied intelligent device control method based on the present application.

[0085] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the multi-modal embodied intelligent device control method.

[0086] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.

[0087] In one embodiment, when the processor 10 executes the multimodal embodied intelligent device control program 40 in the memory 20, the following steps are implemented:

[0088] Obtaining a burned smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model;

[0089] Acquire voice information input by a user, input the voice information into the target multimodal model for alignment processing, and output the alignment result;

[0090] The alignment result is transmitted to the embodied intelligent device, the embodied intelligent device is controlled to operate according to the alignment result, and the operation result of the embodied intelligent device is obtained.

[0091] The step of obtaining a burned smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model specifically includes:

[0092] Obtain an initial smart card inserted into a PC and select AI bitstream data, wherein the AI ​​bitstream data is used to deploy a multimodal model;

[0093] Burning the AI ​​bit stream data into the initial smart card to obtain the smart card driver;

[0094] Constructing a multimodal model, and controlling the smart card driver to deploy the multimodal model to the PC according to the AI ​​bitstream data;

[0095] Historical recognition data is extracted from a knowledge base, a training set is constructed according to the historical recognition data, and the multimodal model is trained using the training set to obtain a target multimodal model.

[0096] The acquiring of voice information input by a user, inputting the voice information into the target multimodal model for alignment processing, and outputting an alignment result specifically includes:

[0097] Acquire voice information input by the user, analyze the voice information through the prompt word writing framework, and obtain target voice information;

[0098] The target speech information is input into the target multimodal model, and the target multimodal model performs alignment processing on the target speech information and outputs an alignment result.

[0099] The step of inputting the target speech information into the target multimodal model, wherein the target multimodal model performs alignment processing on the target speech information and outputs an alignment result, specifically includes:

[0100] Inputting the target speech information into the target multimodal model, the target multimodal model extracting visual features and language features of the target speech information and generating positive sample pairs and negative sample pairs;

[0101] The target multimodal model performs alignment processing on the visual features and the language features according to the image-text association relationship in the target speech information to obtain an initial alignment result;

[0102] The target multimodal model minimizes the distance between the positive sample pairs to obtain the minimum positive sample pair distance, and maximizes the distance between the negative sample pairs to obtain the maximum negative sample pair distance;

[0103] The target multimodal model optimizes the initial alignment result according to the minimum positive sample pair distance and the maximum negative sample pair distance, and outputs the alignment result of the target speech information.

[0104] Wherein, the embodied intelligent device includes: any one of a humanoid robot, a mechanical arm, a robot and a mechanical dog;

[0105] The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device specifically includes:

[0106] Acquire the address of the humanoid robot, the robotic arm, the robot or the robotic dog, and transmit the alignment result to the humanoid robot, the robotic arm, the robot or the robotic dog according to the address;

[0107] The humanoid robot, the robotic arm, the robot or the robotic dog is controlled to perform a task according to the alignment result, and a task result is received.

[0108] The method of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device, further includes:

[0109] Extracting the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot from the historical recognition data;

[0110] Inputting the first historical task result into the target multimodal model for training to obtain a ring generalized linear model, and controlling the smart card driver to deploy a text-to-speech conversion tool into the humanoid robot;

[0111] Inputting the second historical task result into the target multimodal model for training to obtain an environmental object recognition model;

[0112] The third historical task result is input into the target multimodal model for training to obtain a terminal-side language model.

[0113] Wherein, controlling the humanoid robot, the mechanical arm, the robot or the mechanical dog to perform a task according to the alignment result and receiving the task result specifically includes:

[0114] If the control target is the humanoid robot, the alignment result is input into the ring generalized linear model, and a first control instruction is generated. According to the first control instruction, the text-to-speech conversion tool is controlled to convert the voice information into text information and output a semantic response until the conversation with the user is completed, wherein the semantic response is generated by the humanoid robot according to the text information;

[0115] If the control target is the robot arm, the alignment result is input into the environmental object recognition model, and a second control instruction is generated, the robot arm is controlled to perform a second target task according to the second control instruction, and the execution result of the second target task is obtained;

[0116] If the control target is the robot, input the alignment result into the terminal-side language model, generate a third control instruction, control the robot to perform a third target task according to the third control instruction, and obtain the execution result of the third target task;

[0117] If the control target is the robot dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. According to the fourth control instruction, the robot dog is controlled to perform a fourth target task, and the execution result of the fourth target task is obtained.

[0118] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multimodal-based embodied intelligent device control program, and when the multimodal-based embodied intelligent device control program is executed by a processor, the steps of the multimodal-based embodied intelligent device control method as described above are implemented.

[0119] In summary, the present invention provides a multimodal-based embodied intelligent device control method and related equipment, the method comprising: obtaining a burned smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model; obtaining voice information input by a user, and inputting the voice information into the target multimodal model for alignment processing, and outputting an alignment result; transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to run according to the alignment result, and obtaining the running result of the embodied intelligent device. When training a multimodal model, the present invention converts different modal data into the same scale or distribution after preprocessing the data, and aligns the data through a graph structure, thereby improving the efficiency of model training and the performance of the model after training.

[0120] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or terminal including the element.

[0121] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.

[0122] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A multi-modal embodied intelligent device control method, characterized in that: The multi-modal embodied intelligent device control method includes: Obtaining a burned smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model; Acquire voice information input by a user, input the voice information into the target multimodal model for alignment processing, and output the alignment result; The alignment result is transmitted to the embodied intelligent device, the embodied intelligent device is controlled to operate according to the alignment result, and the operation result of the embodied intelligent device is obtained.

2. The multimodal embodied intelligent device control method according to claim 1, characterized in that: The step of obtaining a burned smart card driver, controlling the smart card driver to deploy a multimodal model, and training the multimodal model to obtain a target multimodal model specifically includes: Obtain an initial smart card inserted into a PC and select AI bitstream data, wherein the AI ​​bitstream data is used to deploy a multimodal model; Burning the AI ​​bit stream data into the initial smart card to obtain the smart card driver; Constructing a multimodal model, and controlling the smart card driver to deploy the multimodal model to the PC according to the AI ​​bitstream data; Historical recognition data is extracted from a knowledge base, a training set is constructed according to the historical recognition data, and the multimodal model is trained using the training set to obtain a target multimodal model.

3. The multimodal embodied intelligent device control method according to claim 1, characterized in that: The acquiring of voice information input by the user, inputting the voice information into the target multimodal model for alignment processing, and outputting the alignment result specifically includes: Acquire voice information input by the user, analyze the voice information through the prompt word writing framework, and obtain target voice information; The target speech information is input into the target multimodal model, and the target multimodal model performs alignment processing on the target speech information and outputs an alignment result.

4. The multimodal embodied intelligent device control method according to claim 3, characterized in that: The step of inputting the target speech information into the target multimodal model, wherein the target multimodal model performs alignment processing on the target speech information and outputs an alignment result, specifically includes: Inputting the target speech information into the target multimodal model, the target multimodal model extracting visual features and language features of the target speech information and generating positive sample pairs and negative sample pairs; The target multimodal model performs alignment processing on the visual features and the language features according to the image-text association relationship in the target speech information to obtain an initial alignment result; The target multimodal model minimizes the distance between the positive sample pairs to obtain the minimum positive sample pair distance, and maximizes the distance between the negative sample pairs to obtain the maximum negative sample pair distance; The target multimodal model optimizes the initial alignment result according to the minimum positive sample pair distance and the maximum negative sample pair distance, and outputs the alignment result of the target speech information.

5. The multimodal embodied intelligent device control method according to claim 1, characterized in that: The embodied intelligent device includes: any one of a humanoid robot, a robotic arm, a robot, and a robotic dog; The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device specifically includes: Acquire the address of the humanoid robot, the robotic arm, the robot or the robotic dog, and transmit the alignment result to the humanoid robot, the robotic arm, the robot or the robotic dog according to the address; The humanoid robot, the robotic arm, the robot or the robotic dog is controlled to perform a task according to the alignment result, and a task result is received.

6. The multimodal embodied intelligent device control method according to claim 5, characterized in that: The step of transmitting the alignment result to the embodied intelligent device, controlling the embodied intelligent device to operate according to the alignment result, and obtaining the operating result of the embodied intelligent device, further comprises: Extracting the first historical task result of the humanoid robot, the second historical task result of the robotic arm, and the third historical task result of the robot from the historical recognition data; Inputting the first historical task result into the target multimodal model for training to obtain a ring generalized linear model, and controlling the smart card driver to deploy a text-to-speech conversion tool into the humanoid robot; Inputting the second historical task result into the target multimodal model for training to obtain an environmental object recognition model; The third historical task result is input into the target multimodal model for training to obtain a terminal-side language model.

7. The multimodal embodied intelligent device control method according to claim 6, characterized in that: The controlling the humanoid robot, the mechanical arm, the robot or the mechanical dog to perform a task according to the alignment result and receiving the task result specifically includes: If the control target is the humanoid robot, the alignment result is input into the ring generalized linear model, and a first control instruction is generated. According to the first control instruction, the text-to-speech conversion tool is controlled to convert the voice information into text information and output a semantic response until the conversation with the user is completed, wherein the semantic response is generated by the humanoid robot according to the text information; If the control target is the robot arm, the alignment result is input into the environmental object recognition model, and a second control instruction is generated, the robot arm is controlled to perform a second target task according to the second control instruction, and the execution result of the second target task is obtained; If the control target is the robot, input the alignment result into the terminal-side language model, generate a third control instruction, control the robot to perform a third target task according to the third control instruction, and obtain the execution result of the third target task; If the control target is the robot dog, the alignment result is input into the target multimodal model, and a fourth control instruction is generated. According to the fourth control instruction, the robot dog is controlled to perform a fourth target task, and the execution result of the fourth target task is obtained.

8. A multi-modal embodied intelligent device control system, characterized in that: The multi-modal embodied intelligent device control system includes: A model training module is used to obtain a burned smart card driver, control the smart card driver to deploy a multimodal model, and train the multimodal model to obtain a target multimodal model; An information processing module, used for acquiring voice information input by a user, inputting the voice information into the target multimodal model for alignment processing, and outputting an alignment result; The device control module is used to transmit the alignment result to the embodied intelligent device, control the embodied intelligent device to operate according to the alignment result, and obtain the operation result of the embodied intelligent device.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a multimodal embodied intelligent device control program stored in the memory and executable on the processor. When the multimodal embodied intelligent device control program is executed by the processor, the steps of the multimodal embodied intelligent device control method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multimodal embodied intelligent device control program, and when the multimodal embodied intelligent device control program is executed by a processor, the steps of the multimodal embodied intelligent device control method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Intelligent multi-modal evaluation implementation method and universe intelligent evaluation system

    CN117875918A