Humanoid robot decision interaction method, system and equipment and medium

By combining multimodal large models and industrial small models, the problem of perception bias in humanoid robots in complex industrial environments is solved, enabling real-time and accurate generation of industrial environment interaction strategies and improving the robot's decision-making efficiency.

CN120974368APending Publication Date: 2025-11-18广州里工实业有限公司

Patent Information

Application Number
CN202511051103.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing technologies, humanoid robots have difficulty effectively processing multi-source heterogeneous data such as industrial vision, hearing, and touch, leading to industrial perception biases, difficulty in generating safe and effective industrial interaction strategies, and risks to production safety and decision-making errors.

Method used

A multimodal large model is used for environmental assessment and anomaly identification, combined with an industrial small model for in-depth processing to generate comprehensive industrial environment interaction strategies. Through cross-modal feature alignment and strategy decoding, a balance between industrial real-time performance and accuracy is achieved.

Benefits of technology

It improves the decision-making and interaction efficiency of humanoid robots, reduces unnecessary computational overhead, and ensures real-time and accurate interaction in industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974368A_ABST
    Figure CN120974368A_ABST
Patent Text Reader

Abstract

The invention discloses a humanoid robot decision interaction method, system and device and a medium. The method comprises the steps of obtaining industrial sensor data; inputting the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result; inputting the industrial sensor data into an anomaly recognition model according to the industrial environment evaluation result to obtain an anomaly recognition result; inputting the industrial sensor data and the anomaly recognition result into the multi-modal large model for strategy generation processing to obtain an interaction strategy; and performing interaction control on the humanoid robot according to the interaction strategy. The decision interaction efficiency of the humanoid robot can be improved, and the method can be widely applied to the technical field of intelligent robots.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent robots, and in particular to a humanoid robot decision interaction method, system, device and medium. BACKGROUND

[0002] In industrial manufacturing, industrial inspection and other industrial scenarios, the application of humanoid robots is becoming more and more widespread. In the related art, industrial robots mostly use a single industrial sensor or a perception model that fuses simple industrial sensors, which is difficult to handle the collaborative understanding of industrial visual, industrial auditory, industrial tactile and other industrial multi-source heterogeneous data, and is prone to industrial perception deviation in complex industrial scenarios, resulting in difficulty in generating safe and effective industrial interaction strategies, and existence of industrial production safety risks or industrial decision errors.

[0003] To sum up, the technical problems existing in the related art need to be improved. SUMMARY

[0004] The main purpose of the embodiments of the present application is to provide a humanoid robot decision interaction method, system, device and medium, which can improve the interaction efficiency of the humanoid robot.

[0005] To achieve the above-mentioned purpose, one aspect of the embodiments of the present application provides a humanoid robot decision interaction method, which comprises:

[0006] obtaining industrial sensor data;

[0007] inputting the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result;

[0008] inputting the industrial sensor data into an anomaly recognition model according to the industrial environment evaluation result to obtain an anomaly recognition result;

[0009] inputting the industrial sensor data and the anomaly recognition result into the multi-modal large model for strategy generation processing to obtain an interaction strategy;

[0010] controlling the humanoid robot according to the interaction strategy.

[0011] In some embodiments, the inputting the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result comprises:

[0012] obtaining an industrial preset question text;

[0013] inputting the industrial preset question text and the industrial sensor data into the multi-modal large model;

[0014] performing feature extraction processing on the industrial sensor data to obtain multi-modal perception features;

[0015] performing word segmentation and position coding processing on the industrial preset problem text to obtain a first semantic vector;

[0016] performing cross-modal fusion processing on the multi-modal perception feature and the first semantic vector to obtain a first fusion feature;

[0017] performing industrial environment anomaly judgment processing on the first fusion feature to obtain the industrial environment evaluation result.

[0018] In some embodiments, the inputting the industrial sensor data into an anomaly recognition model according to the industrial environment evaluation result to obtain an anomaly recognition result comprises:

[0019] performing anomaly analysis processing on the industrial environment evaluation result to obtain an anomaly type set;

[0020] mapping and selecting a plurality of industrial small models according to the anomaly type set to obtain an anomaly recognition model;

[0021] performing anomaly recognition processing on the industrial sensor data according to the anomaly recognition model to obtain the anomaly recognition result.

[0022] In some embodiments, the inputting the industrial sensor data and the anomaly recognition result into the multi-modal large model to perform strategy generation processing to obtain an interaction strategy comprises:

[0023] performing feature extraction processing on the industrial sensor data to obtain a multi-modal perception feature;

[0024] performing word embedding and position coding processing on the anomaly recognition result to obtain a second semantic vector;

[0025] performing feature fusion processing on the multi-modal perception feature and the second semantic vector according to an industrial cross-modal attention mechanism to obtain a second fusion feature;

[0026] performing strategy decoding processing on the second fusion feature to obtain the interaction strategy.

[0027] In some embodiments, the performing strategy decoding processing on the second fusion feature to obtain the interaction strategy comprises:

[0028] performing feature enhancement processing on the second fusion feature to obtain an enhanced feature;

[0029] performing strategy semantic mapping processing on the enhanced feature to obtain a strategy semantic vector;

[0030] performing action instruction generation processing on the strategy semantic vector to obtain the interaction strategy.

[0031] In some embodiments, the interaction control of the humanoid robot according to the interaction strategy comprises:

[0032] generating a motion trajectory and an interaction voice according to the interaction strategy;

[0033] controlling the movement of the humanoid robot according to the motion trajectory;

[0034] making voice alarm through the humanoid robot according to the interaction voice.

[0035] In some embodiments, the industrial environment abnormality judgment processing of the first fusion feature comprises:

[0036] performing evaluation space mapping processing on the first fusion feature to obtain an abnormal probability vector;

[0037] performing abnormality judgment processing on the abnormal probability vector according to a preset abnormal threshold matrix to obtain the industrial environment evaluation result.

[0038] To achieve the above-mentioned purposes, another aspect of the embodiment of the present application proposes a humanoid robot decision interaction system, which comprises:

[0039] a data acquisition module configured to acquire industrial sensor data;

[0040] an environment evaluation module configured to input the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result;

[0041] an abnormality recognition module configured to input the industrial sensor data into an abnormality recognition model according to the industrial environment evaluation result to obtain an abnormality recognition result;

[0042] a strategy generation module configured to input the industrial sensor data and the abnormality recognition result into the multi-modal large model for strategy generation processing to obtain an interaction strategy;

[0043] an interaction control module configured to control the interaction of a humanoid robot according to the interaction strategy.

[0044] To achieve the above-mentioned purposes, another aspect of the embodiment of the present application proposes an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the method described above when executing the computer program.

[0045] To achieve the above object, another aspect of the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method described above.

[0046] To achieve the above object, another aspect of the embodiment of the present application provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the method described above

[0047] The embodiment of the present application at least has the following beneficial effects: the present application provides a humanoid robot decision interaction method, system, device and medium, which inputs industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result, can introduce an industrial initial environment evaluation mechanism, trigger a small industrial model for deep processing only when an industrial abnormal situation is detected, and reduce unnecessary industrial computing overhead. The scheme also inputs the industrial sensor data into an abnormality recognition model according to the industrial environment evaluation result to obtain an abnormality recognition result. The industrial sensor data and the abnormality recognition result are input into the multi-modal large model for strategy generation processing to obtain an interaction strategy, which can use the multi-modal large model to fuse the output result of the small industrial model and the original industrial sensor data, generate a comprehensive industrial environment interaction strategy through industrial cross-modal feature alignment and industrial strategy decoding, and balance industrial real-time performance and industrial accuracy. The embodiment of the present application uses an abnormality recognition model to quickly extract industrial special features, a multi-modal large model to realize industrial cross-modal semantic fusion and industrial strategy reasoning, forms an industrial hierarchical decision mode, and improves the decision interaction efficiency of the humanoid robot. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 is a flowchart of a humanoid robot decision interaction method provided by the embodiment of the present application;

[0049] Figure 2 is an industrial interaction decision schematic diagram provided by the embodiment of the present application;

[0050] Figure 3 is a structural schematic diagram of a humanoid robot decision interaction system provided by the embodiment of the present application;

[0051] Figure 4 is a hardware structural schematic diagram of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION

[0052] For the purposes of the present application, the technical solutions and advantages thereof are more clearly apparent, the following will be further described in detail in conjunction with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application, and are not intended to limit the present application. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following example embodiments do not represent all implementations consistent with embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0053] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "when" or "in response to determining".

[0054] The terms "at least one", "multiple", "each", "any" and the like used in the present application include one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0056] In industrial manufacturing, industrial inspection and other industrial scenarios, humanoid robots are increasingly widely used. However, the industrial environment is complex and variable, and there are various industrial equipment, industrial workpieces, industrial operators, etc., which brings many challenges to the environmental interaction decision of humanoid robots.

[0057] Related industrial robots mostly use single industrial sensors or simple industrial sensor fusion perception models, which are difficult to handle the collaborative understanding of industrial multi-source heterogeneous data such as industrial vision, industrial hearing, industrial touch, etc. In complex industrial scenarios (such as industrial production line multi-station cooperation, industrial warehouse dynamic goods handling), industrial perception deviation is prone to occur. Although the end-to-end industrial model based on deep learning has certain industrial environment understanding ability, the model has large number of parameters and high computational complexity, which is difficult to meet the real-time industrial decision-making needs of industrial humanoid robots, especially when running on industrial low-power edge devices, which is inefficient.

[0058] In the related art, the industrial system lacks a rapid identification and hierarchical processing mechanism for industrial environment abnormal conditions. When encountering an untrained industrial emergency scenario (such as abnormal operation of industrial equipment or falling of industrial workpieces), it is difficult to generate a safe and effective industrial interaction strategy, and there is a risk of industrial production safety or industrial decision-making errors.

[0059] Therefore, in the embodiments of the present application, a humanoid robot decision interaction method, system, device and medium are provided. The scheme inputs industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result. An industrial initial environment evaluation mechanism can be introduced, and only when an industrial abnormal condition is detected, a small industrial model is triggered for deep processing to reduce unnecessary industrial computing overhead. The scheme further inputs the industrial sensor data into an abnormality recognition model according to the industrial environment evaluation result to obtain an abnormality recognition result. The industrial sensor data and the abnormality recognition result are input into the multi-modal large model for strategy generation processing to obtain an interaction strategy. The multi-modal large model can be used to fuse the output result of the small industrial model and the original industrial sensor data, and through industrial cross-modal feature alignment and industrial strategy decoding, a comprehensive industrial environment interaction strategy is generated, taking into account industrial real-time performance and industrial accuracy. The embodiments of the present application use an abnormality recognition model to quickly extract industrial special features, and a multi-modal large model to realize industrial cross-modal semantic fusion and industrial strategy reasoning, forming an industrial hierarchical decision-making mode, and improving the decision interaction efficiency of the humanoid robot.

[0060] The humanoid robot decision interaction method provided by the embodiments of the present application relates to the technical field of intelligent robots. The humanoid robot decision interaction method provided by the embodiments of the present application can be applied to a robot, can be applied to a server, and can also be software running in a robot or a server. In some embodiments, the server end can be configured as an independent physical server, can be configured as a server cluster or a distributed system composed of multiple physical servers, can be configured as a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform, and the server can also be a node server in a blockchain network; the software can be an application that implements a humanoid robot decision interaction method, and the like, but is not limited to the above forms.

[0061] The application is operable in a variety of general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that can be suitable for use with the application include personal computers, server computers, handheld or laptop devices, tablet devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like. The application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like, that perform particular tasks or implement particular abstract data types. The application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0062] Figure 1 is an optional flowchart of a humanoid robot decision interaction method provided by an embodiment of the application, Figure 1 The method in the above embodiment can include, but is not limited to, steps S101 to S105.

[0063] Step S101, obtaining industrial sensor data;

[0064] Step S102, inputting the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result;

[0065] Step S103, inputting the industrial sensor data into an anomaly recognition model according to the industrial environment evaluation result to obtain an anomaly recognition result;

[0066] Step S104, inputting the industrial sensor data and the anomaly recognition result into the multi-modal large model for strategy generation processing to obtain an interaction strategy;

[0067] Step S105, performing interaction control on a humanoid robot according to the interaction strategy.

[0068] The steps S101 to S105 shown in the embodiments of the present application can detect industrial abnormal conditions by acquiring industrial sensor data and inputting the industrial sensor data into a multi-modal large model for environment evaluation, trigger a small industrial model for deep processing according to the detected industrial abnormal conditions, use the small industrial model as an abnormality recognition model, input the industrial sensor data into the abnormality recognition model according to the industrial environment evaluation result to obtain an abnormality recognition result, perform parallel and lightweight processing on multi-sensor data by means of multiple industrial special small models such as an industrial target detection model and an industrial equipment abnormal voiceprint recognition model, quickly extract industrial basic features, and obtain the abnormality recognition result. Finally, the multi-modal large model fuses the abnormality recognition result output by the abnormality recognition model and the original industrial sensor data to make a decision, aligns industrial cross-modal features, decodes industrial strategies, generates a comprehensive industrial environment interaction strategy, and can balance industrial real-time performance and industrial accuracy.

[0069] In some embodiments, the industrial sensor data is input into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result, including:

[0070] Obtaining an industrial preset question text;

[0071] Inputting the industrial preset question text and the industrial sensor data into the multi-modal large model;

[0072] Performing feature extraction processing on the industrial sensor data to obtain multi-modal perception features;

[0073] Performing word segmentation and position encoding processing on the industrial preset question text to obtain a first semantic vector;

[0074] Performing cross-modal fusion processing on the multi-modal perception features and the first semantic vector to obtain a first fusion feature;

[0075] Performing industrial environment abnormality judgment processing on the first fusion feature to obtain the industrial environment evaluation result.

[0076] In the embodiment of the present application, the industrial preset question text is obtained, which can be obtained by pre-input or by text recognition and extraction of the question document. The text data can be "whether the industrial equipment has abnormal vibration?" or "whether the temperature of the industrial equipment is abnormal?" and the like. The industrial preset question text and the industrial sensor data are input into the multi-modal large model. The industrial multi-modal large model adopts a four-layer progressive architecture, which specifically includes: an industrial visual feature extraction layer based on a lightweight CNN, which includes 15 convolution modules (including depth separable convolution) and 3 pooling layers, for feature extraction of industrial visual images. An industrial text semantic understanding layer adopting a 6-layer Transformer encoder structure, including a self-attention mechanism and a feedforward neural network, for processing the industrial preset question text; a cross-modal interaction fusion layer composed of 4 cross-modal attention modules, supporting bidirectional alignment of visual features and text features; and an industrial strategy evaluation decoding layer including 2 fully connected layers and a softmax classifier, outputting the industrial environment evaluation result.

[0077] In a feasible embodiment, by obtaining image data in the industrial sensor data, the image data is standardized, and edge or texture features are extracted through a convolution module, and multi-modal perception features are obtained through dimension reduction processing by a pooling layer. The preset question text is processed by word segmentation, and a first semantic vector is obtained by word embedding and position encoding. Then, the multi-modal perception features and the first semantic vector are processed by cross-modal fusion, and the expression of the fusion calculation is as follows:

[0078] Attention(Q,K,V)=softmax((QK^T) / √d_k)V;

[0079] In the formula, Q is a query vector from the text semantic understanding layer, used to query related information in the visual features; K is a key vector from the visual feature extraction layer, used to represent the key information of the visual features; V is a value vector from the visual feature extraction layer, used to provide the specific content of the visual features; d_k represents the dimension of the key vector, used to scale the dot product to avoid gradient disappearance. Then, the first fusion feature is processed by industrial environment anomaly judgment to obtain the industrial environment evaluation result.

[0080] In some embodiments, the first fusion feature is processed by industrial environment anomaly judgment to obtain the industrial environment evaluation result, including:

[0081] The first fusion feature is processed by evaluation space mapping to obtain an anomaly probability vector;

[0082] The anomaly probability vector is processed by anomaly judgment according to a preset anomaly threshold matrix to obtain the industrial environment evaluation result.

[0083] In the embodiment of the present application, the first fusion feature is mapped to the evaluation space through the full connection layer to obtain an anomaly probability vector. Through a preset anomaly threshold matrix including temperature, vibration, voiceprint, and position anomaly thresholds, anomaly judgment processing is performed on the anomaly probability vector according to the preset anomaly threshold matrix. When the score of any type of anomaly exceeds the threshold, an anomaly is triggered, for example, an anomaly condition of "temperature anomaly accompanied by abnormal noise" is triggered, and an industrial environment evaluation result is obtained.

[0084] In some embodiments, the industrial sensor data is input into an anomaly recognition model according to the industrial environment evaluation result, and an anomaly recognition result is obtained, including:

[0085] Performing anomaly analysis processing on the industrial environment evaluation result to obtain an anomaly type set;

[0086] Mapping and selecting processing is performed on a plurality of industrial small models according to the anomaly type set to obtain an anomaly recognition model;

[0087] According to the anomaly recognition model, anomaly recognition processing is performed on the industrial sensor data to obtain the anomaly recognition result.

[0088] In the embodiment of the present application, anomaly analysis processing is performed on the industrial environment evaluation result, and an anomaly type set such as temperature, vibration, voiceprint, and position anomaly types can be obtained. The embodiment of the present application can pre-construct an anomaly type-model mapping table. According to different anomaly types, the corresponding industrial small models can be activated. For example, according to the temperature anomaly, an industrial component temperature perception model and an industrial target detection model can be mapped and selected. According to the vibration anomaly, an industrial action recognition model and an industrial equipment abnormal voiceprint recognition model can be mapped and selected. According to the object anomaly, an industrial target detection model and an industrial action recognition model can be mapped and selected. Finally, the selected model is used as an anomaly recognition model. According to the anomaly recognition model, anomaly recognition processing is performed on the industrial sensor data. For example, visual data is input into an industrial target detection small model to recognize the industrial object categories and positions such as "motor" and "conveyor belt". The detection process can be represented as: Y=f det (X vis ), wherein X vis is industrial visual data, f det is an industrial target detection model, and Y is a detection result. Industrial visual data is simultaneously input into an industrial action recognition small model to determine whether the running action of the motor is normal. The action recognition function is: A=f act (X vis ), wherein A is the action intention. Voiceprint data is input into an industrial equipment abnormal voiceprint recognition small model to analyze the motor abnormal voiceprint features. The voiceprint processing formula is: S=f 声纹 (X aud ), wherein Xaud For auditory data, S is a voiceprint feature. The infrared data is input into an industrial component temperature perception model to obtain temperature data of each part of the motor, and the temperature perception is represented as: T = f temp (X ir ), where X ir is industrial infrared data, T is temperature data, and an abnormality recognition result is obtained.

[0089] In some embodiments, the industrial sensor data and the abnormality recognition result are input into the multi-modal large model for strategy generation processing to obtain an interaction strategy, including:

[0090] The industrial sensor data is subjected to feature extraction processing to obtain multi-modal perception features;

[0091] The abnormality recognition result is subjected to word embedding and position encoding processing to obtain a second semantic vector;

[0092] The multi-modal perception features and the second semantic vector are subjected to feature fusion processing according to an industrial cross-modal attention mechanism to obtain a second fusion feature;

[0093] The second fusion feature is subjected to strategy decoding processing to obtain the interaction strategy.

[0094] In the embodiments of the present application, the industrial sensor data is subjected to feature extraction processing by a multi-modal large model, the original visual data is subjected to industrial visual feature extraction by a CNN network, the voiceprint data is subjected to industrial acoustic feature extraction by a speech feature extraction model, and the infrared data is subjected to industrial temperature feature extraction. The embodiments of the present application also generate an industrial semantic vector from the abnormality recognition result, i.e., the output result of the industrial small model (such as "industrial motor / temperature anomaly / abnormal voiceprint / position coordinates"), through text processing. The text processing process includes word embedding and position encoding. The industrial cross-modal attention mechanism is used to align the visual features, acoustic features, industrial temperature features, and industrial semantic vector, and the industrial control unit is used to fuse the industrial multi-modal features. According to the industrial cross-modal attention mechanism, the multi-modal perception features and the second semantic vector are subjected to feature fusion processing to obtain a second fusion feature, and then the second fusion feature is subjected to strategy decoding processing to obtain an interaction strategy. A feasible interaction strategy is: "move to the motor side, further detect the motor vibration condition through the tactile sensor, send a device anomaly warning to the central control system, and suggest shutdown for maintenance".

[0095] In some embodiments, the second fusion feature is subjected to strategy decoding processing to obtain the interaction strategy, including:

[0096] The second fusion feature is subjected to feature enhancement processing to obtain an enhanced feature;

[0097] The enhanced feature is subjected to policy semantic mapping processing to obtain a policy semantic vector.

[0098] The policy semantic vector is subjected to action instruction generation processing to obtain the interaction policy.

[0099] In the embodiment of the application, the policy evaluation decoding layer of the multi-modal large model adopts a three-layer progressive structure, including a feature enhancement layer, a policy semantic mapping layer and an action instruction generation layer. The feature enhancement layer includes a 2-layer fully connected network for improving the representation ability of the fused features. The policy semantic mapping layer is used to map the enhanced features to a policy semantic space, and the action instruction generation layer is used to generate specific action instructions according to the policy semantics.

[0100] In some embodiments, the interaction control of the humanoid robot according to the interaction policy comprises:

[0101] generating a motion trajectory and an interaction voice according to the interaction policy;

[0102] controlling the movement of the humanoid robot according to the motion trajectory;

[0103] performing voice alarm through the humanoid robot according to the interaction voice.

[0104] In the embodiment of the application, the interaction policy can include a motion trajectory, which can generate a path point sequence according to a trajectory generation network and perform speed planning to generate a motion instruction, and the motion trajectory is obtained according to the motion instruction, for example, a trajectory to avoid high-temperature equipment when the temperature is abnormal. Similarly, through semantic understanding, text generation and speech synthesis processing of the interaction policy, an interaction voice can be obtained, for example, a pre-warning voice: "detecting abnormal temperature of the motor, please check immediately".

[0105] Next, the scheme of the embodiment of the application will be described and explained in detail in combination with a specific application example:

[0106] The embodiment of the application can be applied to the field of intelligent robot technology and is suitable for industrial warehouse, industrial quality inspection and other industrial scenes. For example, in the industrial warehouse scene, an industrial goods detection module is added through an industrial small model, and an industrial goods handling strategy is generated by the industrial multi-modal large model in combination with an industrial warehouse map and industrial real-time sensor data. In the industrial quality inspection scene, an industrial component defect recognition module is added through an industrial small model, and an industrial quality inspection scheme is generated by the industrial multi-modal large model in combination with industrial component images, industrial quality inspection standards and industrial sensor data. Please refer to Figure 2The embodiment of the application inputs the acquired industrial multi-sensor data into multiple industrial special small models to obtain multiple small model output results; inputs the multiple small model output results and the industrial multi-sensor data into a multi-modal large model to generate an industrial environment interaction strategy, thereby realizing real-time and accurate interaction of the humanoid robot in an industrial complex environment, and improving the understanding ability and safety response speed of the robot to an industrial dynamic scene.

[0107] Please refer to Figure 3 The embodiment of the application further provides a humanoid robot decision interaction system, which can implement the above method. The system comprises:

[0108] The data acquisition module 301 is configured to acquire industrial sensor data.

[0109] The environment evaluation module 302 is configured to input the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result.

[0110] The anomaly recognition module 303 is configured to input the industrial sensor data into an anomaly recognition model according to the industrial environment evaluation result to obtain an anomaly recognition result.

[0111] The strategy generation module 304 is configured to input the industrial sensor data and the anomaly recognition result into the multi-modal large model for strategy generation processing to obtain an interaction strategy.

[0112] The interaction control module 305 is configured to perform interaction control on the humanoid robot according to the interaction strategy.

[0113] It can be understood that the contents in the above method embodiments are all applicable to the present system embodiment, the present system embodiment specifically implements the same functions as the above method embodiments, and achieves the same beneficial effects as the above method embodiments.

[0114] The embodiment of the application further provides an electronic device, which comprises a memory and a processor. The memory stores a computer program, and the processor implements the above method when executing the computer program. The electronic device can be any intelligent terminal, such as a tablet computer or a vehicle-mounted computer.

[0115] It can be understood that the contents in the above method embodiments are all applicable to the present device embodiment, the present device embodiment specifically implements the same functions as the above method embodiments, and achieves the same beneficial effects as the above method embodiments.

[0116] Please refer to Figure 4 , Figure 4 The hardware structure of the electronic device of another embodiment is illustrated, which comprises:

[0117] The processor 401 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided by the embodiments of the present application.

[0118] The memory 402 can be implemented by a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory), etc. The memory 402 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present application are implemented by software or firmware, relevant program codes are stored in the memory 402 and are called and executed by the processor 401 to implement the above-mentioned method of the embodiments of the present application.

[0119] The input / output interface 403 is configured to implement information input and output.

[0120] The communication interface 404 is configured to implement communication interaction between the device and other devices, and can realize communication through a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).

[0121] The bus 405 is configured to transmit information between various components (for example, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404) of the device.

[0122] The processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are connected to each other through the bus 405 to realize communication connection between the device.

[0123] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the above-mentioned method.

[0124] It can be understood that the above-mentioned method embodiments are applicable to the present storage medium embodiments, the present storage medium embodiments specifically implement the same functions as the above-mentioned method embodiments, and achieve the same beneficial effects as the above-mentioned method embodiments.

[0125] The embodiments of the present application further provide a computer program product, which includes a computer program. The computer program is executed by a processor to implement the above-mentioned method.

[0126] It can be understood that the contents in the above method embodiments are all applicable to the program product embodiments, the program product embodiments specifically implement the functions same as those of the above method embodiments, and achieve the same beneficial effects as those of the above method embodiments.

[0127] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory disposed remotely relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0128] The embodiment of the present application provides a humanoid robot decision interaction method, system, device and medium, which can introduce an industrial initial environment evaluation mechanism, trigger a small industrial model for deep processing only when an industrial abnormal situation is detected, and reduce unnecessary industrial computing overhead. The scheme further inputs the industrial sensor data into an abnormality recognition model according to the industrial environment evaluation result to obtain an abnormality recognition result. The industrial sensor data and the abnormality recognition result are input into a multi-modal large model for strategy generation processing to obtain an interaction strategy. The multi-modal large model can fuse the output result of the small industrial model and the original industrial sensor data, and generate a comprehensive industrial environment interaction strategy through industrial cross-modal feature alignment and industrial strategy decoding, which takes into account industrial real-time performance and industrial accuracy. The embodiment of the present application can quickly extract industrial special features through an abnormality recognition model, realize industrial cross-modal semantic fusion and industrial strategy reasoning through a multi-modal large model, form an industrial hierarchical decision mode, and improve the decision interaction efficiency of the humanoid robot.

[0129] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0130] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than those shown in the figures, or combine certain steps or different steps.

[0131] The system embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.

[0132] Those skilled in the art can understand that all or some steps in the method disclosed above, and the functional modules / units in the system, the device can be implemented as software, firmware, hardware and appropriate combinations thereof.

[0133] The terms "first", "second", "third", "fourth" and the like in the description of the application and in the claims of the foregoing drawings, if any, are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so

[0134] It should be understood that in this application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0135] In several embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other manners. For example, the system embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, and electrical, mechanical or other forms.

[0136] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0137] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0138] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program storage media.

[0139] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A humanoid robot decision interaction method, characterized by, The method comprises the following steps: acquiring industrial sensor data; inputting the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result; inputting the industrial sensor data into an anomaly recognition model according to the industrial environment evaluation result to obtain an anomaly recognition result; inputting the industrial sensor data and the anomaly recognition result into the multi-modal large model for strategy generation processing to obtain an interaction strategy; controlling a humanoid robot according to the interaction strategy.

2. The method of claim 1, wherein, The step of inputting the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result comprises the following steps: acquiring industrial preset question text; inputting the industrial preset question text and the industrial sensor data into the multi-modal large model; performing feature extraction processing on the industrial sensor data to obtain multi-modal perception features; performing word segmentation and position encoding processing on the industrial preset question text to obtain a first semantic vector; performing cross-modal fusion processing on the multi-modal perception features and the first semantic vector to obtain first fusion features; performing industrial environment anomaly judgment processing on the first fusion features to obtain the industrial environment evaluation result.

3. The method of claim 1, wherein, The step of inputting the industrial sensor data into an anomaly recognition model according to the industrial environment evaluation result to obtain an anomaly recognition result comprises the following steps: performing anomaly analysis processing on the industrial environment evaluation result to obtain an anomaly type set; performing mapping selection processing on a plurality of industrial small models according to the anomaly type set to obtain an anomaly recognition model; performing anomaly recognition processing on the industrial sensor data according to the anomaly recognition model to obtain the anomaly recognition result.

4. The method of claim 1, wherein, The step of inputting the industrial sensor data and the anomaly recognition result into the multi-modal large model for strategy generation processing to obtain an interaction strategy comprises the following steps: performing feature extraction processing on the industrial sensor data to obtain multi-modal perception features; performing word embedding and position encoding processing on the anomaly recognition result to obtain a second semantic vector; performing feature fusion processing on the multi-modal perception features and the second semantic vector according to an industrial cross-modal attention mechanism to obtain second fusion features; performing strategy decoding processing on the second fusion features to obtain the interaction strategy.

5. The method of claim 4, wherein, The step of performing strategy decoding processing on the second fusion features to obtain the interaction strategy comprises the following steps: performing feature enhancement processing on the second fusion features to obtain enhanced features; performing strategy semantic mapping processing on the enhanced features to obtain a strategy semantic vector; performing action instruction generation processing on the strategy semantic vector to obtain the interaction strategy.

6. The method according to any one of claims 1 to 5, characterized in that, The step of controlling a humanoid robot according to the interaction strategy comprises the following steps: generating a motion trajectory and an interaction voice according to the interaction strategy; controlling the humanoid robot to move according to the motion trajectory; performing voice alarm through the humanoid robot according to the interaction voice.

7. The method of claim 2, wherein, The step of performing industrial environment anomaly judgment processing on the first fusion features to obtain the industrial environment evaluation result comprises the following steps: performing evaluation space mapping processing on the first fusion features to obtain an anomaly probability vector; The abnormal probability vector is subjected to an abnormality judgment process according to a preset abnormal threshold matrix, and the industrial environment evaluation result is obtained.

8. A humanoid robot decision interaction system characterized by, The system comprises: a data acquisition module configured to acquire industrial sensor data; an environment evaluation module configured to input the industrial sensor data into a multi-modal large model for environment evaluation processing to obtain an industrial environment evaluation result; an anomaly identification module configured to input the industrial sensor data into an anomaly identification model according to the industrial environment evaluation result to obtain an anomaly identification result; a strategy generation module configured to input the industrial sensor data and the anomaly identification result into the multi-modal large model for strategy generation processing to obtain an interaction strategy; an interaction control module configured to control the humanoid robot according to the interaction strategy.

9. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sweeping robot, voice interaction method and device thereof and storage medium

    CN118924197A

  • Visual interaction method and system for robot with body, terminal and medium

    CN119036461A

  • Abnormal data processing scheme determination method and device, equipment, storage medium and program product

    CN119668920A

  • Equipment fault early warning method and system based on AI algorithm, and storage medium

    CN119885046A

  • Unmanned agricultural production system based on agricultural scene and inspection robot control method

    CN119991337A

Cited By

  • Environment interaction method, system, equipment and product of humanoid robot

    CN121809538A