Man-machine interaction method and man-machine interaction system based on large model assistance

By adopting large-model-assisted human-computer interaction methods in human-computer interaction, we build target speech recognition models, intention fusion recognition and task decomposition, and solve the problem of low efficiency of traditional human-computer interaction and achieve efficient and accurate human-computer interaction operations.

CN120199245AActive Publication Date: 2025-06-24启元实验室

Patent Information

Application Number
CN202510670641.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The traditional human-computer interaction process relies on complex operation interfaces and cumbersome command input, resulting in high difficulty, low flexibility and accuracy of user operations, thereby reducing the testing efficiency of human-computer interaction.

Method used

Using a human-computer interaction method based on large-model assistance, machine control instructions are generated to achieve efficient human-computer interaction by building a target speech recognition model, intent fusion recognition, task decomposition and planning.

Benefits of technology

The testing efficiency of human-computer interaction is improved, and by accurately identifying user voice commands and achieving dynamic adaptation of general semantic understanding and vertical domain knowledge, the accuracy of instruction analysis and operation accuracy are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199245A_ABST
    Figure CN120199245A_ABST
Patent Text Reader

Abstract

The invention provides a man-machine interaction method and a man-machine interaction system based on large model assistance, and relates to the technical field of man-machine interaction. The man-machine interaction method comprises the following steps: constructing a target speech recognition model based on an initial instruction set; recognizing a user voice instruction from a user based on the target voice recognition model, so as to determine a user text instruction according to the user voice instruction; based on the user text instruction, intention fusion recognition is carried out on the user text instruction through a language large model and a preset language small model to obtain a user instruction intention; performing task decomposition and planning on the user instruction intention to obtain a high-level task sequence; and generating a machine control instruction according to the high-level task sequence, and executing man-machine interaction according to the machine control instruction. According to the application, the accuracy of user instruction analysis can be improved, and the accuracy of man-machine interaction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of human - computer interaction. Specifically, it relates to a human - computer interaction method and a human - computer interaction system assisted by a large model. Background Art

[0002] Human - computer interaction refers to the technology of realizing the interaction between humans and computers in an effective way through computer input or output devices. Traditional human - computer interaction experiments (also known as human - machine hybrid experiments) mainly rely on the real - time manual operations of operators. This method is not only redundant and cumbersome but also has high requirements for the professional skills and experience of operators.

[0003] Currently, human - computer interaction technology can obtain natural - language text instructions issued by users through voice wake - up, collection, and accurate speech recognition technology, and then use natural - language understanding technology to convert these instructions into machine - executable action control signals, thus realizing human - computer interaction.

[0004] However, the traditional human - computer interaction process relies on complex operation interfaces and cumbersome instruction inputs, which not only increase the operation difficulty of users but also limit the flexibility and accuracy of human - computer interaction, resulting in a reduction in the test efficiency of human - computer interaction.

[0005] The human - computer interaction technology assisted by a large model has attracted much attention due to its powerful reasoning ability and natural - language understanding technology. Although the human - machine hybrid experiment technology assisted by a large model has significant advantages in theory, a series of key technical problems still need to be overcome in practical applications.

[0006] The content in the background art section is only the technology known to the applicant and does not necessarily represent the prior art in this field. Summary of the Invention

[0007] This application provides a human - computer interaction method and a human - computer interaction system assisted by a large model to solve the problem of low human - computer interaction efficiency.

[0008] According to one aspect of this application, a human - computer interaction method assisted by a large model is provided, including: constructing a target speech recognition model based on an initial instruction set; recognizing a user speech instruction from a user based on the target speech recognition model to determine a user text instruction according to the user speech instruction; performing intention fusion recognition on the user text instruction through a language large model and a preset language small model based on the user text instruction to obtain a user instruction intention; decomposing and planning the user instruction intention to obtain a high - level task sequence; generating a machine control instruction according to the high - level task sequence to perform human - computer interaction according to the machine control instruction.

[0009] According to some embodiments of the present application, constructing a target speech recognition model based on an initial instruction set includes: receiving an interaction instruction set from the field of human-computer interaction to construct an initial instruction set; performing semantic expansion processing on the initial instruction set based on a large language model to generate a synonymous instruction set; performing speech synthesis recording on the initial instruction set and the synonymous instruction set to obtain a training speech instruction data set; and fine-tuning a preset large speech recognition model based on the training speech instruction data set through a preset model fine-tuning technique to obtain a target speech recognition model.

[0010] According to some embodiments of the present application, performing intention fusion recognition on a user text instruction based on a large language model and a preset small language model to obtain the user's instruction intention includes: constructing a preset small language model based on a sample set including a specific task; obtaining first intention information based on the user text instruction according to the large language model; obtaining second intention information based on the user text instruction according to the preset small language model; and performing intention fusion recognition on the first intention information and the second intention information to obtain the user instruction intention.

[0011] According to some embodiments of the present application, constructing a preset small language model based on a sample set including a specific task includes: constructing a sample set including a specific task; and training a text detection small model of machine learning based on the sample set to obtain a preset small language model.

[0012] According to some embodiments of the present application, obtaining first intention information based on a user text instruction according to a large language model includes: constructing a prompt template based on the user text instruction and preset interaction task information; generating a guiding prompt word based on the prompt template; inputting the guiding prompt word and the user text instruction into the large language model for semantic understanding to obtain key elements in the user text instruction; and performing semantic alignment and intention recognition on the guiding prompt word and the key elements to obtain the first intention information.

[0013] According to some embodiments of the present application, obtaining second intention information based on a user text instruction according to a preset small language model includes: processing the user text instruction to obtain intention data information; and processing the intention data information based on the preset small language model to obtain the second intention information.

[0014] According to some embodiments of the present application, performing intention fusion recognition on the first intention information and the second intention information to obtain the user instruction intention includes: determining whether the first intention information and the second intention information are consistent; if so, outputting the user instruction intention; if not, supplementing and correcting the user intention through a multi-source information integration mechanism based on the first intention information and the second intention information to obtain the user instruction intention.

[0015] According to some embodiments of the present application, the user instruction intention is decomposed and planned to obtain a high-level task sequence, including: decomposing the user instruction intention to obtain at least one candidate action information; predicting the relevant probability of the candidate action information for realizing the high-level instruction through a language large model to obtain a prediction probability; determining a candidate subtask with guiding significance based on the prediction probability; estimating the probability of the execution success rate of the candidate subtask according to the current interaction environment state data and historical data to obtain the feasibility score of the candidate subtask under the current conditions; determining the subtask planning information based on the prediction probability and the feasibility score; and determining the high-level task sequence based on the subtask planning information.

[0016] According to some embodiments of the present application, a machine control instruction is generated according to the high-level task sequence, and the human-computer interaction is executed according to the machine control instruction, including: determining the subtask execution parameters based on the subtask planning information; establishing a mapping relationship between the application programming interface and the subtask execution parameters to obtain the machine control instruction; executing the human-computer interaction based on the machine control instruction; and the subtask execution parameters may at least include target parameters, input parameters, output parameters, and execution constraint conditions.

[0017] According to another aspect of the present application, the present application provides a human-computer interaction system assisted by a large model, including a speech recognition module, an intention understanding module, a task planning module, and an instruction execution module. The speech recognition module constructs a target speech recognition model based on an initial instruction set, and recognizes a user speech instruction from the user based on the target speech recognition model to determine a user text instruction according to the user speech instruction; the intention understanding module performs intention fusion recognition on the user text instruction through a language large model and a preset language small model based on the user text instruction to obtain a user instruction intention; the task planning module decomposes and plans the user instruction intention to obtain a high-level task sequence; the instruction execution module generates a machine control instruction according to the high-level task sequence and executes the human-computer interaction according to the machine control instruction.

[0018] According to another aspect of the present application, the present application further provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the human-computer interaction method as described above.

[0019] According to another aspect of the present application, the present application further provides a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the human-computer interaction method as described above can be implemented.

[0020] According to another aspect of the present application, the present application further provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, when the program instructions are executed by a computer, causing the computer to execute the human-computer interaction method as described above.

[0021] Beneficial technical effects: The present application constructs a target speech recognition model through an initial instruction set, enabling the target speech recognition model to simulate various simulated speech scenarios in a human-computer interaction experiment, so that the present application can more accurately recognize user speech instructions to determine user text instructions. The present application performs intention fusion recognition on user text instructions through a large language model and a preset small language model, and can achieve dynamic adaptation of general semantic understanding and vertical domain knowledge, so that the present application can accurately obtain the user instruction intention. Then, the present application decomposes and plans the user instruction intention to determine the execution steps with the highest execution rate, and thus combines them to obtain a high-level task sequence. Finally, the present application generates machine control instructions through the high-level task sequence to perform human-computer interaction, so that the present application can accurately execute human-computer interaction operations, thereby improving the efficiency of human-computer interaction testing. Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0023] Figure 1 A flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 2 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 3 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 4 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 5 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 6 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 7 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 8 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 9 Another flowchart showing the human - computer interaction method according to an embodiment of the present application; Figure 10 A schematic structural diagram showing the human - computer interaction system according to an embodiment of the present application.

[0024] Explanation of reference numerals: 10. Human - computer interaction system; 11. Speech recognition module; 12. Intention understanding module; 13. Task planning module; 14. Instruction execution module. Detailed implementation manners

[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Identical reference numerals in the figures denote identical or similar parts, and thus their repetitive description will be omitted.

[0026] The features, structures, or characteristics described may be combined in one or more embodiments in any suitable manner. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of these specific details, or can be implemented in other ways, components, materials, devices, etc. In these cases, well - known structures, methods, devices, implementations, materials, or operations will not be shown or described in detail.

[0027] In addition, the terms "including" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0028] The terms "first", "second", etc. in the description and claims of the present application and the above - mentioned drawings are used to distinguish different objects, rather than to describe a specific order.

[0029] The technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0030] According to one aspect of the present application, the present application provides a human-computer interaction method assisted by a large model. Figure 1 FIG. 1 shows a schematic flow diagram of the human-computer interaction method according to an embodiment of the present application. As Figure 1 shown, the human-computer interaction method may include steps S100-S500.

[0031] Exemplarily, the human-computer interaction method may be executed by a human-computer interaction system with computing capabilities.

[0032] According to an exemplary embodiment, in step S100, the human-computer interaction system constructs a target speech recognition model based on an initial instruction set.

[0033] For example, the human-computer interaction system receives an interaction instruction set from the field of human-computer interaction and constructs an initial instruction set. The human-computer interaction system performs semantic expansion processing on the initial instruction set through a language large model to generate a synonymous instruction set with diverse expressions. The human-computer interaction system performs speech synthesis recording on the initial instruction set and the synonymous instruction set to obtain a training speech instruction data set. The human-computer interaction system is based on the training speech instruction data set and fine-tunes a preset speech recognition large model through a preset model fine-tuning technique to obtain a target speech recognition model.

[0034] Exemplarily, the interaction instruction set may be an instruction set composed of operation terms, such as "start arc welding", etc. The human-computer interaction system constructs an initial instruction set according to different operation terms.

[0035] Exemplarily, the human-computer interaction system performs semantic expansion processing on "start arc welding" through a language large model to generate a synonymous instruction set with diverse expressions, such as "start welding operation", "activate electrode welding", etc.

[0036] The human-computer interaction system can perform speech synthesis recording on the speech data of industrial noise and interaction instructions such as "start arc welding", "start welding operation", "activate electrode welding", etc. to obtain a training speech instruction data set. For example, according to the operation term set containing industrial noise, a training speech instruction data set is obtained.

[0037] The human-computer interaction system inputs the training speech instruction data set into a preset speech recognition large model and fine-tunes the preset speech recognition large model through LoRA (Low-Rank Adaptation, a model fine-tuning technique). It can reduce computational complexity and memory requirements to obtain an efficient target speech recognition model.

[0038] According to an exemplary embodiment, in step S200, the human-computer interaction system recognizes a user speech instruction from the user based on the target speech recognition model to determine a user text instruction according to the user speech instruction.

[0039] For example, the user voice command may include, but is not limited to, an operation command (such as starting a welding program), a multimodal composite command (such as grasping this part), or a complex task command (such as alarming if the pressure exceeds 50 MPa), etc.

[0040] The human-computer interaction system performs speech-to-text processing on the user voice command through the target speech recognition model to determine the corresponding user text command.

[0041] Exemplarily, the human-computer interaction system receives a multimodal composite command, such as "go to area A to get a flange part". The human-computer interaction system performs speech-to-text processing on the multimodal composite command through the target speech recognition model, and fine-tunes the multimodal composite command through the target speech recognition model to obtain a user text command, such as "go to area A to grasp the flange".

[0042] According to the example embodiment, in step S300, the human-computer interaction system performs intention fusion recognition on the user text command through the language large model and the preset language small model based on the user text command to obtain the user command intention.

[0043] For example, the language large model can process the text through natural language processing technology.

[0044] For example, the human-computer interaction system constructs a prompt template through the user text command and the preset interaction task information, and then generates a guiding prompt word through the prompt template. The human-computer interaction system inputs the guiding prompt word and the user text command into the language large model for semantic understanding to obtain the key elements in the user text command. Then the human-computer interaction system performs semantic alignment and intention recognition on the guiding prompt word and the key elements to obtain the first intention information.

[0045] The human-computer interaction system constructs a sample set containing specific tasks, and trains a text detection small model of machine learning based on the sample set to obtain a preset language small model.

[0046] The human-computer interaction system processes the user text command to obtain intention data information. The human-computer interaction system processes the intention data information based on the preset language small model to obtain the second intention information.

[0047] The human-computer interaction system performs intention fusion recognition on the first intention information and the second intention information to obtain the user command intention.

[0048] According to the example embodiment, in step S400, the human-computer interaction system performs task decomposition and planning on the user command intention to obtain a high-level task sequence.

[0049] For example, the human-computer interaction system decomposes the user instruction intention to obtain at least one candidate action information. The human-computer interaction system predicts the relevant probability of the candidate action information for implementing the high-level instruction through a language large model, and obtains the prediction probability of the candidate action information. The human-computer interaction system determines a candidate subtask with guiding significance based on the prediction probability. The human-computer interaction system estimates the probability of the execution success rate of the candidate subtask according to the current interaction environment state data and historical data, and obtains the feasibility score of the candidate subtask under the current conditions. The human-computer interaction system determines the subtask planning information based on the prediction probability and the feasibility score. The human-computer interaction system determines the high-level task sequence based on the subtask planning information.

[0050] For example, the high-level task sequence can be a multi-level task structure driven by a goal, and a systematic execution path for achieving the goal is realized through step-by-step decomposition and dynamic adjustment.

[0051] According to the exemplary embodiment, in step S500, the human-computer interaction system generates a machine control instruction according to the high-level task sequence to perform human-computer interaction according to the machine control instruction.

[0052] For example, the human-computer interaction system determines the subtask execution parameters based on the subtask planning information. The subtask execution parameters can at least include target parameters, input parameters, output parameters, and execution constraint conditions. The human-computer interaction system establishes a mapping relationship between the application programming interface and the subtask execution parameters to obtain the machine control instruction. Then the human-computer interaction system performs human-computer interaction based on the machine control instruction.

[0053] Through the above embodiments, the present application can construct a target speech recognition model through an initial instruction set, so that the target speech recognition model can simulate various simulated speech scenarios in the human-computer interaction experiment, enabling the present application to more accurately recognize the user speech instruction to determine the user text instruction. The present application performs intention fusion recognition on the user text instruction through a language large model and a preset language small model, and can achieve the dynamic adaptation of general semantic understanding and vertical domain knowledge, enabling the present application to accurately obtain the user instruction intention. Then the present application decomposes and plans the user instruction intention to determine the execution step with the highest execution rate, and thus combines to obtain the high-level task sequence. Finally, the present application generates a machine control instruction through the high-level task sequence to perform human-computer interaction, enabling the present application to accurately perform the human-computer interaction operation, thereby improving the human-computer interaction test efficiency.

[0054] Figure 2 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 3 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 4 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 5Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 6 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 7 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 8 Another flowchart showing the human-computer interaction method according to an embodiment of the present application; Figure 9 Another flowchart showing the human-computer interaction method according to an embodiment of the present application.

[0055] Optionally, as Figure 2 shown, step S100 may further include steps S110 - S140.

[0056] In step S110, the human-computer interaction system receives an interaction instruction set from the field of human-computer interaction to construct an initial instruction set.

[0057] For example, the interaction instruction set from the field of human-computer interaction may include, but is not limited to, device operation instructions or parameter adjustment instructions, etc. The human-computer interaction system constructs an initial instruction set by aggregating different device operation instructions or parameter adjustment instructions.

[0058] Exemplarily, the interaction instruction set may be a device operation instruction. As an embodiment, the device operation instruction may be "Go to area A to grab a flange". As an embodiment, the device operation instruction may be "Start the flange grabbing program".

[0059] The interaction instruction set may be a parameter adjustment instruction. As an embodiment, the parameter adjustment instruction may be "Adjust the moving speed of the robotic arm to 0.3 m / s".

[0060] In step S120, the human-computer interaction system performs semantic extension processing on the initial instruction set based on a language large model to generate a synonymous instruction set.

[0061] For example, the semantic extension processing may include, but is not limited to, synonym generation, professional term conversion, or dialect adaptation, etc.

[0062] Exemplarily, in the case where the semantic extension processing is synonym generation, the initial instruction is "Go to area A to pick up a flange part", and after synonym generation by the human-computer interaction system, synonymous instructions such as "Slide to area A to pick up a flange part" or "Transmit to area A to pick up a flange part" can be obtained.

[0063] In the case where the semantic extension processing is professional term conversion, the initial instruction is "Go to area A to pick up a flange part", and after professional term conversion by the human-computer interaction system, "Activate the pickup protocol for area A" can be obtained.

[0064] In the case where semantic expansion processing is for dialect adaptation, the initial instruction is "Go to area A to pick up the flange part". After dialect adaptation by the human-computer interaction system, instructions such as "Go to area A to pick up the flange part" or "Move to area A to pick up the flange part" can be obtained.

[0065] In step S130, the human-computer interaction system performs voice synthesis recording on the initial instruction set and the synonymous instruction set to obtain a training voice instruction data set.

[0066] For example, the voice synthesis recording process can include but is not limited to multi-timbre generation, environment adaptation processing, or stress emphasis processing, etc. The human-computer interaction system generates a training voice instruction data set through the initial instruction set and the synonymous instruction set after synthesis processing.

[0067] Exemplarily, the initial instruction is "Emergency stop the robotic arm in area A", and the synonymous instructions are "Immediately shut down the robot in section A" and "Activate the emergency stop protocol in area A".

[0068] In the case where the voice synthesis recording process is for multi-timbre generation, the human-computer interaction system can generate 8 timbres through WaveNet (WaveNet), including male and female voices and different age groups.

[0069] In the case where the voice synthesis recording process is for environment adaptation processing, as an embodiment, the human-computer interaction system simulates the factory environment by superimposing 85dB white noise. As an embodiment, the human-computer interaction system simulates the factory environment by adding mechanical impact sounds.

[0070] In the case where the voice synthesis recording process is for stress emphasis processing, the human-computer interaction system increases the volume of the instruction keyword "stop" by 6dB.

[0071] In step S140, the human-computer interaction system fine-tunes a preset large speech recognition model based on the training voice instruction data set through a preset model fine-tuning technique to obtain a target speech recognition model.

[0072] Exemplarily, the preset large speech recognition model can be the Whisper model, and the preset model fine-tuning technique can be the LoRA technique. The human-computer interaction system fine-tunes the preset large speech recognition model through the LoRA technique. The human-computer interaction system introduces low-rank matrices inside each layer of the model, which can effectively reduce the number of model parameters and resource requirements while maintaining the performance of the model, enabling the model to recognize more accurately.

[0073] Through the above embodiments, the present application can ensure a high degree of matching between training data and industrial scenarios by receiving real instructions in the human-computer interaction scenario. Subsequently, through semantic expansion, the present application can improve the recognition accuracy of the model for colloquial expressions and dialect variants. By performing multi-environment simulation recordings on the expanded instruction set, the present application can further improve the recognition ability of the model. By fine-tuning the model, the present application can improve the response speed of the target speech recognition model to speech instructions, thereby improving the test efficiency of human-computer interaction.

[0074] Optionally, as Figure 3 shown, step S300 may further include steps S310-S340.

[0075] In step S310, the human-computer interaction system constructs a preset language small model based on a sample set containing specific tasks.

[0076] For example, the human-computer interaction system constructs a sample set containing specific tasks, and then the human-computer interaction system trains a text detection small model of machine learning based on the sample set to obtain a preset language small model.

[0077] The sample set containing specific tasks may include data composition and adversarial training samples. The text detection small model may have a factory-specific vocabulary, a noise suppression processing module, and a hybrid preference optimization.

[0078] Exemplarily, the data composition may be a voice instruction based on factory goods picking, and the voice instruction may include a noise environment (such as mechanical noise) and a dialect mixture sample (such as Mandarin mixed with Cantonese and Minnan dialect). The adversarial training sample may be voice data with synthetic mixed noise, such as generating an adversarial sample with a dynamic signal-to-noise ratio by mixing mechanical arm mechanical noise and background human voice interference.

[0079] The factory-specific vocabulary may include "robot arm coding", "PLC control instruction", etc. The noise suppression processing module may be noise suppression, and the noise suppression filters the characteristics of mechanical roars. The human-computer interaction system identifies the instruction keywords in the voice instruction through the factory-specific vocabulary and the noise suppression processing module.

[0080] The hybrid preference optimization may adopt a negative supervision correction algorithm to reduce misrecognition caused by dialect pronunciation, such as the acoustic feature confusion between "emergency stop" and "machine stop".

[0081] In step S320, the human-computer interaction system obtains the first intention information based on the user text instruction according to the large language model.

[0082] For example, the human-computer interaction system constructs a prompt template based on the user's text instruction and preset interaction task information. The human-computer interaction system generates guiding prompt words based on the prompt template. The human-computer interaction system inputs the guiding prompt words and the user's text instruction into a large language model for semantic understanding to obtain the key elements in the user's text instruction. The human-computer interaction system performs semantic alignment and intention recognition on the guiding prompt words and the key elements to obtain the first intention information.

[0083] The preset interaction task information may include, but is not limited to, device control instructions or safety warning processing, etc. The generation of guiding prompt words can be dynamic parameter filling. Semantic understanding can include identifying and processing the polysemy and ambiguity of vocabulary and understanding the meaning of the text. Key element extraction can include entity recognition and relationship extraction.

[0084] Semantic alignment and intention recognition may include, but are not limited to, multi-dimensional verification and intention decision trees. Multi-dimensional verification may include, but is not limited to, integrity checks and conflict detection, etc.

[0085] Exemplarily, the user's text instruction can be to go to area A to grab a flange. The device control instruction can be to identify the robotic arm operation instruction and map it to the PLC (Programmable Logic Controller) control protocol. Safety warning processing can be to detect abnormal instructions and trigger an emergency response mechanism, where the abnormal instructions are such as "Emergency stop grabbing" and "High temperature alarm".

[0086] Dynamic parameter filling can include basic templates or scenario adaptation. For example, the basic template can be to please specify the device number to be operated (such as Area A / Number B) and the specific instruction (start / speed adjustment / emergency stop). For example, scenario adaptation can be that if the user input contains noise (such as "Why doesn't the robotic arm in Area B move?"), an enhanced prompt "Please confirm the device number of the robotic arm on Line B and its current operating status" is generated.

[0087] Entity recognition can include device number ("Area A"), operation type ("pause"), target parameter ("0m / s"). Relationship extraction can be to establish a triple of "Area A - Speed_Adjust - 0m / s" and map it to the PLC control instruction.

[0088] Integrity check can be to verify whether the combination of "device number + operation type" is included. If it is missing, a follow-up question is triggered: "Please specify the specific device area". Conflict detection can be that when the instruction parameter exceeds the limit (such as "Adjust the speed to 5m / s" exceeds the device maximum limit of 3.5m / s), it is marked as a "red risk instruction".

[0089] In step S330, the human-computer interaction system obtains the second intention information based on the user's text instruction according to the preset small language model.

[0090] For example, a human-machine interaction system processes user text instructions to obtain intent data information. The human-machine interaction system processes the intent data information based on a preset language model to obtain second intent information.

[0091] Intent data processing may include but is not limited to preprocessing in a noisy environment and extracting intent elements, etc. Preprocessing in a noisy environment may include but is not limited to dialect normalization and noise suppression, etc.

[0092] Exemplarily, dialect normalization may be converting the Cantonese instruction "Why doesn't point B move?" to the standard instruction "The robotic arm of line B has an abnormal shutdown". Noise suppression may be filtering out the mechanical rumbling sound of the robotic arm using a dual-path attention mechanism.

[0093] The human-machine interaction system extracts intent elements from "The robotic arm of line B has an abnormal shutdown" to obtain "Check the equipment number and current operating status of the robotic arm on line B". The human-machine interaction system performs semantic integrity screening on the instruction "Check the equipment number and current operating status of the robotic arm on line B" and executes a full-dimensional verification.

[0094] In step S340, the human-machine interaction system performs intent fusion recognition on the first intent information and the second intent information to obtain the user instruction intent.

[0095] For example, the human-machine interaction system determines whether the first intent information and the second intent information are consistent. If so, the human-machine interaction system outputs the user instruction intent; if not, the human-machine interaction system supplements and corrects the user intent through a multi-source information integration mechanism based on the first intent information and the second intent information to obtain the user instruction intent.

[0096] The multi-source information integration mechanism may be real-time data fusion.

[0097] Exemplarily, the first intent information is "Please confirm the equipment number and current operating status of the robotic arm on line B", and the second intent information is "Check the equipment number and current operating status of the robotic arm on line B". The first intent information and the second intent information are consistent, and the human-machine interaction system outputs the user instruction intent.

[0098] The first intent information is "Please restart the equipment number and current operating status of the robotic arm on line B", and the second intent information is "Check the equipment number and current operating status of the robotic arm on line B". The first intent information and the second intent information are inconsistent. The human-machine interaction system reads the robotic arm status code through real-time data fusion. The human-machine interaction system intercepts the "restart" instruction in the first intent information and extracts the "check" instruction in the second intent information. The human-machine interaction system fuses the real-time data of the equipment to generate an enhanced instruction and finally outputs "Check the equipment number and current operating status of the robotic arm on line B".

[0099] Through the above embodiments, the present application can parse the deep semantics of user instructions from the massive pre-training data of the language model, and support the context-related reasoning of complex instructions. The present application can construct a preset small language model based on a specific task sample set to achieve the domain focus of the small model and reduce the semantic divergence risk of the large model in professional scenarios. The present application can obtain the accurate user instruction intention through intention fusion recognition, thereby improving the accuracy of instruction analysis and further improving the test efficiency of human-computer interaction.

[0100] Optionally, as Figure 4 shown, step S310 may further include steps S311 - S312.

[0101] In step S311, the human-computer interaction system constructs a sample set containing specific tasks.

[0102] For example, the sample set containing specific tasks may include data composition and adversarial training samples.

[0103] Exemplarily, the data composition may be voice instructions based on factory goods clamping, and the voice instructions may include noise environments (such as mechanical noise) and dialect mixed samples (such as Mandarin mixed with Cantonese and Minnan dialect). The adversarial training samples may be voice data synthesized with mixed noise, such as generating adversarial samples with dynamic signal-to-noise ratios by mixing the mechanical noise of the robotic arm and background human voice interference.

[0104] In step S312, the human-computer interaction system processes the intention data information based on the preset small language model to obtain the second intention information.

[0105] For example, the preset small language model may be a text detection small model, and the text detection small model may be built-in with a factory-specific vocabulary, a noise reduction processing module, and a mixed preference optimization.

[0106] Exemplarily, the factory-specific vocabulary may include "robotic arm coding", "PLC control instructions", etc. The noise reduction processing module may be noise suppression, and the noise suppression filters the mechanical roar characteristics. The human-computer interaction system identifies the instruction keywords in the voice instructions through the factory-specific vocabulary and the noise reduction processing module.

[0107] The mixed preference optimization may adopt a negative supervision correction algorithm to reduce misrecognition caused by dialect pronunciation, such as the acoustic feature confusion between "emergency stop" and "machine stop".

[0108] Through the above embodiments, the present application can cover typical factory interference scenarios in the training stage of the small model by constructing a sample set mixed with mechanical noise and dialects. The present application can further improve the accuracy of instruction analysis through the preset small language model, thereby improving the test efficiency of human-computer interaction.

[0109] Optionally, asFigure 5 As shown, step S320 may further include steps S321 - S324.

[0110] In step S321, the human - machine interaction system constructs a prompt template based on the user text instruction and the preset interaction task information.

[0111] For example, the preset interaction task information may include, but is not limited to, device control instructions or safety warning processing, etc.

[0112] Exemplarily, the user text instruction may be to go to area A to grab a flange. The device control instruction may be to identify the robotic arm operation instruction and map it to the PLC (Programmable Logic Controller) control protocol. The safety warning processing may be to detect abnormal instructions and trigger an emergency response mechanism, where the abnormal instructions are such as "Emergency stop grabbing", "High - temperature alarm".

[0113] In step S322, the human - machine interaction system generates guiding prompt words based on the prompt template.

[0114] For example, the generation of guiding prompt words can be dynamic parameter filling.

[0115] Exemplarily, the dynamic parameter filling may include a basic template or scenario adaptation. For example, the basic template may be to please specify the device number to be operated (such as Area A / Number B) and the specific instruction (start / speed adjustment / emergency stop). For example, the scenario adaptation may be that if the user input contains noise (such as "Why doesn't the robotic arm in Area B move?"), an enhanced prompt "Please confirm the device number of the robotic arm in Line B and its current operating status" is generated.

[0116] In step S323, the human - machine interaction system inputs the guiding prompt words and the user text instruction into the language large - model for semantic understanding to obtain the key elements in the user text instruction.

[0117] For example, semantic understanding may include identifying and processing the polysemy and ambiguity of words and understanding the meaning of the text. Key element extraction may include entity recognition and relationship extraction.

[0118] Exemplarily, entity recognition may include the device number ("Area A"), the operation type ("pause"), and the target parameter ("0m / s"). Relationship extraction may be to establish a triple "Area A - Speed_Adjust - 0m / s" and map it to the PLC control instruction.

[0119] In step S324, the human - machine interaction system performs semantic alignment and intention recognition on the guiding prompt words and the key elements to obtain the first intention information.

[0120] For example, semantic alignment and intent recognition may include, but are not limited to, multi-dimensional verification and intent decision trees. Multi-dimensional verification may include, but is not limited to, integrity checks and conflict detection, etc.

[0121] Exemplarily, the integrity check may be to verify whether the combination of "equipment number + operation type" is included. If it is missing, a follow-up question is triggered: "Please specify the specific equipment area." Conflict detection may be when the instruction parameters exceed the limit (such as "adjust the speed to 5 m / s" exceeding the device maximum limit of 3.5 m / s), it is marked as a "red risk instruction".

[0122] Through the above embodiments, the present application can construct a prompt template by presetting interaction task information, enabling the language large model to focus on domain key elements. The present application can also trigger the domain knowledge activation mechanism of the large model through guiding prompt words, improving the semantic understanding accuracy of the language large model. The present application can obtain more accurate first intent information through semantic alignment and intent recognition, thereby improving the test efficiency of human-computer interaction.

[0123] Optionally, as Figure 6 shown, step S330 may further include steps S331 - S332.

[0124] In step S331, the human-computer interaction system processes the user text instruction to obtain intent data information.

[0125] For example, intent data processing may include, but is not limited to, preprocessing in a noisy environment and extracting intent elements, etc. Preprocessing in a noisy environment may include, but is not limited to, dialect normalization and noise suppression, etc.

[0126] Exemplarily, dialect normalization may be to convert the Cantonese instruction "Why doesn't B move?" to the standard instruction "The robotic arm of B line has abnormal shutdown". Noise suppression may be to filter the mechanical roar of the robotic arm using a dual-path attention mechanism.

[0127] In step S332, the human-computer interaction system processes the intent data information based on a preset language small model to obtain second intent information.

[0128] Exemplarily, the human-computer interaction system extracts intent elements from "The robotic arm of B line has abnormal shutdown" to obtain "Check the equipment number and current operating status of the robotic arm on B line". The human-computer interaction system performs semantic integrity screening on the instruction "Check the equipment number and current operating status of the robotic arm on B line" and executes full-dimensional verification.

[0129] Through the above embodiments, the present application can process user text instructions through a preset language model, and perform similarity matching between user intention data and semantic vectors, so as to obtain more accurate second intention information, thereby improving the test efficiency of human-computer interaction.

[0130] Optionally, as Figure 7 shown, step S340 may further include steps S341 - S332.

[0131] In step S341, the human-computer interaction system determines whether the first intention information and the second intention information are consistent. If so, the human-computer interaction system outputs the user instruction intention.

[0132] Exemplarily, the first intention information is "Please confirm the equipment number and current operating status of the robotic arm on line B", and the second intention information is "Check the equipment number and current operating status of the robotic arm on line B". The first intention information and the second intention information are consistent, and the human-computer interaction system outputs the user instruction intention.

[0133] In step S342, the human-computer interaction system determines whether the first intention information and the second intention information are consistent. If not, based on the first intention information and the second intention information, the human-computer interaction system corrects and supplements the user intention through a multi-source information integration mechanism to obtain the user instruction intention.

[0134] For example, the multi-source information integration mechanism can be real-time data fusion.

[0135] Exemplarily, the first intention information is "Please restart the equipment number and current operating status of the robotic arm on line B", and the second intention information is "Check the equipment number and current operating status of the robotic arm on line B". The first intention information and the second intention information are inconsistent. The human-computer interaction system reads the robotic arm status code through real-time data fusion. The human-computer interaction system intercepts the "restart" instruction in the first intention information and extracts the "check" instruction in the second intention information. The human-computer interaction system fuses the device real-time data to generate an enhanced instruction, and finally outputs "Check the equipment number and current operating status of the robotic arm on line B".

[0136] Through the above embodiments, the present application can improve the intention parsing accuracy in complex scenarios through multi-model collaboration and multi-source information integration mechanism.

[0137] Optionally, as Figure 8 shown, step S400 may further include steps S410 - S460.

[0138] In step S410, the human-computer interaction system decomposes the user instruction intention to obtain at least one candidate action information.

[0139] For example, a human-machine interaction system can convert voice commands into structured commands, such as extracting objects, actions, starting points, and ending points in the user command intention. The human-machine interaction system can predict the paths of actions, starting points, and ending points to obtain candidate action information.

[0140] Exemplarily, the user command intention can be "move the display from area A to area B", object: display; starting point: area A; ending point: area B; operation type: handling.

[0141] The human-machine interaction system generates a list of candidate actions: the robotic arm grabs the display; the AGV (Automated Guided Vehicle) is scheduled to area A; path dynamic planning; real-time obstacle avoidance; display placement attitude calibration.

[0142] In step S420, the human-machine interaction system predicts the relevant probabilities of candidate action information pairs for implementing high-level commands through a language large model to obtain prediction probabilities.

[0143] For example, the human-machine interaction system evaluates the correlation between each action and the core command through a language large model-driven decision-making technology. The human-machine interaction system combines the human-machine collaboration rule base to exclude actions with low compliance.

[0144] Exemplarily, the operation of the robotic arm grabbing the directly operated object has a correlation of 0.95; the AGV scheduling as a necessary condition for transportation has a correlation of 0.93; the execution of path planning will affect transportation efficiency, with a correlation of 0.88; real-time obstacle avoidance as a safety requirement has a correlation of 0.82; the manual handling backup plan belongs to an action with low compliance, with a correlation of 0.12.

[0145] The human-machine interaction system combines the human-machine collaboration rule base to directly exclude the "direct manual handling" with low correlation.

[0146] In step S430, the human-machine interaction system determines candidate subtasks with guiding significance based on the prediction probability.

[0147] Exemplarily, the human-machine interaction system retains candidate actions with a correlation probability ≥ 0.8, such as robotic arm grabbing, AGV scheduling, path planning, or obstacle avoidance.

[0148] In step S440, the human-machine interaction system estimates the probability of the execution success rate of the candidate subtasks based on the current interaction environment state data and historical data to obtain the feasibility score of the candidate subtasks under the current conditions.

[0149] For example, the interactive environmental state data can be the working state of the device, and the historical data can be the operating state of the device. The human-machine interaction system uses the environmental state data and historical data to confirm the probability of the execution success rate of the candidate subtasks, and comprehensively obtains the feasibility score of the candidate subtasks under the current conditions.

[0150] Exemplarily, the human-machine interaction system confirms the availability of the robotic arm in area A (idle state score 0.95), sufficient AGV battery power (score 0.92), and unobstructed placement space in area B (score 0.89) through the environmental state data. The human-machine interaction system calls the process database. If the probability of electrostatic damage to "monitor handling" in history is > 10%, an anti-static packaging subtask is added (feasibility score +0.15).

[0151] In step S450, the human-machine interaction system determines the subtask planning information based on the prediction probability and feasibility score.

[0152] For example, the human-machine interaction system can perform weighted scoring by combining the prediction probability and feasibility score, obtain the comprehensive score of each candidate subtask, arrange the priorities of the candidate subtasks according to the comprehensive score, and determine the subtask planning information.

[0153] Exemplarily, the correlation ratio accounts for 60%, and the feasibility score accounts for 40%. For robotic arm grasping, it is 0.95×0.6 = 0.57, 0.95×0.4 = 0.38, and the comprehensive score is 0.95; the calculation method is as shown above, which will not be elaborated here, and the results are directly given. The comprehensive score of AGV scheduling is 0.93; the comprehensive score of dynamic path planning is 0.89; the comprehensive score of real-time obstacle avoidance is 0.83.

[0154] The human-machine interaction system executes robotic arm grasping - AGV scheduling - path planning - obstacle avoidance according to the priorities.

[0155] In step S460, the human-machine interaction system determines the high-level task sequence based on the subtask planning information.

[0156] For example, according to the timing logic, the high-level task sequence is determined.

[0157] Exemplarily, the human-machine interaction system can generate the optimal path through on-site scanning. Then the human-machine interaction system schedules the AGV to area A and controls the robotic arm to grasp the monitor. The human-machine interaction system controls the AGV to move along the planned path. The real-time obstacle avoidance module of the human-machine interaction system is activated. The human-machine interaction system controls the placement attitude calibration of the robotic arm and safely releases the monitor to area B.

[0158] Through the above embodiments, the present application can accurately screen core subtasks through intent parsing and relevance prediction based on a large language model. The present application can improve the success rate of task execution through a dynamic evaluation model that integrates real-time environmental perception and historical execution data. The present application can significantly improve the reliability and execution efficiency of task planning by making a weighted decision on the relevance probability and feasibility score.

[0159] Optionally, as Figure 9 shown, step S500 may further include steps S510 - S530.

[0160] In step S510, the human - machine interaction system determines subtask execution parameters based on the subtask planning information. The subtask execution parameters may at least include target parameters, input parameters, output parameters, and execution constraint conditions.

[0161] For example, the target parameters may include, but are not limited to, navigation - type tasks or process - type tasks, etc. The input parameters may be sensor data, the output parameters may be action feedback criteria, and the execution constraint conditions may be physical resource limitations.

[0162] Exemplarily, in the case where the target parameter is a navigation - type task, the human - machine interaction system may define the target coordinates as the shelf coordinates in area B. In the case where the target parameter is a process - type task, the human - machine interaction system may set the clamping force threshold at the end of the robotic arm to 25N ± 2N.

[0163] The input parameter may be to configure the lidar scanning frequency to 20Hz, the output parameter may be that the speed fluctuation of the AGV < 0.2m / s and the path deviation < 10cm, and the execution constraint condition may be to limit the CPU occupancy rate of the path planning algorithm ≤ 35%.

[0164] In step S520, the human - machine interaction system establishes a mapping relationship between the application programming interface and the subtask execution parameters to obtain machine control instructions.

[0165] Exemplarily, the human - machine interaction system calls an execution program through an API (Application Programming Interface) to execute the subtask and obtain machine control instructions.

[0166] In step S530, the human - machine interaction system performs human - machine interaction based on the machine control instructions.

[0167] Exemplarily, the human - machine interaction system may complete a clamping - moving - releasing action sequence according to the API instructions.

[0168] Through the above embodiments, the present application can accelerate instruction generation by constructing semantic-driven API mapping rules in combination with a resource awareness mechanism. In the execution phase, the present application realizes precise and efficient machine control through real-time verification of output parameters and dynamic adjustment of constraint conditions, and through parametric modeling and a dynamic interface mapping mechanism.

[0169] Figure 10 The structural schematic diagram of the human-computer interaction system according to the embodiment of the present application is shown.

[0170] According to another aspect of the present application, the present application provides a human-computer interaction system assisted by a large model. As Figure 10 shown, the human-computer interaction system 10 includes a speech recognition module 11, an intent understanding module 12, a task planning module 13, and an instruction execution module 14.

[0171] According to an exemplary embodiment, the speech recognition module 11 constructs a target speech recognition model based on an initial instruction set, and recognizes a user speech instruction from a user based on the target speech recognition model to determine a user text instruction according to the user speech instruction.

[0172] For example, the speech recognition module 11 receives an interaction instruction set from the field of human-computer interaction and constructs an initial instruction set. The speech recognition module 11 performs semantic expansion processing on the initial instruction set through a language large model to generate a synonymous instruction set with diverse expressions. The speech recognition module 11 performs speech synthesis recording on the initial instruction set and the synonymous instruction set to obtain a training speech instruction data set. The speech recognition module 11 is based on the training speech instruction data set and fine-tunes a preset speech recognition large model through a preset model fine-tuning technique to obtain a target speech recognition model.

[0173] Exemplarily, the interaction instruction set can be an instruction set composed of operation terms, such as "start arc welding", etc. The speech recognition module 11 constructs an initial instruction set according to different operation terms.

[0174] Exemplarily, the speech recognition module 11 performs semantic expansion processing on "start arc welding" through a language large model to generate a synonymous instruction set with diverse expressions, such as "start welding operation", "activate electrode welding", etc.

[0175] The speech recognition module 11 can perform speech synthesis recording on the speech data of industrial noise and interaction instructions such as "start arc welding", "start welding operation", "activate electrode welding", etc. to obtain a training speech instruction data set. For example, according to an operation term set containing industrial noise, a training speech instruction data set is obtained.

[0176] The speech recognition module 11 inputs the training speech instruction dataset into a preset large speech recognition model and fine-tunes the preset large speech recognition model through LoRA (Low-Rank Adaptation, a model fine-tuning technique). This can reduce computational complexity and memory requirements, resulting in an efficient target speech recognition model.

[0177] Optionally, the speech recognition module 11 receives an interaction instruction set from the field of human-computer interaction to construct an initial instruction set.

[0178] For example, the interaction instruction set from the field of human-computer interaction may include, but is not limited to, device operation instructions or parameter regulation instructions, etc. The speech recognition module 11 combines different device operation instructions or parameter regulation instructions to construct the initial instruction set.

[0179] Exemplarily, the interaction instruction set may be a device operation instruction. As an example, the device operation instruction may be "Go to area A to grab a flange". As an example, the device operation instruction may be "Start the flange grabbing program".

[0180] The interaction instruction set may be a parameter regulation instruction. As an example, the parameter regulation instruction may be "Adjust the moving speed of the robotic arm to 0.3 m / s".

[0181] Optionally, the speech recognition module 11 performs semantic expansion processing on the initial instruction set based on a large language model to generate a synonymous instruction set.

[0182] For example, the semantic expansion processing may include, but is not limited to, synonym generation, professional term conversion, or dialect adaptation, etc.

[0183] Exemplarily, in the case where the semantic expansion processing is synonym generation, the initial instruction is "Go to area A to get a flange part". After synonym generation by the speech recognition module 11, synonymous instructions such as "Slide to area A to get a flange part" or "Transmit to area A to get a flange part" can be obtained.

[0184] In the case where the semantic expansion processing is professional term conversion, the initial instruction is "Go to area A to get a flange part". After professional term conversion by the speech recognition module 11, "Activate the pickup protocol for area A" can be obtained.

[0185] In the case where the semantic expansion processing is dialect adaptation, the initial instruction is "Go to area A to get a flange part". After dialect adaptation by the speech recognition module 11, "Go to area A to get a flange part" or "Move to area A to get a flange part" etc. can be obtained.

[0186] Optionally, the speech recognition module 11 performs speech synthesis recording on the initial instruction set and the synonymous instruction set to obtain a training speech instruction dataset.

[0187] For example, the voice synthesis recording process may include, but is not limited to, multi-timbre generation, environment adaptation processing, or stress emphasis processing, etc. The voice recognition module 11 generates a training voice instruction data set through the initial instruction set and the synonymous instruction set after synthesis processing.

[0188] Exemplarily, the initial instruction is "Emergency stop the robotic arm in Area A", and the synonymous instructions are "Immediately shut down the robot in Section A" and "Activate the emergency stop protocol in Area A".

[0189] In the case where the voice synthesis recording process is multi-timbre generation, the voice recognition module 11 can generate 8 timbres through WaveNet (WaveNet), including male and female voices, and different age groups.

[0190] In the case where the voice synthesis recording process is environment adaptation processing, as an embodiment, the voice recognition module 11 simulates the factory environment by superimposing 85dB white noise. As an embodiment, the voice recognition module 11 simulates the factory environment by adding mechanical impact sounds.

[0191] In the case where the voice synthesis recording process is stress emphasis processing, the voice recognition module 11 increases the volume of the instruction keyword "stop" by 6dB.

[0192] Optionally, the voice recognition module 11 fine-tunes a preset large voice recognition model based on the training voice instruction data set through a preset model fine-tuning technique to obtain a target voice recognition model.

[0193] Exemplarily, the preset large voice recognition model can be the Whisper model, and the preset model fine-tuning technique can be the LoRA technique. The voice recognition module 11 fine-tunes the preset large voice recognition model through the LoRA technique. The voice recognition module 11 introduces low-rank matrices inside each layer of the model, which can effectively reduce the number of model parameters and resource requirements, while maintaining the performance of the model, enabling the model to recognize more accurately.

[0194] Through the above embodiments, the present application can ensure a high degree of matching between the training data and the industrial scenario by receiving real instructions in the human-computer interaction scenario. Then, through semantic extension, the present application can improve the recognition accuracy of the model for colloquial expressions and dialect variants. Through multi-environment simulation recording of the extended instruction set, the present application can further improve the recognition ability of the model. Through fine-tuning the model, the present application can improve the response speed of the target voice recognition model to voice instructions, and further improve the test efficiency of human-computer interaction.

[0195] Optionally, the voice recognition module 11 determines a user text instruction based on the recognized user voice instruction from the user.

[0196] For example, the user voice instruction can include, but is not limited to, an operation instruction (such as starting a welding program), a multi-modal composite instruction (such as grasping this part), or a complex task instruction (such as alarming if the pressure exceeds 50 MPa), etc.

[0197] The voice recognition module 11 performs speech-to-text processing on the user voice instruction through the target voice recognition model to determine the corresponding user text instruction.

[0198] Exemplarily, the voice recognition module 11 receives a multi-modal composite instruction, such as "go to area A to get a flange part". The voice recognition module 11 performs speech-to-text processing on the multi-modal composite instruction through the target voice recognition model, and fine-tunes the multi-modal composite instruction through the target voice recognition model to obtain a user text instruction, such as "go to area A to grasp the flange".

[0199] According to the example embodiment, the intent understanding module 12 performs intent fusion recognition on the user text instruction based on the user text instruction through a large language model and a preset small language model to obtain the user instruction intent.

[0200] For example, the large language model can process the text through natural language processing technology.

[0201] For example, the intent understanding module 12 constructs a prompt template through the user text instruction and the preset interaction task information, and then generates a guiding prompt word through the prompt template. The intent understanding module 12 inputs the guiding prompt word and the user text instruction into the large language model for semantic understanding to obtain the key elements in the user text instruction. Then, the intent understanding module 12 performs semantic alignment and intent recognition on the guiding prompt word and the key elements to obtain the first intent information.

[0202] The intent understanding module 12 obtains a preset small language model by training a text detection small model of machine learning based on a sample set containing specific tasks.

[0203] The intent understanding module 12 processes the user text instruction to obtain intent data information. The intent understanding module 12 processes the intent data information based on the preset small language model to obtain the second intent information.

[0204] The intent understanding module 12 performs intent fusion recognition on the first intent information and the second intent information to obtain the user instruction intent.

[0205] Optionally, the intent understanding module 12 constructs a preset small language model based on a sample set containing specific tasks.

[0206] For example, the intent understanding module 12 constructs a sample set containing specific tasks, and then the intent understanding module 12 uses the sample set to train a small text detection model of machine learning to obtain a preset language small model.

[0207] The sample set containing specific tasks may include data composition and adversarial training samples. The small text detection model may be built-in with a factory-specific vocabulary, a noise suppression processing module, and a hybrid preference optimization.

[0208] Exemplarily, the data composition may be a voice command based on the picking of factory goods, and the voice command may include a noisy environment (such as mechanical noise) and a dialect mixture sample (such as Mandarin mixed with Cantonese and Minnan dialect). The adversarial training sample may be voice data with synthetic mixed noise, such as mixing the mechanical noise of a robotic arm with background human voice interference to generate adversarial samples with dynamic signal-to-noise ratios.

[0209] The factory-specific vocabulary may include "robotic arm coding", "PLC control instructions", etc. The noise suppression processing module may be noise suppression, which filters out the characteristics of mechanical roars. The intent understanding module 12 identifies the command keywords in the voice command through the factory-specific vocabulary and the noise suppression processing module.

[0210] The hybrid preference optimization may adopt a negative supervision correction algorithm to reduce misidentifications caused by dialect pronunciations, such as the acoustic feature confusion between "emergency stop" and "machine stop".

[0211] Optionally, the intent understanding module 12 obtains first intent information based on the user's text instruction and according to the large language model.

[0212] For example, the intent understanding module 12 constructs a prompt template based on the user's text instruction and the preset interaction task information. The intent understanding module 12 generates guiding prompt words based on the prompt template. The intent understanding module 12 inputs the guiding prompt words and the user's text instruction into the large language model for semantic understanding to obtain the key elements in the user's text instruction. The intent understanding module 12 performs semantic alignment and intent recognition on the guiding prompt words and the key elements to obtain the first intent information.

[0213] The preset interaction task information may include but is not limited to device control instructions or safety warning processing, etc. The generation of guiding prompt words may be dynamic parameter filling. Semantic understanding may include identifying and processing the polysemy and ambiguity of vocabulary and understanding the meaning of the text. Key element extraction may include entity recognition and relationship extraction.

[0214] Semantic alignment and intent recognition may include but are not limited to multi-dimensional verification and an intent decision tree. Multi-dimensional verification may include but is not limited to integrity checking and conflict detection, etc.

[0215] Exemplarily, the user text instruction can be to go to area A to grab a flange, and the device control instruction can be to identify the robotic arm operation instruction and map it to the PLC (Programmable Logic Controller) control protocol. The security warning handling can be to detect abnormal instructions and trigger an emergency response mechanism, where the abnormal instructions are such as "Emergency stop grabbing" and "High temperature alarm".

[0216] The dynamic parameter filling can include a basic template or scenario adaptation. For example, the basic template can be to specify the device number to be operated (such as in area A / number B) and the specific instruction (start / speed adjustment / emergency stop). For example, the scenario adaptation can be that if the user input contains noise (such as "Why doesn't the robotic arm in area B move?"), an enhanced prompt "Please confirm the device number of the robotic arm on line B and its current operating status" is generated.

[0217] Entity recognition can include the device number ("area A"), the operation type ("pause"), and the target parameter ("0 m / s"). The relationship extraction can be to establish a triple of "area A - Speed_Adjust - 0 m / s" and map it to the PLC control instruction.

[0218] The integrity check can be to verify whether the combination of "device number + operation type" is included. If it is missing, a follow-up question is triggered: "Please specify the specific device area". The conflict detection can be that when the instruction parameter exceeds the limit (such as "Adjust the speed to 5 m / s" exceeds the device maximum limit of 3.5 m / s), it is marked as a "red risk instruction".

[0219] Optionally, the intent understanding module 12 obtains the second intent information based on the user text instruction according to the preset language small model.

[0220] For example, the intent understanding module 12 processes the user text instruction to obtain the intent data information. The intent understanding module 12 processes the intent data information based on the preset language small model to obtain the second intent information.

[0221] The intent data processing can include but is not limited to preprocessing in a noisy environment and extracting intent elements, etc. The preprocessing in a noisy environment can include but is not limited to dialect normalization and noise suppression, etc.

[0222] Exemplarily, the dialect normalization can be to convert the Cantonese instruction "Why doesn't the robotic arm in area B move?" to the standard instruction "The robotic arm in area B has an abnormal shutdown". The noise suppression can be to use a dual-path attention mechanism to filter out the mechanical roar of the robotic arm.

[0223] The intent understanding module 12 extracts the intent elements from "The robotic arm in area B has an abnormal shutdown" to obtain "Check the device number of the robotic arm on line B and its current operating status". The intent understanding module 12 performs semantic integrity screening on the instruction of "Check the device number of the robotic arm on line B and its current operating status" and executes a full-dimensional verification.

[0224] Optionally, the intent understanding module 12 performs intent fusion recognition on the first intent information and the second intent information to obtain the user instruction intent.

[0225] For example, the intent understanding module 12 determines whether the first intent information and the second intent information are consistent. If so, the intent understanding module 12 outputs the user instruction intent; if not, the intent understanding module 12 supplements and corrects the user intent through a multi-source information integration mechanism based on the first intent information and the second intent information to obtain the user instruction intent.

[0226] The multi-source information integration mechanism can be real-time data fusion.

[0227] Exemplarily, the first intent information is "Please confirm the device number and current operating status of the robotic arm on line B", and the second intent information is "Check the device number and current operating status of the robotic arm on line B". The first intent information and the second intent information are consistent, and the intent understanding module 12 outputs the user instruction intent.

[0228] The first intent information is "Please restart the device number and current operating status of the robotic arm on line B", and the second intent information is "Check the device number and current operating status of the robotic arm on line B". The first intent information and the second intent information are inconsistent. The intent understanding module 12 reads the robotic arm status code through real-time data fusion. The intent understanding module 12 intercepts the "restart" instruction in the first intent information and extracts the "check" instruction in the second intent information. The intent understanding module 12 fuses the device real-time data to generate an enhanced instruction and finally outputs "Check the device number and current operating status of the robotic arm on line B".

[0229] Through the above embodiments, the present application can parse the deep semantics of user instructions through the massive pre-training data of the language model, and can support the context-related reasoning of complex instructions. The present application can construct a preset language small model based on a specific task sample set to achieve the domain focus of the small model and reduce the semantic divergence risk of the large model in professional scenarios. The present application can obtain accurate user instruction intents through intent fusion recognition, thereby improving the accuracy of instruction analysis and further improving the test efficiency of human-computer interaction.

[0230] Optionally, the intent understanding module 12 constructs a sample set including specific tasks.

[0231] For example, the sample set including specific tasks may include data composition and adversarial training samples.

[0232] Exemplarily, the data composition can be a voice command based on the factory goods picking composition, and the voice command can include a noisy environment (such as mechanical noise) and a dialect mixture sample (such as Mandarin mixed with Cantonese and Minnan dialect). The adversarial training sample can be voice data synthesizing mixed noise, such as mixing the mechanical noise of the robotic arm and the background human voice interference to generate an adversarial sample with a dynamic signal-to-noise ratio.

[0233] Optionally, the intent understanding module 12 processes the intent data information based on a preset language small model to obtain second intent information.

[0234] For example, the preset language small model can be a text detection small model, and the text detection small model can be built-in with a factory-specific vocabulary, a noise suppression processing module, and a mixed preference optimization.

[0235] Exemplarily, the factory-specific vocabulary can include "robotic arm coding", "PLC control instruction", etc. The noise suppression processing module can be noise suppression, and the noise suppression filters the mechanical roar characteristics. The intent understanding module 12 identifies the instruction keywords in the voice command through the factory-specific vocabulary and the noise suppression processing module.

[0236] The mixed preference optimization can adopt a negative supervision correction algorithm to reduce the misrecognition caused by dialect pronunciation, such as the acoustic feature confusion between "emergency stop" and "machine stop".

[0237] Through the above embodiments, the present application can cover typical factory interference scenarios in the training stage of the small model by constructing a sample set containing a mixture of mechanical noise and dialects. Through the preset language small model, the present application can further improve the accuracy of instruction analysis, thereby improving the test efficiency of human-computer interaction.

[0238] Optionally, the intent understanding module 12 constructs a prompt template based on the user text instruction and the preset interaction task information.

[0239] For example, the preset interaction task information can include but is not limited to device control instructions or safety warning processing, etc.

[0240] Exemplarily, the user text instruction can be to go to area A to pick up a flange, the device control instruction can be to identify the robotic arm operation instruction and map it to the PLC (programmable logic controller) control protocol. The safety warning processing can be to detect abnormal instructions and trigger an emergency response mechanism, where the abnormal instructions are such as "emergency stop picking", "high temperature alarm".

[0241] Optionally, the intent understanding module 12 generates guiding prompt words based on the prompt template.

[0242] For example, the generation of guiding prompt words can be dynamic parameter filling.

[0243] Exemplarily, dynamic parameter filling may include basic template or scenario adaptation. For example, the basic template may be to specify the device number to be operated (such as Area A / Number B) and specific instructions (start / speed adjustment / emergency stop). For example, scenario adaptation may be that if the user input contains noise (such as "Why doesn't the robotic arm No. B move?"), an enhanced prompt "Please confirm the device number of the robotic arm on Line B and its current operating status" is generated.

[0244] Optionally, the intent understanding module 12 inputs the guiding prompt words and the user text instruction into the language large model for semantic understanding to obtain the key elements in the user text instruction.

[0245] For example, semantic understanding may include identifying and processing the polysemy and ambiguity of words and understanding the meaning of the text. Key element extraction may include entity recognition and relationship extraction.

[0246] Exemplarily, entity recognition may include device number ("Area A"), operation type ("pause"), target parameter ("0m / s"). Relationship extraction may be to establish a triple of "Area A - Speed_Adjust - 0m / s" and map it to the PLC control instruction.

[0247] Optionally, the intent understanding module 12 performs semantic alignment and intent recognition on the guiding prompt words and the key elements to obtain the first intent information.

[0248] For example, semantic alignment and intent recognition may include but are not limited to multi-dimensional verification and intent decision tree. Multi-dimensional verification may include but are not limited to integrity check and conflict detection, etc.

[0249] Exemplarily, the integrity check may be to verify whether the combination of "device number + operation type" is included. If it is missing, a follow-up question is triggered: "Please specify the specific device area." Conflict detection may be when the instruction parameter exceeds the limit (such as "adjust the speed to 5m / s" exceeds the device maximum limit of 3.5m / s), it is marked as a "red risk instruction".

[0250] Through the above embodiments, the present application can construct a prompt template by presetting interaction task information, enabling the language large model to focus its attention on domain key elements. The present application can also trigger the domain knowledge activation mechanism of the large model through guiding prompt words, improving the semantic understanding accuracy of the language large model. The present application can obtain more accurate first intent information through semantic alignment and intent recognition, thereby improving the test efficiency of human-computer interaction.

[0251] Optionally, the intent understanding module 12 processes the user text instruction to obtain intent data information.

[0252] For example, intent data processing may include, but is not limited to, preprocessing in a noisy environment and extracting intent elements, etc. Preprocessing in a noisy environment may include, but is not limited to, dialect normalization and noise suppression, etc.

[0253] Exemplarily, dialect normalization may be converting the Cantonese instruction "Why doesn't point B move?" to the standard instruction "The robotic arm at point B has an abnormal shutdown". Noise suppression may be filtering out the mechanical rumbling sound of the robotic arm using a dual-path attention mechanism.

[0254] Optionally, the intent understanding module 12 processes the intent data information based on a preset language small model to obtain second intent information.

[0255] Exemplarily, the intent understanding module 12 extracts intent elements from "The robotic arm at point B has an abnormal shutdown" to obtain "Check the equipment number and current operating status of the robotic arm on line B". The intent understanding module 12 performs semantic integrity screening on the instruction "Check the equipment number and current operating status of the robotic arm on line B" and executes a full-dimensional verification.

[0256] Through the above embodiments, the present application can process user text instructions through a preset language small model and perform similarity matching between user intent data and semantic vectors, so as to be able to obtain more accurate second intent information, thereby improving the test efficiency of human-computer interaction.

[0257] Optionally, the intent understanding module 12 determines whether the first intent information and the second intent information are consistent. If so, the intent understanding module 12 outputs the user instruction intent.

[0258] Exemplarily, the first intent information is "Please confirm the equipment number and current operating status of the robotic arm on line B", the second intent information is "Check the equipment number and current operating status of the robotic arm on line B", the first intent information and the second intent information are consistent, and the intent understanding module 12 outputs the user instruction intent.

[0259] Optionally, the intent understanding module 12 determines whether the first intent information and the second intent information are consistent. If not, based on the first intent information and the second intent information, the intent understanding module 12 supplements and corrects the user intent through a multi-source information integration mechanism to obtain the user instruction intent.

[0260] For example, the multi-source information integration mechanism may be real-time data fusion.

[0261] Exemplarily, the first intent information is "Please restart the robotic arm device number and current operating status of Line B", and the second intent information is "Check the robotic arm device number and current operating status of Line B". The first intent information and the second intent information are inconsistent. The intent understanding module 12 reads the robotic arm status code through real-time data fusion. The intent understanding module 12 intercepts the "restart" instruction in the first intent information and extracts the "check" instruction in the second intent information. The intent understanding module 12 fuses the real-time device data to generate an enhanced instruction, and finally outputs "Check the robotic arm device number and current operating status of Line B".

[0262] Through the above embodiments, the present application can improve the intent parsing accuracy in complex scenarios through a multi-model collaboration and multi-source information integration mechanism.

[0263] According to the exemplary embodiment, the task planning module 13 decomposes and plans the user instruction intent to obtain a high-level task sequence.

[0264] For example, the task planning module 13 decomposes the user instruction intent to obtain at least one candidate action information. The task planning module 13 predicts the relevant probability of the candidate action information for realizing the high-level instruction through a language large model, and obtains the prediction probability of the candidate action information. The task planning module 13 determines a candidate subtask with guiding significance based on the prediction probability. The task planning module 13 estimates the probability of the execution success rate of the candidate subtask according to the current interaction environment state data and historical data, and obtains the feasibility score of the candidate subtask under the current conditions. The task planning module 13 determines the subtask planning information based on the prediction probability and the feasibility score. The task planning module 13 determines the high-level task sequence based on the subtask planning information.

[0265] For example, the high-level task sequence can be a multi-level task structure driven by goals, and realizes a systematic execution path for achieving the goals through step-by-step decomposition and dynamic adjustment.

[0266] Optionally, the task planning module 13 decomposes the user instruction intent to obtain at least one candidate action information.

[0267] For example, the task planning module 13 can convert the voice instruction into a structured instruction, such as extracting the object, action, starting point, and ending point in the user instruction intent. The task planning module 13 can predict the paths of the action, starting point, and ending point to obtain the candidate action information.

[0268] Exemplarily, the user instruction intent can be "Move the display from Area A to Area B", object: display; starting point: Area A; ending point: Area B; operation type: handling.

[0269] The task planning module 13 generates a list of candidate actions: the robotic arm grasps the display; the AGV (Automated Guided Vehicle) is scheduled to area A; path dynamic planning; real-time obstacle avoidance; display placement attitude calibration.

[0270] Optionally, the task planning module 13 predicts the relevant probability of candidate action information pairs for implementing high-level instructions through a language large model to obtain a prediction probability.

[0271] For example, the task planning module 13 evaluates the relevance of each action to the core instruction through a language large model-driven decision-making technique. The task planning module 13 combines the human-machine collaboration rule base to exclude actions with low compliance.

[0272] Exemplarily, the operation of the robotic arm grasping the direct operation object has a relevance of 0.95; the AGV scheduling as a necessary condition for transportation has a relevance of 0.93; the execution of path planning affects the transportation efficiency, with a relevance of 0.88; real-time obstacle avoidance as a safety requirement has a relevance of 0.82; the manual handling alternative is a low-compliance action with a relevance of 0.12.

[0273] The task planning module 13 combines the human-machine collaboration rule base to directly exclude the low-relevance "direct manual handling".

[0274] Optionally, the task planning module 13 determines a guiding candidate subtask based on the prediction probability.

[0275] Exemplarily, the task planning module 13 retains candidate actions with a relevance probability ≥ 0.8, such as robotic arm grasping, AGV scheduling, path planning, or obstacle avoidance.

[0276] Optionally, the task planning module 13 estimates the probability of the execution success rate of the candidate subtask according to the current interaction environmental state data and historical data to obtain the feasibility score of the candidate subtask under the current conditions.

[0277] For example, the interaction environmental state data can be the working state of the instrument, and the historical data can be the operating state of the instrument. The task planning module 13 confirms the probability of the execution success rate of the candidate subtask through the environmental state data and historical data, and comprehensively obtains the feasibility score of the candidate subtask under the current conditions.

[0278] Exemplarily, the task planning module 13 confirms the availability of the robotic arm in area A (idle state score 0.95), sufficient power of the AGV (score 0.92), and unobstructed placement space in area B (score 0.89) through the environmental status data. By calling the process database, if the probability of electrostatic damage to "monitor handling" in history is > 10%, the task planning module 13 adds an anti-static packaging subtask (feasibility score +0.15).

[0279] Optionally, the task planning module 13 determines the subtask planning information based on the prediction probability and feasibility score.

[0280] For example, the task planning module 13 can perform weighted scoring by combining the prediction probability and feasibility score, obtain the comprehensive score of each candidate subtask, arrange the priorities of the candidate subtasks according to the comprehensive score, and determine the subtask planning information.

[0281] Exemplarily, the correlation ratio accounts for 60%, the feasibility score accounts for 40%, for robotic arm grasping, 0.95×0.6 = 0.57, 0.95×0.4 = 0.38, and the comprehensive score is 0.95; the calculation method is as shown above, which will not be elaborated here, and the results are directly given. The comprehensive score of AGV scheduling is 0.93; the comprehensive score of dynamic path planning is 0.89; the comprehensive score of real-time obstacle avoidance is 0.83.

[0282] The task planning module 13 executes robotic arm grasping - AGV scheduling - path planning - obstacle avoidance according to the priorities.

[0283] Optionally, the task planning module 13 determines the high-level task sequence based on the subtask planning information.

[0284] For example, according to the timing logic, the high-level task sequence is determined.

[0285] Exemplarily, the task planning module 13 can generate the optimal path through on-site scanning. Then the task planning module 13 schedules the AGV to area A and controls the robotic arm to grasp the monitor. The task planning module 13 controls the AGV to move along the planned path. The real-time obstacle avoidance module of the task planning module 13 is activated. The task planning module 13 controls the robotic arm to calibrate the placement posture and safely release the monitor to area B.

[0286] Through the above embodiments, the present application can accurately screen the core subtasks through intention parsing and correlation prediction based on the language large model. The present application can improve the task execution success rate through a dynamic evaluation model that fuses real-time environment perception and historical execution data. The present application can significantly improve the reliability and execution efficiency of task planning through weighted decision-making on the correlation probability and feasibility score.

[0287] According to an exemplary embodiment, the instruction execution module 14 generates machine control instructions according to a high-level task sequence to perform human-machine interaction according to the machine control instructions.

[0288] For example, the instruction execution module 14 determines sub-task execution parameters based on sub-task planning information. The sub-task execution parameters may at least include target parameters, input parameters, output parameters, and execution constraint conditions. The instruction execution module 14 establishes a mapping relationship between the application programming interface and the sub-task execution parameters to obtain machine control instructions. Then, the instruction execution module 14 performs human-machine interaction based on the machine control instructions.

[0289] Through the above embodiments, the present application can build a target speech recognition model through an initial instruction set, enabling the target speech recognition model to simulate various simulated speech scenarios in human-machine interaction experiments, so that the present application can more accurately recognize user speech instructions to determine user text instructions. The present application performs intent fusion recognition on user text instructions through a language large model and a preset language small model, and can achieve dynamic adaptation of general semantic understanding and vertical domain knowledge, so that the present application can accurately obtain the intent of user instructions. Then, the present application decomposes and plans the user instruction intent to determine the execution steps with the highest execution rate, thereby combining to obtain a high-level task sequence. Finally, the present application generates machine control instructions through the high-level task sequence to perform human-machine interaction, enabling the present application to accurately perform human-machine interaction operations, thereby improving the efficiency of human-machine interaction testing.

[0290] Optionally, the instruction execution module 14 determines sub-task execution parameters based on sub-task planning information. The sub-task execution parameters may at least include target parameters, input parameters, output parameters, and execution constraint conditions.

[0291] For example, the target parameters may include, but are not limited to, navigation tasks or process tasks, etc. The input parameters may be sensor data, the output parameters may be action feedback criteria, and the execution constraint conditions may be physical resource limitations.

[0292] Exemplarily, when the target parameter is a navigation task, the instruction execution module 14 may define the target coordinates as the shelf coordinates of area B. When the target parameter is a process task, the instruction execution module 14 may set the clamping force threshold at the end of the robotic arm to 25N ± 2N.

[0293] The input parameter may be to configure the lidar scanning frequency to 20Hz, the output parameter may be that the speed fluctuation of the AGV < 0.2m / s and the path deviation < 10cm, and the execution constraint condition may be to limit the CPU occupancy rate of the path planning algorithm ≤ 35%.

[0294] Optionally, the instruction execution module 14 establishes a mapping relationship between the application programming interface and the subtask execution parameters to obtain a machine control instruction.

[0295] Exemplarily, the instruction execution module 14 invokes an execution program through an API (Application Programming Interface) to execute a subtask and obtain a machine control instruction.

[0296] Optionally, the instruction execution module 14 performs human-machine interaction based on the machine control instruction.

[0297] Exemplarily, the instruction execution module 14 can complete a gripping-moving-releasing action sequence according to the API instruction.

[0298] Through the above embodiments, the present application can accelerate instruction generation by constructing semantic-driven API mapping rules in combination with a resource awareness mechanism. In the execution phase, the present application realizes precise and efficient machine control through real-time verification of output parameters and dynamic adjustment of constraint conditions, and through parametric modeling and dynamic interface mapping mechanisms.

[0299] According to another aspect of the present application, the present application also provides an electronic device. The electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors can implement the human-machine interaction method as described above.

[0300] According to another aspect of the present application, the present application also provides a non-volatile computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it can implement the human-machine interaction method as described above.

[0301] According to another aspect of the present application, the present application also provides a computer program product. The computer program product includes: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is made to execute the human-machine interaction method as described above.

[0302] Finally, it should be noted that the above are only the preferred embodiments of the present application and are not used to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions of the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A human-computer interaction method assisted by a large model, characterized in that Including: Construct a target speech recognition model based on an initial instruction set; Recognize a user speech instruction from a user based on the target speech recognition model to determine a user text instruction according to the user speech instruction; Based on the user text instruction, perform intent fusion recognition on the user text instruction through a large language model and a preset small language model to obtain a user instruction intent; Decompose and plan the user instruction intent to obtain a high-level task sequence; Generate a machine control instruction according to the high-level task sequence to perform human-machine interaction according to the machine control instruction.

2. The human-computer interaction method according to claim 1, wherein The constructing the target speech recognition model based on the initial instruction set includes: Receive an interaction instruction set from the field of human-machine interaction to construct an initial instruction set; Based on the large language model, perform semantic extension processing on the initial instruction set to generate a synonymous instruction set; Perform speech synthesis recording on the initial instruction set and the synonymous instruction set to obtain a training speech instruction data set; Based on the training speech instruction data set, fine-tune a preset large speech recognition model through a preset model fine-tuning technique to obtain the target speech recognition model.

3. The human-computer interaction method according to claim 1, characterized in that The performing intent fusion recognition on the user text instruction through a large language model and a preset small language model based on the user text instruction to obtain a user instruction intent includes: Construct the preset small language model based on a sample set including specific tasks; Based on the user text instruction, obtain first intent information according to the large language model; Based on the user text instruction, obtain second intent information according to the preset small language model; Perform intent fusion recognition on the first intent information and the second intent information to obtain the user instruction intent.

4. The human-computer interaction method according to claim 3, characterized in that, The constructing the preset small language model based on the sample set including specific tasks includes: Construct the sample set including specific tasks; Train a text detection small model of machine learning based on the sample set to obtain the preset small language model.

5. The human-computer interaction method according to claim 3, characterized in that The obtaining first intent information according to the large language model based on the user text instruction includes: Construct a prompt template based on the user text instruction and preset interaction task information; Generate a guiding prompt word based on the prompt template; Input the guiding prompt word and the user text instruction into the large language model for semantic understanding to obtain key elements in the user text instruction; Perform semantic alignment and intent recognition on the guiding prompt word and the key elements to obtain the first intent information.

6. The human-computer interaction method according to claim 3, wherein The obtaining second intent information according to the preset small language model based on the user text instruction includes: Process the user text instruction to obtain intent data information; Process the intent data information based on the preset small language model to obtain the second intent information.

7. The human-computer interaction method according to claim 3, wherein The performing intent fusion recognition on the first intent information and the second intent information to obtain the user instruction intent includes: Judge whether the first intent information and the second intent information are consistent, If so, output the user instruction intent; Otherwise, based on the first intent information and the second intent information, supplement and correct the user intent through a multi-source information integration mechanism to obtain the user instruction intent.

8. The human-computer interaction method according to claim 1, wherein, The task decomposition and planning of the user instruction intent to obtain a high-level task sequence includes: Decompose the user instruction intent to obtain at least one candidate action information; Predict the relevant probability of the candidate action information for realizing the high-level instruction through the language large model to obtain a prediction probability; Determine a candidate subtask with guiding significance based on the prediction probability; Estimate the probability of the execution success rate of the candidate subtask according to the current interaction environment state data and historical data to obtain the feasibility score of the candidate subtask under the current conditions; Determine the subtask planning information based on the prediction probability and the feasibility score; Determine the high-level task sequence based on the subtask planning information.

9. The human-computer interaction method according to claim 8, wherein The generation of a machine control instruction according to the high-level task sequence and the execution of the human-machine interaction according to the machine control instruction includes: Determine the subtask execution parameters based on the subtask planning information; Establish a mapping relationship between the application programming interface and the subtask execution parameters to obtain the machine control instruction; Execute the human-machine interaction based on the machine control instruction; The subtask execution parameters may at least include target parameters, input parameters, output parameters, and execution constraint conditions.

10. A human-computer interaction system assisted by a large model, characterized in that, The human-machine interaction system executes the human-machine interaction method according to any one of claims 1-9. The human-machine interaction system includes: A speech recognition module that constructs a target speech recognition model based on an initial instruction set and recognizes a user speech instruction from the user based on the target speech recognition model to determine a user text instruction according to the user speech instruction; An intent understanding module that performs intent fusion recognition on the user text instruction through a language large model and a preset language small model based on the user text instruction to obtain a user instruction intent; A task planning module that performs task decomposition and planning on the user instruction intent to obtain a high-level task sequence; An instruction execution module that generates a machine control instruction according to the high-level task sequence and executes the human-machine interaction according to the machine control instruction.

Citation Information

Patent Citations

  • Pre-training large model assisted end-to-end network intention decomposition and decision-making method

    CN118158119A

  • Computer and large model interaction system based on manual dictation command

    CN119479651A

  • Dual-prevention intelligent interaction system based on large language model

    CN119579365A

  • User intention alignment robot task planning method based on large language model

    CN119658692A

  • Interactive intention understanding and fast learning system based on multi-mode information fusion

    CN119940369A

Cited By

  • Intention recognition method, intention recognition device and intelligent glasses

    CN120783736A

  • Quadruped robot voice interaction method and system based on cooperation of double large models

    CN121459816A