Human-machine natural interaction method and system based on body intelligence
By using dynamic segmentation technology based on motion trajectory turning points and voice pauses, combined with DTW and cosine similarity algorithms, the problems of insufficient interaction naturalness and cross-scenario adaptability in traditional human-computer interaction methods are solved, achieving more efficient interaction fluency and adaptability.
Patent Information
- Application Number
- CN202510846008.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Traditional human-computer interaction methods rely on users to learn specific commands, resulting in insufficient naturalness in interaction and poor cross-scenario adaptability. It is difficult to accurately capture the semantic boundaries of users' continuous actions or complex commands, and the error correction mechanism for abnormal commands relies on fixed template matching and cannot be intelligently reorganized based on context and user habits, resulting in low interaction efficiency.
By collecting user interaction information, dynamically segmenting based on motion trajectory turning points and voice pauses, combining dynamic time warping (DTW) and cosine similarity algorithms, filtering out abnormal segmentation instructions, reorganizing interaction instructions, and building a recognition-verification-reorganization-execution closed loop, it supports personalized adjustment and incremental learning, thereby improving system adaptability.
It improves the accuracy of command segmentation, solves the ambiguity of semantic segmentation of continuous actions, reduces the burden of secondary input for users, improves the smoothness of interaction and the efficiency of repairing abnormal commands, and enhances the system's adaptability and natural interaction.
Smart Images

Figure CN120669865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information interaction and communication technology, and in particular to a method and system for natural human-computer interaction based on embodied intelligence. Background Art
[0002] With the development of artificial intelligence and sensor technology, human-computer interaction systems based on embodied intelligence are gradually moving from laboratories to practical applications.
[0003] According to the patent application with publication number CN112462940A, a smart home multimodal human-computer natural interaction system and method are disclosed, which include a gesture recognition model pre-training module, which uses a gesture data set that meets the scenario to train the constructed network model and save the trained gesture recognition model; a speech recognition model pre-training module, which uses a Chinese speech data set to train the acoustic model and language model in sequence and save the trained speech recognition model; a gesture recognition module, which uses the saved gesture recognition model to predict the collected gestures; a speech recognition module, which calls the saved speech recognition model to recognize the collected audio; and a multimodal fusion module, which fuses the two modal results of the gesture recognition module and the speech recognition module to obtain the final instruction.
[0004] Traditional human-computer interaction methods (such as keyboard, mouse, and touch) rely on users to learn specific commands, and have problems such as insufficient naturalness of interaction and poor cross-scenario adaptability. At the same time, they lack dynamic adaptability to the segmented processing of continuous user actions or complex commands, and it is difficult to accurately capture pause points and semantic boundaries. Secondly, the error correction mechanism for abnormal commands relies on fixed template matching and cannot be intelligently reorganized based on context and user habits, resulting in low interaction efficiency. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a human-computer natural interaction method and system based on embodied intelligence, which solves the problem of lack of dynamic adaptability in segmented processing of user continuous actions or complex instructions.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a human-computer natural interaction method based on embodied intelligence, the method specifically comprising the following steps: Collect user interaction information, determine whether it can be recognized, and generate an interaction recognition signal or a secondary interaction input signal; Analyze the interaction recognition signal, obtain the user's error feedback to generate the interaction correction analysis signal, and then segment the interaction instructions according to the turning points and match them with the pre-stored instructions to obtain the abnormal segmented instructions; The instructions with the maximum similarity between the abnormal segmented instructions and the pre-stored instructions are recorded as pre-selected instructions, and the interaction results or secondary interaction input signals are generated in combination with user feedback; Analyze the secondary interactive input signal, segment the secondary input interactive information to obtain secondary interactive segment instructions, and match and filter with pre-stored combination instruction information to obtain the combination instruction to be analyzed; Obtain the segmented instructions that do not match the secondary interaction segmented instructions in the combination instructions to be analyzed, calculate their similarity with the corresponding pre-stored instructions, and sum up the similarity mean. Then, interact with the combination to be analyzed with the maximum similarity mean as the standard to generate the interaction result.
[0007] As a further solution of the present invention, the specific method of analyzing the interactive recognition signal is: The parsed interaction instructions are displayed to the user in a visual form, and user feedback is collected. If the user feedback is correct, interaction analysis is performed based on the interaction instructions and the corresponding interaction results are generated. Otherwise, an interaction correction analysis signal is generated.
[0008] As a further solution of the present invention, the specific method of obtaining the abnormal segmentation instruction is: Obtain interactive instructions, and obtain segmented instructions based on the turning points in the user's action trajectory, labeled i, where i=1, 2, ..., j, where j represents the number of segmented instructions. Then, match the obtained segmented instructions i with pre-stored instructions, and the pre-stored instructions are segmented instructions after training. Filter out segmented instructions that cannot be recognized and record them as abnormal segmented instructions.
[0009] As a further solution of the present invention, the specific method of generating the interaction result or the secondary interaction input signal by combining the user feedback is: The abnormal segmented instruction is labeled as n, and n=1, 2, ..., m, where m represents the number of abnormal segmented instructions. The similarity between the abnormal segmented instruction and all pre-stored instructions is calculated, and the pre-stored instruction with the largest similarity value is selected and recorded as the pre-selected instruction. The pre-selected instructions and the segmented instructions in the interaction instructions are recombined to obtain the recombined interaction instructions. Then, user feedback is obtained and it is judged whether the recombined interaction instructions are correct. If correct, the interaction is performed based on the instructions and the interaction results are generated. Otherwise, if it is wrong, a secondary interaction input signal is generated.
[0010] As a further solution of the present invention, the specific method of obtaining the combined instruction to be analyzed is: Obtain the secondary input interaction information and segment it to obtain secondary interaction segmented instructions and label them as a, and a=1, 2, ..., b, where b represents the number of secondary interaction segmented instructions. Then calculate the similarity between the secondary interaction segmented instructions a and the pre-stored instructions, and sort them from large to small according to the corresponding pre-stored instruction similarity. At the same time, obtain all the combination instruction information corresponding to the pre-stored instructions, and filter out the combination instructions to be analyzed that have an intersection with the secondary interaction segmented instructions.
[0011] As a further solution of the present invention, the specific method of generating the interaction result is: Obtain a combination instruction to be analyzed and label it as k, where k=1, 2, ..., p, and p is the type of the combination instruction to be analyzed. Match the pre-stored instructions in the combination instruction k to be analyzed with the secondary interactive segmented instruction a, and filter out the unmatched segmented instructions in the secondary interactive segmented instruction a. At the same time, calculate the similarity between the segmented instructions and the pre-stored instructions in the corresponding combination instruction k to be analyzed, and calculate the mean similarity. The combined instruction to be analyzed with the largest mean similarity value is selected as the standard to generate a standard combined instruction, and then the interaction is performed based on the standard to generate the corresponding interaction result.
[0012] A human-computer natural interaction system based on embodied intelligence, the system comprising: The interaction information recognition module is used to collect the user's interaction information, determine whether it can be recognized, generate an interaction recognition signal or a secondary interaction input signal, and transmit both simultaneously; The interaction recognition and analysis module is used to obtain and analyze interaction recognition signals, obtain error feedback corresponding to user feedback, analyze and generate interaction correction analysis signals, then segment the interaction instructions according to turning points to obtain segmented instructions, and match and filter with pre-stored instructions to obtain abnormal segmented instructions; The instructions with the maximum similarity between the abnormal segmented instructions and the pre-stored instructions are recorded as pre-selected instructions, and the interaction results or secondary interaction input signals are generated in combination with user feedback; The secondary interaction analysis module is used to analyze the acquired secondary interaction input signal, obtain the secondary input interaction information and segment it to obtain secondary interaction segmented instructions, calculate its similarity with pre-stored instructions, and obtain combined instruction information, and filter it with the secondary interaction segmented instructions to obtain the combined instruction to be analyzed; The unmatched instructions in the secondary interactive segmented instructions are obtained based on the combination instruction to be analyzed and recorded as segmented instructions. The similarity between the unmatched instructions and the pre-stored instructions in the combination instruction to be analyzed is calculated and the sum is used to calculate the average. Then, the combination instruction to be analyzed with the largest average is used as the standard for interaction, and the interaction results are generated and transmitted to the interaction information output module. The interaction information output module is used to display the generated interaction results to the corresponding interactive users.
[0013] The present invention provides a method and system for natural human-computer interaction based on embodied intelligence. Compared with the existing technology, it has the following advantages: The present invention supports personalized adjustment, improves the accuracy of instruction segmentation, and solves the ambiguity of semantic segmentation of continuous actions based on dynamic segmentation of motion trajectory turning points and voice pauses. It introduces a multi-layer matching algorithm of dynamic time warping + cosine similarity, gives priority to screening pre-stored instructions in the same position, supports combined instruction reorganization, improves the efficiency of abnormal instruction repair, reduces the burden of secondary input for users, and improves the fluency of interaction. At the same time, it constructs an "identification-verification-reorganization-execution" closed loop, selects the preferred combined instructions through similarity mean evaluation, supports incremental learning, and improves system adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a diagram of the steps and methods of the present invention; Figure 2 This is a system block diagram of the present invention. DETAILED DESCRIPTION
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0016] Example 1: Please refer to Figure 1 , the present application provides a human-computer natural interaction method based on embodied intelligence, the method specifically comprising the following steps: Step S1: Collect multi-source sensor data from the user. Here, the obtained multi-source sensor data is cleaned and fused accordingly. Specifically, methods such as Kalman filtering and wavelet transform are used to eliminate sensor drift (such as camera jitter and electromagnetic interference). Then, multimodal data is synchronized through timestamps (error ≤ 5ms), and a unified coordinate mapping is established in space (such as projecting gesture video and depth data into the same three-dimensional space). Interaction information is identified and interaction recognition is performed on the interaction information. If recognition is successful, an interaction recognition signal is generated. Otherwise, a secondary interaction input signal is generated.
[0017] Step S2: Analyze the obtained interaction recognition signal and obtain the corresponding interaction instruction. Feedback the interaction instruction to the corresponding user and make a judgment based on the user's feedback. Here, the user feedback specifically responds to the correctness of the interaction instruction. If the interaction instruction is correct after feedback, it means that the user's feedback is correct. Then, interaction analysis is performed based on the interaction instruction, and a corresponding interaction result is generated. The interaction result is displayed to the corresponding user. On the contrary, if the interaction instruction is incorrect after feedback, it means that the user's feedback is wrong, and an interaction correction analysis signal is generated. The obtained interaction correction analysis signal is processed to obtain interaction instructions, which are then segmented. The segmentation here is based on the pause points in the interaction instructions. The specific pause points are represented by the turning points of the user's movement trajectory. For example, when shaking the head, the head will swing left and right, and turning points will be generated during the swing. The turning points are recorded as pause points. For example, if the user's continuous action is "drawing a circle with the left hand → pointing to the window → waving", the system will divide it into three segments according to the turning points: Instruction 1: Draw a circle with your left hand (possibly meaning "loop" or "adjust"); Command 2: Point to the window (the positioning target is the window); Command 3: Wave (possibly meaning "off" or "stop").
[0018] The obtained segmented instruction is labeled as i, and i=1, 2, ..., j, where j represents the number of segmented instructions. The obtained segmented instruction i is then matched with the pre-stored instructions, and the pre-stored instructions here are also the segmented instructions. Specifically, the pre-stored instructions are the trained instructions stored by the operator, and abnormal segmented instructions are screened out. The abnormal segmented instructions here are segmented instructions that cannot be recognized by the pre-stored instructions. Specifically, dynamic time warping (DTW) or cosine similarity is used to calculate the matching degree between the segmented instructions and the pre-stored template. For example, the user instruction "draw a circle with your left hand → point to the window" matches the second segment of the pre-stored instruction "raise your hand → point to the window" with a matching degree of 90%, but the first segment does not match and is marked as an abnormal segment. Segmented instructions with a specific matching degree lower than a threshold (such as 70%) are judged to be abnormal; For example, a user in a smart home environment says "turn off" and quickly points to a lamp. The system initially interprets it as a "turn off lamp" command and asks the user, who shakes his head to indicate an error. The system generates correction signals and segments the user's action "quickly pointing → shaking head" into turning points; The first segment "Quick Pointing" matches the pre-stored instruction "Pointing to the Lighting Fixture", but the second segment "Shaking Head" does not match and is marked as an exception; The system combines historical data to infer that the user may have made an incorrect operation and prompts: "Do you want to turn on the light?" The user nods to confirm and then executes the correct command.
[0019] Step S3, obtain all abnormal segmented instructions and label them as n, and n=1, 2, ..., m, where m represents the number of abnormal segmented instructions, then calculate the similarity between the abnormal segmented instructions and all pre-stored instructions, and select the pre-stored instruction with the largest similarity value, recorded as the pre-selected instruction, and the pre-selected instruction here is represented as belonging to the same motion position as the abnormal segmented instruction, and at the same time, recombine the obtained pre-selected instruction with the segmented instruction to obtain a reorganized interaction instruction, then obtain user feedback, and determine whether the reorganized interaction instruction is correct. If correct, use it as a standard for interaction to generate an interaction result. Otherwise, if it is wrong, generate a secondary interaction input signal.
[0020] The overlap of motion trajectories (e.g., hand motion path similarity) is calculated using the coordinates of key skeletal points. The dynamic time warping (DTW) algorithm is then used to match the temporal rhythm of the motions, prioritizing the selection of pre-stored instructions with the same starting position as the abnormal instruction. For example, if the user's abnormal motion is "swinging the right hand from the left to the right," only those action templates starting with "right hand from the left" are selected from the pre-stored instruction library. A similarity threshold is set (e.g., ≥0.7). If the similarity is below this threshold, no pre-selected instructions are generated, and the secondary interaction proceeds directly. Replace the original abnormal segment with the preselected instruction and keep other correct segments. For example: The original segmented command sequence: [wave → point to TV → clench fist] (where "clench fist" is the abnormal segment); Pre-selected command: [Click action] (highest similarity, and the starting position is close to "clenching fist"); Reorganize the instructions: [wave → point to the TV → click action].
[0021] Step S4: Analyze the generated secondary interactive input signal to obtain secondary input interactive information, and perform segmentation processing on it. The segmentation method here is the same as the processing method in step S2. The secondary interactive segmented instructions are obtained and labeled as a, and a=1, 2, ..., b, where b represents the number of secondary interactive segmented instructions. Then calculate the similarity between the secondary interactive segmented instructions a and the pre-stored instructions, and sort them from large to small according to the corresponding pre-stored instruction similarity. Here, first, the pre-stored instructions with a similarity greater than 70% are screened according to the similarity, and then the pre-stored instructions obtained by the screening are sorted from large to small. Specifically, the pre-stored instructions corresponding to all secondary interactive segmented instructions are sorted. For example, if the secondary interactive segmented instruction a=1, then first calculate its similarity with the pre-stored instructions, and screen the pre-stored instructions with a similarity greater than 70%, and then sort the screened pre-stored instructions from large to small according to the similarity. At the same time, obtain all combined instruction information corresponding to the pre-stored instructions, and the combined instruction information here represents a complete interactive instruction that can be recognized by the interactive system; Then, the secondary interactive segmented instruction a is screened based on the combination instruction information to obtain the combination instruction to be analyzed, and the screening method here is based on the secondary interactive segmented instruction a as the standard. If the secondary interactive segmented instruction exists in the combination instruction information, it is recorded as the combination instruction to be analyzed. Otherwise, if it does not exist, it is not selected and is labeled k, and k=1, 2, ..., p, where p represents the type of the combination instruction to be analyzed. Then, the pre-stored instructions in the combination instruction k to be analyzed are matched with the secondary interactive segmented instruction a, and the unmatched segmented instructions in the secondary interactive segmented instruction a are screened. The segmented instructions here represent the secondary interactive segmented instructions that are different from the pre-stored instructions. At the same time, the similarity between the segmented instructions and the pre-stored instructions in the corresponding combination instruction k to be analyzed is calculated, and the mean of the similarities is calculated. Specifically, all the segmented instructions in the combination instruction to be analyzed are obtained, and the similarity between the segmented instructions and their corresponding pre-stored instructions is calculated. Then, the mean of all similarities is calculated. For example, constructing a quadratic interactive segmented instruction matrix M a×n , (a is the number of segments, n is the number of pre-stored instructions), calculate the matching matrix P of each combination instruction k k , element p u,o Represents the similarity between the u-th segment and the o-th pre-stored instruction, where the positions of u and o correspond to each other. Calculate the mean similarity of the combination instruction k.
[0022] Similarly, the mean similarity of the segmented instructions of all the combination instructions k to be analyzed is calculated, and the combination instruction to be analyzed with the largest mean similarity is selected as the standard to generate the standard combination instruction, and then the interaction is performed based on it as the standard to generate the corresponding interaction result.
[0023] For example, if a user makes an unusual interaction in a smart home environment, the system triggers a secondary interaction process as follows: User input: voice "Turn down...this" + gesture (left hand swipe down → point to the air conditioner); Segmentation results: a=1 (voice "lower"), a=2 (voice "this" + gesture).
[0024] Segment a=1 matches the pre-stored commands: "lower the temperature" (similarity 0.85), "lower the volume" (similarity 0.72); Segment a=2 matches the pre-stored commands: "point to the air conditioner" (similarity 0.91), "point to the light" (similarity 0.68, filtered); Combination commands to be analyzed: k=1: [lower temperature → point to air conditioner], similarity average 0.88; k=2: [lower volume → point to air conditioner], similarity average 0.77; Select k=1 as the standard combination instruction.
[0025] Example 2: Please refer to Figure 2, the present application provides a human-computer natural interaction system based on embodied intelligence, the system includes an interaction information recognition module, an interaction recognition analysis module, a secondary interaction analysis module and an interaction information output module; The interaction information recognition module is used to collect the user's interaction information, determine whether it can be recognized, generate an interaction recognition signal or a secondary interaction input signal, and transmit both simultaneously. The specific processing method is the same as the processing process of step S1 in embodiment 1; The interaction recognition and analysis module is used to obtain and analyze the interaction recognition signal, obtain error feedback corresponding to the user feedback, analyze and generate an interaction correction analysis signal, then segment the interaction instruction according to the turning point to obtain segmented instructions, and match and filter with pre-stored instructions to obtain abnormal segmented instructions. The specific processing method is similar to the processing process of step S2 in Example 1; The instruction with the maximum similarity between the abnormal segmented instruction and the pre-stored instruction is recorded as the pre-selected instruction, and the interaction result or the secondary interaction input signal is generated in combination with the user feedback. The specific processing method is the same as the processing process of step S3 in embodiment 1; The secondary interaction analysis module is used to analyze the acquired secondary interaction input signal, obtain the secondary input interaction information and segment it to obtain secondary interaction segmented instructions, calculate its similarity with pre-stored instructions, and obtain combined instruction information, and filter it with the secondary interaction segmented instructions to obtain the combined instruction to be analyzed; The unmatched instructions in the secondary interactive segmented instructions are obtained based on the combination instruction to be analyzed as the standard and recorded as segmented instructions. The similarity between the unmatched instructions and the pre-stored instructions in the combination instruction to be analyzed is calculated, and the sum is calculated to calculate the average. Then, the combination instruction to be analyzed with the largest average is used as the standard for interaction, and the interaction result is generated and transmitted to the interaction information output module. The specific processing method is similar to the processing process of step S4 in Example 1. The interaction information output module is used to display the generated interaction results to the corresponding interactive users.
[0026] Some of the data in the above formulas are calculated based on their numerical values and are not substituted into parameter units for calculation. At the same time, the contents not described in detail in this specification belong to the existing technology known to those skilled in the art.
[0027] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A natural human-computer interaction method based on embodied intelligence, characterized in that: The method specifically comprises the following steps: Collect user interaction information, determine whether it can be recognized, and generate an interaction recognition signal or a secondary interaction input signal; Analyze the interaction recognition signal, obtain the user's error feedback to generate the interaction correction analysis signal, and then segment the interaction instructions according to the turning points and match them with the pre-stored instructions to obtain the abnormal segmented instructions; The instructions with the maximum similarity between the abnormal segmented instructions and the pre-stored instructions are recorded as pre-selected instructions, and the interaction results or secondary interaction input signals are generated in combination with user feedback; Analyze the secondary interactive input signal, segment the secondary input interactive information to obtain secondary interactive segment instructions, and match and filter with pre-stored combination instruction information to obtain the combination instruction to be analyzed; Obtain the segmented instructions that do not match the secondary interaction segmented instructions in the combination instructions to be analyzed, calculate their similarity with the corresponding pre-stored instructions, and sum up the similarity mean. Then, interact with the combination to be analyzed with the maximum similarity mean as the standard to generate the interaction result.
2. The method of natural human-computer interaction based on embodied intelligence according to claim 1, characterized in that: The specific method of analyzing the interactive recognition signal is as follows: The parsed interaction instructions are displayed to the user in a visual form, and user feedback is collected. If the user feedback is correct, interaction analysis is performed based on the interaction instructions and the corresponding interaction results are generated. Otherwise, an interaction correction analysis signal is generated.
3. The method of natural human-computer interaction based on embodied intelligence according to claim 1, characterized in that: The specific method of obtaining the abnormal segmentation instruction is: Obtain interactive instructions, and obtain segmented instructions based on the turning points in the user's action trajectory, labeled i, where i=1, 2, ..., j, where j represents the number of segmented instructions. Then, match the obtained segmented instructions i with pre-stored instructions, and the pre-stored instructions are segmented instructions after training. Filter out segmented instructions that cannot be recognized and record them as abnormal segmented instructions.
4. The method of natural human-computer interaction based on embodied intelligence according to claim 1, characterized in that: The specific method of generating the interaction result or the secondary interaction input signal by combining the user feedback is: The abnormal segmented instruction is labeled as n, and n=1, 2, ..., m, where m represents the number of abnormal segmented instructions. The similarity between the abnormal segmented instruction and all pre-stored instructions is calculated, and the pre-stored instruction with the largest similarity value is selected and recorded as the pre-selected instruction. The pre-selected instructions and the segmented instructions in the interaction instructions are recombined to obtain the recombined interaction instructions. Then, user feedback is obtained and it is judged whether the recombined interaction instructions are correct. If correct, the interaction is performed based on the instructions and the interaction results are generated. Otherwise, if it is wrong, a secondary interaction input signal is generated.
5. The method of natural human-computer interaction based on embodied intelligence according to claim 1, characterized in that: The specific method of obtaining the combined instruction to be analyzed is: Obtain the secondary input interaction information and segment it to obtain secondary interaction segmented instructions and label them as a, and a=1, 2, ..., b, where b represents the number of secondary interaction segmented instructions. Then calculate the similarity between the secondary interaction segmented instructions a and the pre-stored instructions, and sort them from large to small according to the corresponding pre-stored instruction similarity. At the same time, obtain all the combination instruction information corresponding to the pre-stored instructions, and filter out the combination instructions to be analyzed that have an intersection with the secondary interaction segmented instructions.
6. The method of natural human-computer interaction based on embodied intelligence according to claim 1, characterized in that: The specific method of generating the interaction result is: Obtain a combination instruction to be analyzed and label it as k, where k=1, 2, ..., p, and p is the type of the combination instruction to be analyzed. Match the pre-stored instructions in the combination instruction k to be analyzed with the secondary interactive segmented instruction a, and filter out the unmatched segmented instructions in the secondary interactive segmented instruction a. At the same time, calculate the similarity between the segmented instructions and the pre-stored instructions in the corresponding combination instruction k to be analyzed, and calculate the mean similarity. The combined instruction to be analyzed with the largest mean similarity value is selected as the standard to generate a standard combined instruction, and then the interaction is performed based on the standard to generate the corresponding interaction result.
7. A human-computer natural interaction system based on embodied intelligence, used to execute the human-computer natural interaction method according to any one of claims 1 to 6, characterized in that: The system includes: The interaction information recognition module is used to collect the user's interaction information, determine whether it can be recognized, generate an interaction recognition signal or a secondary interaction input signal, and transmit both simultaneously; The interaction recognition and analysis module is used to obtain and analyze interaction recognition signals, obtain error feedback corresponding to user feedback, analyze and generate interaction correction analysis signals, then segment the interaction instructions according to turning points to obtain segmented instructions, and match and filter with pre-stored instructions to obtain abnormal segmented instructions; The instructions with the maximum similarity between the abnormal segmented instructions and the pre-stored instructions are recorded as pre-selected instructions, and the interaction results or secondary interaction input signals are generated in combination with user feedback; The secondary interaction analysis module is used to analyze the acquired secondary interaction input signal, obtain the secondary input interaction information and segment it to obtain secondary interaction segmented instructions, calculate its similarity with pre-stored instructions, and obtain combined instruction information, and filter it with the secondary interaction segmented instructions to obtain the combined instruction to be analyzed; The unmatched instructions in the secondary interactive segmented instructions are obtained based on the combination instruction to be analyzed and recorded as segmented instructions. The similarity between the unmatched instructions and the pre-stored instructions in the combination instruction to be analyzed is calculated and the sum is used to calculate the average. Then, the combination instruction to be analyzed with the largest average is used as the standard for interaction, and the interaction results are generated and transmitted to the interaction information output module. The interaction information output module is used to display the generated interaction results to the corresponding interactive users.
Citation Information
Patent Citations
Smart home multi-mode man-machine natural interaction system and method thereof
CN112462940A
Extensible audio recognition method based on man-machine interaction
CN101923857A
Intelligent conversation device and feedback intelligent voice control system and method
CN107220292A
Interaction method and system of pension service robot
CN117456995A
Interaction control method and system based on intelligent ring
CN118708068A