Multi-modal intelligent interaction optimization method and system

By combining speech, vision and tactile perception and introducing a confidence evaluation mechanism, adjusting the weight of each perception in the decision model, the shortcomings of information reliability judgment in multimodal intelligent device systems are solved, and the accuracy and reliability of task execution are improved.

CN119937787AInactive Publication Date: 2025-05-06BEIJING SANSHILIUXING TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510006956.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing multimodal intelligent device systems lack effective judgment on the reliability of multimodal information, resulting in limited application potential in complex and dynamic environments.

Method used

By combining speech, vision, and tactile perception and introducing confidence evaluation mechanisms, adjusting the weight of each perception in the decision model, ensuring the accuracy and reliability of task execution strategies.

Benefits of technology

Comprehensive evaluation and dynamic adjustment of multimodal information are realized, and the task execution capabilities and reliability of smart devices in complex environments are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937787A_ABST
    Figure CN119937787A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal intelligent interaction optimization method and system, and the method comprises the steps: obtaining text instruction data according to a voice instruction of a user, and calculating the voice confidence; identifying and positioning an object in an image shot by the camera to obtain visual data, and calculating visual confidence; tactile data of the intelligent device and the contacted object are obtained according to the tactile sensor, and the tactile confidence coefficient is calculated; and then, adjusting the weight of each perception in a preset decision model according to the confidence of the three perceptions. And finally, inputting the text instruction data, the visual data and the tactile data into the decision model after weight adjustment to obtain a task execution strategy. According to the method and the system disclosed by the invention, voice, vision and touch perception can be combined, meanwhile, a confidence evaluation mechanism is introduced, multiple verification is performed on task execution, and dynamic adjustment is ensured to be made according to accuracy and reliability of different perception data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing, and specifically relates to a multimodal intelligent interaction optimization method and system. Background Art

[0002] The current interactive systems of smart devices mostly rely on a single perception modality, such as speech recognition or visual perception or tactile feedback. These single-modality technologies can work effectively in specific scenarios, but their limitations are particularly evident when facing complex environments. Systems with a single perception modality usually lack comprehensive environmental understanding capabilities and cannot integrate multiple information sources to accurately judge user intentions, thereby limiting the intelligence level and task execution capabilities of the device. For example, although speech recognition technology can convert user voice commands into text, its recognition accuracy may drop significantly in the presence of background noise, accent differences, or fuzzy speech. Similarly, visual perception can obtain environmental information through image recognition and scene analysis, but it is very sensitive to factors such as light and obstructions, and it is difficult to ensure stability in dynamic or complex scenes. Although tactile feedback can provide information about physical contact, it is difficult to independently judge the overall environment of task execution or user intentions. Therefore, the limitations of a single modality directly lead to the lack of capabilities of smart devices in multi-task collaboration, real-time adjustment, and adaptation to complex environments.

[0003] Multimodal perception technology provides smart devices with a more comprehensive environmental understanding capability by integrating multiple sensory information such as voice, vision, and touch. However, current multimodal information fusion technology still has significant problems, the most important of which is the lack of effective judgment ability on information reliability. In a multimodal system, the quality of input data of various sensory modalities may vary significantly. If the system directly uses low-quality or unreliable data to make decisions, it may lead to incorrect execution of tasks and even cause adverse consequences. Therefore, the lack of information quality judgment ability in current technology not only reduces the reliability of multimodal fusion, but also limits the application potential of smart devices in complex and dynamic environments. Summary of the invention

[0004] The present invention provides a multimodal intelligent interaction optimization method and system to solve the problem in the prior art of lacking effective judgment on the reliability of multimodal information.

[0005] In order to solve the above technical problems, the embodiments of the present invention disclose the following technical solutions:

[0006] One aspect of the present invention provides a multimodal intelligent interaction optimization method, which is applied to an intelligent device having a camera and a tactile sensor, comprising:

[0007] Obtaining text instruction data according to the user's voice instruction and calculating the voice confidence;

[0008] Identify and locate objects in the images taken by the camera, obtain visual data, and calculate visual confidence;

[0009] Acquire tactile data of the smart device and the contacted object according to the tactile sensor, and calculate the tactile confidence;

[0010] Adjust the weight of each perception in the preset decision model according to the confidence of the above three perceptions;

[0011] The text instruction data, visual data and tactile data are input into the decision model after adjusting the weights to obtain a task execution strategy, which at least includes a moving path of the smart device and an intensity value of the force applied to the object.

[0012] Optionally, obtaining text instruction data according to the user's voice instruction and calculating the voice confidence includes:

[0013] Use a preset speech recognition algorithm to convert voice commands into text;

[0014] Processing the text based on a preset natural language understanding model, and obtaining text instruction data according to a semantic analysis result output by the model, wherein the semantic analysis result includes at least one word and a corresponding probability value;

[0015] The speech confidence is calculated based on the probability value of each word in the semantic analysis result.

[0016] Optionally, the identifying and locating the object in the image captured by the camera, obtaining visual data, and calculating visual confidence includes:

[0017] The image is processed using a preset computer vision algorithm, and the category, position and shape of each object in the image are determined based on the algorithm results as visual data;

[0018] Get the probability values ​​of each object category, position and shape in the algorithm results respectively;

[0019] Calculate the visual confidence based on all the above probability values.

[0020] Optionally, acquiring tactile data of the smart device and the contacted object according to the tactile sensor and calculating the tactile confidence includes:

[0021] Acquire the pressure value between the smart device and the contacted object and the acceleration during the contact process according to the output result of the sensor as the tactile data;

[0022] The tactile confidence is calculated according to the difference between the measured value of each tactile data and the preset value, and the change of each tactile data after a preset time interval.

[0023] Optionally, adjusting the weight of each perception in the preset decision model according to the confidence of the above three perceptions includes:

[0024] The weight of each perception in the decision model is calculated according to the voice confidence, visual confidence and tactile confidence, so that the weight of each perception is proportional to the confidence; the input data of the decision model includes text instruction data, visual data and tactile data, and the output data is the task execution strategy.

[0025] Optionally, before the smart device adopts the strategy to perform the task, the method further includes:

[0026] Determine whether there is perception data with a confidence level less than a preset lower limit.

[0027] If yes, send the task execution strategy to the user, and after receiving the user's confirmation information, control the smart device to execute the task according to the strategy; or send a message to the user to re-issue the voice command.

[0028] Optionally, before the smart device adopts the strategy to perform the task, the method further includes:

[0029] Determine whether the text instruction data contains the object category name,

[0030] If so, take the object as the target object, obtain the category probability value and position probability value of the target object in the algorithm result, and when the category probability value or the position probability value is less than a preset threshold, send a message to the user to re-issue the voice command, or obtain the category probability value and position probability value of the target object in the new image and re-compare them with the preset threshold.

[0031] Optionally, the method further includes:

[0032] When the smart device executes the strategy, the strategy is adjusted according to the newly acquired visual data, including:

[0033] Tracking the dynamic position of the target object in the visual data and calculating the distance between the dynamic position and the end position in the path planning;

[0034] When the distance increases, a query message is sent to the user to confirm whether the current task execution strategy is correct.

[0035] Optionally, the method further includes:

[0036] When the smart device executes the strategy, adjusting the strategy according to the newly acquired tactile data includes:

[0037] Determine whether the text command data contains words related to strength or movement speed.

[0038] If so, searching a preset pressure-speed correspondence table for a target pressure value and / or a target acceleration corresponding to the force intensity and / or the movement speed;

[0039] Calculating a pressure difference between a pressure value in the tactile data and a target pressure value, and / or calculating an acceleration difference between an acceleration in the tactile data and a target acceleration;

[0040] When the pressure difference exceeds the first set difference, the pressure sum applied by the smart device to the target object is adjusted to the target pressure value, and / or, when the acceleration difference exceeds the second set difference, the acceleration sum applied by the smart device to the target object is adjusted to the target acceleration.

[0041] Another aspect of the present invention provides a multimodal intelligent interaction optimization system, which is applied to an intelligent device having a camera and a tactile sensor, comprising:

[0042] A voice confidence module is configured to obtain text instruction data according to the user's voice instruction and calculate the voice confidence;

[0043] A visual confidence module is configured to identify and locate objects in the image captured by the camera, obtain visual data, and calculate visual confidence;

[0044] A tactile confidence module is configured to obtain tactile data of the smart device and the contacted object according to the tactile sensor, and calculate the tactile confidence;

[0045] A model weight adjustment module, configured to adjust the weight of each perception in the preset decision model according to the confidence of the above three perceptions;

[0046] The task execution strategy acquisition module is configured to input text instruction data, visual data and tactile data into a decision model with adjusted weights to obtain a task execution strategy, which at least includes a moving path of the smart device and an intensity value of the force applied to the object.

[0047] The present invention discloses a multimodal intelligent interaction optimization method and system. First, text instruction data is obtained according to the user's voice instruction, and the voice confidence is calculated; the object in the image captured by the camera is identified and located, visual data is obtained, and visual confidence is calculated; the tactile data of the intelligent device and the contacted object is obtained according to the tactile sensor, and the tactile confidence is calculated. Then, the weight of each perception in the preset decision model is adjusted according to the confidence of the above three perceptions. Finally, the text instruction data, visual data and tactile data are input into the decision model with adjusted weights to obtain the task execution strategy. The method and system disclosed in the present invention can combine voice, vision and tactile perception. At the same time, a confidence evaluation mechanism is introduced to perform multiple verifications on task execution to ensure dynamic adjustments based on the accuracy and reliability of different perception data.

[0048] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.

[0050] Figure 1 A schematic diagram of a multi-modal intelligent interaction optimization method provided by an embodiment of the present invention;

[0051] Figure 2 An implementation provided for an embodiment of the present invention Figure 1 Schematic diagram of the process of step S100;

[0052] Figure 3 An implementation provided for an embodiment of the present invention Figure 1 Schematic diagram of the process of step S200;

[0053] Figure 4 An implementation provided for an embodiment of the present invention Figure 1 Schematic diagram of the process of step S300;

[0054] Figure 5 A schematic diagram of the structure of a multi-modal intelligent interactive optimization system provided by an embodiment of the present invention;

[0055] Figure 6 A schematic diagram of a combination of a multimodal intelligent interaction optimization method and system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0057] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0058] Figure 1 A schematic diagram of a multimodal intelligent interaction optimization method provided by an embodiment of the present invention is applied to an intelligent device with a camera and a tactile sensor, such as Figure 1 As shown, the method comprises the following steps:

[0059] Step S100: obtaining text instruction data according to the user's voice instruction and calculating the voice confidence.

[0060] In one embodiment disclosed in the present invention, Figure 2 As shown, the following sub-steps may be used to implement step S100:

[0061] Step S101: converting voice instructions into text using a preset voice recognition algorithm.

[0062] Using a pre-selected speech recognition model (such as DeepSpeech, Whisper, Kaldi, etc.), the voice command is mapped into a probability distribution, the phonemes at each time step are predicted, and decoded into a text sequence through connectionist temporal classification (CTC) or Transformer-based architecture.

[0063] In the embodiments disclosed in the present invention, a language model (such as BERT or N-gram) can be combined to improve the fluency and accuracy of text output, and Weiner filtering or convolutional neural network (CNN) can be used for noise reduction, as well as gain adjustment and speech separation for low-quality audio.

[0064] After the above process, a preliminary text result is generated, along with the confidence distribution of each phoneme.

[0065] Step S102: Process the text based on a preset natural language understanding model, and obtain text instruction data according to the semantic analysis result output by the model.

[0066] (1) Preprocess the text

[0067] Split the text into words or phrases and mark the parts of speech using word segmentation tools (such as NLTK, spaCy). Apply a spelling check algorithm to fix spelling errors that may be caused by speech recognition.

[0068] (2) Use natural language understanding models to process text

[0069] The preprocessed text is input into a preset natural language understanding model (such as BERT, GPT, RoBERTa) to obtain text semantic vectors and intent classification results. The main purpose of the user text (such as "play music", "turn on the light") is identified, and key entities (such as date, location, device name) are extracted from the text.

[0070] The probability value of each word is calculated through the softmax distribution of the model output. The semantic analysis results include the user's target action, the core words and attributes related to the instruction, and the probability value of each word or phrase of the model.

[0071] Step S103: Calculate the speech confidence according to the probability value of each word in the semantic analysis result.

[0072] Summarize the probability value of each word in the semantic analysis results and calculate the speech confidence:

[0073]

[0074] Among them, P(w i ) is the probability value of the i-th word; weight (w i ) is the weight value of the ith word, which is pre-assigned according to the importance of the word's part of speech, for example, verbs and nouns have higher weights than prepositions.

[0075] Different confidence thresholds are set according to the application scenario. If the confidence is lower than the threshold, the user is asked to confirm or repeat the input.

[0076] Step S200: Identify and locate objects in the image captured by the camera, obtain visual data, and calculate visual confidence.

[0077] In one embodiment disclosed in the present invention, Figure 3 As shown, the following sub-steps may be used to implement step S200:

[0078] Step S201: Use a preset computer vision algorithm to process the image, and determine the category, position and shape of each object in the image according to the algorithm result as visual data.

[0079] When processing images, a combination of deep learning and traditional computer vision can be used. First, the images captured by the camera are preprocessed, including noise reduction, normalization, and resolution adjustment, to adapt to the model input requirements. Subsequently, the object category and bounding box position in the image are identified through the target detection algorithm, that is, the computer vision algorithm (such as YOLO, Faster R-CNN, or SSD), and each bounding box is instance segmented to extract the specific shape information of the object.

[0080] During object detection, the model generates image features through feature extractors (such as ResNet and EfficientNet), predicts the category through the classification head, and predicts the position and size of the object through the regression head. For instance segmentation tasks, Mask R-CNN can be combined to accurately separate shape information. The output results include the category label, location coordinates (such as the upper left corner and lower right corner coordinates), and shape of each object.

[0081] Step S202: Obtain the probability value of each object category, position and shape in the algorithm result respectively.

[0082] The category probability value is the predicted score for an object belonging to a certain category, and the category corresponding to the maximum value is selected as the output; the position probability value is the overlap between the predicted box output by the regression model and the real box (such as IoU, intersection over union ratio), and the higher the value, the more accurate the positioning; the shape probability value is calculated by the pixel overlap between the shape segmentation area and the real area (such as F1 score). These probability values ​​reflect the reliability of the model in different dimensions.

[0083] The above probability values ​​can all be obtained from the algorithm results.

[0084] Step S203: Calculate the visual confidence according to all the above probability values.

[0085] In the embodiments disclosed in the present invention, a weighted average or fusion strategy can be used to calculate visual confidence. The probability values ​​of category, position and shape are weighted according to their importance, such as category and position with higher weights and shape with slightly lower weights, and then the weighted probability values ​​are comprehensively calculated. For example, the formula of visual confidence C can be expressed as:

[0086] C=w1·P class +w2·P position +w3·P shape

[0087] Among them, w1 is the preset category weight, w2 is the preset position weight, w3 is the preset shape weight, P class is the probability value of the category, P position is the probability value of the position, P shape is the probability value of the shape, and w1+w2+w3=1. This method can dynamically adjust the importance of various factors in different scenarios to improve the robustness of the system. If the combined result of all confidence levels is higher than the preset threshold, the visual data is considered credible and can be used as input for subsequent operations; if it is lower than the threshold, the system can be triggered to perform secondary analysis or prompt the user to intervene.

[0088] Step S300: Acquire tactile data of the smart device and the contacted object according to the tactile sensor, and calculate the tactile confidence.

[0089] In one embodiment disclosed in the present invention, Figure 4 As shown, the following sub-steps may be used to implement step S300:

[0090] Step S301: Obtain the pressure value between the smart device and the contacted object and the acceleration during the contact process as tactile data according to the output result of the sensor.

[0091] The tactile sensors of smart devices can collect pressure values ​​and acceleration data generated during the contact process. The sensors are installed on the surface of the device and can sense the pressure changes under the action of external forces in real time and convert them into digital signals to represent the pressure values. At the same time, the accelerometer generates acceleration data by detecting the changes in the motion state of the device during the contact process (such as vibration, sliding or collision during touch). These data are recorded in the form of time series to ensure that the dynamic characteristics of the contact behavior can be captured. To improve accuracy, the collected data needs to be filtered and denoised, such as using Kalman filtering or bandpass filtering to remove environmental interference and retain the main components of the tactile signal.

[0092] Step S302: Calculate the tactile confidence according to the difference between the measured value of each tactile data and the preset value, and the change of each tactile data after a preset time interval.

[0093] The calculation of tactile confidence is based on the acquired tactile data, and is judged by comparing the difference between the measured value and the preset value, and analyzing the trend of the tactile data over time. First, the pressure value is compared with the preset reference pressure range. If the measured value falls within the reference range, the tactile data is preliminarily considered to be credible; otherwise, the confidence is reduced. Secondly, the change in acceleration during the contact process is also used as an evaluation basis. For example, if the acceleration signal is stable and conforms to the expected contact mode (such as light touch, sliding), the confidence is increased; if the acceleration signal fluctuates violently, it may indicate that the contact is unstable or abnormal, and the confidence is reduced. In addition, by comparing the changes in the same tactile parameter after a preset time length, such as the rate of increase or decrease of pressure, and the smoothness of the acceleration curve, the reliability of the contact can be further confirmed.

[0094] The calculation of tactile confidence can be completed by weighted fusion of pressure value and acceleration confidence. For example, let pressure confidence be P pressure , the acceleration confidence is P acceleration , the overall tactile confidence C can be expressed as:

[0095] C=w1·P prseeure +w2·P acceleration

[0096] Among them, w1 and w2 are pressure value and acceleration weight respectively, and the specific values ​​can be adjusted according to the actual scenario. If the comprehensive result is higher than the set threshold, the tactile data is considered reliable, otherwise it may trigger a mechanism to re-collect or adjust the device contact behavior. In this way, smart devices can effectively evaluate the reliability of tactile data and ensure the accuracy and safety of subsequent operations.

[0097] Step S400: adjusting the weight of each perception in the preset decision model according to the confidence of the above three perceptions.

[0098] In one embodiment disclosed in the present invention, step S400 may be implemented in the following manner:

[0099] The weight of each perception in the decision model is calculated according to the speech confidence, visual confidence and tactile confidence, so that the weight of each perception is proportional to the confidence. The input data of the decision model includes text instruction data, visual data and tactile data, and the output data is the task execution strategy.

[0100] First, the calculation results of speech confidence, visual confidence and tactile confidence are obtained, which represent the reliability of each sensory input. For example, assuming that the speech confidence is C speech , visual confidence is C vision , tactile confidence is C tactile, these three confidence levels are values ​​between 0 and 1, and the closer they are to 1, the higher the credibility of the corresponding perception.

[0101] Then, the weight of each perception is calculated according to the confidence value, so that the weight is proportional to the confidence. The method for calculating the weight is as follows:

[0102] The three confidence values ​​are normalized to the ratio of the weight values ​​using the following normalization formula:

[0103]

[0104] Among them, w i is the weight of speech, vision or touch in the decision model, satisfying that the sum of the weights of the three is equal to 1, and the weight ratio of each perception reflects the contribution of its confidence in the whole.

[0105] In the embodiments disclosed in the present invention, the weights are dynamically adjusted as the perceived confidence changes. For example, when the visual confidence is high, the visual weight will account for a larger proportion, thereby giving it a higher importance in the decision model.

[0106] Then, the adjusted weights are applied to the decision model. The input data of the decision model includes: text instruction data (results from speech recognition and natural language processing), visual data (results from image recognition and positioning), and tactile data (perception results such as pressure value and acceleration).

[0107] Finally, based on the fused data, the decision model outputs a task execution strategy. When the voice confidence and visual confidence are high, the decision model may prioritize the instructions provided by voice and vision; when the tactile confidence is high, the decision model may rely more on tactile perception to judge the external environment.

[0108] Step S500: input text instruction data, visual data and tactile data into the decision model after adjusting the weights to obtain a task execution strategy.

[0109] The task execution strategy at least includes the moving path of the smart device and the intensity value of the force applied to the object.

[0110] The text instruction data, visual data and tactile data are input into the decision model with adjusted weights for comprehensive analysis and calculation. The text instruction data is parsed through the voice recognition and natural language processing modules to analyze the user's specific needs. For example, the instruction may contain descriptions such as "move to the side of the table and gently push the chair away"; the visual data is captured by the camera and processed by the computer vision algorithm to provide key information such as the category, position and shape of objects in the environment, such as identifying the positional relationship between the table and the chair and obstacles in the environment; the tactile data is monitored in real time by the sensors of the smart device, and the feedback of the pressure value or motion acceleration when the device contacts the target object is used to optimize the accuracy and safety of the operation.

[0111] In the process of generating task execution strategies, the decision model will analyze based on the weight distribution of different perception data. The weight setting is usually based on the confidence of each perception data, and data with higher confidence will play a greater role in the generation of the final strategy. Specifically, visual data provides the spatial position of the target object for path planning; text instruction data parses the user's required operation method, such as "nudge" will be converted into a specific target force value; tactile data provides real-time feedback to ensure that the applied force and movement state meet the target requirements.

[0112] The decision model first calculates the optimal movement path of the smart device through a path planning algorithm (such as the Dijkstra algorithm) combined with the target position and environmental information in the visual data. This path will avoid obstacles and shorten the distance as much as possible to improve task efficiency. At the same time, the model uses tactile data and the operation description in the text instruction to determine the intensity value of the force applied. For example, if the instruction requires a "gentle push", the model will look for a preset force intensity range and correct the device force in real time through tactile feedback.

[0113] The resulting task execution strategy is output in a structured form, including the device's movement path and operating parameters. The task execution strategy is sent directly to the control module of the smart device, driving the device to move along the planned path and perform the corresponding operation.

[0114] In one embodiment disclosed in the present invention, before the smart device adopts a strategy to execute a task, the following steps need to be performed:

[0115] (1) Determine whether there is perception data with a confidence level less than a preset lower limit. If so, send a task execution strategy to the user, and after receiving the user's confirmation information, control the smart device to execute the task according to the strategy; or send a message to the user to re-issue the voice command.

[0116] In the embodiments disclosed in the present invention, the confidence of the three types of perception data needs to be verified to ensure the reliability of the strategy. If the confidence of a certain perception data is detected to be lower than the preset lower limit, the user interaction mechanism will be triggered to prompt the user to check the policy content and make confirmation. The task strategy is sent to the user through a user interface (such as a smart device screen or a mobile application), displaying detailed information such as the movement path and operation strength, and the user can confirm the strategy by touch, voice or other means. Once the user confirms, the device executes the task according to the strategy.

[0117] If the user denies the policy or does not confirm it, the system will prompt the user to re-enter a clear voice command, and then recalculate and generate the task policy. If the confidence of all perception data is higher than the lower limit, the device can directly execute the task without user confirmation.

[0118] In another embodiment disclosed by the present invention, before the smart device adopts a strategy to execute a task, the following steps need to be performed:

[0119] (1) Determine whether the text instruction data contains the name of the object category. If so, take the object as the target object, obtain the category probability value and position probability value of the target object in the algorithm result, and send a message to the user to re-issue the voice instruction when the category probability value or the position probability value is less than a preset threshold, or obtain the category probability value and position probability value of the target object in the new image and re-compare them with the preset threshold.

[0120] In order to further improve the accuracy of smart device task execution, in the embodiments disclosed in the present invention, it is necessary to perform specific verification and processing on the text instruction data and the perception results. First, the system extracts the user's instruction content from the text instruction data to determine whether it contains a clear object category name, such as "table", "cup" or "door". If the object category is clearly specified in the text instruction, the object category is used as the target object, and the verification process for the target object is started.

[0121] The system extracts the category probability value and position probability value of the target object through the output results of the computer vision algorithm, which respectively represent the degree of confidence in identifying the category of the object and the accuracy of locating its spatial position. These two probability values ​​will be compared with the preset thresholds to evaluate the reliability of the recognition results. If the category probability value or the position probability value is less than the threshold, it means that there may be uncertainty or error in the current perception data. For example, the target object may not be clearly identified due to insufficient ambient light, or the positioning information may be biased due to occlusion. In this case, the system will send a prompt message to the user through the user interface or voice interaction module, requiring the user to re-issue the voice command to provide a clearer task goal.

[0122] If the user does not want to re-enter the voice command, the system will adopt another solution, which is to obtain new image data to re-identify and locate the target. The system captures new environmental images through the camera and runs the visual algorithm again to obtain the latest category probability value and position probability value of the target object. Subsequently, these new data will be compared with the preset threshold. If the threshold requirements are met, the recognition result of the target object is confirmed to be valid and the system continues to perform the task; if the threshold requirements are still not met, the system will prompt the user again until sufficiently reliable data is obtained.

[0123] In one embodiment disclosed in the present invention, when the smart device executes the strategy, it is also necessary to adjust the strategy according to the newly acquired visual data, including the following steps:

[0124] (1) Track the dynamic position of the target object in the visual data and calculate the distance between the dynamic position and the end position in the path planning.

[0125] (2) When the distance increases, a query message is sent to the user to confirm whether the current task execution strategy is correct.

[0126] To ensure the accuracy and dynamic adaptability of smart devices in executing tasks, it is necessary to adjust the task execution strategy in real time during the task execution process. Vision data is continuously acquired through sensor devices such as cameras, and the dynamic position of the target object in the visual data is tracked in real time. The dynamic position of the target object is usually achieved through target detection and tracking algorithms, such as using KCF, SORT or deep learning models combined with optical flow methods to locate the current coordinates of the target object from continuous video frames.

[0127] During the tracking process, the system will compare the dynamic position of the target object with the pre-planned path endpoint position in the task execution strategy, and calculate the Euclidean distance or other suitable distance measurement between the two. If the position of the target object is offset, such as being moved by other objects or moving itself, causing the distance to increase, it may indicate that the task execution strategy needs to be adjusted. At this time, the system will trigger the user interaction mechanism and send an inquiry message to the user through voice broadcast, mobile notification or other means. The inquiry information includes the latest position of the current target object, the original endpoint position and the increase in the distance between them, and requests the user to confirm whether the task execution strategy is still applicable.

[0128] If the user confirms that the current strategy is still correct, the system will continue to execute the task according to the original strategy; if the user indicates that the strategy needs to be adjusted, the system will re-plan the path and update the task execution strategy based on the latest visual data. For example, when the target object is moved to a new location, the system can recalculate the movement route of the smart device through the path planning algorithm to ensure that the device can accurately reach the location of the target object and complete the task.

[0129] In one embodiment disclosed in the present invention, when the smart device executes the strategy, it is also necessary to adjust the strategy according to the newly acquired tactile data, including the following steps:

[0130] (1) Determine whether the text instruction data contains words related to force intensity or movement speed. If so, search for a target pressure value and / or target acceleration corresponding to the force intensity and / or movement speed in a preset pressure-speed correspondence table.

[0131] (2) Calculating a pressure difference between a pressure value in the tactile data and a target pressure value, and / or calculating an acceleration difference between an acceleration in the tactile data and a target acceleration.

[0132] (3) When the pressure difference exceeds a first set difference, the pressure sum applied by the smart device to the target object is adjusted to a target pressure value, and / or, when the acceleration difference exceeds a second set difference, the acceleration sum applied by the smart device to the target object is adjusted to a target acceleration.

[0133] In the embodiments disclosed in the present invention, it is also necessary to optimize and adjust the force intensity and movement speed according to the tactile data obtained in real time. First, the text instruction data is parsed to see whether it contains words related to force intensity or movement speed, such as "gentle push", "fast movement" or "hard press". If the instruction explicitly contains these descriptions, the system will use the preset pressure-speed correspondence table to find the target pressure value and / or target acceleration value that matches it according to the description in the instruction.

[0134] During the execution of the task, the tactile sensor continuously monitors the pressure value applied by the smart device on the target object and the acceleration value when the device moves. The system determines whether the current force applied by the device meets the instruction requirements by calculating the pressure difference between the actual pressure value and the target pressure value. At the same time, the system also calculates the acceleration difference between the actual acceleration and the target acceleration to evaluate the accuracy of the moving speed. If the pressure difference exceeds the first set difference (for example, the instruction requires a light push but the actual force is too large), the system will adjust the execution force mechanism (such as reducing the motor output power or adjusting the force feedback system) so that the actual pressure applied by the device gradually approaches the target pressure value. Similarly, when the acceleration difference exceeds the second set difference (for example, the instruction requires slow movement but the actual acceleration is high), the system will adjust the actual acceleration to the target acceleration by adjusting the motion control module (such as reducing the speed or increasing friction control).

[0135] Figure 5 A schematic diagram of the structure of a multimodal intelligent interactive optimization system provided by the present invention is applied to an intelligent device with a camera and a tactile sensor, such as Figure 5 As shown, the system includes the following modules:

[0136] A speech confidence module 1 is configured to obtain text instruction data according to a user's speech instruction and calculate speech confidence;

[0137] A visual confidence module 2 is configured to identify and locate objects in the image captured by the camera, obtain visual data, and calculate visual confidence;

[0138] A tactile confidence module 3 is configured to obtain tactile data of the smart device and the contacted object according to the tactile sensor, and calculate the tactile confidence;

[0139] A model weight adjustment module 4 is configured to adjust the weight of each perception in the preset decision model according to the confidence of the above three perceptions;

[0140] The task execution strategy acquisition module 5 is configured to input text instruction data, visual data and tactile data into the decision model with adjusted weights to obtain a task execution strategy, which at least includes the moving path of the smart device and the intensity value of the force applied to the object.

[0141] Figure 6 A schematic diagram of a method combined with a system provided in an embodiment of the present invention is used as a graphic illustration of the above embodiment.

[0142] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A multimodal intelligent interaction optimization method, applied to an intelligent device with a camera and a tactile sensor, characterized in that: include: Obtaining text instruction data according to the user's voice instruction and calculating the voice confidence; Identify and locate objects in the images taken by the camera, obtain visual data, and calculate visual confidence; Acquire tactile data of the smart device and the contacted object according to the tactile sensor, and calculate the tactile confidence; Adjust the weight of each perception in the preset decision model according to the confidence of the above three perceptions; The text instruction data, visual data and tactile data are input into the decision model after adjusting the weights to obtain a task execution strategy, which at least includes a moving path of the smart device and an intensity value of the force applied to the object.

2. The optimization method according to claim 1, characterized in that: The obtaining of text instruction data according to the user's voice instruction and calculating the voice confidence includes: Use a preset speech recognition algorithm to convert voice commands into text; Processing the text based on a preset natural language understanding model, and obtaining text instruction data according to a semantic analysis result output by the model, wherein the semantic analysis result includes at least one word and a corresponding probability value; The speech confidence is calculated based on the probability value of each word in the semantic analysis result.

3. The optimization method according to claim 1, characterized in that: The process of identifying and locating an object in an image captured by a camera, obtaining visual data, and calculating visual confidence includes: The image is processed using a preset computer vision algorithm, and the category, position and shape of each object in the image are determined based on the algorithm results as visual data; Get the probability values ​​of each object category, position and shape in the algorithm results respectively; Calculate the visual confidence based on all the above probability values.

4. The optimization method according to claim 1, characterized in that: The step of acquiring tactile data of the smart device and the contacted object according to the tactile sensor and calculating the tactile confidence includes: Acquire the pressure value between the smart device and the contacted object and the acceleration during the contact process according to the output result of the sensor as the tactile data; The tactile confidence is calculated according to the difference between the measured value of each tactile data and the preset value, and the change of each tactile data after a preset time interval.

5. The optimization method according to claim 1, characterized in that: The step of adjusting the weight of each perception in the preset decision model according to the confidence of the above three perceptions includes: The weight of each perception in the decision model is calculated according to the voice confidence, visual confidence and tactile confidence, so that the weight of each perception is proportional to the confidence; the input data of the decision model includes text instruction data, visual data and tactile data, and the output data is the task execution strategy.

6. The optimization method according to claim 1, characterized in that: Before the smart device adopts the strategy to perform the task, the method further includes: Determine whether there is perception data with a confidence level less than a preset lower limit. If yes, send the task execution strategy to the user, and after receiving the user's confirmation information, control the smart device to execute the task according to the strategy; or send a message to the user to re-issue the voice command.

7. The optimization method according to claim 3, characterized in that: Before the smart device adopts the strategy to perform the task, the method further includes: Determine whether the text instruction data contains the object category name, If so, take the object as the target object, obtain the category probability value and position probability value of the target object in the algorithm result, and when the category probability value or the position probability value is less than a preset threshold, send a message to the user to re-issue the voice command, or obtain the category probability value and position probability value of the target object in the new image and re-compare them with the preset threshold.

8. The optimization method according to claim 7, characterized in that: The method further comprises: When the smart device executes the strategy, the strategy is adjusted according to the newly acquired visual data, including: Tracking the dynamic position of the target object in the visual data and calculating the distance between the dynamic position and the end position in the path planning; When the distance increases, a query message is sent to the user to confirm whether the current task execution strategy is correct.

9. The optimization method according to claim 1, characterized in that: The method further comprises: When the smart device executes the strategy, adjusting the strategy according to the newly acquired tactile data includes: Determine whether the text command data contains words related to strength or movement speed. If so, searching a preset pressure-speed correspondence table for a target pressure value and / or a target acceleration corresponding to the force intensity and / or the movement speed; Calculating a pressure difference between a pressure value in the tactile data and a target pressure value, and / or calculating an acceleration difference between an acceleration in the tactile data and a target acceleration; When the pressure difference exceeds the first set difference, the pressure sum applied by the smart device to the target object is adjusted to the target pressure value, and / or, when the acceleration difference exceeds the second set difference, the acceleration sum applied by the smart device to the target object is adjusted to the target acceleration.

10. A multimodal intelligent interaction optimization system, applied to an intelligent device with a camera and a tactile sensor, characterized in that: include: A voice confidence module is configured to obtain text instruction data according to the user's voice instruction and calculate the voice confidence; A visual confidence module is configured to identify and locate objects in the image captured by the camera, obtain visual data, and calculate visual confidence; A tactile confidence module is configured to obtain tactile data of the smart device and the contacted object according to the tactile sensor, and calculate the tactile confidence; A model weight adjustment module, configured to adjust the weight of each perception in the preset decision model according to the confidence of the above three perceptions; The task execution strategy acquisition module is configured to input text instruction data, visual data and tactile data into a decision model with adjusted weights to obtain a task execution strategy, which at least includes a moving path of the smart device and an intensity value of the force applied to the object.

Citation Information

Patent Citations

  • Interest point identifying and positioning method, instruction processor and intelligent equipment

    CN114691073A

  • Active interaction system of multi-mode large model based on AR glasses

    CN118585071A

Cited By

  • Humanoid robot multi-modal data processing method, system and equipment and medium

    CN120687743A

  • A humanoid robot multi-modal data processing method, system, device and medium

    CN120687743B