Multi-modal context-based intention acquisition method and system
Through the multimodal context intention acquisition method, integrating position, gesture and voice information, and using the improved Markov random field model to evaluate massage intention, solving the problem of insufficient intelligence and interactive experience of existing massage robots, and improving the intelligence and safety of massage robots.
Patent Information
- Application Number
- CN202510318100.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
The existing massage robots have limited intelligence, lack real-time perception and dynamic adjustment capabilities, poor human-computer interaction experience, and safety and reliability need to be improved.
The multimodal context intent acquisition method is adopted to fuse position, gesture and voice information, and massage intention is evaluated through the improved Markov random field model, combining propagation update and weighted fusion algorithm to improve the accuracy of user intent extraction.
It achieves more accurate capture and response to user needs, improves the system's interactive effect and user experience, adapts to the physiological changes of the elderly, and enhances safety and reliability.
Smart Images

Figure CN120256892A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-modal intent acquisition, and specifically to a method and system for intent acquisition based on multi-modal context. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and robotics technologies, massage robots have gradually entered people's lives and are favored by more and more elderly people due to their convenient operation and diverse functions. Compared with traditional manual massage, massage robots can provide more precise force control, richer massage modes, and massage experiences at any time, effectively alleviating problems such as muscle soreness and fatigue of the elderly and improving their quality of life.
[0003] Currently, the massage robots on the market are mainly implemented by the following technologies:
[0004] 1. Manipulator and sensor technology: By using high-precision manipulators and sensors such as pressure and position sensors, simulating the massage actions of human hands, and realizing various massage techniques such as kneading, knocking, and pushing.
[0005] 2. Artificial intelligence algorithms: Using machine learning algorithms to analyze information such as users' body data and massage preferences, and providing personalized massage solutions.
[0006] 3. Speech recognition and interaction technology: Users can control the massage robot through voice commands to achieve a more convenient operation experience.
[0007] Although the technology of massage robots has developed rapidly, there are still some deficiencies:
[0008] 1. Limited intelligence level: Most of the existing massage robots rely on preset programs, lack the ability to perceive the user's physical condition in real time and make dynamic adjustments, and it is difficult to achieve "personalized" massage in the true sense.
[0009] 2. The human-computer interaction experience needs to be improved: The operation interfaces of some massage robots are complex, and the speech recognition accuracy is not high, which has a relatively high learning cost for the elderly.
[0010] 3. The safety and reliability still need to be strengthened: Massage robots are in direct contact with the human body, and their safety and reliability are crucial. Currently, the relevant technical standards and specifications are not yet perfect, and there are certain potential safety hazards.
[0011] Therefore, there is an urgent need for a method and system for intent acquisition based on multi-modal context to solve the above problems. Summary of the Invention
[0012] The object of the present invention is to provide a method and system for obtaining an intention based on multi-modal context, which improves the accuracy of user intention extraction and enables the system to more accurately capture and respond to user needs.
[0013] To achieve the above object, the present invention is realized through the following technical solutions:
[0014] On the one hand, a method for obtaining an intention based on multi-modal context is provided, including the following steps:
[0015] Initialize the position confidence of each acupoint of the user;
[0016] Construct a method that combines gesture recognition confidence and the association degree between gestures and acupoints to calculate gesture confidence;
[0017] Through the method of propagation and update, make the confidence of each acupoint reflect the information of the entire human acupoint structure;
[0018] Construct an improved Markov random field evaluation model to evaluate the massage intention.
[0019] Preferably, the initialization of the position confidence of each acupoint of the user includes:
[0020] The weighted distance Dis between the user-specified position and each acupoint, define the position distance as the comprehensive distance in the multi-dimensional space, and assign different weights to the distances in different dimensions according to the error sensitivity in different directions:
[0021]
[0022] Among them, (x i , y i , z i ) is the coordinate of acupoint i, (x0, y0, z0) is the coordinate of the user-specified position; w x , w y , w z are the weights in the corresponding directions respectively, reflecting the error sensitivity in this direction
[0023] Determine the weight of each direction by recording the number and amplitude of the user's adjustments in different directions and statistically analyzing the influence of the adjustments in different directions on the final massage effect;
[0024] The position confidence C pos (i) is initialized to the reciprocal of its distance:
[0025]
[0026] Preferably, the construction of the method that combines gesture recognition confidence and the association degree between gestures and acupoints includes:
[0027] Define the confidence of gesture recognition algorithm output gesture H as C gesture (H), and the correlation degree between the gesture and the acupoint;
[0028] According to the gesture intention table and the user's historical information, determine the correlation degree A gesture (H, i) between the gesture H and the acupoint i;
[0029] By calculating the correlation degree between each gesture and the acupoint, construct a correlation degree matrix A, where the rows in the matrix A represent different gestures H, the columns represent different acupoints i, and the matrix element A H,i represents the correlation degree between the gesture H and the acupoint i, and A H,i ∈[0, 1], 0 means completely uncorrelated, 1 means completely correlated:
[0030]
[0031] Among them, the number of joint occurrences of the gesture H and the acupoint i is N H,i , the number of times the gesture H appears in total is N H , combining the gesture recognition confidence and the correlation degree between the gesture and the acupoint, calculate the confidence of the gesture H corresponding to the acupoint i:
[0032] C gesture (H, i) = C gesture (H) · A gesture (H, i).
[0033] Preferably, the method of constructing based on the combination of gesture recognition confidence and the correlation degree between the gesture and the acupoint further includes:
[0034] Sum all the recognized gestures H, and calculate the comprehensive gesture confidence of each acupoint i:
[0035]
[0036] Calculate the voice confidence based on the result of voice recognition and the correlation degree with the acupoint:
[0037] C voice (V, i) = C voice (V) · A voice (V, i)
[0038] Perform weighted fusion on the confidences of position, gesture and voice to obtain the comprehensive confidence of each acupoint:
[0039] C(i) = w pos · C pos (i) + w gesture · ∑ H C gesture (H, i) + wvoice ·∑ V C vooce (V,i).
[0040] Preferably, the method of updating through propagation enables the confidence of each acupoint to reflect the information of the entire human acupoint structure, specifically as follows:
[0041] For each acupoint i, perform weighted averaging based on the confidence of its neighboring acupoints:
[0042]
[0043] where N(i) represents the set of neighbors of acupoint i, and w ij represents the weight between acupoint i and neighboring acupoint j;
[0044] Repeat the propagation update and normalization until the confidence of the acupoints converges, and take the acupoint with the highest confidence as the possible massage position selected by the user, that is, the multimodal context sub-intention set M.
[0045] Preferably, the construction of the improved Markov random field evaluation model includes:
[0046] The joint probability distribution of the improved Markov random field is:
[0047]
[0048] where P, G, V, T, H respectively represent position information, gesture information, voice information, time information, and historical information in the intelligent massage system, M represents the multimodal context intention, φ represents the potential function between the node and the massage intention, and Z is a normalization constant used to ensure the normalization of probabilities;
[0049] Calculate the conditional probability distribution of the massage intention M:
[0050] P(M|P,G,V,T,H) ∝ exp(-(ψ P (P,M) + ψ G (G,M) + ψ V (V,M) + ψ T (T,M) + ψ H (H,M)))
[0051] Define the information entropy H(E) and the recognition credibility C(E):
[0052]
[0053] The recognition credibility is expressed as the confidence of unimodal recognition:
[0054] C(E) = confidence(E)
[0055] The posterior probability of the massage intention is as follows:
[0056]
[0057] Preferably, the construction of the improved Markov random field evaluation model further includes:
[0058] Set θ(t) to represent the dynamic threshold:
[0059] θ(t) = α·θ0·f(t)
[0060] where θ0 is the initial threshold and f(t) is a time-related function
[0061] If P(M) ≥ θ(t), then M is taken as the executable intention E; if P(M) ≤ θ(t), then the true intention M does not meet the confidence requirement.
[0062] On the other hand, a multi-modal context-based intention acquisition system is provided, including:
[0063] A data initialization module for initializing the position confidence of each acupoint of the user;
[0064] A method construction module for constructing a method that combines gesture recognition confidence and gesture-acupoint association degree to calculate gesture confidence;
[0065] A data processing module for enabling the confidence of each acupoint to reflect the information of the entire human acupoint structure through the method of propagation and update;
[0066] A model construction module for constructing an improved Markov random field evaluation model for evaluating massage intentions.
[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0068] 1. By fusing multi-modal information of position, gesture, and voice, calculating the confidence of each modality, and through propagation and update, the confidence of each acupoint can reflect the information of the entire human acupoint structure. This method improves the accuracy of user intention extraction and enables the system to capture and respond to user needs more precisely;
[0069] 2. An active human-machine collaboration algorithm based on massage intention trust: Aiming at physiological changes such as forgetfulness and slow movement in the elderly, a method for evaluating intention trust based on an improved Markov random field (MRF) evaluation model is proposed. This model combines time factors, historical factors, single-modal information entropy, and single-modal recognition credibility. Through the dynamic threshold method and the method of actively prompting the user to input, it realizes the precise capture and response to the user's intention, thereby enhancing the interaction effect and user experience of the system. Brief Description of the Drawings
[0070] Figure 1 is the flowchart of the method of the present invention;
[0071] Figure 2 is the schematic structural diagram of the system of the present invention. Detailed Embodiments
[0072] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the present application.
[0073] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationship of each component or element of the present invention, and do not specifically refer to any component or element of the present invention, and should not be construed as a limitation to the present invention.
[0074] In the present invention, terms such as "fixed connection", "connected", "connected" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For those skilled in the relevant scientific research or technology in this field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances, and should not be construed as a limitation to the present invention.
[0075] Embodiment:
[0076] As Figure 1 shown, this embodiment provides a method for obtaining an intention based on multi-modal context, including:
[0077] Initializing the position confidence of each acupoint of the user;
[0078] Constructing a method combining gesture recognition confidence and gesture-acupoint association degree for calculating gesture confidence;
[0079] Making the confidence of each acupoint reflect the information of the entire human acupoint structure through a propagation update method;
[0080] Constructing an improved Markov random field evaluation model for evaluating massage intention.
[0081] Initialize the position confidence of each acupoint. This process includes the weighted distance Dis between the user's indicated position and each acupoint. In this embodiment, the position distance is defined as the comprehensive distance in the multidimensional space, and different weights are given to the distances in different dimensions according to the error sensitivity in different directions:
[0082]
[0083] Where (x i ,y i ,z i ) is the coordinate of acupoint i, (x0, y0, z0) is the coordinate of the position indicated by the user, and w x ,w y ,w z are the weights of the corresponding directions, reflecting the error sensitivity of the direction. The weight of each direction is determined by recording the number and amplitude of adjustments made by the user in different directions and by counting the impact of adjustments in different directions on the final massage effect. The position confidence of the acupoint C pos (i) Initialized as the inverse of the distance to the acupoint to more accurately reflect the relative proximity between the acupoint and the location information:
[0084]
[0085] In order to scientifically calculate the gesture confidence and accurately reflect the user's intention, this embodiment proposes a method based on combining the gesture recognition confidence and the gesture-acupoint association:
[0086] First, define the confidence level of the gesture H output by the gesture recognition algorithm as C gesture (H), the confidence reflects the predicted probability of gesture H by the gesture recognition model;
[0087] Secondly, define the association between gestures and acupoints: According to the gesture intention table and user history information, the association A between gesture H and acupoint i can be determined gesture (H,i), for example, gesture H-1 (only the index finger is extended, indicating a change in massage position) may be associated with multiple acupoints, while gesture H-4 (four fingers together and upward, increasing strength) may be associated with a specific acupoint. By calculating the correlation between each gesture and the acupoint, we can construct a correlation matrix A, where the rows of the matrix A represent different gestures H, and the columns represent different acupoints i. The matrix elements A H,i Indicates the correlation between gesture H and acupoint i, correlation A H,i The value of is usually between 0 and 1, with 0 indicating no correlation and 1 indicating perfect correlation:
[0088]
[0089] The number of times gesture H and acupoint i appear together is N. H,i, the total number of times gesture H appears is N H . By combining the gesture recognition confidence and the association degree between the gesture and the acupoint, the confidence of gesture H corresponding to acupoint i can be calculated. The specific calculation formula is as follows:
[0090] C gesture (H, i) = C gesture (H) · A gesture (H, i) (4)
[0091] To obtain the comprehensive gesture confidence of each acupoint i, sum all the recognized gestures H. The calculation formula is as follows:
[0092]
[0093] Similarly, the voice confidence based on the voice recognition result and the association degree with the acupoint can be obtained:
[0094] C voice (V, i) = C voice (V) · A voice (V, i) (6)
[0095] Weightedly fuse the confidences of position, gesture, and voice to obtain the comprehensive confidence of each acupoint:
[0096]
[0097] After obtaining the initial comprehensive confidence, through the method of propagation and update, the confidence of each acupoint can reflect the information of the entire human acupoint structure. For each acupoint i, perform weighted averaging according to the confidence of its neighboring acupoints. The calculation formula is as follows:
[0098]
[0099] Among them, N(i) represents the neighbor set of acupoint i, and w ij represents the weight between acupoint i and neighboring acupoint j. Next, normalize the confidence of all acupoints according to formula 9 to ensure that their sum is 1:
[0100]
[0101] After that, repeat the propagation update and normalization until the confidence of the acupoints converges. Finally, take the acupoint with the highest confidence as the possible massage position for the elderly, that is, the multimodal context sub-intention set M.
[0102] In order to more accurately evaluate the user's massage intention in the intelligent massage system, and perform self-repair and avoid wrong intentions, this embodiment proposes an improved Markov random field (MRF) evaluation model. This model combines time factors, historical factors, unimodal information entropy, and unimodal recognition credibility, making the evaluation results more objective and accurate, and providing a scientific decision-making basis for harmonious human-computer interaction. The joint probability distribution of the improved Markov random field is shown in formula (10).
[0103]
[0104] Where P, G, V, T, H represent position information, gesture information, voice information, time information, historical information in the intelligent massage system respectively, and M represents multimodal context intention. φ represents the potential function between the node and the massage intention, and Z is the normalization constant, which is used to ensure the normalization of probability.
[0105] To calculate the conditional probability distribution of the massage intention M, we use Bayes' formula 3.2.2 for reasoning:
[0106]
[0107] To further improve the model, this embodiment comprehensively considers unimodal information entropy and unimodal recognition credibility. Define information entropy H(E) and recognition credibility C(E):
[0108]
[0109] The recognition credibility can be expressed as the confidence of unimodal recognition:
[0110] C(E) = confidence(E) (13)
[0111] After comprehensively considering unimodal information entropy and unimodal recognition credibility, the posterior probability of the massage intention is:
[0112]
[0113] Finally, use formula (14) to calculate the credibility P(M) of the multimodal context intention M.
[0114] In the intelligent massage system, in order to effectively evaluate and determine the user's massage intention M, a reasonable threshold is required to judge whether the trust degree is high enough. If the trust degree is lower than the threshold, the system will request more information for reverse active fusion. θ(t) represents the dynamic threshold. The dynamic intention threshold means that the thresholds of different intentions may be different at the same moment, and the thresholds of the same intention may also be different at different moments. It can be adjusted according to the changes of time, historical data, and user habits. To ensure the dynamics and adaptability of the threshold, we define the following formula:
[0115] θ(t) = α·θ0·f(t) (15)
[0116] Wherein, θ0 is the initial threshold, and the initial threshold θ0 can be set by analyzing historical data and user feedback. For example, assuming that the average value μ and standard deviation σ of the user intention trust degree are obtained from historical data, the initial threshold can be set as: μ - σ; α is an adjustment factor used to adjust according to user feedback and system performance. f(t) is a time-related function representing the change law of the threshold at different times, which can be set according to actual situations. In the embodiment, an exponential decay function is used to represent it. If P(M) ≥ θ(t), then M is taken as the executable intention E; if P(M) ≤ θ(t), it means that the true intention M does not meet the trust degree requirement, and the system will actively remind the user to interact again and will extract and evaluate the intention again, and finally extract the executable intention E.
[0117] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for obtaining an intention based on multi-modal context, characterized in that, It includes the following steps: Initialize the position confidence of each acupoint of the user; Construct a method based on the combination of gesture recognition confidence and the association degree between gesture and acupoint to calculate gesture confidence; Through the method of propagation update, make the confidence of each acupoint reflect the information of the whole human acupoint structure; Construct an improved Markov random field evaluation model to evaluate the massage intention.
2. The method for obtaining an intention based on multimodal context according to claim 1, wherein The initialization of the position confidence of each acupoint of the user includes: The weighted distance Dis between the user-indicated position and each acupoint. Define the position distance as the comprehensive distance in the multi-dimensional space, and assign different weights to the distances in different dimensions according to the error sensitivity in different directions: Among them, (x i , y i , z i ) are the coordinates of acupoint i, and (x0, y0, z0) are the coordinates of the position indicated by the user; w x , w y , w z are the weights in the corresponding directions, respectively, reflecting the error sensitivity in that direction Determine the weight of each direction by recording the number and amplitude of adjustments of the user in different directions and statistically analyzing the influence of adjustments in different directions on the final massage effect; Position confidence C of acupoints pos (i) Initialize it to the reciprocal of its distance:
3. The method for obtaining an intention based on multimodal context according to claim 1, wherein The construction of the method based on the combination of gesture recognition confidence and the association degree between gesture and acupoint includes: Define the confidence of the gesture recognition algorithm outputting gesture H as C gesture (H), as well as the correlation between the gesture and the acupoints; Determine the association degree A between the gesture H and the acupoint i according to the gesture intention table and the user's historical information gesture (H, i); By calculating the correlation degree between each gesture and acupoints, a correlation degree matrix A is constructed. The rows in the matrix A represent different gestures H, the columns represent different acupoints i, and the matrix element A H,i represents the correlation degree between gesture H and acupoint i, and A H,i ∈[0, 1], where 0 means completely uncorrelated and 1 means completely correlated: Among them, the number of times the gesture H and the acupoint i appear jointly is N H,i , and the total number of times the gesture H appears is N H , combining the gesture recognition confidence and the association degree between the gesture and the acupoint, calculate the confidence of the gesture H corresponding to the acupoint i: C gesture (H, i) = C gesture (H) · A gesture (H, i).
4. The method for obtaining an intention based on multimodal context according to claim 3, wherein, The construction of the method based on the combination of gesture recognition confidence and the association degree between gesture and acupoint also includes: Sum all the recognized gestures H to calculate the comprehensive gesture confidence of each acupoint i: Calculate the speech confidence based on the result of speech recognition and the association degree with the acupoint: C voice (V,i) = C voice (V)·A voice (V,i) Perform weighted fusion on the confidences of position, gesture and speech to obtain the comprehensive confidence of each acupoint: C(i) = w pos ·C pos (i) + w gesture ·∑ H C gesture (H, i) + w voice ·∑ V C voice (V, i).
5. The method for obtaining an intention based on multimodal context according to claim 4, wherein The method of making the confidence of each acupoint reflect the information of the whole human acupoint structure through propagation update is specifically: For each acupoint i, perform weighted average according to the confidence of its neighboring acupoints: Among them, N(i) represents the set of neighbors of acupoint i, and w ij represents the weight between acupoint i and its neighboring acupoint j; Repeat propagation update and normalization until the confidence of the acupoints converges, and take the acupoint with the highest confidence as the possible massage position selected by the user, that is, the multi-modal context sub-intention set M.
6. The method for obtaining an intention based on multimodal context according to claim 1, wherein The construction of the improved Markov random field evaluation model includes: The joint probability distribution of the improved Markov random field is: Among them, P, G, V, T, H respectively represent the position information, gesture information, speech information, time information, historical information in the intelligent massage system, M represents the multi-modal context intention, φ represents the potential function between the node and the massage intention, and Z is the normalization constant used to ensure the normalization of the probability; Calculate the conditional probability distribution of the massage intention M: P(M|P,G,V,T,H) ∝ exp(-(ψ P (P,M) + ψ G (G,M) + ψ V (V,M) + ψ T (T,M) + ψ H (H,M))) Define the information entropy H(E) and the recognition credibility C(E): The recognition credibility is expressed as the confidence of unimodal recognition: C(E)=confidence(E) The posterior probability of the massage intention is:
7. The method for obtaining an intention based on multimodal context according to claim 6, characterized in that, The construction of the improved Markov random field evaluation model also includes: Set θ(t) to represent the dynamic threshold: θ(t)=α·θ0·f(t) Among them, θ0 is the initial threshold, and f(t) is a time-related function If P(M)≥θ(t), then take M as the executable intention E; if P(M)≤θ(t), then the true intention M does not meet the trust requirement.
8. An intention acquisition system based on multimodal context, characterized in that, It includes: A data initialization module for initializing the position confidence of each acupoint of the user; A method construction module for constructing a method based on the combination of gesture recognition confidence and the association degree between gesture and acupoint to calculate gesture confidence; A data processing module, which is used to make the confidence of each acupoint reflect the information of the whole human acupoint structure by means of propagation update; A model construction module, which is used to construct an improved Markov random field evaluation model for evaluating massage intention.