Semantic representation method and system for multi-mode operation instruction
Through multimodal feature extraction and semantic mapping of WSABIE model, the problem that multimodal instructions are difficult to convert into LLM understandable semantics is solved, and high-precision multimodal instructions semantic analysis and alignment are realized, improving user experience and system performance.
Patent Information
- Application Number
- CN202510159680.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art is difficult to convert non-directly semantic multimodal instructions (such as images, gestures, and touch) into semantic information that can be understood by large language models (LLM), limiting their application in multimodal scenarios.
Through multimodal feature extraction, semantic label generation, semantic mapping based on WSABIE model and correlation quantization evaluation, semantic alignment and precise analysis of multimodal instructions are realized.
It realizes high-precision, high-adaptive and efficient semantic analysis of multimodal instructions, significantly improves user experience and system performance, and provides strong technical support for multimodal interactive applications.
Smart Images

Figure CN120071046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic representation, and particularly to a semantic representation method and system for multimodal operation instructions. Background Art
[0002] With the rapid development of artificial intelligence technology, multimodal interaction has gradually become an important direction of human-computer interaction. The input forms of multimodal instructions include diverse ways such as voice, text, image, gesture, touch, etc. However, the instruction forms of modalities such as image, gesture, and touch do not directly contain explicit semantic content, which limits their direct application in the scenario understanding of large language models (LLMs).
[0003] Therefore, how to convert these non-directly semantic instructions into semantic information that can be understood by LLMs and achieve semantic alignment of multimodal operation instructions has become a key issue in the current technological development. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a semantic representation method and system for multimodal operation instructions. Through multimodal feature extraction, semantic label generation, semantic mapping based on the WSABIE model, and correlation quantification evaluation, semantic alignment and precise parsing of multimodal instructions are successfully achieved, which has the advantages of high precision, high adaptability, and high efficiency, provides strong technical support for multimodal interaction applications, and significantly improves the user experience and system performance.
[0005] To achieve the above purpose, the present invention provides the following solutions:
[0006] A semantic representation method for multimodal operation instructions, comprising:
[0007] Obtaining instruction features of different modalities according to the characteristics of multimodal instruction data;
[0008] Determining semantic labels according to the abstract description of the user's intention;
[0009] Inputting the instruction features of different modalities and the semantic labels into a WSABIE model for training to determine the optimal mapping relationship between the instruction features of different modalities and the semantic labels;
[0010] Based on the optimal mapping relationship, quantifying and evaluating the correlation between the instruction features of different modalities and the semantic labels to precisely parse the multimodal instruction data at the semantic level and obtain the semantic representation of the multimodal instruction data.
[0011] Preferably, the multimodal instruction data includes: image data, gesture data, and touch data.
[0012] Preferably, instruction features of different modalities are obtained according to the characteristics of the multimodal instruction data, including:
[0013] Performing image feature extraction on the image data to obtain the instruction features of the image data; the instruction features of the image data include: pixel values, color channel distribution, object morphology, surface texture, and semantic scene;
[0014] Performing gesture recognition and sensor feature extraction on the gesture data to obtain the instruction features of the gesture data; the instruction features of the gesture data include: the movement trajectory, speed, acceleration, duration, direction, and static pose features of the hand joints.
[0015] Performing feature extraction on the touch data based on a sensor signal processing algorithm to obtain the instruction features of the touch data; the instruction features of the touch data include: the coordinate position of the touch point, touch area, pressure, duration, and the execution order of touch actions.
[0016] Preferably, performing image feature extraction on the image data to obtain the instruction features of the image data, including:
[0017] Performing filtering preprocessing on the image data to obtain a filtered image;
[0018] Segmenting the filtered image to obtain a target region and a background region;
[0019] Performing pixel value feature extraction, color channel distribution feature extraction, object morphology feature extraction, surface texture feature extraction, and semantic scene feature extraction on the image at the pixel positions in the target region to obtain the instruction features of the image data.
[0020] Preferably, performing filtering preprocessing on the image data to obtain a filtered image, including:
[0021] Using a filtering window to detect noise points on the image data to obtain noise points to be processed;
[0022] When the number of noise points to be processed in the filtering window is greater than a preset threshold, denoising the image in the corresponding filtering window;
[0023] Sliding the filtering window until the entire image data is traversed to obtain the filtered image.
[0024] Preferably, the expression of the semantic label is:
[0025]
[0026] Among them, S is the semantic tag, Embed(Intent) is the intent keyword extracted by parsing the user's voice or text input through natural language processing technology; C i is the i-th feature of the context, γ i is the weight of the i-th feature, n is the number of features, Action is the specific operation expressed by the multi-modal instruction data, and Action is the result obtained by feature extraction through the CNN network; Object is the detection result obtained by the object detection algorithm based on the image, β 1 、β 2 and β 3 are dynamic weights respectively, and the sum of β 1 、β 2 and β 3 is 1.
[0027] Preferably, based on the optimal mapping relationship, the correlation between the instruction features of different modalities and the semantic tag is quantitatively evaluated to accurately parse the multi-modal instruction data at the semantic level and obtain the semantic representation of the multi-modal instruction data, including:
[0028] Determine the correlation between the instruction feature and the semantic tag according to the mutual information calculation method; the expression of the correlation MI(F, S) is: where P(f, s): the joint probability distribution of the instruction feature F and the semantic tag S, P(f) is the marginal probability distribution of the instruction feature F, and P(s) is the marginal probability distribution of the semantic tag S;
[0029] Automatically learn the importance weight of different instruction features for the semantic tag according to the attention mechanism in deep learning; the expression of the attention mechanism is: where e i = Score(F i , S) = W a ·(F i ⊙ S) + b a , e i represents the correlation score between the i-th feature F i and the semantic tag S; α i is the importance weight of the i-th feature, and after normalization, it satisfies W a and b a represent the trainable parameters of the attention mechanism; ⊙ is the element-wise product;
[0030] Determine the semantic representation of the multi-modal instruction data according to the correlation and the importance weight.
[0031] Preferably, the formula for the semantic representation is:
[0032]
[0033] Among them, w i is the final weight, w i = λ 1 ·MI(F i , S) + λ 2 ·α i , λ 1 and λ 2 are extremely heavy adjustment coefficients, which are used to balance the contributions of mutual information and the attention mechanism.
[0034] A semantic representation system for multimodal operation instructions includes:
[0035] A feature acquisition unit, which is used to obtain instruction features of different modalities according to the characteristics of multimodal instruction data;
[0036] A label acquisition unit, which is used to determine semantic labels according to the abstract description of the user's intention;
[0037] A mapping determination unit, which is used to input the instruction features of different modalities and the semantic labels into a WSABIE model for training to determine the optimal mapping relationship between the instruction features of different modalities and the semantic labels;
[0038] A semantic representation unit, which is used to quantitatively evaluate the correlation between the instruction features of different modalities and semantic labels based on the optimal mapping relationship, so as to accurately analyze multimodal instruction data at the semantic level and obtain the semantic representation of the multimodal instruction data.
[0039] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0040] The present invention provides a semantic representation method and system for multimodal operation instructions, including: obtaining instruction features of different modalities according to the characteristics of multimodal instruction data; determining semantic labels according to the abstract description of the user's intention; inputting the instruction features of different modalities and the semantic labels into a WSABIE model for training to determine the optimal mapping relationship between the instruction features of different modalities and the semantic labels; quantitatively evaluating the correlation between the instruction features of different modalities and semantic labels based on the optimal mapping relationship, so as to accurately analyze multimodal instruction data at the semantic level and obtain the semantic representation of the multimodal instruction data. Through multimodal feature extraction, semantic label generation, semantic mapping based on the WSABIE model, and correlation quantitative evaluation, the present invention successfully realizes semantic alignment and accurate analysis of multimodal instructions, has the advantages of high precision, high adaptability, and high efficiency, provides strong technical support for multimodal interaction applications, and significantly improves the user experience and system performance. Description of the Drawings
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 It is the flowchart of the method provided by the embodiment of the present invention;
[0043] Figure 2 It is the schematic diagram of the system structure provided by the embodiment of the present invention. Detailed implementation manners
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0045] The purpose of the present invention is to provide a semantic representation method and system for multimodal operation instructions. Through multimodal feature extraction, semantic label generation, semantic mapping based on the WSABIE model, and correlation quantification evaluation, the semantic alignment and accurate parsing of multimodal instructions are successfully realized, with the advantages of high precision, high adaptability, and high efficiency, providing strong technical support for multimodal interaction applications and significantly improving the user experience and system performance.
[0046] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will further describe the present invention in detail in conjunction with the drawings and specific implementation manners.
[0047] Figure 1 It is the flowchart of the method provided by the embodiment of the present invention. As Figure 1 shown, the present invention provides a semantic representation method for multimodal operation instructions, including:
[0048] Step 100: Obtain instruction features of different modalities according to the characteristics of multimodal instruction data;
[0049] Step 200: Determine semantic labels according to the abstract description of the user's intention;
[0050] Step 300: Input the instruction features of different modalities and semantic labels into the WSABIE model for training to determine the optimal mapping relationship between the instruction features of different modalities and semantic labels;
[0051] Step 400: Based on the optimal mapping relationship, quantitatively evaluate the correlation between instruction features of different modalities and semantic tags, so as to accurately analyze multi-modal instruction data at the semantic level and obtain the semantic representation of multi-modal instruction data.
[0052] Preferably, the multi-modal instruction data includes: image data, gesture data, and touch data.
[0053] Specifically, due to its inherent high complexity and diversity, multi-modal data presents unique feature representations, and each modality has its irreplaceable characteristics. In the image modality, the information contained is extensive, ranging from the distribution of pixel values and color channels at the basic level to the object shape, surface texture, and overall semantic scene at a higher level in multiple dimensions. The gesture modality contains rich dynamic features, such as the movement trajectory, speed, acceleration, duration, direction of hand joints in space, and the features of static postures at the start and end of the gesture. The touch modality covers multiple aspects of touch interaction, including the specific coordinate position of the touch point, the touch area, the applied touch pressure, the touch duration, and the execution order of touch actions. As the core for parsing multi-modal information, semantic tags play a role in highly abstractly describing the user's intention. In order to thoroughly understand the deep connection between multi-modal information and semantic tags, the present invention needs to carefully reveal the corresponding relationship between different modality features and specific semantics through extensive data analysis and in-depth domain knowledge. For example, the shape and position of an object in an image may be closely related to the semantic tag of "activating the relevant application of the object"; while a specific gesture movement pattern may be mapped to specific semantics such as "confirm" or "cancel".
[0054] Preferably, instruction features of different modalities are obtained according to the characteristics of multi-modal instruction data, including:
[0055] Extract image features from the image data to obtain the instruction features of the image data; the instruction features of the image data include: pixel values, color channel distribution, object shape, surface texture, and semantic scene;
[0056] Perform gesture recognition and sensor feature extraction on the gesture data to obtain the instruction features of the gesture data; the instruction features of the gesture data include: the movement trajectory, speed, acceleration, duration, direction of hand joints, and static posture features.
[0057] Extract features from the touch data based on a sensor signal processing algorithm to obtain the instruction features of the touch data; the instruction features of the touch data include: the coordinate position of the touch point, the touch area, the pressure, the duration, and the execution order of touch actions.
[0058] Optionally, the WSABIE annotation technology endows the present invention with the powerful ability to process multi-modal data annotation and its semantic mapping. First, in terms of feature extraction of multi-modal data, the present invention designs a dedicated feature extraction module. For image data, by combining the deep feature extraction ability of convolutional neural network (CNN) with traditional image feature extraction technology, a comprehensive representation of the image is achieved. In the field of gesture recognition, the present invention adopts a strategy that combines sensor data feature extraction algorithm and vision-based gesture recognition technology to ensure the accuracy of gesture description. As for touch features, the present invention effectively obtains the feature expression of touch data through the processing algorithm of touch sensor signals. Projecting multi-modal features and semantic labels into a shared latent embedding space constitutes the core step of the research. When designing this latent embedding space, the dimensionality differences of different modal features and the diversity of semantics must be fully considered. By training the WSABIE model and optimizing its weights and parameters, the present invention can ensure that the multi-modal feature vectors and semantic label vectors find the optimal mapping relationship in this space. During the training process, the present invention uses a large-scale multi-modal data set, which contains instruction data generated by users in various scenarios, so as to improve the generalization performance of the model.
[0059] Preferably, performing image feature extraction on the image data to obtain the instruction features of the image data includes:
[0060] Performing filtering preprocessing on the image data to obtain a filtered image;
[0061] Segmenting the filtered image to obtain a target region and a background region;
[0062] Performing pixel value feature extraction, color channel distribution feature extraction, object morphology feature extraction, surface texture feature extraction, and semantic scene feature extraction on the image at the pixel positions in the target region to obtain the instruction features of the image data.
[0063] Preferably, performing filtering preprocessing on the image data to obtain a filtered image includes:
[0064] Using a filtering window to detect noise points on the image data to obtain noise points to be processed;
[0065] When the number of noise points to be processed in the filtering window is greater than a preset threshold, denoising the image in the corresponding filtering window;
[0066] Sliding the filtering window until the entire image data is traversed to obtain the filtered image.
[0067] Preferably, the expression of the semantic label is:
[0068]
[0069] Among them, S is the semantic label, Embed(Intent) is the intent keyword extracted by parsing the user's voice or text input through natural language processing technology; C i is the i-th feature of the context, γ i is the weight of the i-th feature, n is the number of features, Action is the specific operation expressed by the multi-modal instruction data, and Action is the result obtained by feature extraction through the CNN network; Object is the detection result obtained by the object detection algorithm based on the image, β 1 、β 2 and β 3 are dynamic weights respectively, and the sum of β 1 、β 2 and β 3 is 1.
[0070] Exemplarily, the determination of the dynamic weights β 1 、β 2 and β 3 is automatically completed through the model training process. In the model initialization stage, the weights can be set to equal values (for example, β 1 = β 2 = β 3 ) to ensure the balance of the contributions of each part in the initial state. During the training process, a target function (such as a loss function based on semantic matching) is defined, and the weights are optimized through the backpropagation algorithm. The model will dynamically adjust the weight values according to the matching effect between the multi-modal features (Intent, Action, Object) and the semantic label, so that each weight can reflect the actual contribution of its corresponding feature to the semantic label. The adjustment of the dynamic weights depends on the correlation and importance between the features and the semantic label. By using the mutual information calculation method to quantify the correlation between the features and the semantic label, and combining the attention mechanism to capture the dynamic relationship between the features and the semantic label, the model will dynamically adjust β 1 、β 2 and β 3 . For example, if a certain feature (such as Intent) makes a greater contribution to the semantic label, the model will automatically increase the weight related to this feature (such as β 1 ). This dynamic adjustment mechanism ensures that the weight distribution can adapt to different scenarios and user needs, while maintaining the normalization constraint of β 1 + β 2 + β 3 = 1.
[0071] Specifically, in this embodiment, the generation of semantic tags is achieved through the processing and feature extraction of multimodal data. First, natural language processing techniques (such as BERT or Word2Vec) are used to parse the user's speech or text input to extract the user's core intent keywords (Intent). At the same time, in combination with context information (such as device status, user historical behavior, etc.), by assigning weights to multiple features of the context, the contribution of each feature to the semantic tag is dynamically adjusted. The weights of the context features are automatically optimized by model training to ensure that they can accurately reflect the semantic requirements of the current scenario. Then, the generation of semantic tags is further improved through the feature extraction of multimodal data. Specifically, the Action feature extracts key behavior features from the user's gestures, touches, or other operations through a convolutional neural network (CNN); the Object feature identifies the target object of the user's operation from the image through an object detection algorithm (such as Faster R-CNN). Finally, the Intent, Action, and Object features are combined with dynamic weights for fusion to generate a complete semantic tag, ensuring that it can accurately express the user's intent and provide support for subsequent semantic parsing and instruction execution.
[0072] Optionally, step 300 of this embodiment includes:
[0073] 1. First, prepare a multimodal dataset, including instruction features of modalities such as images, gestures, touches, etc., and corresponding semantic tags. Preprocess the instruction features of each modality. For example, normalize the image data, perform time series segmentation on the gesture data, extract path features from the touch data, etc., to ensure that the features of all modalities can be uniformly represented. At the same time, perform embedding processing on the semantic tags (such as using Word2Vec or BERT) to convert them into high-dimensional vector representations for alignment with the multimodal features in the same space.
[0074] 2. Then, design and initialize the WSABIE model. The core of the WSABIE model is a shared latent embedding space for mapping multimodal instruction features and semantic tags into the same semantic space. The input of the model includes multimodal feature vectors and semantic tag vectors, and the output is the similarity score between the two. By defining an objective function (such as a ranking-based loss function), the model can learn how to maximize the similarity between the correct instruction features and semantic tags while minimizing the similarity of incorrect matches.
[0075] 3. During the model training process, the WSABIE model is optimized using the supervised learning method. The multi-modal features and their corresponding semantic labels are used as positive samples, and at the same time, some randomly generated mis-matched negative samples are added. Through the contrastive learning of positive and negative sample pairs, the model can gradually adjust the mapping relationship in the embedding space, making the similarity score of positive samples higher than that of negative samples. During the training process, the gradient descent algorithm can be used to optimize the model parameters to ensure that the model can converge to the optimal state.
[0076] 4. Finally, the performance of the model is verified and optimized. By evaluating the accuracy, recall rate, and F1 score of the model on the validation set, it is judged whether the WSABIE model can accurately match the multi-modal features with the semantic labels. If the performance is not ideal, the hyperparameters of the model (such as the dimension of the embedding space, learning rate, etc.) can be adjusted or the diversity of the training data can be increased. The optimized WSABIE model can generate the optimal mapping relationship, providing efficient and accurate support for the semantic parsing of multi-modal instructions.
[0077] In order to accurately parse instructions at the semantic level, the present invention needs to accurately quantify the correlation between instruction features and semantic labels. In addition to using traditional statistical correlation analysis methods, such as mutual information calculation, Pearson correlation coefficient calculation, etc., the present invention will also introduce the attention mechanism in deep learning. This mechanism can automatically learn the importance weights of different features for semantic labels. For example, in a multi-modal instruction that combines images and gestures, the attention mechanism is capable of identifying the image regions related to the gesture semantics and assigning higher weights to these regions, thereby achieving a more accurate measurement of the correlation between the overall instruction and the semantic label. In addition, the present invention constructs a dynamic correlation evaluation model, which can continuously update the correlation measurement between features and semantic labels in real time as new multi-modal instruction data is continuously input. This real-time update can be achieved through an online learning algorithm, ensuring that the system can adapt to new user behaviors and instruction patterns.
[0078] Preferably, based on the optimal mapping relationship, the correlation between the instruction features of different modalities and the semantic labels is quantitatively evaluated to accurately parse the multi-modal instruction data at the semantic level and obtain the semantic representation of the multi-modal instruction data, including:
[0079] Determine the correlation between the instruction feature and the semantic label according to the mutual information calculation method; the expression of the correlation MI(F,S) is: where P(f,s): the joint probability distribution of the instruction feature F and the semantic label S, P(f) is the marginal probability distribution of the instruction feature F, and P(s) is the marginal probability distribution of the semantic label S;
[0080] Automatically learn the importance weights of different instruction features for the semantic label according to the attention mechanism in deep learning; the expression of the attention mechanism is: where e i =Score(F i ,S)=W a ·(F i ⊙S)+b a , e i represents the correlation score between the i-th feature F i and the semantic label S; α i is the importance weight of the i-th feature, and after normalization, it satisfies W a and b a represent the trainable parameters of the attention mechanism; ⊙ is the element-wise product;
[0081] Determine the semantic representation of the multimodal instruction data according to the correlation and the importance weight.
[0082] Preferably, the formula for the semantic representation is:
[0083]
[0084] where w i is the final weight, w i =λ 1 ·MI(F i ,S)+λ 2 ·α i , λ 1 and λ 2 are extremely important adjustment coefficients used to balance the contributions of mutual information and the attention mechanism.
[0085] Specifically, step 400 of this embodiment includes:
[0086] 1. Calculate the correlation by mutual information
[0087] First, quantify the correlation between the instruction feature and the semantic label through the mutual information calculation method. Mutual information is used to measure the dependence between the instruction feature and the semantic label and can reflect the contribution of the feature to the semantic label. In specific implementation, the joint probability distribution of the multimodal instruction feature and the semantic label and their respective marginal probability distributions are statistically analyzed. By analyzing these probability distributions, the correlation score between each modal feature (such as image feature, gesture feature, etc.) and the semantic label can be calculated. The higher the mutual information value, the greater the contribution of the feature to the semantic label, providing a basis for subsequent weight assignment.
[0088] 2. Learn the weight by the attention mechanism
[0089] Next, the attention mechanism in deep learning is used to automatically learn the importance weights of different modal features for semantic labels. The attention mechanism dynamically adjusts the importance of features by calculating the correlation scores between each feature and the semantic label. In specific implementation, a trainable scoring function (such as dot product or multi-layer perceptron) is used to calculate the correlation scores between each feature and the semantic label, and a weight distribution is generated through softmax normalization. The trainable parameters of the attention mechanism are continuously optimized through model training, enabling the system to focus on the feature regions most relevant to the semantic label, thereby improving the accuracy of semantic parsing.
[0090] 3. Generate semantic representations by integrating correlation and weights
[0091] After obtaining the mutual information correlation scores and the attention mechanism weights, the two are combined to generate the final feature weights. By dynamically adjusting the contribution ratios of mutual information and the attention mechanism, it is ensured that the feature weights can comprehensively reflect the relationship between features and semantic labels. Mutual information provides a statistical-level quantification of correlation, while the attention mechanism captures the dynamic relationship between features and semantic labels through deep learning. The combination of the two makes the final feature weights both statistically meaningful and capable of adapting to complex multi-modal scenarios.
[0092] 4. Generate semantic representations of multi-modal instructions
[0093] Finally, based on the comprehensively obtained feature weights, all modal features are weighted and aggregated to generate the semantic representation of multi-modal instruction data. The semantic representation is a high-dimensional vector that can accurately express the semantic information of multi-modal instructions. By dynamically adjusting the contribution ratios of mutual information and the attention mechanism, the semantic representation can adapt to different scenarios and user requirements. The generated semantic representation can not only be used for semantic parsing but also be passed as input to subsequent task modules (such as instruction execution or semantic retrieval), thereby realizing the efficient processing and execution of multi-modal instructions.
[0094] Exemplarily, in the initialization of the model, an initial value (such as a uniform distribution or an empirical value) is set for each extremely heavy adjustment coefficient to ensure that the contributions of different parts are balanced in the initial stage. Subsequently, during the training process, by defining an objective function (such as a loss function for semantic matching or classification accuracy), optimization algorithms such as gradient descent are used to iteratively update the extremely heavy adjustment coefficients. The model will dynamically adjust the values of these coefficients according to the matching effect between multi-modal features and semantic labels, enabling it to balance the contributions of mutual information and the attention mechanism, thereby maximizing the accuracy and robustness of the semantic representation. In addition, methods such as cross-validation or grid search can be used to test different coefficient combinations on the validation set and select the parameter configuration with the best performance to ensure that the extremely heavy adjustment coefficients can adapt to different scenarios and task requirements.
[0095] The present invention aims to construct an integrated technical solution based on an efficient system architecture. This architecture encompasses key modules such as multi-modal data acquisition, feature extraction, semantic mapping based on WSABIE, relevance measurement, and the interface with the LLM. To optimize system performance, this solution adopts distributed computing technology to process large multi-modal data sets, thereby significantly enhancing computing efficiency. In addition, by compressing and quantifying the model, the memory requirements and computing resource consumption can be effectively reduced, ensuring that the system can operate stably on various hardware platforms. This technical solution aims to achieve fast and accurate processing of the semantics of multi-modal operation instructions, providing strong technical support for multi-modal interaction applications.
[0096] The following is an implementation case of the present invention in an intelligent interaction system, which aims to provide users with a convenient and efficient multi-modal interaction experience, and achieve accurate understanding and execution of the semantics of multi-modal operation instructions.
[0097] 1. System Environment and Data Preparation
[0098] The system developed in the present invention is implemented on an intelligent terminal integrated with multiple interaction devices. This system integrates devices such as touch screens, cameras, and sensors, aiming to collect multi-modal data such as touch, images, and gestures. To further train and optimize the system, the present invention has collected a large number of multi-modal instruction samples. These samples reflect the interaction operations of different users in diverse application scenarios such as games, office work, and entertainment. For example, in the game application scenario, users may use gesture instructions to control the movement of characters and execute skill releases through touch instructions; while in the office application scenario, users may use gesture instructions to switch document pages and open specified files through image recognition technology.
[0099] 2. Multi-modal Data Acquisition and Preprocessing
[0100] During the process of user interaction with the intelligent terminal, the system starts to collect multi-modal data. In the recording of touch data, the sensor accurately records key information such as the coordinates, pressure, and duration of the touch point. The camera captures the user's gesture actions in real time and extracts the image sequence of the gestures. For image data, the system not only acquires the images of the user on the operation interface but also includes the images of the surrounding environment (in specific scenarios, this may be related to the instruction semantics). The system conducts preliminary processing on the collected data, which includes steps such as data cleaning and normalization. Taking the touch pressure value as an example, the system normalizes it to a specific interval range to eliminate the influence of differences between different devices on the data.
[0101] 3. Semantic Mapping Based on WSABIE
[0102] In the feature extraction stage, for touch data, the system of the present invention extracts the dynamic path characteristics and pressure change characteristics of touch points. For gesture images, the present invention uses a pre-trained convolutional neural network model to extract characteristics such as the shape and movement direction of gestures. For the feature extraction of image data, the present invention comprehensively uses object recognition and scene classification models to obtain key object and scene information in the image. Next, these multi-modal features and semantic labels are input into the annotation model based on WSABIE. In the trained latent embedding space, the system can accurately map the multi-modal features to the corresponding semantic labels.
[0103] 4. Correlation Measurement and Instruction Understanding
[0104] The present invention uses the attention mechanism and statistical correlation analysis method to quantitatively evaluate the correlation between instruction features and semantic labels. Taking "open the notification bar" as an example, the attention mechanism will focus on two key features: the upward sliding action of the gesture and the lower right corner of the touch screen, and assign them higher weights because these two features have a significant correlation with the semantics of "open the notification bar" in the training data. Further, by calculating the statistical correlation indicators of these features and semantic labels, the semantics of the instruction are further confirmed.
[0105] 5. Integration and Execution with the Large Language Model (LLM)
[0106] Once the system accurately grasps the semantic content of the multi-modal instruction, it translates it into a format that can be recognized by the language model (LLM) and transmits it to the LLM. Then, the LLM executes the corresponding command at the application or operating system level according to the received semantic information. For example, if the semantic instruction is "open the notification bar", the LLM will send a specific instruction to the system to activate the display function of the notification bar.
[0107] 6. Continuous Optimization and Feedback
[0108] During the operation of the system, new multi-modal instruction data is continuously aggregated and incorporated into the feedback mechanism. This batch of new data will be used to further refine the WSABIE model and the correlation evaluation system. For example, in the face of the emergence of a novel combination of gestures and touches, the system can update the model based on online learning algorithms to ensure correct interpretation of the semantics of such new instructions. This process continuously improves the accuracy and robustness of the system in understanding the semantics of multi-modal operation instructions, thereby bringing a smoother and more natural multi-modal interaction experience to users.
[0109] Corresponding to the above method, this embodiment also provides a semantic representation system for multi-modal operation instructions, including:
[0110] A feature acquisition unit for obtaining instruction features of different modalities according to the characteristics of multimodal instruction data;
[0111] A label acquisition unit for determining semantic labels according to the abstract description of the user's intention;
[0112] A mapping determination unit for inputting the instruction features of different modalities and the semantic labels into a WSABIE model for training to determine the optimal mapping relationship between the instruction features of different modalities and the semantic labels;
[0113] A semantic representation unit for quantitatively evaluating the correlation between the instruction features of different modalities and semantic labels based on the optimal mapping relationship to accurately analyze multimodal instruction data at the semantic level and obtain the semantic representation of the multimodal instruction data.
[0114] The beneficial effects of the present invention are as follows:
[0115] (1) By mapping the features of multimodal data (such as images, gestures, touches, etc.) and semantic labels to a shared latent embedding space, the present invention solves the problems of different feature dimensions and inconsistent semantic expressions between multimodal data; enables non-directly semantic multimodal instructions (such as gestures and touches) to be understood by large language models (LLMs), thereby achieving semantic alignment of multimodal instructions.
[0116] (2) By calculating mutual information and using an attention mechanism to quantitatively evaluate the correlation between multimodal features and semantic labels, the present invention can automatically identify key feature regions related to semantics and assign higher weights to them; this precise correlation quantification method of the present invention ensures more accurate parsing of multimodal instructions at the semantic level and reduces the possibility of misunderstanding or misparsing.
[0117] (3) Based on the large-scale training of the WSABIE model and using diverse instruction samples in the multimodal dataset, the present invention improves the model's adaptability to different scenarios and user behaviors; through a dynamic correlation evaluation model and an online learning algorithm, the system can update the correlation between features and semantic labels in real time to adapt to new user behaviors and instruction patterns, further enhancing the generalization ability and robustness of the system.
[0118] (4) By extracting features and performing semantic mapping on multimodal data such as images, gestures, and touches, the present invention can support complex multimodal interaction scenarios, such as games, office work, entertainment, etc.; in these scenarios, users can perform complex operations through multimodal instructions (such as a combination of gestures and touches), and the system can accurately understand and execute corresponding commands.
[0119] (5) The present invention uses distributed computing technology to process large-scale multimodal data, significantly improving the computing efficiency of the system; through model compression and quantization processing, the memory requirements and computing resource consumption are reduced, enabling the system to operate stably on various hardware platforms.
[0120] (6) The present invention can be continuously optimized, and the semantic understanding ability of new instructions is continuously improved through a feedback mechanism and an online learning algorithm, providing users with a more fluent and natural multimodal interaction experience; users can achieve intuitive and efficient operations through multimodal instructions (such as gestures, touches, voices, etc.), reducing the learning cost and operation complexity.
[0121] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same and similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.
[0122] Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manner and application scope according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A semantic representation method for multimodal operation instructions, characterized in that: include: Obtaining command features of different modes according to the characteristics of multi-modal command data; Determine semantic tags based on the abstract description of user intent; Inputting the command features of the different modalities and the semantic labels into the WSABIE model for training to determine the optimal mapping relationship between the command features of the different modalities and the semantic labels; Based on the optimal mapping relationship, the correlation between the instruction features of the different modalities and the semantic labels is quantitatively evaluated to accurately parse the multimodal instruction data at the semantic level and obtain the semantic representation of the multimodal instruction data.
2. The semantic representation method of multimodal operation instructions according to claim 1, characterized in that: The multimodal instruction data includes: image data, gesture data and touch data.
3. The semantic representation method of multimodal operation instructions according to claim 2, characterized in that: According to the characteristics of multi-modal instruction data, instruction features of different modes are obtained, including: Extracting image features from the image data to obtain instruction features of the image data; the instruction features of the image data include: pixel value, color channel distribution, object morphology, surface texture and semantic scene; Performing gesture recognition and sensor feature extraction on the gesture data to obtain instruction features of the gesture data; the instruction features of the gesture data include: motion trajectory, speed, acceleration, duration, direction and static posture features of hand joints. The touch data is feature extracted based on a sensor signal processing algorithm to obtain instruction features of the touch data; the instruction features of the touch data include: coordinate position of the touch point, touch area, pressure, duration and execution order of touch actions.
4. The semantic representation method of multimodal operation instructions according to claim 3, characterized in that: Extracting image features from the image data to obtain instruction features of the image data includes: Performing filtering preprocessing on the image data to obtain a filtered image; Segmenting the filtered image to obtain a target area and a background area; The image at the pixel position of the target area is subjected to pixel value feature extraction, color channel distribution feature extraction, object morphology feature extraction, surface texture feature extraction and semantic scene feature extraction to obtain the instruction feature of the image data.
5. The semantic representation method of multimodal operation instructions according to claim 4, characterized in that: Performing filtering preprocessing on the image data to obtain a filtered image includes: Using a filter window to detect noise points on the image data to obtain noise points to be processed; When the number of noise points to be processed in the filtering window is greater than a preset threshold, denoising the image in the corresponding filtering window; The filtering window is slid until the entire image data is traversed to obtain the filtered image.
6. The semantic representation method of multimodal operation instructions according to claim 1, characterized in that: The expression of the semantic tag is: Wherein, S is the semantic tag, Embed(Intent) is the intent keyword extracted by parsing the user's voice or text input through natural language processing technology; C i is the i-th feature of the context, γ i is the weight of the i-th feature, n is the number of features, Action is the specific operation expressed by the multimodal instruction data, and Action is the result of feature extraction through the CNN network; Object is the detection result obtained by the image-based target detection algorithm, β1, β2 and β3 are dynamic weights respectively, and the sum of β1, β2 and β3 is 1.
7. The semantic representation method of multimodal operation instructions according to claim 1, characterized in that: Based on the optimal mapping relationship, the correlation between the command features of the different modalities and the semantic labels is quantitatively evaluated to accurately parse the multimodal command data at the semantic level and obtain the semantic representation of the multimodal command data, including: The correlation between the instruction feature and the semantic label is determined according to the mutual information calculation method; the expression of the correlation MI(F, S) is: Where, P(f,s): joint probability distribution of instruction feature F and semantic label S, P(f) is the marginal probability distribution of instruction feature F, and P(s) is the marginal probability distribution of semantic label S; The importance weights of different instruction features to the semantic labels are automatically learned according to the attention mechanism in deep learning; the expression of the attention mechanism is: Among them, e i =Score(F i ,S)=W a ·(F i ⊙S)+b a , e i represents the i-th feature F i The relevance score with the semantic label S; α i is the importance weight of the i-th feature, which satisfies after normalization W a and b a represents the trainable parameters of the attention mechanism; ⊙ is the element-wise product; A semantic representation of the multimodal instruction data is determined based on the relevance and the importance weight.
8. The semantic representation method of multimodal operation instructions according to claim 7, characterized in that: The formula for the semantic representation is: Among them, w i is the final weight, w i =λ1·MI(F i ,S)+λ2·α i , λ1 and λ2 are weight adjustment coefficients used to balance the contribution of mutual information and attention mechanism.
9. A semantic representation system for multimodal operation instructions, characterized in that: include: A feature acquisition unit, used to obtain command features of different modes according to the characteristics of the multi-modal command data; A tag acquisition unit, used to determine a semantic tag based on an abstract description of the user's intention; A mapping determination unit, configured to input the command features of the different modalities and the semantic labels into a WSABIE model for training to determine an optimal mapping relationship between the command features of the different modalities and the semantic labels; The semantic representation unit is used to quantitatively evaluate the correlation between the instruction features of the different modalities and the semantic labels based on the optimal mapping relationship, so as to accurately parse the multimodal instruction data at the semantic level and obtain the semantic representation of the multimodal instruction data.
Citation Information
Patent Citations
Multi-modal interaction method and device, robot and storage medium
CN116149468A
Visual interaction system based on multiple modes
CN118535023A
Big language model fusion method and device based on multi-modal data and medium
CN118965283A
Method, device and storage medium for training model based on multi-modal data joint learning
US20220327809A1