Earphone control method and system based on multi-mode interactive gesture recognition

Through the method of multimodal interactive gesture recognition, a dynamic gesture library is screened and constructed, combined with individual user adaptation models and incremental learning, the problems of complex design of traditional headphone gestures and high error rate are solved, and more efficient and convenient headphone operation is achieved.

CN120010812AInactive Publication Date: 2025-05-16SHENZHEN XINKE MEIDA COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510473005.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional headphone gesture design is complex, and users need to remember a large number of gestures and function correspondence, resulting in high learning costs and high error touch rate. The existing solutions do not filter high-frequency gestures based on the actual data used by users, resulting in the core function operation path being diluted by redundant gestures.

Method used

By collecting behavioral data of group users, filtering high-frequency usage gestures, building a multimodal data feature library and a basic model of group users, and obtaining a dynamic gesture library. Obtain the historical usage data of individual users, conduct deviation analysis with the group user basic model, build an individual user adaptation model, dynamically adjust the multimodal weight through incremental learning, and output the final gesture recognition result.

Benefits of technology

It reduces the user's cognitive load, improves the accuracy of gesture recognition, reduces the error touch rate, and improves the convenience and efficiency of headphone operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010812A_ABST
    Figure CN120010812A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of earphone control, and discloses an earphone control method and system based on multi-modal interactive gesture recognition, and the method comprises the following steps: collecting behavior data of group users, carrying out the screening analysis of the collected behavior data through the number of times of gesture use and the Miller's law, and outputting a high-frequency use gesture set; and based on the high-frequency use gesture set, constructing a multi-modal data feature library and a group user basic model to obtain a dynamic gesture library. According to the method, by counting the use times of the gestures, the attention can be focused on the gestures frequently used by the user, the gestures with low use frequency are eliminated, the workload of subsequent processing is reduced, and the screening efficiency is improved; the Miller's law is introduced for screening, it is guaranteed that the gestures and the functions are within the range where the user can easily memorize, the cognitive load of the user is reduced, the user does not need to spend too much energy to memorize the complex corresponding relation between the gestures and the functions, and therefore operation of the earphone is simpler and more convenient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of earphone control technology, and in particular to an earphone control method and system based on multimodal interactive gesture recognition. Background Art

[0002] Earphones (Earphones, Headphones, Head-sets, Earpieces) are a pair of conversion units that receive electrical signals from a media player or receiver and convert them into audible sound waves using speakers close to the ears. Earphones can be separated from the media player and connected using only a plug, so that you can listen to the audio alone without disturbing others. Earphones were originally used for telephones and radios, but with the popularity of portable electronic devices, earphones are mostly used in mobile phones, walkmans, radios, portable video games, and digital audio players.

[0003] However, the design of traditional headphone gestures mostly relies on the subjective experience of engineers, and the number of gestures often exceeds the short-term memory capacity of humans. For example, a certain sports headset has set 12 gestures to cover all-scenario operations, including double-clicking the left ear to answer a call, triple-clicking the right ear to switch to noise reduction mode, sliding to adjust the volume, etc. Users need to remember the complex "gesture-function" correspondence, resulting in high learning costs. In high-frequency interaction scenarios such as sports and commuting, the user's false touch rate is as high as 30% due to fuzzy memory or hasty operation, which seriously affects the smoothness of interaction. More importantly, the existing solution does not filter high-frequency gestures based on the actual user usage data, and designs low-frequency functions such as "equalizer adjustment" and "find headphones" (usage frequency <50 times / month) and "play / pause" (usage frequency >1000 times / month) as independent gestures of the same level, resulting in the core function operation path being diluted by redundant gestures. Users need to "retrieve" the target operation in a complex system, thereby reducing efficiency. Summary of the invention

[0004] The object of the present invention is to provide a headset control method and system based on multimodal interactive gesture recognition to solve at least one of the above-mentioned problems in the prior art.

[0005] In a first aspect, the present invention provides a headset control method based on multimodal interactive gesture recognition, comprising the following steps:

[0006] Collect the behavioral data of group users, filter and analyze the collected behavioral data based on the number of gestures used and Miller's law, and output a set of high-frequency gestures;

[0007] Based on the frequently used gesture collection, a multimodal data feature library and a group user basic model are constructed to obtain a dynamic gesture library;

[0008] Obtain historical usage data of individual users, perform deviation analysis on historical data and the basic model of group users to obtain differential features, and build an individual user adaptation model based on the differential features;

[0009] Based on the individual user adaptation model, the multimodal weights are dynamically adjusted through incremental learning to output the final gesture recognition results.

[0010] In a second aspect, the present invention provides an earphone control system based on multimodal interactive gesture recognition, the system comprising:

[0011] High-frequency gesture screening module: collects the behavioral data of group users, screens and analyzes the collected behavioral data based on the number of gestures used and Miller's law, and outputs a set of high-frequency gestures;

[0012] Gesture library construction module: Based on the frequently used gesture collection, a multimodal data feature library and a group user basic model are constructed to obtain a dynamic gesture library;

[0013] Individual deviation analysis module: obtains historical usage data of individual users, performs deviation analysis on the historical data and the basic model of group users to obtain difference features, and builds an individual user adaptation model based on the difference features;

[0014] Gesture recognition output module: Based on the individual user adaptation model, the multimodal weights are dynamically adjusted through incremental learning to output the final gesture recognition results.

[0015] Beneficial effects of the present invention:

[0016] 1. The present invention can focus on gestures frequently used by users by counting the number of times gestures are used, and exclude gestures with low frequency of use, thereby reducing the workload of subsequent processing and improving screening efficiency. The Miller's law is introduced for screening to ensure that gestures and functions are within the range that users can easily remember, reducing the user's cognitive load. Users do not need to spend too much energy to remember the complex correspondence between gestures and functions, thereby making the operation of the headset simpler and more convenient.

[0017] 2. The present invention solves the problem of individual operation habit adaptation deviation caused by the one-size-fits-all general model. By calculating the deviation between individual gesture features and the group mean, a personalized adaptation model is constructed. Compared with the general model, the individual user adaptation model can better adapt to the unique gesture characteristics of different users, thereby significantly improving the accuracy of gesture recognition;

[0018] 3. The present invention absorbs new gesture data of users in real time through an incremental learning mechanism, automatically optimizes multimodal weights and synchronizes them to a dynamic gesture library without the need for manual configuration by the user, thereby improving the model's ability to continuously optimize as habits evolve. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] Figure 1 It is a flow chart of a headset control method based on multimodal interactive gesture recognition of the present invention;

[0021] Figure 2 It is a structural schematic diagram of an earphone control system based on multimodal interactive gesture recognition of the present invention. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0023] Embodiment 1

[0024] like Figure 1 As shown, an embodiment of the present invention provides an earphone control method based on multimodal interactive gesture recognition, which specifically includes S1-S2, and the process is as follows:

[0025] S1: Collect the behavior data of group users, filter and analyze the collected behavior data, and output a set of high-frequency gestures;

[0026] In S1, data on the headphone usage behavior of several groups of users is collected and analyzed. The cognitive psychology theory of Miller's Law is introduced to analyze the various usage functions of the headphones. After screening and statistics, the frequently used gestures in the group of users are identified, and finally a set of high-frequency gestures is output;

[0027] It should be explained that Miller's law states that human short-term memory can usually only hold 7±2 pieces of information. Headphones have many functions. If the design of operation gestures and functions exceeds the scope of user short-term memory, it will greatly increase the user's cognitive load. By introducing Miller's law, we can analyze whether the operation gestures corresponding to each headphone function are within the range that users can easily remember based on scientific cognitive theory. For example, if the headset is originally designed with 10 different gestures corresponding to 10 different functions, according to Miller's law, it is difficult for users to remember so many complex correspondences, resulting in difficulty in operation. Through analysis, the corresponding design of gestures and functions can be simplified to ensure that the operation gestures of key functions are simple and easy to remember, reducing the cognitive burden of users during use.

[0028] Therefore, Miller's law can be used to conduct an in-depth analysis of the various functions of the headset from the perspective of cognitive load and information processing. The analysis results will provide a basis for the subsequent screening of high-frequency gestures, and help to accurately identify gestures that are both in line with the user's cognitive habits and are frequently used;

[0029] Furthermore, a preferred implementation of this step includes:

[0030] Determine the group of users according to the target user group characteristics of the headset, wherein the group characteristics include but are not limited to: age, gender, occupation;

[0031] For example, if the earphones are positioned as sports earphones, the group of users may be people who exercise regularly;

[0032] Set the collection time to collect the headphone behavior data of group users in different usage scenarios, where the headphone behavior data includes: gestures and usage functions;

[0033] Specifically, gesture actions include but are not limited to: touch actions, such as clicking or sliding, and motion actions, such as nodding or shaking the head; usage functions include but are not limited to: play, pause, adjust volume, and switch songs;

[0034] For example, if the collection time is set to one month, for users who often exercise, such as those wearing sports earphones, gesture data is collected through the high-precision accelerometer, gyroscope and other sensors built into the earphones in different sports scenarios such as running, fitness, and cycling. At the same time, the user's operation of the earphones such as playing, pausing, adjusting the volume, switching songs, etc. is recorded;

[0035] Preprocessing the collected behavioral data includes but is not limited to: removing abnormal data points caused by sensor failure, data transmission errors, etc., and filling missing data with appropriate methods, such as mean filling and interpolation; standardizing data in different formats and units to ensure data consistency and comparability, for example, converting gesture action data collected by different sensors into a unified coordinate system and time scale;

[0036] Build a gesture-function association matrix based on the processed behavioral data and count the number of times each gesture is used;

[0037] Setting a usage limit, wherein the usage limit is summarized and set by those skilled in the art based on the usage scenario of the headset, and the usage limit may be different for different usage scenarios;

[0038] For each gesture, compare it with the set usage limit, retain the gestures with a usage limit greater than the limit, mark them as high-frequency candidate gestures, and build a preliminary high-frequency candidate set based on the high-frequency candidate gestures;

[0039] For example, if gesture A is used 20 times in the play function, 15 times in the pause function, and 10 times in the volume adjustment function, the total number of times gesture A is used is 45 times; if the set usage limit is 30 times, gesture A is retained;

[0040] Based on the preliminary high-frequency candidate set, Miller's law is introduced to perform secondary screening on each high-frequency candidate gesture. The specific process is as follows:

[0041] From the perspective of cognitive load, the preliminary high-frequency candidate set is evaluated to assess whether the total number of high-frequency candidate gestures exceeds the scope of the group user's short-term memory, wherein the scope of the group user's short-term memory is the total number of gestures exceeding 7±2. In this embodiment, 7 can be used as the limit;

[0042] If it exceeds the scope of the group user's short-term memory, a correlation analysis is performed on the used functions to determine whether the used gestures can be merged. The specific process is as follows:

[0043] Use natural language processing technology to perform semantic analysis on the description of each function. For example, convert the function descriptions of "play", "pause", "previous song", and "next song" into word vectors.

[0044] The semantic correlation between the word vectors corresponding to any two functional combinations is calculated using the cosine similarity algorithm, and the calculated cosine similarity is used as the association value;

[0045] Setting a correlation threshold, and comparing the correlation value with the correlation threshold;

[0046] Arrange the function combinations corresponding to the correlation threshold values ​​from large to small, and merge the gestures in order of ranking until the total number of gestures does not exceed the short-term memory range of the group users, thus obtaining a high-frequency gesture set.

[0047] It needs to be explained that the meaning of the combined representation is: using the same gesture to correspond to the two usage functions;

[0048] For example, the gesture for combining the two functions of play and pause is to click the earphone once, that is, to play once and to pause again; the gesture for combining the two functions of previous song and next song is to slide the earphone left and right, that is, to slide left for the previous song and to slide right for the next song;

[0049] It should be noted that if the total number of gestures used still exceeds the short-term memory of the group of users after all strongly related functional combinations have been merged, the gestures used with the smallest number of times will be filtered out in turn;

[0050] By counting the number of times gestures are used, attention can be focused on gestures that are frequently used by users, and those gestures that are used less frequently can be excluded, reducing the workload of subsequent processing and improving screening efficiency. The Miller Law is introduced for screening to ensure that gestures and functions are within the range that users can easily remember, reducing the user's cognitive load. Users do not need to spend too much energy to remember the complex correspondence between gestures and functions, making the operation of the headset simpler and more convenient.

[0051] S2: Based on the frequently used gesture set, a multimodal data feature library and a group user basic model are constructed to obtain a dynamic gesture library;

[0052] In this step, for each gesture in the high-frequency gesture set, historical multimodal data of group users in different usage scenarios are collected;

[0053] Specifically, multimodal data includes visual data, such as gesture action videos captured by a camera, audio data, such as the sound produced when operating the headset, such as the sound of clicking the headset, voice command audio, and sensor data, such as the acceleration and angular velocity of the gesture action feedback from the built-in accelerometer and gyroscope of the headset;

[0054] Extract features from the collected historical multimodal data; fuse the features extracted from different modal data to construct a feature vector of historical multimodal data;

[0055] The method of feature fusion includes weighted fusion, which assigns different weights to feature vectors of different modes and then performs weighted summation, wherein the weights are set by those skilled in the art according to the importance or reliability of the features;

[0056] Specifically, for visual data, image recognition technology is used to extract features of gestures in the video, including but not limited to: the shape, trajectory, and amplitude of the gesture; for audio data, audio signal processing algorithms are used to extract features such as the frequency, amplitude, and duration of the sound; sensor data is filtered and then feature extracted, including but not limited to: acceleration peak value and angular velocity change rate;

[0057] Using the fused historical multimodal data feature vectors, a multimodal data feature library is constructed;

[0058] Select the recurrent neural network RNN ​​to train the group user basic model. The specific process is as follows:

[0059] Normalize the multimodal data to unify the data of different modalities into the same scale range; divide the processed data into training set, validation set and test set, where the division ratio is 70%, 15% and 15%;

[0060] Determine the number of layers in the network; set the number of neurons in each layer;

[0061] Define loss functions and optimizers. The loss function measures the difference between the model prediction result and the true label. Loss functions include but are not limited to: cross entropy loss function; optimizers include but are not limited to: stochastic gradient descent (SGD), Adam optimizer;

[0062] Divide the training set data into several batches, each batch contains a certain number of samples;

[0063] For each batch of data, the feature vector is input into the RNN model and the prediction result is output;

[0064] Calculate the loss between the predicted result and the true label using the selected loss function;

[0065] Through the back-propagation algorithm, the loss value is back-propagated from the output layer to each layer of the network, and the gradient of each parameter is calculated; the gradient represents the rate of change of the loss function to the parameter, and the direction of parameter update can be determined by the gradient;

[0066] Use the optimizer to update the model parameters based on the calculated gradients;

[0067] After each round of training, the model is evaluated using the validation set data, and the feature vector of the validation set is input into the currently trained model to obtain the prediction result;

[0068] Calculate the evaluation indicators on the validation set, including but not limited to: accuracy, recall, F1 value, etc.; measure the performance of the model on the validation set;

[0069] Perform model tuning based on the evaluation results of the validation set;

[0070] Adjust the model's hyperparameters, such as learning rate, batch size, etc., retrain and evaluate, and observe changes in model performance until a good set of hyperparameter combinations is found so that the model has good performance on the validation set;

[0071] Repeat the above training cycle and validation set evaluation and tuning steps for multiple rounds of training until the model performance on the validation set reaches a satisfactory level or the preset training stop condition is reached, thus obtaining a group user basic model;

[0072] Combine the multimodal data feature library and the group user basic model to obtain a dynamic gesture library;

[0073] The technical solution of this embodiment is: by collecting the headphone usage behavior data (including gesture actions and functional operations) of the target group users, firstly screen out high-frequency candidate gestures based on the number of times used, and construct a preliminary high-frequency candidate set; then introduce Miller's law, evaluate whether the total number of candidate gestures exceeds the user's short-term memory range from the perspective of cognitive load, analyze the functional semantic relevance of the exceeding part through natural language processing and cosine similarity algorithm, sort and merge gestures with strongly related functions by correlation, and finally output a high-frequency gesture set that conforms to the user's cognitive habits. For high-frequency gestures, collect historical multimodal data of group users, construct a multimodal data feature library after feature extraction and fusion, and combine the recurrent neural network to train the group user basic model. Finally, combine the feature library with the model to form a dynamic gesture library, so as to provide basic data and model support for subsequent individual user adaptation and real-time recognition.

[0074] Therefore, the total number of gestures can be controlled and related functions can be merged through Miller's Law to ensure that the operation gestures are simple and easy to remember, reduce the memory difficulties caused by too many gestures, and improve the convenience of user operation; gestures can be screened based on the high-frequency behavior data of group users to ensure that the core functions of the headset match the real needs of users, reduce low-frequency function interference, and improve the efficiency of function use.

[0075] Embodiment 2

[0076] Based on the above embodiments, Figure 1 As shown, an earphone control method based on multimodal interactive gesture recognition provided by an embodiment of the present invention also includes S3-S4, and the process is as follows:

[0077] S3: Obtain historical usage data of individual users, perform deviation analysis on the historical data and the basic model of group users to obtain differential features, and build an individual user adaptation model based on the differential features;

[0078] In this step, historical usage data of individual users is obtained based on the group user basic model, wherein the historical usage data of individual users is the same as the headphone behavior data obtained by the group users;

[0079] Preprocess the acquired historical data of individual users, including format unification, exception handling, and missing value filling;

[0080] Based on the preprocessed historical data of individual users, the same features as those used in the training of the group user basic model are extracted to construct the multimodal data feature vector of individual users;

[0081] Match and compare the multimodal data feature vector of individual users with the mapping relationship between the multimodal data features and corresponding labels learned by the group user basic model;

[0082] As a preferred method of this embodiment: the multimodal feature vector of an individual user is: , the group user feature vector is ;in, , i=1,2,…,m,n are feature dimensions;

[0083] Use distance measurement method to calculate the deviation between individual user historical usage data and group user basic model;

[0084] For the multimodal feature vector X of an individual user and a feature vector in the group user basic model , the distance between the two is calculated by Euclidean distance : ;

[0085] Among them, the calculated multimodal feature vector X and all feature vectors in the group user basic model are The Euclidean distance of ;

[0086] Take the mean of all distances in the distance set: ; It is used to measure the comprehensive difference between all characteristic dimensions of individual users and group users, reflecting the overall deviation between individuals and groups;

[0087] For each individual feature dimension, calculate the deviation between it and the mean of the distance of the corresponding all features as the feature deviation value;

[0088] It should be noted that the feature deviation value is the absolute value of the deviation between the mean of the distance between the individual feature and the corresponding group feature;

[0089] Set a deviation threshold. If the feature deviation value is greater than the deviation threshold, the corresponding feature dimension is marked as a difference feature. Otherwise, the corresponding feature dimension is marked as a non-difference feature.

[0090] For example, it is assumed that the gesture data of an individual user includes three features: gesture amplitude, gesture speed and gesture duration, that is, the feature vector , there are 100 feature vectors in the group user base model, namely ;

[0091] For each individual feature dimension, calculate the difference between the individual user's feature value and the mean value of the feature dimension of the group users. If the mean value of gesture amplitude is 30 and the gesture amplitude of the individual user is 40, then the feature deviation value is 10; if the mean value of gesture speed is 20 and the gesture speed of the individual user is 22, then the feature deviation value is 2; if the mean value of gesture duration is 1.5 and the gesture duration of the individual user is 1.8, then the feature deviation value is 0.3;

[0092] As another preferred method of this embodiment, the cosine similarity method can be used to measure the similarity of the two vector directions;

[0093] Based on the difference features, combined with the historical usage data of individual users and the corresponding usage functions, a data set for training the individual user adaptation model is constructed;

[0094] Select the same recurrent neural network as the group user basic model for model training; the specific training process is the same as the group user basic model training process, which will not be repeated here;

[0095] Train to obtain an individual user adaptation model;

[0096] By building an individual user adaptation model, the operation habits and feature differences of individual users are taken into account, so that the model can more accurately identify the gestures of individual users. Compared with the general model, the individual user adaptation model can better adapt to the unique gesture characteristics of different users, thereby significantly improving the accuracy of gesture recognition and providing users with a smoother and more efficient interactive experience. For example, for users with larger gestures, the model can more accurately identify their gestures and reduce misjudgments.

[0097] By finding out the difference features through deviation analysis and constructing a training data set based on the difference features, the training of the individual user adaptation model is more targeted. Compared with the group user basic model, the individual user adaptation model can better capture the characteristic information of individual users, thereby improving the performance and generalization ability of the model;

[0098] S4: Based on the individual user adaptation model, the multimodal weights are dynamically adjusted through incremental learning to output the final gesture recognition results;

[0099] In this step, multimodal data is collected in real time and preprocessed. The preprocessing includes: formatting, denoising, and outlier removal of real-time data to ensure that the data quality is consistent with the feature dimensions of group and individual model training;

[0100] Extract features from the collected real-time multimodal data; fuse the features extracted from different modal data to construct a real-time multimodal data feature vector;

[0101] Input the real-time multimodal feature vector into the individual user adaptation model and output the predicted usage function;

[0102] Calculate the deviation between the real-time feature vector and the historical feature distribution of the individual user adaptation model, set the incremental learning threshold, and trigger the incremental learning mechanism if the deviation exceeds the incremental learning threshold; otherwise, it will not be triggered.

[0103] Add real-time multimodal data to the historical data set of individual users to form an incremental training set ;

[0104] The gradient descent algorithm is used to dynamically adjust the weights of multimodal features with real-time recognition accuracy as the objective function. The specific process is as follows:

[0105] Calculate the recognition loss under the current weight: ; Where L represents the recognition loss, which is used to measure the difference between the model prediction result and the true label. Indicates the use of function, represents the prediction value of the individual user adaptation model;

[0106] Back propagate the loss and calculate the weight gradient: ;in, represents the gradient of weight w, which represents the rate of change of loss function L to weight w. Represents the partial derivative of weight w;

[0107] Update the weights using an optimizer such as Adam: ;in, represents the learning rate, represents the weight after the t+1th update, Indicates the current weight for the tth time

[0108] Synchronize verified real-time gesture data and its optimized weight parameters to the dynamic gesture library to ensure that the model adapts to the real-time habit changes of individual users;

[0109] The recognized gesture is used as the final gesture recognition result and mapped to the usage function (i.e., the headphone control function), and the control command is sent to the headphone via Bluetooth or wired connection;

[0110] The technical solution of this embodiment is as follows: by obtaining the historical usage data of individual users, extracting the same features as the basic model of group users after preprocessing, constructing the individual multimodal data feature vector, calculating the deviation between individual and group features through Euclidean distance, screening out the difference features and constructing a training data set accordingly, using recurrent neural network training to obtain an individual user adaptation model, realizing personalized fitting of individual operation habits, collecting multimodal data in real-time scenarios, constructing real-time feature vectors after preprocessing and inputting them into the prediction function of the individual adaptation model; at the same time, triggering the incremental learning mechanism by calculating the deviation between the real-time features and the historical distribution, integrating the new data into the training set, optimizing the multimodal weights through gradient descent, and finally mapping the recognition results to the headphone control function and sending instructions, realizing the dynamic adaptation of the model to the real-time habits of users and scene changes;

[0111] By analyzing differential features, the gesture differences between individual users and groups are captured, allowing the model to adapt to personalized operations such as light touches by the elderly and quick swipes by athletes, solving the recognition bias caused by the "one-size-fits-all" approach of traditional general models. The incremental learning mechanism absorbs new user gesture data in real time, automatically updates model weights and dynamic gesture libraries, and does not require manual configuration by the user. As usage time increases, the model's fit to user habits continues to improve.

[0112] Embodiment 3

[0113] Based on the above embodiments, Figure 2 As shown, an earphone control system based on multimodal interactive gesture recognition provided by an embodiment of the present invention specifically includes:

[0114] High-frequency gesture screening module: collects the behavior data of group users, screens and analyzes the collected behavior data, and outputs a set of high-frequency gestures;

[0115] In this embodiment, the group of users is determined according to the target user group characteristics of the headset, wherein the group characteristics include but are not limited to: age, gender, occupation;

[0116] Set the collection time to collect the headphone behavior data of group users in different usage scenarios, where the headphone behavior data includes: gestures and usage functions;

[0117] Preprocess the collected behavioral data; construct a gesture and function association matrix based on the processed behavioral data, and count the number of times each gesture is used;

[0118] Set a usage limit; for each gesture, compare it with the set usage limit, retain gestures with a value greater than the usage limit, mark them as high-frequency candidate gestures, and construct a preliminary high-frequency candidate set based on the high-frequency candidate gestures;

[0119] Based on the preliminary high-frequency candidate set, Miller's law is introduced to perform secondary screening on each high-frequency candidate gesture. The specific process is as follows:

[0120] From the perspective of cognitive load, the preliminary high-frequency candidate set is evaluated to determine whether the total number of high-frequency candidate gestures exceeds the short-term memory range of the group users, where the short-term memory range of the group users is when the total number of gestures exceeds 7±2;

[0121] If it exceeds the scope of the group user's short-term memory, a correlation analysis is performed on the used functions to determine whether the used gestures can be merged. The specific process is as follows:

[0122] Use natural language processing technology to perform semantic analysis on the description of each usage function;

[0123] The semantic correlation between the word vectors corresponding to any two functional combinations is calculated using the cosine similarity algorithm, and the calculated cosine similarity is used as the association value;

[0124] Setting a correlation threshold, and comparing the correlation value with the correlation threshold;

[0125] Arrange the function combinations corresponding to the correlation threshold values ​​from large to small, and merge the gestures in order of ranking until the total number of gestures does not exceed the short-term memory range of the group users, thus obtaining a high-frequency gesture set.

[0126] Gesture library construction module: Based on the frequently used gesture collection, a multimodal data feature library and a group user basic model are constructed to obtain a dynamic gesture library;

[0127] In this embodiment, for each gesture in the high-frequency gesture set, historical multimodal data of group users in different usage scenarios are collected;

[0128] Extract features from the collected historical multimodal data; fuse the features extracted from different modal data to construct a feature vector of historical multimodal data;

[0129] Using the fused historical multimodal data feature vectors, a multimodal data feature library is constructed;

[0130] Normalize the multimodal data to unify the data of different modalities into the same scale range; divide the processed data into training set, validation set and test set;

[0131] Determine the number of layers in the network; set the number of neurons in each layer;

[0132] Define loss functions and optimizers. The loss function measures the difference between the model prediction result and the true label. Loss functions include but are not limited to: cross entropy loss function; optimizers include but are not limited to: stochastic gradient descent (SGD), Adam optimizer;

[0133] Divide the training set data into several batches, each batch contains a certain number of samples;

[0134] For each batch of data, the feature vector is input into the RNN model and the prediction result is output;

[0135] Calculate the loss between the predicted result and the true label using the selected loss function;

[0136] Through the back-propagation algorithm, the loss value is back-propagated from the output layer to each layer of the network, and the gradient of each parameter is calculated; the gradient represents the rate of change of the loss function to the parameter, and the direction of parameter update can be determined by the gradient;

[0137] Use the optimizer to update the model parameters based on the calculated gradients;

[0138] After each round of training, the model is evaluated using the validation set data, and the feature vector of the validation set is input into the currently trained model to obtain the prediction result;

[0139] Calculate the evaluation indicators on the validation set; perform model tuning based on the evaluation results of the validation set;

[0140] Adjust the model's hyperparameters, retrain and evaluate, and observe changes in model performance until a good set of hyperparameter combinations is found so that the model has good performance on the validation set;

[0141] Repeat the above training cycle and validation set evaluation and tuning steps for multiple rounds of training until the model performance on the validation set reaches a satisfactory level or the preset training stop condition is reached, thus obtaining a group user basic model;

[0142] Combine the multimodal data feature library and the group user basic model to obtain a dynamic gesture library;

[0143] Individual deviation analysis module: obtains historical usage data of individual users, performs deviation analysis on the historical data and the basic model of group users to obtain difference features, and builds an individual user adaptation model based on the difference features;

[0144] In this embodiment, historical usage data of individual users is obtained based on the group user basic model;

[0145] Preprocess the acquired historical data of individual users;

[0146] Based on the preprocessed historical data of individual users, the same features as those used in the training of the group user basic model are extracted to construct the multimodal data feature vector of individual users;

[0147] Match and compare the multimodal data feature vector of individual users with the mapping relationship between the multimodal data features and corresponding labels learned by the group user basic model;

[0148] Use distance measurement method to calculate the deviation between individual user historical usage data and group user basic model;

[0149] For each individual feature dimension, calculate the deviation between it and the mean of the distance of the corresponding all features as the feature deviation value;

[0150] Set a deviation threshold. If the feature deviation value is greater than the deviation threshold, the corresponding feature dimension is marked as a difference feature. Otherwise, the corresponding feature dimension is marked as a non-difference feature.

[0151] Based on the difference features, combined with the historical usage data of individual users and the corresponding usage functions, a data set for training the individual user adaptation model is constructed;

[0152] Select the same recurrent neural network as the group user base model for model training;

[0153] Train to obtain an individual user adaptation model;

[0154] Gesture recognition output module: Based on the individual user adaptation model and real-time scenario, it dynamically adjusts the multimodal weights through incremental learning and outputs the final gesture recognition results;

[0155] In this embodiment, multimodal data is collected in real time and preprocessed;

[0156] Extract features from the collected real-time multimodal data; fuse the features extracted from different modal data to construct a real-time multimodal data feature vector;

[0157] Input the real-time multimodal feature vector into the individual user adaptation model and output the predicted usage function;

[0158] Calculate the deviation between the real-time feature vector and the historical feature distribution of the individual user adaptation model, set the incremental learning threshold, and trigger the incremental learning mechanism if the deviation exceeds the incremental learning threshold; otherwise, it will not be triggered.

[0159] Add real-time multimodal data to the historical data set of individual users to form an incremental training set;

[0160] The gradient descent algorithm is used to dynamically adjust the weights of multimodal features with real-time recognition accuracy as the objective function. The specific process is as follows:

[0161] Calculate the recognition loss under the current weight: ; Where L represents the recognition loss, which is used to measure the difference between the model prediction result and the true label. Indicates the use of function, represents the prediction value of the individual user adaptation model;

[0162] Back propagate the loss and calculate the weight gradient: ;in, represents the gradient of weight w, which represents the rate of change of loss function L to weight w. Represents the partial derivative of weight w;

[0163] Update the weights using the optimizer: ;in, represents the learning rate, represents the weight after the t+1th update, Indicates the current weight for the tth time

[0164] Synchronize verified real-time gesture data and its optimized weight parameters to the dynamic gesture library to ensure that the model adapts to the real-time habit changes of individual users;

[0165] The recognized gesture is used as the final gesture recognition result output and mapped to the usage function, and the control command is sent to the headset via Bluetooth or wired connection.

[0166] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0167] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A headset control method based on multimodal interactive gesture recognition, characterized in that: The following steps are involved: Collect the behavioral data of group users, filter and analyze the collected behavioral data based on the number of gestures used and Miller's law, and output a set of high-frequency gestures; Based on the frequently used gesture collection, a multimodal data feature library and a group user basic model are constructed to obtain a dynamic gesture library; Obtain historical usage data of individual users, perform deviation analysis on historical data and the basic model of group users to obtain differential features, and build an individual user adaptation model based on the differential features; Based on the individual user adaptation model, the multimodal weights are dynamically adjusted through incremental learning to output the final gesture recognition results.

2. The headphone control method based on multimodal interactive gesture recognition according to claim 1, characterized in that: The process of screening and analyzing the collected behavioral data is as follows: Collect gestures and usage functions of group users; analyze the number of times gestures are used, screen high-frequency candidate gestures, and construct a preliminary high-frequency candidate set; Based on the preliminary high-frequency candidate set, the number and correlation of gesture actions are analyzed in combination with Miller's law, and high-frequency gestures are screened out, and a high-frequency gesture set is constructed.

3. The headphone control method based on multimodal interactive gesture recognition according to claim 2, characterized in that: The process of acquiring the high-frequency candidate gestures is as follows: Preprocess the gesture action and function data; construct a correlation matrix between gestures and functions, and count the number of times each gesture is used during the collection time; For each gesture, the gestures with a usage count greater than a limit are marked as high-frequency candidate gestures.

4. The headset control method based on multimodal interactive gesture recognition according to claim 2, characterized in that: The process of high-frequency gesture use includes: Based on the preliminary high-frequency candidate set, Miller's law is introduced to evaluate whether the total number of high-frequency candidate gestures exceeds the short-term memory of group users from the perspective of cognitive load; If it exceeds the short-term memory of the group of users, a correlation analysis is performed on the used functions to determine whether the used gestures can be merged.

5. The headphone control method based on multimodal interactive gesture recognition according to claim 4, characterized in that: The process of determining whether the gestures can be merged is as follows: Get the association value of the gestures, arrange the function combinations corresponding to the ones greater than the correlation threshold in order from large to small, and merge the gestures in order of ranking until the total number of gestures does not exceed the short-term memory range of the group users, and get a high-frequency gesture set.

6. The headphone control method based on multimodal interactive gesture recognition according to claim 4, characterized in that: Use natural language processing technology to perform semantic analysis on the description of usage functions; The semantic correlation between the word vectors corresponding to any two functional combinations is calculated using the cosine similarity algorithm, and the cosine similarity is used as the association value.

7. The headphone control method based on multimodal interactive gesture recognition according to claim 1, characterized in that: The dynamic gesture library includes a multimodal data feature library and a group user basic model.

8. The headphone control method based on multimodal interactive gesture recognition according to claim 1, characterized in that: For each gesture in the set of frequently used gestures, historical multimodal data of group users is collected; Perform feature extraction on the collected historical multimodal data, fuse the features extracted from different modal data, obtain the feature vector of historical multimodal data, and build a multimodal data feature library; A recurrent neural network is selected based on historical multimodal data to train a group user base model.

9. The headphone control method based on multimodal interactive gesture recognition according to claim 1, characterized in that: The process of performing deviation analysis to obtain difference features is as follows: Based on the group user basic model, obtain the historical usage data of individual users; Extract the same features as those used in training the group user base model and construct the multimodal data feature vector of individual users; Use distance measurement method to calculate the deviation between individual user historical usage data and group user basic model; For each individual feature dimension, calculate the deviation between it and the mean of the distance of the corresponding all features as the feature deviation value; If the feature deviation value is greater than the deviation threshold, the corresponding feature dimension is marked as a difference feature.

10. The headphone control method based on multimodal interactive gesture recognition according to claim 1, characterized in that: The process of constructing the individual user adaptation model is as follows: Based on the difference features, combined with the historical usage data of individual users and the corresponding usage functions, a data set for training the individual user adaptation model is constructed; Select a recurrent neural network for model training to obtain an individual user adaptation model.

11. The headphone control method based on multimodal interactive gesture recognition according to claim 1, characterized in that: The process of outputting the final gesture recognition result is as follows: Collect multimodal data in real time to build real-time multimodal data feature vectors; then input the individual user adaptation model and output the predicted usage function; Calculate the deviation between the real-time feature vector and the historical feature distribution of the individual user adaptation model. If the deviation exceeds the incremental learning threshold, the incremental learning mechanism is triggered. Add real-time multimodal data to the historical data set of individual users to form an incremental training set; The gradient descent algorithm is used to dynamically adjust the weights of multimodal features with real-time recognition accuracy as the objective function.

12. A headset control system based on multimodal interactive gesture recognition, characterized in that: The system is used to execute the method described in any one of claims 1 to 11, and the system comprises: High-frequency gesture screening module: collects the behavioral data of group users, screens and analyzes the collected behavioral data based on the number of gestures used and Miller's law, and outputs a set of high-frequency gestures; Gesture library construction module: Based on the frequently used gesture collection, a multimodal data feature library and a group user basic model are constructed to obtain a dynamic gesture library; Individual deviation analysis module: obtains historical usage data of individual users, performs deviation analysis on the historical data and the basic model of group users to obtain difference features, and builds an individual user adaptation model based on the difference features; Gesture recognition output module: Based on the individual user adaptation model, the multimodal weights are dynamically adjusted through incremental learning to output the final gesture recognition results.