Machine learning model training method, device and equipment for target tasks

By constructing a sample data set of multimodal data and using a combination method of feature extraction and classification networks, the problem of insufficient recognition accuracy of single modal data is solved, and higher recognition accuracy and robustness are achieved.

CN115049077BActive Publication Date: 2025-05-20BEIJING BEYONCA INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210635268.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-05-20
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

The prior art relies solely on single modal data in personnel emotion recognition, driver driving status recognition and personnel health status recognition, resulting in insufficient information and easy to be interfered with by external information, and the identification accuracy is difficult to guarantee.

Method used

By constructing a sample data set of multimodal data, multiple feature extraction networks are used to extract data of different modes, and these feature vectors are input into the classification network to obtain the classification prediction results of the target task, and adjust the model parameters based on the label of the sample data.

Benefits of technology

The accuracy of target task recognition results is improved, and the robustness and accuracy of the recognition process are enhanced through the analysis of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115049077B_ABST
    Figure CN115049077B_ABST
Patent Text Reader

Abstract

A method, device and apparatus for training a machine learning model for a target task are provided. The method comprises: obtaining a sample data set for a target task; performing the following operations on each sample data in the sample data set: extracting a feature vector of each sub-sample data in a plurality of sub-sample data of the sample data from the sub-sample data through a feature extraction network corresponding to the modality of the sub-sample data in a plurality of feature extraction networks; obtaining a classification result of the sample data through a classification network based on the feature vectors corresponding to the plurality of sub-sample data of the sample data; and adjusting the parameters of the classification network and the parameters of the plurality of feature extraction networks based on the classification result and the target task category label of the sample data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and particularly to a machine learning model training method, apparatus, computer device, vehicle, computer-readable storage medium, and computer program product for a target task. Background Art

[0002] With the increasing popularity of new energy intelligent vehicles in China, intelligent perception in the cockpit has become a very important direction. Intelligent perception and control in the vehicle field play a crucial role in the user experience. Especially in automatic trigger scenarios, through intelligent perception, it is possible to adaptively adjust relevant control parameters of the vehicle based on the habits of different users and the perception of different users, thereby enhancing the user experience. Summary of the Invention

[0003] According to one aspect of the present disclosure, there is provided a machine learning model training method for a target task, including: obtaining a sample data set for the target task, wherein each sample data in the sample data set includes a plurality of sub-sample data and a class label corresponding to the sample data, the plurality of sub-sample data respectively have a plurality of corresponding different modalities, the machine learning model includes a plurality of feature extraction networks and a classification network, and the plurality of feature extraction networks respectively correspond to the plurality of modalities; for each sample data in the sample data set, performing operations including the following: for each of the plurality of sub-sample data of the sample data, extracting a feature vector of the sub-sample data from the sub-sample data through the feature extraction network corresponding to the modality of the sub-sample data among the plurality of feature extraction networks; obtaining a classification result of the sample data through the classification network based on the feature vectors corresponding to the plurality of sub-sample data of the sample data; and adjusting the parameters of the classification network and the parameters of the plurality of feature extraction networks based on the classification result and the target task class label of the sample data.

[0004] According to another aspect of the present disclosure, there is provided a recognition method for a target task, wherein the target task includes any one of personnel emotion recognition, driver driving state recognition, and personnel health state recognition, and the method includes: obtaining a plurality of data for the target task; and using a machine learning model to recognize the plurality of data to obtain recognition results of the plurality of data output by the machine learning model, wherein the machine learning model is trained according to the above-mentioned machine learning model training method for the target task, and wherein the plurality of data respectively have a plurality of corresponding different modalities.

[0005] According to another aspect of the present disclosure, there is provided a machine learning model training apparatus for a target task, including: a first acquisition unit configured to acquire a sample data set for the target task, wherein each sample data in the sample data set includes a plurality of sub-sample data and a class label corresponding to the sample data, the plurality of sub-sample data respectively have a plurality of corresponding different modalities, and the machine learning model includes a plurality of feature extraction networks and a classification network, and the plurality of feature extraction networks respectively correspond to the plurality of modalities; an execution unit configured to perform operations including the following sub-units on each sample data in the sample data set, wherein the execution unit includes: an extraction sub-unit configured to, for each of the plurality of sub-sample data of the sample data, extract a feature vector of the sub-sample data from the sub-sample data through the feature extraction network corresponding to the modality of the sub-sample data among the plurality of feature extraction networks; a first acquisition sub-unit configured to obtain a classification result of the sample data through the classification network based on the feature vectors corresponding to the plurality of sub-sample data of the sample data; and an adjustment sub-unit configured to adjust the parameters of the classification network and the parameters of the plurality of feature extraction networks based on the classification result and the target task class label of the sample data.

[0006] According to still another aspect of the present disclosure, there is provided an identification apparatus for a target task, wherein the target task includes any one of personnel emotion identification, driver driving state identification, and personnel health state identification. The apparatus includes: a second acquisition unit configured to acquire a plurality of data for the target task; and an identification unit configured to use a machine learning model to identify the plurality of data to obtain an identification result of the plurality of data output by the machine learning model, wherein the machine learning model is trained according to the above-mentioned machine learning model training method for the target task, and wherein the plurality of data respectively have a plurality of corresponding different modalities.

[0007] According to another aspect of the present disclosure, there is provided a computer device, including: at least one processor; and at least one memory storing a computer program thereon, wherein when the computer program is executed by the at least one processor, the at least one processor is caused to execute the above-mentioned machine learning model training method for the target task.

[0008] According to another aspect of the present disclosure, there is provided a vehicle including the above-mentioned identification apparatus for the target task or the above-mentioned computer device.

[0009] According to another aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program thereon, and when the computer program is executed by a processor, the processor is caused to execute the above-mentioned machine learning model training method for the target task or the identification method for the target task.

[0010] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, causes the processor to execute the machine learning model training method for a target task or the recognition method for a target task as described above.

[0011] According to the embodiments described hereinafter, these and other aspects of the present disclosure will be apparent and will be elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features and advantages of the present disclosure are disclosed, in which:

[0013] Figure 1 is a schematic diagram illustrating an example system in which various methods described herein can be implemented according to an exemplary embodiment;

[0014] Figure 2 is a flowchart illustrating a machine learning model training method 200 for a target task according to an exemplary embodiment;

[0015] Figure 3 is a flowchart illustrating a method 300 for obtaining a sample data set for a target task according to an exemplary embodiment;

[0016] Figure 4 is a schematic diagram illustrating the network structure of an emotion recognition model according to an exemplary embodiment;

[0017] Figure 5 is a flowchart illustrating a method 500 for generating a sample data set based on a plurality of first sample data and their first sample labels according to another exemplary embodiment;

[0018] Figure 6 is a schematic diagram illustrating the network structure of a machine learning model for a target task (such as human emotion recognition) according to an exemplary embodiment;

[0019] Figure 7 is a flowchart illustrating a method 700 for obtaining a classification result of the sample data through a classification network according to an exemplary embodiment;

[0020] Figure 8 is a flowchart illustrating a recognition method 800 for a target task according to another exemplary embodiment;

[0021] Figure 9 is a schematic block diagram illustrating a machine learning model training apparatus 900 for a target task according to an exemplary embodiment;

[0022] Figure 10FIG. 0 is a schematic block diagram of an identification device 1000 for a target task according to an exemplary embodiment;

[0023] Figure 11 FIG. 1 is a block diagram of an exemplary computer device that can be applied to an exemplary embodiment. DETAILED DESCRIPTION

[0024] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the description of the context, they may also refer to different instances.

[0025] In the description of the various examples in the present disclosure, the terms used are only for the purpose of describing a specific example and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. As used herein, the term "plurality" means two or more, and the term "based on" should be interpreted as "at least partially based on". In addition, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations.

[0026] In the related art, the identification and analysis of target tasks such as personnel emotion recognition, driver driving state recognition, and personnel health state recognition are usually based on single-modal data, so as to identify the personnel emotion, driving state, health state, etc. in real time and provide corresponding services for users. However, using only single-modal data for the identification and analysis of target tasks usually has problems such as insufficient information, being prone to overgeneralization, and being easily interfered by external information, and it is difficult to guarantee the accuracy of the identification.

[0027] According to an embodiment of the present disclosure, a machine learning model training method for a target task (such as personnel emotion recognition, driver driving state recognition, personnel health state recognition, etc.) is provided. By constructing a sample data set with multi-modal data, feature extraction is performed on each modality of data in the sample data based on a feature extraction network corresponding to different modalities, and each feature vector is simultaneously input into a classification network, so as to obtain a classification prediction result of the target task and complete model training in combination with the label corresponding to the sample data. Thus, the trained model can analyze multi-modal data to obtain an identification result during the identification process of the target task, thereby improving the accuracy of the target task identification result.

[0028] The exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0029] Figure 1 FIG. is a schematic diagram showing an example system 100 in which various methods described herein can be implemented according to an exemplary embodiment.

[0030] Referring to Figure 1 , the system 100 includes an in-vehicle system 110, a server 120, and a network 130 communicatively coupling the in-vehicle system 110 and the server 120.

[0031] The in-vehicle system 110 includes a display 114 and an application program (APP) 112 that can be displayed via the display 114. The application program 112 can be an application program pre-installed by default for the in-vehicle system 110 or downloaded and installed by the user 102, or a mini-program as a lightweight application program. In the case where the application program 112 is a mini-program, the user 102 can directly run the application program 112 on the in-vehicle system 110 without installing the application program 112 by searching for the application program 112 in the host application (e.g., by the name of the application program 112, etc.) or scanning the graphic code of the application program 112 (e.g., barcode, QR code, etc.). In some embodiments, the in-vehicle system 110 can include one or more processors and one or more memories (not shown), and the in-vehicle system 110 is implemented as an in-vehicle computer. In some embodiments, the in-vehicle system 110 can include more or fewer displays 114 (e.g., without including the display 114), and / or one or more speakers or other human-computer interaction devices. In some embodiments, the in-vehicle system 110 may not communicate with the server 120.

[0032] The server 120 can represent a single server, a cluster of multiple servers, a distributed system, or a cloud server providing basic cloud services such as cloud databases, cloud computing, cloud storage, and cloud communication. It will be understood that although Figure 1 shows the server 120 communicating with only one in-vehicle system 110, the server 120 can provide background services for multiple in-vehicle systems simultaneously.

[0033] The network 130 allows wireless communication and information exchange between vehicle-X ("X" means vehicle, road, pedestrian, or Internet, etc.) according to agreed communication protocols and data interaction standards. Examples of the network 130 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. The network 130 can be a wired or wireless network. In one example, the network 130 can be an in-vehicle network, an inter-vehicle network, and / or an in-vehicle mobile Internet.

[0034] For the purposes of the embodiments of the present disclosure, in Figure 1In the example, the application 112 can be one or more of a personnel emotion recognition program based on multimodal data, a driver driving state recognition program, and a personnel health state recognition program. The above applications can respectively provide functions such as emotion recognition, driving state recognition, or personnel health state recognition of personnel based on multimodal data. Correspondingly, the server 120 can be a server used together with the above applications. The server 120 can send the updated and online recognition model based on multimodal data to the application 112 running in the vehicle-mounted system 110 to update the model applied by the application 112. Alternatively, the server 120 can also send the above model to the vehicle-mounted system 110, and the vehicle-mounted system 110 will complete the deployment of the model and support the call of the application 112.

[0035] Figure 2 FIG. shows a flowchart of a machine learning model training method 200 for a target task according to an exemplary embodiment. The method 200 can be executed at a vehicle-mounted system (e.g., Figure 1 the vehicle-mounted system 110 shown in ), that is, the execution entity of each step of the method 200 can be Figure 1 the vehicle-mounted system 110 shown in. In some embodiments, the method 200 can be executed at a server (e.g., Figure 1 the server 120 shown in ). In some embodiments, the method 200 can be executed in combination by a vehicle-mounted system (e.g., the vehicle-mounted system 110) and a server (e.g., the server 120). Hereinafter, taking the execution entity as the vehicle-mounted system 110 as an example, each step of the method 200 will be described in detail.

[0036] Referring to Figure 2 , the method 200 includes steps S210 to S240.

[0037] Step S210: Obtain a sample data set for a target task. Each sample data in the sample data set includes a plurality of sub-sample data and a class label corresponding to the sample data. The plurality of sub-sample data respectively have a plurality of corresponding different modalities. The machine learning model includes a plurality of feature extraction networks and a classification network, and the plurality of feature extraction networks respectively correspond to the plurality of modalities.

[0038] For each sample data in the sample data set, perform an operation including the following steps:

[0039] Step S220: For each sub-sample data in the plurality of sub-sample data of the sample data, extract the feature vector of the sub-sample data from the sub-sample data through the feature extraction network corresponding to the modality of the sub-sample data in the plurality of feature extraction networks;

[0040] Step S230: Obtain the classification result of the sample data through a classification network based on the feature vectors corresponding to multiple sub-sample data of the sample data; and

[0041] Step S240: Adjust the parameters of the classification network and the parameters of multiple feature extraction networks based on the classification result and the target task category label of the sample data.

[0042] According to the above embodiments, by constructing a sample data set with multi-modal data, performing feature extraction on each modality of data in the sample data based on feature extraction networks corresponding to different modalities, and simultaneously inputting each feature vector into the classification network, the classification prediction result of the target task can be obtained, and the model training can be completed in combination with the label corresponding to the sample data. Thus, the trained model can analyze multi-modal data to obtain the recognition result during the recognition process for the target task, thereby improving the accuracy of the target task recognition result.

[0043] In some embodiments, the target task may include any one of personnel emotion recognition, driver driving state recognition, and personnel health state recognition.

[0044] It can be understood that those skilled in the relevant art can also set the target task according to the actual classification requirements based on the above method of the present disclosure, and train the corresponding model based on the target task, which is not limited herein. In the following, the method of the present disclosure will be specifically described by taking emotion recognition as an example.

[0045] As used herein, the term "modality" may refer to the source, medium, or form of information describing the target object. The source of information includes, for example, touch, hearing, vision, and smell; the medium of information includes, for example, speech, video, and text; the form of information includes, for example, radar signals, infrared signals, and acceleration signals. In some embodiments, the multiple different-modal sub-sample data included in each sample data may include, for example, face visual data, pose visual data, speech data, text data, biological signal data, etc.

[0046] Among them, face visual data and pose visual data can be obtained by collecting video data through an RGB camera device, speech data and text data can be obtained by collecting audio data, and biological signal data can be electrocardiogram (ECG) signals and photoplethysmography (PPG) signals collected by an infrared camera or radar, etc. After collecting the above data through different devices respectively, it is first necessary to align the above data in terms of time and space. The alignment in space can adopt the method of three-dimensional space calibration to eliminate the parallax problem of the camera device caused by position deviation.

[0047] After aligning the data, each frame of the video data can be further processed to extract face image and pose visual data (for example, by identifying and marking the human skeleton to extract limb movements), thereby obtaining corresponding face visual data and pose visual data. At the same time, the audio data can be further processed for speech recognition to extract semantic information in the audio and save it in the form of text data. Subsequently, the above-mentioned collected data are respectively annotated to obtain a sample data set for the target task.

[0048] It is understandable that relevant technical personnel can also set the types of different modality data obtained according to actual needs, which is not limited here.

[0049] Figure 3 The flowchart of method 300 for obtaining a sample data set for a target task according to an exemplary embodiment is shown.

[0050] Reference Figure 3 , method 300 includes step S310 to step S340. The specific steps are as follows:

[0051] Step S310, obtain a plurality of first sample data, where each first sample data in the plurality of first sample data includes a plurality of unlabeled first sub-sample data, and among the plurality of first sub-sample data, each has a plurality of modalities;

[0052] For each first sample data in the plurality of first sample data, perform an operation including the following steps:

[0053] Step S320, for each first sub-sample data in the plurality of first sub-sample data in this first sample data, obtain a first annotation result of this first sub-sample data through a pre-trained target task classification model corresponding to the modality of this first sub-sample data; and

[0054] Step S330, unify the first annotation results corresponding to the plurality of first sub-sample data in this first sample data to obtain a first sample label corresponding to this first sample data, where the first sample label is one of a plurality of preset labels; and

[0055] Step S340, generate a sample data set based on the plurality of first sample data and the plurality of first sample labels respectively corresponding to the plurality of first sample data.

[0056] In some embodiments, first, a plurality of sample data of different modalities can be collected in the above manner, and data alignment and extraction of face visual data, pose visual data, and text data can be performed. Subsequently, the above multi-modal data can be stored according to the corresponding relationship in time and space to obtain a plurality of first sample data.

[0057] In some embodiments, for data of different modalities, emotion recognition models corresponding to each modality can be constructed respectively. For example, for face visual data and pose visual data, emotion recognition models can be constructed based on networks such as CNN, VGGNet, and R-CNN. For audio data, biosignal data, text data, etc., emotion recognition models can be constructed based on network structures such as UBM-GMM, SVM, DNN, CNN, LSTM, Conformer, and TDNN. For text data, emotion recognition models can also be constructed based on network structures such as TextCNN and DPCNN. It is understandable that relevant technical personnel can select the network structures for constructing emotion recognition models for each modality of data by themselves, and no limitation is made here.

[0058] Figure 4 A schematic diagram showing the network structure of an emotion recognition model according to an exemplary embodiment is shown.

[0059] As Figure 4 shown, emotion recognition models of different modalities can all include an input layer, multiple hidden layers (hidden layer h1, hidden layer h2, …, hidden layer hn), and an output layer. After constructing the emotion recognition models for each modality of data, relevant publicly available datasets can be applied respectively for training the emotion recognition models, so that they respectively have the ability to recognize emotions based on the data of the corresponding modality, thereby obtaining pre-trained emotion recognition models corresponding to the modality (that is, pre-trained target task classification models).

[0060] The unlabeled sample data of each modality collected above are respectively used for emotion recognition through the emotion recognition models of the corresponding modality, so as to obtain the emotion labels (that is, the first annotation results) of each sample data of each modality.

[0061] Generally, since the types and quantities of emotion labels of sample data in the publicly available datasets corresponding to each modality are different, the emotion labels obtained by the emotion recognition models trained based on these publicly available datasets for annotating the data of each modality also have different types and quantities of types. Therefore, in order to construct an emotion recognition model for training based on multi-modal data, it is necessary to unify the emotion labels corresponding to each modality of data. For example, the emotion labels corresponding to each modality of data can be unified into six emotion labels such as fear, sadness, anger, pleasure, surprise, and disgust (that is, multiple preset labels), and the emotion labels corresponding to each modality of data are adjusted adaptively accordingly. It is understandable that relevant technical personnel can set the types and quantities of emotion labels based on actual needs by themselves, and no limitation is made here.

[0062] In some cases, for data of different modalities with spatio-temporal correspondence (e.g., data of different modalities at the same time and corresponding to the same user), there may be differences in the emotion recognition results. Therefore, based on a voting mechanism, the above emotion recognition results can be unified. For example, for the five-modal data corresponding to the above spatio-temporal, emotion recognition is performed respectively on them. If 3 of the obtained emotion labels are "pleased" and 2 are "surprised", then based on the voting mechanism, the emotion label corresponding to the five-modal data is unified as "pleased".

[0063] Thus, the annotation of each first sample data is achieved based on the above method. Based on the multiple annotated first sample data, a sample data set for the target task can be generated. Among them, each sample data in the sample data set corresponds to each first sample data respectively, and the class label corresponding to the sample data corresponds to the emotion label of the first sample data obtained through the above method.

[0064] Figure 5 The flowchart of a method 500 for generating a sample data set based on multiple first sample data and their first sample labels according to another exemplary embodiment is shown.

[0065] Reference Figure 5 , method 500 includes steps S510 to S540. Among them, the multiple first sample data can be temporally continuous and arranged in time sequence. The specific steps of method 500 include:

[0066] Step S510: Based on a preset sliding step and a preset time sliding window, divide the multiple first sample data into multiple first sample data subsets;

[0067] For each first sample data subset in the multiple first sample data subsets, perform an operation including the following steps:

[0068] Step S520: Based on multiple first sub-sample data in each modality of the first sample data subset, generate second sub-sample data corresponding to the first sample data subset and in this modality; and

[0069] Step S530: Based on the multiple first sample labels respectively corresponding to the multiple first sample data in the first sample data subset, determine the second sample label corresponding to the first sample data subset; and

[0070] Step S540: Based on the multiple second sub-sample data corresponding to each first sample data subset in the multiple first sample data subsets and the multiple second sample labels respectively corresponding to the multiple first sample data subsets, generate a sample data set.

[0071] Thus, by further organizing and merging the labeled first sample data and its labels, the temporal continuity of the sample data is ensured, and the accuracy of the sample data and its labels is further improved; applying the sample data processed as described above for model training can further improve the classification accuracy of the model.

[0072] In some embodiments, based on a preset sliding step and a preset time sliding window, multiple first sample data within the preset time sliding window can be integrated. For example, the window length of the preset time sliding window can be 5 seconds, and the preset sliding step can be 1 second. Based on the above sliding window and step, multiple first sample data can be respectively integrated into multiple first sample data subsets, where the multiple first sample data subsets respectively correspond to the first sample data within the first 5 seconds, the first sample data within the second to sixth seconds... and so on for the first sample data within 5 seconds in sequence.

[0073] In some embodiments, for data of different modalities in each first sample data subset, the corresponding first sub-sample data of each modality can be spliced according to time sequence to generate second sub-sample data corresponding to each modality.

[0074] In some embodiments, multiple first sample data in each first sample data subset may respectively correspond to different emotion labels. Therefore, based on a voting mechanism, these different emotion labels can be unified, and the emotion label with the largest number is used as the second sample label corresponding to the first sample data subset. For example, a first sample data subset includes 5 first sample data, among which 3 first sample data have the label "pleased", and 2 first sample data have the label "surprised". Then, based on the voting mechanism, the emotion label corresponding to this first sample data subset is unified as "pleased".

[0075] Based on the above method, after integrating the first sample data and its labels, a sample data set can be further generated. Among them, the sub-sample data of different modalities in each sample data in the sample data set respectively correspond to the multiple second sub-sample data in the above first sample data subsets, and the label of this sample data corresponds to the second sample label corresponding to this first sample data subset.

[0076] In some embodiments, after performing corresponding preprocessing on sample data of different modalities, the preprocessed data can be respectively input into a machine learning model for a target task for model training. Among them, the preprocessing of audio data can include extracting its acoustic features (such as Mel frequency cepstral parameter features, constant Q transform cepstral parameter features, etc.), and performing mean normalization and differential processing on the above acoustic features; the preprocessing of text data can include operations such as word segmentation on the text data.

[0077] Figure 6 FIG. shows a schematic diagram of the network structure of a machine learning model for a target task (such as human emotion recognition) according to an exemplary embodiment.

[0078] In some embodiments, as Figure 6 shown, the model includes a feature extraction network corresponding to each modality and a classification network. Among them, the feature extraction network of each modality can be respectively obtained based on the pre-trained target task classification model of the corresponding modality. In the example, the output layer of the pre-trained target task classification model can be removed, and its input layer and multiple hidden layers can be retained as the feature extraction network. In some embodiments, the output layer of the pre-trained target task classification model and one or more hidden layers adjacent to the output layer can also be removed, and its input layer and the remaining multiple hidden layers can be retained as the feature extraction network.

[0079] Through the above-mentioned feature extraction networks corresponding to different modalities, the input sub-sample data is subjected to feature extraction, and the feature vectors corresponding to each modality can be obtained accordingly.

[0080] Figure 7 FIG. shows a flowchart of a method 700 for obtaining the classification result of the sample data through a classification network according to an exemplary embodiment.

[0081] Referring to Figure 7 , method 700 includes steps S710 to S730. The specific steps include:

[0082] Step S710: Obtain the target feature value in each feature vector of the corresponding feature vector;

[0083] Step S720: Perform numerical normalization on the feature vector based on the target feature value in each feature vector of the corresponding feature vector; and

[0084] Step S730: Input the normalized corresponding feature vector into the classification network to obtain the classification result of the sample data output by the classification network.

[0085] Since the numerical ranges of the feature vectors extracted by the feature extraction networks of different modalities often vary greatly, the feature vectors with larger numerical values often have a greater impact on the model, and features with too small numerical values may even be ignored by the model. Therefore, to avoid this problem, before inputting the feature vectors of each modality into the classification network, the feature vectors of each modality are normalized to eliminate the above differences, thereby improving the performance of the model.

[0086] In some embodiments, the target eigenvalue can be the difference between the maximum value and the minimum value of each eigenvector, and the eigenvector is numerically normalized based on this difference so that each eigenvalue of the eigenvector is unified to the interval [0, 1]. In some embodiments, it is also possible to first obtain the difference between the global maximum value and the global minimum value of the eigenvectors corresponding to each modality as the target eigenvalue, and perform a numerical normalization operation on the eigenvectors of the corresponding modality based on this value.

[0087] In some embodiments, the normalized multiple eigenvectors corresponding to each sample data can be concatenated to obtain a fused eigenvector, and this vector is input into a classification network (which can be a Softmax layer, for example), so as to obtain the classification result corresponding to this sample data, and the parameters of each feature extraction network and the classification network are respectively adjusted based on this classification result and the label of this sample data, thereby completing the training of the above model.

[0088] In some embodiments, the machine learning model training method for the target task may further include: determining the learning rate of the feature extraction network corresponding to each modality based on the amount of sample data in each modality in the sample dataset, and wherein, adjusting the parameters of the classification network and the parameters of the multiple feature extraction networks based on the classification result and the target task category label of this sample data may include: for each feature extraction network among the multiple feature extraction networks, adjusting the parameters of this feature extraction network based on the classification result, the target task category label of this sample data, and the learning rate of this feature extraction network.

[0089] In some embodiments, since there are differences in the amount of data of different modalities that can be collected, for example, the amount of visual data and audio data is often higher than that of other modality data, and the feature extraction network corresponding to the modality with a large amount of data converges faster than the feature extraction network corresponding to the modality with a small amount of data. To balance the training speed of each network, different learning rates can be set for different networks respectively. For example, the feature extraction networks corresponding to face visual data, pose visual data, and audio data are set to a smaller learning rate, and the feature extraction network corresponding to text data is set to a larger learning rate, so as to synchronize the training progress of each network and improve the training efficiency of the model.

[0090] Figure 8 FIG. shows a flowchart of an identification method 800 for a target task according to another exemplary embodiment.

[0091] Reference Figure 8 , method 800 includes steps S810 to step S820. Among them, the target task includes any one of personnel emotion recognition, driver driving state recognition, and personnel health state recognition. The specific steps of method 800 include:

[0092] Step S810: Obtain multiple data for a target task; and

[0093] Step S820: Use a machine learning model to identify the multiple data to obtain an identification result of the multiple data output by the machine learning model, where the machine learning model is trained according to the above-mentioned machine learning model training method for the target task, and where the multiple data respectively have corresponding different multiple modalities.

[0094] Thus, the model obtained based on the above training method can, during the identification process for the target task, analyze multi-modal data to obtain an identification result, thereby improving the accuracy of the target task identification result.

[0095] Although each operation is depicted in the drawings as being in a specific order, this should not be construed as requiring that these operations must be performed in the specific order shown or in a sequential order, nor should it be construed as requiring that all the operations shown must be performed to obtain the desired result.

[0096] Figure 9 FIG. is a schematic block diagram showing a machine learning model training apparatus 900 for a target task according to an exemplary embodiment.

[0097] As Figure 9 shown, the machine learning model training apparatus 900 for a target task may include:

[0098] A first acquisition unit 910, configured to acquire a sample data set for a target task, where each sample data in the sample data set includes multiple sub-sample data and a class label corresponding to the sample data, the multiple sub-sample data respectively have corresponding different multiple modalities, the machine learning model includes multiple feature extraction networks and a classification network, and the multiple feature extraction networks respectively correspond to the multiple modalities;

[0099] An execution unit 920, configured to perform operations including the following respective sub-units on each sample data in the sample data set, where the execution unit 920 includes:

[0100] An extraction sub-unit 921, configured to, for each of the multiple sub-sample data of the sample data, extract a feature vector of the sub-sample data from the sub-sample data through the feature extraction network corresponding to the modality of the sub-sample data among the multiple feature extraction networks;

[0101] A first acquisition sub-unit 922, configured to obtain a classification result of the sample data through the classification network based on the feature vectors corresponding to the multiple sub-sample data of the sample data; and

[0102] An adjustment subunit 923, configured to adjust parameters of a classification network and parameters of a plurality of feature extraction networks based on a classification result and a target task category label of the sample data.

[0103] Wherein, Figure 9 Each unit of the device 900 shown in Figure 2 can correspond to each step in the method 200 described in the reference. Thus, the operations, features, and advantages described above for the method 200 also apply to the device 900 and its included units. For the sake of brevity, certain operations, features, and advantages are not described herein again.

[0104] Figure 10 is a schematic block diagram showing an identification device 1000 for a target task according to an exemplary embodiment.

[0105] As Figure 10 shown, the identification device 1000 for a target task may include:

[0106] A second acquisition unit 1010, configured to acquire a plurality of data for a target task; and

[0107] An identification unit 1020, configured to identify the plurality of data by using a machine learning model to obtain an identification result of the plurality of data output by the machine learning model, wherein the target task includes any one of personnel emotion recognition, driver driving state recognition, and personnel health state recognition, the machine learning model is trained according to the above-mentioned machine learning model training method for a target task, and wherein the plurality of data respectively have corresponding different modalities.

[0108] Wherein, Figure 10 each unit of the device 1000 shown in Figure 8 can correspond to each step in the method 800 described in the reference. For the sake of brevity, certain operations, features, and advantages are not described herein again.

[0109] Although specific functions have been discussed above with reference to specific modules, it should be noted that the functions of the various modules discussed herein can be divided into multiple modules, and / or at least some functions of multiple modules can be combined into a single module. The actions performed by the specific modules discussed herein include the specific module itself performing the action, or alternatively the specific module calling or otherwise accessing another component or module that performs the action (or performs the action in combination with the specific module). Thus, a specific module that performs an action can include the specific module itself that performs the action and / or another module that the specific module calls or otherwise accesses and that performs the action. As used herein, the phrase "entity A initiates action B" can mean that entity A issues an instruction to perform action B, but entity A itself does not necessarily perform action B.

[0110] It should also be understood that various technologies may be described herein in the general context of software-hardware elements or program modules. As described above with respect to Figure 9 and 10 each unit described may be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units may be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these units may be implemented as hardware logic / circuits. The SoC may include an integrated circuit chip (which includes one or more components of a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or other circuits), and may optionally execute the received program code and / or include embedded firmware to perform functions.

[0111] According to one aspect of the present disclosure, there is provided a computer device including at least one memory, at least one processor, and a computer program stored on the at least one memory. The at least one processor is configured to execute the computer program to implement the steps of any of the method embodiments described above.

[0112] According to one aspect of the present disclosure, there is provided a vehicle including the device or computer device as described above.

[0113] According to one aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the steps of any of the method embodiments described above.

[0114] According to one aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, it implements the steps of any of the method embodiments described above.

[0115] Hereinafter, illustrative examples of such computer devices, non-transitory computer-readable storage media, and computer program products will be described in conjunction with Figure 11 An example configuration of a computer device 1100 that may be used to implement the methods described herein is shown. By way of example,

[0116] Figure 11 Figure 1 Figure 1The server 120 and / or the vehicle system 110 shown in may include an architecture similar to that of the computer device 1100. The above-mentioned device 900 or device 1000 may also be implemented in whole or at least in part by the computer device 1100 or a similar device or system.

[0117] The computer device 1100 may include at least one processor 1102, a memory 1104, (multiple) communication interfaces 1106, a display device 1108, other input / output (I / O) devices 1110, and one or more mass storage devices 1112 that can communicate with each other, such as via a system bus 1114 or other suitable connections.

[0118] The processor 1102 may be a single processing unit or multiple processing units, and all processing units may include a single or multiple computing units or multiple cores. The processor 1102 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operation instructions. Among other capabilities, the processor 1102 may be configured to obtain and execute computer-readable instructions stored in the memory 1104, the mass storage device 1112, or other computer-readable media, such as the program code of an operating system 1116, the program code of an application 1118, the program code of other programs 1120, etc.

[0119] The memory 1104 and the mass storage device 1112 are examples of computer-readable storage media for storing instructions that are executed by the processor 1102 to implement the various functions described above. For example, the memory 1104 generally may include both volatile and non-volatile memories (such as RAM, ROM, etc.). In addition, the mass storage device 1112 generally may include a hard disk drive, a solid-state drive, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical discs (such as CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. The memory 1104 and the mass storage device 1112 may both be collectively referred to as memory or computer-readable storage media herein, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code that can be executed by the processor 1102 as a specific machine configured to implement the operations and functions described in the examples herein.

[0120] Multiple programs can be stored on the mass storage device 1112. These programs include an operating system 1116, one or more application programs 1118, other programs 1120, and program data 1122, and they can be loaded into the memory 1104 for execution. Examples of such application programs or program modules can include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functions: Method 200 and / or Method 800 (including any suitable steps of Methods 200, 800), and / or additional embodiments described herein.

[0121] Although illustrated as being stored in the memory 1104 of the computer device 1100 in Figure 11 , the modules 1116, 1118, 1120, and 1122 or portions thereof can be implemented using any form of computer-readable medium accessible by the computer device 1100. As used herein, "computer-readable medium" includes at least two types of computer-readable media, namely computer-readable storage media and communication media.

[0122] Computer-readable storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computer device. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism. Computer-readable storage media as defined herein does not include communication media.

[0123] One or more communication interfaces 1106 are used to exchange data with other devices, such as via a network, direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., network interface card (NIC)), wired or wireless (such as IEEE 802.11 wireless LAN (WLAN)) wireless interface, Worldwide Interoperability for Microwave Access (WiMAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth TMCommunication interfaces, such as a near field communication (NFC) interface, etc. The communication interface 1106 can facilitate communication within a variety of network and protocol types, including wired networks (such as LAN, cable, etc.) and wireless networks (such as WLAN, cellular, satellite, etc.), the Internet, etc. The communication interface 1106 can also provide communication with external storage devices (not shown) such as those in storage arrays, network attached storage, storage area networks, etc.

[0124] In some examples, a display device 1108, such as a monitor, etc., can be included for displaying information and images to a user. Other I / O devices 1110 can be devices that receive various inputs from a user and provide various outputs to the user, and can include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, etc.

[0125] The techniques described herein can be supported by these various configurations of the computer device 1100 and are not limited to the specific examples of the techniques described herein. For example, the functionality can also be implemented in whole or in part on a "cloud" using a distributed system. The cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud. Resources can include applications and / or data that can be used when performing computational processing on servers remote from the computer device 1100. Resources can also include services provided over the Internet and / or over a subscriber network such as a cellular or Wi-Fi network. The platform can abstract the resources and functionality to connect the computer device 1100 with other computer devices. Thus, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality can be implemented partially on the computer device 1100 and partially through a platform that abstracts the functionality of the cloud.

[0126] Although the present disclosure has been illustrated and described in detail in the accompanying drawings and the foregoing description, such illustration and description should be considered illustrative and schematic rather than restrictive; the present disclosure is not limited to the disclosed embodiments. By studying the drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments when practicing the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps not listed, the indefinite article "a" or "an" does not exclude a plurality, the term "plurality" means two or more, and the term "based on" should be interpreted as "at least partially based on". The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

Claims

1. A method for training a machine learning model for a target task, comprising: A sample data set for the target task is obtained, wherein each sample data in the sample data set includes a plurality of sub-sample data and a category label corresponding to the sample data, the plurality of sub-sample data respectively have a plurality of corresponding different modalities, the plurality of sub-sample data include face visual data, posture visual data, voice data, text data and bio-signal data, and the plurality of sub-sample data correspond in time and space, the machine learning model includes a plurality of feature extraction networks and a classification network, the plurality of feature extraction networks respectively correspond to the plurality of modalities, the target task includes any one of personnel emotion recognition, driver driving state recognition and personnel health state recognition, and obtaining the sample data set for the target task includes: Acquire a plurality of first sample data, wherein each of the plurality of first sample data includes a plurality of unlabeled first sub-sample data, wherein the plurality of first sub-sample data respectively have the plurality of modalities; For each of the plurality of first sample data, performing the following operations: For each first sub-sample data in the plurality of first sub-sample data in the first sample data, Obtaining a first labeling result of the first sub-sample data by using a pre-trained target task classification model corresponding to the modality of the first sub-sample data; and unifying the first annotation results corresponding to the multiple first sub-sample data in the first sample data to obtain a first sample label corresponding to the first sample data, wherein the first sample label is one of multiple preset labels, and the multiple first annotation results corresponding to the multiple first sub-sample data are unified through a voting mechanism to obtain the first sample label; and Based on the plurality of first sample data and the plurality of first sample labels respectively corresponding to the plurality of first sample data, generating the sample data set, wherein the plurality of first sample data have temporal continuity and are arranged in time sequence, and the generating the sample data set based on the plurality of first sample data and the plurality of first sample labels respectively corresponding to the plurality of first sample data comprises: Based on a preset sliding step size and a preset time sliding window, the plurality of first sample data are divided into a plurality of first sample data subsets; For each of the plurality of first sample data subsets, performing operations including the following: Based on the multiple first sub-sample data in each modality in the first sample data subset, splicing is performed in time sequence to generate second sub-sample data in the modality corresponding to the first sample data subset; and Determine a second sample label corresponding to the first sample data subset based on a plurality of first sample labels respectively corresponding to a plurality of first sample data in the first sample data subset, and unify the plurality of first sample labels through a voting mechanism to obtain the second sample label; and Generate the sample data set based on a plurality of second sub-sample data corresponding to each of the plurality of first sample data subsets and a plurality of second sample labels respectively corresponding to the plurality of first sample data subsets, wherein the plurality of second sub-sample data corresponding to the first sample data subset corresponds to a plurality of sub-sample data in the sample data, and the second sample label corresponding to the first sample data subset corresponds to a category label of the sample data; For each sample data in the sample data set, perform the following operations: For each sub-sample data of the plurality of sub-sample data of the sample data, extracting a feature vector of the sub-sample data from the sub-sample data by using a feature extraction network corresponding to a mode of the sub-sample data among the plurality of feature extraction networks; Based on the feature vectors corresponding to the plurality of sub-sample data of the sample data, obtaining a classification result of the sample data through the classification network; and Based on the classification result and the target task category label of the sample data, the parameters of the classification network and the parameters of the multiple feature extraction networks are adjusted.

2. The method according to claim 1, wherein: The obtaining the classification result of the sample data through the classification network based on the feature vectors corresponding to the plurality of sub-sample data of the sample data comprises: Obtaining a target eigenvalue in each of the corresponding eigenvectors; performing numerical normalization on each of the corresponding feature vectors based on a target feature value in the feature vector; and The normalized corresponding feature vector is input into the classification network to obtain the classification result of the sample data output by the classification network.

3. The method according to claim 1, further comprising: Based on the amount of sample data under each modality in the sample data set, determining the learning rate of the feature extraction network corresponding to each modality, Wherein, adjusting the parameters of the classification network and the parameters of the plurality of feature extraction networks based on the classification result and the target task category label of the sample data includes: For each feature extraction network in the multiple feature extraction networks, the parameters of the feature extraction network are adjusted based on the classification result, the target task category label of the sample data, and the learning rate of the feature extraction network.

4. A method for identifying a target task, wherein: The target task includes any one of person emotion recognition, driver driving status recognition, and person health status recognition, and the method includes: Acquire a plurality of data for the target task; and Using a machine learning model to identify the multiple data to obtain identification results of the multiple data output by the machine learning model, Wherein, the machine learning model is trained according to the method according to any one of claims 1 to 3, and Wherein, the multiple data respectively have the corresponding different multiple modalities.

5. A machine learning model training device for a target task, comprising: A first acquisition unit is configured to acquire a sample data set for the target task, wherein each sample data in the sample data set includes a plurality of sub-sample data and a category label corresponding to the sample data, the plurality of sub-sample data respectively have a plurality of corresponding different modalities, the plurality of sub-sample data include face visual data, posture visual data, voice data, text data and bio-signal data, and the plurality of sub-sample data correspond in time and space, the machine learning model includes a plurality of feature extraction networks and a classification network, the plurality of feature extraction networks correspond to the plurality of modalities respectively, the target task includes any one of personnel emotion recognition, driver driving state recognition and personnel health state recognition, and the acquisition of the sample data set for the target task includes: Acquire a plurality of first sample data, wherein each of the plurality of first sample data includes a plurality of unlabeled first sub-sample data, wherein the plurality of first sub-sample data respectively have the plurality of modalities; For each of the plurality of first sample data, performing the following operations: For each first sub-sample data in the plurality of first sub-sample data in the first sample data, Obtaining a first labeling result of the first sub-sample data by using a pre-trained target task classification model corresponding to the modality of the first sub-sample data; and unifying the first annotation results corresponding to the multiple first sub-sample data in the first sample data to obtain a first sample label corresponding to the first sample data, wherein the first sample label is one of multiple preset labels, and the multiple first annotation results corresponding to the multiple first sub-sample data are unified through a voting mechanism to obtain the first sample label; and Based on the plurality of first sample data and the plurality of first sample labels respectively corresponding to the plurality of first sample data, generating the sample data set, wherein the plurality of first sample data have temporal continuity and are arranged in time sequence, and the generating the sample data set based on the plurality of first sample data and the plurality of first sample labels respectively corresponding to the plurality of first sample data comprises: Based on a preset sliding step size and a preset time sliding window, the plurality of first sample data are divided into a plurality of first sample data subsets; For each of the plurality of first sample data subsets, performing operations including the following: Based on the multiple first sub-sample data in each modality in the first sample data subset, splicing is performed in time sequence to generate second sub-sample data in the modality corresponding to the first sample data subset; and Determine a second sample label corresponding to the first sample data subset based on a plurality of first sample labels respectively corresponding to a plurality of first sample data in the first sample data subset, and unify the plurality of first sample labels through a voting mechanism to obtain the second sample label; and Generate the sample data set based on a plurality of second sub-sample data corresponding to each of the plurality of first sample data subsets and a plurality of second sample labels respectively corresponding to the plurality of first sample data subsets, wherein the plurality of second sub-sample data corresponding to the first sample data subset corresponds to a plurality of sub-sample data in the sample data, and the second sample label corresponding to the first sample data subset corresponds to a category label of the sample data; The execution unit is configured to execute the operations of the following sub-units for each sample data in the sample data set, wherein the execution unit includes: an extraction subunit, configured to extract, for each sub-sample data of the plurality of sub-sample data of the sample data, a feature vector of the sub-sample data from the sub-sample data by using a feature extraction network corresponding to a mode of the sub-sample data among the plurality of feature extraction networks; A first acquisition subunit is configured to acquire a classification result of the sample data through the classification network based on the feature vectors corresponding to the plurality of sub-sample data of the sample data; and The adjustment subunit is configured to adjust the parameters of the classification network and the parameters of the multiple feature extraction networks based on the classification result and the target task category label of the sample data.

6. A recognition device for a target task, wherein: The target task includes any one of personnel emotion recognition, driver driving status recognition, and personnel health status recognition, and the device includes: a second acquisition unit, configured to acquire a plurality of data for the target task; and an identification unit, configured to identify the plurality of data using a machine learning model to obtain identification results of the plurality of data output by the machine learning model, Wherein, the machine learning model is trained according to the method according to any one of claims 1 to 3, and Wherein, the multiple data respectively have the corresponding different multiple modalities.

7. A computer device comprising: at least one processor; as well as at least one memory having a computer program stored thereon, Wherein, when the computer program is executed by the at least one processor, the at least one processor executes the method according to any one of claims 1 to 4.

8. A vehicle comprising the apparatus as claimed in claim 6 or the computer device as claimed in claim 7.

9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 4.

10. A computer program product, comprising a computer program, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Video classification method, device and equipment based on multi-modal representation, and storage medium

    CN113762322A