Model training method and multi-modal data feature recognition method

By introducing a timing interactive attention mechanism and feedback mechanism into the feature recognition network, the problems of low feature recognition accuracy and poor generalization ability of multimodal data in the prior art are solved, and a more efficient and reliable feature recognition effect is achieved.

CN120030331APending Publication Date: 2025-05-23ZHONGSHAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226027.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to fully mine complementary information in multimodal data, ignore timing features, resulting in low feature recognition accuracy and inability to achieve effective feedback updates in complex scenarios, affecting generalization capabilities and extraction performance.

Method used

A model training method is proposed to capture the interaction dependence of multimodal data through the timing interactive attention mechanism, and introduce a feedback mechanism to update the model network parameters, and build a feature recognition network including feature pre-identification module, timing interactive attention mechanism and feature fine recognition module.

Benefits of technology

Through the timing interactive attention mechanism and feedback mechanism, the comprehensiveness, robustness and reliability of feature recognition are enhanced, and the accuracy and generalization ability of multimodal data feature recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030331A_ABST
    Figure CN120030331A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and a multi-modal data feature recognition method, and the model training method comprises the steps: constructing a feature recognition network which comprises a feature pre-recognition module, a time sequence interaction attention mechanism and a feature fine recognition module, and carrying out the pre-recognition of multi-modal data in the feature pre-recognition module, the pre-recognized feature set is calculated through a time sequence interaction attention mechanism to obtain a multi-modal coupling attention weight, and finally, a feature fine recognition module processes the pre-recognized feature set based on the time sequence interaction attention weight to obtain a fine recognition feature set; according to the method, a closed-loop feedback mechanism is introduced in the model training process, a comparison result is obtained by comparing the pre-recognition feature set and the fine recognition feature set, network parameters of the feature recognition network are updated according to the comparison result, dynamic optimization of the model is achieved, and comprehensiveness, robustness and reliability of feature recognition are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data feature processing technology, and in particular to a model training method and a multimodal data feature recognition method. Background Art

[0002] With the advent of the big data era, various data sources are highly heterogeneous and diverse. Traditional single-modal data processing methods are often unable to fully tap the complementary information between different modalities when faced with multiple data sources such as images, text, and voice. Many existing data feature recognition methods usually only focus on single-channel data and cannot fully utilize the rich information contained in multi-channel data. They also ignore time series features and are difficult to adapt to dynamic and time-varying scenarios, affecting feature recognition accuracy. Secondly, many existing methods are simple one-step recognition modes and cannot provide feedback updates based on feature recognition results and actual application conditions, affecting the generalization ability and extraction performance of feature recognition algorithms in complex scenarios. Summary of the invention

[0003] Based on this, the present invention aims to propose a model training method and a multimodal data feature recognition method, which captures the interactive dependencies between different time steps of multimodal data with sequence characteristics through a temporal interactive attention mechanism, and introduces a feedback mechanism to update the model network parameters.

[0004] In a first aspect, the present invention provides a model training method, wherein a model trained by the method is used for feature recognition of multimodal data, comprising:

[0005] Obtain multimodal data as a training sample set;

[0006] Construct a feature recognition network including a feature pre-recognition module, a temporal interactive attention mechanism and a feature fine recognition module. The multimodal data in the training sample set is processed by the feature pre-recognition module to obtain a pre-recognition feature set. The pre-recognition feature set is calculated by the temporal interactive attention mechanism to obtain a multimodal coupling attention weight. The feature fine recognition module processes the pre-recognition feature set based on the multimodal coupling attention weight to obtain a fine recognition feature set.

[0007] The pre-recognition feature set and the precise recognition feature set are compared to obtain a comparison result, the network parameters of the feature recognition network are updated according to the comparison result, and the trained feature recognition network is output as a multimodal feature recognition model.

[0008] Furthermore, the feature pre-recognition module includes at least two modality recognition channels corresponding to different modalities, and the multimodal data is passed through the feature pre-recognition module to obtain a pre-recognition feature set including:

[0009] Each modal data in the multimodal data is processed through the modality and its corresponding modal recognition channel to obtain a modal pre-recognition feature, and the modal pre-recognition features recognized by each modal recognition channel are combined into a pre-recognition feature set.

[0010] Furthermore, the feature pre-recognition module uses a convolutional neural network as the network architecture, and the modality pre-recognition feature is expressed as:

[0011] ,

[0012] Where n represents the modality recognition channel identifier, is the modal pre-recognition feature identified by modal recognition channel n at the t-th time slot, is the number of modal pre-recognition features identified by modal recognition channel n at the t-th time slot, is the first time slot modal identification channel n Features, is the convolutional neural network output function, is the activation function, is the attention coefficient of the feature pre-recognition module, ,in is the initial attention coefficient, represents the multimodal coupled attention matrix of the tth time slot, which is built-in as 1 in the feature pre-recognition module. Representation Matrix The diagonal matrix of the inner elements, Represents the bias term of the convolutional neural network.

[0013] Furthermore, the pre-recognized feature set is calculated through the temporal interactive attention mechanism to obtain the multimodal coupling attention weights including:

[0014] Calculate the similarity between any two modal pre-recognition features obtained through the same modal recognition channel, record it as the first similarity, and divide the modal pre-recognition features of the same modal recognition channel into several modal feature groups according to the first similarity;

[0015] Calculate the similarity between two modal feature groups belonging to different modal recognition channels, recorded as the second similarity, and divide the modal feature groups of each modal recognition channel into several interactive feature groups according to the second similarity. Any two modal feature groups in each interactive feature group do not belong to the same modal recognition channel.

[0016] The temporal interaction attention weight is calculated according to the number of features in the interactive feature group, and the channel self-attention weight is calculated according to the feature mapping of the pre-recognized feature set. The multimodal coupling attention weight is calculated by weighting the temporal interaction attention weight and the channel self-attention weight.

[0017] Furthermore, the first similarity calculation is expressed as follows:

[0018]

[0019] in, For the The time slot modal identification channel n and The similarity of features, , and Respectively The time slot modal identification channel n and The features of channel pre-identification, and is the weight coefficient, is the operator of waveform similarity between features, is the operator of temporal similarity between features.

[0020] Furthermore, the modal recognition features of the same modal recognition channel are divided into a number of modal feature groups according to the first similarity, including:

[0021] The modal recognition features whose first similarity is greater than a set threshold are divided into the same modal feature group.

[0022] Further, comparing the pre-recognition feature set and the precise recognition feature set to obtain a comparison result, and updating the network parameters of the feature recognition network according to the comparison result includes:

[0023] Compare the pre-identification feature set and the fine identification feature set, record the feature that exists in the fine identification feature set but not in the pre-identification feature set as the first feature, record the feature that exists in the pre-identification feature set but not in the fine identification feature set as the second feature, and record the recognition result of the first feature and / or the second feature as the comparison result;

[0024] When the first feature exists, the first feature is added to the pre-recognition feature set for updating, and the updated pre-recognition feature set is subjected to the temporal interactive attention mechanism to calculate the attention weight of the first feature. When the attention weight of the first feature is greater than the attention weight of the pre-recognition feature set, the network parameters of the feature pre-recognition module are updated; otherwise, the correlation between the first feature and the set feature triggering event is calculated, and the network parameters of the temporal interactive attention mechanism are updated according to the correlation;

[0025] When the second feature exists, the recognition error between the pre-recognition feature set and the fine recognition feature set is calculated, and the network parameters of the feature fine recognition module are updated according to the recognition error.

[0026] In a second aspect, the present invention provides a method for multimodal data feature recognition, comprising:

[0027] Acquire multimodal data to be identified;

[0028] The multimodal data is input into the multimodal feature recognition model trained in the first aspect for recognition to obtain multimodal features.

[0029] In a third aspect, the present invention provides a model training device for training a multimodal feature recognition model, comprising:

[0030] A sample acquisition module is used to acquire multimodal data as a training sample set;

[0031] A network construction training module is used to construct a feature recognition network including a feature pre-recognition module, a temporal interactive attention mechanism and a feature fine recognition module. The multimodal data in the training data set is processed by the feature pre-recognition module to obtain a pre-recognition feature set. The pre-recognition feature set is calculated by the temporal interactive attention mechanism to obtain a multimodal coupling attention weight. The feature fine recognition module processes the pre-recognition feature set based on the temporal interactive attention weight to obtain a fine recognition feature set.

[0032] The network update module is used to compare the pre-recognition feature set and the precise recognition feature set to obtain a comparison result, update the network parameters of the feature recognition network according to the comparison result, and output the trained feature recognition network as a multimodal feature recognition model.

[0033] In a fourth aspect, the present invention provides a multimodal data feature recognition device, comprising:

[0034] A data acquisition module, used to acquire multimodal data to be identified;

[0035] The feature recognition module is used to input the multimodal data into the multimodal feature recognition model trained by the device of the third aspect for recognition, so as to obtain multimodal features.

[0036] In a fifth aspect, the present invention provides an electronic device comprising a memory storing computer-executable instructions and a processor, wherein when the computer-executable instructions are executed by the processor, the device executes the model training method provided in the first aspect and / or each step of the multimodal feature recognition method provided in the second aspect.

[0037] In a sixth aspect, the present invention provides a readable storage medium storing a computer executable program, which, when executed, can implement the model training method provided in the first aspect and / or the various steps of the multimodal data feature recognition method provided in the second aspect.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The present invention proposes a model training method and a multimodal data feature recognition method, constructs a feature recognition network including a feature pre-recognition module, a temporal interactive attention mechanism and a feature fine recognition module, multimodal data is pre-recognized in the feature pre-recognition module, and the pre-recognized feature set is calculated by the temporal interactive attention mechanism to obtain a multimodal coupling attention weight, and finally the feature fine recognition module processes the pre-recognition feature set based on the temporal interactive attention weight to obtain a fine recognition feature set; the present invention introduces a closed-loop feedback mechanism in the model training process, obtains a comparison result by comparing the pre-recognition feature set and the fine recognition feature set, updates the network parameters of the feature recognition network according to the comparison result, realizes dynamic optimization of the model, and enhances the comprehensiveness, robustness and reliability of feature recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0041] Figure 1 It is the network architecture of the feature recognition model provided by the embodiment of the present invention;

[0042] Figure 2 is a flow chart for implementing the model training method provided by an embodiment of the present invention;

[0043] Figure 3 is a flowchart of a method for realizing multimodal data feature recognition provided by an embodiment of the present invention;

[0044] Figure 4 is a structural diagram of a model training device provided by an embodiment of the present invention;

[0045] Figure 5 is a structural diagram of a multimodal data feature recognition device provided by an embodiment of the present invention;

[0046] Figure 6 This is a diagram of the electronic device architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] The following embodiments of the present invention will provide a training method for a multimodal feature recognition model and its application in a multimodal data feature recognition process.

[0049] See also Figure 1 An embodiment of the present invention provides a network architecture of a multimodal feature recognition model, including a feature pre-recognition module 102, a temporal interactive attention mechanism 104 and a feature precision recognition module 106.

[0050] Specifically, the feature pre-recognition module 102 includes at least two modality recognition channels corresponding to different modalities, each modality recognition channel is independently responsible for recognizing a modality data, and each modality data in the multimodal data is processed through the modality recognition channel corresponding to the modality to obtain the modality pre-recognition feature. Figure 1 Illustratively, taking multimodal data including images and voiceprints as an example, the feature pre-recognition module 102 includes an image feature pre-recognition channel and a voiceprint feature pre-recognition channel. After the multimodal data enters the feature pre-recognition module, the image modal data and the voiceprint modal data respectively enter their corresponding modal recognition channels for feature pre-recognition.

[0051] The temporal interaction attention mechanism 104 includes a channel self-attention mechanism and a temporal interaction attention mechanism. The channel self-attention mechanism includes a channel attention network corresponding to each modality. The modal pre-recognition feature calculates the self-attention weight through the channel attention network corresponding to the modality. At the same time, the multi-modal pre-recognition feature set calculates the temporal interaction attention weight through the temporal interaction attention mechanism. The multi-modal coupling attention weight is calculated based on the self-attention weight and the temporal interaction attention weight, which is used to update the attention parameters of the feature recognition module.

[0052] At the feature precision recognition module 106 , the pre-recognition feature set is processed based on the multimodal coupling attention weights to fully exploit the temporal features and cross-modal information of the multimodal data to obtain precision recognition features.

[0053] The multimodal feature recognition model provided by the embodiment of the present invention also introduces a closed-loop feedback mechanism 108, which can obtain a comparison result by comparing the pre-recognition feature set and the precise recognition feature set, and update the network parameters of the feature recognition network according to the comparison result.

[0054] See also Figure 2 For the model architecture provided in the above-mentioned embodiment, an embodiment of the present invention provides a model training method, wherein the trained model is used for feature recognition of multimodal data, and comprises the following steps:

[0055] Step S210: Acquire multimodal data as a training sample set.

[0056] Step S220. Construct a feature recognition network including a feature pre-recognition module, a temporal interactive attention mechanism and a feature fine recognition module. The multimodal data in the training sample set is processed by the feature pre-recognition module to obtain a pre-recognition feature set. The pre-recognition feature set is calculated by the temporal interactive attention mechanism to obtain a multimodal coupling attention weight. The feature fine recognition module processes the pre-recognition feature set based on the multimodal coupling attention weight to obtain a fine recognition feature set.

[0057] Specifically, the feature pre-recognition module includes at least two modality recognition channels corresponding to different modalities, and each modality recognition channel receives a modality data stream of a corresponding modality. The image data and voiceprint data sets input in the time slot are and , where Indicates the number of image data streams, Indicates the number of voiceprint data streams. Indicates Time slot input image data stream, Indicates Time slot input A voiceprint data stream.

[0058] In a further embodiment, the feature pre-recognition module uses a convolutional neural network as the network architecture, and the modality pre-recognition features recognized by each modality recognition channel are expressed as:

[0059] ,

[0060] Where n represents the modality recognition channel identifier, is the modal pre-recognition feature identified by modal recognition channel n at the t-th time slot, is the number of modal pre-recognition features identified by modal recognition channel n at the t-th time slot, is the first time slot modal identification channel n Features, is the convolutional neural network output function, is the activation function, is the attention coefficient of the feature pre-recognition module, ,in is the initial attention coefficient, represents the multimodal coupling attention matrix of the t-th time slot, which is built-in as 1 in the feature pre-recognition module. Representation Matrix The diagonal matrix of the inner elements, Represents the bias term of the convolutional neural network.

[0061] Then the pre-recognition feature set can be expressed as .

[0062] Furthermore, the pre-identified feature set in step S220 is calculated by the temporal interactive attention mechanism to obtain the multimodal coupling attention weight, which includes the following steps:

[0063] Step S221. Calculate the similarity between any two modal recognition features obtained through the same modal recognition channel, record it as the first similarity, and divide the modal recognition features of the same modal recognition channel into several modal feature groups according to the first similarity.

[0064] In this step, for the modal recognition features obtained from the same modal recognition channel, the similarity between any two features is calculated, and similar features are divided into the same modal feature group.

[0065] Specifically, the first similarity calculation is expressed as follows:

[0066]

[0067] in, For the The time slot modal identification channel n and The similarity of features, , and Respectively The time slot modal identification channel n and The features of channel pre-identification, and is the weight coefficient, is the operator of waveform similarity between features, is the operator of temporal similarity between features.

[0068] The embodiment of the present invention divides the modal identification features whose first similarity is greater than a set threshold into the same modal feature group, that is, the similarity between two features is greater than the threshold Right now , then the two features are considered similar. For a series of features, if any two features are similar, they constitute a feature group.

[0069] Furthermore, if , then the first, second, and third features of the modal recognition channel n are considered to belong to the same feature group. Indicates mean calculation.

[0070] Step S222. Calculate the similarity between two modal feature groups belonging to different modal recognition channels, recorded as the second similarity, and divide the modal feature groups of each modal recognition channel into several interactive feature groups according to the second similarity. Any two modal feature groups in each interactive feature group do not belong to the same modal recognition channel.

[0071] Specifically, the second cross-modal similarity calculation is similar to the feature similarity calculation of the same modality in the aforementioned step. The features in the interactive feature group are similar to each other and can be considered as the same feature group.

[0072] Step S223. Calculate the temporal interaction attention weight according to the number of features in the interactive feature group, calculate the channel self-attention weight according to the feature mapping of the pre-identified features, and weight the temporal interaction attention weight and the channel self-attention weight to obtain the multimodal coupling attention weight.

[0073] Specifically, the temporal interaction attention weight is calculated as follows:

[0074]

[0075] in, Features The temporal interaction attention weights, Features The temporal interaction attention weight coefficient, Interaction feature group The number of features in represents the hth interactive feature group, feature Interaction feature group The characteristics in The set of temporal interaction attention weights of slot modality recognition channel n is .

[0076] The higher the value, the more channels have recognized the feature within a specific time range, that is, the feature is representative and typical, so the feature attention is high. The minimum value is 1, that is, only one channel recognizes this unique feature in the tth time slot, thus forming a feature group containing only this feature. Since this feature is bursty, the feature attention is low.

[0077] The channel self-attention weight calculation is expressed as:

[0078]

[0079] In the formula, For the The set of self-attention weights of slotted modality recognition channel n, , For the The time slot modal identification channel n The self-attention weight of each channel, In order to facilitate the Softmax function of network back propagation gradient calculation and smooth the result to the range of 0-1, the dimension Scaling is performed to avoid excessively large dot products and low gradients. , , Respectively The query matrix, key matrix and value matrix of slotted modal identification channel n.

[0080] The multimodal coupling attention weight obtained by weighting the temporal interaction attention weight and the channel self-attention weight can be expressed as follows:

[0081]

[0082] in, , represents the multi-channel coupled attention weight of the kth channel in the modality recognition channel n of the tth time slot.

[0083] Based on the multimodal coupling attention weights obtained in the above steps, the attention coefficient of the feature recognition module is corrected.

[0084] Specifically, the feature pre-recognition module and the feature fine recognition module can use the same network structure, and the feature extraction at the feature fine recognition module can be regarded as a feature extraction process after the attention coefficient of the network is corrected by the multi-channel coupled attention weight obtained in the above steps, which can be specifically expressed as follows:

[0085]

[0086] In the formula, is the corrected attention coefficient of the kth channel of modality recognition channel n, is the initial attention coefficient of the kth channel of modality recognition channel n, For the A precise identification feature set for time-slot multimodal data, For the The feature accurate recognition result of the i-th multimodal data in the time slot, I(t) is the number of features accurately recognized in the t-th time slot.

[0087] Step S230. Compare the pre-recognition feature set and the refined recognition feature set to obtain a comparison result, update the network parameters of the feature recognition network according to the comparison result, and output the trained feature recognition network as a multimodal feature recognition model.

[0088] Further, step S230 includes the following steps:

[0089] Compare the pre-identification feature set and the fine identification feature set, record the feature that exists in the fine identification feature set but not in the pre-identification feature set as the first feature, record the feature that exists in the pre-identification feature set but not in the fine identification feature set as the second feature, and record the recognition result of the first feature and / or the second feature as the comparison result;

[0090] When the first feature exists, the first feature is added to the pre-recognition feature set for updating, and the updated pre-recognition feature set is subjected to the temporal interactive attention mechanism to calculate the attention weight of the first feature. When the attention weight of the first feature is greater than the attention weight of the pre-recognition feature set, the network parameters of the feature pre-recognition module are updated; otherwise, the correlation between the first feature and the set feature triggering event is calculated, and the network parameters of the temporal interactive attention mechanism are updated according to the correlation;

[0091] When the second feature exists, the recognition error between the pre-recognition feature set and the fine recognition feature set is calculated, and the network parameters of the feature fine recognition module are updated according to the recognition error.

[0092] Specifically, the pre-recognition feature set is recorded as , the precise recognition feature set is recorded as If the feature recognition module recognizes a feature that the feature pre-recognition module does not recognize, , then add the feature to the pre-recognition feature set To update, , and re-input into the temporal interactive attention mechanism for attention recognition.

[0093] If the mode identification channel Medium Features The self-attention weight of is greater than the mean self-attention weight of the pre-recognized feature set, then the feature is considered The attention of the dominant position, namely:

[0094]

[0095] in, For modal identification channels Medium Features The self-attention weight, For modal identification channels The mean self-attention weight of the set of pre-recognized features in .

[0096] Update the network parameters of the feature pre-recognition module. Specifically, the multimodal coupling attention matrix Identification channel Features The corresponding self-attention weight is updated as .

[0097] If the feature Attention does not take precedence, that is , but based on the historical feature recognition results, the feature trigger event is set after the feature appears will occur, and no event is recognized The data features of For failure The potential features of the model are increased by adjusting the network parameters of the temporal interaction attention mechanism. attention, thereby increasing awareness of potential events Accuracy of early warning.

[0098] Specifically, we first calculate the features With events Correlation between occurrences , expressed as

[0099]

[0100] in, is the counting function, For events The set of historical feature recognition results that occurred, For events The total number of elements in the historical feature recognition result set that occurred. This formula means feature In Events The higher the frequency of occurrence in the historical feature recognition results, the higher the feature With events The greater the correlation between occurrences.

[0101] Furthermore, the network parameters of the temporal interaction attention mechanism are adjusted, which is expressed as

[0102]

[0103] In the formula, Features The temporal interaction attention weight coefficient, Features The temporal interaction attention weight coefficient, is the correlation threshold, is the update step size.

[0104] This formula indicates that when the characteristic With events Correlation between occurrences Greater than threshold At the time, Time slot increase feature The temporal interaction attention weight coefficient of attention.

[0105] If the feature pre-recognition module identifies a feature that occupies a dominant position , which is repeatedly filtered out by the feature precision recognition module, i.e., the precision recognition feature set Features not included If the number of times exceeds the set threshold, the network parameters of the feature recognition module are updated, which is expressed as

[0106]

[0107] In the formula, Represents the network architecture of the feature recognition module, including convolution kernel, bias, activation function, pooling layer parameters, weights and bias of the fully connected layer, etc. Representing image features Caused by identification errors.

[0108] See also Figure 3 An embodiment of the present invention provides a multimodal data feature recognition method, comprising the following steps:

[0109] Step S310: Acquire multimodal data to be identified.

[0110] Step S320: Input the multimodal data into the trained multimodal feature recognition model for recognition to obtain multimodal features.

[0111] The above method disclosed can be implemented by using various forms of equipment, so the present invention also discloses a device corresponding to the above method, and specific embodiments are given below for detailed description.

[0112] like Figure 4 As shown, an embodiment of the present invention provides a model training device for training a multimodal feature recognition model, comprising:

[0113] A sample acquisition module 402 is used to acquire multimodal data as a training sample set;

[0114] A network construction training module 404 is used to construct a feature recognition network including a feature pre-recognition module, a temporal interactive attention mechanism and a feature fine recognition module. The multimodal data in the training data set is processed by the feature pre-recognition module to obtain a pre-recognition feature set. The pre-recognition feature set is calculated by the temporal interactive attention mechanism to obtain a multimodal coupling attention weight. The feature fine recognition module processes the pre-recognition feature set based on the temporal interactive attention weight to obtain a fine recognition feature set.

[0115] The network update module 406 is used to compare the pre-recognition feature set and the precise recognition feature set to obtain a comparison result, update the network parameters of the feature recognition network according to the comparison result, and output the trained feature recognition network as a multimodal feature recognition model.

[0116] See also Figure 5 An embodiment of the present invention provides a multimodal data feature recognition device, comprising:

[0117] A data acquisition module 502 is used to acquire multimodal data to be identified;

[0118] The feature recognition module 504 is used to input the multimodal data into the multimodal feature recognition model trained by the aforementioned model training device for recognition, so as to obtain multimodal features.

[0119] The implementation principle and technical effects of the device provided in the embodiments of the present application are the same as those of the aforementioned method embodiments. For the sake of brief description, for matters not mentioned in the device embodiments, reference may be made to the corresponding contents in the aforementioned method embodiments.

[0120] The methods and related devices mentioned in the above embodiments are described with reference to the method flow charts and / or structural diagrams provided in the embodiments of the present application. Specifically, each process and / or block in the method flow charts and / or structural diagrams, as well as the combination of processes and / or blocks in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the process. Figure 1 A process or multiple processes and / or structures Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or structures Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A flow or multiple flows and / or structures illustrate the steps of the functions specified in one block or multiple blocks.

[0121] The following embodiments are described by taking the method applied to a computer device as an example. It can be understood that the computer device can be any device with computing and processing functions, and can be but not limited to a server or a personal laptop computer, etc. In one embodiment, the computer device can be an application server, and the application server can be a server for running an application to be tested.

[0122] See also Figure 6 , which shows a hardware block diagram of an electronic device, the electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0123] like Figure 6 As shown, the electronic device includes: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4;

[0124] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;

[0125] The processor 1 may be a central processing unit CPU, or an application-specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;

[0126] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), etc., such as at least one disk memory;

[0127] Among them, the memory stores a program, and the processor can call the program stored in the memory, and the program is used to: implement the model training method provided in the aforementioned embodiment, and / or each processing flow of the multimodal feature recognition method.

[0128] An embodiment of the present invention also provides a readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the model training method provided in the above-mentioned embodiment and / or any possible implementation method in combination with the embodiment, and / or the various processing flows of the multimodal feature recognition method.

[0129] The above-mentioned embodiments have described the present invention in particular detail with respect to possible scenarios, and those skilled in the art will recognize that the present invention can be practiced through other embodiments. The specific naming of components, the capitalization of terms, attributes, data structures, or any other programming or structural aspects are not mandatory or important, and the mechanisms or features of the present invention may have different names, forms, or procedures. The system may be implemented by a combination of hardware and software (as described), entirely by hardware elements, or entirely by software elements. The specific division of functions between the various system components described herein is exemplary only and not mandatory; on the contrary, the functions performed by a single system component may be performed by multiple components, or the functions performed by multiple components may be performed by a single component.

[0130] Those skilled in the art should understand that the various steps of the above disclosed method can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, optionally, they can be implemented with program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the embodiments of the present invention are not limited to any specific combination of hardware and software.

[0131] These computing device executable programs (also referred to as programs, software, software applications, or code) include machine instructions for programmable processors, and these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0132] Certain aspects of the present invention include process steps and instructions described herein in the form of algorithms. It should be noted that the process steps and instructions of the present invention can be implemented in software, firmware and / or hardware, and when implemented by software, it can be downloaded, stored on different platforms used by various operating systems and operated from the platforms.

[0133] Those skilled in the art will understand that the structures shown in the accompanying drawings are merely block diagrams of partial structures related to the scheme of the present application, and do not constitute a limitation on the terminal device to which the scheme of the present application is applied. The specific terminal device may include more or fewer components than shown in the figures, or combine certain components, or have a different arrangement of components.

[0134] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "possible design" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0135] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0136] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A model training method, characterized in that: The model trained by the method is used for feature recognition of multimodal data, including: Obtain multimodal data as a training sample set; Constructing a feature recognition network including a feature pre-recognition module, a temporal interactive attention mechanism and a feature fine recognition module, wherein the training sample set is passed through the feature pre-recognition module to obtain a pre-recognition feature set, the pre-recognition feature set is calculated through the temporal interactive attention mechanism to obtain a multimodal coupling attention weight, and the feature fine recognition module processes the pre-recognition feature set based on the multimodal coupling attention weight to obtain a fine recognition feature set; The pre-recognition feature set and the precise recognition feature set are compared to obtain a comparison result, network parameters of the feature recognition network are updated according to the comparison result, and the trained feature recognition network is output as a multimodal feature recognition model.

2. The method according to claim 1, characterized in that The feature pre-recognition module includes at least two modality recognition channels corresponding to different modalities, and the training sample set is passed through the feature pre-recognition module to obtain a pre-recognition feature set including: Each modality data in the training sample set is processed through the modality and its corresponding modality recognition channel to obtain a modality pre-recognition feature, and the modality pre-recognition features recognized by each modality recognition channel are combined into a pre-recognition feature set.

3. The method according to claim 2, characterized in that The feature pre-recognition module uses a convolutional neural network as the network architecture, and the modality pre-recognition feature is expressed as: , Where n represents the modality recognition channel identifier, is the modal pre-recognition feature identified by the modal recognition channel n at the t-th time slot, is the number of modal pre-recognition features identified by modal recognition channel n at the t-th time slot, is the first time slot modal identification channel n Features, is the convolutional neural network output function, is the activation function, is the attention coefficient of the feature pre-recognition module, ,in is the initial attention coefficient, represents the multimodal coupled attention matrix of the tth time slot, which is built-in as 1 in the feature pre-recognition module. Representation Matrix The diagonal matrix of the inner elements, Represents the bias term of the convolutional neural network.

4. The method according to claim 2, characterized in that: The multimodal coupling attention weights obtained by calculating the pre-recognized feature set through the temporal interactive attention mechanism include: Calculating the similarity between any two modal pre-recognition features obtained through the same modal recognition channel, recorded as a first similarity, and dividing the modal pre-recognition features of the same modal recognition channel into a plurality of modal feature groups according to the first similarity; Calculate the similarity between two modal feature groups belonging to different modal identification channels, recorded as a second similarity, and divide the modal feature groups of each modal identification channel into a plurality of interactive feature groups according to the second similarity, wherein any two modal feature groups in each interactive feature group do not belong to the same modal identification channel; The temporal interaction attention weight is calculated according to the number of features in the interactive feature group, the channel self-attention weight is calculated according to the feature mapping of the pre-identified feature set, and the multimodal coupling attention weight is obtained by weighted calculation of the temporal interaction attention weight and the channel self-attention weight.

5. The method according to claim 4, characterized in that The first similarity calculation is expressed as follows: in, For the The time slot modal identification channel n and The similarity of features, , and Respectively The time slot modal identification channel n and The features pre-identified by each channel, and is the weight coefficient, is the operator of waveform similarity between features, is the operator of temporal similarity between features.

6. The method according to claim 4, characterized in that The dividing the modal pre-recognition features of the same modal recognition channel into a plurality of modal feature groups according to the first similarity comprises: The modal pre-recognition features whose first similarity is greater than a set threshold are divided into the same modal feature group.

7. The method according to claim 1, characterized in that The comparing the pre-recognition feature set and the precise recognition feature set to obtain a comparison result, and updating the network parameters of the feature recognition network according to the comparison result includes: Comparing the pre-identification feature set with the fine identification feature set, recording the feature that exists in the fine identification feature set but does not exist in the pre-identification feature set as a first feature, recording the feature that exists in the pre-identification feature set but does not exist in the fine identification feature set as a second feature, and recording the recognition result of the first feature and / or the second feature as a comparison result; When the first feature exists, the first feature is added to the pre-recognition feature set for updating, and the updated pre-recognition feature set is subjected to the temporal interactive attention mechanism to calculate the attention weight of the first feature. When the attention weight of the first feature is greater than the attention weight of the pre-recognition feature set, the network parameters of the feature pre-recognition module are updated; otherwise, the correlation between the first feature and the set feature triggering event is calculated, and the network parameters of the temporal interactive attention mechanism are updated according to the correlation; When the second feature exists, the recognition error between the pre-recognition feature set and the fine recognition feature set is calculated, and the network parameters of the feature fine recognition module are updated according to the recognition error.

8. A multimodal data feature recognition method, characterized in that: include: Acquire multimodal data to be identified; The multimodal data is input into the multimodal feature recognition model trained by the model training method described in any one of claims 1 to 7 for recognition to obtain multimodal features.

9. A model training device, characterized in that: The device is used to train a multimodal feature recognition model, comprising: A sample acquisition module is used to acquire multimodal data as a training sample set; A network construction training module is used to construct a feature recognition network including a feature pre-recognition module, a temporal interactive attention mechanism and a feature fine recognition module. The multimodal data in the training data set is processed by the feature pre-recognition module to obtain a pre-recognition feature set. The pre-recognition feature set is calculated by the temporal interactive attention mechanism to obtain a multimodal coupling attention weight. The feature fine recognition module processes the pre-recognition feature set based on the temporal interactive attention weight to obtain a fine recognition feature set. The network update module is used to compare the pre-recognition feature set and the precise recognition feature set to obtain a comparison result, update the network parameters of the feature recognition network according to the comparison result, and output the trained feature recognition network as a multimodal feature recognition model.

10. A multimodal data feature recognition device, characterized in that: include: A data acquisition module, used to acquire multimodal data to be identified; The feature recognition module is used to input multimodal data into a multimodal feature recognition model trained by the device as described in claim 9 for recognition to obtain multimodal features.