Multimodal audiovisual video positioning method, device and storage medium
Through a multimodal audio-visual video positioning method, combining visual and auditory semantic representations, and using loss modulation coefficients to optimize model parameters, the problem of inaccurate positioning caused by insufficient visual data is solved, and the accuracy of video positioning is improved.
Patent Information
- Application Number
- CN202310880951.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-17
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-07-17
AI Technical Summary
Existing video positioning methods rely on visual data, resulting in low positioning accuracy in conditions of mixed lighting or background scenes.
Through a multimodal audio-visual video positioning method, combining visual and auditory semantic representations, using loss modulation coefficients to optimize model parameters, the semantic representation modeling of different modalities is balanced to improve positioning accuracy.
The positioning accuracy of weak semantic modal data in the video is improved, solving the problem of inaccurate positioning caused by insufficient visual data.
Smart Images

Figure CN117011764B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a multimodal audio-visual video positioning method, device, and storage medium. Background Art
[0002] In recent years, with the rapid development of internet technology, the widespread adoption of diverse multimedia data acquisition devices, and the continued growth in the number of internet users, more and more people are participating in the filming and production of online videos, leading to a surge in the amount of multimedia video data available online. These videos often contain a wealth of information, and mining and understanding this information can help promote the rapid progress and development of human life and production. However, due to the inherent multimodal structure and high-dimensional spatiotemporal properties of videos, as well as the massive amount of video data available on online platforms, relying solely on human effort to mine and understand this information presents significant challenges. Therefore, there is an urgent need to design efficient and reliable intelligent video content understanding models.
[0003] Currently, most video content understanding models are designed based on visual data. These methods primarily mine the spatiotemporal dimensions of visual modal data within videos, as well as the correlations between visual modal data across different videos, to help the models improve their understanding of the visual modal information within the videos. However, in situations where lighting or background noise are mixed, visual data is often insufficient to capture clear and reliable content. Consequently, model predictions based on this visual information are often unreliable or even erroneous, resulting in very low video localization accuracy. Summary of the Invention
[0004] The embodiments of the present application provide a multimodal audio-visual video positioning method, device and storage medium to solve the technical problem of low video positioning accuracy caused by poor visual data in the related art.
[0005] In a first aspect, an embodiment of the present application provides a multimodal audiovisual video positioning method, comprising:
[0006] Access to audiovisual videos;
[0007] The audio-visual video is input into a multimodal audio-visual video positioning model to obtain a video positioning result output by the multimodal audio-visual video positioning model; the multimodal audio-visual video positioning model is obtained through training based on audio-visual video samples in a training set, semantic category labels, and a loss modulation coefficient determined by the degree of imbalance of semantic representations between different modalities.
[0008] In some embodiments, the step of training the multimodal audio-visual video localization model includes:
[0009] Based on the audiovisual video samples and semantic category labels, obtain visual semantic representation and auditory semantic representation, and determine the visual classification prediction loss value and the auditory classification prediction loss value;
[0010] determining a degree of imbalance of the semantic representation based on a semantic difference between the visual semantic representation and the auditory semantic representation;
[0011] determining a loss modulation coefficient based on a degree of imbalance of the semantic representation;
[0012] Based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value, the model parameters are optimized to complete the training.
[0013] In some embodiments, determining the degree of imbalance of the semantic representation based on the semantic difference between the visual semantic representation and the auditory semantic representation includes:
[0014] Determining a visual classification prediction result and an auditory classification prediction result based on the visual semantic representation and the auditory semantic representation;
[0015] The imbalance degree of the semantic representation is determined based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label.
[0016] In some embodiments, determining the degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label includes:
[0017] Determining an average loss value of visual classification prediction and an average loss value of auditory classification prediction based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label;
[0018] The imbalance degree of the semantic representation is calculated according to the average loss value of the visual classification prediction and the average loss value of the auditory classification prediction.
[0019] In some embodiments, determining the loss modulation coefficient based on the imbalance degree of the semantic representation includes:
[0020] Based on the imbalance degree of the semantic representation, an activation function is used to determine a loss modulation coefficient.
[0021] In some embodiments, the optimizing model parameters based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value to complete the training includes:
[0022] Based on the loss modulation coefficient, the visual classification prediction loss value and the auditory classification prediction loss value are balanced modulated to determine a loss function;
[0023] Based on the loss function, the model parameters are optimized to complete the training.
[0024] In some embodiments, obtaining a visual semantic representation and an auditory semantic representation based on the audiovisual video sample and the semantic category label, and determining a visual classification prediction loss value and an auditory classification prediction loss value includes:
[0025] Obtaining a visual semantic representation and an auditory semantic representation of the audiovisual video sample;
[0026] A visual classification prediction loss value and an auditory classification prediction loss value are determined based on the visual semantic representation, the auditory semantic representation, and the semantic category label.
[0027] In some embodiments, obtaining the visual semantic representation and the auditory semantic representation of the audiovisual video sample includes:
[0028] Obtaining the visual semantic representation based on a visual semantic representation encoder;
[0029] Based on the auditory semantic representation encoder, the auditory semantic representation is obtained.
[0030] In a second aspect, an embodiment of the present application further provides a multimodal audio-visual video positioning device, comprising:
[0031] A first acquisition module is used to acquire audio-visual videos;
[0032] The second acquisition module is used to input the audio-visual video into a multimodal audio-visual video positioning model to obtain the video positioning result output by the multimodal audio-visual video positioning model; the multimodal audio-visual video positioning model is obtained through training based on the audio-visual video samples in the training set, semantic category labels and loss modulation coefficients determined by the degree of imbalance of semantic representations between different modalities.
[0033] In some embodiments, the second acquisition module includes a first determination submodule, a second determination submodule, a third determination submodule, and a first optimization submodule, wherein:
[0034] The first determination submodule is used to obtain a visual semantic representation and an auditory semantic representation based on the audiovisual video sample and the semantic category label, and determine a visual classification prediction loss value and an auditory classification prediction loss value;
[0035] The second determining submodule is used to determine the degree of imbalance of the semantic representation based on the semantic difference between the visual semantic representation and the auditory semantic representation;
[0036] The third determination submodule is used to determine a loss modulation coefficient based on the imbalance degree of the semantic representation;
[0037] The first optimization submodule is used to optimize model parameters based on the loss modulation coefficient, the visual classification prediction loss value and the auditory classification prediction loss value to complete training.
[0038] In some embodiments, the second determining submodule includes a first determining unit and a second determining unit, wherein:
[0039] The first determining unit is used to determine a visual classification prediction result and an auditory classification prediction result based on the visual semantic representation and the auditory semantic representation;
[0040] The second determining unit is configured to determine a degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label.
[0041] In some embodiments, the first determining unit includes a first determining subunit and a first calculating subunit:
[0042] The first determining subunit is used to determine an average loss value of visual classification prediction and an average loss value of auditory classification prediction based on the visual classification prediction result, the auditory classification prediction result and the semantic category label;
[0043] The first calculation subunit calculates the imbalance degree of the semantic representation according to the visual classification prediction average loss value and the auditory classification prediction average loss value.
[0044] In some embodiments, the third determining submodule includes a third determining unit:
[0045] The third determining unit is configured to determine a loss modulation coefficient using an activation function based on the degree of imbalance of the semantic representation.
[0046] In some embodiments, the first optimization submodule includes a fourth determination unit and a first optimization unit, wherein:
[0047] The fourth determining unit is configured to perform balanced modulation on the visual classification prediction loss value and the auditory classification prediction loss value based on the loss modulation coefficient to determine a loss function;
[0048] The first optimization unit is used to optimize model parameters based on the loss function to complete training.
[0049] In some embodiments, the first determining submodule includes a first acquiring unit and a fifth determining unit, wherein:
[0050] The first acquisition unit is used to acquire the visual semantic representation and the auditory semantic representation of the audio-visual video sample;
[0051] The fifth determining unit is used to determine a visual classification prediction loss value and an auditory classification prediction loss value based on the visual semantic representation, the auditory semantic representation and the semantic category label.
[0052] In some embodiments, the first acquisition unit includes a first acquisition subunit and a second acquisition subunit, wherein:
[0053] The first acquisition subunit is used to acquire the visual semantic representation based on the visual semantic representation encoder;
[0054] The second acquisition subunit is used to acquire the auditory semantic representation based on the auditory semantic representation encoder.
[0055] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multimodal audio-visual video positioning method as described above is implemented.
[0056] In a fourth aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the multimodal audio-visual video positioning methods described above.
[0057] In a fifth aspect, an embodiment of the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the multimodal audiovisual video positioning methods described above.
[0058] The multimodal audiovisual video positioning method, device and storage medium provided in the embodiments of the present application measure the imbalance between the visual semantic representation and the auditory semantic representation in the audiovisual video, and perform loss modulation on the multimodal audiovisual video positioning model based on the measurement results to optimize the model parameters, thereby improving the positioning accuracy of weak semantic modal data in the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0060] Figure 1 Schematic diagram of the process of the multimodal audio-visual video positioning method provided in the embodiment of the present application;
[0061] Figure 2 This is a framework diagram of the multimodal audio-visual video positioning method provided in an embodiment of the present application;
[0062] Figure 3 Schematic diagram of the structure of a multimodal audio-visual video positioning device provided in an embodiment of the present application;
[0063] Figure 4 It is a schematic diagram of the physical structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0064] Recently, with the rapid development of deep learning technology and the growing demand for understanding audiovisual multimodal data in real-world scenarios, a large number of audiovisual video understanding tasks have been proposed and have achieved remarkable research progress. In these research tasks, to optimize the parameters of multimodal models, a unified learning objective and training strategy are often used to learn semantic representation modeling networks across different modalities during training. However, no audiovisual video parsing method has considered the imbalance in semantic representation modeling between modalities caused by differences in semantic salience between different modalities during model learning.
[0065] To address the above-mentioned issues, an embodiment of the present application provides a multimodal audio-visual video positioning method, which measures the imbalance between the visual semantic representation and the auditory semantic representation in the audio-visual video, and performs loss modulation on the multimodal audio-visual video positioning model based on the measurement results to optimize the model parameters, thereby improving the positioning accuracy of weak semantic modal data in the video.
[0066] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0067] Figure 1 FIG. 1 is a flow chart of a multimodal audiovisual video positioning method provided in an embodiment of the present application, such as Figure 1 As shown, the embodiment of the present application provides a multimodal audio-visual video positioning method, including:
[0068] Step 101: Acquire audio-visual video.
[0069] Step 102: Input the audio-visual video into a multimodal audio-visual video positioning model to obtain a video positioning result output by the multimodal audio-visual video positioning model; the multimodal audio-visual video positioning model is obtained through training based on audio-visual video samples in a training set, semantic category labels, and a loss modulation coefficient determined by the degree of imbalance of semantic representations between different modalities.
[0070] Specifically, after completing the training of the multimodal audiovisual video localization model, the audiovisual video can be input into the trained multimodal audiovisual video localization model. The multimodal audiovisual video localization model can obtain the degree of imbalance between the visual semantic representation and the auditory semantic representation based on the visual data and auditory data in the audiovisual video, and perform loss modulation based on this imbalance to determine the classification of the video. Then, based on the video classification, a video localization result is output. The video localization result can be a temporal localization result of the video, such as a certain time point or a certain time period in the video.
[0071] The multimodal audio-visual video localization method provided in the embodiment of the present application measures the imbalance between the visual semantic representation and the auditory semantic representation in the audio-visual video, and performs loss modulation on the multimodal audio-visual video localization model based on the measurement results to optimize the model parameters, thereby improving the localization accuracy of weak semantic modal data in the video.
[0072] In some embodiments, the step of training the multimodal audio-visual video localization model includes:
[0073] Based on the audiovisual video samples and semantic category labels, obtain visual semantic representation and auditory semantic representation, and determine the visual classification prediction loss value and the auditory classification prediction loss value;
[0074] determining a degree of imbalance of the semantic representation based on a semantic difference between the visual semantic representation and the auditory semantic representation;
[0075] determining a loss modulation coefficient based on a degree of imbalance of the semantic representation;
[0076] Based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value, the model parameters are optimized to complete the training.
[0077] Specifically, the training steps of the multimodal audio-visual video localization model may include: obtaining the visual modality and auditory modality semantic information of the video based on the audio-visual video samples, modeling the semantic representation of the visual and auditory data of the videos in the training set, and determining the visual classification prediction loss value and the auditory classification prediction loss value; determining the imbalance of semantic representation learning between the visual and auditory modalities based on the visual classification prediction loss value and the auditory classification prediction loss value, and measuring the imbalance of multimodal semantic representation modeling; determining the optimized loss modulation coefficient of the semantic representation modeling network in different modalities based on the imbalance degree of multimodal semantic representation modeling obtained by measurement; using the loss modulation coefficient to modulate the classification loss function terms of different modalities, and optimizing the parameters of the semantic representation modeling network in different modalities in a balanced manner according to the modulated loss function to complete the training. The audio-visual video samples can be any video with visual and auditory information.
[0078] The multimodal audio-visual video positioning method provided by the embodiment of the application can improve the positioning accuracy of weak semantic modal data in a video by measuring the imbalance between visual semantic representation and auditory semantic representation in the audio-visual video, and modulating the loss of the multimodal audio-visual video positioning model according to the measurement result, and optimizing the model parameters.
[0079] In some embodiments, based on the audio-visual video sample and the semantic category label, the visual semantic representation and the auditory semantic representation are obtained, and the visual classification prediction loss value and the auditory classification prediction loss value are determined, including:
[0080] The visual semantic representation and the auditory semantic representation of the audio-visual video sample are obtained.
[0081] The visual classification prediction loss value and the auditory classification prediction loss value are determined according to the visual semantic representation, the auditory semantic representation and the semantic category label.
[0082] Specifically, for each video sample containing audio-visual modal data in the training data set, the video segment level visual semantic representation and the video segment level auditory semantic representation are extracted using the visual semantic representation encoder E v (·;θ v ) and the auditory semantic representation encoder E a (·;θ a ), and the video segment level visual semantic representation and the video segment level auditory semantic representation are integrated using the time sequence pooling strategy to obtain the video level visual semantic representation v and the video level auditory semantic representation a.
[0083] Then, the video level classification prediction loss value of different modalities is calculated using the video level visual, auditory and audio-visual modal labels as supervision information:
[0084]
[0085]
[0086]
[0087] Here, and represent the category labels of the auditory, visual and video audio-visual modalities, respectively, and correspondingly, and represent the classification prediction probabilities of the corresponding auditory, visual and video audio-visual modalities, and N and C represent the number of samples in each training batch and the total number of classification categories, respectively. On this basis, the model comprehensively considers the classification loss in different modalities, and updates the model parameters of the semantic representation modeling network in different modalities.
[0088] The multimodal audio-visual video positioning method provided in the embodiments of the present application extracts semantic representations of different modalities in a video sample through a semantic representation encoder, and calculates classification prediction loss values of different modalities according to semantic category labels, so that the imbalance between visual semantic representations and auditory semantic representations in a video can be measured.
[0089] In some embodiments, the obtaining of the visual semantic representation and the auditory semantic representation of the audio-visual video sample comprises:
[0090] The visual semantic representation is obtained based on a visual semantic representation encoder.
[0091] The auditory semantic representation is obtained based on an auditory semantic representation encoder.
[0092] Specifically, for each video sample containing audio-visual modality data in a training data set, the video segment level visual semantic representation and the video segment level auditory semantic representation of the video sample are extracted using a visual semantic representation encoder E v (·;θ v ) and an auditory semantic representation encoder E a (·;θ a ).
[0093] The multimodal audio-visual video positioning method provided in the embodiments of the present application extracts semantic representations of different modalities in a video sample through a semantic representation encoder, and calculates classification prediction loss values of different modalities according to semantic category labels, so that the imbalance between visual semantic representations and auditory semantic representations in a video can be measured.
[0094] In some embodiments, the determining of the imbalance degree of the semantic representations based on the semantic difference between the visual semantic representation and the auditory semantic representation comprises:
[0095] The visual classification prediction result and the auditory classification prediction result are determined based on the visual semantic representation and the auditory semantic representation.
[0096] The imbalance degree of the semantic representations is determined based on the visual classification prediction result, the auditory classification prediction result and the semantic category label.
[0097] Specifically, the multimodal audio-visual video positioning model can map the semantic representations of different modalities to corresponding video level classification prediction results P a and P v according to an auditory semantic classification mapping function f a (·) and a visual semantic category mapping function f v . And the imbalance degree of the semantic representations is determined according to the classification prediction results Pa and P v And the category label vector Y of the corresponding mode a and Y v Calculate the average loss of classification prediction of different modalities. Finally, according to the calculated average loss PL of category prediction in different modalities a and PL v To calculate the imbalance degree of semantic representation modeling between different modalities and
[0098]
[0099]
[0100] if, The larger the value, the higher the PL v The larger the PL a The smaller it is, the lower the accuracy of visual modality classification prediction is, and the higher the accuracy of auditory modality classification prediction is. This indicates that the auditory modality semantic representation learning is better and the visual modality semantic representation modeling performance is poor.
[0101] The multimodal audiovisual video localization method provided in the embodiment of the present application can reflect the differences in semantic representation modeling performance between different modalities by measuring the imbalance between visual semantic representation and auditory semantic representation in audiovisual videos, thereby effectively balancing the training processes of different modalities.
[0102] In some embodiments, determining the degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label includes:
[0103] Determining an average loss value of visual classification prediction and an average loss value of auditory classification prediction based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label;
[0104] The imbalance degree of the semantic representation is calculated according to the average loss value of the visual classification prediction and the average loss value of the auditory classification prediction.
[0105] Specifically, the multimodal audio-visual video localization model can be mapped according to the auditory semantic classification function f a (·) and the visual semantic category mapping function f v (·), mapping the semantic representations of different modalities into corresponding video-level classification prediction results P a and P v :
[0106] P a =f a (a) = W aE a (a; θ a )+b a
[0107] P v =f v (v) = W v E v (v;θ v )+b v
[0108] Here W a 、b a 、W v and b v Represent the weight and offset coefficient matrices in different modal classification mapping functions, θ a and θ v They represent the parameters of different modal semantic representation encoders.
[0109] Then, according to the classification prediction results P of different modes a and P v And the category label vector Y of the corresponding mode a and Y v Calculate the average loss of classification predictions for different modalities:
[0110]
[0111]
[0112] Here, P[c] and Y[c] represent the classification prediction probability and the corresponding class label value corresponding to the cth class, respectively. Different classes can be represented by letters or arbitrary symbols.
[0113] Finally, the average loss PL of category prediction in different modalities is calculated a and PL v To calculate the imbalance degree of semantic representation modeling between different modalities and
[0114] The multimodal audiovisual video localization method provided in the embodiment of the present application can reflect the differences in semantic representation modeling performance between different modalities by measuring the imbalance between visual semantic representation and auditory semantic representation in audiovisual videos, thereby effectively balancing the training processes of different modalities.
[0115] In some embodiments, determining the loss modulation coefficient based on the imbalance degree of the semantic representation includes:
[0116] Based on the imbalance degree of the semantic representation, an activation function is used to determine a loss modulation coefficient.
[0117] Specifically, based on the degree of imbalance in the multimodal semantic representation modeling, the optimized loss modulation coefficients of the semantic representation modeling networks in different modalities can be determined. In some embodiments, the optimized loss modulation coefficients can be selected from activation functions such as tanh(·) and sigmoid(·).
[0118] For example, when using the sigmoid(·) activation function, the imbalance measurement method of modeling the different modal semantic representations and the calculated degree of imbalance between modalities are used. and The optimized loss modulation coefficient ρ of the above different modes v and ρ a as follows:
[0119]
[0120]
[0121] Here, σ(·) represents the sigmoid activation function, which is used to normalize the imbalance between different modalities to between 0 and 1, so as to further modulate the importance weights of different modal loss functions in the model optimization process.
[0122] The multimodal audio-visual video localization method provided in the embodiment of the present application measures the imbalance between the visual semantic representation and the auditory semantic representation in the audio-visual video, and modulates the loss of the multimodal audio-visual video localization model according to the measurement results, thereby effectively balancing the training processes of different modalities.
[0123] In some embodiments, the optimizing model parameters based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value to complete the training includes:
[0124] Based on the loss modulation coefficient, the visual classification prediction loss value and the auditory classification prediction loss value are balanced modulated to determine a loss function;
[0125] Based on the loss function, the model parameters are optimized to complete the training.
[0126] Specifically, first, the model further balances the prediction losses of the visual and auditory semantic classification networks based on the optimized loss modulation coefficients of the semantic representation modeling networks of different modalities. The overall loss function of the modulated model is as follows:
[0127] L=ρ a L a +ρ v L v +L av
[0128] Then the model optimizes the parameters of the semantic representation modeling network in different modalities based on the modulated loss function. The model uses the optimization loss modulation coefficient p v and p a Balanced optimization of semantic representation networks in different modalities is achieved, thereby facilitating full learning of the parameters of the networks in different modalities.
[0129] The multi-modal audio-visual video positioning method provided by the embodiments of the present application measures the imbalance between the visual semantic representation and the auditory semantic representation in the audio-visual video, and modulates the loss of the multi-modal audio-visual video positioning model according to the measurement result, so as to effectively balance the training process of different modalities and make the loss functions of different modalities very close after balancing adjustment.
[0130] The method in the above embodiments will be further described below with specific examples.
[0131] Figure 2 is a framework diagram of the multi-modal audio-visual video positioning method provided by the embodiments of the present application. The method is modeled by using a deep convolutional neural network, and the imbalance degree of semantic representation modeling between different modalities is measured in a unified audio-visual video analysis framework, and a scheme of balanced semantic representation modeling between modalities is designed according to the imbalance of semantic representation learning between different modalities. As shown in Figure 2 , the method includes four main parts: (1) video semantic representation modeling; (2) inter-modal imbalance semantic representation modeling measurement; (3) loss function modulation; (4) video positioning decision.
[0132] Specifically, the method includes the following steps:
[0133] Step S1, obtaining the visual modality and auditory modality semantic information of a video, modeling the semantic representation of the visual and auditory data of the video in the training set; the video can be any video with visual and auditory information, and the public video crawled on an online video website can be used.
[0134] The step S1 further includes the following steps:
[0135] Step S11: for each video sample containing audio-visual modality data in the training data set, using the visual semantic representation encoder E v (·; θ v ) and the auditory semantic representation encoder E a (·; θ a ) to extract the video segment level visual semantic representation and auditory semantic representation thereof, and using a time sequence pooling strategy to integrate the video segment level visual semantic representation and auditory semantic representation and obtain the video level visual semantic representation v and auditory semantic representation a;
[0136] Step S12: Using the video-level visual, auditory, and audiovisual modality labels as supervision information, calculate the loss values of the video-level classification prediction corresponding to different modalities:
[0137]
[0138]
[0139]
[0140] here, and Represent the category labels of auditory, visual and video audiovisual modalities respectively, and accordingly, and represents the classification prediction probability for the corresponding auditory, visual, and video audiovisual modalities, while N and C represent the number of samples in each training batch and the total number of classification categories, respectively. Based on this, the model comprehensively considers the classification losses across different modalities and uses this to update the model parameters of the semantic representation modeling networks in each modality.
[0141] In step S2, the imbalance in semantic representation learning between visual and auditory modalities is considered, and a measurement method for the imbalance in multimodal semantic representation modeling is determined.
[0142] The step S2 further comprises the following steps:
[0143] Step S21: Determine the auditory semantic classification mapping function f for the video-level auditory semantic representation a and visual semantic representation v extracted in step S11 a (·) and the visual semantic category mapping function f v (·), mapping the semantic representations of different modalities into corresponding video-level classification prediction results P a and P v Here, the video-level auditory classification prediction and video-level visual classification prediction results are calculated as follows:
[0144] P a =f a (a) = W a E a (a; θ a )+b a
[0145] P v =f v (v) = W v E v (v;θ v )+b v
[0146] Here Wa 、b a 、W v and b v Represent the weight and offset coefficient matrices in different modal classification mapping functions, θ a and θ v They represent the parameters of the encoders for semantic representation of different modalities.
[0147] Step S22: Classification prediction results P based on different modalities a and P v And the category label vector Y of the corresponding mode a and Y v Calculate the average loss of classification predictions for different modalities:
[0148]
[0149]
[0150] Here, P[c] and Y[c] represent the classification prediction probability corresponding to the c-th category and the corresponding category label value, respectively.
[0151] Step S23: The average loss PL of the category prediction in different modes calculated in step S22 is a and PL v To calculate the imbalance degree of semantic representation modeling between different modalities and
[0152]
[0153]
[0154] here, The larger the value, the higher the PL v The larger the PL a The smaller it is, the lower the accuracy of visual modality classification prediction is, and the higher the accuracy of auditory modality classification prediction is. This indicates that the auditory modality semantic representation learning is better and the visual modality semantic representation modeling performance is poor.
[0155] Step S3: Determine the optimal loss modulation coefficients for the semantic representation modeling networks in different modalities based on the degree of imbalance in the multimodal semantic representation modeling measured in step S2. In some embodiments, the optimal loss modulation coefficients may be selected from activation functions such as tanh(·) and sigmoid(·).
[0156] The measurement method for imbalance of modeling based on the semantic representation of different modalities determined in step S23 and the degree of imbalance between modalities calculated and The optimized loss modulation coefficient ρ of the above different modes v and ρ a as follows:
[0157]
[0158]
[0159] Here, σ(·) represents the sigmoid activation function, which is used to normalize the imbalance between different modalities to between 0 and 1, so as to further modulate the importance weights of different modal loss functions in the model optimization process.
[0160] In step S4, the classification loss function terms of different modalities are modulated using the loss modulation coefficient, and the semantic representation modeling network parameters in different modalities are optimized in a balanced manner according to the modulated loss function.
[0161] The step S4 further comprises the following steps:
[0162] Step S41: First, the model further balances the prediction losses of the visual and auditory semantic classification networks based on the optimized loss modulation coefficients of the semantic representation modeling networks of different modalities. The overall loss function of the modulated model is as follows:
[0163] L=ρ a L a +ρ v L v +L av
[0164] Step S42: The model optimizes the parameters of the semantic representation modeling network in different modalities based on the above-mentioned modulated loss function. The specific update process of the network parameters of different modalities is as follows. For simplicity, the modal subscripts of the model parameters in the following formula are omitted:
[0165]
[0166]
[0167] At this point, the model uses the optimized loss modulation coefficient ρ of different modes v and ρ a This method achieves balanced optimization of semantic representation networks across different modalities, thereby promoting full learning of network parameters across different modalities. Results show that after balancing, the optimization losses across different modalities are nearly identical after model convergence, demonstrating that this method can effectively balance the training process across different modalities.
[0168] The multimodal audio-visual video localization method provided in the embodiment of the present application measures the imbalance between the visual semantic representation and the auditory semantic representation in the audio-visual video, and performs loss modulation on the multimodal audio-visual video localization model based on the measurement results to optimize the model parameters, thereby improving the localization accuracy of weak semantic modal data in the video.
[0169] Figure 3 is a structural diagram of a multimodal audio-visual video positioning device provided in an embodiment of the present application, such as Figure 3 As shown, the multimodal audio-visual video positioning device provided in the embodiment of the present application includes a first acquisition module 301 and a second acquisition module 302, wherein:
[0170] A first acquisition module 301 is used to acquire audio-visual videos;
[0171] The second acquisition module 302 is used to input the audio-visual video into a multimodal audio-visual video positioning model to obtain the video positioning result output by the multimodal audio-visual video positioning model; the multimodal audio-visual video positioning model is obtained through training based on the audio-visual video samples in the training set, semantic category labels and loss modulation coefficients determined by the degree of imbalance of semantic representations between different modalities.
[0172] In some embodiments, the second acquisition module includes a first determination submodule, a second determination submodule, a third determination submodule, and a first optimization submodule, wherein:
[0173] The first determination submodule is used to obtain a visual semantic representation and an auditory semantic representation based on the audiovisual video sample and the semantic category label, and determine a visual classification prediction loss value and an auditory classification prediction loss value;
[0174] The second determining submodule is used to determine the degree of imbalance of the semantic representation based on the semantic difference between the visual semantic representation and the auditory semantic representation;
[0175] The third determination submodule is used to determine a loss modulation coefficient based on the imbalance degree of the semantic representation;
[0176] The first optimization submodule is used to optimize model parameters based on the loss modulation coefficient, the visual classification prediction loss value and the auditory classification prediction loss value to complete training.
[0177] In some embodiments, the second determining submodule includes a first determining unit and a second determining unit, wherein:
[0178] The first determining unit is used to determine a visual classification prediction result and an auditory classification prediction result based on the visual semantic representation and the auditory semantic representation;
[0179] The second determining unit is configured to determine a degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label.
[0180] In some embodiments, the first determining unit includes a first determining subunit and a first calculating subunit:
[0181] The first determining subunit is used to determine an average loss value of visual classification prediction and an average loss value of auditory classification prediction based on the visual classification prediction result, the auditory classification prediction result and the semantic category label;
[0182] The first calculation subunit calculates the imbalance degree of the semantic representation according to the visual classification prediction average loss value and the auditory classification prediction average loss value.
[0183] In some embodiments, the third determining submodule includes a third determining unit:
[0184] The third determining unit is configured to determine a loss modulation coefficient using an activation function based on the degree of imbalance of the semantic representation.
[0185] In some embodiments, the first optimization submodule includes a fourth determination unit and a first optimization unit, wherein:
[0186] The fourth determining unit is configured to perform balanced modulation on the visual classification prediction loss value and the auditory classification prediction loss value based on the loss modulation coefficient to determine a loss function;
[0187] The first optimization unit is used to optimize model parameters based on the loss function to complete training.
[0188] In some embodiments, the first determining submodule includes a first acquiring unit and a fifth determining unit, wherein:
[0189] The first acquisition unit is used to acquire the visual semantic representation and the auditory semantic representation of the audio-visual video sample;
[0190] The fifth determining unit is used to determine a visual classification prediction loss value and an auditory classification prediction loss value based on the visual semantic representation, the auditory semantic representation and the semantic category label.
[0191] In some embodiments, the first acquisition unit includes a first acquisition subunit and a second acquisition subunit, wherein:
[0192] The first acquisition subunit is used to acquire the visual semantic representation based on the visual semantic representation encoder;
[0193] The second acquisition subunit is used to acquire the auditory semantic representation based on the auditory semantic representation encoder.
[0194] Specifically, the above-mentioned multimodal audio-visual video positioning device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned multimodal audio-visual video positioning method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0195] Figure 4 is a schematic diagram of the physical structure of the electronic device provided in the embodiment of the present application, such as Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the multimodal audio-visual video positioning method, which includes:
[0196] Access to audiovisual videos;
[0197] The audio-visual video is input into a multimodal audio-visual video positioning model to obtain a video positioning result output by the multimodal audio-visual video positioning model; the multimodal audio-visual video positioning model is obtained through training based on audio-visual video samples in a training set, semantic category labels, and a loss modulation coefficient determined by the degree of imbalance of semantic representations between different modalities.
[0198] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0199] In some embodiments, the step of training the multimodal audio-visual video localization model includes:
[0200] Based on the audiovisual video samples and semantic category labels, obtain visual semantic representation and auditory semantic representation, and determine the visual classification prediction loss value and the auditory classification prediction loss value;
[0201] determining a degree of imbalance of the semantic representation based on a semantic difference between the visual semantic representation and the auditory semantic representation;
[0202] determining a loss modulation coefficient based on a degree of imbalance of the semantic representation;
[0203] Based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value, the model parameters are optimized to complete the training.
[0204] In some embodiments, determining the degree of imbalance of the semantic representation based on the semantic difference between the visual semantic representation and the auditory semantic representation includes:
[0205] Determining a visual classification prediction result and an auditory classification prediction result based on the visual semantic representation and the auditory semantic representation;
[0206] The imbalance degree of the semantic representation is determined based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label.
[0207] In some embodiments, determining the degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label includes:
[0208] Determining an average loss value of visual classification prediction and an average loss value of auditory classification prediction based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label;
[0209] The imbalance degree of the semantic representation is calculated according to the average loss value of the visual classification prediction and the average loss value of the auditory classification prediction.
[0210] In some embodiments, determining the loss modulation coefficient based on the imbalance degree of the semantic representation includes:
[0211] Based on the imbalance degree of the semantic representation, an activation function is used to determine a loss modulation coefficient.
[0212] In some embodiments, the optimizing model parameters based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value to complete the training includes:
[0213] Based on the loss modulation coefficient, the visual classification prediction loss value and the auditory classification prediction loss value are balanced modulated to determine a loss function;
[0214] Based on the loss function, the model parameters are optimized to complete the training.
[0215] In some embodiments, obtaining a visual semantic representation and an auditory semantic representation based on the audiovisual video sample and the semantic category label, and determining a visual classification prediction loss value and an auditory classification prediction loss value includes:
[0216] Obtaining a visual semantic representation and an auditory semantic representation of the audiovisual video sample;
[0217] A visual classification prediction loss value and an auditory classification prediction loss value are determined based on the visual semantic representation, the auditory semantic representation, and the semantic category label.
[0218] In some embodiments, obtaining the visual semantic representation and the auditory semantic representation of the audiovisual video sample includes:
[0219] Obtaining the visual semantic representation based on a visual semantic representation encoder;
[0220] Based on the auditory semantic representation encoder, the auditory semantic representation is obtained.
[0221] Specifically, the above-mentioned electronic device provided in the embodiment of the present application can implement all the method steps implemented in the method embodiment in which the execution subject is the electronic device, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0222] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the multimodal audiovisual video positioning method provided by the above methods, the method comprising:
[0223] Access to audiovisual videos;
[0224] The audio-visual video is input into a multimodal audio-visual video positioning model to obtain a video positioning result output by the multimodal audio-visual video positioning model; the multimodal audio-visual video positioning model is obtained through training based on audio-visual video samples in a training set, semantic category labels, and a loss modulation coefficient determined by the degree of imbalance of semantic representations between different modalities.
[0225] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal audiovisual video localization method provided by the above methods, the method comprising:
[0226] Access to audiovisual videos;
[0227] The audio-visual video is input into a multimodal audio-visual video positioning model to obtain a video positioning result output by the multimodal audio-visual video positioning model; the multimodal audio-visual video positioning model is obtained through training based on audio-visual video samples in a training set, semantic category labels, and a loss modulation coefficient determined by the degree of imbalance of semantic representations between different modalities.
[0228] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0229] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0230] It should also be noted that the terms "first," "second," and the like in the embodiments of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein. Furthermore, the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects. For example, the first object can be one or more.
[0231] In the embodiments of this application, the term "and / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0232] In the embodiments of the present application, the term "plurality" refers to two or more than two, and other quantifiers are similar.
[0233] In this application, "determine B based on A" means that factor A must be considered when determining B. This is not limited to "determine B based solely on A" and should also include: "determine B based on A and C", "determine B based on A, C, and E", "determine C based on A, and further determine B based on C", etc. It can also include using A as a condition for determining B, for example, "when A meets the first condition, use the first method to determine B"; another example, "when A meets the second condition, determine B"; another example, "when A meets the third condition, determine B based on the first parameter", etc. Of course, it can also be a condition that uses A as a factor in determining B, for example, "when A meets the first condition, use the first method to determine C, and further determine B based on C", etc.
[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A multimodal audio-visual video positioning method, characterized in that: include: Access to audiovisual videos; Inputting the audio-visual video into a multimodal audio-visual video positioning model, and obtaining a video positioning result output by the multimodal audio-visual video positioning model; The multimodal audio-visual video localization model is obtained through training based on audio-visual video samples in a training set, semantic category labels, and a loss modulation coefficient determined by the degree of imbalance of semantic representations between different modalities; The training steps of the multimodal audio-visual video localization model include: Based on the audiovisual video samples and semantic category labels, obtain visual semantic representation and auditory semantic representation, and determine the visual classification prediction loss value and the auditory classification prediction loss value; determining a degree of imbalance of the semantic representation based on a semantic difference between the visual semantic representation and the auditory semantic representation; The determining the degree of imbalance of the semantic representation based on the semantic difference between the visual semantic representation and the auditory semantic representation includes: Determining a visual classification prediction result and an auditory classification prediction result based on the visual semantic representation and the auditory semantic representation; Determining a degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label; The determining the degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label includes: Determining an average loss value of visual classification prediction and an average loss value of auditory classification prediction based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label; The imbalance degree of the semantic representation is calculated according to the average loss value of the visual classification prediction and the average loss value of the auditory classification prediction.
2. The multimodal audio-visual video positioning method according to claim 1, characterized in that: The step of training the multimodal audio-visual video localization model further includes: determining a loss modulation coefficient based on a degree of imbalance of the semantic representation; Based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value, the model parameters are optimized to complete the training.
3. The multimodal audio-visual video positioning method according to claim 2, characterized in that: The determining of the loss modulation coefficient based on the imbalance degree of the semantic representation includes: Based on the imbalance degree of the semantic representation, an activation function is used to determine a loss modulation coefficient.
4. The multimodal audio-visual video positioning method according to claim 2, characterized in that: The optimizing model parameters based on the loss modulation coefficient, the visual classification prediction loss value, and the auditory classification prediction loss value to complete the training includes: Based on the loss modulation coefficient, the visual classification prediction loss value and the auditory classification prediction loss value are balanced modulated to determine a loss function; Based on the loss function, the model parameters are optimized to complete the training.
5. The multimodal audio-visual video positioning method according to claim 1, characterized in that: The obtaining of visual semantic representation and auditory semantic representation based on the audiovisual video sample and the semantic category label, and determining the visual classification prediction loss value and the auditory classification prediction loss value, includes: Obtaining a visual semantic representation and an auditory semantic representation of the audiovisual video sample; A visual classification prediction loss value and an auditory classification prediction loss value are determined based on the visual semantic representation, the auditory semantic representation, and the semantic category label.
6. The multimodal audio-visual video positioning method according to claim 5, characterized in that: The obtaining of the visual semantic representation and the auditory semantic representation of the audiovisual video sample includes: Obtaining the visual semantic representation based on a visual semantic representation encoder; Based on the auditory semantic representation encoder, the auditory semantic representation is obtained.
7. A multimodal audio-visual video positioning device, characterized in that: include: A first acquisition module is used to acquire audio-visual videos; a second acquisition module, configured to input the audio-visual video into a multimodal audio-visual video localization model and obtain a video localization result output by the multimodal audio-visual video localization model; the multimodal audio-visual video localization model is obtained through training based on audio-visual video samples in a training set, semantic category labels, and a loss modulation coefficient determined by the degree of imbalance in semantic representations between different modalities; The second acquisition module includes a first determination submodule and a second determination submodule, wherein: The first determination submodule is used to obtain a visual semantic representation and an auditory semantic representation based on the audiovisual video sample and the semantic category label, and determine a visual classification prediction loss value and an auditory classification prediction loss value; The second determining submodule is used to determine the degree of imbalance of the semantic representation based on the semantic difference between the visual semantic representation and the auditory semantic representation; The second determining submodule includes a first determining unit and a second determining unit, wherein: The first determining unit is used to determine a visual classification prediction result and an auditory classification prediction result based on the visual semantic representation and the auditory semantic representation; The second determining unit is configured to determine a degree of imbalance of the semantic representation based on the visual classification prediction result, the auditory classification prediction result, and the semantic category label; The second determining unit includes a first determining subunit and a first calculating subunit: The first determining subunit is used to determine an average loss value of visual classification prediction and an average loss value of auditory classification prediction based on the visual classification prediction result, the auditory classification prediction result and the semantic category label; The first calculation subunit calculates the imbalance degree of the semantic representation according to the visual classification prediction average loss value and the auditory classification prediction average loss value.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the multimodal audio-visual video positioning method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Audio-visual video analysis device and method based on multi-scale semantic network
CN114519809A
Audio-visual event positioning method and device, model training method and device, equipment and medium
CN116246214A