Multi-view video matching model training method, multi-view video matching method and related equipment
By integrating video, audio and text modal features, the problem of insufficient accuracy in cross-view video matching is solved, and higher recognition accuracy and consistency are achieved.
Patent Information
- Application Number
- CN202510485181.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, due to the differences in visual characteristics of different perspectives, the accuracy of cross-view video data matching is low, especially in the identification of local information and global information.
By extracting video and audio features and fusing to generate video modal audio-visual representations, the audio similarity in the first and third perspectives is used, and text modal audio-visual representations are combined as semantic anchors, the distance between matching video pairs in the feature space is shortened through comparative learning, semantic consistency is enhanced, and viewing angle differences are reduced.
The accuracy of cross-view video matching is improved, and the accuracy of recognition and matching in downstream tasks such as scene recognition and action comprehension is enhanced.
Smart Images

Figure CN120472362A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a multi-view video matching model training method, a multi-view video matching method and related equipment. Background Art
[0002] The ability to depict the activities of others from one's own perspective is an innate human skill known as cross-view correlation. In neural networks, this cross-view video correlation capability is represented by a cross-view video representation space, in which the resulting video feature vectors simultaneously incorporate information from different viewpoints. In practical applications, the use of a cross-view video representation space allows insufficient information from one viewpoint to be supplemented by correlating information from another. Therefore, matching video data from different viewpoints is essential.
[0003] In related technologies, matching video data from different perspectives within the same scene at the same time in a video database typically involves constructing a cross-perspective video representation space based on visual information from multiple video data sets, thereby matching and associating first-perspective video data with third-perspective video data. However, due to differences in visual features between different perspectives—first-perspective video data focuses more on local scene recognition, while third-perspective video data focuses more on global scene recognition—this visual information-based cross-perspective video data matching and association method results in low matching accuracy. Summary of the Invention
[0004] The embodiments of the present application provide a multi-view video matching model training method, a multi-view video matching method and related equipment, which can improve the matching accuracy of multi-view video data.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present application provides a multi-view video matching model training method, the method comprising:
[0006] Acquire multiple sets of video sample data and voice sample data and video description text corresponding to each of the video sample data, wherein the video sample data includes a first-perspective video and a third-perspective video, and the voice sample data includes a first-perspective voice and a third-perspective voice;
[0007] Inputting the first-perspective video and the first-perspective speech into a video and audio encoder in a multi-perspective video matching model for processing to obtain a first video modality audiovisual representation, and inputting the third-perspective video and the third-perspective speech into the video and audio encoder for processing to obtain a third video modality audiovisual representation;
[0008] Inputting the video description text corresponding to the video sample data into a text encoder in a multi-view video matching model for processing to obtain a text modality audiovisual representation;
[0009] A multi-view modality loss is calculated based on the first video modality audio-visual representation, the third video modality audio-visual representation and the text modality audio-visual representation, and the multi-view video matching model is updated based on the multi-view modality loss to obtain a trained multi-view video matching model.
[0010] In some embodiments, the video and audio encoder includes a video encoder and an audio encoder, and inputting the first-view video and the first-view speech into the video and audio encoder in the multi-view video matching model for processing to obtain the first video modality audiovisual representation includes:
[0011] Inputting the first-perspective video into the video encoder for representation extraction to obtain a first-perspective video representation;
[0012] Inputting the first-perspective speech into the audio encoder for representation extraction to obtain a first-perspective audio representation;
[0013] Cross-attention processing is performed on the first-perspective video representation and the first-perspective audio representation to obtain the first video modality audiovisual representation.
[0014] In some embodiments, inputting the first-perspective speech into the audio encoder for representation extraction to obtain the first-perspective audio representation includes:
[0015] Inputting the first-perspective speech data into the audio encoder for representation extraction to obtain an initial first-perspective speech representation;
[0016] Inputting the third-perspective speech data into the audio encoder for representation extraction to obtain an initial third-perspective speech representation;
[0017] Dynamic time warping is performed on the initial first-perspective speech representation and the initial third-perspective speech representation to obtain the first-perspective audio representation and the third-perspective audio representation.
[0018] In some embodiments, performing cross-attention processing on the first-perspective video representation and the first-perspective audio representation to obtain the first video modality audiovisual representation includes:
[0019] Using the first-view video representation as a first key vector and a first value vector, and using the first-view audio representation as a first query vector;
[0020] Cross-attention processing is performed based on the first key vector, the first value vector, and the first query vector to obtain an audiovisual representation of the first video modality.
[0021] In some embodiments, inputting the video description text corresponding to the video sample data into a text encoder in a multi-view video matching model for processing to obtain a text modality audiovisual representation includes:
[0022] Inputting the video description text into a large language model for audio description conversion to obtain an audio description text;
[0023] Inputting the video description text into the text encoder for representation extraction to obtain a video description text representation;
[0024] Inputting the audio description text into the text encoder for representation extraction to obtain an audio description text representation;
[0025] Cross-attention processing is performed on the video description text representation and the audio description text representation to obtain the text modality audiovisual representation.
[0026] In some embodiments, the step of inputting the video description text into a large language model for audio description conversion to obtain the audio description text includes:
[0027] Obtaining a video-audio conversion prompt, and inputting the video-audio conversion prompt into the large language model;
[0028] Obtaining a video-audio conversion example description, and inputting the video-audio conversion example description into the large language model;
[0029] The video description text is input into a large language model for audio description conversion to obtain an audio description text.
[0030] In some embodiments, performing cross-attention processing on the video description text representation and the audio description text representation to obtain the text modality audiovisual representation includes:
[0031] Using the video description text representation as a second key vector and a second value vector, and using the audio description text representation as a second query vector;
[0032] Cross-attention processing is performed based on the second key vector, the second value vector, and the second query vector to obtain the text modality audiovisual representation.
[0033] In some embodiments, the calculating the multi-view modality loss based on the first video modality audiovisual representation, the third video modality audiovisual representation, and the text modality audiovisual representation includes:
[0034] Calculating a first-perspective loss based on the first video modality audiovisual representation and the text modality audiovisual representation;
[0035] Calculating a third perspective loss based on the third video modality audiovisual representation and the text modality audiovisual representation;
[0036] The multi-view modal loss is obtained by accumulating the first view loss and the third view loss.
[0037] In some embodiments, the calculating the first perspective loss based on the first video modality audiovisual representation and the text modality audiovisual representation includes:
[0038] Get temperature parameters;
[0039] Calculating a first perspective similarity between the first video modality audiovisual representation and the text modality audiovisual representation, and performing exponential processing based on a ratio of the first perspective similarity to the temperature parameter to obtain a first perspective index similarity;
[0040] Accumulating the first viewing angle index similarities corresponding to all the multi-view video sample data to obtain a first viewing angle index test value;
[0041] Based on the ratio of the first viewing index similarity and the first viewing index test value, logarithmic processing is performed to obtain the first viewing angle sample loss;
[0042] All first-view sample losses are averaged to obtain the first-view loss.
[0043] In some embodiments, updating the multi-view video matching model based on the multi-view modality loss includes:
[0044] The video encoder and the text encoder in the multi-view video matching model are updated based on the multi-view modality loss.
[0045] To achieve the above-mentioned purpose, a second aspect of the embodiments of the present application provides a multi-view video matching method, the method comprising:
[0046] In response to a multi-view video matching task, obtaining view video data to be matched;
[0047] Based on the multi-perspective video matching task, the perspective video data to be matched is input into the multi-perspective video matching model trained by the multi-perspective video matching model training method described in the first aspect for video data matching processing to obtain target matching perspective video data, wherein the perspective video data to be matched and the target matching perspective video data are video data of different perspectives.
[0048] To achieve the above objectives, a third aspect of the embodiments of the present application provides a multi-view video matching model training device, the device comprising:
[0049] A sample data acquisition module is used to acquire multiple sets of video sample data and voice sample data and video description text corresponding to each of the video sample data, wherein the video sample data includes first-perspective video and third-perspective video, and the voice sample data includes first-perspective voice and third-perspective voice;
[0050] a video representation calculation module, configured to input the first-view video and the first-view speech into a video and audio encoder in a multi-view video matching model for processing to obtain a first video modality audiovisual representation, and input the third-view video and the third-view speech into the video and audio encoder for processing to obtain a third video modality audiovisual representation;
[0051] a text representation calculation module, configured to input the video description text corresponding to the video sample data into a text encoder in a multi-view video matching model for processing to obtain a text modality audiovisual representation;
[0052] A model updating module is used to calculate the multi-view modality loss based on the first video modality audio-visual representation, the third video modality audio-visual representation and the text modality audio-visual representation, and to update the multi-view video matching model based on the multi-view modality loss to obtain a trained multi-view video matching model.
[0053] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the multi-view video matching model training method as described in the first aspect or the multi-view video matching method as described in the second aspect.
[0054] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium, and the storage medium stores a computer program. When the computer program is executed by the processor, it implements the multi-view video matching model training method described in the first aspect or the multi-view video matching method described in the second aspect.
[0055] The multi-view video matching model training method, multi-view video matching method and related equipment proposed in the embodiments of the present application include: first, obtaining multiple groups of video sample data and voice sample data and video description text corresponding to each video sample data, the video sample data including first-view video and third-view video, and the voice sample data including first-view voice and third-view voice; second, inputting the first-view video and the first-view voice into the video and audio encoder in the multi-view video matching model for processing to obtain a first video modality audio-visual representation, and inputting the third-view video and the third-view voice into the video and audio encoder in the multi-view video matching model for processing to obtain a third video modality audio-visual representation; next, inputting the video description text corresponding to the video sample data into the text encoder in the multi-view video matching model for processing to obtain a text modality audio-visual representation; finally, calculating a multi-view modality loss based on the first video modality audio-visual representation, the third video modality audio-visual representation and the text modality audio-visual representation, and updating the multi-view video matching model based on the multi-view modality loss to obtain a trained multi-view video matching model. The embodiment of the present application simultaneously extracts video and audio features and fuses them to generate a video modal audio-visual representation, utilizes the fact that the audio in the first and third perspectives are similar, makes up for the lack of visual information from a single perspective, and then combines the text modal audio-visual representation as a semantic anchor. Through comparative learning, the distance between matching video pairs in the feature space is shortened, the semantic consistency of videos from different perspectives is enhanced, and the abstract semantics of the text modality are used to guide the multi-perspective video matching model to focus on the core features of events shared across perspectives, so as to reduce the interference of perspective differences, thereby utilizing the trained multi-perspective video matching model to further improve the accuracy of recognition and matching in downstream tasks such as cross-perspective matching, scene recognition, and action understanding.
[0056] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a flowchart of a multi-view video matching model training method provided in one embodiment of the present application.
[0058] Figure 2 yes Figure 1 Flowchart of step 102 in FIG.
[0059] Figure 3 yes Figure 2 Flowchart of step 202 in FIG.
[0060] Figure 4 yes Figure 2Flowchart of step 203 in FIG.
[0061] Figure 5 yes Figure 1 Flowchart of step 103 in FIG.
[0062] Figure 6 yes Figure 5 Flowchart of step 501 in FIG.
[0063] Figure 7 yes Figure 5 Flowchart of step 504 in FIG.
[0064] Figure 8 yes Figure 1 Flowchart of step 104 in FIG.
[0065] Figure 9 yes Figure 8 Flowchart of step 801 in FIG.
[0066] Figure 10 This is a schematic diagram of the training process of a multi-view video matching model provided in another embodiment of the present application.
[0067] Figure 11 This is a schematic block diagram of the training process of a multi-view video matching model provided by another embodiment of the present application.
[0068] Figure 12 This is a flowchart of a multi-view video matching method provided by another embodiment of the present application.
[0069] Figure 13 This is a structural diagram of a multi-view video matching model training device provided in another embodiment of the present application.
[0070] Figure 14 This is a schematic diagram of the hardware structure of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0071] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0072] It should be noted that although the functional modules are divided in the device schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flowchart.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0074] The ability to depict the activities of others from one's own perspective is an innate human skill known as cross-view correlation. In neural networks, this cross-view video correlation capability is represented by a cross-view video representation space, in which the resulting video feature vectors simultaneously incorporate information from different viewpoints. In practical applications, the use of a cross-view video representation space allows insufficient information from one viewpoint to be supplemented by correlating information from another. Therefore, matching video data from different viewpoints is essential.
[0075] Among them, cross-perspectives usually include the first perspective and the third perspective. The first perspective refers to the video captured from the participant's own perspective through a wearable device (such as a head-mounted camera, smart glasses), and its core feature is to simulate the subjective observation experience of humans. This perspective focuses on the operator's local movements (such as hand movements, object interactions), but is limited by the position of the device and may lack global environmental information (such as head movements, background scenes) or some limb movements (such as facial expressions). The third perspective refers to the video recorded from the perspective of a bystander through an external fixed device (such as a surveillance camera, a drone), and its core feature is to provide global scene information (such as spatial layout, multi-target interaction). This perspective can capture the complete motion trajectory and environmental context, but is easily affected by occlusion (such as the body occluding hand movements).
[0076] In related technologies, matching video data from different perspectives within the same scene at the same time in a video database typically involves constructing a cross-perspective video representation space based on visual information from multiple video data sets, thereby matching and associating first-perspective video data with third-perspective video data. However, due to differences in visual features between different perspectives—first-perspective video data focuses more on local scene recognition, while third-perspective video data focuses more on global scene recognition—this visual information-based cross-perspective video data matching and association method results in low matching accuracy.
[0077] In order to improve the matching accuracy of multi-perspective video data, the embodiment of the present application simultaneously extracts video and audio features and fuses them to generate video modal audio-visual representation, uses the fact that the audio in the first and third perspectives are similar to make up for the lack of visual information in a single perspective, and then combines the text modal audio-visual representation as a semantic anchor. Through comparative learning, the distance between matching video pairs in the feature space is shortened, the semantic consistency of videos from different perspectives is enhanced, and the abstract semantics of the text modality are used to guide the multi-perspective video matching model to focus on the core features of events shared across perspectives, so as to reduce the interference of perspective differences, thereby using the trained multi-perspective video matching model to further improve the accuracy of recognition matching in downstream tasks such as cross-perspective matching, scene recognition, and action understanding.
[0078] The following will further describe the multi-view video matching model training method, multi-view video matching method and related equipment provided by the embodiment of the present application. First, the multi-view video matching model training method in the embodiment of the present application is described. Figure 1 , which is an optional flowchart of the multi-view video matching model training method provided in an embodiment of the present application, Figure 1 The method may include but is not limited to steps 101 to 104. It is also understood that this embodiment is for Figure 1 The order of steps 101 to 104 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs. The multi-view video matching model training method provided in the embodiment of the present application can be applied to any server or smart terminal equipped with certain computing resources, etc.
[0079] Step 101: Acquire multiple groups of video sample data and the voice sample data and video description text corresponding to each video sample data.
[0080] Step 101 is described in detail below.
[0081] In some embodiments, in order to train a multi-view video matching model that can accurately match and identify multi-view video data in practical applications, it is necessary to pre-acquire a large number of multi-view training samples so that these multi-view training samples can be used to train and update the multi-view video matching model. These multi-view training samples typically include video sample data, corresponding voice sample data, and video description text. The video sample data includes first-view video and third-view video, and the voice sample data also includes first-view voice and third-view voice.
[0082] These multi-view training samples can be collected from commonly used image, text, and speech datasets on the internet. These datasets are processed to remove invalid and dirty data before being used as model training sets. These datasets typically include MSCOCO, SBU, OPENIMAGE, GOOGLECC, FLICKER30K, and GQA, totaling approximately 9.6 million multi-view and multi-modal data pairs.
[0083] The first-perspective voice may be audio extracted from the first-perspective video; similarly, the third-perspective voice may be audio extracted from the third-perspective video.
[0084] Step 102: Input the first-perspective video and the first-perspective speech into the video and audio encoder in the multi-perspective video matching model for processing to obtain a first video modality audio-visual representation, and input the third-perspective video and the third-perspective speech into the video and audio encoder for processing to obtain a third video modality audio-visual representation.
[0085] Step 102 is described in detail below.
[0086] In some embodiments, after obtaining a large number of multi-perspective training samples, the video and audio encoder in the multi-perspective video matching model is used to process each first-perspective video and the corresponding first-perspective speech to obtain a first video modality audio-visual representation of the first-perspective video modality; similarly, the video and audio encoder in the multi-perspective video matching model is used to process each third-perspective video and the third-perspective speech to obtain a third video modality audio-visual representation of the third-perspective video modality.
[0087] The video and audio encoder includes a video encoder and an audio encoder. Moreover, the processing flow for obtaining the first video modality audiovisual representation and the third video modality audiovisual representation is similar, and the following will be described using how to obtain the first video modality audiovisual representation as an example.
[0088] Reference Figure 2 , inputting the first-perspective video and the first-perspective voice into the video and audio encoder in the multi-perspective video matching model for processing to obtain a first video modality audiovisual representation, including the following steps 201 to 203.
[0089] Step 201: Input the first-perspective video into a video encoder for representation extraction to obtain a first-perspective video representation.
[0090] Step 202: Input the first-perspective speech into an audio encoder for representation extraction to obtain a first-perspective audio representation.
[0091] Steps 201 to 202 are described in detail below.
[0092] In some embodiments, for first-person video and first-person speech, to extract visual features from the video, the video encoder TimeSformer-B is integrated into the visual branch of the CLIP model to generate a video encoder. The first-person video is then input into the video encoder to extract a video representation, thereby generating an output first-person video representation. Similarly, the third-person video is also input into the video encoder to extract a video representation, thereby generating an output third-person video representation.
[0093] Next, we use the BEATs model as an audio encoder to extract representations of the first-person and third-person voices, respectively, to obtain corresponding first-person and third-person audio representations. However, since the audio produced by videos from different perspectives may have different durations, corresponding time processing operations are required, as described below.
[0094] Reference Figure 3 , inputting the first-perspective speech into the audio encoder for representation extraction to obtain the first-perspective audio representation, including the following steps 301 to 303.
[0095] Step 301: Input the first-perspective speech data into an audio encoder for representation extraction to obtain an initial first-perspective speech representation.
[0096] Step 302: Input the third-perspective speech data into an audio encoder for representation extraction to obtain an initial third-perspective speech representation.
[0097] Step 303: Perform dynamic time warping on the initial first-perspective speech representation and the initial third-perspective speech representation to obtain a first-perspective audio representation and a third-perspective audio representation.
[0098] Steps 301 to 303 are described in detail below.
[0099] In some embodiments, the first-perspective speech data and the third-perspective speech data are respectively characterized and extracted using an existing audio encoder to obtain the corresponding initial first-perspective speech representation A. ego and the initial third-perspective speech representation A exo .
[0100] Next, the Dynamic Time Warping (DTW) algorithm is used to represent the initial first-person perspective speech A of the matched video pair. ego and the initial third-perspective speech representation A exoDynamic time warping is performed to obtain the first-perspective audio representation and the third-perspective audio representation as shown in the following formula (1), so that the two audio features have the same time length and are more consistent in the occurrence time of the same event.
[0101]
[0102] The Dynamic Time Warping (DTW) algorithm is used to align time series data (such as audio and action sequences) of different lengths or inconsistent time axes. It uses dynamic programming to find the optimal matching path, so that two sequences can be nonlinearly aligned in the time dimension.
[0103] Through the above steps 301 to 303, the time offset is eliminated through the nonlinear alignment of DTW to address the problem of inconsistent audio time axes between the first and third perspectives, ensuring that the audio features of the same event strictly match in the time dimension; 2) Cross-perspective information complementarity enables the aligned audio representation to integrate acoustic details from different perspectives, enhance the capture of the complete semantics of the event, and make the fusion of audio features with visual and textual modal features more coordinated, providing a temporally consistent cross-modal association foundation for subsequent cross-perspective comparative learning.
[0104] Step 203: Perform cross-attention processing on the first-person perspective video representation and the first-person perspective audio representation to obtain a first video modality audiovisual representation.
[0105] Step 203 is described in detail below.
[0106] In some embodiments, next, the first-perspective video representation and the first-perspective audio representation are cross-attention processed to obtain a first video modality audio-visual representation that is a mixture of the video features and audio features of the first perspective; similarly, the third-perspective video representation and the third-perspective audio representation are cross-attention processed to obtain a third video modality audio-visual representation that is a mixture of the video features and audio features of the third perspective. How to perform cross-attention processing of video representation and audio representation will be further described below.
[0107] Reference Figure 4 , cross-attention processing is performed on the first-perspective video representation and the first-perspective audio representation to obtain a first video modality audiovisual representation, including the following steps 401 to 402.
[0108] Step 401: Use the first-perspective video representation as a first key vector and a first value vector, and use the first-perspective audio representation as a first query vector.
[0109] Step 402: Perform cross-attention processing based on the first key vector, the first value vector, and the first query vector to obtain a first video modality audiovisual representation.
[0110] Steps 401 to 402 are described in detail below.
[0111] In some embodiments, for video representation (including first-view video representation and third-view video representation) and audio representation (including first-view audio representation and third-view audio representation), the video representation is used as the first key vector K=F V and the first value vector V = F V , take the audio representation as the first query vector Q = F A .
[0112] Then, a cross-attention process CrossAttn() is performed based on the first key vector, the first value vector, and the first query vector to obtain the video modality audiovisual representation V VA As shown in the following formula (2).
[0113] V VA =CrossAttn(Q=F A , K=F V , V=F V ) (2)
[0114] Among them, the video modality audiovisual representation corresponds to the first video modality audiovisual representation of the first perspective: The third video modality audiovisual representation corresponding to the third perspective is Where i is the sample data index in the multi-view training sample, ego represents the first perspective, and exo represents the third perspective.
[0115] Through the above steps 201 to 203, and steps 401 to 402, with audio features as query vectors and video features as keys and values, the audio information dynamically filters visual clips related to acoustic events, enhances the ability to capture local key actions, and adaptively adjusts the temporal correlation between video and audio through attention weights to alleviate the problem of partial visual information loss from a single perspective (for example, when hand movements are blocked in the first perspective, audio features guide the model to focus on the video context of the corresponding time period). The fused video modality audio-visual representation retains both the spatiotemporal details of the video and the temporal event clues of the audio, providing a semantically aligned multimodal basis for cross-perspective comparative learning, so as to facilitate the subsequent use of the video modality audio-visual representation to train the multi-perspective video matching model, thereby improving the hybrid learning of audio features and video features by the multi-perspective video matching model.
[0116] Step 103: Input the video description text corresponding to the video sample data into the text encoder in the multi-view video matching model for processing to obtain a text modality audiovisual representation.
[0117] Step 103 is described in detail below.
[0118] In some embodiments, while performing video modality audiovisual representation processing, the video description text corresponding to the video sample data is also input into the text encoder in the multi-view video matching model for relevant processing to obtain a text modality audiovisual representation from a text modality perspective.
[0119] In this embodiment, a text encoder of the CLIP model is used to encode the video description text to obtain a representation of the text, which is described in detail as follows.
[0120] Reference Figure 5 , inputting the video description text corresponding to the video sample data into the text encoder in the multi-view video matching model for processing to obtain a text modality audiovisual representation, including the following steps 501 to 504.
[0121] Step 501: Input the video description text into the large language model for audio description conversion to obtain audio description text.
[0122] Step 501 is described in detail below.
[0123] In some embodiments, in order to effectively utilize the audio information in the video, the existing video description text is input into a large language model (LLM) for audio description conversion to generate corresponding audio description text, as described below.
[0124] Reference Figure 6 , input the video description text into the large language model for audio description conversion to obtain the audio description text, including the following steps 601 to 603.
[0125] Step 601: Obtain a video-audio conversion prompt, and input the video-audio conversion prompt into a large language model.
[0126] Step 602: Obtain a video-audio conversion example description, and input the video-audio conversion example description into a large language model.
[0127] Step 603: Input the video description text into the large language model for audio description conversion to obtain audio description text.
[0128] Steps 601 to 603 are described in detail below.
[0129] In some embodiments, the primary task in this process is prompt design, which includes three key elements: context, requirements, and examples. Below, we explain and provide examples for each design element. Regarding context, using a video-to-speech prompt, the large language model (LLM) is assigned a role, limiting its output to the audio description of the video (for example, the video-to-speech prompt is: "You are an expert in the audio and visual description of the video"). This helps the large language model (LLM) limit its generated content to relevant audio descriptions. Next, more detailed requirements are provided using the video-to-audio example description (for example, the video-to-audio example description is: "Ensure that the audio description is concise, accurate, and does not repeat the visual description") to specify the details of the audio description that the model should generate. Finally, an example integrating the above elements is provided, providing a specific reference for the large language model and clarifying the format of the output audio description (for example, "washing dishes" or "splashing water and clinking glasses"). Based on the prompts designed based on the three elements above, the video description text is input into the large language model for audio description conversion to obtain the corresponding audio description text.
[0130] Through the above steps 601 to 603, the video-audio conversion prompt is used to define the context (such as "audio description expert") and the video-audio conversion example description is used to perform example constraints, ensuring that the generated audio description text (such as "the sound of glasses colliding") and the video description text (such as "the action of washing dishes") complement each other, avoiding information redundancy and enhancing the semantic consistency of multimodal association. Combined with the requirement instructions (such as "concise and accurate") and example references, the audio description generated by the large language model focuses on the core features of the acoustic event and adapts to the cross-view association task's need to capture key details, thereby facilitating the use of the generated audio description text as a cross-modal anchor point, providing structured semantic input for the construction of subsequent text modality audio-visual representation, and supporting the robust alignment of cross-view video and audio features in the comparative learning process.
[0131] Step 502: Input the video description text into a text encoder for representation extraction to obtain a video description text representation.
[0132] Step 503: Input the audio description text into a text encoder for representation extraction to obtain an audio description text representation.
[0133] Step 504: Perform cross-attention processing on the video description text representation and the audio description text representation to obtain a text modality audiovisual representation.
[0134] Steps 502 to 504 are described in detail below.
[0135] In some embodiments, the text encoder is used to extract representations of the video audio conversion example description and the audio description text to obtain the corresponding video description text representation T V and audio description text representation T A .
[0136] Next, the video description text representation T V and audio description text representation T A Perform cross-attention processing to obtain a text modality audiovisual representation T that is a mixture of video description features and audio description features VA , as described below.
[0137] Reference Figure 7 , cross-attention processing is performed on the video description text representation and the audio description text representation to obtain a text modality audiovisual representation, including the following steps 701 to 702.
[0138] Step 701: Use the video description text representation as a second key vector and a second value vector, and use the audio description text representation as a second query vector.
[0139] Step 702: Perform cross-attention processing based on the second key vector, the second value vector, and the second query vector to obtain a text modal audiovisual representation.
[0140] Steps 701 to 702 are described in detail below.
[0141] In some embodiments, for a video description text representation T V and audio description text representation T A , the video description text is represented as the second key vector K = T V and the second value vector V = T V , the audio description text representation T A As the second query vector Q=T A .
[0142] Then, a cross-attention process CrossAttn() is performed based on the second key vector, the second value vector, and the second query vector to obtain the text modality audiovisual representation T VA , as shown in the following formula (3).
[0143] T VA =CrossAttn(Q=T A ,K=T V ,V=T V ) (3)
[0144] Through the above steps 501 to 504, and steps 701 to 702, with the audio description text as the query vector (Q) and the video description text as the key (K) and value (V), the semantics of acoustic events and visual actions are dynamically associated through attention weights to generate fine-grained aligned text modality representations, and cross-attention is used to screen video text features that are highly relevant to the audio description (such as ignoring redundant scene descriptions in the video text), enhance the semantic expression of core events shared across perspectives, and use the fused text modality audio-visual representation as a semantic anchor for comparative learning to provide a unified semantic alignment target for cross-perspective video modality features, guide the model to focus on common features independent of perspective, and improve the robustness and interpretability of cross-perspective associations.
[0145] Step 104: Calculate a multi-view modality loss based on the first video modality audio-visual representation, the third video modality audio-visual representation, and the text modality audio-visual representation, and update the multi-view video matching model based on the multi-view modality loss to obtain a trained multi-view video matching model.
[0146] Step 104 is described in detail below.
[0147] In some embodiments, after obtaining the first video modality audiovisual representation corresponding to the first perspective Audiovisual representation of the third video modality corresponding to the third perspective and textual modal audiovisual representation T VA Afterwards, the multi-view modality loss L between the three is calculated total , thereby utilizing the multi-view modality loss L total The video encoder and text encoder in the multi-view video matching model are updated.
[0148] The following will further describe how to calculate the multi-view modality loss L total .
[0149] Reference Figure 8 , calculating the multi-view modality loss based on the first video modality audiovisual representation, the third video modality audiovisual representation and the text modality audiovisual representation, includes the following steps 801 to 803.
[0150] Step 801: Calculate the first perspective loss based on the first video modality audiovisual representation and the text modality audiovisual representation.
[0151] Step 801 is described in detail below.
[0152] In some embodiments, for the first perspective, the audio-visual text features of the text modality are used as anchors to align the cross-view video audio-visual representations. This reduces the distance between the matching video representations and constructs a cross-view association space. That is, in this embodiment, based on the audio-visual representation of the first video modality and textual modal audiovisual representation T VA , calculate the first-view loss L ego , as described below.
[0153] Reference Figure 9 , based on the first video modality audiovisual representation and the text modality audiovisual representation, the first perspective loss is calculated, including the following steps 901 to 905.
[0154] Step 901: Acquire temperature parameters.
[0155] Step 902: Calculate the first perspective similarity between the first video modality audiovisual representation and the text modality audiovisual representation, and perform exponential processing based on the ratio of the first perspective similarity and the temperature parameter to obtain the first perspective index similarity.
[0156] Step 903: Accumulate the first viewing angle index similarities corresponding to all multi-view video sample data to obtain a first viewing angle index test value.
[0157] Step 904: Based on the ratio of the first-view index similarity and the first-view index test value, logarithmic processing is performed to obtain the first-view sample loss.
[0158] Step 905: average all first-perspective sample losses to obtain the first-perspective loss.
[0159] Steps 901 to 905 are described in detail below.
[0160] In some embodiments, for the i-th video sample data in the multi-view training sample, the corresponding temperature parameter τ is determined, and then the similarity function sim() is used to calculate the first video modality audiovisual representation and textual modal audiovisual representation T VA First-person perspective similarity Where j is the index value of the video description text corresponding to the i-th video sample data in the multi-view training sample.
[0161] Then, based on the ratio of the first-view similarity and the temperature parameter, the index processing is performed to obtain the first-view index similarity And, accumulate the first-view index similarity corresponding to all multi-view video sample data to obtain the first-view index test value Based on the ratio of the first perspective index similarity and the first perspective index test value, logarithmic processing is performed to obtain the first perspective sample loss corresponding to the first perspective of the i-th video sample data As shown in the following formula (4).
[0162]
[0163] Here, m refers to the total number of sample data of multi-view training samples.
[0164] Finally, the first-view sample loss of all multi-view video sample data Perform average processing to obtain the first-perspective loss L ego As shown in the following formula (5).
[0165]
[0166] Step 802: Calculate a third perspective loss based on the third video modality audiovisual representation and the text modality audiovisual representation.
[0167] Step 803: Accumulate the first-view loss and the third-view loss to obtain the multi-view modal loss.
[0168] Steps 801 to 803 are described in detail below.
[0169] In some embodiments, the first viewing angle loss L corresponding to the first viewing angle is obtained. ego Similarly, based on the third video modality audiovisual representation corresponding to each training data sample and textual modal audiovisual representation T VA , the third perspective loss L is calculated by the following formula (6) and formula (7): exo .
[0170]
[0171] After obtaining the first-perspective loss L corresponding to each training data sample ego and the third perspective loss L exo Afterwards, the losses of the two views are summed to obtain the multi-view modality loss, as shown in the following formula (8).
[0172] L total =L ego +L exo (8)
[0173] Through the above steps 801 to 803, and steps 901 to 905, the similarity distribution is dynamically scaled by the temperature coefficient, the learning weights of positive and negative samples are balanced, and the overfitting or underfitting problem of the model caused by excessive similarity difference is alleviated. The similarity-related loss function is used to maximize the similarity between the audio-visual representation of the video modality and the audio-visual representation of the text modality, and by independently calculating the first-view loss and the third-view loss and jointly optimizing them, the model's collaborative perception ability of local actions (first view) and global scenes (third view) is enhanced, and a perspective-independent semantic alignment space is constructed. The generalization performance of cross-modal retrieval and downstream tasks (such as action recognition in occluded scenes) is improved, so that when the obtained multi-view modal loss is used to train the multi-view video matching model, the recognition matching degree of the multi-view video matching model for videos of different perspectives can be greatly improved.
[0174] In some embodiments, after obtaining the multimodal loss L total Afterwards, the multi-view video matching model can update and optimize the video encoder and text encoder through stochastic gradient descent (SGD) or other optimization algorithms.
[0175] Reference Figure 10 , is a schematic diagram of the training process of a multi-view video matching model provided in an embodiment of the present application. Figure 10 As shown in the figure, the video, audio, and video / audio description text generated by LLM are respectively extracted through dedicated encoders (video encoder, audio encoder, text encoder) to generate corresponding video representation, audio representation, and text representation. The dynamic interaction between video and audio modalities is achieved through the cross-attention mechanism (two cross-attention modules in the figure), and the multi-source representations are fused into video modality audio-visual representation and text modality audio-visual representation. Finally, the relevant parameters of the video encoder and text encoder in the multi-view video matching model are optimized through contrastive learning (loss function calculation), achieving robust association when visual information is missing in cross-view scenarios. This process embodies the core innovation of audio-driven feature alignment strategy and unified modeling of multimodal representation.
[0176] Reference Figure 11 , is a schematic diagram of the training process of a multi-view video matching model provided in an embodiment of the present application. Figure 11As shown in [1], an end-to-end video processing system architecture based on multimodal fusion and cross-view association is presented, covering both training and inference stages. During the training phase, the first-person view video, third-person view video, and corresponding audio are respectively generated through a video encoder, an audio encoder, and a large language model to generate video features, audio features, and text descriptions (including automatically generated "video descriptions" and "audio descriptions"). Multi-layer cross-attention modules enable dynamic interaction between the video and audio modalities, and the dynamic time-programming algorithm (DTW) is used to optimize temporal alignment. During the inference phase, the system uses a joint representation generated by contrastive learning (LContrast) to perform cross-view retrieval (based on similarity calculation), scene classification (Task 1), and action recognition (Task 3). Scene classification outputs a "scene label" through a fully connected layer, and action classification outputs an "action classification result." The overall architecture highlights the core innovations of audio modality-driven feature alignment and multi-source information fusion, achieving robust association in the absence of visual information in cross-view scenarios through contrastive learning.
[0177] In addition, the embodiment of the present application proposes a multi-view video matching method. Figure 12 , which is an optional flowchart of the multi-view video matching method provided in an embodiment of the present application, Figure 12 The method may include but is not limited to steps 1201 to 1204. It is also understood that this embodiment is for Figure 12 The order of steps 1201 to 1204 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs. The method can be applied to a server or intelligent terminal equipped with a knowledge graph and a large language model.
[0178] Step 1201: In response to a multi-view video matching task, obtain view video data to be matched.
[0179] Step 1202: Based on the multi-view video matching task, the view video data to be matched is input into the multi-view video matching model trained by the multi-view video matching model training method to perform video data matching processing to obtain the target matching view video data.
[0180] Steps 1201 to 1202 are described in detail below.
[0181] In some embodiments, after obtaining the trained multi-perspective video matching model through the training process of the multi-perspective video matching model mentioned above, in actual applications, in response to an actual multi-perspective video data matching task, the to-be-matched perspective video data corresponding to the multi-perspective video data matching task is input into the multi-perspective video matching model to perform corresponding video data matching processing, and the target matching perspective video data of another perspective of the multimodal recognition processing task is obtained.
[0182] The multi-perspective video matching model training method, multi-perspective video matching method and related equipment proposed in the embodiment of the present application include: first, obtaining multiple groups of video sample data and voice sample data and video description text corresponding to each video sample data, the video sample data including first-perspective video and third-perspective video, and the voice sample data including first-perspective voice and third-perspective voice; second, inputting the first-perspective video into a video encoder for characterization extraction to obtain a first-perspective video representation, inputting the first-perspective voice data into an audio encoder for characterization extraction to obtain an initial first-perspective voice representation, inputting the third-perspective voice data into an audio encoder for characterization extraction to obtain an initial third-perspective voice representation, and performing characterization extraction on the initial first-perspective voice. The representation and the initial third-perspective speech representation are dynamically time-warped to obtain the first-perspective audio representation and the third-perspective audio representation, the first-perspective video representation is used as the first key vector and the first value vector, the first-perspective audio representation is used as the first query vector, and cross-attention processing is performed based on the first key vector, the first value vector and the first query vector to obtain the first video modality audio-visual representation, the third-perspective video and the third-perspective speech are input into the video and audio encoder in the multi-perspective video matching model for processing to obtain the third video modality audio-visual representation; next, the video and audio conversion prompt is obtained, and the video and audio conversion prompt is input into the large language model, the video and audio conversion example description is obtained, and the video and audio conversion example description is input into the large language model. Language model, input the video description text into the large language model for audio description conversion to obtain audio description text, input the video description text into the text encoder for representation extraction to obtain video description text representation, input the audio description text into the text encoder for representation extraction to obtain audio description text representation, use the video description text representation as the second key vector and the second value vector, use the audio description text representation as the second query vector, perform cross-attention processing based on the second key vector, the second value vector and the second query vector, and obtain text modal audiovisual representation; finally, obtain the temperature parameter, calculate the first perspective similarity of the first video modal audiovisual representation and the text modal audiovisual representation, and based on the ratio of the first perspective similarity and the temperature parameter , then perform exponential processing to obtain the first-view index similarity, accumulate the first-view index similarities corresponding to all multi-view video sample data, and obtain the first-view index test value. Based on the ratio of the first-view index similarity and the first-view index test value, perform logarithmic processing to obtain the first-view sample loss, average all first-view sample losses to obtain the first-view loss, and calculate the third-view loss based on the third video modality audio-visual representation and the text modality audio-visual representation. Accumulate the first-view loss and the third-view loss to obtain the multi-view modality loss, and update the video encoder and text encoder in the multi-view video matching model based on the multi-view modality loss to obtain the trained multi-view video matching model.
[0183] The embodiment of the present application simultaneously extracts video and audio features and fuses them to generate video modal audio-visual representation, uses the fact that the audio in the first and third perspectives are similar to make up for the lack of visual information from a single perspective, and then combines the text modal audio-visual representation as a semantic anchor. Through comparative learning, the distance between matching video pairs in the feature space is shortened, the semantic consistency of videos from different perspectives is enhanced, and the abstract semantics of the text modality are used to guide the multi-perspective video matching model to focus on the core features of events shared across perspectives, so as to reduce the interference of perspective differences, thereby using the trained multi-perspective video matching model to further improve recognition matching in downstream tasks such as cross-perspective matching, scene recognition, and action understanding. Accuracy; in addition, in order to solve the problem of inconsistent audio time axes between the first and third perspectives, the time offset is eliminated through DTW nonlinear alignment to ensure that the audio features of the same event are strictly matched in the time dimension; 2) Cross-perspective information complementarity enables the aligned audio representation to integrate acoustic details from different perspectives, enhance the capture of the complete semantics of the event, and make the fusion of audio features with visual and textual modal features more coordinated, providing a temporally consistent cross-modal association basis for subsequent cross-perspective comparative learning; and, using audio features as query vectors and video features as keys and values, the audio information dynamically filters visual clips related to acoustic events, enhancing the ability to capture local key actions. The fusion of audio and video features can be used to improve the learning process of the video and audio, and to improve the learning process of the audio and video features. The generated audio description text (such as "the sound of glasses colliding") and the video-audio conversion example description are constrained to ensure that the generated audio description text (such as "glass clash") complements the video description text (such as "dishwashing action"), avoiding information redundancy and enhancing the semantic consistency of multimodal association. Combined with the requirement instructions (such as "concise and accurate") and example references, the audio description generated by the large language model focuses on the core features of the acoustic event and adapts to the cross-view association task's need to capture key details. This makes it easier to use the generated audio description text as a cross-modal anchor, providing structured semantic input for the subsequent construction of text modality audio-visual representation, and supporting the robust alignment of cross-view video and audio features in the contrastive learning process;In addition, the audio description text is used as the query vector (Q) and the video description text as the key (K) and value (V), and the semantics of acoustic events and visual actions are dynamically associated through attention weights to generate fine-grained aligned text modality representations. Cross-attention is used to screen video text features that are highly relevant to the audio description (such as ignoring redundant scene descriptions in the video text), enhance the semantic expression of core events shared across perspectives, and use the fused text modality audio-visual representation as a semantic anchor for comparative learning to provide a unified semantic alignment target for cross-perspective video modality features, guide the model to focus on common features independent of perspective, and improve the robustness and interpretability of cross-perspective associations. In addition, the temperature coefficient is used to dynamically scale the similarity distribution to balance positive and negative The learning weights of the samples alleviate the model's overfitting or underfitting problems caused by large similarity differences. Using a similarity-related loss function, the similarity between the audiovisual representations of the video modality and the audiovisual representations of the text modality is maximized. By independently calculating the first-view loss and the third-view loss and jointly optimizing them, the model's collaborative perception of local actions (first-view) and global scenes (third-view) is enhanced, a perspective-independent semantic alignment space is constructed, and the generalization performance of cross-modal retrieval and downstream tasks (such as action recognition in occluded scenes) is improved. Therefore, when using this multi-view modality loss to train a multi-view video matching model, the recognition matching degree of the multi-view video matching model for videos from different perspectives can be greatly improved.
[0184] The embodiment of the present application also provides a multi-view video matching model training device, which can implement the multi-view video matching model training method described above. Figure 13 , the apparatus 1300 comprises:
[0185] The sample data acquisition module 1310 is used to acquire multiple sets of video sample data and voice sample data and video description text corresponding to each video sample data. The video sample data includes first-perspective video and third-perspective video, and the voice sample data includes first-perspective voice and third-perspective voice.
[0186] The video representation calculation module 1320 is configured to input the first-view video and the first-view speech into the video and audio encoder in the multi-view video matching model for processing to obtain a first video modality audiovisual representation, and input the third-view video and the third-view speech into the video and audio encoder for processing to obtain a third video modality audiovisual representation;
[0187] A text representation calculation module 1330 is configured to input the video description text corresponding to the video sample data into a text encoder in the multi-view video matching model for processing to obtain a text modality audiovisual representation;
[0188] The model updating module 1340 is used to calculate the multi-view modality loss based on the first video modality audio-visual representation, the third video modality audio-visual representation and the text modality audio-visual representation, and update the multi-view video matching model based on the multi-view modality loss to obtain a trained multi-view video matching model.
[0189] In some embodiments, the video representation calculation module 1320 is further configured to:
[0190] Inputting the first-perspective video into a video encoder for representation extraction to obtain a first-perspective video representation;
[0191] Inputting the first-perspective speech into the audio encoder for representation extraction to obtain the first-perspective audio representation;
[0192] Cross-attention processing is performed on the first-person perspective video representation and the first-person perspective audio representation to obtain the first video modality audiovisual representation.
[0193] In some embodiments, the video representation calculation module 1320 is further configured to:
[0194] Inputting the first-person perspective speech data into an audio encoder for representation extraction to obtain an initial first-person perspective speech representation;
[0195] Input the third-perspective speech data into the audio encoder for representation extraction to obtain an initial third-perspective speech representation;
[0196] Dynamic time warping is performed on the initial first-perspective speech representation and the initial third-perspective speech representation to obtain a first-perspective audio representation and a third-perspective audio representation.
[0197] In some embodiments, the video representation calculation module 1320 is further configured to:
[0198] Using the first-person perspective video representation as a first key vector and a first value vector, and using the first-person perspective audio representation as a first query vector;
[0199] Cross-attention processing is performed based on the first key vector, the first value vector, and the first query vector to obtain a first video modality audiovisual representation.
[0200] In some embodiments, the text representation calculation module 1330 is further configured to:
[0201] Input the video description text into the large language model for audio description conversion to obtain audio description text;
[0202] Input the video description text into the text encoder for representation extraction to obtain the video description text representation;
[0203] Input the audio description text into a text encoder for representation extraction to obtain an audio description text representation;
[0204] Cross-attention processing is performed on the video description text representation and the audio description text representation to obtain a textual modality audiovisual representation.
[0205] In some embodiments, the text representation calculation module 1330 is further configured to:
[0206] Obtain video-audio conversion prompts and input the video-audio conversion prompts into a large language model;
[0207] Obtain a description of a video-audio conversion example, and input the description of the video-audio conversion example into a large language model;
[0208] The video description text is input into the large language model for audio description conversion to obtain audio description text.
[0209] In some embodiments, the text representation calculation module 1330 is further configured to:
[0210] Using the video description text representation as a second key vector and a second value vector, and using the audio description text representation as a second query vector;
[0211] Cross-attention processing is performed based on the second key vector, the second value vector, and the second query vector to obtain a text modal audiovisual representation.
[0212] In some embodiments, the model updating module 1340 is further configured to:
[0213] Based on the first video modality audiovisual representation and the text modality audiovisual representation, a first-perspective loss is calculated;
[0214] Based on the third video modality audiovisual representation and the text modality audiovisual representation, the third perspective loss is calculated;
[0215] The multi-view modality loss is obtained by accumulating the first-view loss and the third-view loss.
[0216] In some embodiments, the model updating module 1340 is further configured to:
[0217] Get temperature parameters;
[0218] Calculating the first-perspective similarity between the audiovisual representation of the first video modality and the audiovisual representation of the text modality, and performing exponential processing based on the ratio of the first-perspective similarity and the temperature parameter to obtain the first-perspective index similarity;
[0219] Accumulate the first-view index similarities corresponding to all multi-view video sample data to obtain the first-view index test value;
[0220] Based on the ratio of the first-view index similarity and the first-view index test value, logarithmic processing is performed to obtain the first-view sample loss;
[0221] The first-view loss is obtained by averaging all first-view sample losses.
[0222] In some embodiments, the model updating module 1340 is further configured to:
[0223] The video encoder and text encoder in the multi-view video matching model are updated based on the multi-view modality loss.
[0224] In the above embodiments, the description of each embodiment has its own focus. For the part that is not described in detail in a certain embodiment, the specific implementation method of the multi-view video matching model training device is basically the same as the specific implementation method of the above-mentioned multi-view video matching model training method, and will not be repeated here.
[0225] In the embodiment of the present application, the multi-view video matching model training device simultaneously extracts video and audio features and fuses them to generate video modal audio-visual representation, uses the fact that the audio in the first and third perspectives are similar to make up for the lack of visual information from a single perspective, and then combines the text modal audio-visual representation as a semantic anchor. Through comparative learning, the distance between matching video pairs in the feature space is shortened, the semantic consistency of videos from different perspectives is enhanced, and the abstract semantics of the text modality is used to guide the multi-view video matching model to focus on the core features of events shared across perspectives, so as to reduce the interference of perspective differences, thereby using the trained multi-view video matching model to perform downstream tasks such as cross-perspective matching, scene recognition, and action understanding. In the service, the accuracy of recognition and matching is further improved; in addition, in order to solve the problem of inconsistent audio time axes between the first and third perspectives, the time offset is eliminated through the nonlinear alignment of DTW to ensure that the audio features of the same event are strictly matched in the time dimension; 2) Cross-perspective information is complementary, so that the aligned audio representation can integrate the acoustic details of different perspectives, enhance the capture of the complete semantics of the event, and make the fusion of audio features with visual and textual modal features more coordinated, providing a temporally consistent cross-modal association basis for subsequent cross-perspective comparative learning; and, using audio features as query vectors and video features as keys and values, the audio information dynamically filters the visual clips related to the acoustic event, enhancing the local focus. The fusion of audio and video features can improve the ability to capture key movements, and adaptively adjust the time correlation between video and audio through attention weights to alleviate the problem of partial visual information loss from a single perspective (for example, when the hand movement is blocked in the first perspective, the audio features guide the model to focus on the video context of the corresponding time period). The fused video modality audio-visual representation retains the spatiotemporal details of the video and the temporal event clues of the audio, providing a semantically aligned multimodal basis for cross-perspective comparative learning, so that when the video modality audio-visual representation is subsequently used to train the multi-perspective video matching model, the hybrid learning of audio features and video features of the multi-perspective video matching model can be improved; in addition, the video and audio conversion prompts are used for context definition (such as "audio The system uses example constraints on the generated audio description text (such as "the sound of glasses colliding") and the video-audio conversion example description to ensure that the generated audio description text (such as "the sound of glasses colliding") complements the video description text (such as "the action of washing dishes"), avoids information redundancy, and enhances the semantic consistency of multimodal association. Combined with requirement instructions (such as "concise and accurate") and example references, the audio description generated by the large language model focuses on the core features of the acoustic event and adapts to the cross-view association task's need to capture key details. This makes it easier to use the generated audio description text as a cross-modal anchor, providing structured semantic input for the subsequent construction of text modality audiovisual representations, and supporting the robust alignment of cross-view video and audio features in the contrastive learning process;In addition, the audio description text is used as the query vector (Q) and the video description text as the key (K) and value (V), and the semantics of acoustic events and visual actions are dynamically associated through attention weights to generate fine-grained aligned text modality representations. Cross-attention is used to screen video text features that are highly relevant to the audio description (such as ignoring redundant scene descriptions in the video text), enhance the semantic expression of core events shared across perspectives, and use the fused text modality audio-visual representation as a semantic anchor for comparative learning to provide a unified semantic alignment target for cross-perspective video modality features, guide the model to focus on common features independent of perspective, and improve the robustness and interpretability of cross-perspective associations. In addition, the temperature coefficient is used to dynamically scale the similarity distribution to balance positive and negative The learning weights of the samples alleviate the model's overfitting or underfitting problems caused by large similarity differences. Using a similarity-related loss function, the similarity between the audiovisual representations of the video modality and the audiovisual representations of the text modality is maximized. By independently calculating the first-view loss and the third-view loss and jointly optimizing them, the model's collaborative perception of local actions (first-view) and global scenes (third-view) is enhanced, a perspective-independent semantic alignment space is constructed, and the generalization performance of cross-modal retrieval and downstream tasks (such as action recognition in occluded scenes) is improved. Therefore, when using this multi-view modality loss to train a multi-view video matching model, the recognition matching degree of the multi-view video matching model for videos from different perspectives can be greatly improved.
[0226] An embodiment of the present application further provides an electronic device, including:
[0227] at least one memory;
[0228] at least one processor;
[0229] at least one program;
[0230] The program is stored in the memory, and the processor executes at least one program to implement the multi-view video matching model training method implemented in this application. The electronic device can be any smart terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0231] See also Figure 14 , Figure 14 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0232] The processor 1401 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0233] The memory 1402 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device or RAM (Random Access Memory). The memory 1402 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1402 and is called by the processor 1401 to execute the multi-view video matching model training method of the embodiments of this application;
[0234] Input / output interface 1403, used to implement information input and output;
[0235] Communication interface 1404, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0236] Bus 1405 , which transmits information between various components of the device (e.g., processor 1401 , memory 1402 , input / output interface 1403 , and communication interface 1404 );
[0237] The processor 1401 , the memory 1402 , the input / output interface 1403 and the communication interface 1404 are connected to each other in communication within the device via a bus 1405 .
[0238] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned multi-view video matching model training method is implemented.
[0239] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0240] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0241] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0242] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0243] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0244] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0245] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0246] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0247] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0248] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0249] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.
[0250] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A multi-view video matching model training method, characterized in that: The method comprises: Acquire multiple sets of video sample data and voice sample data and video description text corresponding to each of the video sample data, wherein the video sample data includes a first-perspective video and a third-perspective video, and the voice sample data includes a first-perspective voice and a third-perspective voice; Inputting the first-perspective video and the first-perspective speech into a video and audio encoder in a multi-perspective video matching model for processing to obtain a first video modality audiovisual representation, and inputting the third-perspective video and the third-perspective speech into the video and audio encoder for processing to obtain a third video modality audiovisual representation; Inputting the video description text corresponding to the video sample data into a text encoder in a multi-view video matching model for processing to obtain a text modality audiovisual representation; A multi-view modality loss is calculated based on the first video modality audio-visual representation, the third video modality audio-visual representation and the text modality audio-visual representation, and the multi-view video matching model is updated based on the multi-view modality loss to obtain a trained multi-view video matching model.
2. The multi-view video matching model training method according to claim 1, characterized in that: The video and audio encoder includes a video encoder and an audio encoder, and the step of inputting the first-view video and the first-view speech into the video and audio encoder in the multi-view video matching model for processing to obtain a first video modality audiovisual representation includes: Inputting the first-perspective video into the video encoder for representation extraction to obtain a first-perspective video representation; Inputting the first-perspective speech into the audio encoder for representation extraction to obtain a first-perspective audio representation; Cross-attention processing is performed on the first-perspective video representation and the first-perspective audio representation to obtain the first video modality audiovisual representation.
3. The multi-view video matching model training method according to claim 2, characterized in that: The step of inputting the first-perspective speech into the audio encoder for characterization extraction to obtain the first-perspective audio representation includes: Inputting the first-perspective speech data into the audio encoder for representation extraction to obtain an initial first-perspective speech representation; Inputting the third-perspective speech data into the audio encoder for representation extraction to obtain an initial third-perspective speech representation; Dynamic time warping is performed on the initial first-perspective speech representation and the initial third-perspective speech representation to obtain the first-perspective audio representation and the third-perspective audio representation.
4. The multi-view video matching model training method according to claim 2, characterized in that: The performing cross-attention processing on the first-perspective video representation and the first-perspective audio representation to obtain the first video modality audiovisual representation includes: Using the first-view video representation as a first key vector and a first value vector, and using the first-view audio representation as a first query vector; Cross-attention processing is performed based on the first key vector, the first value vector, and the first query vector to obtain the first video modality audiovisual representation.
5. The multi-view video matching model training method according to claim 1, characterized in that: The step of inputting the video description text corresponding to the video sample data into a text encoder in a multi-view video matching model for processing to obtain a text modality audiovisual representation includes: Inputting the video description text into a large language model for audio description conversion to obtain an audio description text; Inputting the video description text into the text encoder for representation extraction to obtain a video description text representation; Inputting the audio description text into the text encoder for representation extraction to obtain an audio description text representation; Cross-attention processing is performed on the video description text representation and the audio description text representation to obtain the text modality audiovisual representation.
6. The multi-view video matching model training method according to claim 5, characterized in that: The step of inputting the video description text into a large language model for audio description conversion to obtain the audio description text comprises: Obtaining a video-audio conversion prompt, and inputting the video-audio conversion prompt into the large language model; Obtaining a video-audio conversion example description, and inputting the video-audio conversion example description into the large language model; The video description text is input into a large language model for audio description conversion to obtain an audio description text.
7. The multi-view video matching model training method according to claim 5, characterized in that: The performing cross-attention processing on the video description text representation and the audio description text representation to obtain the text modality audiovisual representation includes: Using the video description text representation as a second key vector and a second value vector, and using the audio description text representation as a second query vector; Cross-attention processing is performed based on the second key vector, the second value vector, and the second query vector to obtain the text modality audiovisual representation.
8. The multi-view video matching model training method according to claim 1, characterized in that: The calculating of the multi-view modality loss based on the first video modality audiovisual representation, the third video modality audiovisual representation, and the text modality audiovisual representation includes: Calculating a first-perspective loss based on the first video modality audiovisual representation and the text modality audiovisual representation; Calculating a third perspective loss based on the third video modality audiovisual representation and the text modality audiovisual representation; The multi-view modal loss is obtained by accumulating the first view loss and the third view loss.
9. The multi-view video matching model training method according to claim 8, characterized in that: The calculating the first perspective loss based on the first video modality audiovisual representation and the text modality audiovisual representation includes: Get temperature parameters; Calculating a first perspective similarity between the first video modality audiovisual representation and the text modality audiovisual representation, and performing exponential processing based on a ratio of the first perspective similarity to the temperature parameter to obtain a first perspective index similarity; Accumulating the first viewing angle index similarities corresponding to all the multi-view video sample data to obtain a first viewing angle index test value; Based on the ratio of the first viewing index similarity and the first viewing index test value, logarithmic processing is performed to obtain the first viewing angle sample loss; All first-view sample losses are averaged to obtain the first-view loss.
10. The multi-view video matching model training method according to claim 1, characterized in that: The updating of the multi-view video matching model based on the multi-view modality loss includes: The video encoder and the text encoder in the multi-view video matching model are updated based on the multi-view modality loss.
11. A multi-view video matching method, characterized in that: The method comprises: In response to a multi-view video matching task, obtaining view video data to be matched; Based on the multi-perspective video matching task, the perspective video data to be matched is input into the multi-perspective video matching model trained by the multi-perspective video matching model training method as described in claim 1 for video data matching processing to obtain target matching perspective video data, and the perspective video data to be matched and the target matching perspective video data are video data of different perspectives.
12. A multi-view video matching model training device, characterized in that: The device comprises: A sample data acquisition module is used to acquire multiple sets of video sample data and voice sample data and video description text corresponding to each of the video sample data, wherein the video sample data includes first-perspective video and third-perspective video, and the voice sample data includes first-perspective voice and third-perspective voice; a video representation calculation module, configured to input the first-view video and the first-view speech into a video and audio encoder in a multi-view video matching model for processing to obtain a first video modality audiovisual representation, and input the third-view video and the third-view speech into the video and audio encoder for processing to obtain a third video modality audiovisual representation; a text representation calculation module, configured to input the video description text corresponding to the video sample data into a text encoder in a multi-view video matching model for processing to obtain a text modality audiovisual representation; A model updating module is used to calculate the multi-view modality loss based on the first video modality audio-visual representation, the third video modality audio-visual representation and the text modality audio-visual representation, and to update the multi-view video matching model based on the multi-view modality loss to obtain a trained multi-view video matching model.
13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, it implements the multi-view video matching model training method according to any one of claims 1 to 10 or the multi-view video matching method according to claim 11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-view video matching model training method according to any one of claims 1 to 10 or the multi-view video matching method according to claim 11 is implemented.