Video analysis method, device and equipment based on multi-mode graph generation text large model

Through the video analysis method based on the multi-modal graphics and text big model, the target scenes in the short video and the image content description text is generated, and the problem of difficult to understand short videos without text information in the prior art is solved, and the efficiency and accuracy of video analysis are improved.

CN119992425AActive Publication Date: 2025-05-13BEIJING ZHIHUI XINGGUANG INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510201429.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The existing natural language processing methods are difficult to accurately understand short video contents of short videos without specific text information, resulting in many difficulties in subsequent processing such as public opinion analysis.

Method used

Using a video analysis method based on a multimodal graphical text model, the target scene in the video is identified through the object detection model, image description task instructions are created, and image content description text of the video is generated through the neural network encoder and decoder.

Benefits of technology

It improves the efficiency and timeliness of video analysis, prevents the generation of invalid text information, enhances the generalization ability of the large-scale picture and text model, and achieves more accurate video content understanding and public opinion analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992425A_ABST
    Figure CN119992425A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video analysis, and discloses a video analysis method, device and equipment based on a multi-modal graph-to-text large model. According to the method, a target image description task instruction is created by training a target detection model and taking a target scene as priori knowledge; the method focuses on key scenes in the video to better generate picture description needing to be focused, prevents invalid text information from being generated by combining with a target detection mode, improves video analysis efficiency and timeliness, improves generalization ability of picture-to-text big model training by adding matrix-level noise disturbance, and improves video analysis efficiency and timeliness. Meanwhile, mapping of image description task instructions and image features is increased by utilizing cross attention, so that the model can perform image description more accurately, and the model can better understand the instructions and better generate text description by inputting, fusing and aligning two modals and fusing a text sequence and an output matrix after the cross attention; and the video content understanding accuracy of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video analysis technology, and in particular to a video analysis method, device and equipment based on a multimodal graph-text model. Background Art

[0002] In the Internet era, short videos of self-media are becoming more and more popular. There is no specific text information in short videos. They are just a video uploaded to the Internet. The existing natural language processing methods are unable to analyze such videos. After extracting frames from short videos, optical character recognition (OCR) and automatic speech recognition (ASR) are performed to convert them into text information. There is no specific text content after conversion. Since it is impossible to accurately understand the video content of such videos, subsequent processing of such videos, such as public opinion analysis, is fraught with difficulties. Summary of the invention

[0003] In view of this, the present invention provides a video analysis method, device, and equipment based on a multimodal graph-text model to solve the problem that the existing natural language processing methods in the relevant technology are difficult to accurately understand the video content of videos without specific text information.

[0004] In a first aspect, the present invention provides a video analysis method based on a multimodal graph-based model, the method comprising:

[0005] Obtaining an image with scene annotations of the video analysis target to train the target detection model, obtain the identified target scene, and create an image description task instruction based on the target scene;

[0006] After word segmentation processing is performed on the image description task instruction, a text sequence is obtained through a layer of neural network encoder, and feature vectors are extracted from each image block after the image segmentation processing through a layer of residual network to obtain an image block sequence, and position information is added to the text sequence and the image block sequence;

[0007] fusing and aligning the text sequence and the image block sequence to obtain an input vector;

[0008] The matrix output by the N-layer neural network encoder of the input vector is subjected to matrix-level perturbation to add noise to obtain a first output matrix;

[0009] Performing cross attention processing on the first output matrix and the text sequence to obtain a second output matrix;

[0010] The first output matrix and the second output matrix are fused and input into a feedforward neural network, and then decoded by a neural network decoder to obtain a picture description text of the picture;

[0011] Based on the image description text and the preset video analysis target description annotation corresponding to the image, a model training is performed using a cross entropy loss function to obtain a trained image-to-text large model;

[0012] Each frame image of the video to be analyzed is input into the trained target detection model to obtain the identified target scene, and the frame image and image description task instructions corresponding to the target scene are input into the trained image-to-text model to obtain the image content description text of the video to be analyzed based on the target scene.

[0013] In an optional implementation, the text sequence and the image block sequence are fused and aligned using the following formula to obtain an input vector:

[0014] s n =λ w ew n *std(ew n )+λ m em n *std(em n )+el n

[0015] Among them, s n Represents the nth vector element of the input vector, ew n Represents the nth text element in a text sequence, em n Represents the nth image element of the image block sequence, el n Indicates the position information corresponding to the nth image element and the nth text element, λ w Represents the weight empirical parameter of the text sequence, λ m Represents the weighted empirical parameter of the image block sequence, std(ew n ) represents the standard deviation of the nth text element in the text sequence, std(em n ) represents the standard deviation of the nth image element of the image block sequence.

[0016] In an optional implementation, the matrix output by the input vector through the N-layer neural network encoder is subjected to matrix-level perturbation to add noise to obtain a first output matrix, including:

[0017] The matrix output by the input vector through the N-layer neural network encoder is [w1,w2,...,w S ], where S represents the number of parameter matrices in the model, and the formula for adding noise to each parameter matrix by matrix-level perturbation is as follows:

[0018]

[0019] Among them, w snew represents the nth parameter matrix in the first output matrix, Indicates from arrive The noise is uniformly distributed in the range, λ represents the hyperparameter controlling the noise intensity, std(w s ) indicates w S The standard deviation of .

[0020] In an optional implementation, performing cross attention processing on the first output matrix and the text sequence to obtain a second output matrix includes:

[0021] The text sequence is used as the query Q and the features of the image in the first output matrix are used as the key K and the value V, and the second output matrix is ​​obtained by the following formula:

[0022]

[0023] Z=score(Q,K,V)*V

[0024] Where Z represents the second output matrix, QK T The dot product of the query Q and the key K represents the similarity between the two sequences at different positions, d k represents the dimension of key K, softmax is the normalized exponential function, and score(Q,K,V) represents the attention weight of query Q for each key K.

[0025] In an optional implementation, the first output matrix and the second output matrix are fused using the following formula:

[0026] Znew n =λ1Z1 n *std(Z1 n )+(1-λ1)Z2 n *std(Z2 n )

[0027] Among them, Znew n Represents the nth element of the fused matrix, Z1 n represents the nth element of the first output matrix, Z2 n represents the nth element of the second output matrix, std(Z1 n ) indicates Z1 n The standard deviation, std(Z2 n ) represents Z2 n, λ1 represents the fusion weight empirical parameter of the first output matrix.

[0028] In an optional implementation, the target detection model is a yolov11 model.

[0029] In an optional embodiment, the method further includes:

[0030] Public opinion analysis is performed based on the image content description text of the target scene in the video to be analyzed to obtain a public opinion analysis result of the video to be analyzed.

[0031] In a second aspect, the present invention provides a video analysis device based on a multimodal graph-based model, comprising:

[0032] The first processing module is used to obtain a picture with a scene annotation of a video analysis target to train a target detection model, obtain an identified target scene, and create an image description task instruction based on the target scene;

[0033] The second processing module is used to obtain a text sequence by performing word segmentation processing on the image description task instruction through a layer of neural network encoder, extract feature vectors from each image block after the image segmentation processing through a layer of residual network to obtain an image block sequence, and add position information to the text sequence and the image block sequence;

[0034] A third processing module is used to fuse and align the text sequence and the image block sequence to obtain an input vector;

[0035] A fourth processing module is used to add noise to the matrix output by the input vector through the N-layer neural network encoder by means of matrix-level perturbation to obtain a first output matrix;

[0036] a fifth processing module, configured to perform cross attention processing on the first output matrix and the text sequence to obtain a second output matrix;

[0037] A sixth processing module, configured to fuse the first output matrix and the second output matrix, input the matrix into a feedforward neural network, and then decode the matrix using a neural network decoder to obtain a picture description text of the picture;

[0038] A seventh processing module is used to perform model training based on the picture description text and the preset video analysis target description annotation corresponding to the picture using a cross entropy loss function to obtain a trained picture-to-text large model;

[0039] The eighth processing module is used to input each frame image of the video to be analyzed into the trained target detection model to obtain the identified target scene, and input the frame image and image description task instructions corresponding to the target scene into the trained image-to-text model to obtain the image content description text of the video to be analyzed based on the target scene.

[0040] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method provided in the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0041] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method provided in the first aspect or any corresponding embodiment thereof.

[0042] Beneficial effects:

[0043] The present invention trains a target detection model to identify the target scene of the video to be analyzed, and uses the target scene as prior knowledge to create a target image description task instruction, thereby focusing on the key scenes in the video to better generate picture descriptions that need to be focused on. By combining the target detection method, invalid text information is prevented from being generated, the efficiency and timeliness of video analysis are improved, and the generalization ability of the image-to-text large model training is improved by adding matrix-level noise disturbance. At the same time, cross-attention is used to increase the mapping of image description task instructions and picture features, so that the model can more accurately describe the image. In addition, by fusing and aligning the two modal text sequences and the image block sequence, and fusing the text sequence with the output matrix after cross-attention, the model can better understand the instructions and better generate text descriptions, further improving the accuracy of the image-to-text large model in understanding the video content. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0045] Figure 1 is a flow chart of a video analysis method based on a multimodal graph-text model according to an embodiment of the present invention;

[0046] Figure 2is a workflow diagram of a video analysis system based on a multimodal graph-based model according to an embodiment of the present invention;

[0047] Figure 3 is a structural schematic diagram of a video analysis device based on a multimodal graph-text model according to an embodiment of the present invention;

[0048] Figure 4 is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0050] In the Internet era, short videos of self-media are becoming more and more popular. There is no specific text information in short videos. They are just videos uploaded to the Internet. The existing natural language processing methods cannot analyze and judge such videos. Moreover, after extracting frames from short videos, OCR and ASR are performed to convert them into text information. After conversion, there is no specific text content. Such short videos are becoming more and more common in the Internet era. Therefore, the present invention proposes a video analysis method based on multimodal graph-based text, which mainly extracts key frames from videos, performs scene recognition on key frames, identifies target scenes, treats target scenes as prior knowledge, and inputs them into a large graph-based text model. The description of the picture is generated with the target scene as the theme, so as to realize the subsequent application of public opinion judgment and analysis of business scenarios that are focused on in short videos.

[0051] According to an embodiment of the present invention, an embodiment of a video analysis method based on a multimodal graph-based model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in an order different from that shown here.

[0052] Based on the above problems, in this embodiment, a video analysis method based on a multimodal image-based model is provided, which is applied to computer devices such as CPU, single-chip microcomputer, etc. Figure 1 Flow chart of a video analysis method based on a multimodal graph-based text model according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0053] Step S101, obtain a picture with a scene annotation of a video analysis target to train a target detection model, obtain an identified target scene, and create an image description task instruction based on the target scene.

[0054] Among them, the video analysis target can be set according to business needs, such as: fire, police, police car, fire truck, car accident, fight, etc. In an embodiment of the present invention, the target detection model is a yolov11 model. By collecting and organizing the key scenes of the above-mentioned video analysis targets required by the business, such as fire scenes, by annotating a large number of pictures, and using the YOLO (You Only Look Once) target detection algorithm to train the yolov11 model, the above-mentioned key target scenes can be accurately identified and passed to the downstream for use. The purpose is to provide the downstream model with instructions on what target content needs to be described in the image, which is also prior knowledge.

[0055] Furthermore, the image description task instruction is to input the prompt word prompt, which needs to instruct the image-generated text model of the specific task type currently being trained, for example: prompt: Image description task: xxxxxx. Secondly, based on upstream target detection, the identified target scene is described in the picture, such as prompt: Image description task: Please describe the scene content related to [fire] in the picture. If there is no [fire], describe the overall picture content. Finally, if there are multiple target scenes in the picture, they can be input together to the image-generated text model for description, such as prompt: Image description task: Please describe the scene content related to [fire] [fire truck] in the picture. If there is no [fire] [fire truck], describe the overall picture content. The overall image description task instruction format is: Task type: Description words + [target scene].

[0056] Step S102, after word segmentation processing is performed on the image description task instruction, a text sequence is obtained through a layer of neural network encoder, and feature vectors are extracted from each image block after the image is segmented through a layer of residual network to obtain an image block sequence, and position information is added to the text sequence and the image block sequence.

[0057] Specifically, for the text input part of the image description task instruction, byte-pair encoding (BPE) is used for subword level segmentation. BPE reduces the size of the vocabulary by gradually merging the most common characters or character sequences, so as to more efficiently process and represent text data. For example, the word "example" will be segmented into subwords such as "ex", "am", and "ple". For example, "hello" will be segmented into "you", "good", and other words; Special tags: In addition to the regular subword tags, some special tags are added, such as "<|begin_of_text|>" and "<|end_of_text|>", etc., which are used to indicate the beginning and end of the text. The input text is converted into word vectors through a layer of transformer encoder, and the input is aligned to obtain the text sequence text encoder: (ew1, ew2,...ew n ).

[0058] Furthermore, for the image input part, the image is directly divided into multiple image patches, and then a linear projection and image segmentation are performed, which can effectively reduce the sequence length of the image representation. For example, an image with a resolution of 256×256 is represented as an image sequence with a length of 16×16. In this way, the image can be similar to a text sequence, and sequence alignment and contextual semantic analysis can also be performed. For each patch block, a feature vector is extracted through a layer of resnet network, and finally the patch sequence is aligned to generate an image block sequence image encoder: (em1, em2, ...., em n Then, for the above text sequence and image block sequence, we need to add position information position encoder: (el1, el2, ..., el n ).

[0059] Step S103: fusing and aligning the text sequence and the image block sequence to obtain an input vector.

[0060] Specifically, the text sequence and the image block sequence are fused and aligned using the following formula to obtain the input vector:

[0061] s n =λ w ew n *std(ew n )+λ m em n *std(em n )+el n

[0062] Among them, s nRepresents the nth vector element of the input vector, ew n Represents the nth text element in a text sequence, em n Represents the nth image element of the image block sequence, el n Indicates the position information corresponding to the nth image element and the nth text element, λ w Represents the weight empirical parameter of the text sequence, λ m Represents the weighted empirical parameter of the image block sequence, std(ew n ) represents the standard deviation of the nth text element in the text sequence, std(em n ) represents the standard deviation of the nth image element of the image block sequence.

[0063] The final input vector can be expressed as sum:(s1,s2...s n ), thereby adding weights and standard deviations in the above fusion process, on the one hand, adjusting the input knowledge ratio, and on the other hand, standardizing the input vector, providing a standardized data basis for the subsequent training and prediction of the graph-to-text model.

[0064] Step S104, adding noise to the matrix output by the input vector after passing through the N-layer neural network encoder by means of matrix-level perturbation to obtain a first output matrix.

[0065] Specifically, in order to improve the generalization ability of the large image-to-text model and ensure better output effects, the matrix-level noise perturbation capability is introduced here to add noise to the matrix output after the input text encoder and image encoder are fused and passed through N layers of transformer encoder. The details are as follows:

[0066] The noise is increased by matrix-wise perturbing method. The matrix output by the input vector through the N-layer neural network encoder is [w1,w2,…,w S ], where S represents the number of parameter matrices in the model, and the formula for adding noise to each parameter matrix by matrix-level perturbation is as follows:

[0067]

[0068] Among them, w snew represents the nth parameter matrix in the first output matrix, Indicates from arrive The noise is uniformly distributed in the range, λ represents the hyperparameter controlling the noise intensity, std(w s ) indicates w S The standard deviation of .

[0069] Step S105, performing cross attention processing on the first output matrix and the text sequence to obtain a second output matrix.

[0070] Step S106, the first output matrix and the second output matrix are fused and input into a feedforward neural network, and then decoded by a neural network decoder to obtain a picture description text of the picture.

[0071] Specifically, the first output matrix is ​​subjected to a cross attention with the above text encoder. The main function of the cross attention is to capture the dependency between the two inputs. The cross attention mechanism uses the text instruction description as the query and the features in the image as the key and value. This can help the image-to-text model generate more accurate image descriptions. The cross attention mechanism is based on the calculation of query, key, and value. The specific process is as follows:

[0072] The text sequence is used as the query Q and the features of the image in the first output matrix are used as the key K and value V, and the second output matrix is ​​obtained by the following formula:

[0073]

[0074] Z=score(Q,K,V)*V

[0075] Where Z represents the second output matrix, QK T The dot product of the query Q and the key K represents the similarity between the two sequences at different positions, d k represents the dimension of key K, softmax is the normalized exponential function, and score(Q,K,V) represents the attention weight of query Q for each key K.

[0076] Thus, by calculating the similarity between the query and the key: dot product the query Q with the key K to get the correlation score between the two inputs; then convert these similarities into probability distributions through the softmax function, representing the attention weight of the query on each key; finally, apply these attention weights to the value V, and finally get the output vector Z, which is the second output matrix mentioned above. This is equivalent to extracting the information of interest from the value sequence, and obtaining the second output matrix can be input into the next Feedforward network.

[0077] Then, the two encoder output matrices are fused and input to the next Feed forward, which is then output to the decoder for decoding. Specifically, the first output matrix and the second output matrix are fused using the following formula:

[0078] Znew n =λ1Z1n *std(Z1 n )+(1-λ1)Z2 n *std(Z2 n )

[0079] Among them, Znew n Represents the nth element of the fused matrix, Z1 n represents the nth element of the first output matrix, Z2 n represents the nth element of the second output matrix, std(Z1 n ) indicates Z1 n The standard deviation, std(Z2 n ) represents Z2 n , λ1 represents the fusion weight empirical parameter of the first output matrix.

[0080] Step S107, based on the picture description text and the preset video analysis target description annotation corresponding to the picture, model training is performed using a cross entropy loss function to obtain a trained picture-to-text large model.

[0081] Among them, the loss function of the graph-to-text model uses cross-entropy loss (Cross-Entropy Loss), and the specific formula is:

[0082]

[0083] Among them, H(p,q) is the value of the cross entropy loss function, p(x) is the true probability distribution (in the large language model, the true probability distribution of the next word can be regarded as a distribution in which all other words are 0 except this word is 1), and q(x) is the probability distribution predicted by the large graph-based model.

[0084] Specifically, the YOLO target detection model is trained by annotating the target scene image data. Secondly, the images and their corresponding descriptions, as well as the input prompts, are organized to build and organize the training data set. The hyperparameters of the large image-to-text model can be set normally, and the image description task training is performed. The model with the highest accuracy and precision values ​​for the three tasks is selected and saved as the optimal model.

[0085] Step S108, input each frame image of the video to be analyzed into the trained target detection model to obtain the identified target scene, input the frame image and image description task instructions corresponding to the target scene into the trained image-to-text model to obtain the image content description text of the video to be analyzed based on the target scene.

[0086] Specifically, by loading the optimal model trained in the above step S107, YOLO scene recognition is performed on the new image to identify the scene type, and then a command prompt is constructed, and the image and prompt are input into the trained image-to-text model to generate a description of the image content based on the target scene.

[0087] By executing the above steps, the present invention trains the target detection model to identify the target scene of the video to be analyzed, and uses the target scene as prior knowledge to create the target image description task instruction, so as to focus on the key scenes in the video to better generate the picture description that needs to be focused on, and prevents the generation of invalid text information by combining the target detection method, thereby improving the efficiency and timeliness of video analysis, and by adding matrix-level noise disturbance, the generalization ability of the large image-to-text model training is improved, and at the same time, the mapping of image description task instructions and image features is increased by cross-attention, so that the model can more accurately describe the image, and in addition, by fusing and aligning the two modal text sequences and the image block sequence, and fusing the text sequence with the output matrix after cross-attention, the model can better understand the instructions and better generate text descriptions, further improving the accuracy of the large image-to-text model in understanding the video content.

[0088] The present invention innovatively generates text descriptions based on key scenes in focused pictures by constructing and combining yolo target detection and adjusting the network structure to prevent invalid information from being generated. It realizes the generation of text descriptions based on the focus range by combining target detection; improves the generalization ability of training by fusion method of two modal inputs and adding matrix-level noise disturbance; increases the mapping of input prompt and picture features by cross-attention, so that the model can describe the image more accurately. By fusing the input prompt, the output matrix after the picture encoder and the cross-attention, the model can better understand the instructions and better generate text descriptions.

[0089] In practical applications, the video analysis method based on the multimodal image-to-text model provided by the embodiment of the present invention further includes the following steps:

[0090] Step S109, performing public opinion analysis based on the image content description text of the target scene in the video to be analyzed, and obtaining a public opinion analysis result of the video to be analyzed.

[0091] Among them, the specific analysis scheme of using image content description text to perform public opinion analysis can be implemented using the public opinion analysis method of the existing technology, which will not be described in detail here.

[0092] Specifically, taking the short video public opinion analysis application as an example, the existing implementation solutions mainly fall into the following categories:

[0093] (1) Analysis of short video subscriptions based on keywords.

[0094] For short videos with text information, including those converted from OCR and ASR to text information, the analysis of short videos and the judgment of public opinion can be completed through key point subscription retrieval. This method cannot recall short videos without text information, and sensitive short videos often have no specific information.

[0095] (2) Short video analysis based on image search.

[0096] By building a picture library for existing hot events, we can retrieve videos similar to the pictures in the picture library by searching pictures, so as to recall the videos and conduct public opinion analysis on the videos. We can only recall similar short videos if we have pictures first, and we can only find them after the fact, which is not in line with advance discovery.

[0097] (3) Scene recognition short video analysis.

[0098] Through scene recognition technology, we can identify key scenes in short videos and label them. We can recall short video data through scene labels and conduct public opinion analysis on the videos. However, we only label key scenes, but lack a specific description of key scenes, whether they are negative scenes or general scenes.

[0099] Based on the problems existing in the above-mentioned prior art, the video analysis solution based on the multimodal image-based text model provided by the present invention can be used to build a video analysis system based on the multimodal image-based text model. The entire system includes four major modules: scene recognition module, image-text vocabulary normalization module, image-based text model training module, and image-based text model prediction module. The specific working process of the above-mentioned steps S101 to S108 is performed through these four modules. The overall workflow of the system is shown in the figure below. Figure 2 As shown. Thus, the scene labels of key focus are identified through scene recognition, and the labels are input into the large model of image-based text to generate the theme description of the key scenes of the short video key frames. Then, the public opinion is analyzed and judged by the generated description information. Compared with the existing short video public opinion analysis solutions, the technical solution provided by the present invention has the following advantages:

[0100] 1. Focus on scene recognition.

[0101] There is too much short video data, and less than 10% of it is truly sensitive information with public opinion value. The massive amount of worthless data has brought huge challenges to public opinion analysis. Through scene recognition, we can identify the scene data that needs to be focused on to improve the timeliness and efficiency of subsequent public opinion analysis.

[0102] 2. Focus on the core topic and generate descriptive information.

[0103] In the case of complex scenes, it is difficult for the generated image description to focus on the subject that needs attention, and the generated image description may even be irrelevant to the subject that needs attention. In order to better generate image descriptions that need to be focused on, the scene information obtained by scene recognition is input into the model to generate scene-based image descriptions.

[0104] 3. Convert into text information to facilitate public opinion analysis.

[0105] The key frames of short videos are converted into text information through scene recognition and image-to-text, and natural processing is used to process the text information, so as to fully grasp the public opinion information of the short video.

[0106] Other existing solutions are basically based on keyword subscription and scene recognition, and are based on existing text information or identifying scene tags. However, for short videos with no text information, no OCR information, and no ASR information, the existing solutions cannot understand the video content, that is, they cannot make public opinion judgments. The technical solution provided by the present invention can solve the problem that short videos with no text information, no OCR information, and no ASR information cannot generate accurate and understandable text information. Through scene recognition and image description generation based on scene information, a text that best represents the semantics of the short video is generated, so that the public opinion of the short video can be grasped by analyzing the text generated by accurate judgment.

[0107] In the embodiment of the present invention, a video analysis device based on a multimodal graph model is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0108] This embodiment provides a video analysis device based on a multi-modal graph-based model. Figure 3 As shown, the device comprises:

[0109] The first processing module 301 is used to obtain a picture with a scene annotation of a video analysis target to train a target detection model, obtain an identified target scene, and create an image description task instruction based on the target scene;

[0110] The second processing module 302 is used to obtain a text sequence by performing word segmentation processing on the image description task instruction through a layer of neural network encoder, extract feature vectors from each image block after the image segmentation processing through a layer of residual network to obtain an image block sequence, and add position information to the text sequence and the image block sequence;

[0111] The third processing module 303 is used to fuse and align the text sequence and the image block sequence to obtain an input vector;

[0112] The fourth processing module 304 is used to add noise to the matrix output by the N-layer neural network encoder through matrix-level perturbation to obtain a first output matrix;

[0113] A fifth processing module 305 is used to perform cross attention processing on the first output matrix and the text sequence to obtain a second output matrix;

[0114] A sixth processing module 306 is used to fuse the first output matrix and the second output matrix, input them into a feed-forward neural network, and then decode them by a neural network decoder to obtain a picture description text of the picture;

[0115] The seventh processing module 307 is used to perform model training based on the preset video analysis target description annotation corresponding to the picture description text and the picture using a cross entropy loss function to obtain a trained picture-to-text large model;

[0116] The eighth processing module 308 is used to input each frame image of the video to be analyzed into the trained target detection model to obtain the identified target scene, and input the frame image and image description task instructions corresponding to the target scene into the trained image-to-text model to obtain the image content description text of the video to be analyzed based on the target scene.

[0117] The video analysis device based on the multimodal graph-to-text large model provided by the embodiment of the present invention trains the target detection model to identify the target scene of the video to be analyzed, and uses the target scene as prior knowledge to create the target image description task instruction, so as to focus on the key scenes in the video and better generate the picture description that needs to be focused on. By combining the target detection method, it is prevented that all invalid text information is generated, thereby improving the efficiency and timeliness of video analysis, and by adding matrix-level noise disturbance, the generalization ability of the graph-to-text large model training is improved. At the same time, cross-attention is used to increase the mapping of image description task instructions and picture features, so that the model can more accurately describe the image. In addition, by fusing and aligning the two modal text sequences and the image block sequence, and fusing the text sequence with the output matrix after cross-attention, the model can better understand the instructions and better generate text descriptions, further improving the accuracy of the graph-to-text large model in understanding the video content.

[0118] The video analysis device based on the multimodal graph-based model in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0119] The further functional description of each of the above modules and units is the same as that of the above corresponding method embodiments and will not be repeated here.

[0120] See also Figure 4 , Figure 4 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 4 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.

[0121] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0122] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0123] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0124] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0125] The computer device further comprises a communication interface 30 for the control unit to communicate with other devices or a communication network.

[0126] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0127] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A video analysis method based on a multimodal graph-text model, characterized in that: The method comprises: Obtaining an image with scene annotations of the video analysis target to train the target detection model, obtain the identified target scene, and create an image description task instruction based on the target scene; After word segmentation processing is performed on the image description task instruction, a text sequence is obtained through a layer of neural network encoder, and feature vectors are extracted from each image block after the image segmentation processing through a layer of residual network to obtain an image block sequence, and position information is added to the text sequence and the image block sequence; fusing and aligning the text sequence and the image block sequence to obtain an input vector; The matrix output by the N-layer neural network encoder of the input vector is subjected to matrix-level perturbation to add noise to obtain a first output matrix; Performing cross attention processing on the first output matrix and the text sequence to obtain a second output matrix; The first output matrix and the second output matrix are fused and input into a feedforward neural network, and then decoded by a neural network decoder to obtain a picture description text of the picture; Based on the image description text and the preset video analysis target description annotation corresponding to the image, a model training is performed using a cross entropy loss function to obtain a trained image-to-text large model; Each frame image of the video to be analyzed is input into the trained target detection model to obtain the identified target scene, and the frame image and image description task instructions corresponding to the target scene are input into the trained image-to-text model to obtain the image content description text of the video to be analyzed based on the target scene.

2. The method according to claim 1, characterized in that The text sequence and the image block sequence are fused and aligned using the following formula to obtain an input vector: s n =λ w oh n *std(ew n )+λ m yes n *std(em n )+el n Among them, s n Represents the nth vector element of the input vector, ew n Represents the nth text element in a text sequence, em n Represents the nth image element of the image block sequence, el n Indicates the position information corresponding to the nth image element and the nth text element, λ w Represents the weight empirical parameter of the text sequence, λ m Represents the weighted empirical parameter of the image block sequence, std(ew n ) represents the standard deviation of the nth text element in the text sequence, std(em n ) represents the standard deviation of the nth image element of the image block sequence.

3. The method according to claim 1, characterized in that The step of adding noise to the matrix output by the N-layer neural network encoder of the input vector by matrix-level perturbation to obtain a first output matrix includes: The matrix output by the input vector through the N-layer neural network encoder is [w1,w2,...,w S ], where S represents the number of parameter matrices in the model, and the formula for adding noise to each parameter matrix by matrix-level perturbation is as follows: Among them, w snew represents the nth parameter matrix in the first output matrix, Indicates from arrive The noise is uniformly distributed in the range, λ represents the hyperparameter controlling the noise intensity, std(w s ) indicates w S The standard deviation of .

4. The method according to claim 1, characterized in that: The cross-attention processing is performed on the first output matrix and the text sequence to obtain a second output matrix, including: The text sequence is used as the query Q and the features of the image in the first output matrix are used as the key K and the value V, and the second output matrix is ​​obtained by the following formula: Z=score(Q,K,V)*V Where Z represents the second output matrix, QK T The dot product of the query Q and the key K represents the similarity between the two sequences at different positions, d k represents the dimension of key K, softmax is the normalized exponential function, and score(Q,K,V) represents the attention weight of query Q for each key K.

5. The method according to claim 1, characterized in that The first output matrix and the second output matrix are fused by the following formula: Znew n =λ1Z1 n *std(Z1 n )+(1-λ1)Z2 n *std(Z2 n ) Among them, Znew n Represents the nth element of the fused matrix, Z1 n represents the nth element of the first output matrix, Z2 n represents the nth element of the second output matrix, std(Z1 n ) indicates Z1 n The standard deviation, std(Z2 n ) represents Z2 n , λ1 represents the fusion weight empirical parameter of the first output matrix.

6. The method according to claim 1, characterized in that The target detection model is the yolov11 model.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Public opinion analysis is performed based on the image content description text of the target scene in the video to be analyzed to obtain a public opinion analysis result of the video to be analyzed.

8. A video analysis device based on a multimodal graph-text model, characterized in that: include: The first processing module is used to obtain a picture with a scene annotation of a video analysis target to train a target detection model, obtain an identified target scene, and create an image description task instruction based on the target scene; The second processing module is used to obtain a text sequence by performing word segmentation processing on the image description task instruction through a layer of neural network encoder, extract feature vectors from each image block after the image segmentation processing through a layer of residual network to obtain an image block sequence, and add position information to the text sequence and the image block sequence; A third processing module is used to fuse and align the text sequence and the image block sequence to obtain an input vector; A fourth processing module is used to add noise to the matrix output by the N-layer neural network encoder of the input vector by matrix-level perturbation to obtain a first output matrix; a fifth processing module, configured to perform cross attention processing on the first output matrix and the text sequence to obtain a second output matrix; A sixth processing module, configured to fuse the first output matrix and the second output matrix, input the matrix into a feedforward neural network, and then decode the matrix using a neural network decoder to obtain a picture description text of the picture; A seventh processing module is used to perform model training based on the picture description text and the preset video analysis target description annotation corresponding to the picture using a cross entropy loss function to obtain a trained picture-to-text large model; The eighth processing module is used to input each frame image of the video to be analyzed into the trained target detection model to obtain the identified target scene, and input the frame image and image description task instructions corresponding to the target scene into the trained image-to-text model to obtain the image content description text of the video to be analyzed based on the target scene.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-model comprehensive remote sensing image scene description method

    CN113610025A

  • Terminal suitability judgment method and device, electronic equipment and storage medium

    CN115734029A

  • Traffic scene generation type image description method

    CN117173450A

  • Multi-modal large model scene understanding method based on scene graph enhancement

    CN119418339A

  • Text-to-image generation method, apparatus and device, and storage medium

    WO2025010950A1

Cited By

  • Method and device for analyzing video and electronic equipment

    CN121074744A