Text evaluation method and device and storage medium

Through the multimodal evaluation model, the problem of low accuracy in comments or barrage recognition in the prior art is solved, and efficient and accurate text evaluation and safe risk control are achieved.

CN120257968APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410008787.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the identification accuracy of texts such as comments or barrage is low, making it difficult to effectively identify illegal content.

Method used

By obtaining the image features and text features of the text to be evaluated for fusion, using a multimodal evaluation model for text evaluation, combining pre-joint training of the image encoder and the multimodal evaluation model, the real-time computing needs are reduced and evaluation accuracy and efficiency are improved.

Benefits of technology

It realizes efficient and accurate identification of comments or barrage texts, reduces calculation time and improves the timeliness and reliability of safe risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257968A_ABST
    Figure CN120257968A_ABST
Patent Text Reader

Abstract

The invention discloses a text evaluation method and device and a storage medium, which can be applied to various scenes such as cloud technology, artificial intelligence, smart traffic, video playing and the like. Obtaining image features of a target video picture matched with the to-be-evaluated text from a feature cache set; extracting text features of the to-be-evaluated text through a multi-modal evaluation model, and fusing the text features and the image features to obtain fused features; according to the fusion features, performing text evaluation through the multi-modal evaluation model to obtain a text evaluation result; wherein the image features in the feature cache set are obtained by encoding video pictures of the target video through an image encoder, and the image encoder and the multi-modal evaluation model are obtained through pre-joint training. And the accuracy and efficiency of text evaluation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and particularly to a text evaluation method, apparatus, and storage medium, where the storage medium is a computer-readable storage medium. Background Art

[0002] With the development of Internet technologies, the applications of the Internet have become more and more extensive. For example, live broadcasts or short videos can be transmitted or released through the Internet. In video stream scenarios such as live broadcasts or short videos, there is usually a function of posting comments or bullet screens. For security risk control requirements, it is usually necessary to identify the comments or bullet screens to be posted to determine whether there are any violations.

[0003] Currently, the existing methods for identifying texts such as comments or bullet screens generally perform semantic analysis on the texts of comments or bullet screens, and determine whether there are any violations based on the analysis results, so as to take control measures as needed. However, due to the influence of factors such as ambiguity in some texts, the accuracy of the identification results of this identification method is relatively low. Summary of the Invention

[0004] Embodiments of this application provide a text evaluation method, apparatus, and storage medium, which can improve the accuracy and efficiency of text evaluation.

[0005] To solve the above technical problems, the embodiments of this application provide the following technical solutions:

[0006] Embodiments of this application provide a text evaluation method, including:

[0007] Obtaining the text to be evaluated corresponding to the target video that needs to be text-evaluated, and obtaining the image features of the target video frame matching the text to be evaluated from the feature cache set;

[0008] Extracting the text features of the text to be evaluated through a multimodal evaluation model, and fusing the text features and the image features to obtain fused features;

[0009] Performing text evaluation through the multimodal evaluation model according to the fused features to obtain a text evaluation result;

[0010] Among them, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multimodal evaluation model are jointly trained in advance.

[0011] According to one aspect of this application, there is also provided a text evaluation apparatus, including:

[0012] An acquisition unit, configured to acquire the text to be evaluated corresponding to the target video that needs to be textually evaluated, and acquire the image features of the target video frame matching the text to be evaluated from the feature cache set;

[0013] A fusion unit, configured to extract the text features of the text to be evaluated through a multi-modal evaluation model, and fuse the text features and the image features to obtain fused features;

[0014] An evaluation unit, configured to perform text evaluation through the multi-modal evaluation model according to the fused features to obtain a text evaluation result;

[0015] Wherein, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multi-modal evaluation model are jointly trained in advance.

[0016] In some embodiments, the text evaluation device further includes:

[0017] A sampling unit, configured to perform image sparse sampling on the target video at a preset sampling frequency to obtain sampled video frames;

[0018] An encoding unit, configured to encode the sampled video frames through the image encoder to obtain image features;

[0019] A storage unit, configured to store the image features in the feature cache set.

[0020] In some embodiments, the sampling unit is specifically configured to:

[0021] Determine the video scene of the target video, and determine the preset sampling frequency of the target video according to the video scene;

[0022] Perform image sparse sampling on the target video at the preset sampling frequency to obtain sampled video frames.

[0023] In some embodiments, the storage unit is specifically configured to:

[0024] Acquire the video identifier of the target video;

[0025] Store the video identifier and the image features in the feature cache set in the form of key-value pairs.

[0026] In some embodiments, there are multiple sampled video frames, and the encoding unit is specifically configured to: encode each sampled video frame through the image encoder to obtain multiple image features, and each image feature corresponds to the image feature of each sampled video frame;

[0027] The storage unit is specifically configured to: when historical image features are stored in the feature cache set, store the video identifier and the multiple image features in the form of key-value pairs, and replace the historical image features in the feature cache set.

[0028] In some embodiments, the text evaluation device further includes:

[0029] A sample acquisition unit, configured to acquire training samples, where the training samples include text evaluation labels, sample video frames, and sample evaluation texts;

[0030] A sample encoding unit, configured to encode the sample video frames through an image encoder to obtain sample image features;

[0031] A sample fusion unit, configured to extract sample text features of the sample evaluation text through a multimodal evaluation model, and fuse the sample image features and the sample text features to obtain sample fusion features;

[0032] A prediction unit, configured to perform text evaluation prediction through the multimodal evaluation model according to the sample fusion features to obtain a prediction result;

[0033] A calculation unit, configured to calculate the difference between the text evaluation label and the prediction result to obtain a loss value;

[0034] An adjustment unit, configured to adjust the parameters of the image encoder and the multimodal evaluation model according to the loss value until a preset stop condition is met.

[0035] In some embodiments, the text evaluation device further includes:

[0036] An execution unit, configured to determine a corresponding security risk control policy according to the text evaluation result, and perform corresponding operations according to the security risk control policy.

[0037] In some embodiments, the execution unit is specifically configured to:

[0038] When the text evaluation result indicates that the text to be evaluated is abnormal, determine that the corresponding security risk control policy is not to output the text to be evaluated, and perform the operation of not outputting the text to be evaluated;

[0039] When the text evaluation result indicates that the text to be evaluated is normal, determine that the corresponding security risk control policy is to output the text to be evaluated, and perform the operation of outputting the text to be evaluated.

[0040] According to an aspect of the present application, there is also provided a computer device, including a processor and a memory. A computer program is stored in the memory. When the processor calls the computer program in the memory, it executes any one of the text evaluation methods provided by the embodiments of the present application.

[0041] According to an aspect of the present application, there is also provided a storage medium for storing a computer program, which is loaded by a processor to execute any one of the text evaluation methods provided by the embodiments of the present application.

[0042] According to an aspect of the present application, there is also provided a computer program product, including a computer program, which is loaded by a processor to execute any one of the text evaluation methods provided by the embodiments of the present application.

[0043] The present application can obtain the text to be evaluated corresponding to the target video that needs to be text-evaluated, and obtain the image features of the target video frame matching the text to be evaluated from the feature cache set. Then, the text features of the text to be evaluated are extracted through a multimodal evaluation model, and the text features and image features are fused to obtain fused features. At this time, text evaluation can be performed through the multimodal evaluation model according to the fused features to obtain a text evaluation result. Among them, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multimodal evaluation model are jointly trained in advance. This solution can pre-encode the video frames of the target video to obtain image features and store them in the feature cache set, so that the image features of the target video frame matching the text to be evaluated can be quickly obtained from the feature cache set, without the need to frequently pull the video stream for calculation in real time, reducing the calculation time-consuming, and accurate multimodal text evaluation can be performed based on the fused features obtained by fusing the text features of the text to be evaluated and the image features of the matching target video frame, improving the accuracy and efficiency of text evaluation. Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic diagram of the scenario to which the text evaluation method provided by the embodiments of the present application is applied;

[0046] Figure 2 It is a schematic flowchart of the text evaluation method provided by the embodiments of the present application;

[0047] Figure 3 It is another schematic flowchart of the text evaluation method provided by an embodiment of the present application;

[0048] Figure 4 It is a schematic flowchart of the overall architecture of the text evaluation provided by an embodiment of the present application;

[0049] Figure 5 It is a schematic diagram of the text evaluation device provided by an embodiment of the present application;

[0050] Figure 6 It is a schematic diagram of the structure of the computer device provided by an embodiment of the present application. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0052] In the following description of the present application, the term "some embodiments" is involved, which describes a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0053] In the following description of the present application, the terms "first / second", etc. are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0055] An embodiment of the present application provides a text evaluation method, device and storage medium.

[0056] Please refer to Figure 1 , Figure 1FIG. 0 is a schematic diagram of the scenario to which the text evaluation method provided by the embodiments of the present application is applied. The text evaluation method can be applied to a text evaluation system, and the text evaluation system can include a server 10 and a terminal 20, etc. The server 10 can be integrated with the text evaluation device provided by the present application, or the terminal 20 can be integrated with the text evaluation device provided by the present application. Hereinafter, the case where the text evaluation device is integrated in the server 10 will be described in detail. Among them, the server 10 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto.

[0057] The server 10 and the terminal 20 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this. The terminal 20 can be a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc.

[0058] Among them, the server 10 can be used to obtain a target video such as a live video or a short video played by the terminal 20, and obtain the text to be evaluated corresponding to the target video that needs to be text-evaluated, such as text such as bullet screens or comments, and obtain the image features of the target video frame matching the text to be evaluated from the feature cache set. The image features in the feature cache set are pre-obtained by encoding the video frames of the target video through an image encoder. Then, the server 10 can extract the text features of the text to be evaluated through a multimodal evaluation model, and fuse the text features and the image features to obtain a fused feature. At this time, the server 10 can perform text evaluation through the multimodal evaluation model according to the fused feature to obtain a text evaluation result, and the text evaluation result can include information such as whether the text to be evaluated meets the release conditions. After obtaining the text evaluation result, the server 10 can send the text evaluation result to the terminal 20, or the server 10 can determine the corresponding security risk control strategy according to the text evaluation result, and control the terminal 20 to perform corresponding operations according to the security risk control strategy. Since the image features can be obtained by pre-encoding the video frames of the target video and stored in the feature cache set, the image features of the target video frame matching the text to be evaluated can be quickly obtained from the feature cache set, without the need to frequently pull the video stream for calculation in real time, reducing the calculation time-consuming, and accurate multimodal text evaluation can be performed based on the fused feature obtained by fusing the text features of the text to be evaluated and the image features of the matching target video frame, improving the accuracy and efficiency of text evaluation.

[0059] It should be noted that Figure 1 The schematic diagram of the application scenario of the text evaluation method shown is only an example. The application and scenario of the text evaluation method described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the application of the text evaluation method and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0060] In the present application, an artificial intelligence learning method can be adopted to achieve multi-modal text evaluation, improving the accuracy and efficiency of text evaluation. It should be noted that artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0061] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include machine learning (ML) technology. Among them, deep learning (DL) is a new research direction in machine learning. It is introduced into machine learning to make it closer to the original goal, that is, artificial intelligence. Currently, deep learning is mainly applied in fields such as machine vision, speech processing technology, and natural language processing.

[0062] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as intentional degree theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pretrained models are the latest development results of deep learning, which integrate the above technologies.

[0063] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0064] In this embodiment, the description will be made from the perspective of a text evaluation device, which can be specifically integrated in computer devices such as servers or terminals.

[0065] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a text evaluation method provided by an embodiment of the present application. The text evaluation method may include:

[0066] S101. Obtain the text to be evaluated corresponding to the target video that needs to be text-evaluated, and obtain the image features of the target video frame matching the text to be evaluated from the feature cache set.

[0067] Among them, the target video may include videos such as live broadcasts, short videos, movies, and TV dramas. The text to be evaluated may include texts such as comments or bullet screens for the target video, and may also include texts obtained by converting the speech to be published. For security risk control requirements, during the playback of the video, texts such as comments or bullet screens to be published can be evaluated to determine whether there are violations. To achieve the accuracy of text evaluation, multimodal text evaluation can be performed by combining the text features of the text to be evaluated and the image features of the target video frame in the target video.

[0068] Specifically, the text to be evaluated corresponding to the target video that needs to be text-evaluated can be obtained. For example, during the playback of the target video by the playback terminal, texts such as comments or bullet screens for the target video can be received to obtain the text to be evaluated. Another example is that texts such as comments or bullet screens for the target video can be obtained from the text cache library to obtain the text to be evaluated. Of course, the text to be evaluated can also be obtained through other means, which is not limited here.

[0069] Moreover, the image features of the target video frame matching the text to be evaluated can be obtained from the feature cache set. The feature cache set is used to store one or more image features obtained by asynchronous calculation. The image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder. Since the image features can be calculated and stored asynchronously, when the image features are needed, they can be quickly obtained directly from the feature cache set to improve the efficiency of text evaluation, without the need to frequently pull video frames from the video stream and perform feature calculations. Therefore, in the case of live broadcasts or short videos with a large amount of text such as comments or bullet screens, where the time interval between the publication of texts is short and the corresponding frames of the previous and subsequent texts change little, the frequent pulling of video frames can be avoided, which may result in redundant information, and the real-time feature calculation can be avoided, which may consume a large amount of time.

[0070] To reduce information redundancy and improve the efficiency of asynchronous acquisition of image features, the target video can be sparsely sampled and then encoded to obtain image features, which are stored in the feature cache set. In some embodiments, before obtaining the image features of the target video frame matching the text to be evaluated from the feature cache set, the text evaluation method further includes:

[0071] Sparsely sample the target video at a preset sampling frequency to obtain the sampled video frame;

[0072] Encode the sampled video frame through an image encoder to obtain image features;

[0073] Store the image features in the feature cache set.

[0074] Among them, the preset sampling frequency can be flexibly set according to actual needs and is not limited here. For example, the preset sampling frequency of the target video can be set according to the video scene of the target video, the key frames of the target video, or empirical knowledge, etc. After determining the preset sampling frequency, the target video can be sparsely sampled at the preset sampling frequency to obtain the sampled video frame, which can include one or more video frames (i.e., images). Then, the sampled video frame can be encoded through an image encoder (Image Encoder) to obtain image features, that is, the image features are extracted from the sampled video frame through the image encoder. Among them, the image encoder can be called an image feature encoder, and the specific type and structure of the image encoder can be flexibly set according to actual needs and are not limited here. For example, the image encoder can be a Bootstrapping Language Image Pre-training (BLIP) model that can be uniformly understood and generated.

[0075] After obtaining the image features, the image features can be stored in the feature cache set. For example, the image features can be associated with the video identifier of the target video and stored in the feature cache set. Subsequently, the image features can be retrieved from the feature cache set based on the video identifier. Another example is that the playback timestamp of the sampled video frame can be obtained, and the playback timestamp, the image features, and the video identifier of the target video can be associated and stored in the feature cache set. Subsequently, the image features matching the playback timestamp can be retrieved from the feature cache set based on the video identifier. Storing the image features obtained by performing image sparse sampling and encoding on the target video in the feature cache set improves the flexibility and reliability of image feature storage. Moreover, the required image features can be quickly retrieved from the feature cache set, improving the efficiency of image feature retrieval.

[0076] During the process of performing image sparse sampling on the target video, to improve the accuracy of image sparse sampling on the target video, the preset sampling frequency can be determined based on the video scene of the target video for sampling. In some embodiments, performing image sparse sampling on the target video according to the preset sampling frequency to obtain the sampled video frame includes:

[0077] Determine the video scene of the target video, and determine the preset sampling frequency of the target video according to the video scene;

[0078] Perform image sparse sampling on the target video according to the preset sampling frequency to obtain the sampled video frame.

[0079] Specifically, first, the video scene of the target video can be determined. For example, the video scene of the target video can be determined according to the video content or video identifier of the target video. The video scene can include a live broadcast scene and a short video scene, etc. Then, the preset sampling frequency of the target video can be determined according to the video scene, so that image sparse sampling can be performed on the target video according to the preset sampling frequency to obtain the sampled video frame. For example, in the live broadcast scene, generally, the content of the video frame does not change much, and the sampling frequency can be set to sample one video frame (i.e., an image) every 10 seconds. Another example is that in a scene where the video frame changes rapidly, such as a short video, a shorter sampling interval can be set, such as setting the sampling frequency to sample one video frame every 1 second. By performing image sparse sampling on the target video based on the video scene of the target video to determine the preset sampling frequency, the accuracy of image sparse sampling on the target video is improved.

[0080] During the process of storing the image features, to improve the convenience of storage, the image features can be stored in the form of key-value pairs. In some embodiments, storing the image features in the feature cache set includes:

[0081] Obtain the video identifier of the target video;

[0082] Store the video identifier and the image features in the feature cache set in the form of key-value pairs.

[0083] Specifically, the video identifier of the target video can be obtained from the video file. This video identifier is used to uniquely identify the target video and can be composed of at least one of numbers, letters, text, and symbols, which is not limited herein. For example, the video identifier can include the live room number or the short video ID (Identity document), etc. Then, the video identifier and the image features can be stored in the feature cache set in the form of key-value pairs (i.e., key-value pairs, which can be abbreviated as kv). The key-value pair can specifically be a map or a dictionary, etc., which is not limited herein. Among them, the key can be used to store the video identifier, and the value can be used to store the image features. By storing the image features in the form of key-value pairs, not only the convenience of storing the image features is improved, but also the convenience and efficiency of subsequent image feature search are improved.

[0084] During the process of obtaining the image features, in order to avoid missing the key information in the video stream due to factors such as too fast change of the video stream or poor timing of video frame capture, the image features of multiple video frames can be obtained for storage so that the image features of multiple video frames can be used subsequently. In some embodiments, there are multiple sampled video frames. The sampled video frames are encoded by an image encoder to obtain the image features, including: encoding each sampled video frame by the image encoder respectively to obtain multiple image features, and each image feature corresponds to the image feature of each sampled video frame.

[0085] Storing the video identifier and the image features in the feature cache set in the form of key-value pairs includes: when there are historical image features stored in the feature cache set, storing the video identifier and multiple image features in the form of key-value pairs and replacing the historical image features in the feature cache set.

[0086] Specifically, when there are multiple sampled video frames, each sampled video frame can be encoded by the image encoder respectively to obtain the image features of each sampled video frame, and thus multiple image features can be obtained. The multiple image features can be formed into an image feature list, and the video identifier and the image feature list are stored in the feature cache set in the form of key-value pairs. At this time, the value in the key-value pair is used to store multiple image features, such as an image feature list with a length of n.

[0087] In order to save cache space, when storing the currently acquired image features, it is possible to determine whether historical image features are already stored in the feature cache set. If no historical image features are stored in the feature cache set, the video identifier and multiple image features can be directly stored in the feature cache set in the form of key-value pairs; if historical image features are stored in the feature cache set, the video identifier and multiple image features can be stored in the form of key-value pairs and replace the historical image features in the feature cache set, that is, clear the historical image features in the feature cache set, and only store the currently newly acquired image features in the feature cache set.

[0088] Since the image features of the target video are pre-stored in the feature cache set, when image features are needed, one or n image features can be quickly obtained directly from the feature cache set, without the need to pull the video frame in real time, nor the need to perform feature calculation in real time. This can reduce the time for data transmission and the time for image feature calculation, make full use of the redundancy of video features, reduce the computing power overhead and unnecessary operations of pulling the video frame, and use an asynchronous image feature calculation method to overcome the bottleneck of the most time-consuming video frame pulling link, so as to achieve fast multi-modal text evaluation and improve the efficiency of text evaluation.

[0089] S102. Extract the text features of the text to be evaluated through a multi-modal evaluation model, and fuse the text features and image features to obtain fused features.

[0090] Among them, the model type, model structure, etc. of the multi-modal evaluation model can be flexibly set according to actual needs and will not be limited here. For example, the multi-modal evaluation model can include a text encoder (Text Encoder), a feature fusion module (Fusion Module), and a task head. The text encoder can be a Bidirectional Encoder Representation from Transformers (BERT), which is used to extract text features. The feature fusion module is used to fuse features (such as feature splicing), and the task head includes several linear layers (Linear Layer) for classification. The deployment of the image encoder can be called an asynchronous service, and the deployment of the multi-modal evaluation model can be called a real-time service. In order to improve the efficiency of model inference (such as text evaluation), the model formats of the image encoder and the multi-modal evaluation model can be converted into a more efficient inference format such as TensorRT through a format conversion tool to evaluate the text through this format.

[0091] In the process of applying the multimodal evaluation model, the text features of the text to be evaluated can be extracted through the text encoder of the multimodal evaluation model, and the text features and image features can be fused (such as splicing) through the feature fusion module of the multimodal evaluation model to obtain fused features. Through several linear layers in the task head of the multimodal evaluation model, text evaluation is performed based on the fused features to obtain the output results corresponding to each linear layer, and the output result corresponding to the linear layer with the highest output confidence among the several linear layers is used as the final text evaluation result, thereby realizing multimodal text evaluation and improving the accuracy of text evaluation.

[0092] In order to improve the accuracy of text evaluation, the image encoder and the multimodal evaluation model may be jointly trained in advance. In some embodiments, the text evaluation method further includes:

[0093] Obtain training samples, where the training samples include text evaluation labels, sample video images, and sample evaluation texts;

[0094] Encode the sample video picture through an image encoder to obtain sample image features;

[0095] Extracting sample text features of the sample evaluation text through a multimodal evaluation model, and fusing sample image features and sample text features to obtain sample fusion features;

[0096] According to the sample fusion features, the text evaluation prediction is performed through the multimodal evaluation model to obtain the prediction result;

[0097] Calculate the difference between the text evaluation label and the predicted result to get the loss value;

[0098] The parameters of the image encoder and the multimodal evaluation model are adjusted according to the loss value until the preset stopping condition is met.

[0099] Specifically, first, training samples can be obtained from a sample library, or training samples sent by a terminal for collecting samples can be received. Of course, training samples can also be obtained in other ways, which are not limited here. Among them, the training samples can include information such as text evaluation labels, sample video images, and sample evaluation texts. The text evaluation labels can be used to indicate the real category to which the sample evaluation text belongs, such as whether the sample evaluation text is a violation or not, etc. The sample video images can include one or more images, and the sample evaluation text includes text that comments on or expresses opinions about the sample video images.

[0100] After obtaining the training samples, the sample video images in the training samples can be encoded by an image encoder to obtain sample image features, that is, the sample image features of the sample video images can be extracted by the image encoder. Also, the sample text features of the sample evaluation text in the training samples can be extracted by the multimodal evaluation model, and the sample image features and the sample text features can be fused by the multimodal evaluation model to obtain sample fusion features. For example, the sample image features and the sample text features can be spliced ​​by the multimodal evaluation model to obtain sample fusion features.

[0101] Then, based on the sample fusion features, a multimodal evaluation model can be used to perform text evaluation prediction to obtain a prediction result, which may include a prediction label and other information. The prediction label can be used to indicate the prediction category to which the predicted sample evaluation text belongs, such as whether the sample evaluation text is a violation or not, and its corresponding score.

[0102] After obtaining the prediction result, the difference between the text evaluation label and the prediction result can be calculated through a loss function (such as a cross entropy loss function) to obtain a loss value. At this time, the parameters of the image encoder and the multimodal evaluation model can be adjusted according to the loss value until the preset stop condition is met. Among them, the preset stop condition can be flexibly set according to actual needs. For example, the preset stop condition can be minimization of loss or the number of iterations reaches a preset number, etc., which is not limited here. The preset number can be flexibly set according to actual needs and is not limited here.

[0103] After completing the training of the image encoder and the multimodal evaluation model, the image features of the video screen can be extracted through the image encoder, and the text features of the text to be evaluated can be extracted through the multimodal evaluation model. The text features and image features are fused through the multimodal evaluation model to obtain fused features, so as to perform text evaluation based on the fused features.

[0104] S103. Perform text evaluation through a multimodal evaluation model according to the fusion features to obtain a text evaluation result, wherein the image features in the feature cache set are obtained by encoding the video screen of the target video through an image encoder, and the image encoder and the multimodal evaluation model are pre-jointly trained.

[0105] After fusing text features and image features to obtain fused features, a multimodal evaluation model can be used to perform text evaluation based on the fused features to obtain a text evaluation result (which can be called a multimodal recognition result). The text evaluation result is used to indicate the target category to which the text to be evaluated belongs, for example, whether the text to be evaluated is a violation or not, and its corresponding confidence level, etc.

[0106] After obtaining the text evaluation result, security risk control can be performed on target video playback scenarios such as live broadcasts or short videos. In some embodiments, according to the fusion features, text evaluation is performed through a multi-modal evaluation model. After obtaining the text evaluation result, the text evaluation method further includes:

[0107] Determine the corresponding security risk control strategy according to the text evaluation result, and perform corresponding operations according to the security risk control strategy.

[0108] Specifically, after obtaining the text evaluation result, the corresponding security risk control strategy can be determined according to the text evaluation result, and corresponding operations can be performed according to the security risk control strategy. By quickly obtaining the text evaluation result, the service response time can be significantly improved, providing support for the real-time video stream scenario security risk control strategy. For example, for the to-be-evaluated text (such as comments) that violates the regulations, it is not displayed, and for the to-be-evaluated text that does not violate the regulations, it is displayed; for another example, the corresponding security level can be determined according to the text evaluation result. When the security level is the preset level, it indicates that the to-be-evaluated text seriously violates the regulations. At this time, the account that publishes the to-be-evaluated text can be obtained, and the account that continuously publishes violation texts within a period of time can be blocked, warned, or muted, etc. By using the method of image feature reuse and image feature asynchronous calculation, the time for feature calculation and data pulling is significantly reduced, realizing multi-modal real-time text evaluation, and timely taking corresponding measures based on the security risk control strategy, improving the timeliness and reliability of security risk control for scenarios such as live broadcasts or short videos.

[0109] In some embodiments, determining the corresponding security risk control strategy according to the text evaluation result and performing corresponding operations according to the security risk control strategy includes:

[0110] When the text evaluation result is that the to-be-evaluated text is abnormal, determine that the corresponding security risk control strategy is not to output the to-be-evaluated text, and perform the operation of not outputting the to-be-evaluated text;

[0111] When the text evaluation result is that the to-be-evaluated text is not abnormal, determine that the corresponding security risk control strategy is to output the to-be-evaluated text, and perform the operation of outputting the to-be-evaluated text.

[0112] Specifically, when the text evaluation result indicates that there is an abnormality in the text to be evaluated, it means that the text to be evaluated is in violation of the regulations and does not meet the publication conditions. At this time, it can be determined that the corresponding security risk control strategy is not to output the text to be evaluated, and the operation of not outputting the text to be evaluated is executed. When the text evaluation result indicates that there is no abnormality in the text to be evaluated, it means that the text to be evaluated is not in violation of the regulations and meets the publication conditions. At this time, it can be determined that the corresponding security risk control strategy is to output the text to be evaluated, and the operation of outputting the text to be evaluated is executed, such as displaying the text to be evaluated. By quickly and accurately obtaining the real-time text evaluation results of multiple modalities and taking corresponding measures in a timely manner based on the security risk control strategy, the timeliness and reliability of processing the text to be evaluated are improved.

[0113] Embodiments of the present application can obtain the text to be evaluated corresponding to the target video that needs to be text-evaluated, and obtain the image features of the target video frame matching the text to be evaluated from the feature cache set. Then, the text features of the text to be evaluated are extracted through a multi-modal evaluation model, and the text features and image features are fused to obtain fused features. At this time, text evaluation can be performed through the multi-modal evaluation model according to the fused features to obtain a text evaluation result. Among them, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multi-modal evaluation model are jointly trained in advance. This solution can pre-encode the video frames of the target video to obtain image features and store them in the feature cache set, so that the image features of the target video frame matching the text to be evaluated can be quickly obtained from the feature cache set, without the need to frequently pull the video stream for calculation in real time, reducing the calculation time-consuming, and accurate multi-modal text evaluation can be performed based on the fused features obtained by fusing the text features of the text to be evaluated and the image features of the matching target video frame, improving the accuracy and efficiency of text evaluation.

[0114] According to the method described in the above embodiments, the following will give further detailed examples for illustration.

[0115] In this embodiment, it is assumed that the text evaluation device is integrated in the server. Please refer to Figure 3 , Figure 3 which is a schematic flowchart of the text evaluation method provided by the embodiments of the present application. The method process may include:

[0116] S201. Determine the video scene of the currently played target video, and determine the preset sampling frequency of the target video according to the video scene.

[0117] For example Figure 4As shown, the server can calculate image features asynchronously. First, it can obtain a real-time video stream such as a live broadcast or a short video, and use this real-time video stream as the target video being played currently. Then, the server can determine the video scene of the target video based on the video content or video identifier of the target video. This video scene can include a live broadcast scene or a short video scene, etc. At this time, the server can determine the preset sampling frequency of the target video according to the video scene. By determining the preset sampling frequency based on the video scene of the target video, the accuracy of sampling frequency determination can be improved.

[0118] S202. Perform image sparse sampling on the target video according to the preset sampling frequency to obtain the sampled video frame.

[0119] After determining the preset sampling frequency of the target video, the server can perform image sparse sampling on the target video according to the preset sampling frequency to obtain the sampled video frame, which can include one or more video frames (i.e., images). For example, in a live broadcast scene, the content of the general video frame changes little, and the sampling frequency can be set to sample one video frame every 10 seconds. Another example is that in a scene with fast-changing images such as a short video, a shorter sampling interval can be set, such as setting the sampling frequency to sample one video frame every 1 second. By determining the preset sampling frequency based on the video scene of the target video and performing image sparse sampling on the target video, the accuracy of image sparse sampling of the target video is improved.

[0120] S203. Encode the sampled video frame through an image encoder to obtain image features.

[0121] As Figure 4 shown, the server can encode the sampled video frame through the image encoder of the asynchronous service to obtain image features, that is, extract the image features of the sampled video frame through the image encoder. By performing image sparse sampling on the target video and then encoding, information redundancy is reduced, and the efficiency of asynchronous acquisition of image features is improved.

[0122] S204. Store the image features and the video identifier of the target video in the feature cache set in the form of key-value pairs.

[0123] As Figure 4 shown, after obtaining the image features, the server can cache the image features. For example, the server can obtain the video identifier of the target video and store the image features and the video identifier of the target video in the feature cache set in the form of key-value pairs (i.e., key-value pairs). This feature cache set can be a database for storing image features. By storing image features in the form of key-value pairs, not only the flexibility and convenience of image feature storage are improved, but also the convenience and efficiency of subsequent image feature search are improved.

[0124] It should be noted that in order to avoid missing key information in the video stream due to factors such as too fast video stream changes or poor video frame capture timing, a single video frame may miss the key information in the video stream. Therefore, the server can obtain and store the image features of multiple video frames so that the image features of multiple video frames can be used to fuse text features for accurate multi-modal text evaluation later. For example, the server can sample multiple sampled video frames, encode each sampled video frame through an image encoder to obtain the image features of each sampled video frame, and thus obtain multiple image features. Then the server can form a list of image features from the multiple image features and store the video identifier and the list of image features in the feature cache set in the form of key-value pairs. At this time, the value in the key-value pair represents the storage of multiple image features, such as an image feature list of length n, and the key represents the video identifier. Later, multiple image features can be obtained from the feature cache set through the video identifier.

[0125] In addition, in order to save cache space, when storing the currently obtained image features, the server can determine whether historical image features have been stored in the feature cache set. If no historical image features are stored in the feature cache set, the server can directly store the video identifier and multiple image features in the feature cache set in the form of key-value pairs; if historical image features are stored in the feature cache set, the server can store the video identifier and multiple image features in the form of key-value pairs and replace the historical image features in the feature cache set, that is, clear the historical image features in the feature cache set, and only store the currently newly obtained image features in the feature cache set.

[0126] S205. Obtain the text to be evaluated corresponding to the target video that needs to be text-evaluated.

[0127] As Figure 4 shown, during the playback of the target video, the server can obtain the text to be evaluated through a real-time channel and perform text evaluation. Specifically, during the playback of the target video, when receiving comment text such as comments or bullet screens published by the user for the currently played video frame, the server can obtain the comment text such as comments or bullet screens and use this comment text as the text to be evaluated that needs to be text-evaluated. Or, the server can obtain the voice uttered for the currently played video frame in depth and convert the text obtained from the voice, and use this text as the text to be evaluated.

[0128] S206. Extract the text features of the text to be evaluated through the text encoder of the multi-modal evaluation model.

[0129] As Figure 4As shown in the figure, after obtaining the text to be evaluated, the server can extract the text features of the text to be evaluated through the multi-modal evaluation model of the real-time service. For example, the text encoder of the multi-modal evaluation model can be used to extract the features of the text to be evaluated to obtain the text features of the text to be evaluated.

[0130] It should be noted that in order to improve the accuracy of text evaluation, the image encoder and the multi-modal evaluation model can be jointly trained in advance. For example, the server can obtain training samples, which include text evaluation labels, sample video frames, and sample evaluation texts; then encode the sample video frames through the image encoder to obtain sample image features; and extract the sample text features of the sample evaluation texts through the multi-modal evaluation model, and fuse the sample image features and the sample text features to obtain sample fusion features; at this time, according to the sample fusion features, text evaluation prediction is performed through the multi-modal evaluation model to obtain a prediction result; finally, the difference between the text evaluation label and the prediction result is calculated to obtain a loss value, and the parameters of the image encoder and the multi-modal evaluation model are adjusted according to the loss value until the preset stop conditions such as minimizing the loss or reaching the preset number of iterations are met, and the training of the image encoder and the multi-modal evaluation model is completed. After the training of the image encoder and the multi-modal evaluation model is completed, the server can accurately perform text evaluation through the image encoder and the multi-modal evaluation model.

[0131] S207. Obtain the image features of the target video frame matching the text to be evaluated from the feature cache set according to the video identifier of the target video.

[0132] Since the image features have been calculated asynchronously and stored in the cache set, when the image features are needed, the server can directly obtain the required image features of the target video frame matching the text to be evaluated from the feature cache set according to the video identifier of the target video, without frequently pulling video frames from the video stream and performing feature calculations, which can reduce the data transmission time and the image feature calculation time. Therefore, in the case of live broadcasts or short videos with a large amount of text such as comments or bullet screens, where the time interval between the publication of texts is short and the corresponding frames of the front and back texts change little, the problem of redundant information caused by frequently pulling video frames and the large time consumption caused by real-time feature calculation can be avoided. By means of asynchronous image feature calculation, the efficiency of obtaining image features is improved, and thus the efficiency of multi-modal text evaluation is improved.

[0133] S208. Fuse the text features and the image features through the feature fusion module of the multi-modal evaluation model to obtain fusion features.

[0134] Such as Figure 4As shown, after obtaining the text features and the image features, the server can fuse the text features and the image features (such as feature splicing) through the feature fusion module of the multimodal evaluation model to obtain the fused features.

[0135] S209, performing text evaluation based on fusion features through several linear layers in the task head of the multimodal evaluation model to obtain a text evaluation result.

[0136] like Figure 4 As shown, after obtaining the fused features, the server can perform text evaluation based on the fused features through several linear layers in the task head of the multimodal evaluation model, obtain the output results corresponding to each linear layer (such as the output classification results), and use the output result corresponding to the linear layer with the highest output confidence among the several linear layers as the final text evaluation result. The text evaluation result is used to indicate the target category to which the text to be evaluated belongs, for example, whether the text to be evaluated is a violation or not, thereby realizing multimodal text evaluation and improving the accuracy of text evaluation.

[0137] S210: Determine a corresponding security risk control strategy based on the text evaluation result, and perform corresponding operations on the text to be evaluated based on the security risk control strategy.

[0138] After obtaining the text evaluation results, the server can perform security risk control on target video playback scenarios such as live broadcasts or short videos. For example, the server can determine the corresponding security risk control strategy based on the text evaluation results, and perform corresponding operations on the text to be evaluated based on the security risk control strategy.

[0139] By quickly obtaining text evaluation results, the service response time can be greatly improved, providing support for real-time video streaming scenario security risk control strategies. For example, illegal texts to be evaluated (such as comments) are not displayed, while non-illegal texts to be evaluated are displayed; for another example, the account that posted the text to be evaluated can be obtained, and the account that continuously posted illegal texts over a period of time can be blocked, warned, or banned.

[0140] Embodiments of the present application can determine a preset sampling frequency of a target video based on the video scene of the target video to perform image sparse sampling on the target video, and extract image features of the sampled video frames through an image encoder and store them in a feature cache set, realizing pre-caching of image features through asynchronous calculation of image features. Thus, image features of the target video frame matching the text to be evaluated can be quickly obtained from the feature cache set as needed, without the need to frequently pull the video stream for calculation in real time, significantly reducing the time for feature calculation and data pulling. Moreover, when the text to be evaluated is obtained, the text features of the text to be evaluated can be extracted through a multi-modal evaluation model, and the text features are fused with the image features obtained from the feature cache set to obtain fused features, and text evaluation is performed based on the fused features to obtain a text evaluation result, realizing real-time text evaluation in multiple modalities, quickly and accurately obtaining the real-time text evaluation result in multiple modalities, and timely taking corresponding measures based on the security risk control strategy determined by the text evaluation result, improving the timeliness and reliability of security risk control for video scenarios such as live broadcasts or short videos.

[0141] To facilitate better implementation of the text evaluation method provided by the embodiments of the present application, the embodiments of the present application also provide an apparatus based on the above text evaluation method. The meanings of the nouns are the same as those in the above text evaluation method, and the specific implementation details can refer to the description in the method embodiments.

[0142] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of the text evaluation apparatus provided by the embodiments of the present application. The text evaluation apparatus 300 may include an acquisition unit 301, a fusion unit 302, an evaluation unit 303, etc.

[0143] Among them, the acquisition unit 301 is used to acquire the text to be evaluated corresponding to the target video that needs to be text-evaluated, and acquire the image features of the target video frame matching the text to be evaluated from the feature cache set.

[0144] The fusion unit 302 is used to extract the text features of the text to be evaluated through a multi-modal evaluation model, and fuse the text features and the image features to obtain fused features.

[0145] The evaluation unit 303 is used to perform text evaluation through a multi-modal evaluation model based on the fused features to obtain a text evaluation result.

[0146] Among them, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multi-modal evaluation model are jointly trained in advance.

[0147] In some embodiments, the text evaluation apparatus 300 further includes:

[0148] A sampling unit for sparsely sampling images of a target video at a preset sampling frequency to obtain a sampled video frame;

[0149] An encoding unit for encoding the sampled video frame through an image encoder to obtain image features;

[0150] A storage unit for storing the image features in a feature cache set.

[0151] In some embodiments, the sampling unit is specifically configured to:

[0152] Determine the video scene of the target video, and determine the preset sampling frequency of the target video according to the video scene;

[0153] Sparsely sample images of the target video at the preset sampling frequency to obtain a sampled video frame.

[0154] In some embodiments, the storage unit is specifically configured to:

[0155] Obtain the video identifier of the target video;

[0156] Store the video identifier and the image features in the form of key-value pairs in the feature cache set.

[0157] In some embodiments, there are multiple sampled video frames. The encoding unit is specifically configured to: encode each sampled video frame through an image encoder respectively to obtain multiple image features, and each image feature corresponds to the image feature of each sampled video frame;

[0158] The storage unit is specifically configured to: when historical image features are stored in the feature cache set, store the video identifier and the multiple image features in the form of key-value pairs, and replace the historical image features in the feature cache set.

[0159] In some embodiments, the text evaluation device 300 further includes:

[0160] A sample acquisition unit for acquiring training samples, where the training samples include text evaluation labels, sample video frames, and sample evaluation texts;

[0161] A sample encoding unit for encoding the sample video frame through an image encoder to obtain sample image features;

[0162] A sample fusion unit for extracting sample text features of the sample evaluation text through a multimodal evaluation model, and fusing the sample image features and the sample text features to obtain sample fusion features;

[0163] A prediction unit for performing text evaluation prediction through a multimodal evaluation model according to the sample fusion features to obtain a prediction result;

[0164] A calculation unit, configured to calculate the difference between the text evaluation label and the prediction result to obtain a loss value;

[0165] An adjustment unit, configured to adjust the parameters of the image encoder and the multimodal evaluation model according to the loss value until a preset stop condition is met.

[0166] In some embodiments, the text evaluation device 300 further includes:

[0167] An execution unit, configured to determine a corresponding security risk control policy according to the text evaluation result, and perform corresponding operations according to the security risk control policy.

[0168] In some embodiments, the execution unit is specifically configured to:

[0169] When the text evaluation result indicates that the text to be evaluated is abnormal, determine that the corresponding security risk control policy is not to output the text to be evaluated, and perform the operation of not outputting the text to be evaluated;

[0170] When the text evaluation result indicates that the text to be evaluated is not abnormal, determine that the corresponding security risk control policy is to output the text to be evaluated, and perform the operation of outputting the text to be evaluated.

[0171] In an embodiment of the present application, the acquisition unit 301 can acquire the text to be evaluated corresponding to the target video that needs to be text-evaluated, and acquire the image features of the target video frame matching the text to be evaluated from the feature cache set. Then, the fusion unit 302 extracts the text features of the text to be evaluated through the multimodal evaluation model, and fuses the text features and the image features to obtain fused features. At this time, the evaluation unit 303 can perform text evaluation through the multimodal evaluation model according to the fused features to obtain a text evaluation result. Among them, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multimodal evaluation model are jointly trained in advance. This solution can encode the video frames of the target video in advance to obtain image features and store them in the feature cache set, so that the image features of the target video frame matching the text to be evaluated can be quickly acquired from the feature cache set, without the need to frequently pull the video stream for calculation in real time, reducing the calculation time-consuming, and can accurately perform multimodal text evaluation based on the fused features obtained by fusing the text features of the text to be evaluated and the image features of the matching target video frame, improving the accuracy and efficiency of text evaluation.

[0172] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0173] The embodiments of the present application further provide a computer device, which can be a terminal, a server, etc. For example, Figure 6 as shown, which shows a schematic structural diagram of the computer device involved in the embodiments of the present application. Specifically:

[0174] The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art can understand that Figure 6 the computer device structure shown in does not constitute a limitation on the computer device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0175] The processor 401 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it executes various functions of the computer device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401.

[0176] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0177] The computer device further includes a power supply 403 for powering each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0178] The computer device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0179] Although not shown, the computer device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:

[0180] Obtain the text to be evaluated corresponding to the target video that needs to be textually evaluated, and obtain the image features of the target video frame matching the text to be evaluated from the feature cache set; extract the text features of the text to be evaluated through a multimodal evaluation model, and fuse the text features and the image features to obtain fused features; according to the fused features, perform text evaluation through the multimodal evaluation model to obtain a text evaluation result; among them, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multimodal evaluation model are jointly trained in advance.

[0181] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not described in detail in a certain embodiment, reference may be made to the detailed description of the text evaluation method above, and details will not be repeated here.

[0182] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementation manners in the above embodiments.

[0183] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by computer instructions, or by controlling related hardware through computer instructions. The computer instructions can be stored in a computer-readable storage medium (i.e., the storage medium) and loaded and executed by a processor. For this reason, an embodiment of the present application provides a storage medium, in which a computer program is stored. The computer program may include computer instructions, and the computer program can be loaded by a processor to execute any one of the text evaluation methods provided in the embodiments of the present application.

[0184] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, and details will not be repeated here.

[0185] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0186] Since the instructions stored in the storage medium can execute the steps in any one of the text evaluation methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any one of the text evaluation methods provided in the embodiments of the present application can be realized. For details, reference may be made to the previous embodiments, and details will not be repeated here.

[0187] The above has introduced in detail a text evaluation method, device and storage medium provided in the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A text evaluation method, characterized in that, Including: Obtaining the text to be evaluated corresponding to the target video that needs text evaluation, and obtaining the image features of the target video frame matching the text to be evaluated from the feature cache set; Extracting the text features of the text to be evaluated through a multi-modal evaluation model, and fusing the text features and the image features to obtain fused features; Performing text evaluation through the multi-modal evaluation model according to the fused features to obtain a text evaluation result; Wherein, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multi-modal evaluation model are jointly trained in advance.

2. The text evaluation method according to claim 1, characterized in that Before obtaining the image features of the target video frame matching the text to be evaluated from the feature cache set, the text evaluation method further includes: Performing image sparse sampling on the target video according to a preset sampling frequency to obtain sampled video frames; Encoding the sampled video frames through the image encoder to obtain image features; Storing the image features in the feature cache set.

3. The text evaluation method according to claim 2, wherein Performing image sparse sampling on the target video according to a preset sampling frequency to obtain sampled video frames includes: Determining the video scene of the target video, and determining the preset sampling frequency of the target video according to the video scene; Performing image sparse sampling on the target video according to the preset sampling frequency to obtain sampled video frames.

4. The text evaluation method according to claim 2, characterized in that Storing the image features in the feature cache set includes: Obtaining the video identifier of the target video; Storing the video identifier and the image features in the feature cache set in the form of key-value pairs.

5. The text evaluation method according to claim 4, characterized in that, There are multiple sampled video frames, and encoding the sampled video frames through the image encoder to obtain image features includes: Encoding each sampled video frame through the image encoder respectively to obtain multiple image features, and each image feature corresponds to the image feature of each sampled video frame; Storing the video identifier and the image features in the feature cache set in the form of key-value pairs includes: When historical image features are stored in the feature cache set, storing the video identifier and the multiple image features in the form of key-value pairs, and replacing the historical image features in the feature cache set.

6. The text evaluation method according to claim 1, characterized in that, The text evaluation method further includes: Obtaining training samples, where the training samples include text evaluation labels, sample video frames, and sample evaluation texts; Encoding the sample video frames through an image encoder to obtain sample image features; Extracting the sample text features of the sample evaluation text through a multi-modal evaluation model, and fusing the sample image features and the sample text features to obtain sample fused features; Performing text evaluation prediction through the multi-modal evaluation model according to the sample fused features to obtain a prediction result; Calculating the difference between the text evaluation label and the prediction result to obtain a loss value; Adjust the parameters of the image encoder and the multi-modal evaluation model according to the loss value until a preset stop condition is met.

7. The text evaluation method according to any one of claims 1 to 6, characterized in that, After obtaining the text evaluation result by performing text evaluation on the fused features through the multi-modal evaluation model, the text evaluation method further includes: Determine a corresponding security risk control strategy according to the text evaluation result, and perform corresponding operations according to the security risk control strategy.

8. The text evaluation method according to claim 7, wherein The determining a corresponding security risk control strategy according to the text evaluation result and performing corresponding operations according to the security risk control strategy includes: When the text evaluation result indicates that the text to be evaluated is abnormal, determine that the corresponding security risk control strategy is not to output the text to be evaluated, and perform the operation of not outputting the text to be evaluated; When the text evaluation result indicates that the text to be evaluated is normal, determine that the corresponding security risk control strategy is to output the text to be evaluated, and perform the operation of outputting the text to be evaluated.

9. A text evaluation device, characterized in that, Includes: An acquisition unit, configured to acquire the text to be evaluated corresponding to the target video that needs to be text-evaluated, and acquire the image features of the target video frame matching the text to be evaluated from the feature cache set; A fusion unit, configured to extract the text features of the text to be evaluated through a multi-modal evaluation model, and fuse the text features and the image features to obtain fused features; An evaluation unit, configured to perform text evaluation through the multi-modal evaluation model according to the fused features to obtain a text evaluation result; Wherein, the image features in the feature cache set are obtained by encoding the video frames of the target video through an image encoder, and the image encoder and the multi-modal evaluation model are jointly trained in advance.

10. A storage medium, characterized in that, The storage medium is used to store a computer program, and the computer program is loaded by a processor to execute the text evaluation method according to any one of claims 1 to 8.