Resource auditing method, device and equipment based on large model and storage medium
Through the resource audit method based on the big model, the large model reviews the graphic data of the resources to be reviewed, and the problem of low manual audit efficiency in the existing technology is solved, achieving more efficient and accurate content security audits.
Patent Information
- Application Number
- CN202411845497.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-13
AI Technical Summary
The existing technology relies on manual audit in the field of content audit, which is inefficient and difficult to cope with the audit requirements of multiple modal data.
The resource audit method based on the big model is adopted, and the target graphic data of the resources to be reviewed is retrieved, and the positive and negative sample knowledge bases are retrieved, and the pre-trained big model is used to review whether the resources are violated.
Improve the accuracy and efficiency of content security audits, enable automated processing, and reduce the time and cost of manual audits.
Smart Images

Figure CN119988856A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, specifically to the field of artificial intelligence technology such as data security and big models, and more particularly to a resource audit method, device, equipment and storage medium based on big models. Background Art
[0002] Content security review is a key link in ensuring the healthy and orderly development of online platforms. The content that needs to be reviewed often includes multiple modalities such as text, pictures, and videos.
[0003] In the current content review field, manual review is still one of the main means. This is a traditional review method, where reviewers need to check the content to be reviewed item by item according to established standards and guidelines to determine whether the content is compliant. Summary of the invention
[0004] The present invention provides a resource audit method, device, equipment and storage medium based on a large model.
[0005] According to one aspect of the present disclosure, a resource audit method based on a large model is provided, wherein the method comprises:
[0006] Get the target graphic and text data of the resource to be reviewed;
[0007] Based on the target graphic data, respectively retrieving a target positive sample and a target negative sample that are most relevant to the target graphic data from a pre-created positive sample knowledge base and a pre-created negative sample knowledge base;
[0008] Based on the target graphic data, the target positive samples and the target negative samples, a pre-trained large model is used to review whether the resource to be reviewed is in violation of regulations.
[0009] According to another aspect of the present disclosure, a resource audit device based on a large model is provided, wherein the device comprises:
[0010] A data acquisition module is used to obtain target graphic data of resources to be reviewed;
[0011] A sample retrieval module is used to retrieve, based on the target graphic data, a target positive sample and a target negative sample that are most relevant to the target graphic data from a pre-created positive sample knowledge base and a pre-created negative sample knowledge base respectively;
[0012] The content review module is used to review whether the resource to be reviewed violates the regulations based on the target graphic data, the target positive sample and the target negative sample, using a pre-trained large model.
[0013] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any possible implementation manner and the aspects described above.
[0017] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the above-mentioned aspect and any possible implementation manner.
[0018] According to yet another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements the above-mentioned aspects and any possible implementation method.
[0019] According to the technology disclosed in the present invention, the accuracy and efficiency of content security review can be effectively improved.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0022] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;
[0023] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0024] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0025] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;
[0026] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure;
[0027] Figure 6 The block diagram is a block diagram of an electronic device for implementing the method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0029] Obviously, the described embodiments are only part of the embodiments of the present disclosure, but not all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present disclosure.
[0030] It should be noted that the terminal devices involved in the embodiments of the present disclosure may include but are not limited to mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers and other smart devices; display devices may include but are not limited to personal computers, televisions and other devices with display functions.
[0031] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0032] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure; Figure 1 As shown, this embodiment provides a resource audit method based on a large model, which may specifically include the following steps:
[0033] S101, obtaining target graphic and text data of resources to be reviewed;
[0034] S102, based on the target image and text data, respectively retrieve the target positive sample and target negative sample that are most relevant to the target image and text data from a pre-created positive sample knowledge base and a negative sample knowledge base;
[0035] S103. Based on the target graphic data, target positive samples and target negative samples, a pre-trained large model is used to review whether the resources to be reviewed are in violation of regulations.
[0036] The execution subject of the resource audit method based on a large model in this embodiment is a resource audit device based on a large model, which can be an independent electronic entity or a software integrated application. When used, it is installed on a computer or other device to audit the resources to be audited. In this embodiment, it is mainly used to audit whether the content of the resources to be audited violates the regulations and verify the security of the resources to be audited, which can also be called a security audit of the resources to be audited.
[0037] Considering that the audit of resource content in plain text form is relatively easy to implement, this embodiment mainly performs security audits on resources including graphic data. That is, in this embodiment, the resources to be audited may be resources including images and text. For example, the resources to be audited may be web pages, documents, etc. in graphic form including images and text, and may also be resources in video form. For resources in graphic form, the images and text therein may be directly obtained as target graphic data. For resources in video form, each video frame may be obtained as a picture, and the corresponding text may be obtained based on each video frame to form the target graphic data together.
[0038] In this embodiment, the target positive sample refers to the positive sample in the positive sample knowledge base that is most relevant to the target image and text data. The target positive sample is a sample that does not violate the rules and can also be called a safe sample. The target negative sample refers to the negative sample in the negative sample knowledge base that is most relevant to the target image and text data. The target negative sample is a sample that violates the rules and can also be called an unsafe sample.
[0039] In this embodiment, the pre-trained big model can review whether the resources to be reviewed are in violation of regulations based on the target graphic data, target positive samples and target negative samples of the resources to be reviewed. The target positive samples and target negative samples can provide an effective basis for the big model, which can improve the accuracy of the big model in conducting security reviews on the content of the resources to be reviewed.
[0040] The resource audit method based on a large model of this embodiment can effectively improve the accuracy and efficiency of resource security audits by adopting a large model to audit whether the resources to be audited are in violation of regulations based on target graphic data, target positive samples, and target negative samples. Moreover, the technical solution of this embodiment can be automatically implemented, which saves time and effort compared to traditional manual audits, and can effectively improve the efficiency of audits.
[0041] Figure 2 is a schematic diagram of a second embodiment of the present disclosure; the resource audit method based on a large model of this embodiment, in the above Figure 1 Based on the technical solutions of the embodiments shown in the figure, the technical solutions of the present disclosure are further described in more detail. Figure 2 As shown, the resource audit method based on the big model of this embodiment may specifically include the following steps:
[0042] S201: When the resource to be reviewed includes a video, obtaining at least two key frames in the video;
[0043] In this embodiment, the resource to be reviewed includes a video as an example. In actual application scenarios, each frame in the video can be used as a key frame, but since there is content redundancy between many consecutive adjacent frames in the video, using each video frame as a key frame will result in a large amount of repeated reviews. Therefore, in order to reduce the cost of review and improve review efficiency, in this embodiment, at least two key frames in the video can be obtained to participate in subsequent security reviews, which can effectively reduce the amount of review and improve review efficiency.
[0044] In order to improve the accuracy of at least two key frames in the acquired video, step 201 can be implemented by the following steps:
[0045] (1) Divide the video into at least two scene video segments according to the scene;
[0046] (2) According to a preset extraction strategy, at least one key frame is extracted from each scene video clip to obtain at least two key frames.
[0047] Since the video content to be reviewed is usually edited from multiple video clips, the camera switches frequently, the scenes are diverse, and the number of key frames cannot be determined in advance, the frames containing security risks may appear for a very short time in the video. In order to improve the accuracy of video content review, in this embodiment, it is required to extract at least one key frame for each scene; thereby, at least two key frames extracted from the entire video can accurately cover each scene in the video.
[0048] Based on the above ideas, the video can be first divided into at least two scene video segments according to the scene, and then at least one key frame can be extracted from each scene video segment, so that at least two key frames involved in the review can cover the information of each scene, thereby improving the comprehensiveness and accuracy of the video review.
[0049] In this embodiment, the preset extraction strategy may require extracting a key frame every preset frame. Alternatively, a multi-level judgment may be used. For example, if the duration of the scene video segment is greater than or equal to the first preset duration and less than the second preset duration, a first number of key frames are extracted; if the scene video segment is greater than or equal to the second preset duration and less than the third preset duration, a second number of key frames are extracted; if the scene video segment is greater than or equal to the third preset duration, a third number of key frames are extracted. The third preset duration is greater than the second preset duration, and the second preset duration is greater than the first preset duration. The first number is less than the second number, and the second number is less than the third number. The specific lengths of each preset duration and the values of each number are set according to the needs of the specific scene and are not limited here. In this embodiment, three-level judgment is taken as an example. In actual applications, only two-level judgments may be included, or multi-level judgments of more than four levels may be included, which are not limited here.
[0050] Optionally, in an embodiment of the present disclosure, dividing the video into at least two scene video segments according to scenes may include: acquiring a scene switching video frame in the video; and then dividing the video into at least two scene video segments based on the scene switching video frame.
[0051] The scene switching video frame of this embodiment can be used as the index frame of scene switching and appear at the beginning of the corresponding scene video segment. Using the index frame of scene switching as the basis for segmenting the scene video segment, the video can be accurately segmented according to the scene, and at least two key frames of the video can be efficiently and accurately obtained.
[0052] S202, obtaining text data aligned with each key frame;
[0053] In this embodiment, the text data aligned with each key frame can be considered as the text content before the key frame in the video that has not been reviewed. For example, in the content review process, when the acquisition and review of the target graphic data for review are performed at the granularity of key frames, the text data aligned with each key frame can refer to the text data corresponding to at least one continuous video frame before the key frame and after the previous key frame. In this embodiment, the key frame before includes the key frame itself. If the current key frame is the first key frame in the video, it is sufficient to obtain at least one continuous video frame before the current key frame. If the current key frame is the first frame in the video, only one video frame is included at this time.
[0054] The specific implementation may include the following:
[0055] (a1) for each key frame, obtaining at least one continuous video frame in the video that has not been reviewed before the key frame;
[0056] (b1) extracting text content of the at least one continuous video frame to obtain key-frame aligned text data.
[0057] For example, in a specific implementation, the text in each video frame of at least one continuous video frame can be extracted first to obtain the first text data; this method mainly corresponds to extracting subtitles from video frames. Specifically, image processing techniques such as edge detection, binarization, connected domain analysis, etc. can be used to distinguish subtitles from backgrounds, filter out frames without subtitles, and identify video frames with subtitles. Then, the consecutive frames are compared to filter out duplicate and redundant video frames. Then, optical character recognition (OCR) technology is used to process and identify the subtitles in the remaining video frames to obtain text information with the timestamp of the video frame, and finally, the text information extracted from the adjacent video frames is deduplicated and filtered again. The timestamp may refer to the timestamp corresponding to each video frame, which can realize the alignment of text information with the video frame.
[0058] Then, speech recognition is performed on the speech corresponding to at least one continuous video frame to obtain second text data. Specifically, audio information in at least one continuous video frame is first extracted, and then the audio information is converted into text information with a time stamp using automatic speech recognition (Automatic Speech Recognition; ASR) technology to obtain the second text data.
[0059] Finally, the first text data and the second text data are concatenated to obtain the key-frame aligned text data.
[0060] It should be noted that in the above implementation, the text data is also acquired in segments. In practical applications, all the first text data and all the second text data of the video can also be acquired in the above manner of acquiring the first text data and the second text data, and all the text data of the video can be obtained by splicing them before and after the timestamps. When in use, the text data can be segmented according to the timestamps carried in the text data to obtain text data aligned with each key frame. For example, if the timestamp of the first key frame in the video is t1, then the corresponding text data acquired is all text data before the timestamp t1, including the text data with the timestamp t1. If the timestamp of the second key frame is t2, then the corresponding text data acquired is all text data before the timestamp t2 and after the timestamp t1, including the text data with the timestamp t2, but excluding the text data with the timestamp t1, and so on, and the text data aligned with each key frame can be acquired.
[0061] Further optionally, the above-mentioned at least two key frames are extracted according to the scene video clip. After the text data of all video frames are obtained, the text data of each scene video clip can also be segmented based on the timestamp of the scene switching video frame. When the text data aligned with each key frame is obtained, it is possible to obtain only all unreviewed text data before the timestamp of the key frame in the current scene video clip. For example, in a scene video clip, the timestamp of the first key frame is t3, then the corresponding text data obtained includes all text data before the timestamp t3 in the text data of the current scene video clip. The timestamp of the second key frame is t4, then the corresponding text data obtained is all text data before the timestamp t4 and after the timestamp t3 in the text data of the current scene video clip, and so on. The text data aligned with each key frame in the scene video clip can be obtained.
[0062] Based on the above, it can be known that the target graphic data of the video of this embodiment includes at least two key frames and text data aligned with each key frame. The target graphic data corresponding to each key frame may include the key frame and text data aligned with the key frame.
[0063] By adopting the above method, at least two key frames in the video to be reviewed and the text data aligned with each key frame can be accurately, reasonably and comprehensively obtained.
[0064] S203: for each key frame, based on the key frame and the text data corresponding to the key frame, retrieve the most relevant target positive sample and target negative sample from a pre-created positive sample knowledge base and a negative sample knowledge base;
[0065] Optionally, before step S203, the following steps may be included:
[0066] (a2) Collect multiple positive samples; configure positive sample labels; build a positive sample knowledge base; each positive sample includes a positive sample image and a positive sample text, as well as a positive sample label; the positive sample label is used to indicate that the positive sample does not contain any illegal information; that is, the positive sample image and the positive sample text of each positive sample do not contain any illegal information; the positive sample labels of all positive samples in the positive sample knowledge base are the same;
[0067] (b3) Collect multiple negative samples; configure negative sample labels; build a negative sample knowledge base; each negative sample includes a negative sample image, a negative sample text, and a negative sample label; the negative sample label is used to identify that the negative sample includes illegal information; that is, the negative sample image and / or negative sample text of each negative sample includes illegal information. The negative sample labels of all negative samples in the negative sample knowledge base are the same.
[0068] In this embodiment, during retrieval, for each key frame, the feature vector of the key frame is first obtained, and then the feature vector of the text data corresponding to the key frame is obtained, and then the two feature vectors are projected into a vector space for fusion to obtain the feature vector of the target graphic data corresponding to the key frame.
[0069] Correspondingly, for the positive sample knowledge base and the negative sample knowledge base, the feature vector of each positive sample and the feature vector of each negative sample can also be calculated in the above manner to obtain the positive sample feature vector library and the negative sample feature vector library. In this embodiment, the above manner can be used to accurately and effectively construct the positive sample knowledge base and the negative sample knowledge base, providing effective support for subsequent resource audits.
[0070] During retrieval, the similarity between the feature vector of the target image and text data corresponding to the key frame and the feature vector of each positive sample in the positive sample feature vector library can be calculated respectively, and the one with the greatest similarity can be obtained as the feature vector of the target positive sample; similarly, the one with the greatest similarity to the feature vector of the target image and text data corresponding to the key frame can be obtained from the negative sample feature vector library as the feature vector of the target negative sample; thus, the target positive sample and the target negative sample can be obtained.
[0071] The above method of this embodiment is to obtain the target positive sample and the target negative sample by vector retrieval. In practical applications, the most semantically relevant target positive sample and target negative sample can also be retrieved from the positive sample knowledge base and the negative sample knowledge base directly based on the target graphic data corresponding to the key frame, that is, the key frame and the corresponding text data. In practical applications, other methods can also be used for retrieval, which are not limited here.
[0072] S204: Based on each key frame and the text data corresponding to the key frame, the target positive sample and the target negative sample, a pre-trained large model is used to review whether the video to be reviewed violates the regulations.
[0073] When conducting content review, this embodiment is subject to the maximum input limit of the large model, and can review only one or a preset number of keyframes corresponding to the graphic data at a time. The graphic data corresponding to a keyframe includes a keyframe and the text data corresponding to the keyframe.
[0074] In this embodiment, it is taken as an example that only one key frame's graphic data is audited each time. During the specific audit, the key frame and the text data corresponding to the key frame can be input into the large model, and the target positive sample and the target negative sample can be used as prompts and also input into the large model, so that the large model can use the target positive sample and the target negative sample as examples to audit whether the key frame and the text data key frame corresponding to the key frame violate the rules; if the rules are violated, the corresponding video is considered to be in violation. Only when the target graphic data of all key frames in the video are not in violation after detection by the large model, can it be determined that the corresponding video is not in violation.
[0075] The target positive samples and target negative samples in this embodiment act as cases and are embedded in a specific Prompt template as few-shots, which helps the large model generate accurate and expected outputs.
[0076] In this embodiment, the feature vectors of the target graphic data corresponding to the above-mentioned key frames, the feature vectors and corresponding labels of the target positive samples, and the feature vectors and corresponding negative sample labels of the target negative samples can also be directly input into the large model, so that the large model can directly perform a security review of whether there is a violation based on the input feature vectors.
[0077] In this embodiment, a combination of one positive and one negative case is used to provide a few-shot to the large model. By comparing the positive and negative cases, the large model can more robustly respond to subtle changes in the input data and enhance the robustness of the model output.
[0078] The step S203 in this embodiment can be implemented by using a retrieval-augmented generation (RAG) module.
[0079] In the field of content security review, illegal content is diverse, and some users will confront it based on their experience and generate new illegal content. In this embodiment, the RAG module can be controlled to flexibly and efficiently absorb recent dynamic data into the positive sample knowledge base and the negative sample knowledge base, thereby ensuring the timeliness of the data and allowing the large model to more easily cope with new types of violations.
[0080] The large model-based resource review method of this embodiment can be widely used in multiple fields such as Internet content platforms, social media, advertising review, copyright protection, etc., and can realize content review of multiple modal data including text, pictures, videos, etc.
[0081] The big model-based resource audit method of this embodiment can accurately obtain target positive samples and target negative samples by adopting the above-mentioned method, provide effective cases for the security audit of the big model, assist the big model in performing security audits on each key frame in the video and the corresponding text data, and effectively improve the accuracy of resource audits.
[0082] Moreover, the technical solution of this embodiment can perform security audits on the content of image, text, and video modalities; for example, when auditing resources in video modality, it can efficiently extract picture and text information, reduce the redundancy of video information input, and improve the accuracy and efficiency of resource audits.
[0083] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure; the resource audit method based on the large model of this embodiment, in the above Figure 2 Based on the technical solutions of the embodiments shown in the figure, the technical solutions of the present disclosure are further described in more detail. Figure 3 As shown, in this embodiment, the method for obtaining a scene switching video frame in a video may specifically include the following steps:
[0084] S301, sampling a video according to a first downsampling rate to obtain a plurality of first video frames included in the video;
[0085] S302, performing scene switching detection on each first video frame, and obtaining an identifier of the first video frame where the scene switching occurs;
[0086] S303: Based on the identifier of the first video frame, mark the corresponding video frame in the video as a scene switching video frame;
[0087] In practical applications, scene switching includes direct conversion and gradual transition. For scenes that are directly switched, adjacent video frames can be detected to detect them. Considering that the repetition of continuous video frames is high, in order to improve the detection efficiency, in this embodiment, the video can be sampled at a first downsampling rate to obtain multiple first video frames included in the video. The first downsampling rate of this embodiment can be 2 or 3. Then, the identifier of the first video frame where the scene switching occurs is detected and obtained from the multiple first video frames. Since the first video frame also comes from the original video, the identifier of the first video frame where the scene switching occurs is the scene switching video frame in the corresponding video.
[0088] For example, when step S302 of this embodiment is specifically implemented, it may include the following steps:
[0089] (a4) converting each first video frame from a first color space (i.e., RGB color space) marked by red (Red; R), green (Green; G), and blue (Blue; B) to a second color space (i.e., HSL color space) marked by hue (Hue; H), saturation (Saturation; S), and brightness (Lightness; L);
[0090] (b4) for each first video frame in the second color space, calculating a histogram difference of the first video frame relative to a first video frame of a previous nearest neighbor as the histogram difference of the first video frame;
[0091] For details, please refer to the calculation method of the histogram difference between two related video frames, which will not be described in detail here.
[0092] (c4) for each first video frame in the second color space, obtaining a plurality of adjacent video frames included in a front nearest neighbor sliding window and a rear nearest neighbor sliding window of the first video frame;
[0093] The size of the sliding window in this embodiment can be set according to actual experience and needs. For example, the size of a sliding window can be set to 2 frames or 3 frames. Of course, it can also be set to other numbers of frames according to actual needs. During use, the sliding window can move back and forth in the video, so that the front nearest neighbor sliding window and the back nearest neighbor sliding window corresponding to each video frame can be obtained, and then multiple adjacent video frames can be obtained.
[0094] (d4) obtaining an average value of the histogram difference of the first video frame based on the histogram difference of the first video frame and the histogram difference of each adjacent video frame in the plurality of adjacent video frames;
[0095] The histogram difference of each adjacent video frame can be implemented by referring to the above step (b4). Specifically, the sum of the histogram difference of the current first video frame and the histogram differences of all adjacent video frames in multiple adjacent video frames can be taken, and then the average is taken as the average value of the histogram difference of the current first video frame. Steps (c4) and (d4) are used to accurately and effectively obtain the average value of the histogram difference of the current first video frame.
[0096] (e4) for each first video frame in the second color space, detecting whether a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame is greater than or equal to a preset ratio threshold;
[0097] The preset ratio threshold of this embodiment can be set according to experience or demand, for example, it can be 2, 3, 4 or other values, which are not limited here.
[0098] (f4) in response to a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame being greater than or equal to a preset ratio threshold, determining that a scene switch occurs in the first video frame;
[0099] Specifically, if the ratio of the histogram difference of the first video frame to the average value of the histogram difference of the first video frame is greater than a preset ratio threshold, it indicates that a scene switch occurs in the first video frame relative to the previous adjacent video frame.
[0100] (g4) Obtain the identifier of the first video frame.
[0101] Through this round of downsampling, the scene switching frames in the video can be accurately and effectively detected and marked.
[0102] Optionally, in actual application scenarios, the video may also include a transition scene of a gradual in and gradual out type, which cannot be detected by using the above steps S301-S303. In this case, the following steps S304-S306 may be further used for detection.
[0103] S304, sampling the video according to a second downsampling rate to obtain a plurality of second video frames included in the video; the second downsampling rate is greater than the first downsampling rate;
[0104] For example, the second downsampling rate may be 10, that is, 10 frames are sampled into 1 frame. Alternatively, it may be 9, 11, or other values, which are not limited here.
[0105] S305, performing scene switching detection on each second video frame, and obtaining an identifier of a second video frame in which a scene switching occurs;
[0106] S306: Based on the identifier of the second video frame, mark the corresponding video frame in the video as a scene switching video frame.
[0107] The implementation method of step S305 also adopts the method of the above steps (a4)-(e4), which will not be repeated here.
[0108] Since the transition scene of the gradual in and out type changes slowly, steps S304-S306 effectively detect the scene switching mode by setting a larger downsampling rate, and then obtain the corresponding scene switching video frame.
[0109] It should be noted that in actual scenes, there are often direct scene changes in the video. In this case, steps S301-S303 are directly adopted to mark the scene switching video frames in the video. In order to obtain the scene switching video frames more accurately and comprehensively, steps S304-S306 can be adopted as a supplement to detect whether there is a transition scene with a gradual in and out style. If so, the corresponding scene switching video frames are also marked in the video. If there is no gradual in and out transition scene change in a certain video, steps S304-S306 may not mark the scene switching video frames.
[0110] In this embodiment, the method for obtaining scene switching video frames in the video, through the above-mentioned two-round downsampling detection method, the total processing time is shorter than the time for detecting all frames in the video frame by frame, but the number of valid scenes detected is more than one round of frame-by-frame detection, which can efficiently, accurately and comprehensively mark the scene switching video frames in the video, and then accurately obtain scene video clips, improve the accuracy of key frames in the obtained video, and provide effective support for the security review of subsequent videos.
[0111] Based on the above method of this embodiment, the scene switching video frames in the video can be obtained efficiently and accurately, and then based on the scene switching video frames, the video can be efficiently and accurately divided into at least two scene video segments; then, according to the preset extraction strategy, at least one key frame can be extracted from each scene video segment to obtain at least two key frames, and the image data in the video to be reviewed can be obtained efficiently and accurately, providing effective support for the security review of the video.
[0112] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure; Figure 4 As shown, this embodiment provides a resource audit device 400 based on a large model, including:
[0113] Data acquisition module 401, used to acquire target graphic data of resources to be reviewed;
[0114] A sample retrieval module 402 is used to retrieve, based on the target graphic data, a target positive sample and a target negative sample that are most relevant to the target graphic data from a pre-created positive sample knowledge base and a pre-created negative sample knowledge base respectively;
[0115] The content review module 403 is used to review whether the resource to be reviewed violates the regulations based on the target graphic data, the target positive sample and the target negative sample, using a pre-trained large model.
[0116] The big model-based resource audit device 400 of this embodiment implements the implementation principle and technical effect of big model-based resource audit by adopting the above-mentioned modules, which is the same as the implementation of the above-mentioned related method embodiments. For details, please refer to the records of the above-mentioned related method embodiments, which will not be repeated here.
[0117] Figure 5 is a schematic diagram of a fifth embodiment of the present disclosure; the resource audit device 500 based on a large model of this embodiment, in the above Figure 4 Based on the technical solutions of the embodiments shown in the figure, the technical solutions of the present disclosure are further described in more detail. Figure 5 As shown, the resource audit device 500 based on the large model of this embodiment includes the above Figure 4 The modules with the same name and function are shown as: data acquisition module 501 , sample retrieval module 502 and content review module 503 .
[0118] like Figure 5 As shown, in this embodiment, the data acquisition module 501:
[0119] The picture data acquisition unit 5011 is used to acquire at least two key frames in the video when the resource to be reviewed includes a video;
[0120] The text data acquisition unit 5012 is used to acquire the text data aligned with each of the key frames.
[0121] Further optionally, in one embodiment of the present disclosure, the picture data acquiring unit 5011 is used to:
[0122] Dividing the video into at least two scene video segments according to the scene;
[0123] According to a preset extraction strategy, at least one key frame is extracted from each of the scene video clips to obtain the at least two key frames.
[0124] Further optionally, in one embodiment of the present disclosure, the picture data acquiring unit 5011 is used to:
[0125] Obtaining a scene switching video frame in the video;
[0126] Based on the scene switching video frame, the video is divided into at least two scene video segments.
[0127] Further optionally, in one embodiment of the present disclosure, the picture data acquiring unit 5011 is used to:
[0128] Sampling the video at a first downsampling rate to obtain a plurality of first video frames included in the video;
[0129] Performing scene switching detection on each of the first video frames to obtain an identifier of the first video frame in which the scene switching occurs;
[0130] Based on the identifier of the first video frame, a corresponding video frame is marked in the video as a scene switching video frame.
[0131] Further optionally, in one embodiment of the present disclosure, the picture data acquiring unit 5011 is used to:
[0132] Converting each first video frame from a first color space labeled with red, green, and blue to a second color space labeled with hue, saturation, and brightness;
[0133] For each of the first video frames in the second color space, calculating a histogram difference of the first video frame relative to a first video frame of a previous nearest neighbor as a histogram difference of the first video frame;
[0134] For each of the first video frames in the second color space, detecting whether a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame is greater than or equal to a preset ratio threshold;
[0135] In response to a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame being greater than or equal to a preset ratio threshold, determining that a scene switch occurs in the first video frame;
[0136] Obtain an identifier of the first video frame.
[0137] Further optionally, in one embodiment of the present disclosure, the picture data acquiring unit 5011 is further configured to:
[0138] For each of the first video frames in the second color space, obtaining a plurality of adjacent video frames included in a front nearest neighbor sliding window and a rear nearest neighbor sliding window of the first video frame;
[0139] Based on the histogram difference of the first video frame and the histogram difference of each adjacent video frame in the plurality of adjacent video frames, an average value of the histogram difference of the first video frame is obtained.
[0140] Further optionally, in one embodiment of the present disclosure, the picture data acquiring unit 5011 is used to:
[0141] An average value of the histogram difference of the first video frame and the histogram differences of each adjacent video frame in the plurality of adjacent video frames is taken as the average value of the histogram difference of the first video frame.
[0142] Further optionally, in one embodiment of the present disclosure, the picture data acquiring unit 5011 is further configured to:
[0143] Sampling the video at a second downsampling rate to obtain a plurality of second video frames included in the video; the second downsampling rate is greater than the first downsampling rate;
[0144] Performing scene switching detection on each of the second video frames to obtain an identifier of the second video frame in which the scene switching occurs;
[0145] Based on the identifier of the second video frame, a corresponding video frame is marked in the video as a scene switching video frame.
[0146] Further optionally, in one embodiment of the present disclosure, the text data acquisition unit 5012 is used to:
[0147] For each of the key frames, obtaining at least one continuous video frame in the video that has not been previously reviewed for the key frame;
[0148] The text content of the at least one continuous video frame is extracted to obtain the key frame aligned text data.
[0149] Further optionally, in one embodiment of the present disclosure, the text data acquisition unit 5012 is used to:
[0150] Extracting text from each of the at least one continuous video frame to obtain first text data;
[0151] Performing speech recognition on the speech corresponding to the at least one continuous video frame to obtain second text data;
[0152] The first text data and the second text data are concatenated.
[0153] Further optionally, if Figure 5 As shown, in one embodiment of the present disclosure, the content security audit device 500 based on the large model further includes: a construction module 504, which is used to:
[0154] Collect multiple positive samples; configure positive sample labels; build the positive sample knowledge base; each positive sample includes a positive sample image and a positive sample text, as well as the positive sample label; the positive sample label is used to indicate that the corresponding positive sample does not include illegal information;
[0155] Collect multiple negative samples; configure negative sample labels; build the negative sample knowledge base; each of the negative samples includes a negative sample image and a negative sample text, as well as the negative sample label; the negative sample label is used to identify that the corresponding negative sample includes violation information.
[0156] The big model-based resource audit device 500 of this embodiment implements the implementation principle and technical effect of big model-based resource audit by adopting the above-mentioned modules, which is the same as the implementation of the above-mentioned related method embodiments. For details, please refer to the records of the above-mentioned related method embodiments, which will not be repeated here.
[0157] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0158] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0159] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0160] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0161] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0162] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the above-mentioned methods of the present disclosure. For example, in some embodiments, the above-mentioned methods of the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the above-mentioned methods of the present disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the above-mentioned methods of the present disclosure in any other appropriate manner (e.g., by means of firmware).
[0163] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0164] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0165] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0167] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0168] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0169] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0170] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A resource audit method based on a large model, wherein: The method comprises: Get the target graphic and text data of the resource to be reviewed; Based on the target graphic data, respectively retrieving a target positive sample and a target negative sample that are most relevant to the target graphic data from a pre-created positive sample knowledge base and a pre-created negative sample knowledge base; Based on the target graphic data, the target positive samples and the target negative samples, a pre-trained large model is used to review whether the resource to be reviewed is in violation of regulations.
2. The method according to claim 1, wherein: The step of obtaining target graphic and text data of the resource to be reviewed includes: When the resource to be reviewed includes a video, obtaining at least two key frames in the video; The text data aligned with each of the key frames is obtained.
3. The method according to claim 2, wherein: The obtaining of at least two key frames in the video includes: Dividing the video into at least two scene video segments according to the scene; According to a preset extraction strategy, at least one key frame is extracted from each of the scene video clips to obtain the at least two key frames.
4. The method according to claim 3, wherein: The step of dividing the video into at least two scene video segments according to the scene includes: Obtaining a scene switching video frame in the video; Based on the scene switching video frame, the video is divided into at least two scene video segments.
5. The method according to claim 4, wherein: The obtaining of the scene switching video frame in the video includes: Sampling the video at a first downsampling rate to obtain a plurality of first video frames included in the video; Performing scene switching detection on each of the first video frames to obtain an identifier of the first video frame in which the scene switching occurs; Based on the identifier of the first video frame, a corresponding video frame is marked in the video as a scene switching video frame.
6. The method according to claim 5, wherein: The performing scene switching detection on each of the first video frames to obtain an identifier of the first video frame where the scene switching occurs includes: Converting each first video frame from a first color space labeled with red, green, and blue to a second color space labeled with hue, saturation, and brightness; For each of the first video frames in the second color space, calculating a histogram difference of the first video frame relative to a first video frame of a previous nearest neighbor as a histogram difference of the first video frame; For each of the first video frames in the second color space, detecting whether a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame is greater than or equal to a preset ratio threshold; In response to a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame being greater than or equal to a preset ratio threshold, determining that a scene switch occurs in the first video frame; Obtain an identifier of the first video frame.
7. The method according to claim 6, wherein: Before detecting, for each of the first video frames in the second color space, whether a ratio of a histogram difference of the first video frame to an average value of the histogram difference of the first video frame is greater than or equal to a preset ratio threshold, the method further includes: For each of the first video frames in the second color space, obtaining a plurality of adjacent video frames included in a front nearest neighbor sliding window and a rear nearest neighbor sliding window of the first video frame; Based on the histogram difference of the first video frame and the histogram difference of each adjacent video frame in the plurality of adjacent video frames, an average value of the histogram difference of the first video frame is obtained.
8. The method according to claim 7, wherein: The obtaining, based on the histogram difference of the first video frame and the histogram difference of each adjacent video frame in the plurality of adjacent video frames, an average value of the histogram difference of the first video frame includes: An average value of the histogram difference of the first video frame and the histogram differences of each adjacent video frame in the plurality of adjacent video frames is taken as the average value of the histogram difference of the first video frame.
9. The method according to any one of claims 4 to 8, wherein: The obtaining of the scene switching video frame in the video further includes: Sampling the video at a second downsampling rate to obtain a plurality of second video frames included in the video; the second downsampling rate is greater than the first downsampling rate; Performing scene switching detection on each of the second video frames to obtain an identifier of the second video frame in which the scene switching occurs; Based on the identifier of the second video frame, the corresponding video frame is marked in the video as a scene switching video frame.
10. The method according to any one of claims 2 to 8, wherein: The step of obtaining the text data aligned with each of the key frames comprises: For each of the key frames, obtaining at least one continuous video frame in the video that has not been previously reviewed for the key frame; The text content of the at least one continuous video frame is extracted to obtain the key frame aligned text data.
11. The method according to claim 10, wherein: The step of extracting text content of the at least one continuous video frame to obtain the key frame aligned text data includes: Extracting text from each of the at least one continuous video frame to obtain first text data; Performing speech recognition on the speech corresponding to the at least one continuous video frame to obtain second text data; The first text data and the second text data are concatenated.
12. The method according to any one of claims 1 to 8 and 11, wherein: The method further comprises: Collect multiple positive samples; configure positive sample labels; build the positive sample knowledge base; each positive sample includes a positive sample image and a positive sample text, as well as the positive sample label; the positive sample label is used to indicate that the corresponding positive sample does not include illegal information; Collect multiple negative samples; configure negative sample labels; build the negative sample knowledge base; each of the negative samples includes a negative sample image and a negative sample text, as well as the negative sample label; the negative sample label is used to identify that the corresponding negative sample includes violation information.
13. A resource audit device based on a large model, wherein: The device comprises: A data acquisition module is used to obtain target graphic data of resources to be reviewed; A sample retrieval module, used for retrieving, based on the target graphic data, a target positive sample and a target negative sample that are most relevant to the target graphic data from a pre-created positive sample knowledge base and a pre-created negative sample knowledge base; The content review module is used to review whether the resource to be reviewed violates the regulations based on the target graphic data, the target positive sample and the target negative sample, using a pre-trained large model.
14. The device according to claim 13, wherein: The data acquisition module comprises: A picture data acquisition unit, configured to acquire at least two key frames in a video when the resource to be reviewed includes a video; The text data acquisition unit is used to acquire the text data aligned with each of the key frames.
15. The device according to claim 14, wherein: The picture data acquisition unit is used to: Dividing the video into at least two scene video segments according to the scene; According to a preset extraction strategy, at least one key frame is extracted from each of the scene video clips to obtain the at least two key frames.
16. The device according to claim 15, wherein: The picture data acquisition unit is used to: Obtaining a scene switching video frame in the video; Based on the scene switching video frame, the video is divided into at least two scene video segments.
17. The device according to claim 16, wherein: The picture data acquisition unit is used to: Sampling the video at a first downsampling rate to obtain a plurality of first video frames included in the video; Performing scene switching detection on each of the first video frames to obtain an identifier of the first video frame in which the scene switching occurs; Based on the identifier of the first video frame, a corresponding video frame is marked in the video as a scene switching video frame.
18. The device according to claim 17, wherein: The picture data acquisition unit is used to: Converting each first video frame from a first color space labeled with red, green, and blue to a second color space labeled with hue, saturation, and brightness; For each of the first video frames in the second color space, calculating a histogram difference of the first video frame relative to a first video frame of a previous nearest neighbor as a histogram difference of the first video frame; For each of the first video frames in the second color space, detecting whether a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame is greater than or equal to a preset ratio threshold; In response to a ratio of a histogram difference of the first video frame to an average value of the histogram differences of the first video frame being greater than or equal to a preset ratio threshold, determining that a scene switch occurs in the first video frame; Obtain an identifier of the first video frame.
19. The device according to claim 18, wherein: The picture data acquisition unit is further used for: For each of the first video frames in the second color space, obtaining a plurality of adjacent video frames included in a front nearest neighbor sliding window and a rear nearest neighbor sliding window of the first video frame; Based on the histogram difference of the first video frame and the histogram difference of each adjacent video frame in the plurality of adjacent video frames, an average value of the histogram difference of the first video frame is obtained.
20. The device according to claim 19, wherein The picture data acquisition unit is used to: An average value of the histogram difference of the first video frame and the histogram differences of each adjacent video frame in the plurality of adjacent video frames is taken as the average value of the histogram difference of the first video frame.
21. The device according to any one of claims 16 to 20, wherein: The picture data acquisition unit is further used for: Sampling the video at a second downsampling rate to obtain a plurality of second video frames included in the video; the second downsampling rate is greater than the first downsampling rate; Performing scene switching detection on each of the second video frames to obtain an identifier of the second video frame in which the scene switching occurs; Based on the identifier of the second video frame, the corresponding video frame is marked in the video as a scene switching video frame.
22. The device according to any one of claims 14 to 20, wherein: The text data acquisition unit is used to: For each of the key frames, obtaining at least one continuous video frame in the video that has not been previously reviewed for the key frame; The text content of the at least one continuous video frame is extracted to obtain the key frame aligned text data.
23. The device according to claim 22, wherein: The text data acquisition unit is used to: Extracting text from each of the at least one continuous video frame to obtain first text data; Performing speech recognition on the speech corresponding to the at least one continuous video frame to obtain second text data; The first text data and the second text data are concatenated.
24. The device according to any one of claims 13 to 20 and 24, wherein: The device further comprises: a building module, configured to: Collect multiple positive samples; configure positive sample labels; build the positive sample knowledge base; each positive sample includes a positive sample image and a positive sample text, as well as the positive sample label; the positive sample label is used to indicate that the corresponding positive sample does not include illegal information; Collect multiple negative samples; configure negative sample labels; build the negative sample knowledge base; each of the negative samples includes a negative sample image and a negative sample text, as well as the negative sample label; the negative sample label is used to identify that the corresponding negative sample includes violation information.
25. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.
26. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
27. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.
Citation Information
Cited By
Method, computing program product and device for identifying specific type of user address
CN120499150A