Data auditing method, terminal equipment and computer readable storage medium

By integrating the image text and audio text information of video data, combining multimodal large models and audit rule engines, the problem of low audit accuracy of financial short videos is solved, and efficient and accurate compliance judgment is achieved.

CN120751175AActive Publication Date: 2025-10-03SHENZHEN XIAOYING INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511204752.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-03
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

The existing technology for reviewing financial short videos has low accuracy, making it difficult to meet complex and diverse compliance requirements and unable to achieve accurate semantic reasoning.

Method used

By extracting image text information and audio text information from video data and fusing them into text fusion information, a comprehensive audit is conducted using a multimodal large model and an audit rule engine to detect risk warnings, user privacy, product pricing and other information, generating comprehensive and accurate audit results.

Benefits of technology

It improves the review accuracy of financial short videos, ensures that the video content meets compliance requirements, provides high-quality review basis, and can comprehensively and accurately judge the compliance of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751175A_ABST
    Figure CN120751175A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of data auditing, and provides a data auditing method, terminal equipment and a computer readable storage medium, and the method comprises the steps: obtaining first video data; extracting first image text information corresponding to the first video data; extracting first audio text information corresponding to the first video data; fusing the first image text information with the first audio text to obtain text fusion information; performing data auditing based on the text fusion information to obtain a first auditing result; and determining the compliance of the first video data according to the first auditing result. According to the method, the auditing precision of the short video data can be improved, and complex and multivariate video compliance requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data audit technology, and in particular relates to a data audit method, a terminal device, and a computer-readable storage medium. Background Art

[0002] With the development of multi-channel marketing scenarios, financial short videos have also been widely used and have become an important form of marketing and promotion for financial institutions. They must undergo compliance review before being released to meet compliance requirements in the financial field.

[0003] Most related technologies rely on rule engines and single-modal models, which cannot achieve accurate semantic reasoning, resulting in low review accuracy for financial short videos and difficulty in meeting the complex and diverse compliance requirements of financial videos. Summary of the Invention

[0004] The embodiments of the present application provide a data review method, apparatus, terminal device, and storage medium, which can improve the review accuracy of short videos and meet complex and diverse video compliance requirements.

[0005] In a first aspect, an embodiment of the present application provides a data audit method, comprising: Acquire first video data; Extracting first image text information corresponding to the first video data; Extracting first audio text information corresponding to the first video data; Fusing the first image text information with the first audio text to obtain text fusion information; Perform data review based on text fusion information to obtain a first review result; The compliance of the first video data is determined according to the first review result.

[0006] In an embodiment of the present application, image text information and audio text information are extracted from video data, and the image text information and audio text information are fused into text fusion information. The text fusion information combines the advantages of images and audio, can make up for the defects of single text extraction, can improve the integrity and accuracy of the review information, can adapt to the multimodal content characteristics of video data, can obtain more accurate and comprehensive text fusion information, and provide a high-quality basis for the review. Finally, the review is conducted based on the fused text fusion information, which can more comprehensively and accurately judge the compliance of the video content, improve the reliability of the review results, and help ensure that the video data meets relevant specifications and requirements before it is released.

[0007] In a possible implementation of the first aspect, extracting first image text information corresponding to the first video data includes: Extracting a single frame from the first video data according to a preset frequency to obtain a plurality of first images; For each first image, extract first information corresponding to the first image; wherein the first information includes first text content, first text type, and first text position corresponding to the first text content; The first information corresponding to each of the plurality of first images is combined to obtain first image text information corresponding to the first video data.

[0008] In an embodiment of the present application, by extracting a single frame of the video at a preset frequency and extracting the text information of the frame (including content, type and location), and then combining the text information of all frames, all image texts in the video can be obtained comprehensively, accurately and completely, providing comprehensive and reliable basic data for subsequent audio text correction, information fusion and compliance review based on image text.

[0009] In a possible implementation of the first aspect, extracting first audio text information corresponding to the first video data includes: extracting first audio data from the first video data; Transcribing the first audio data into text data corresponding to each of the multiple time series; The text data is combined according to a time sequence to obtain first audio text information corresponding to the first video data.

[0010] In an embodiment of the present application, by extracting video audio and transcribing it into text of the corresponding time series, and then combining it by time to form audio text information, the temporal logic of the audio content can be completely preserved, providing accurate voice dimension data support for subsequent fusion with image text and compliance review.

[0011] In a possible implementation of the first aspect, fusing the first image text information with the first audio text to obtain text fusion information includes: Performing text supplementation and correction on the first audio text information according to the first image text information to obtain supplemented and corrected second audio text information; The first image text information and the second audio text information are fused to obtain text fusion information.

[0012] In an embodiment of the present application, by supplementing and correcting the audio text information with image text information and fusing the two, it is possible to compensate for possible errors in audio transcription, supplement key content not covered by the audio, form complete and accurate text fusion information, and provide comprehensive and accurate text basis for subsequent data review.

[0013] In a possible implementation of the first aspect, performing data review based on the text fusion information to obtain a first review result includes: Detect risk warning information in the text fusion information to obtain a first sub-result; Detecting user privacy data in the text fusion information to obtain a second sub-result; Detect product pricing information in the text fusion information to obtain the third sub-result; A first audit result is obtained according to the first sub-result, the second sub-result and the third sub-result.

[0014] In the embodiment of the present application, through special detection and result integration of three core information types, namely risk warnings, user privacy, and product pricing, in text fusion information, a first audit result covering key dimensions can be accurately generated, providing a comprehensive and clear basis for compliance judgment at the text level of video data.

[0015] In a possible implementation of the first aspect, determining compliance of the first video data according to the first review result includes: detecting prohibited content in the first image to obtain a fourth sub-result; Detecting compliance of the logo content contained in the first image to obtain a fifth sub-result; The second audit result is obtained according to the fourth and fifth sub-results; The compliance of the first video data is determined according to the first review result and the second review result.

[0016] In an embodiment of the present application, by detecting prohibited content and logo compliance in an image and integrating the results to generate a second review result, and then combining it with the first review result at the text level, the compliance of the video data can be comprehensively determined from the dual dimensions of "image + text", providing an accurate and complete basis for the compliance review of video content.

[0017] In a possible implementation of the first aspect, determining compliance of the first video data according to the first review result and the second review result includes: If the first review result or the second review result indicates that the review has failed, the first video data is non-compliant; If both the first review result and the second review result indicate that the review is passed, the first video data is compliant.

[0018] In the embodiment of the present application, by clarifying the logic that "if any dimension of text or image fails the review, it is judged as non-compliant, and it is judged as compliant only if both dimensions pass", the final judgment of the compliance of video data can be completed strictly and efficiently, ensuring that only videos with both text and images that meet the requirements are deemed compliant, providing a clear and rigorous standard for the compliance control of video content.

[0019] In a possible implementation of the first aspect, the method further includes: Acquire second information; wherein the second information includes data information that has not been reviewed and location information corresponding to the data information that has not been reviewed in the first video data; Generate a data audit report based on the second information and issue an alarm prompt.

[0020] In the embodiment of the present application, by obtaining data that has not passed the review and its location information, generating reports and alarms, it is possible to accurately locate illegal content, intuitively present the review results, and promptly remind relevant personnel to handle it, providing efficient support for the rectification and risk prevention and control of illegal videos.

[0021] In a second aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a data audit method as described in any one of the first aspects above is implemented.

[0022] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the data audit method as described in any one of the first aspects above.

[0023] In a fourth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute any one of the data audit methods in the first aspect.

[0024] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] Figure 1 This is a flow chart of a data review method provided in one embodiment of the present application; Figure 2 This is a schematic diagram of the process of extracting image text information provided by an embodiment of the present application; Figure 3 This is a flowchart of extracting audio text information provided by an embodiment of the present application; Figure 4 This is a schematic diagram of the process of extracting fused text information provided by an embodiment of the present application; Figure 5This is a schematic diagram of the structure of the data audit rules provided in the embodiment of the present application; Figure 6 This is a flow chart of the data review process provided in the embodiment of the present application; Figure 7 This is a flowchart of data compliance judgment provided by an embodiment of the present application; Figure 8 This is a schematic diagram of the overall structure of the data review process provided in the embodiment of the present application; Figure 9 This is a structural diagram of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0029] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0030] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0031] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0032] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.

[0033] With the development of multi-channel marketing scenarios, financial short videos have also been widely used and have become an important form of marketing and promotion for financial institutions. They must undergo compliance review before being released to meet compliance requirements in the financial field.

[0034] Most related technologies rely on rule engines and single-modal models, which cannot achieve accurate semantic reasoning, resulting in low review accuracy for financial short videos and difficulty in meeting the complex and diverse compliance requirements of financial videos.

[0035] In order to solve the problems in the above-mentioned related technologies, the embodiments of the present application provide a data audit method, a terminal device and a computer-readable storage medium. The present application is applicable to the compliance audit of short videos or pictures for financial marketing. The method receives the video or picture link and metadata and encapsulates it into a task, which is processed by the audit process in the Kubernetes cluster; obtains information by extracting frames, extracting audio, etc., extracts text using multimodal large models, ASR models, etc., and generates unified subtitles through fusion and completion of the text large model; then combines multiple audit rules, calls related models, etc. for analysis, and outputs the audit conclusion. At the same time, it has distributed execution, multi-rule parallelism, fault tolerance and monitoring mechanisms. The core advantage lies in the integration of multi-source information and large model reasoning capabilities. The above method can improve the audit accuracy of short videos and meet complex and diverse video compliance requirements.

[0036] See also Figure 1 , is a flow chart of a data audit method provided in an embodiment of the present application. As an example and not a limitation, the method may include the following steps: S101, obtaining first video data.

[0037] In an embodiment of the present application, the first video data is submitted by the caller, packaged by the system, and stored in the task queue and database as one of the basic information for the subsequent review process. For example, in the financial field, the first video data mainly refers to the basic information used for the compliance review of financial marketing short videos or pictures, including metadata such as video or picture links, distribution channels, and content types. The caller submits the video or picture link and its metadata information (distribution channels, content types, etc.) to the review system. After receiving it, the system packages this information into a review task and adds it to the task queue. At the same time, the relevant metadata, creation time, and review status of the task are recorded in the database.

[0038] The entire system runs on Kubernetes (a container orchestration platform), and all audit-related computing resources (such as Pods) are centrally managed by Kubernetes. When the task volume surges (such as the centralized delivery of financial marketing videos) or decreases, resulting in an increase in idle Pods, excess Pods are automatically destroyed (for example, reducing the number to two) to avoid resource waste. When the queue backlog exceeds a threshold, Kubernetes automatically creates new Pods (for example, expanding from three to ten). Each Pod is an independent task processing unit, internally running an "audit process" containing a multi-threaded pool (for example, 10 threads). This multi-threaded pool allows a single Pod to handle multiple audit tasks simultaneously (for example, thread 1 downloads Task A, thread 2 extracts frames for Task B, and thread 3 merges subtitles for Task C), rather than processing them one by one sequentially. This significantly improves the task processing efficiency of a single Pod.

[0039] The Pod listens to the task queue through the thread in the thread pool. When a "pending" task is detected, it sends a request to the queue and obtains the task (the queue will mark the task as "received" to prevent duplicate processing). The thread parses the "video / image link" in the task and calls the built-in download tool (supporting HTTP / HTTPS protocol) to download the resource (first video data) from the corresponding platform (such as Douyin and Tencent Advertising); during downloading, the validity of the link is verified (such as whether it has expired or whether the format is compliant).

[0040] S102: Extract first image text information corresponding to the first video data.

[0041] In this embodiment of the present application, "first image text information" refers to all visible text extracted by identifying image frames of the first video data (a short financial marketing video to be reviewed). This text is directly derived from the video footage and is one of the core criteria for video content compliance review.

[0042] In one embodiment, see Figure 2 , is a flow chart of extracting image text information provided by an embodiment of the present application, such as Figure 2 As shown, step S102 includes: S201 : extracting a single frame from first video data according to a preset frequency to obtain a plurality of first images.

[0043] In the embodiment of the present application, "extracting a single frame from the first video data to obtain multiple first images" is the basic link in video content analysis. Its core is to extract image frames at fixed intervals from the financial marketing short video to be reviewed (first video data) to provide visual materials for subsequent text recognition and image compliance review.

[0044] Specifically, the system obtains the first video data (including the video link) from the task queue. After the Pod downloads the video resource to obtain the first video data, it performs frame extraction on the video, extracting image frames at a specified frequency, i.e., a preset frequency (e.g., 1 frame every 0.5 seconds). At the same time, it focuses on extracting the last frame (about 2 seconds of content, where financial videos often place compliance statements) to obtain an image sequence for recognition (i.e., the "first image").

[0045] S202 : For each first image, extract first information corresponding to the first image; wherein the first information includes first text content, a first text type, and a first text position corresponding to the first text content.

[0046] In the embodiment of the present application, text information (first information) is extracted from a single-frame image (first image) obtained by extracting frames from a video. The essence of this is to identify the text in the image through technical means, and to classify and mark the position of the text, thereby providing structured data for subsequent compliance audits.

[0047] Extracting the text content (first text content) from each image (first image) refers to the actual text content identified from the first image, which is the "semantics itself" of the text. For example, if the first image is an interest rate promotion screen in a video, the extracted text may be "annualized interest rate as low as 3.5%." If there is a risk warning in the corner of the screen, the extracted text may be "Investment is risky, enter the market with caution."

[0048] After extracting the text content corresponding to each image, the text is classified and categorized (first text type) to distinguish the nature and purpose of the text. For example, common types of text in financial video data include subtitle text, which is the explanatory text displayed dynamically in the video (such as product introduction subtitles); logo text, which is the text accompanying the logo of a financial institution (such as the brand name "XX Bank"); and other text, such as the descriptive text in product promotional images ("Loan limit up to 200,000 yuan"). After the text content is categorized, the specific location of the text in the image is located (first text location).

[0049] Specifically, the above-mentioned first information (first text content, first text type and first text position) can be completed using a multimodal large model (such as Qwen2.5-VL). The multimodal large model receives the first image as input, locates the text area in the picture through image recognition technology, performs OCR (optical character recognition) on the text area, and obtains the "first text content". Based on the position, style, and semantics of the text (such as "risk warning" is often located in a fixed corner and has a small font), it is automatically classified as the "first text type", and the coordinates or area of ​​the text in the image are synchronously recorded to generate the "first text position", which is finally output in a structured format (such as JSON).

[0050] S203: Combine the first information corresponding to the plurality of first images to obtain first image text information corresponding to the first video data.

[0051] In this embodiment of the application, the first information of all first images is integrated in a time series (frame order), and repeated text content is deduplicated and merged (e.g., if a subtitle appears in multiple frames, it is recorded only once, but all occurrence periods are annotated). Ultimately, a complete text information set covering the entire video (i.e., the first image text information) is generated. This set includes the content, type, appearance time, and location of all visible text in the video, providing a comprehensive textual basis for subsequent compliance audits.

[0052] In the above method, by extracting single-frame images of the video at a preset frequency and extracting the text information of the images (including content, type and location), and then combining the text information of all images, all image texts in the video can be obtained comprehensively, accurately and completely, providing comprehensive and reliable basic data for subsequent audio text correction, information fusion and compliance review based on image text.

[0053] S103: Extract first audio text information corresponding to the first video data.

[0054] In this embodiment of the present application, "first audio text information" refers to text content extracted from the audio track of the first video data to be reviewed (e.g., a short financial marketing video), primarily using Automatic Speech Recognition (ASR) technology. For details on this process, see steps S301-S303.

[0055] In one embodiment, see Figure 3 , is a flowchart of extracting audio text information provided by an embodiment of the present application, such as Figure 3 As shown, step S103 includes: S301: Extract first audio data from first video data.

[0056] In an embodiment of the present application, extracting first audio data from first video data (such as a financial marketing short video to be reviewed) is essentially a process of separating the audio track in the video and obtaining the original audio material.

[0057] Specifically, the system first downloads the complete video resource of the first video data through the task queue, then calls the video processing tool (such as a component based on FFmpeg) to parse the video, extracts the independent audio stream from the video file, and converts it into a standard audio format (such as WAV, MP3). This separated, processed and stored audio stream is the first audio data.

[0058] S302: transcribe the first audio data into text data corresponding to each of a plurality of time series.

[0059] In an embodiment of the present application, the first audio data extracted from the first video data (i.e., the audio track of the video) is subjected to speech recognition transcription by calling a custom ASR model, and the continuous audio is divided into multiple ordered time segments according to time (e.g., each segment corresponds to the duration of a sentence in the audio), and corresponding text content (e.g., a text record of this sentence) is generated for each time segment, ultimately obtaining a set of structured data of "time series + corresponding text" (i.e., speech transcribed text with timestamps).

[0060] For example, the audio from 00:00 to 00:03 is transcribed as "Welcome to learn about this financial product," and the audio from 00:04 to 00:08 is transcribed as "Annualized interest rate as low as 3.5%." These chronologically arranged text data are the transcription results. Because a custom ASR model is used, recognition of hot financial terms (such as "annualized interest rate") is more accurate.

[0061] S303: Combine the text data according to the time sequence to obtain first audio text information corresponding to the first video data.

[0062] In an embodiment of the present application, text data corresponding to different time segments extracted from the first audio data (the audio track of the first video data) (such as the text corresponding to each sentence transcribed by the ASR model) are integrated and arranged in the order of their actual appearance in the audio, and finally form a text collection that completely records all the audio content of the video, namely the first audio text information.

[0063] The first audio text information not only contains all the text of the audio transcription, but also retains the temporal logic of the content through time series association (for example, a certain text corresponds to the voice in the video period of 00:05-00:10), which can fully reflect the semantics and time distribution of the video audio, and provide a complete basis for the audio dimension for subsequent fusion with image text and compliance review.

[0064] In the above method, by extracting video audio and transcribing it into text of the corresponding time series, and then combining it by time to form audio text information, the temporal logic of the audio content can be completely preserved, providing accurate voice dimension data support for subsequent fusion with image text and compliance review.

[0065] S104: Fusing the first image text information with the first audio text to obtain text fusion information.

[0066] In the embodiment of the present application, although the subtitles recognized by ASR (the first audio text information) can completely correspond to the voice content in a time series, they are limited by the voice clarity and the difficulty of professional vocabulary recognition, and are prone to typos (such as misidentifying "urgently need money" as "before getting married"), and can only cover audio-related text; while the subtitles recognized by image (the first image text information) can extract all visible text in the picture (including dynamic subtitles, fixed risk warnings, logo text, etc.), and the information is richer, but it may contain irrelevant content such as background decorative text.

[0067] Therefore, the text fusion information finally obtained by combining the two through text fusion has both temporal accuracy and content integrity and accuracy, avoiding the defects of a single source.

[0068] In one embodiment, see Figure 4 , is a flow chart of extracting fused text information provided by an embodiment of the present application, such as Figure 4 As shown, step S104 includes: S401: supplement and correct the first audio text information according to the first image text information to obtain supplemented and corrected second audio text information.

[0069] In an embodiment of the present application, the advantages of the first image text information (subtitles, risk warnings, and other text extracted from the video screen) are utilized to optimize the first audio text information (text transcribed by ASR speech): first, the contents of the two in the same time period are aligned through timestamps (such as the audio transcribed text of the video 00:10-00:15 corresponds to the screen subtitles of the same time period); then, the clear and accurate content in the image text is used to supplement the missing or ambiguous parts of the audio text (for example, the "loan amount of up to 200,000" that ASR did not fully recognize is supplemented by the screen subtitles), and at the same time, the recognition errors of the audio text are corrected (such as the "before marriage" misjudged by ASR is corrected to "urgent need money" based on the image subtitles); and finally, the second audio text information is generated.

[0070] The second audio text information generated above not only retains the time series logic of the audio text, but also makes up for the defects of ASR with the help of the accuracy of the image text, becoming more reliable audio-related text data.

[0071] S402: Fusing the first image text information and the second audio text information to obtain text fusion information.

[0072] In an embodiment of the present application, the first image text information (subtitles, risk warnings, logo text, etc. extracted from the video screen, including text type and position) is fused with the second audio text information (audio transcribed text supplemented and corrected by the image text, with a timestamp) to obtain text fusion information, and the process is an integration process achieved through temporal alignment and semantic association.

[0073] Specifically, based on the timestamp of the second audio text, the image text of the same time period (such as the subtitles in the screen from 00:05 to 00:10) is associated with the audio text to ensure that the two match in the time dimension; retain key information in the image text that is not covered by the audio (such as the fixed risk warning "Investment is risky" on the screen, which may not be mentioned in the voice), and use the audio text to supplement the voice content not displayed in the image (such as the "limited-time offer" mentioned in the commentary does not appear in the subtitles); eliminate irrelevant information in the image text (such as background decorative text) and error residues in the audio text that have been corrected by the image text to avoid duplication or invalid content.

[0074] The final generated text fusion information is a more accurate and structured set of complete subtitles, which not only contains complete semantics (voice + screen text) arranged in time series, but also marks the text source type (such as "audio transcription", "image subtitles", "risk warning") and key location information, providing a comprehensive and accurate text basis for subsequent compliance audits (such as checking the consistency of interest rate statements and the completeness of risk warnings).

[0075] In the above method, the audio text information is supplemented and corrected by the image text information and the two are integrated, which can make up for the possible errors in audio transcription, supplement the key content not covered by the audio, form complete and accurate text fusion information, and provide comprehensive and accurate text basis for subsequent data review.

[0076] S105: Perform data review based on the text fusion information to obtain a first review result.

[0077] In the examples of this application, see Figure 5 , is a schematic diagram of the structure of the data audit rules provided in the embodiment of the present application, such as Figure 5 As shown, the system packages core data such as the integrated subtitle information (text fusion information), the image extracted from the video frame (the first image), and the video metadata (basic information such as the delivery channel and duration included in the first video data), and transmits them uniformly to the review rule engine as input material for the review. The rule engine adopts a strategy model design, which means that it encapsulates different review logics (such as interest rate compliance rules, risk warning rules, and illegal vocabulary screening rules) into independent "rule subclasses", each of which corresponds to a specific review scenario (for example, the "Financial Product Interest Rate Review Subclass" specifically verifies the interest rate statement, and the "Risk Warning Integrity Subclass" checks whether the necessary prompts are included). This design allows for added or modified review rules by simply adjusting the corresponding subclass without changing the overall engine framework, which is extremely flexible.

[0078] Based on the text fusion information (structured data that integrates image text and corrected audio text, including text content, type, timestamp and location), the system will call the audit rule engine to call Figure 5 The large model rule checker in the rule subclass A shown here reviews it to obtain the review result (first review result) corresponding to the text fusion information.

[0079] In one embodiment, see Figure 6 , is a flow chart of the data review process provided in the embodiment of the present application, such as Figure 6 In the embodiment, step S105 includes: S501: Detect risk warning information in text fusion information to obtain a first sub-result.

[0080] In the embodiment of the present application, detecting the risk warning information in the text fusion information and obtaining the first sub-result is a special verification process carried out by the audit rule engine for the specific dimension of "risk warning compliance".

[0081] Specifically, the text fusion information incorporates text extracted from the video (such as the "Investment is risky" slogan in the corner) and corrected audio transcriptions (such as the audio mention of "Financial management requires caution"), including the text's type, timestamp, and location information. The rule engine then invokes the "Risk Warning Detection Subclass" (e.g., Rule Subclass A, which utilizes a large text model) to perform validation based on compliance requirements for financial marketing content (e.g., "must include a clear risk warning statement," "risk warnings must be displayed in a fixed area at the end of the frame," and "must not be merely mentioned in voice without being displayed on screen"). This includes checking for the presence of compliant risk warning text (e.g., whether it contains core phrases such as "Investment is risky" and "Financial management requires caution"); verifying the presentation of the risk warning (e.g., whether it is visible in the image, not just in the audio); and confirming compliance with its location and duration (e.g., whether it appears continuously for at least two seconds in the end of the video and is located in a fixed visible area).

[0082] Finally, based on the generated first sub-result, the verification conclusion of the risk warning will be clearly recorded: if it fully meets the requirements, it will be marked as "Risk Warning Compliant" and the relevant text content, location and duration; if there is a missing (such as no risk warning appears), a form violation (such as only a voice mention) or a position inconsistency (such as not displayed in the last frame), it will be marked as "Risk Warning Non-Compliant" and the specific problem as an important part of the first review result.

[0083] S502: Detect user privacy data in the text fusion information to obtain a second sub-result.

[0084] In the embodiment of the present application, detecting user privacy data in text fusion information and obtaining the second sub-result is a special detection process performed by the audit rule engine for the "privacy compliance" dimension, which is a detailed verification item of the first audit result.

[0085] Specifically, text fusion information includes text in the video screen (such as subtitles, background text) and corrected audio transcription text (such as content mentioned in the voice), and user privacy data may exist in it in various forms, such as the ID number, bank card number, mobile phone number appearing in the screen, or "a user's loan record" and "personal contact information" mentioned in the voice.

[0086] The rule engine will call the "privacy data detection subclass" (special rules designed based on the strategy pattern) and conduct detection through two core methods. For example, it uses the preset privacy data format library (such as the 11-digit rule for mobile phone numbers and the 18-digit encoding rule for ID numbers) and combines the text model to identify sensitive information in the text that meets the format characteristics. The second sub-result generated after the detection is completed will be presented in a structured form: if no privacy data is found, it will be marked as "No user privacy information was detected"; if relevant content is detected, the type of privacy data (such as mobile phone number / ID number), specific content (partially desensitized display, such as "138****5678"), source (image text / audio transcription) and corresponding timestamp will be clearly recorded to provide a basis for subsequent privacy compliance judgments.

[0087] S503: Detect product pricing information in the text fusion information to obtain a third sub-result.

[0088] In the embodiment of the present application, detecting the product pricing information in the text fusion information and obtaining the third sub-result is a special verification performed by the audit rule engine on the "compliance of product pricing statements", which is a sub-item of the first audit result.

[0089] In text-integrated information, product pricing information may appear as image text (e.g., "Annualized interest rate 3.5%" in the caption) or as audio transcripts (e.g., "Monthly loan interest rate 0.8%" in the voice). To verify the accuracy of the interest rate expressed in the captions, the system invokes a rule engine, such as the text model of rule subclass A. This combines the text model with a customized annualized interest rate calculation tool, leveraging the model's tool-calling capabilities to achieve accurate verification. The specific process is as follows: The large text model first performs semantic analysis on interest rate-related expressions in subtitles. Whether it is a direct "annualized interest rate of 3.5%" or an indirect non-standard expression such as "daily interest of 1 yuan for 10,000 yuan" or "borrow 10,000 yuan and repay 0.8 yuan every day", the large model can accurately extract the core parameters through natural language understanding: for example, it can identify "principal 10,000 yuan" and "daily interest 1 yuan" from "10,000 yuan daily interest", and locate "principal 10,000 yuan" and "daily repayment interest 0.8 yuan" from "borrow 10,000 yuan and repay 0.8 yuan every day". This kind of analysis is highly robust. Even if the expression form is flexible and there are colloquial expressions (such as "10,000 yuan with a daily interest of 1 yuan"), it can accurately match the key variables required for interest calculation (principal, interest amount, time unit). The large model then calls a customized annualized interest rate calculation tool, passing in the extracted parameters in the tool's required format (for example, "Principal = 10,000, Daily Interest = 1"). The tool automatically calculates the annualized interest rate based on financial formulas (for example, "10,000 yuan daily interest = 1 yuan" corresponds to an annualized interest rate of 3.65%) and returns the result. The large model then compares the calculated result with the interest rate directly stated in the subtitles (if any), or considers regulatory compliance to determine whether the interest rate statement complies with regulations (for example, whether there are issues such as "as low as X%" without indicating applicable conditions). The model ultimately outputs a verification conclusion on the correctness of the interest rate statement, which serves as the core basis for product pricing information review (the third sub-result).

[0090] S504: Obtain a first review result according to the first sub-result, the second sub-result, and the third sub-result.

[0091] In an embodiment of the present application, the first review result is obtained based on the first sub-result (risk warning information detection result), the second sub-result (user privacy data detection result) and the third sub-result (product pricing information detection result), which is a process of summarizing, grading and integrating the results of each special verification.

[0092] In the above method, through special detection and result integration of three core information types, namely risk warnings, user privacy, and product pricing in text fusion information, the first audit results covering key dimensions can be accurately generated, providing a comprehensive and clear basis for compliance judgment at the text level of video data.

[0093] S106: Determine compliance of the first video data according to the first review result.

[0094] In an embodiment of the present application, determining the compliance of the first video data according to the first audit result is a process of comprehensive judgment based on the verification conclusions of various dimensions (risk warnings, user privacy, product pricing, etc.) integrated in the audit result.

[0095] Specifically, the system evaluates the overall conclusion and detailed issues (such as "whether there are serious violations" and "whether the number of general violations exceeds a threshold") in the first review results against preset compliance criteria. If the first review result shows that all sub-results are compliant (no violations), the video data is deemed "compliant." If there are serious violations such as "detection of user privacy data," or if the number of general violations such as "missing risk warnings" or "incorrect interest rate statements" exceeds a set threshold, the video data is deemed "non-compliant" and the appropriate action is triggered based on the violation level (such as direct blocking or labeling requiring manual review). The final compliance conclusion directly determines whether the first video data can enter subsequent dissemination stages.

[0096] In one embodiment, see Figure 7 , is a flowchart of data compliance judgment provided by an embodiment of the present application, such as Figure 7 As shown, step S106 includes: S601: Detect prohibited content in a first image to obtain a fourth sub-result.

[0097] In an embodiment of the present application, the violation review of the first video data includes not only the compliance review of the text fusion information, but also the review of multiple image frames.

[0098] Specifically, a multimodal large-scale model checker can be used to detect content containing violence, vulgarity, terror, threats, and other objectionable content. This model leverages the model's ability to understand visual features and semantics to accurately identify objectionable content. A multimodal large-scale model (such as the Qwen2.5-VL or GPT-4V, which have image and text understanding capabilities) receives a keyframe (the first image) extracted from the first video data. It first extracts low-level visual features, identifying elements such as character actions (e.g., fighting, violent gestures), scene atmosphere (e.g., gory scenes, dark and terrifying scenes), and clothing or objects (e.g., vulgar and revealing clothing, horror props). Combined with semantic features learned during model training (e.g., "violence" corresponds to scenes of physical conflict and weapon threats, and "vulgarity" corresponds to revealing clothing and indecent gestures), it performs a comprehensive semantic assessment of the image content, resulting in the fourth sub-result.

[0099] S602: Detect compliance of the logo content contained in the first image to obtain a fifth sub-result.

[0100] In the embodiment of the present application, the logo content in the first image may include a company logo (such as a financial institution brand logo), which can be obtained by a self-training model such as Figure 5 The Logo model corresponding to the rule subclass N shown, namely the visual model rule checker, verifies whether the logo is the latest version of the company (to avoid misuse of the old version of the logo), and at the same time checks whether there is any unauthorized use of third-party logos (such as impersonating the logo of a well-known financial institution). The fifth sub-result generated after the detection is completed will be presented in a structured form: if all logos comply with the regulations, it will be marked as "Logo content compliance"; if there is a violation (such as using an old version of the logo or a false qualification logo), the type of the illegal logo will be clearly recorded (such as "Logo version mismatch" or "False qualification logo").

[0101] In addition to logo detection, the fifth sub-result also uses a multimodal model to identify whether the tail frame structure complies with financial advertising standards. The multimodal model combines visual features (layout, elements) and text semantics (content completeness) to comprehensively score the compliance of the tail frame structure. If the tail frame contains a complete risk warning, is prominently positioned, and remains for a sufficient duration, it is judged as "compliant with standards." If the risk warning is obscured or remains for less than 3 seconds, it is judged as "non-compliant with standards" and the specific issues are clearly marked (such as "risk warning location is not prominent" or "staying for insufficient time"). Furthermore, content must be screened for violations of core socialist values ​​or ethical and moral standards.

[0102] S603, obtaining a second audit result according to the fourth sub-result and the fifth sub-result.

[0103] In this embodiment, the second review result is derived from the fourth sub-result (the first image's prohibited content detection result) and the fifth sub-result (the first image's logo content compliance detection result). If both sub-results are consistent, the second review result is labeled "Image content and logo compliance." If a violation is detected, the overall judgment is clearly stated (e.g., "Serious violation" or "Only general violation"), along with a list of specific issues (e.g., "Fourth sub-result: Fake qualification image; Fifth sub-result: Use of old logo"), along with corresponding original evidence such as frame timestamps and image positions.

[0104] S604: Determine compliance of the first video data according to the first review result and the second review result.

[0105] In the embodiment of the present application, the compliance of the first video data is determined based on the first review result (text dimension, covering risk warnings, user privacy, product pricing, etc.) and the second review result (image dimension, covering prohibited content, logo compliance, etc.). This is a comprehensive assessment of the video's "text + image" dual-dimensional review information, and ultimately forms a complete compliance judgment conclusion.

[0106] In the above method, by detecting prohibited content and logo compliance in the image and integrating the results to generate a second review result, and then combining it with the first review result at the text level, the compliance of the video data can be comprehensively determined from the "image + text" dual dimensions, providing an accurate and complete judgment basis for the compliance review of video content.

[0107] In one embodiment, step S604 includes: If the first review result or the second review result indicates that the review has failed, the first video data is non-compliant; if both the first review result and the second review result indicate that the review has passed, the first video data is compliant.

[0108] In an embodiment of the present application, if the first review result (text dimension) shows "review failed" (such as the lack of risk warnings, user privacy leakage and other problems), regardless of whether the second review result (image dimension) is passed, the first video data is directly judged as "non-compliant". If the second review result (image dimension) shows "review failed" (such as the presence of prohibited content, logo violations and other problems), regardless of whether the first review result is passed, the first video data is also judged as "non-compliant". Only when both the first review result and the second review result show "review passed" (no violations in the text dimension and no violations in the image dimension) will the first video data be finally judged as "compliant".

[0109] The above method clarifies the logic that "if either the text or image dimension fails the review, it is considered non-compliant, and if both dimensions pass, it is considered compliant." This allows for a strict and efficient final determination of video data compliance, ensuring that only videos with both text and images meeting the requirements are deemed compliant, providing a clear and rigorous standard for compliance control of video content.

[0110] In one embodiment, the data audit method further includes: Obtain second information; wherein the second information includes data information that has not been reviewed and location information corresponding to the data information that has not been reviewed in the first video data; generate a data review report based on the second information and issue an alarm.

[0111] In this embodiment, the second information is extracted based on the premise that either the first or second audit results failed, including the data information and corresponding location information for the data that failed the audit. The data audit report is a structured presentation of the second information, including the video ID, audit time, audit type (financial advertising audit), a clear indication of "Failed audit, first video data non-compliant," a detailed list of violations, and suggestions for additions and corrections.

[0112] It should also be noted that to ensure the stable execution of audit tasks, the system has designed the following fault-tolerant mechanisms, including a retry mechanism. If the model call fails or times out, the system can retry a specified number of times according to the configuration; the system also includes health monitoring and alarms. If the error rate of a model or interface continues to increase, an alarm will be automatically sent to the operation and maintenance system to assist in timely repairs.

[0113] In the above method, by obtaining the data that has not passed the review and its location information, generating reports and alarms, it can accurately locate the illegal content, intuitively present the review results, and promptly remind relevant personnel to handle it, providing efficient support for the rectification and risk prevention and control of illegal videos.

[0114] See also Figure 8, is a schematic diagram of the overall structure of the data review process provided in an embodiment of the present application. As shown in the figure, the steps include: 1) The caller submits an audit request and the audit data is added to the task queue The caller submits the video or image link to be delivered to the system. The system encapsulates it as a review task and adds it to the task queue based on information such as the video delivery channel (such as Douyin, Tencent) and content type.

[0115] 2) Download videos Each Pod acts as a queue consumer, responsible for pulling tasks and downloading video or image resources.

[0116] 3) Scene segmentation to generate multiple video frames For video tasks, the system identifies all scenes and extracts the last frame scene (about 2 seconds) from them, and then extracts frames at a specified frequency (such as every 0.5 seconds) to generate an image sequence for review (the first image).

[0117] 4) Subtitle fusion to obtain fused subtitles The multimodal model extracts image subtitle information (first image text), and the ASR model extracts audio subtitle information (first audio text information). The image subtitle information is then used to correct the audio subtitle information, and the corrected audio subtitle information is fused with the image subtitle information to obtain fused subtitle information (text fusion information).

[0118] 5) Multi-rule data audit The system uniformly transmits subtitle information (fused text), frame images (first image), and video metadata (first video data) to the review rules engine. The rules engine adopts a strategy model, supports different review rules for different subclasses, and supports parallel review.

[0119] The thread pool contains multiple types of audit rules, including visual rule checkers, multimodal rule checkers, large model rule checkers, and regularization rule checkers. Multiple types of audit rules use different audit rules for different data to comprehensively determine the compliance of video data.

[0120] This application applies to compliance reviews of short videos or images for financial marketing. This method receives video or image links and metadata, encapsulates them into tasks, and processes them in a review process within a Kubernetes cluster. It extracts information by extracting frames and audio, then uses multimodal and ASR models to extract text. This text model is then integrated and completed to generate unified subtitles. This method then combines various review rules, calls relevant models, performs analysis, and outputs review conclusions. It also features distributed execution, multi-rule parallelism, fault tolerance, and monitoring. Its core advantage lies in its ability to integrate multi-source information and reason with large models.

[0121] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0122] Figure 9 This is a schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Figure 9 As shown, the terminal device 9 of this embodiment includes: at least one processor 90 ( Figure 9 Only one is shown in the figure) a processor, a memory 91, and a computer program 92 stored in the memory 91 and executable on at least one processor 90. When the processor 90 executes the computer program 92, the steps of any of the above-mentioned data audit method embodiments are implemented.

[0123] The terminal device can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 9 It is only an example of the terminal device 9 and does not constitute a limitation on the terminal device 9. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0124] The processor 90 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0125] In some embodiments, the memory 91 may be an internal storage unit of the terminal device 9, such as the terminal device 9's hard drive or memory. In other embodiments, the memory 91 may also be an external storage device of the terminal device 9, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the memory 91 may include both the terminal device 9's internal storage unit and an external storage device. The memory 91 is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of computer programs. The memory 91 may also be used to temporarily store data that has been output or is about to be output.

[0126] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0127] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0128] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. Computer-readable media can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunications signals.

[0129] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0130] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0131] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal devices and methods can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0132] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0133] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A data audit method, characterized in that: The method comprises: Acquire first video data; extracting first image text information corresponding to the first video data; extracting first audio text information corresponding to the first video data; Fusing the first image text information with the first audio text to obtain text fusion information; Performing data review based on the text fusion information to obtain a first review result; The compliance of the first video data is determined according to the first review result.

2. The data audit method according to claim 1, wherein: The obtaining of first image text information corresponding to the first video data includes: Extracting a single frame from the first video data according to a preset frequency to obtain a plurality of first images; For each of the first images, extract first information corresponding to the first image; wherein the first information includes first text content, first text type, and a first text position corresponding to the first text content; The first information corresponding to each of the plurality of first images is combined to obtain the first image text information corresponding to the first video data.

3. The data audit method according to claim 1, wherein: The extracting the first audio text information corresponding to the first video data includes: extracting first audio data from the first video data; Transcribing the first audio data into text data corresponding to each of a plurality of time series; The text data are combined according to the time sequence to obtain first audio text information corresponding to the first video data.

4. The data audit method according to claim 1, wherein: The fusing the first image text information with the first audio text to obtain text fusion information includes: Performing text supplementation and correction on the first audio text information according to the first image text information to obtain supplemented and corrected second audio text information; The first image text information and the second audio text information are fused to obtain the text fusion information.

5. The data audit method according to claim 4, wherein: The data review based on the text fusion information is performed to obtain a first review result, including: Detecting risk warning information in the text fusion information to obtain a first sub-result; Detecting user privacy data in the text fusion information to obtain a second sub-result; detecting product pricing information in the text fusion information to obtain a third sub-result; The first audit result is obtained according to the first sub-result, the second sub-result and the third sub-result.

6. The data audit method according to claim 2, wherein: Determining the compliance of the first video data according to the first review result includes: detecting prohibited content in the first image to obtain a fourth sub-result; detecting compliance of identification content contained in the first image to obtain a fifth sub-result; Obtaining a second audit result according to the fourth sub-result and the fifth sub-result; The compliance of the first video data is determined based on the first audit result and the second audit result.

7. The data audit method according to claim 6, wherein: The determining the compliance of the first video data according to the first review result and the second review result includes: If the first review result or the second review result indicates that the review has failed, the first video data is non-compliant; If both the first review result and the second review result indicate that the review is passed, the first video data is compliant.

8. The data audit method according to claim 7, wherein: The method further comprises: Acquire second information; wherein the second information includes data information that has not been reviewed and location information corresponding to the data information that has not been reviewed in the first video data; Generate a data audit report based on the second information and issue an alarm prompt.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Subject teaching and training video auditing method, device, equipment and medium

    CN114842385A

  • Short video auditing method based on multiple modes

    CN115512259A

  • Video data multi-mode compliance detection method, storage medium and compliance detection equipment

    CN116208802A

  • Video auditing method based on multi-modal large model

    CN118968380A

  • Video content text generation method and device, equipment, storage medium and program product

    CN119629380A