Data auditing method, terminal device, and computer-readable storage medium

By extracting and fusing image and audio text information from financial short videos, and combining multimodal big data models and review rules, the problem of low review accuracy in existing technologies has been solved, enabling efficient and accurate compliance judgment of video content.

CN120751175BActive Publication Date: 2025-12-12SHENZHEN XIAOYING INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511204752.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-12-12
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

The current technology for reviewing financial short videos has low accuracy and is difficult to meet complex and diverse compliance requirements. This is mainly because relying on rule engines and single-modal models cannot achieve accurate semantic reasoning.

Method used

By extracting image and audio text information from video data and fusing them into text fusion information, and combining it with multiple review rules for analysis, a multimodal large model is used for review, including multi-rule parallel and distributed execution.

Benefits of technology

It improves the accuracy of short video review, enabling a comprehensive and accurate assessment of video content compliance and ensuring that video data meets relevant standards and requirements before being released.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751175B_ABST
    Figure CN120751175B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of data auditing, and provides a data auditing method, a terminal device and a computer readable storage medium, comprising: acquiring first video data; extracting first image text information corresponding to the first video data; extracting first audio text information corresponding to the first video data; fusing the first image text information and the first audio text to obtain text fusion information; performing data auditing based on the text fusion information to obtain a first auditing result; and determining compliance of the first video data according to the first auditing result. The above method can improve the auditing accuracy of short video data and meet complex and multiple video compliance requirements.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data auditing, and particularly relates to a data auditing method, a terminal device and a computer readable storage medium. BACKGROUND

[0002] With the development of multi-channel marketing scenarios, financial short videos are also widely used and become an important form of marketing promotion for financial institutions. Before being put into use, the financial short videos need to pass compliance auditing to meet the compliance requirements in the financial field.

[0003] In related technologies, a rule engine and a single modal model are mostly relied on, which cannot realize accurate semantic reasoning, resulting in low auditing accuracy of the financial short videos and being difficult to meet the complex and multiple compliance requirements of the financial videos. SUMMARY

[0004] The embodiments of the application provide a data auditing method, device, terminal device and storage medium, which can improve the auditing accuracy of short videos and meet the complex and multiple compliance requirements of videos.

[0005] In a first aspect, the embodiments of the application provide a data auditing method, comprising:

[0006] obtaining first video data;

[0007] extracting first image text information corresponding to the first video data;

[0008] extracting first audio text information corresponding to the first video data;

[0009] fusing the first image text information and the first audio text to obtain text fusion information;

[0010] performing data auditing based on the text fusion information to obtain a first auditing result;

[0011] determining compliance of the first video data according to the first auditing result.

[0012] In the embodiments of the application, the image text information and the audio text information in the video data are extracted, and the image text information and the audio text information are fused into text fusion information. The text fusion information integrates the advantages of images and audio, can make up for the defects of single text extraction, can improve the completeness and accuracy of the auditing information, can adapt to the multi-modal content characteristics of the video data, can obtain more accurate and comprehensive text fusion information, provides a high-quality basis for auditing, and finally performs auditing based on the fused text fusion information, can more comprehensively and accurately judge the compliance of the video content, improves the reliability of the auditing result, and helps to ensure that the video data meets the relevant specifications and requirements before being put into use.

[0013] In a possible implementation manner of the first aspect, the first image text information corresponding to the first video data is extracted, including:

[0014] According to the preset frequency, a single frame of picture in the first video data is extracted to obtain a plurality of first images;

[0015] For each first image, first information corresponding to the first image is extracted; wherein the first information includes first text content, a first text type, and a first text position corresponding to the first text content;

[0016] The first information corresponding to each of the plurality of first images is combined to obtain the first image text information corresponding to the first video data.

[0017] In the embodiments of the present application, by extracting a single frame of picture of a video according to a preset frequency and extracting picture text information (including content, type and position), and then combining the text information of all pictures, all image texts in the video can be comprehensively, accurately and completely obtained, thereby providing comprehensive and reliable basic data for subsequent audio text correction, information fusion and compliance audit based on image texts.

[0018] In a possible implementation manner of the first aspect, the first audio text information corresponding to the first video data is extracted, including:

[0019] The first audio data is extracted from the first video data;

[0020] The first audio data is transcribed into text data corresponding to a plurality of time sequences respectively;

[0021] The text data is combined according to the time sequence to obtain the first audio text information corresponding to the first video data.

[0022] In the embodiments of the present application, by extracting a video audio and transcribing it into text corresponding to a time sequence, and then combining the audio text information according to time, the time sequence logic of the audio content can be completely retained, thereby providing accurate voice dimension data support for subsequent image text fusion and compliance audit.

[0023] In a possible implementation manner of the first aspect, the first image text information and the first audio text are fused to obtain text fusion information, including:

[0024] According to the first image text information, the first audio text information is supplemented and corrected to obtain second audio text information after supplementation and correction;

[0025] The first image text information and the second audio text information are fused to obtain the text fusion information.

[0026] In the embodiments of the present application, the image text information is supplemented and corrected to the audio text information, and the two are fused, so as to make up for possible errors in audio transcription, supplement key content not covered by audio, form complete and accurate text fusion information, and provide comprehensive and accurate text basis for subsequent data auditing.

[0027] In a possible implementation of the first aspect, the data auditing based on the text fusion information obtains a first auditing result, including:

[0028] detecting risk prompt information in the text fusion information obtains a first sub-result;

[0029] detecting user privacy data in the text fusion information obtains a second sub-result;

[0030] detecting product pricing information in the text fusion information obtains a third sub-result;

[0031] the first auditing result is obtained according to the first sub-result, the second sub-result and the third sub-result.

[0032] In the embodiments of the present application, through special detection and result integration of three types of core information of risk prompt, user privacy and product pricing in the text fusion information, the first auditing result covering key dimensions can be accurately generated, and comprehensive and explicit basis for compliance judgment of the video data at the text level is provided.

[0033] In a possible implementation of the first aspect, the compliance of the first video data is determined according to the first auditing result, including:

[0034] detecting prohibited content in the first image obtains a fourth sub-result;

[0035] detecting compliance of identification type content contained in the first image obtains a fifth sub-result;

[0036] the second auditing result is obtained according to the fourth sub-result and the fifth sub-result;

[0037] the compliance of the first video data is determined according to the first auditing result and the second auditing result.

[0038] In the embodiments of the present application, the second auditing result is generated by detecting prohibited content and identification compliance in the image and integrating the results, and then the first auditing result at the text level is combined, so that the compliance of the video data can be comprehensively judged from the “image + text” dual dimensions, and accurate and complete judgment basis for video content compliance auditing is provided.

[0039] In a possible implementation of the first aspect, the compliance of the first video data is determined according to the first auditing result and the second auditing result, including:

[0040] If the first audit result or the second audit result indicates that the audit fails, the first video data is non-compliant;

[0041] If the first audit result and the second audit result both indicate that the audit passes, the first video data is compliant.

[0042] In the embodiments of the present application, by explicitly determining that "any dimension of text or image audit fails, and only both dimensions pass to determine compliance", the final determination of the compliance of the video data can be strictly and efficiently completed, and only the video whose text and image both meet the requirements is determined to be compliant, thereby providing a clear and rigorous standard for the compliance control of the video content.

[0043] In a possible implementation form of the first aspect, the method further includes:

[0044] obtaining second information, wherein the second information includes data information that fails the audit and position information of the data information that fails the audit in the first video data;

[0045] generating a data audit report according to the second information and performing an alarm prompt.

[0046] In the embodiments of the present application, by obtaining the data that fails the audit and the position information thereof, generating a report and performing an alarm, the non-compliant content can be accurately located, the audit result can be intuitively presented, and relevant personnel can be timely reminded to handle, thereby providing efficient support for rectification and risk prevention and control of non-compliant videos.

[0047] In a second aspect, the embodiments of the present application provide a terminal device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the data audit method of any one of the first aspect when executing the computer program.

[0048] In a third aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the data audit method of any one of the first aspect.

[0049] In a fourth aspect, the embodiments of the present application provide a computer program product, which, when executed on a terminal device, causes the terminal device to execute the data audit method of any one of the first aspect.

[0050] It can be understood that the beneficial effects of the above-mentioned second aspect to fourth aspect can be referred to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without any creative effort.

[0052] Figure 1 is a flowchart of a data auditing method provided by an embodiment of the present application;

[0053] Figure 2 is a flowchart of extracting image text information provided by an embodiment of the present application;

[0054] Figure 3 is a flowchart of extracting audio text information provided by an embodiment of the present application;

[0055] Figure 4 is a flowchart of extracting fusion text information provided by an embodiment of the present application;

[0056] Figure 5 is a structural diagram of a data auditing rule provided by an embodiment of the present application;

[0057] Figure 6 is a flowchart of a data auditing process provided by an embodiment of the present application;

[0058] Figure 7 is a flowchart of data compliance judgment provided by an embodiment of the present application;

[0059] Figure 8 is a general structural diagram of a data auditing process provided by an embodiment of the present application;

[0060] Figure 9 is a structural diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0061] In the following description, for the purpose of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.

[0062] It will be understood that the term “includes,” “including,” “has,” “having” or “has” when used in this application and the appended claims indicates non-exclusive inclusion, such that a process, method, article, or apparatus that comprises several elements does not include those elements solely “including” the several elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Other words of similar meaning, in connection with determining, detecting, or including, have similar meanings.

[0063] It should also be understood that the term “and / or” when used in the specification and the appended claims, means any one and / or any combination of the associated listed items can be present or added.

[0064] As used in this application and the appended claims, the term “if’ can be construed to mean “when” or “once” or “in response to determining” or “in response to detecting” depending on the context. Similarly, the phrase “if it is determined” or “if [a described condition or event] is detected” can be construed to mean “once it is determined” or “in response to determining” or “once [the described condition or event] is detected” or “in response to detecting [the described condition or event]” depending on the context.

[0065] In addition, in the description of the application and the appended claims, the terms “first”, “second”, “third”, etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0066] The reference in the specification to “one embodiment” or “some embodiments” or the like means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearance of the phrases “in one embodiment”, “in some embodiments”, “in other some embodiments”, “in yet some embodiments” or the like in various places in the specification is not necessarily all referring to the same embodiment, but means “one or more but not all embodiments”, unless otherwise specifically stated.

[0067] With the development of multi-channel marketing scenarios, financial short videos are also widely used, becoming an important form of marketing and promotion for financial institutions. Before being put into use, they need to go through compliance review to meet the compliance requirements of the financial field.

[0068] In the related art, it is mostly dependent on rule engines and single-modal models, which cannot achieve accurate semantic reasoning, resulting in low accuracy of the review of financial short videos, and it is difficult to meet the complex and diverse compliance requirements of financial videos.

[0069] To solve the problems in the related art, the embodiments of the present application provide a data auditing method, a terminal device and a computer readable storage medium, and the present application is suitable for financial marketing short video or picture compliance auditing. The method receives video or picture links and metadata, encapsulates them into tasks, and processes them by an auditing process in a Kubernetes cluster; information is obtained by frame extraction, audio extraction, etc., text is extracted by using a multi-modal large model and an ASR model, and a unified subtitle is generated by text large model fusion and completion; then, in combination with various auditing rules, relevant models are called for analysis, and an auditing conclusion is output. At the same time, it has mechanisms such as distributed execution, multi-rule parallelism, fault tolerance and monitoring, and the core advantage is the fusion of multi-source information and large model reasoning capability. The above method can improve the auditing accuracy of short videos and meet the complex and multi-element video compliance requirements.

[0070] Referring to Figure 1 is a flowchart of the data auditing method provided by the embodiments of the present application, as an example but not limitation, the method can include the following steps:

[0071] S101, first video data is obtained.

[0072] In the embodiments of the present application, the first video data is submitted by the calling party, and is stored in the task queue and the database after being encapsulated by the system, as one of the basic information for the subsequent auditing process. In the financial field, the first video data mainly refers to the basic information for financial marketing short video or picture compliance auditing, including video or picture links, distribution channels, content types and other metadata. The calling party submits the video or picture links and their metadata information (distribution channels, content types, etc.) to the auditing system, and the system encapsulates these information into auditing tasks (Task) after receiving them, and adds them to the task queue, and records the related metadata, creation time and auditing status of the task in the database.

[0073] The system as a whole runs on Kubernetes (a container orchestration platform), and all audit-related computing resources (such as Pods) are managed by Kubernetes. When the task volume surges (such as the centralized launch of financial marketing videos), when the task volume decreases, and the idle Pods increase, the excess Pods are automatically destroyed (such as being reduced to 2), avoiding resource waste, and when the queue backlog exceeds the threshold, Kubernetes will automatically create new Pods (such as expanding from 3 to 10). A Pod is an independent task processing unit, which internally runs an "audit process" containing a multi-thread pool (such as 10 threads), which allows a single Pod to handle multiple audit tasks simultaneously (such as thread 1 handling task A download, thread 2 handling task B frame extraction, and thread 3 handling task C subtitle fusion), rather than processing them one by one in sequence, greatly improving the task processing efficiency of a single Pod.

[0074] The Pod listens to the task queue through the threads in the thread pool, and when a "to-be-processed" task is detected, it sends a request to the queue and obtains the task (the queue marks the task as "picked up" to prevent repeated processing). The thread parses the "video / image link" in the task, calls the built-in download tool (supports HTTP / HTTPS protocol) to download resources (first video data) from the corresponding platform (such as Douyin, Tencent Advertising), and verifies the link validity (such as whether it is expired or not, and whether the format is compliant) during download.

[0075] S102, extract first image text information corresponding to the first video data.

[0076] In the embodiments of the present application, the "first image text information" refers to all visible text information extracted by identifying the image frames of the first video data (to-be-audited financial marketing short videos). These texts are directly derived from video pictures and are one of the core bases for video content compliance audit.

[0077] In one embodiment, referring to Figure 2 is a flowchart of extracting image text information provided by the embodiments of the present application, as Figure 2 shown, step S102 includes:

[0078] S201, extracting a single frame picture in the first video data according to a preset frequency to obtain a plurality of first images.

[0079] In the embodiments of the present application, "extracting single-frame pictures in the first video data to obtain a plurality of first images" is the basic link of video content analysis, and the core is to extract image frames from the financial marketing short video (first video data) to be audited at fixed intervals, to provide visual materials for subsequent text recognition and picture compliance audit.

[0080] Specifically, the system obtains the first video data (containing a video link) from the task queue, downloads the video resource to obtain the first video data, and then performs frame extraction processing on the video, extracts image frames at a specified frequency, i.e., a preset frequency (such as 1 frame per 0.5 seconds), and focuses on extracting tail frames (about 2 seconds of content, financial videos often place compliance statements here), to obtain an image sequence (i.e., "first images") for recognition.

[0081] S202, for each first image, extracting first information corresponding to the first image; wherein the first information includes first text content, first text type and first text content corresponding first text position.

[0082] In the embodiments of the present application, the single-frame images (first images) obtained by video frame extraction are subjected to text information (first information) extraction, which essentially identifies the text in the image through technical means, and classifies and positions the text, to provide structured data for subsequent compliance audit.

[0083] Extracting text content (first text content) in each image (first image) refers to the actual text content identified from the first image, which is the "semantic itself" of the text. If the first image is a video picture of interest rate promotion, it may extract "annual interest rate as low as 3.5%", and if there is a risk prompt in the corner of the picture, it may extract "investment has risks, be careful when entering the market".

[0084] After extracting the text content corresponding to each image, the text content is classified and classified (first text type) to distinguish the nature and purpose of the text. Common types in financial video data include caption text, i.e., the commentary text (such as product introduction caption) displayed dynamically in the video; Logo text, i.e., the text (such as brand name "XX Bank") attached to the financial institution Logo; other text: such as description text ( "borrowing limit up to 200,000" ) in product promotion picture, etc. After classifying the text content, the specific position of the text in the picture (first text position) is located.

[0085] Specifically, the first information (first text content, first text type, and first text position) can be completed by using a multi-modal large model (such as Qwen2.5-VL). The multi-modal large model receives the first image as input, locates the text region in the picture through image recognition technology, performs OCR (optical character recognition) on the text region, obtains the "first text content", automatically classifies the "first text type" based on the position, style, and semantics of the text (such as "risk prompt" is usually located in a fixed corner and the font is small), synchronously records the coordinates or region of the text in the image to generate the "first text position", and finally outputs in a structured format (such as JSON).

[0086] S203, combining the first information corresponding to each of the plurality of first images to obtain first image text information corresponding to the first video data.

[0087] In the embodiment of the application, the first information of all the first images is integrated in time sequence (frame extraction order), and the repeated text content is de-duplicated and merged (such as the same subtitle appears in multiple frames only once, but all the appearance time periods are recorded). Finally, a complete text information set covering the entire video (i.e., the first image text information) is generated. The set contains the content, type, appearance time, and position of all visible text in the video, providing comprehensive text basis for subsequent compliance review.

[0088] In the above method, by extracting single-frame pictures of the video at a preset frequency and extracting picture text information (including content, type, and position), and then combining the text information of all pictures, all image texts in the video can be comprehensively, accurately, and completely obtained, providing comprehensive and reliable basic data for subsequent audio text correction, information fusion, and compliance review based on image text.

[0089] S103, extracting first audio text information corresponding to the first video data.

[0090] In the embodiment of the application, the "first audio text information" refers to the text content extracted from the audio track of the first video data (such as a financial marketing short video) to be audited, which is mainly realized by automatic speech recognition (ASR) technology. For specific implementation process, see steps S301-S303.

[0091] In one embodiment, referring to Figure 3 is a flowchart of extracting audio text information provided by the embodiment of the application, as shown in Figure 3 S103 includes:

[0092] S301, extracting first audio data from the first video data.

[0093] In the embodiments of the present application, the first audio data is extracted from the first video data (such as the financial marketing short video to be audited), which is essentially the process of separating the audio track in the video and obtaining the original audio material.

[0094] Specifically, the system first downloads the complete video resource of the first video data through the task queue, and then calls the video processing tool (such as the FFmpeg-based component) to analyze the video, and strips the independent audio stream from the video file, and converts it to a standard audio format (such as WAV, MP3). The audio stream separated, processed and stored is the first audio data.

[0095] S302, the first audio data is transcribed into a plurality of time sequence respectively corresponding text data.

[0096] In the embodiments of the present application, the first audio data (i.e. the audio track of the video) extracted from the first video data is transcribed by calling the custom ASR model for speech recognition, the continuous audio is divided into a plurality of ordered time segments (such as the duration of a sentence in each segment of audio), and the corresponding text content (such as the text record of the sentence) is generated for each time segment, and finally a set of structured data (i.e. time-stamped speech transcription text) of “time sequence + corresponding text” is obtained.

[0097] For example, the speech in the 00:00-00:03 period of the audio is transcribed as “Welcome to learn about this financial product”, and the 00:04-00:08 period is transcribed as “The annual interest rate is as low as 3.5%”. These text data arranged in chronological order are the transcription results. Since the custom ASR model is used, the recognition of hot words in the financial field (such as “annual interest rate”) will be more accurate.

[0098] S303, the text data is combined according to the time sequence to obtain the first audio text information corresponding to the first video data.

[0099] In the embodiments of the present application, the text data (such as the text corresponding to each sentence of speech transcribed by the ASR model) corresponding to different time segments extracted from the first audio data (the audio track of the first video data) is integrated and arranged according to the actual time sequence of the audio, and finally a complete set of text recording all audio content of the video, i.e. the first audio text information, is formed.

[0100] The first audio text information not only contains all the text of audio transcription, but also retains the time sequence logic of the content (such as a certain text corresponding to the voice of the video 00:05-00:10 period) through time sequence association, and can completely reflect the semantic and time distribution of the video audio, providing a complete basis for subsequent image text fusion, compliance audit in the audio dimension.

[0101] In the above method, by extracting the video audio and transcribing it into text corresponding to the time sequence, and then combining the audio text information according to time, the time sequence logic of the audio content can be completely retained, providing accurate voice dimension data support for subsequent image text fusion and compliance audit.

[0102] S104, the first image text information and the first audio text are fused to obtain text fusion information.

[0103] In the embodiments of the present application, although the subtitles (first audio text information) recognized by ASR can completely correspond to the voice content according to the time sequence, it is limited by the clarity of the voice and the difficulty of recognizing professional vocabulary, and is prone to have errors (such as misrecognizing “urgent need for money” as “before marriage”), and can only cover audio-related text; while the subtitles (first image text information) recognized by image recognition can extract all visible text in the picture (including dynamic subtitles, fixed risk prompts, logo text, etc.), the information is more abundant, but may contain irrelevant content such as background decoration text.

[0104] Therefore, by text fusion, the final text fusion information obtained by combining the two has both time accuracy and content integrity and accuracy, avoiding the defects of a single source.

[0105] In one embodiment, referring to Figure 4 is a flowchart of extracting and fusing text information provided by the embodiments of the present application, as Figure 4 indicated, step S104 includes:

[0106] S401, according to the first image text information, text supplement and correction is performed on the first audio text information to obtain second audio text information after supplement and correction.

[0107] In the embodiments of the present application, the advantages of the first image text information (caption, risk prompt, etc. extracted from the video screen) are used to optimize the first audio text information (ASR speech transcription text): first, the content of the two in the same period is aligned through the timestamp (such as the audio transcription text of 00:10-00:15 of the video corresponds to the screen caption in the same period); then the clear and accurate content in the image text is used to supplement the missing or blurred part in the audio text (for example, the ASR does not recognize the complete "borrowing limit is 200,000", which is completed by the screen caption), and the recognition error of the audio text is corrected (such as the ASR misjudgment "before marriage" is corrected to "urgent money" according to the image caption); and finally the second audio text information is generated.

[0108] The generated second audio text information not only retains the time sequence logic of the audio text, but also makes up for the defects of ASR with the accuracy of the image text, and becomes more reliable audio-related text data.

[0109] S402, the first image text information and the second audio text information are fused to obtain text fusion information.

[0110] In the embodiments of the present application, the first image text information (caption, risk prompt, logo text, etc. extracted from the video screen, including text type and position) and the second audio text information (audio transcription text supplemented and corrected by image text, with timestamp) are fused to obtain text fusion information, and the process is an integration process realized by time sequence alignment and semantic association.

[0111] Specifically, the image text (such as the caption in the 00:05-00:10 screen) in the same period is associated with the audio text based on the timestamp of the second audio text, to ensure that the two match in the time dimension; the key information in the image text that is not covered by the audio (such as the fixed risk prompt "investment has risks" that may not be mentioned in the voice) is retained, and the voice content not displayed in the image (such as "limited time offer" mentioned in the commentary that does not appear in the caption) is supplemented by the audio text; irrelevant information in the image text (such as background decoration text) and error residues in the audio text that have been corrected by the image text are removed, to avoid repeated or invalid content.

[0112] The finally generated text fusion information is a more accurate and structured complete caption, which contains complete semantics arranged in time sequence (voice + screen text) and also labels the text source type (such as "audio transcription", "image caption", "risk prompt") and key position information, to provide comprehensive and accurate text basis for subsequent compliance audit (such as checking the consistency of interest rate expression and the completeness of risk prompt).

[0113] In the above method, the audio text information is supplemented and corrected by the image text information, and the two are fused, which can make up for possible errors in audio transcription, supplement key content not covered by audio, form complete and accurate text fusion information, and provide comprehensive and accurate text basis for subsequent data auditing.

[0114] S105, auditing data based on the text fusion information to obtain a first auditing result.

[0115] In the embodiments of the present application, referring to Figure 5 , a structural diagram of the data auditing rule provided by the embodiments of the present application is shown, as Figure 5 indicated, the system will package the integrated subtitle information (text fusion information), pictures obtained by video frame extraction (first images), video metadata (first video data including basic information such as delivery channel and duration), and other core data, and uniformly transmit them to the auditing rule engine as input materials for auditing. The rule engine is designed using the strategy pattern, which means that it encapsulates different auditing logics (such as interest rate compliance rules, risk prompt rules, and illegal word screening rules) into independent “rule subclasses”, and each subclass corresponds to a specific auditing scenario (for example, the “financial product interest rate auditing subclass” is specifically used to check interest rate expressions, and the “risk prompt integrity subclass” is used to check whether necessary prompt statements are included). This design allows new or modified auditing rules to be adjusted by corresponding subclasses without changing the overall framework of the engine, which is highly flexible.

[0116] Based on the text fusion information (which integrates structured data of image text and corrected audio text, including text content, type, timestamp, and position), the system will call the auditing rule engine to perform auditing using a large model rule checker in rule subclass A as shown in Figure 5 to obtain an auditing result corresponding to the text fusion information (first auditing result).

[0117] In one embodiment, referring to Figure 6 , a flowchart of the data auditing process provided by the embodiments of the present application is shown, as Figure 6 indicated, step S105 includes:

[0118] S501, detecting risk prompt information in the text fusion information to obtain a first sub-result.

[0119] In the embodiments of the present application, detecting risk prompt information in the text fusion information and obtaining a first sub-result is a special verification process carried out by the auditing rule engine for the specific dimension of “risk prompt compliance”.

[0120] Specifically, the text fusion information has integrated the text extracted from the video picture (such as the corner "investment has risks" slogan) and the corrected audio transcription text (such as the voice mentioned "financial caution"), and contains the type, timestamp and location information of these texts. The rule engine calls the "risk prompt detection subclass" (such as rule subclass A, that is, using a text large model) to check according to the compliance requirements of financial marketing content (for example, "must contain explicit risk prompt sentences", "risk prompt needs to be displayed in the fixed area of the tail frame", "cannot only mention through voice without being presented in the picture"), such as checking whether there is a risk prompt text that meets the specifications (such as whether it contains "investment has risks", "financial caution" and other core expressions); verify the presentation form of the risk prompt (such as whether it is visible in the image, not just in the audio), confirm the compliance of the location and duration (such as whether it is displayed continuously for at least 2 seconds at the tail frame of the video, whether it is located in the fixed visible area of the picture).

[0121] Finally, according to the generated first sub-result, the verification conclusion of the risk prompt will be recorded: if it fully meets the requirements, it will be marked as "risk prompt compliance" and the related text content, location and duration; if there is a lack (such as no risk prompt), a form violation (such as only mentioning through voice) or a location inconsistency (such as not displaying in the tail frame), it will be marked as "risk prompt non-compliance" and the specific problems, as an important part of the first audit result.

[0122] S502, detecting user privacy data in the text fusion information to obtain a second sub-result.

[0123] In the embodiments of the present application, detecting user privacy data in the text fusion information and obtaining a second sub-result is a special detection process performed by the audit rule engine for the "privacy compliance" dimension, which is a subdivision verification item of the first audit result.

[0124] Specifically, the text fusion information contains text in the video picture (such as subtitles, background text) and corrected audio transcription text (such as voice-mentioned content), and user privacy data may exist in various forms, such as identity card numbers, bank card numbers, mobile phone numbers appearing in the picture, or "certain user loan records", "personal contact information" mentioned in the voice.

[0125] The rule engine calls the "privacy data detection subclass" (a special rule designed based on the strategy pattern) to conduct detection in two core ways, such as using a preset privacy data format library (such as a 11-digit mobile phone number rule and an 18-digit ID number coding rule), and combining a text large model to identify sensitive information in the text that meets the format characteristics. The second sub-result generated after detection is presented in a structured form: if no privacy data is found, it is marked as "no user privacy information detected"; if relevant content is detected, the type of privacy data (such as mobile phone number / ID number), specific content (partially desensitized display, such as "138****5678"), source (image text / audio transcription), and corresponding timestamp are recorded, providing a basis for subsequent privacy compliance judgment.

[0126] S503, detecting product pricing information in the text fusion information to obtain a third sub-result.

[0127] In the embodiments of the present application, detecting product pricing information in the text fusion information and obtaining a third sub-result is a special check carried out by the audit rule engine for "product pricing expression compliance", which belongs to a subdivision item of the first audit result.

[0128] In the text fusion information, product pricing information may exist in the form of image text (such as picture subtitles "annual interest rate 3.5%") or audio transcription text (such as voice mentioning "borrowing monthly interest 0.8%"). When verifying the correctness of the interest rate expression in the subtitles, the system calls the rule engine, such as a text large model of rule subclass A, combines the text large model with a customized annual interest rate calculation tool, and uses the tool calling ability of the large model to realize accurate verification. The specific process is as follows:

[0129] The text large model first performs semantic analysis on the interest rate related expressions in the subtitles. Whether it is a direct "annual interest rate 3.5%" or an indirect "10,000 yuan per day interest 1 yuan" or "borrow 10,000 yuan per day 0.8 yuan", the large model can accurately extract the core parameters through natural language understanding ability, such as identifying "principal 10,000 yuan" and "daily interest 1 yuan" from "10,000 yuan per day interest 1 yuan", and locating "principal 10,000 yuan" and "daily repayment interest 0.8 yuan" from "borrow 10,000 yuan per day 0.8 yuan". This analysis has strong robustness, and even if the expression form is flexible and there is a colloquial expression (such as "10,000 yuan per day interest 1 yuan"), it can accurately match the key variables (principal, interest amount, time unit) required for interest calculation

[0130] Subsequently, the large model calls the customized annualized interest rate calculation tool, and passes the extracted parameters into the tool in the format required by the tool (such as inputting "principal = 10000, daily interest = 1"). The tool automatically calculates the annualized interest rate based on financial formulas (for example, the annualized interest rate corresponding to "10000 yuan daily interest 1 yuan" is 3.65%), and returns the calculation result. The large model compares the calculation result with the interest rate directly expressed in the subtitle (if any), or combines the regulatory rules to determine whether the interest rate expression is compliant (such as whether there is a problem of "as low as X%" without labeling the applicable conditions), and finally outputs the verification conclusion of the correctness of the interest rate expression as the core basis for product pricing information review (the third sub-result).

[0131] S504, obtaining a first review result according to the first sub-result, the second sub-result and the third sub-result.

[0132] In the embodiments of the present application, the first review result obtained according to the first sub-result (risk prompt information detection result), the second sub-result (user privacy data detection result) and the third sub-result (product pricing information detection result) is a process of summarizing, classifying and integrating each special verification result.

[0133] In the above method, through special detection and result integration of three types of core information of risk prompt, user privacy and product pricing in the text fusion information, the first review result covering the key dimensions can be accurately generated, providing comprehensive and clear basis for compliance judgment of the video data at the text level.

[0134] S106, determining the compliance of the first video data according to the first review result.

[0135] In the embodiments of the present application, determining the compliance of the first video data according to the first review result is a process of comprehensive judgment based on the integrated verification conclusions (risk prompt, user privacy, product pricing, etc.) in the review result.

[0136] Specifically, the system will refer to the preset compliance judgment standard to evaluate the overall conclusion and subdivided problems (such as "whether there is a serious violation item" and "whether the number of general violation items exceeds the threshold value") in the first review result: if the first review result shows that all sub-results are compliant (without any violation record), it is determined that the video data is "compliant"; if there is a serious violation item such as "user privacy data is detected" or the number of general violation items such as "risk prompt missing" and "interest rate expression error" exceeds the set threshold value, it is determined as "non-compliant", and the corresponding processing (such as direct interception, marking for manual review before deciding) is triggered according to the violation level. The final compliance conclusion directly determines whether the first video data can enter the subsequent dissemination link.

[0137] In one embodiment, referring toFigure 7 is a process schematic diagram of data compliance judgment provided by an embodiment of the present application, as shown in Figure 7 Step S106 includes:

[0138] S601, detecting prohibited content in the first image to obtain a fourth sub-result.

[0139] In the embodiment of the present application, the violation review of the first video data includes not only the compliance review of the text fusion information, but also the review of multiple image frames.

[0140] Specifically, a multi-modal large model inspector can be used to detect whether violent, vulgar, horror, and intimidation bad content appears in the picture. This is achieved by the model's visual feature and semantic understanding ability to accurately identify bad content. The multi-modal large model (such as Qwen2.5-VL, GPT-4V, etc. with image-text understanding ability) receives the key frame pictures (first images) extracted from the first video data, first extracts the underlying visual features of the pictures, identifies the elements such as the actions of the characters in the image (such as fighting, violent gestures), the scene atmosphere (such as bloody pictures, dark and terrifying scenes), clothing or objects (such as vulgar and exposed clothing, horror props), etc. Then, combined with the bad content semantic features (such as “violence” corresponding to physical conflict, weapon threat, etc., “vulgarity” corresponding to exposed clothing, indecent action, etc.) learned during model training, the picture content is comprehensively judged for semantics to obtain the fourth sub-result.

[0141] S602, detecting the compliance of the identification type content contained in the first image to obtain a fifth sub-result.

[0142] In the embodiment of the present application, the identification type content in the first image can include a company logo (such as a financial institution brand logo), which can be verified by a self-trained model such as Figure 5 The logo model corresponding to the rule sub-class N shown in the visual model rule inspector verifies whether the logo is the latest version of the company (to avoid misuse of old logos), and checks whether there is an unauthorized use of a third-party logo (such as using a well-known financial institution logo). After the detection is completed, the fifth sub-result generated will be presented in a structured form: if all the logos comply with the specifications, it is marked as “identification type content is compliant”; if there is a violation (such as using an old logo, false qualification logo), the type of the violating logo (such as “Logo version inconsistency” “false qualification logo”) is explicitly recorded.

[0143] In addition to logo detection, the fifth sub-result can also use a multimodal model to identify whether the tail frame structure complies with financial advertising regulations. The multimodal model combines visual features (layout, elements) and textual semantics (content completeness) to comprehensively score the compliance of the tail frame structure. If the tail frame contains a complete risk warning, is prominently positioned, and has the required dwell time, it is judged as "compliant tail frame structure." If the risk warning is obscured or the dwell time is less than 3 seconds, it is judged as "non-compliant tail frame structure," and specific issues are clearly marked (e.g., "risk warning not prominently positioned" or "insufficient dwell time"). Furthermore, it is necessary to screen whether the content violates socialist core values ​​or ethical bottom lines.

[0144] S603, the second audit result is obtained based on the fourth and fifth sub-results.

[0145] In this embodiment, the second review result is obtained based on the fourth sub-result (the detection result of prohibited content in the first image) and the fifth sub-result (the detection result of compliance of the first image's identifier content). If neither sub-result is found to be in violation, the second review result is marked "The image content and identifier are both compliant". If a violation is found, an overall judgment will be made (e.g., "Serious violation exists" or "General violation only"), and specific issues will be listed item by item (e.g., "Fourth sub-result: Fake qualification image; Fifth sub-result: Use of old logo"), along with the corresponding frame timestamps, image positions, and other original evidence.

[0146] S604, determine the compliance of the first video data based on the first audit result and the second audit result.

[0147] In this application embodiment, the compliance of the first video data is determined based on the first review result (text dimension, covering risk warnings, user privacy, product pricing, etc.) and the second review result (image dimension, covering prohibited content, label compliance, etc.). This is a comprehensive evaluation of the video's "text + image" dual-dimensional review information, ultimately forming a complete compliance judgment conclusion.

[0148] The above method detects prohibited content and compliance of labels in images and integrates the results to generate a second review result. Combined with the first review result at the text level, it can comprehensively determine the compliance of video data from the dual dimensions of "image + text", providing an accurate and complete basis for video content compliance review.

[0149] In one embodiment, step S604 includes:

[0150] If either the first or second review result indicates that the review has failed, then the first video data is non-compliant; if both the first and second review results indicate that the review has passed, then the first video data is compliant.

[0151] In the embodiments of the present application, if the first audit result (text dimension) shows “audit failed” (such as missing risk prompt, user privacy leakage, etc.), whether the second audit result (image dimension) passes or not, the first video data is directly determined as “non-compliant”. If the second audit result (image dimension) shows “audit failed” (such as existence of prohibited content, identification violation, etc.), whether the first audit result passes or not, the first video data is also determined as “non-compliant”. Only when the first audit result and the second audit result both show “audit passed” (no violation in text dimension, no violation in image dimension), the first video data is finally determined as “compliant”.

[0152] In the above method, by explicitly determining that “any one dimension of text or image fails to pass the audit, and both dimensions pass to determine compliance”, the final determination of the compliance of the video data can be strictly and efficiently completed, ensuring that only the video that meets the requirements of both text and image is recognized as compliant, providing a clear and rigorous standard for the compliance control of video content.

[0153] In one embodiment, the data audit method further comprises:

[0154] obtaining second information; wherein the second information includes data information that fails to pass the audit and position information of the data information that fails to pass the audit in the first video data; generating a data audit report according to the second information, and performing an alarm prompt.

[0155] In the embodiments of the present application, the second information is the core data related to violation based on the premise that “the first audit result or the second audit result fails to pass”, including the data information that fails to pass the audit and the corresponding position information. The data audit report is a structured presentation of the second information, including video ID, audit time, audit type (financial advertisement audit), explicit label “fails to pass the audit, the first video data is non-compliant”, detailed list of violations, and suggestions for improvement, etc.

[0156] In addition, it should be noted that in order to ensure the stable execution of the audit task, the system designs the following fault tolerance mechanisms, including a retry mechanism, if the model fails to call or times out, the system can retry a specified number of times according to the configuration; the system also includes health monitoring and alarm, if the error rate of a certain model or interface continues to rise, an alarm will be automatically sent to the operation and maintenance system to assist in timely repair.

[0157] In the above method, by obtaining the data that fails to pass the audit and its position information, generating a report and alarming, the violation content can be accurately located, the audit result can be intuitively presented, and relevant personnel can be timely reminded to handle, thereby providing efficient support for the rectification and risk prevention and control of the violation video.

[0158] Reference is made to Figure 8Fig. 1 is a schematic diagram of a data auditing process according to an embodiment of the present application. As shown in the figure, the steps include:

[0159] 1) The calling party submits an auditing request, and the auditing data is added to the task queue

[0160] The calling party submits the video or picture link to be launched to the system. The system encapsulates the video launch channel (such as Douyin and Tencent) and content type information as an auditing task (Task) and adds it to the task queue.

[0161] 2) Download the video

[0162] Each Pod acts as a queue consumer and is responsible for pulling tasks and downloading video or picture resources.

[0163] 3) Scene segmentation to generate multiple video frames

[0164] For a video task, the system identifies all scenes and extracts the tail frame scene (about 2 seconds) from them, then performs frame extraction at a specified frequency (such as every 0.5s) to generate an image sequence (first image) for auditing.

[0165] 4) Subtitle fusion to obtain fused subtitles

[0166] The multi-modal model extracts image caption information (first image text), the ASR model extracts audio caption information (first audio text information), then corrects the audio caption information using the image caption information, and fuses the corrected audio caption information with the image caption information to obtain fused caption information (text fusion information).

[0167] 5) Multi-rule data auditing

[0168] The system uniformly delivers caption information (fused text), frame extraction pictures (first image), video metadata (first video data), etc. to the auditing rule engine. The rule engine uses a strategy pattern, supports different subclasses to implement different auditing rules, and supports parallelized auditing.

[0169] The thread pool contains multiple types of auditing rules, including visual rule checkers, multi-modal rule checkers, large model rule checkers, and regular rule checkers, etc. Different types of auditing rules use different auditing rules for different data to comprehensively determine the compliance of the video data.

[0170] The application is suitable for compliance review of marketing short videos or pictures in the financial field. After receiving video or picture links and metadata, the method is packaged as a task and processed by a review process in a Kubernetes cluster; information is obtained by frame extraction, audio extraction, etc., text is extracted using multi-modal large models and ASR models, and unified subtitles are generated by text large model fusion and completion; combined with various review rules, relevant models are called for analysis, and review conclusions are output. At the same time, it has distributed execution, multi-rule parallelism, fault tolerance and monitoring mechanisms, and the core advantage is the fusion of multi-source information and large model reasoning capability.

[0171] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or software. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0172] Figure 9 is a structural schematic diagram of a terminal device provided by the embodiments of the application. As shown in Figure 9 the terminal device 9 of the embodiment includes at least one processor 90 (only one processor is shown in the figure), a memory 91, and a computer program 92 stored in the memory 91 and executable on the at least one processor 90, and the processor 90 implements the steps in any of the above data review method embodiments when executing the computer program 92. Figure 9

[0173] The terminal device can be a desktop computer, a notebook computer, a palm computer, and a cloud server, etc. The terminal device can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 9 The terminal device 9 is only an example and does not constitute a limitation on the terminal device 9, which can include more or fewer components than shown, or combine certain components, or different components, for example, it can also include input / output devices, network access devices, etc.

[0174] ​The processor 90 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0175] The memory 91 can be an internal storage unit of the terminal device 9 in some embodiments, for example, a hard disk or a memory of the terminal device 9. The memory 91 can also be an external storage device of the terminal device 9 in other embodiments, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 91 can include both an internal storage unit and an external storage device of the terminal device 9. The memory 91 is used to store an operating system, application programs, a boot loader, data, and other programs, for example, program codes of computer programs, etc. The memory 91 can also be used to temporarily store data that has been output or will be output.

[0176] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps in the above-mentioned various method embodiments.

[0177] The embodiments of the present application provide a computer program product. When the computer program product is run on a terminal device, the terminal device is caused to implement the steps in the above-mentioned various method embodiments.

[0178] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the computer program for instructing the related hardware to complete all or part of the processes in the above-mentioned embodiments can be stored in a computer readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunication signal.

[0179] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0180] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0181] In the embodiments provided in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed each other can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0182] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may also be distributed to multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0183] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A data auditing method, characterized by, The method comprises: acquiring first video data; extracting first image text information corresponding to the first video data; extracting first audio text information corresponding to the first video data; fusing the first image text information and the first audio text information to obtain text fusion information; performing data auditing based on the text fusion information to obtain a first auditing result; determining compliance of the first video data according to the first auditing result; the first image text information and the first audio text information to obtain text fusion information, comprising: performing text supplement correction on the first audio text information according to the first image text information to obtain second audio text information after supplement correction; fusing the first image text information and the second audio text information to obtain the text fusion information; the first image text information and the first audio text information to obtain text fusion information, comprising: detecting risk prompt information in the text fusion information to obtain a first sub-result; detecting user privacy data in the text fusion information to obtain a second sub-result; detecting product pricing information in the text fusion information to obtain a third sub-result; obtaining the first auditing result according to the first sub-result, the second sub-result and the third sub-result; determining compliance of the first video data according to the first auditing result, comprising: detecting prohibited content in the first image to obtain a fourth sub-result; detecting compliance of identification type content contained in the first image to obtain a fifth sub-result; obtaining a second auditing result according to the fourth sub-result and the fifth sub-result; determining compliance of the first video data according to the first auditing result and the second auditing result.

2. The data auditing method of claim 1, wherein, The first image text information corresponding to the first video data is obtained, comprising: extracting single-frame pictures in the first video data according to a preset frequency to obtain a plurality of first images; for each first image, extracting first information corresponding to the first image; wherein the first information includes first text content, a first text type, and a first text position corresponding to the first text content; combining the first information corresponding to each of the plurality of first images to obtain the first image text information corresponding to the first video data.

3. The data auditing method of claim 1, wherein, The first audio text information corresponding to the first video data is extracted, comprising: extracting first audio data from the first video data; transcribing the first audio data into text data corresponding to each of a plurality of time sequences; combining the text data according to the time sequence to obtain the first audio text information corresponding to the first video data.

4. The data auditing method of claim 3, wherein, The first image text information corresponding to the first video data is obtained, comprising: if the first auditing result or the second auditing result indicates that the auditing fails, the first video data is not compliant; if the first auditing result and the second auditing result both indicate that the auditing passes, the first video data is compliant.

5. The data auditing method of claim 4, wherein, The method further comprises: obtaining second information; wherein the second information comprises un-passed data information and position information of the un-passed data information in the first video data; generating a data auditing report according to the second information and performing an alarm prompt.

6. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor implements the method of any one of claims 1 to 5 when executing the computer program.

7. A computer-readable storage medium storing a computer program, wherein the computer program comprises the following steps of: receiving a request for a resource from a client; determining whether the client is authorized to access the resource; and if the client is authorized to access the resource, providing the resource to the client. The computer program, when executed by the processor, implements the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Short video auditing method based on multiple modes

    CN115512259A