Multi-modal real-time compliance auditing method, system and device for investment adviser live broadcast and storage medium
By employing a multimodal real-time compliance review method, combined with audio and video processing and semantic recognition technology fine-tuned by the BERT model, the problem of insufficient accuracy and traceability in traditional manual review during investment advisory live broadcasts has been solved. This method enables real-time and reliable identification and recording of illegal content, and is suitable for high-frequency investment advisory live broadcast scenarios with multiple broadcasters.
Patent Information
- Application Number
- CN202511087469.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional compliance management solutions for investment advisory live streams, which rely on manual review, have significant shortcomings in terms of compliance efficiency, identification accuracy, and regulatory traceability, making it difficult to meet the regulatory requirements for real-time risk control, process monitoring, and compliance record keeping.
A multimodal real-time compliance review method is adopted, which processes the live stream in both audio and video formats, combines image frame and audio slice recognition, and uses semantic recognition technology fine-tuned by the BERT model to perform speech transcription and text analysis, generating a traceable compliance record chain to achieve real-time identification and recording of illegal content.
It significantly improves the ability to identify and review illegal content, achieves speech transcription and semantic analysis with a latency of less than a second, supports multi-anchor concurrent scenarios, has good scalability and adaptability, and meets the real-time risk control requirements of the financial sector.
Smart Images

Figure CN120951137A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of compliance auditing technology, and in particular to a multimodal real-time compliance auditing method, system, device, and storage medium for investment advisory live streaming. Background Technology
[0002] With the rapid development of the digital economy, securities technology and online live streaming are increasingly merging, and securities firms are widely using live streaming to promote investment advisory services and educate investors. However, the surge in live streams has also led to a high incidence of violations related to investment advisory content (such as unauthorized stock recommendations, profit-sharing implications, and false advertising). This not only harms investors' interests but also poses a potential threat to the stability and healthy development of the financial market. Therefore, strengthening compliance supervision of investment advisory live streams is particularly important.
[0003] Currently, compliance measures in the securities investment advisory industry primarily rely on pre-broadcast submission outlines, requiring hosts to submit their content for pre-approval before the broadcast. However, this method has several limitations. First, it struggles to cover improvisation, spontaneous changes in delivery, and evolving violations during the live stream. Second, internet live streaming platforms' own review capabilities are limited to basic content such as pornography, terrorism, and prohibited words, failing to identify the complex and implicit risks of expression in securities investment advisory scenarios. Furthermore, traditional manual review methods also suffer from the following problems: Limited real-time performance and response speed: Relying on manual monitoring methods may result in untimely responses to sudden or covert violations, posing a risk of delayed review.
[0004] Limited recognition dimensions: Manual review mainly relies on auditory judgment, making it difficult to comprehensively recognize non-audio content such as video footage and subtitle information, and it also places high demands on compliance personnel.
[0005] Limited review efficiency: In scenarios with multiple live streamers operating concurrently and long live stream durations, the number and energy of compliance personnel are limited, making it difficult to achieve high-frequency, large-scale, and comprehensive coverage.
[0006] Lack of compliant record-keeping mechanisms: Manual monitoring cannot efficiently record risky content that occurs during live broadcasts, making it difficult to meet the requirements for post-event auditing and evidence tracing.
[0007] In summary, traditional compliance management solutions for investment advisory live streams, which rely on manual review, have significant shortcomings in terms of compliance efficiency, identification accuracy, and regulatory traceability, making it difficult to meet regulatory technical requirements for "real-time risk control, process monitoring, and compliance record-keeping." Therefore, the securities investment advisory industry urgently needs a real-time live stream compliance risk control system that integrates multimodal technologies to adapt to the content characteristics, compliance requirements, and regulatory pressures of the securities scenario. Summary of the Invention
[0008] This invention proposes a multimodal real-time compliance review method, system, and storage medium for investment advisory live streaming, solving the problems mentioned in the background section. The technical solution of this invention is implemented as follows: A multimodal real-time compliance auditing method for the investment advisory live streaming industry includes the following steps: Step S1: Perform dual-track audio and video processing on the live stream, and extract image frames and audio slices; Step S2: Identify the image frames to determine if there are any illegal posters, slogans, or unauthorized product displays. Step S3: Transcribe the audio segments into text in real time and use a natural language processing model to identify potential offensive language. Step S4: Complete speech transcription, image recognition and semantic analysis within a second-level delay, and automatically label the identified risky content, generate timestamps and violation tags to build a traceable compliance record chain.
[0009] Further steps for identifying image frames include: Step S21: Perform optical character recognition on the image frame to extract text information from the image; Step S22: Compare the extracted text information with the preset prohibited text to determine whether there is any prohibited expression; Step S23: Input the extracted text information into a semantic analysis model finely tuned from the investment advisory corpus to further identify potential violations and hidden compliance risks.
[0010] Further steps for real-time transcription of audio slices include: Step S31: The audio is sliced using a sliding window mechanism. Each audio slice is divided according to a fixed length and overlap ratio to ensure semantic integrity. Step S32: Perform speech recognition on the sliced audio and transcribe it into structured text data.
[0011] Furthermore, the steps for identifying potentially offensive statements using natural language processing models include: Step S33: Based on the pre-trained Chinese natural language understanding model, perform semantic analysis on the transcribed text; Step S34: Determine if there are any illegal tendencies such as disguised promises of returns, implicit stock recommendations, or hints of insider information.
[0012] A multimodal real-time compliance auditing system for the investment advisory live streaming industry includes: The live stream data acquisition module is used to access the real-time video stream of investment advisors and anchors; The live stream data parsing module is used to perform multimedia demultiplexing processing on the acquired video stream and extract audio tracks and video image frames; The image compliance module is used to identify image frames and determine whether there is any illegal content, potential illegal tendencies, or hidden compliance risks. The audio speech processing and transcription module is used to transcribe audio segments into text in real time. The text semantic model analysis module is used to perform semantic analysis on the transcribed text to determine whether there is any tendency to violate regulations. The risk tracking module is used to automatically tag identified risky content and generate timestamps and violation labels.
[0013] Furthermore, the image compliance module includes: The image text recognition unit is used to perform optical character recognition on image frames and extract text information; The illegal text detection unit is used to compare the extracted text information with preset illegal text. A semantic analysis model, fine-tuned using corpus data from the investment advisory field, is used to identify potential violations and hidden compliance risks.
[0014] Furthermore, the audio speech processing and transcription module includes: The audio slicing unit is used to slice audio using a sliding window mechanism; The speech recognition unit is used to perform speech recognition on the sliced audio and transcribe it into structured text data.
[0015] Furthermore, the text semantic model analysis module includes: The semantic analysis unit is used to perform semantic analysis on the transcribed text based on a pre-trained Chinese natural language understanding model. The violation judgment unit is used to determine whether there are any violations such as disguised promises of returns, implicit stock recommendations, or hints of insider information.
[0016] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned multimodal real-time compliance audit method for the investment advisory live streaming industry.
[0017] A multimodal real-time compliance auditing device for the investment advisory live streaming industry includes a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the multimodal real-time compliance auditing method for the investment advisory live streaming industry.
[0018] Compared with existing technologies, this solution has the following advantages: 1. Enhanced ability to identify illegal content. By introducing semantic recognition technology based on BERT model fine-tuning, the system can identify implicit illegal content such as "borderline expressions," "vague promises," and "soft stock recommendations," breaking through the limitations of traditional keyword recognition and significantly enhancing the system's ability to identify high-risk semantics. Combining image recognition and text analysis, it can not only identify illegal content in speech but also detect illegal text, qualification displays, and promotional materials in live broadcasts, improving the comprehensiveness and accuracy of illegal content identification. 2. Improve review efficiency and real-time performance. The system can complete speech transcription, image recognition, and semantic analysis within seconds, and automatically generate violation tags to build a traceable compliance record chain, meeting the real-time risk control requirements of the financial sector. Employing an overlapping audio sliding window algorithm avoids omissions and misidentifications caused by missing speech boundary information, significantly improving the accuracy of speech recognition and the continuity of semantic analysis. 3. Enhance the traceability of audits. The system automatically records violations in a structured manner, including the type of violation, timestamp, source identifier, etc., forming a complete and traceable violation record, providing a reliable basis for compliance audits and risk control management. It supports real-time slicing and caching of live streams based on the HLS protocol, allowing compliance personnel to quickly retrieve and review violation live stream video slices for verification and evidence consolidation. 4. Reduced deployment costs and improved scalability. Utilizing lightweight pre-trained language models such as BERT for fine-tuning, compared to complex large language model solutions, deployment costs are lower and controllability is stronger, making it suitable for rapid deployment and iteration in investment advisory scenarios such as securities, funds, and insurance. The system supports multiple mainstream protocols and access methods, adapting to high-frequency investment advisory live streaming scenarios with multiple broadcasters and multi-platform synchronization, exhibiting good scalability and adaptability. 5. Meets regulatory requirements. Through real-time identification and alert mechanisms, timely intervention can be provided when violations occur, avoiding delays in post-event accountability and ensuring the compliant operation and controllable risks of investment advisory live streams. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the core module of a multimodal real-time compliance review system for the investment advisory live streaming industry according to the present invention; Figure 2This is a schematic diagram of the live audio compliance module of a multimodal real-time compliance review system for the investment advisory live streaming industry, as described in this invention. Figure 3 This is a schematic diagram of the live image compliance module of a multimodal real-time compliance review system for the investment advisory live streaming industry, as described in this invention. Figure 4 This is a schematic diagram illustrating the completion process of a multimodal real-time compliance review method for the investment advisory live streaming industry according to the present invention. Detailed Implementation
[0021] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0022] Reference Figures 1-4 This invention provides a multimodal real-time compliance review method for the investment advisory live streaming industry, characterized by the following steps: Step S1: Perform dual-track audio and video processing on the live stream, and extract image frames and audio slices; Step S2: Identify the image frames to determine if there are any illegal posters, slogans, or unauthorized product displays. Step S3: Transcribe the audio segments into text in real time and use a natural language processing model to identify potential offensive language. Step S4: Complete speech transcription, image recognition and semantic analysis within a second-level delay, and automatically label the identified risky content, generate timestamps and violation tags to build a traceable compliance record chain.
[0023] This method is implemented based on a real-time compliance review system for investment advisory live streaming scenarios, which includes: The live stream data acquisition module is used to access the real-time video stream of investment advisors and anchors; The live stream data parsing module is used to perform multimedia demultiplexing processing on the acquired video stream and extract audio tracks and video image frames; The image compliance module is used to identify image frames and determine whether there is any illegal content. The audio speech processing and transcription module is used to transcribe audio segments into text in real time. The text semantic model analysis module is used to perform semantic analysis on the transcribed text to determine whether there is any tendency to violate regulations. The risk tracking module is used to automatically tag identified risky content and generate timestamps and violation labels.
[0024] Furthermore, the image compliance module includes: The image text recognition unit is used to perform optical character recognition on image frames and extract text information; The illegal text detection unit is used to compare the extracted text information with preset illegal text.
[0025] Furthermore, the audio speech processing and transcription module includes: The audio slicing unit is used to slice audio using a sliding window mechanism; The speech recognition unit is used to perform speech recognition on the sliced audio and transcribe it into structured text data.
[0026] Furthermore, the text semantic model analysis module includes: The semantic analysis unit is used to perform semantic analysis on the transcribed text based on a pre-trained Chinese natural language understanding model. The violation judgment unit is used to determine whether there are any violations such as disguised promises of returns, implicit stock recommendations, or hints of insider information.
[0027] This system integrates various artificial intelligence algorithms and methods, such as audio and video data stream acquisition, natural language processing (NLP), speech recognition (ASR), image analysis, text semantic understanding, and machine learning models, to achieve automatic identification, intelligent analysis, and early warning of violations in investment advisory live broadcasts.
[0028] This system can fuse and analyze multimodal information such as audio, image frames, and subtitle text in live video streams. It has advantages such as low latency, high accuracy, and strong scalability. It can perform real-time compliance review and risk intervention on the voice expression and on-screen text during the live broadcast. The overall process includes, but is not limited to: live stream access, image frame capture and recognition, real-time audio transcription, text sensitive content analysis, semantic understanding judgment, violation behavior identification, risk alarm and handling response. It is suitable for high-frequency investment advisory live broadcast compliance scenarios with multiple anchors and multiple platforms.
[0029] Specifically, its main modules and review process are as follows: 1. Live stream data acquisition module The live stream data acquisition module is used to access the real-time video stream from investment advisors and anchors, serving as the input source for system processing. This module supports various mainstream protocols and access methods, including but not limited to: accessing via the SDK provided by the live streaming platform, or pulling live video streams based on streaming media transmission protocols such as RTMP (Real-Time Messaging Protocol), HLS (HTTP Live Streaming), and M3U8. The module can stably receive, buffer, and parse the video stream, providing raw data support for subsequent processing stages such as audio / video separation, image frame extraction, and speech recognition.
[0030] 2. Live Stream Data Parsing Module
[0031] The live stream data parsing module performs multimedia demultiplexing on the acquired video stream, extracting the audio track and video image frames, which serve as input sources for speech recognition and image analysis, respectively. This module supports parsing mainstream streaming media formats such as RTMP, HLS, and M3U8, and performs real-time processing of the live data stream based on multimedia encoding and decoding technologies.
[0032] 3. Live Stream Audio Track and Image Extraction Module
[0033] This module is used to perform real-time decoding of live video streams and periodically extract audio segments and image frames according to a set strategy, serving as input for image compliance recognition and voice content review. Specifically, it includes: Image frame extraction: Based on the video frame stream, the system performs image frame extraction at fixed time intervals (configurable to every 10-60 seconds) to capture key content displayed in the live broadcast. The extracted image frames are then sent to the image compliance recognition module to detect text and other content on the screen.
[0034] Audio segmentation and overlap processing: The system synchronously extracts audio data segments corresponding to the time period. The audio duration is generally consistent with the image frame extraction cycle (e.g., every 15 seconds). However, to avoid semantic segmentation breaks or missed words, an overlap window can be set between audio segments, usually 5 to 10 seconds. The overlap design helps to preserve the complete expressive structure in subsequent speech recognition and contextual semantic analysis.
[0035] The extracted audio slices and image frames will be sent to the audio compliance module and image compliance module respectively for processing, so as to realize segmented and multimodal compliance review of live broadcast content.
[0036] There is temporal overlap between audio segments. This overlapping design effectively avoids semantic breaks and keyword omissions caused by audio slice boundaries, improving the accuracy of subsequent speech recognition and the completeness of the model's understanding of contextual semantics in step 7.
[0037] The technical process of the overlapping sliding window time alignment algorithm is as follows: 1) Audio sampling and buffering processing After receiving video stream data from the live stream, the system first extracts the audio track data into audio frames and stores them in the audio-video buffer. The buffer stores the audio stream in a continuous manner, with the unit being timestamp + audio frame. 2) Window definition and parameter configuration
[0038] Define two core parameters for the audio sliding window: W: Total duration of each audio slice (e.g., 15 seconds) O: The overlap duration between adjacent audio slices (e.g., 5 seconds) The sliding step size S = W - O (i.e., extracting once every S seconds). 3) Sliding window slicing mechanism Starting from the initial time T, the system slides and samples according to the step size S, cutting the window [T, T+W], [T + S, T + S + W], and [T + 2S, T + 2S + W] to form a sequence of continuously overlapping audio slices.
[0039] With a starting time of T = 0 seconds, a total duration of audio slices of W = 15 seconds, an overlap duration of O = 5 seconds, and a sliding step size of S = W - O = 10 seconds, the resulting time series of audio slices are: [0, 15], [10, 25], [20, 35] ...
[0040] 4. Image Compliance Module
[0041] The image compliance module is used to intelligently identify and assess the compliance of image frames captured during investment advisory live streams, in order to identify potentially illegal text information or display non-compliant behavior within the images. This module primarily processes keyframe images extracted from step 2 (video frame capture module), serving as the core processing step in the image content analysis within the risk control process. The main processes of image compliance processing include: 1) Image text recognition (OCR) The module first performs OCR (Optical Character Recognition) processing on the image frames to extract text information from the images. This step is not limited to using a specific OCR engine and can be compatible with general open-source frameworks (such as PaddleOCR and Tesseract) or commercial interfaces (such as Baidu and Tencent OCR services), only requiring that it can output structured text; 2) Detection of illegal copywriting
[0042] The extracted image text will be used as input, and the sensitive word recognition module and the text compliance model (see step 7) will jointly determine whether there are any illegal expressions, such as: This implies a guarantee of returns or stock recommendation behavior. The advertising language used was beyond the permitted scope; Misleading statements (such as "guaranteed profit" or "institutional research"); Unauthorized display of unauthorized content; 3) Verification of live broadcast display requirements
[0043] The system automatically detects whether the image displays information such as the name, license number, or any special text related to investment advisors. If such content is found, the OCR extracts it and matches it against the preset registration information for the live broadcast room in the system for verification. 5. Audio Speech Processing and Transcription Module The audio speech processing and transcription module is used to perform speech recognition (ASR) processing on the live audio segments extracted in step 2, transcribing the broadcaster's speech content into structured text data in real time for subsequent text semantic understanding and compliance analysis. In specific implementation schemes, this invention does not limit the implementation method of the speech recognition model, and can adopt a local speech recognition model or a third-party cloud speech recognition model according to system deployment requirements and performance requirements.
[0044] 6. Sensitive text content recognition
[0045] The text-sensitive content recognition module performs a rapid initial compliance screening of the structured text data output by the audio-to-speech module to determine whether it contains typical prohibited expressions. This module employs a keyword-level sensitive word recognition model, based on a pre-set compliance lexicon, to perform high-performance string matching and rule validation on the text content, identifying whether it contains sensitive terms or suggestive expressions that violate regulatory provisions. Typical sensitive expressions include, but are not limited to: Illegal promises include: "Guaranteed minimum return", "Sure profit", "Only rising prices", "Target price of XX yuan"; Misleading information: such as "insider information", "national team has entered the market", "experts strongly recommend", "price surge is imminent"; Inappropriate statements: such as "suitable for everyone", "the first choice for the elderly", "you can buy it without thinking", etc.
[0046] This module is the system's first line of semantic defense, enabling efficient interception and rapid early warning of obviously illegal content.
[0047] 7. Text Semantic Model Analysis
[0048] The text model analysis module is used to process text content that was not detected by the sensitive word identification module, further assessing whether it has potential violations or suggestive expressions, and improving the system's ability to identify non-obvious violations such as "borderline expressions" and "implicit misleading content".
[0049] This module employs pre-trained language understanding models, such as BERT (Bidirectional Encoder Representations from Transformers), and fine-tunes the training based on a large corpus of language from the financial compliance field and manually annotated corpora to construct a dedicated semantic analysis model suitable for investment advisory live streaming scenarios. Compared to general large language models (LLM), this approach offers lower deployment costs, stronger controllability, and more stable output behavior, avoiding the illusion, overfitting, and uninterpretability issues that large models may encounter in small-sample compliance scenarios. It possesses the following core capabilities: Contextual semantic understanding capability: Supports analysis of the anchor's speech segments to determine whether there are behaviors such as disguised promises of returns, implicit stock recommendations, or hints of insider information; Semantic recognition capabilities in the financial field: It can identify compliance risk points in financial terms, such as expressions that imply investment intentions, such as "position recommendations", "pre-holiday positioning", and "buying at low valuations". Semantic deviation detection mechanism: Analyzes the degree of deviation between the expressed compliance intent and the actual semantics, and identifies "borderline language" that circumvents keywords but expresses illegal meanings.
[0050] The specific implementation process is as follows: 1. Sources and Construction Methods of Training Corpus The model is derived from a high-quality corpus constructed using a combination of the following two methods: Large-scale model annotation (automatic generation of initial labels). Automated analysis and initial annotation of live streaming historical corpora based on a general large-scale model.
[0051] Human-assisted annotation (enhancing data quality). Ambiguous expressions are manually reviewed and adjusted by professional compliance personnel. Typical regulatory cases and penalty notices are obtained from the Internet as positive and negative samples to enhance the model's ability to recognize boundary semantics.
[0052] 2. Model Structure and Training Method
[0053] Model Structure
[0054] Using a pre-trained Chinese natural language understanding model as the backbone model, a classification head is added to output a binary classification, with the classification labels defined as 0 - compliant and 1 - indicating risk.
[0055] Input features
[0056] The input content is a text fragment transcribed from an audio slice, and the sentence length is limited to 256 tokens.
[0057] Based on the above implementation process, the indicators after model fine-tuning are as follows: Assessment accuracy: 0.9750.
[0058] The recall rate for those labeled 0 (compliant) was assessed at 0.9723%.
[0059] The recall rate for the label 1 (a statement indicating risk) was assessed at 0.9776%.
[0060] The accuracy rate of recognition is greater than 95% of the industry average.
[0061] This solution utilizes a finely labeled financial compliance dataset for supervised training, enabling the model to achieve good accuracy with limited computing resources, facilitating deployment and rapid iteration. Compared to general large-scale model solutions, BERT-like models offer greater interpretability, faster response times, and lower review latency, making them suitable for high-frequency business scenarios such as live streaming review.
[0062] 8. Risk Record Keeping
[0063] When the compliance identification module detects potential violations in the images, text, or audio content of a live stream, the system automatically structures and records the violation content and related metadata (including violation type, timestamp, source identifier, etc.) in a data warehouse, forming a complete and traceable violation record. This recording mechanism ensures continuous tracking and subsequent analysis of violation content, providing a reliable basis for compliance audits and risk control management, and also serves as the original corpus data for the compliance big data model.
[0064] 9. Manual review
[0065] Risk control alerts are pushed to compliance personnel in real time via multiple channels, including WeChat, SMS, and email. Compliance personnel can quickly retrieve and review the corresponding violation video clips through the system for in-depth review. Based on the review results, compliance personnel assess the risk level of the violation and implement corresponding management measures for the broadcaster, including real-time warnings and live stream interruption (cutting off the broadcast), ensuring that violations are promptly curbed and effectively handled, and guaranteeing the compliant operation and risk controllable nature of investment advisory live streams.
[0066] The technical highlight of this invention is that compliance personnel can directly review the precisely corresponding violation live video clips, greatly improving the efficiency and accuracy of the review, which is different from the traditional methods that rely solely on manual monitoring or vague post-event review.
[0067] Compared with the prior art, the beneficial effects of the present invention are mainly reflected in the following aspects: 1. Enhance the ability to identify illegal content. Implicit violation identification: By introducing semantic recognition technology based on BERT model fine-tuning, it is possible to identify implicit violations such as "borderline expressions", "vague promises" and "soft stock recommendations", which breaks through the limitations of traditional keyword recognition and significantly enhances the system's ability to identify high-risk semantics.
[0068] Multimodal fusion: Combining image recognition and text analysis, it can not only identify illegal content in speech, but also detect illegal text, qualification displays and promotional materials in live broadcasts, improving the comprehensiveness and accuracy of illegal content identification.
[0069] 2. Improve review efficiency and real-time performance.
[0070] Real-time identification and record keeping: The system can complete speech transcription, image recognition and semantic analysis within a second-level delay, and automatically generate violation tags to build a traceable compliance record keeping chain, meeting the requirements of the financial sector for real-time risk control.
[0071] Sliding window audio slicing mechanism: The overlapping audio sliding window algorithm avoids the problems of missed recognition and false recognition caused by missing speech boundary information, and significantly improves the accuracy of speech recognition and the continuity of semantic analysis.
[0072] 3. Enhance the traceability of audits
[0073] Risk tracking mechanism: The system automatically records violations in a structured manner, including the type of violation, timestamp, source identifier, etc., forming a complete and traceable violation record, providing a reliable basis for compliance audit and risk control management.
[0074] Video slicing and playback mechanism: Supports real-time slicing and caching of live streams based on the HLS protocol. Compliance personnel can quickly retrieve and review violation live stream video slices through the system for verification and evidence preservation.
[0075] 4. Reduce deployment costs and improve scalability
[0076] Lightweight model deployment: Lightweight pre-trained language models such as BERT are used for fine-tuning. Compared with complex large language model solutions, the deployment cost is lower and the controllability is stronger. It is suitable for rapid implementation and iteration in investment advisory scenarios such as securities, funds, and insurance.
[0077] Excellent scalability: The system supports multiple mainstream protocols and access methods, and can adapt to high-frequency investment advisory live streaming scenarios with multiple broadcasters and multiple platforms in sync, demonstrating excellent scalability and adaptability.
[0078] 5. Improve the efficiency of manual review
[0079] Precise violation identification: The system supports structured recording of violation content and marking of corresponding time segments. Compliance personnel can directly review the precise violation-related live video clips, greatly improving the efficiency and accuracy of the review, which is different from traditional methods that rely solely on manual monitoring or vague post-event review.
[0080] 6. Meet regulatory requirements
[0081] Compliance Record Keeping: The system has the ability to automatically mark and trace violations, which can meet the compliance record keeping requirements in the financial field and provide strong support for supervision.
[0082] Real-time intervention: Through real-time identification and alarm mechanisms, timely intervention can be provided when violations occur, avoiding delays in post-event accountability and ensuring compliant operation and controllable risks in investment advisory live broadcasts.
[0083] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multimodal real-time compliance review method for the investment advisory live streaming industry, characterized in that, Includes the following steps: Step S1: Perform dual-track audio and video processing on the live stream, and extract image frames and audio slices; Step S2: Identify the image frames to determine whether there are any illegal posters, slogans, unauthorized product displays, or materials not displayed as required; Step S3: Transcribe the audio segments into text in real time and use a natural language processing model to identify potential offensive language. Step S4: Complete speech transcription, image recognition and semantic analysis within a second-level delay, and automatically label the identified risky content, generate timestamps and violation tags to build a traceable compliance record chain.
2. The multimodal real-time compliance review method for the investment advisory live streaming industry according to claim 1, characterized in that, The steps for identifying image frames include: Step S21: Perform optical character recognition on the image frame to extract text information from the image; Step S22: Match and compare the extracted text information with the preset illegal word library and sensitive expression templates to determine whether there are any illegal expressions; Step S23: Input the extracted text information into a semantic analysis model finely tuned from the investment advisory corpus to further identify potential violations and hidden compliance risks.
3. The multimodal real-time compliance review method for the investment advisory live streaming industry according to claim 1, characterized in that, The steps for real-time transcription of audio slices include: Step S31: The audio is sliced using a sliding window mechanism. Each audio slice is divided according to a fixed length and overlap ratio to ensure semantic integrity. Step S32: Perform speech recognition on the sliced audio and transcribe it into structured text data.
4. The multimodal real-time compliance review method for the investment advisory live streaming industry according to claim 1, characterized in that, The steps for identifying potential offensive statements using a natural language processing model include: Step S33: Based on the pre-trained Chinese natural language understanding model, perform semantic analysis on the transcribed text; Step S34: Determine if there are any illegal tendencies such as disguised promises of returns, implicit stock recommendations, or hints of insider information.
5. A multimodal real-time compliance review system for the investment advisory live streaming industry, characterized in that: include: The live stream data acquisition module is used to access the real-time video stream of investment advisors and anchors; The live stream data parsing module is used to perform multimedia demultiplexing processing on the acquired video stream and extract audio tracks and video image frames; The image compliance module is used to identify image frames and determine whether there is any illegal content, potential illegal tendencies, or hidden compliance risks. The audio speech processing and transcription module is used to transcribe audio segments into text in real time. The text semantic model analysis module is used to perform semantic analysis on the transcribed text to determine whether there is any tendency to violate regulations. The risk tracking module is used to automatically tag identified risky content and generate timestamps and violation labels.
6. The multimodal real-time compliance review system for the investment advisory live streaming industry according to claim 5, characterized in that, The image compliance module includes: The image text recognition unit is used to perform optical character recognition on image frames and extract text information; The illegal text detection unit is used to compare the extracted text information with preset illegal text. A semantic analysis model, fine-tuned using corpus data from the investment advisory field, is used to identify potential violations and hidden compliance risks.
7. The multimodal real-time compliance review system for the investment advisory live streaming industry according to claim 5, characterized in that, The audio speech processing and transcription module includes: The audio slicing unit is used to slice audio using a sliding window mechanism; The speech recognition unit is used to perform speech recognition on the sliced audio and transcribe it into structured text data.
8. The multimodal real-time compliance review system for the investment advisory live streaming industry according to claim 5, characterized in that, The text semantic model analysis module includes: The semantic analysis unit is used to perform semantic analysis on the transcribed text based on a pre-trained Chinese natural language understanding model. The violation judgment unit is used to determine whether there are any violations such as disguised promises of returns, implicit stock recommendations, or hints of insider information.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multimodal real-time compliance audit method for the investment advisory live streaming industry as described in any one of claims 1 to 4.
10. A multimodal real-time compliance review device for the investment advisory live streaming industry, characterized in that, It includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the multimodal real-time compliance audit method for the investment advisory live streaming industry as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Financial live broadcast violation detection method, device and equipment and readable storage medium
CN113038153A
System and method for compliance quality inspection and interactive question answering of live broadcast scene
CN114786035A
Model-based network live broadcast e-commerce intelligent monitoring method and system
CN118972665A
Internet live broadcast illegal behavior monitoring method and system based on image and voice recognition
CN120302088A
Cited By
Financial industry live broadcast-oriented intelligent compliance monitoring method and system
CN122093589A
Real-time message multi-mode content auditing system based on AI
CN122333191A
An AI-based real-time message multi-modal content auditing system
CN122333191B