Audio and video stream processing method and device, computing equipment, storage medium and product

By working together with domain-general and domain-specific models, sensitive information in live audio and video streams can be identified and processed in a tiered manner, solving the problems of low efficiency and high cost in the supervision of live streaming chaos, and achieving efficient and accurate blocking of sensitive information and system self-adaptation.

CN121750890APending Publication Date: 2026-03-27BEIJING 58 INFORMATION TTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In the current technology, it is difficult to prevent the chaos in online live streaming before it occurs, and the manual review method is inefficient and costly, and cannot effectively supervise the spread and dissemination of sensitive information.

Method used

By introducing domain-general and domain-specific models to work together, the system identifies the confidence level of sensitive information in live audio and video streams across multiple domains, determines the target domain, identifies and classifies sensitive information, and adopts corresponding handling strategies to prevent the spread of sensitive information.

Benefits of technology

It enables the detection of sensitive information before the transmission of live audio and video streams, effectively preventing their spread and propagation, improving processing efficiency, reducing costs, and enhancing recognition accuracy and system adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750890A_ABST
    Figure CN121750890A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio and video stream processing method and device, computing equipment, a storage medium and a product. The method comprises the following steps: in response to a live broadcast starting event, acquiring a live broadcast audio and video stream transmitted by an anchor end, and determining confidence degrees that the live broadcast audio and video stream has sensitive information in a plurality of fields by using a field general model; determining at least one target field of which the confidence is greater than a corresponding confidence threshold, and identifying sensitive information and sensitive levels corresponding to the live audio and video streams in each target field by using a field-specific model corresponding to each target field; and determining a risk level of the live broadcast audio and video stream according to the sensitivity level corresponding to each target field, and determining and executing a processing operation of a target processing strategy corresponding to the risk level from grading processing strategies for transmitting the live broadcast audio and video stream to a playing end. According to the scheme provided by the embodiment of the invention, the live broadcast audio and video stream containing the sensitive information can be processed at low cost and high efficiency so as to prevent the propagation and diffusion of the sensitive information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computing device, storage medium and product for processing audio and video streams. Background Technology

[0002] With the rapid development of the online live streaming industry, the chaos in live streaming has become an issue that cannot be ignored. The so-called chaos in live streaming refers to a series of negative phenomena that arise during the process of online live streaming due to the excessive pursuit of traffic by platforms or streamers. These phenomena are mainly manifested in the appearance of vulgar content, inducement to tip, false marketing, data fraud, and other sensitive information during the live streaming process, which seriously disrupt the order of the online ecosystem.

[0003] Currently, the supervision of live streaming chaos is usually carried out by network administrators and inspectors through manual review. This method can only take post-event measures against the platform or the streamer after the chaos occurs, such as shutting down the live streaming room or cutting off the video signal source. It cannot prevent the spread of sensitive information and has the problems of high supervision costs and low efficiency, which urgently need to be improved. Summary of the Invention

[0004] This application provides a method, apparatus, computing device, storage medium, and product for processing audio and video streams, in order to avoid the chaos in live streaming as seen in the prior art.

[0005] Firstly, embodiments of this application provide a method for processing audio and video streams, including: In response to the start of a live broadcast, obtain the live audio and video streams transmitted from the broadcaster's end; Using a domain-general model, the confidence level of the presence of sensitive information in the live audio and video streams across multiple domains is determined; Identify at least one target domain with a confidence level greater than the corresponding confidence threshold; Using domain-specific models corresponding to the at least one target domain, sensitive information is identified in the live audio and video stream to obtain sensitive information and sensitivity levels corresponding to the at least one target domain; the sensitivity level is determined based on the sensitive information of the corresponding target domain. The risk level of the live audio and video stream is determined based on the sensitivity level corresponding to each of the at least one target domain. The target handling strategy corresponding to the risk level is determined from the graded handling strategy for transmitting the live audio and video stream to the playback end, and the processing operation corresponding to the target handling strategy is executed for the live audio and video stream.

[0006] Secondly, embodiments of this application provide an audio / video stream processing apparatus, comprising: The acquisition module is used to acquire the live audio and video streams transmitted by the broadcaster in response to the live broadcast start event; The first model execution module is used to determine the confidence level of the existence of sensitive information in the live audio and video stream in multiple domains using a domain-general model. The domain determination module is used to determine at least one target domain with a confidence level greater than a corresponding confidence threshold; The second model execution module is used to identify sensitive information in the live audio and video stream using domain-specific models corresponding to the at least one target domain, so as to obtain sensitive information and sensitivity levels corresponding to the at least one target domain; the sensitivity level is determined based on the sensitive information of the corresponding target domain. The risk level determination module is used to determine the risk level of the live audio and video stream based on the sensitivity levels corresponding to the at least one target domain. The processing module is used to determine the target handling strategy corresponding to the risk level from the hierarchical handling strategy for transmitting the live audio and video stream to the playback end, and to perform the processing operation corresponding to the target handling strategy for the live audio and video stream.

[0007] Thirdly, this application provides a computing device, including a processing component and a storage component; the storage component stores a computing program; the computer program is invoked and executed by the processing component to implement the audio and video stream processing method of the first aspect described above.

[0008] Fourthly, this application provides a computer storage medium storing a computer program thereon, which, when executed by a processing component, implements the audio and video stream processing method described in the first aspect above.

[0009] Fifthly, this application provides a computer program product, including a computer program or instructions, which, when executed by a processing component, implement the audio and video stream processing method described in the first aspect above.

[0010] In response to a live broadcast launch event, this embodiment acquires the live audio and video stream transmitted from the broadcaster's end. It uses a domain-general model to determine the confidence level of sensitive information in the live audio and video stream across multiple domains. It identifies at least one target domain with a confidence level greater than a corresponding confidence threshold and uses a domain-specific model corresponding to each of these target domains to identify the sensitive information and sensitivity level of the live audio and video stream in each target domain. Based on the sensitivity level corresponding to each of the at least one target domain, it determines the risk level of the live audio and video stream. It then determines the target handling strategy corresponding to this risk level from a tiered handling strategy for transmitting the live audio and video stream to the playback end and executes the target handling strategy. In this embodiment, after acquiring the live audio and video stream transmitted from the broadcaster's end, it collaboratively uses a domain-general model and a domain-specific model to identify sensitive information in the live audio and video stream. Based on the risk level of the identified sensitive information, it employs a corresponding target handling strategy (such as direct transmission, transmission after filtering sensitive data, or stopping transmission) to execute the operation of transmitting the live video stream to the playback end. It can detect live streaming chaos before the live audio and video streams are transmitted from the live streaming terminal to the playback terminal, effectively preventing the spread and dissemination of sensitive information in the live audio and video streams. Compared with manual review, it improves processing efficiency and reduces processing costs.

[0011] Furthermore, this embodiment introduces a domain-wide general model for initial screening of sensitive information across multiple domains and sets confidence thresholds for each domain, fundamentally changing the paradigm of parallel full-scale computation across multiple domains and achieving a paradigm shift to serial condition-triggered computation. This allows the system to skip a large number of unnecessary domain-specific model calculations and perform in-depth sensitive information analysis only on high-confidence domains. This significantly reduces redundant computational overhead and solves efficiency and real-time bottlenecks under high concurrency. In addition, this embodiment employs a division of labor between the domain-wide general model and the domain-specific model, allowing the domain-specific model to focus on identifying in-depth sensitive information within its own domain, avoiding interference from irrelevant domain information, and improving the accuracy of sensitive information identification. Moreover, this embodiment designs graded handling strategies with different intensities for different risk levels, improving the synergistic optimization of handling accuracy and system adaptability. The solution in this embodiment also endows the system with unprecedented elasticity and scalability. By adjusting the confidence thresholds of each domain, it can automatically switch to a more focused domain during peak system loads, ensuring uninterrupted core business operations. This forms a complete autonomous system from perception to decision-making to execution and evolution, significantly reducing operating costs and the risk of live streaming chaos.

[0012] It is particularly important to point out that the above effects are not simply the sum of the effects of various technical features. It is precisely because the domain-general model identifies the confidence level of sensitive information existing in live audio and video streams across multiple domains that it provides a basis for the subsequent selective invocation of domain-specific models. In turn, the sensitive information and sensitivity levels identified by the domain-specific models provide a basis for selecting tiered handling strategies. This interconnected and positively cyclical collaborative design results in a qualitative leap from static rigidity to dynamic intelligence. Therefore, this embodiment is not a simple functional combination of a domain-general model and multiple domain-specific models, but rather a new working paradigm.

[0013] These or other aspects of this application will become more apparent from the description of the following embodiments. Attached Figure Description

[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This application provides a schematic diagram of the architecture of a live streaming system. Figure 2 A flowchart of one embodiment of an audio / video stream processing method provided in this application is shown; Figure 3 A flowchart illustrating the audio and video stream processing method provided in this application for a practical application scenario is shown; Figure 4 A schematic diagram of the structure of an embodiment of an audio / video stream processing apparatus provided in this application is shown; Figure 5 A schematic diagram of the structure of the computing device provided in this application is shown. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0016] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.

[0017] As the background technology introduction shows, current supervision of live streaming irregularities relies heavily on manual review. However, manual review typically only allows for reviewing recordings after the live stream ends or intervention based on user reports. This results in regulatory actions lagging behind the occurrence of live streaming irregularities, thus only enabling post-event handling of platforms or streamers and failing to prevent the spread of sensitive information in live audio and video streams. Furthermore, manual review suffers from low efficiency and high costs.

[0018] With the development of artificial intelligence technology, inventors conceived of introducing AI models to identify sensitive information. Before live audio and video streams are transmitted to the playback terminal, sensitive information (i.e., information that could lead to chaos in the live stream) is identified and filtered out. While this method can, to some extent, prevent the spread of sensitive information, reduce monitoring costs, and improve efficiency, it still suffers from the following problems: First, it suffers from low computational resource efficiency and real-time bottlenecks. This approach typically involves one or more AI models performing sensitive information identification in parallel across multiple domains, resulting in a significant waste of computational resources. In high-concurrency scenarios, this directly leads to increased processing latency and limited system throughput. To ensure real-time performance, model accuracy is often sacrificed, or high hardware costs are incurred. Second, it lacks fine-grained scheduling and dynamic adaptation capabilities. This approach relies on static resource allocation. It cannot intelligently determine the intensity and direction of computational resource allocation for different domains based on the real-time characteristics of the live audio and video stream. It cannot quickly allow normal live audio and video streams to pass, nor can it prioritize in-depth analysis of high-risk content when the system load is too high. The entire system lacks resilience and struggles to cope with challenges posed by traffic fluctuations and sensitive content in new areas.

[0019] Therefore, designing an audio / video stream processing method that can both guarantee and improve the accuracy of sensitive information identification, achieve intelligent on-demand allocation of computing resources, and automatically adapt to changes in content and system status has become the key to solving the problem. To this end, the inventors conducted a series of studies and proposed the solution of the embodiments of this application. The basic idea is as follows: In response to a live broadcast event, the live audio / video stream transmitted by the broadcaster is acquired; a domain-general model is used to determine the confidence level of sensitive information in the live audio / video stream across multiple domains; at least one target domain with a confidence level greater than the corresponding confidence threshold is identified; and a domain-specific model corresponding to each of the at least one target domain is used to identify the sensitive information and sensitivity level of the live audio / video stream in each of the at least one target domains; based on the sensitivity level corresponding to each of the at least one target domains, the risk level of the live audio / video stream is determined; a target handling strategy corresponding to the risk level is determined from the graded handling strategies for transmitting the live audio / video stream to the playback end; and the target handling strategy is executed. This embodiment, after acquiring the live audio and video stream transmitted from the broadcaster, uses a domain-general model and a domain-specific model in collaboration to identify sensitive information in the live audio and video stream. Based on the risk level of the identified sensitive information, it employs corresponding target handling strategies (such as direct transmission, transmission after filtering sensitive data, or stopping transmission) to execute the operation of transmitting the live video stream to the playback end. This enables the detection of live streaming irregularities before the live audio and video stream is transmitted from the broadcaster to the playback end, effectively preventing the spread and propagation of sensitive information in the live audio and video stream. Compared to manual review, it improves processing efficiency and reduces processing costs.

[0020] Furthermore, this embodiment introduces a domain-general model for initial screening of sensitive information across multiple domains and sets confidence thresholds for each domain, fundamentally changing the paradigm of parallel full-scale computation across multiple domains and achieving a paradigm shift to serial condition-triggered computation. This allows the system to skip a large number of unnecessary domain-specific model calculations and perform in-depth sensitive information analysis only on high-confidence domains. This significantly reduces redundant computational overhead and solves the efficiency and real-time bottlenecks under high concurrency. In addition, this embodiment employs a division of labor between the domain-general model and the domain-specific model, allowing the domain-specific model to focus on identifying in-depth sensitive information within its own domain, avoiding interference from irrelevant domain information, and improving the accuracy of sensitive information identification. Moreover, this embodiment designs graded handling strategies with different intensities for different risk levels, improving the synergistic optimization of handling accuracy and system adaptability.

[0021] This embodiment also endows the system with unprecedented elasticity and scalability. By adjusting the confidence thresholds of various domains, it can automatically switch to a more focused domain during peak system loads, ensuring uninterrupted core business operations. This forms a complete autonomous system from perception to decision-making to execution and evolution, significantly reducing operating costs and the risk of live-streaming chaos.

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] Figure 1 This is a schematic diagram of the architecture of a live streaming system provided in an embodiment of this application. The system may include a broadcaster terminal 101, a server terminal 102, and a playback terminal 103.

[0024] The broadcast client 101 and the server 102, as well as the playback client 103 and the server 102, can be connected via a network. The network provides the communication link between the broadcast client 101 and the server 102, and between the playback client 103 and the server 102. The network can include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0025] In a practical application, the broadcaster terminal 101 is responsible for real-time acquisition of the sound and image from the live broadcast location to obtain a live audio and video stream. It can also process the live audio and video stream, such as video processing (e.g., beautification, watermarking), audio and video encoding and compression, and audio and video encapsulation. After processing, the live audio and video stream can be transmitted to the server terminal 102. Live audio and video streams from different broadcasters can be distinguished by live streaming rooms. A broadcaster can first apply for a live streaming room from the server terminal 102 through the broadcaster terminal 101 to record and upload the live audio and video stream in real time. Users (i.e., viewers) can request to enter a specific live streaming room through the playback terminal 103. The server terminal 102 can then send the live audio and video stream of that live streaming room to the playback terminal 103, which can then play the live audio and video stream. Furthermore, as those skilled in the art will understand, the live audio and video stream may need to undergo encoding, transcoding, compression, and other processing before being uploaded to the server terminal 102. Similarly, the playback terminal 103 may need to perform decoding, decompression, and other processing before playing the live audio and video stream.

[0026] Among them, the playback terminal 103 can be configured in electronic devices such as mobile phones, tablets, computers, and smartwatches; the server terminal 102 can be implemented using a CDN (Content Delivery Network) system; and the broadcaster terminal 101 can be composed of electronic devices with acquisition functions and OBS (Open Broadcaster Software) streaming functions, such as smart devices with cameras, such as mobile phones, tablets, and smart wearable devices. Of course, this application is not limited to using the above-mentioned live streaming technology solutions to achieve network live streaming.

[0027] Both the playback client and the broadcaster client can be independent applications or functional modules integrated into other applications.

[0028] The implementation details of the technical solutions in the embodiments of this application are described in detail below.

[0029] Figure 2 This is a flowchart of an embodiment of an audio / video stream processing method provided in this application. The technical solution of this embodiment can be derived from... Figure 1 The server 102 of the live streaming system shown is executed. Figure 2 The audio and video stream processing method shown may include the following steps: S201, in response to the live broadcast start event, obtains the live audio and video stream transmitted by the broadcaster.

[0030] The "live broadcast start event" refers to an event triggered when a streamer begins a live broadcast, marking the official start of the broadcast and its availability for users to watch through the streaming platform. For example, the streamer's client detects when the streamer clicks the "Start Live Broadcast" button on the live broadcast platform, triggering the generation of a live broadcast start event and sending it to the server.

[0031] In practical applications, after the live streaming client triggers a live broadcast start event, the processing terminal responsible for managing the live streaming business logic (which can be the server in this embodiment or another processing terminal) first creates a live streaming room for the broadcaster and assigns a live streaming room address. The broadcaster uses this address to transmit the live audio and video streams generated in real time during the live broadcast to the server. In response to the live broadcast start event, the server queries the live streaming room address corresponding to the live broadcast start event, and then obtains the live audio and video streams transmitted in real time from the broadcaster's address through a pre-configured callback hook mechanism. It should be noted that this embodiment obtains the live audio and video streams generated during the broadcaster's live broadcast in real time, and performs the following processing operations for each obtained live audio and video stream.

[0032] S202, using a domain-general model, determines the confidence level that the live audio and video streams contain sensitive information in multiple domains.

[0033] The domain-specific model can be a general model for initial screening of sensitive information across multiple domains. It can be trained on sample audio and video streams containing sensitive information from multiple domains. Specific training methods will be described in detail in subsequent embodiments.

[0034] The "multi-domain" approach discussed in this article can be based on live-streaming scenarios. For example, it can include, but is not limited to: the e-commerce live-streaming domain, the gaming domain, the daily life domain, the education domain, and the politically sensitive domain, such as the political discussion domain. The confidence level of the presence of sensitive information in each domain can characterize the probability of such information existing in that domain.

[0035] The domain-general models mentioned in this paper, as well as the domain-specific models, audio general models, video general models, audio-specific models, video-specific models, information recognition models, and multimodal recognition models involved in subsequent embodiments, can refer to large-parameter models trained using massive amounts of data and powerful computing capabilities. These are machine learning models with complex structures capable of processing massive amounts of data and completing various complex tasks, such as natural language processing, computer vision, and speech recognition. They can include Large Language Models (LLMs) or Multimodal Large Models (MLMs). The models mentioned in this paper can be pre-trained models, which are then retrained through model fine-tuning to adapt to different processing tasks. This leverages the powerful capabilities of pre-trained models while also adapting to new data distributions, thus ensuring the model's generalization ability and reducing overfitting.

[0036] The domain-wide model of this embodiment can be used in conjunction with prompts to predict the confidence level of sensitive information existing in multiple domains. The prompts may include processing instructions, role information, processing requirements, thought chain information, and / or example data. Therefore, prompts can be generated based on processing instructions, role information, processing requirements, thought chain information, and / or example data. The content items included in the prompts can be set according to the actual situation, etc., and this application does not limit them.

[0037] The processing instructions explicitly tell the model what to do, such as "identify sensitive information from the input live audio and video streams from multiple domains and output the confidence scores of the presence of sensitive information in each domain." Role information can instruct the model to play a specific role, changing its professional nature, such as "You are a professional expert in regulating the chaos of online live streaming." Extraction requirements specify constraints, such as the range of confidence scores, to limit the model's operations. Thought chain information can guide the model to reason step by step. Example data can provide learning samples to help the model understand the operations performed. This embodiment utilizes the powerful semantic understanding and intelligent reasoning capabilities of a domain-general model to extract semantic features from the live audio and video streams. Based on these semantic features, it identifies the sensitive information that may exist in each domain and its confidence score. By combining each identified sensitive information and its confidence score in the domain, it determines the confidence score of the presence of sensitive information in that domain.

[0038] S203, identify at least one target domain with a confidence level greater than the corresponding confidence threshold.

[0039] The confidence thresholds for multiple fields can be pre-set, and the confidence thresholds for different fields can be the same or different.

[0040] In this embodiment, for each domain, it can be determined whether the confidence level determined by the general model of the domain is greater than the confidence level threshold corresponding to that domain. If so, then that domain is taken as the target domain.

[0041] Optionally, in practical application scenarios, there may be situations where the confidence level in a single domain is low and the distribution is relatively even. In such cases, there may be sensitive information that crosses domains or is ambiguous, and the general domain model may fail to identify it accurately. To solve this problem, the process of determining the target domain in this embodiment further includes: if there is no domain with a confidence level greater than the corresponding confidence level threshold, then from at least two domains with the highest confidence levels, the domain with a confidence level difference less than a first value is selected as the target domain.

[0042] Specifically, when the confidence scores of multiple domains are all below the corresponding confidence thresholds, the domains can be sorted in descending order of confidence. Then, the top N domains are selected, and the confidence differences between each pair are calculated. The two domains whose confidence differences are less than a first value (e.g., 0.03) are designated as the target domains. In this embodiment, when multiple domains do not meet the confidence thresholds, domains with similar confidence scores are designated as target domains. This triggers the corresponding domain-specific model to further identify sensitive information, thereby improving the accuracy of sensitive information identification and precisely preventing the spread and propagation of sensitive information in live audio and video streams.

[0043] It should be noted that if there are no domains with a confidence level greater than the corresponding confidence level threshold, and the confidence level difference between the top N domains is greater than the first value, then the live video stream can be considered to contain no sensitive information, and it can be directly transmitted to the playback end.

[0044] Optionally, in practical application scenarios, to further improve the accuracy of target domain determination, this embodiment can adjust the confidence thresholds for multiple domains in conjunction with the current live audio and video stream before performing this step. This can involve identifying whether a content mutation event exists in the live audio and video stream; if so, using an information recognition model, locating the content mutation moment corresponding to the content mutation event in the live audio and video stream, and determining the semantic information of the content mutation moment; and adjusting the confidence thresholds corresponding to each of the multiple domains based on the semantic information of the content mutation moment.

[0045] The content mutation events in the live audio and video streams may include, but are not limited to, sudden changes in speech tone in the audio data and / or sudden switching of image content in the video data. The content mutation moment can be the moment when the content mutation event occurs in the live audio and video stream, for example, it can be the start and / or end moment of the content mutation event.

[0046] This embodiment can detect sudden changes in speech intonation and / or abrupt changes in image content in a live audio / video stream using an event recognition model or audio / video recognition algorithm. If such changes occur, it indicates a content mutation event in the live audio / video stream. At this point, an information recognition model is used to identify the occurrence time of the content mutation event in the live audio / video stream, which is designated as the content mutation moment. Then, the contextual audio / video content at the content mutation moment is obtained and processed using natural language parsing to obtain semantic information about the content mutation moment. The confidence threshold for the semantic information association domain is then adjusted, for example, by increasing the confidence threshold for the semantic information association domain. This embodiment dynamically adjusts the confidence thresholds corresponding to multiple domains based on the contextual audio / video content at the content mutation moment, which can further improve the accuracy of target domain determination compared to static confidence thresholds.

[0047] It should be noted that since content mutation events in live audio and video streams are not inevitable, there may be situations where no content mutation events occur in the live audio and video stream. In such cases, it is not necessary to adjust the confidence thresholds corresponding to each domain separately. For example, the default confidence thresholds corresponding to each domain can be used to determine the target domain.

[0048] In practical applications, some sensitive information is cross-domain. A general domain model may not be able to identify all the domains corresponding to this sensitive information. To further improve the comprehensiveness and accuracy of target domain determination, this embodiment can pre-create a sensitive domain knowledge graph. In this knowledge graph, nodes represent domains, and edges indicate contribution relationships, i.e., the two domains corresponding to a given sensitive information. Since the general domain model inevitably needs to identify sensitive information across multiple domains when predicting the confidence level of sensitive information in multiple domains, it can, based on the sensitive domain knowledge graph, determine whether the sensitive information identified by the general domain model for that target domain also belongs to other domains. If so, those other domains are also included as the target domain.

[0049] S204, using domain-specific models corresponding to at least one target domain, perform sensitive information identification on the live audio and video stream to obtain sensitive information and sensitivity levels corresponding to at least one target domain, wherein the sensitivity level is determined based on the sensitive information of the corresponding target domain.

[0050] This embodiment pre-trains a domain-specific model for each of the multiple domains, designed to identify sensitive information and predict sensitivity levels within that domain. Compared to a general domain model, each domain-specific model can more accurately identify sensitive information within that domain. This domain-specific model can be trained based on sample audio and video streams containing sensitive information specific to the corresponding domain. The detailed training method will be described in detail in subsequent embodiments.

[0051] Sensitivity level is a quantitative indicator used to characterize the severity of illegal or inappropriate content (i.e., sensitive information) in live audio and video streams. It can be used to assess the possibility and degree of harm that it may cause to live stream chaos once it is spread.

[0052] In this embodiment, for each target domain determined in S203, the domain-specific model corresponding to that target domain is invoked to identify sensitive information in the live audio and video stream, and the sensitivity level corresponding to that target domain is determined based on the identified sensitive information.

[0053] The domain-specific model in this embodiment can also be used in conjunction with prompts to predict sensitive information and sensitivity levels corresponding to the target domain. The prompts may include processing instructions, role information, processing requirements, thought chain information, and / or example data. Therefore, prompts can be generated based on processing instructions, role information, processing requirements, thought chain information, and / or example data. The content items included in the prompts can be set according to the actual situation, etc., and this application does not limit them.

[0054] The processing instructions explicitly tell the model what to do, such as "identify sensitive information in the input audio and video stream and determine the sensitivity level based on the identified sensitive information." Role information instructs the model to play a specific role, changing its professional nature, such as "You are a professional expert in regulating the chaos of online live streaming." Extraction requirements specify constraints, such as constraints for different levels of high sensitivity, to limit the model's operations. Thought chain information guides the model to reason step by step. Sample data provides learning samples to help the model understand the operations performed. This embodiment utilizes the powerful semantic understanding and intelligent reasoning capabilities of a domain-specific model to accurately identify sensitive information in live audio and video streams and predict the sensitivity level based on the identified sensitive information.

[0055] Optionally, this embodiment can predict sensitivity levels based on sensitive information in many ways. One approach is to quantify the severity of each identified sensitive information based on its semantics, and use the level corresponding to the highest quantified value as the sensitivity level. Another approach is to quantify sensitivity based on the quantity of identified sensitive information. For example, a pre-defined correspondence between different sensitivity levels and the range of the number of sensitive information can be established, and the sensitivity level corresponding to the range of the number of identified sensitive information can be used as the sensitivity level for the target domain. Alternatively, both approaches can be combined, considering both the semantics and quantity of sensitive information to determine the sensitivity level. For example, the higher sensitivity level from the two approaches can be used as the final sensitivity level. No limitation is imposed on this approach.

[0056] S205, determine the risk level of the live audio and video stream based on the sensitivity level corresponding to at least one target area.

[0057] Among them, the risk level can characterize the overall harmful tendency of live audio and video streams after comprehensively considering the sensitivity levels of various fields. That is, it reflects the potential risk level that it may cause to platform order, user safety or social impact once it is spread from a global perspective.

[0058] Optionally, if there is only one target domain, the sensitivity level corresponding to that target domain can be directly used as the risk level of the live audio and video stream. If there are multiple target domains, the sensitivity level corresponding to the target domain with the highest sensitivity level can be used as the risk level of the live audio and video stream. Alternatively, the sensitivity levels of multiple target domains can be merged to obtain the risk level of the live audio and video stream. For example, if the sensitivity level is a rating, the risk level of the live audio and video stream can be obtained by averaging or weighted averaging the sensitivity level ratings of multiple target domains.

[0059] S206, determine the target handling strategy corresponding to the risk level from the graded handling strategy for transmitting live audio and video streams to the playback end, and perform the processing operation corresponding to the target handling strategy for the live audio and video streams.

[0060] This embodiment pre-sets corresponding handling strategies for different risk levels, thus obtaining a tiered handling strategy. This tiered handling strategy pertains to the transmission strategy for sending live audio and video streams to the playback end. As the risk level increases, the intensity of the corresponding handling strategy also increases accordingly, meaning the handling strategy is upgraded level by level. For example, for the first risk level (i.e., low risk level), the first handling strategy might be to directly transmit the live audio and video stream to the playback end; for the second risk level (i.e., medium risk level), the second handling strategy might be to filter sensitive information before transmitting to the playback end; and for the third risk level (i.e., high risk level), the third handling strategy might be to interrupt the transmission of the live audio and video stream. For the second and third risk levels, the corresponding handling strategies might also include sending risk warning information to the broadcaster.

[0061] Optionally, a target handling strategy corresponding to the risk level of the live audio and video stream can be selected from the pre-set multi-level handling strategies, and the corresponding processing operation can be performed on the live audio and video stream according to the target handling strategy.

[0062] In practical applications, the first risk level is lower than the second risk level, and the second risk level is lower than the third risk level. If the target handling strategy is the first handling strategy corresponding to the first risk level, then the live audio and video stream will be transmitted to the playback end. For example, this can be done through a communication connection between the server and the playback end, or the live audio and video stream can be injected into a playback address negotiated with the playback end, so that the playback end can pull the live audio and video stream from that playback address.

[0063] If the target handling strategy is the second handling strategy corresponding to the second risk level, then the live audio and video streams are desensitized based on sensitive information corresponding to at least one target domain, and the processed live audio and video streams are transmitted to the playback end. For example, sensitive information from each target domain can be filtered out from the live audio and video streams first, and then the processed live audio and video streams without sensitive information can be transmitted to the playback end in a similar manner as described above.

[0064] If the target handling strategy is the third handling strategy corresponding to the third risk level, then the transmission of the live audio and video stream will be interrupted. Specifically, this could mean interrupting the transmission of only the live audio and video stream acquired in this instance, or interrupting the transmission of all live audio and video streams corresponding to this live broadcast event and all subsequent broadcasts. In other words, the broadcast on the broadcaster's end will be stopped (i.e., the broadcaster's live video stream will no longer be acquired, nor will it be transmitted to the audience). Alternatively, the broadcast on the broadcaster's end can be stopped if the number of interruptions exceeds a preset threshold.

[0065] This embodiment adopts a tiered handling strategy of direct transmission, transmission after desensitization, and interruption of transmission. This strategy can accurately deal with content of different risk levels, avoid "one-size-fits-all" approach that may inadvertently harm normal live streaming, and effectively control the spread of high-risk content, thus balancing user experience and platform security.

[0066] This embodiment responds to a live broadcast start event by acquiring the live audio and video stream transmitted from the broadcaster's end. It uses a domain-general model to determine the confidence level of sensitive information in the live audio and video stream across multiple domains. It identifies at least one target domain with a confidence level greater than a corresponding confidence threshold and uses a domain-specific model corresponding to each target domain to identify the sensitive information and sensitivity level of the live audio and video stream in each target domain. Based on the sensitivity level corresponding to each target domain, it determines the risk level of the live audio and video stream. It then selects the target handling strategy corresponding to this risk level from the tiered handling strategies for transmitting the live audio and video stream to the playback end and executes the target handling strategy. In this embodiment, after acquiring the live audio and video stream transmitted from the broadcaster's end, it collaboratively uses a domain-general model and a domain-specific model to identify sensitive information in the live audio and video stream. Based on the risk level of the identified sensitive information, it employs a corresponding target handling strategy (such as direct transmission, transmission after filtering sensitive data, or stopping transmission) to execute the operation of transmitting the live video stream to the playback end. It can detect live streaming chaos before the live audio and video streams are transmitted from the live streaming terminal to the playback terminal, effectively preventing the spread and dissemination of sensitive information in the live audio and video streams. Compared with manual review, it improves processing efficiency and reduces processing costs.

[0067] Furthermore, this embodiment introduces a domain-general model for initial screening of sensitive information across multiple domains and sets confidence thresholds for each domain, fundamentally changing the paradigm of parallel full-scale computation across multiple domains and achieving a paradigm shift to serial condition-triggered computation. This allows the system to skip a large number of unnecessary domain-specific model calculations and perform in-depth sensitive information analysis only on high-confidence domains. This significantly reduces redundant computational overhead and solves the efficiency and real-time bottlenecks under high concurrency. In addition, this embodiment employs a division of labor between the domain-general model and the domain-specific model, allowing the domain-specific model to focus on identifying in-depth sensitive information within its own domain, avoiding interference from irrelevant domain information, and improving the accuracy of sensitive information identification. Moreover, this embodiment designs graded handling strategies with different intensities for different risk levels, improving the synergistic optimization of handling accuracy and system adaptability.

[0068] This embodiment also endows the system with unprecedented elasticity and scalability. By adjusting the confidence thresholds of various domains, it can automatically switch to a more focused domain during peak system loads, ensuring uninterrupted core business operations. This forms a complete autonomous system from perception to decision-making to execution and evolution, significantly reducing operating costs and the risk of live-streaming chaos.

[0069] To more comprehensively and accurately identify sensitive information in live audio and video streams, independent models can be set up for the audio and video data in the live audio and video streams to identify sensitive information. That is, the domain-general model in this embodiment can include a general audio model and a general video model; the domain-specific model corresponding to any domain includes a domain-specific audio model and a domain-specific video model. The functions and training methods of the general audio model and the general video model are similar to those of the domain-general model. The functions and training methods of the domain-specific audio model and the domain-specific video model are similar to those of the domain-specific model.

[0070] Accordingly, when performing the above S202 step, the first confidence level of the audio data in the live audio and video stream containing sensitive information in multiple domains can be determined by using a general audio model, and the second confidence level of the video data in the live audio and video stream containing sensitive information in multiple domains can be determined by using a general video model; based on the first confidence level and the second confidence level corresponding to the multiple domains, the confidence level of the live audio and video stream containing sensitive information in multiple domains can be determined.

[0071] Specifically, the general audio model and the general video model can adopt a similar approach to that described in S202 above, respectively determining the first confidence level of the existence of sensitive information in the audio data across multiple domains, and the second confidence level of the existence of sensitive information in the video data across multiple domains. Then, for each domain, the relatively higher confidence level between the first and second confidence levels for that domain can be used as the confidence level of the live audio / video stream containing sensitive information in that domain; alternatively, the first and second confidence levels can be fused (e.g., by averaging, weighted averaging, etc.) to obtain the confidence level of the live audio / video stream containing sensitive information in that domain.

[0072] Accordingly, when performing the above S204 step, for any target domain, the audio sensitive information in the audio data can be identified using the audio-specific model corresponding to the target domain, and a first level can be determined based on the audio sensitive information; the identification sensitive information in the video data can be identified using the video-specific model corresponding to the target domain, and a second level can be determined based on the video sensitive information; and the sensitivity level corresponding to the target domain can be determined based on the first level and the second level.

[0073] Specifically, the audio-specific model and the video-specific model can adopt a similar approach to that described in S204 above to identify the sensitive information (i.e., audio sensitive information) and the first level corresponding to the audio data in the live audio and video stream, and the sensitive information (i.e., video sensitive information) and the second level corresponding to the video data in the live audio and video stream. For each target domain, the audio sensitive information and video sensitive information of the target domain can be merged and deduplicated to obtain the sensitive information corresponding to the target domain in the live audio and video stream. When determining the sensitivity level based on the first level and the second level, the relatively higher level between the first level and the second level can be used as the sensitivity level corresponding to the target domain. If the first level and the second level are level scores, the level scores of the first level and the second level can also be fused (e.g., by averaging, weighted averaging, etc.) to obtain the sensitivity level corresponding to the target domain.

[0074] This embodiment adopts a two-modal processing approach of audio and video, which can give full play to the professionalism of each modality model, effectively make up for the blind spots of single-modality recognition, and significantly improve the accuracy and robustness of overall sensitive information recognition.

[0075] Considering the potential inconsistencies in the expression of sensitive information between audio and video data, this embodiment's domain-specific model further includes a multimodal recognition model to improve the accuracy of sensitivity levels in live audio and video streams. The multimodal recognition model can be a discriminator capable of comprehensively analyzing audio and video data and determining sensitivity levels. Accordingly, when determining the sensitivity level of a target domain based on a first and second level, it can be determined whether the difference between the first and second levels is greater than a second value. If not, the sensitivity level of the target domain is determined based on the first and second levels; if so, the multimodal recognition model is used to determine the sensitivity level of the target domain based on the live audio and video stream.

[0076] Specifically, when the difference between the first and second levels is less than or equal to the second value (i.e., the difference is small), the first and second levels can be directly fused as described in the above embodiments to determine the sensitivity level corresponding to the target domain, thereby improving the efficiency of sensitivity level determination. If the difference is greater than the second value (i.e., the difference is large), a multimodal recognition model is introduced to jointly analyze the context of audio and video data to more accurately determine the true sensitivity level. This method balances processing efficiency and discrimination accuracy, significantly improving the comprehensive perception capability for complex and hidden sensitive content while reducing false positives and false negatives.

[0077] In practical applications, to facilitate the tracing of sensitive information in live video streams, this embodiment can also store the live audio and video streams and their corresponding sensitive information after the domain-specific model identifies the sensitive information for the target domain. It can also further associate and store information such as the broadcaster's information and the live broadcast time of the live audio and video stream for subsequent querying and evidence collection by relevant organizations. Alternatively, the stored information can be proactively sent to the user terminals of relevant organizations; there are no limitations on this.

[0078] Next, the training process of the domain-general model and the domain-specific model in this embodiment will be described. Specifically, it includes the following sub-steps: Sub-step 1: Obtain the newly added sample audio and video streams, and label the sample sensitive information contained in the newly added sample audio and video streams, as well as the sample domain and sensitivity level labels corresponding to the sample sensitive information.

[0079] Among them, the newly added sample audio and video streams can be live audio and video stream samples added to the existing training data during the iterative training of the domain-general model and the domain-specific model, which are used to supplement or update the model's cognition.

[0080] Optionally, this embodiment can periodically extract audio and video streams that were not used during the previous training of the model from manually reviewed samples, new violation case databases, and audio and video streams reported by regulatory authorities. These newly added audio and video streams are used as new sample audio and video streams. The sensitive information contained in the new sample audio and video streams (i.e., sample sensitive information), the corresponding domain of the sample sensitive information (i.e., sample domain), and the sensitivity level of the corresponding domain (i.e., sensitivity level label) are labeled by manual means, algorithms, or models.

[0081] In practical application scenarios, this embodiment can also involve obtaining processing feedback information for the live audio and video stream. This processing feedback information is used to indicate whether the processing of the live audio and video stream meets expectations. If the processing feedback information indicates that the processing of the live audio and video stream does not meet expectations, then the live audio and video stream is used as a new sample audio and video stream. The processing feedback information can be the feedback given by the network administrators and inspectors of the live streaming platform after reviewing the processed live audio and video stream. It can also be the processing feedback information provided by users on the playback end through actions such as sending bullet comments, comments, or reports, indicating whether the processing of the live audio and video stream meets expectations. For example, if a user reports false information in the live audio and video stream through the playback end, the processing feedback information in this case indicates that the processing of the live audio and video stream does not meet expectations. This processing feedback information can also be the appeal information sent by the live streaming end after the live video stream transmission operation is stopped.

[0082] In this embodiment, after executing the processing operation corresponding to the target handling strategy on the live audio and video stream, if processing feedback information is obtained for the processed live video stream, it is determined whether the processing feedback information does not represent what is expected. That is, the processing of the live audio and video stream is unreasonable due to an error in the identification of sensitive information. For example, the live audio and video stream contains sensitive information but is directly transmitted to the playback end, or the live audio and video stream does not contain sensitive information but is interrupted during transmission. In the case where the processing feedback information does not represent what is expected, the live audio and video stream is used as a new sample audio and video stream. Based on the processing of the live audio and video stream that does not represent what is expected, the domain-general model and the corresponding domain-specific model are continuously fine-tuned, which can effectively make up for the gap between the training data and the actual online data. This mechanism not only improves the recognition accuracy of the domain-specific model and the domain-general model for new, marginal, or disguised sensitive content, but also enhances their adaptability and robustness in complex live streaming environments, realizing closed-loop optimization of "running and evolving at the same time".

[0083] Sub-step 2: Input the newly added sample audio and video stream into the domain general model, determine the first prediction sensitive information of the newly added sample audio and video stream in the sample domain, and adjust the model parameters corresponding to the sample domain in the domain general model according to the first prediction sensitive information and the sample sensitive information.

[0084] Although a domain-specific model is used to predict the confidence level of a live audio / video stream containing sensitive information in multiple domains, the key to improving the accuracy of confidence level prediction lies in guiding the domain-specific model to accurately learn and identify sensitive information in multiple domains. Therefore, in this embodiment, a new sample audio / video stream can be input into the domain-specific model. Since the sample domains corresponding to the sensitive information in the new sample audio / video stream have already been labeled, the domain-specific model can be controlled to identify the sensitive information belonging to the sample domain (i.e., the first predicted sensitive information) for the new sample audio data stream, and calculate the difference between the first predicted sensitive information and the sample sensitive information. Based on this difference, the model parameters related to the sample domain in the domain-specific model can be adjusted. For example, if the domain-specific model contains classification heads corresponding to multiple domains, and each domain classification head is used to identify the sensitive information corresponding to that domain and determine the confidence level, then the parameters of the classification head corresponding to that sample domain can be adjusted.

[0085] This embodiment allows newly added sample audio and video streams to cover sensitive data from multiple domains, thereby improving the accuracy of training domain-specific general models.

[0086] Sub-step 3: Input the newly added sample audio and video stream into the domain-specific model corresponding to the sample domain to obtain the second predicted sensitivity information and the predicted sensitivity level. Based on the difference between the second predicted sensitivity information and the sample sensitivity information, and the difference between the predicted sensitivity level and the sensitivity level label, adjust the model parameters of the domain-specific model.

[0087] This sub-step trains a domain-specific model corresponding to the labeled sample domain of the newly added audio stream. Specifically, it can involve inputting the newly added audio / video stream into the domain-specific model corresponding to the labeled sample domain, obtaining the sensitive information (i.e., second predicted sensitive information) identified by the model for the newly added audio stream in the corresponding sample domain, determining the sensitivity level (i.e., predicted sensitivity level) of the sample domain based on the second predicted sensitive information, and then determining the loss value based on the difference between the second predicted sensitive information and the sample sensitive information, as well as the difference between the predicted sensitivity level and the sensitivity level label. The model parameters of the domain-specific model are then adjusted based on the loss value. Optionally, this embodiment can freeze the bottom-level parameters and adjust the top-level adaptation layer parameters to reduce the computational cost of the model training process, avoid model overfitting, and improve training stability.

[0088] In this embodiment, sub-steps 2 and 3, during the training of the domain-general model and the domain-specific model, can employ online incremental learning or federated learning methods. Incremental training with incremental samples enhances the model's ability to learn new sensitive information and improves training efficiency. Furthermore, it enables timely fine-tuning of the domain-general model and the domain-specific model during online inference, periodically or when feedback representations do not meet expectations, thus resolving the issue of model update lag. In a practical application scenario, combined with Figure 3 This document introduces the audio and video stream processing process in practical application scenarios. Specifically, in the preparation phase, multiple sensitive information domains (i.e., multiple domains) need to be identified based on the live streaming scenario. For each domain, a general domain model and a specialized domain model are pre-trained. Each model can be assigned a corresponding model number, such as "domain-specific model 1." Each specialized domain model can also be tagged with a domain label and associated with its model number for quick and accurate invocation of the required specialized domain model. Furthermore, since the specialized domain model and the general domain model in this embodiment can be large models, standard prompts and instructions need to be compiled for them. To facilitate model invocation, the trained specialized domain model and the general domain model can be packaged into a model library and provided with an external calling interface for subsequent use.

[0089] After completing the above preparations, you can execute the command during live streaming. Figure 3The audio and video stream processing method is shown below. Specifically, when the server detects that the live broadcast has started (S301), it obtains the live broadcast address and, through the installed callback hook mechanism, intercepts the live audio and video stream transmitted in real time from the broadcaster's end (S302). Then, by calling the interface, it calls the corresponding model in the model library to perform the following intelligent recognition operation: First, it uses a general domain model to determine the confidence level of sensitive information in the live audio and video stream across multiple domains. Then, it identifies at least one target domain where the confidence level is greater than the corresponding confidence threshold. Next, it uses the domain-specific model corresponding to each of the at least one target domain to identify sensitive information in the live audio and video stream, thus obtaining the sensitive information and sensitivity level corresponding to each of the at least one target domain. If no sensitive information is identified, the live audio and video stream is considered normal, i.e., it passes the review and can be transmitted to the playback end (S303). At this time, viewers can see healthy live audio and video (S304). If sensitive information is detected, the live audio and video stream is intercepted (S305). Based on the sensitivity levels corresponding to at least one target domain, the risk level of the live audio and video stream is determined, and the target handling strategy corresponding to the risk level is determined from the tiered handling strategy. The processing operation corresponding to the target handling strategy is then executed. Specifically, if the risk level is not very high, such as low to medium, the live audio and video stream can be de-identified based on the sensitive information corresponding to at least one target domain (S307), and the processed live audio and video stream is transmitted to the playback terminal. At this time, viewers can also see healthy live audio and video (S304). If the risk level is high, such as high, the live stream can be shut down (S308).

[0090] To preserve operational evidence, this embodiment can also save sensitive information after detecting tangerine information for submission to relevant authorities (S306). For example, based on the domain of the sensitive information, a domain tag can be set for the sensitive information, and the domain tag, sensitive information, live audio and video stream, live time, and anchor information can be associated and stored in a database. This facilitates subsequent searches using any of the domain tag, sensitive information, anchor information, or live time as an index, and can be submitted to relevant authorities for evidence when necessary. This embodiment's solution can prevent the spread of harmful sensitive information in online live streaming. It also allows security agencies to easily retrieve relevant sensitive information as evidence. Furthermore, this process is model-based, requiring no manual intervention, reducing costs while improving efficiency.

[0091] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.

[0092] It should be noted that some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The sequence numbers of the operations are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should also be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0093] Figure 4 A schematic diagram of an audio / video stream processing apparatus provided for an exemplary embodiment of this application is shown. The apparatus includes: The acquisition module 401 is used to acquire the live audio and video stream transmitted by the broadcaster in response to the live broadcast start event; The first model running module 402 is used to determine the confidence level of the existence of sensitive information in multiple domains using a domain-general model; The domain determination module 403 is used to determine at least one target domain with a confidence level greater than a corresponding confidence threshold; The second model execution module 404 is used to identify sensitive information in the live audio and video stream using domain-specific models corresponding to the at least one target domain, so as to obtain sensitive information and sensitivity level corresponding to the at least one target domain; the sensitivity level is determined based on the sensitive information of the corresponding target domain. The risk level determination module 405 is used to determine the risk level of the live audio and video stream based on the sensitivity levels corresponding to the at least one target domain. The processing module 406 is used to determine the target handling strategy corresponding to the risk level from the hierarchical handling strategy for transmitting the live audio and video stream to the playback end, and to perform the processing operation corresponding to the target handling strategy for the live audio and video stream.

[0094] In some embodiments, the apparatus further includes: a confidence adjustment module, configured to identify whether a content mutation event exists in the live audio and video stream; if so, to use an information recognition model to locate the content mutation time corresponding to the content mutation event in the live audio and video stream, and to determine the semantic information of the content mutation time; and to adjust the confidence thresholds corresponding to the multiple domains based on the semantic information of the content mutation time.

[0095] In some embodiments, the domain determination module 403 is further configured to select, from at least two domains ranked first in terms of confidence, a domain with a confidence difference less than a first value as the target domain if there is no domain with a confidence level greater than the corresponding confidence threshold.

[0096] In some embodiments, the domain-general model includes an audio general model and a video general model; the domain-specific model corresponding to any domain includes an audio specific model and a video specific model. The first model running module 402 is specifically used to determine, using the audio general model, the first confidence level that the audio data in the live audio and video stream contains sensitive information in multiple domains, and to determine, using the video general model, the second confidence level that the video data in the live audio and video stream contains sensitive information in multiple domains; and to determine the confidence level that the live audio and video stream contains sensitive information in multiple domains based on the first confidence level and the second confidence level corresponding to each of the multiple domains. The second model running module 404 is specifically used to identify audio sensitive information in the audio data using an audio-specific model corresponding to the target domain for any target domain, and determine a first level based on the audio sensitive information; and to identify identification sensitive information in the video data using a video-specific model corresponding to the target domain, and determine a second level based on the video sensitive information; and to determine the sensitivity level corresponding to the target domain based on the first level and the second level.

[0097] In some embodiments, the domain-specific model further includes a multimodal recognition model; the level determination module 405 is specifically used to determine whether the difference between the first level and the second level is greater than a second value; if not, determine the sensitivity level corresponding to the target domain based on the first level and the second level; if yes, use the multimodal recognition model to determine the sensitivity level corresponding to the target domain based on the live audio and video stream.

[0098] In some embodiments, the processing module 406 is specifically configured to: transmit the live audio / video stream to the playback end if the target handling strategy is a first handling strategy corresponding to a first risk level; perform desensitization processing on the live audio / video stream based on sensitive information corresponding to at least one target domain if the target handling strategy is a second handling strategy corresponding to a second risk level, and transmit the processed live audio / video stream to the playback end if the target handling strategy is a third handling strategy corresponding to a third risk level; and interrupt the transmission operation of the live audio / video stream if the target handling strategy is a third handling strategy corresponding to a third risk level; wherein the first risk level is lower than the second risk level, and the second risk level is lower than the third risk level.

[0099] In some embodiments, the apparatus further includes a sample acquisition module, configured to acquire newly added sample audio / video streams and label the sample sensitive information contained in the newly added sample audio / video streams, as well as the sample domain and sensitivity level label corresponding to the sample sensitive information; a model training module is configured to input the newly added sample audio / video streams into a domain-specific model, determine the first predicted sensitive information present in the sample domain of the newly added sample audio / video streams, and adjust the model parameters corresponding to the sample domain in the domain-specific model according to the first predicted sensitive information and the sample sensitive information; or, input the newly added sample audio / video streams into a domain-specific model corresponding to the sample domain to obtain second predicted sensitive information and a predicted sensitivity level, and adjust the model parameters of the domain-specific model based on the difference between the second predicted sensitive information and the sample sensitive information, and the difference between the predicted sensitivity level and the sensitivity level label.

[0100] In some embodiments, the sample acquisition module is specifically used to acquire processing feedback information for the live audio and video stream; the processing feedback information is used to characterize whether the processing of the live audio and video stream meets expectations; if the processing feedback information characterizes that the processing of the live audio and video stream does not meet expectations, then the live audio and video stream is used as a new sample audio and video stream.

[0101] Figure 4 The aforementioned audio and video stream processing device can perform... Figure 2 The implementation principle and technical effects of the audio and video stream processing method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the audio and video stream processing apparatus in the above embodiments perform operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0102] Figure 5 This is a schematic diagram of the structure of one embodiment of a computing device provided in this application. Figure 5 As shown, in practice, the computing device may include a storage component 501 and a processing component 502.

[0103] Storage component 501 is used to store computer programs and can be configured to store various other data to support operation on a computing device. Examples of this data include instructions for any application or method used to operate on the computing device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0104] Processing component 502, coupled to storage component 501, is used to execute computer programs in storage component 501 for implementing, etc. Figure 2 The method for processing audio and video streams is shown.

[0105] Furthermore, such as Figure 5 As shown, the computing device may also include other components such as a communication component 503, a display component 504, a power supply component 505, and an audio component 506. Figure 5 The diagram only shows some components and does not mean that the device includes only these components. Figure 5 The components shown. Additionally... Figure 5 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the computing device. The computing device in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT (Internet of Things) device, or a server-side device such as a conventional server, cloud server, or server array. If the computing device in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 5 The components within the dashed box; if the computing device in this embodiment is implemented as a conventional server, cloud server, or server array, etc., it may be omitted. Figure 5 The component within the dashed box.

[0106] The processing component described above includes one or more processors to execute computer instructions to complete all or part of the steps in the method described above. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the method described above.

[0107] The aforementioned storage components can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0108] The aforementioned communication component is configured to facilitate wired or wireless communication between the device housing the communication component and other devices. The device housing the communication component can access wireless networks based on communication standards, such as mobile communication networks, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.

[0109] The aforementioned display components may include a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0110] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0111] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0112] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.

[0113] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0114] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0115] Finally, it should be noted that the above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for processing audio and video streams, characterized in that, include: In response to the start of a live broadcast, obtain the live audio and video streams transmitted from the broadcaster's end; Using a domain-general model, the confidence level of the presence of sensitive information in the live audio and video streams across multiple domains is determined; Identify at least one target domain with a confidence level greater than the corresponding confidence threshold; Using domain-specific models corresponding to the at least one target domain, sensitive information is identified in the live audio and video stream to obtain the sensitive information and sensitivity level corresponding to the at least one target domain. The sensitivity level is determined based on sensitive information in the corresponding target domain; The risk level of the live audio and video stream is determined based on the sensitivity level corresponding to each of the at least one target domain. The target handling strategy corresponding to the risk level is determined from the graded handling strategy for transmitting the live audio and video stream to the playback end, and the processing operation corresponding to the target handling strategy is executed for the live audio and video stream.

2. The method according to claim 1, characterized in that, Before identifying at least one target sensitive area with a confidence level greater than the corresponding confidence threshold, the process also includes: Identify whether there are any content abrupt events in the live audio and video stream; If it exists, the information recognition model is used to locate the content mutation time corresponding to the content mutation event in the live audio and video stream, and to determine the semantic information of the content mutation time. Based on the semantic information at the moment of content mutation, the confidence thresholds corresponding to the multiple domains are adjusted.

3. The method according to claim 1, characterized in that, Also includes: If there is no domain with a confidence level greater than the corresponding confidence level threshold, then select the domain with a confidence level difference less than the first value from at least two domains with the highest confidence level ranking as the target domain.

4. The method according to claim 1, characterized in that, The domain-specific model includes a general audio model and a general video model; the domain-specific model corresponding to any domain includes a specific audio model and a specific video model. The method of using a domain-general model to determine the confidence level of the presence of sensitive information in the live audio and video stream across multiple domains includes: The audio general model is used to determine the first confidence level that the audio data in the live audio and video stream contains sensitive information in multiple domains, and the video general model is used to determine the second confidence level that the video data in the live audio and video stream contains sensitive information in multiple domains. Based on the first and second confidence levels corresponding to the multiple domains respectively, the confidence level of the live audio and video streams in the multiple domains is determined; The step of using domain-specific models corresponding to the at least one target domain to identify sensitive information in the live audio and video stream, so as to obtain the sensitive information and sensitivity level corresponding to the at least one target domain, includes: For any target domain, an audio-specific model corresponding to the target domain is used to identify audio-sensitive information in the audio data and a first level is determined based on the audio-sensitive information; and a video-specific model corresponding to the target domain is used to identify identification-sensitive information in the video data and a second level is determined based on the video-sensitive information. The sensitivity level corresponding to the target area is determined based on the first level and the second level.

5. The method according to claim 4, characterized in that, The domain-specific model further includes a multimodal recognition model; determining the sensitivity level corresponding to the target domain based on the first level and the second level includes: Determine whether the difference between the first level and the second level is greater than the second value; If not, determine the sensitivity level corresponding to the target area based on the first level and the second level; If so, the sensitivity level corresponding to the target domain is determined based on the live audio and video stream using the multimodal recognition model.

6. The method according to claim 1, characterized in that, The processing operations corresponding to the target processing strategy for the live audio and video stream include: If the target handling strategy is the first handling strategy corresponding to the first risk level, then the live audio and video stream will be transmitted to the playback terminal. If the target handling strategy is the second handling strategy corresponding to the second risk level, then the live audio and video stream is desensitized based on the sensitive information corresponding to at least one target domain, and the processed live audio and video stream is transmitted to the playback terminal. If the target handling strategy is the third handling strategy corresponding to the third risk level, then the transmission operation of the live audio and video stream is interrupted; wherein the first risk level is lower than the second risk level, and the second risk level is lower than the third risk level.

7. The method according to claim 1, characterized in that, Also includes: Acquire newly added sample audio and video streams, and label the sample sensitive information contained in the newly added sample audio and video streams, as well as the sample domain and sensitivity level labels corresponding to the sample sensitive information; The newly added sample audio and video streams are input into the domain general model to determine the first prediction sensitive information of the newly added sample audio and video streams in the sample domain, and the model parameters corresponding to the sample domain in the domain general model are adjusted according to the first prediction sensitive information and the sample sensitive information. or, The newly added sample audio and video streams are input into the domain-specific model corresponding to the sample domain to obtain the second predicted sensitivity information and the predicted sensitivity level. Based on the difference between the second predicted sensitivity information and the sample sensitivity information, and the difference between the predicted sensitivity level and the sensitivity level label, the model parameters of the domain-specific model are adjusted.

8. The method according to claim 7, characterized in that, The acquisition of newly added sample audio and video streams includes: Obtain processing feedback information for the live audio and video stream; the processing feedback information is used to characterize whether the processing of the live audio and video stream meets expectations; If the processing feedback information indicates that the processing of the live audio and video stream does not meet expectations, then the live audio and video stream will be used as a new sample audio and video stream.

9. An audio / video stream processing apparatus, characterized in that, include: The acquisition module is used to acquire the live audio and video streams transmitted by the broadcaster in response to the live broadcast start event; The first model execution module is used to determine the confidence level of the existence of sensitive information in the live audio and video stream in multiple domains using a domain-general model. The domain determination module is used to determine at least one target domain with a confidence level greater than a corresponding confidence threshold; The second model running module is used to identify sensitive information in the live audio and video stream by using the domain-specific models corresponding to the at least one target domain, so as to obtain the sensitive information and sensitivity level corresponding to the at least one target domain. The sensitivity level is determined based on sensitive information in the corresponding target area; The risk level determination module is used to determine the risk level of the live audio and video stream based on the sensitivity levels corresponding to the at least one target domain. The processing module is used to determine the target handling strategy corresponding to the risk level from the hierarchical handling strategy for transmitting the live audio and video stream to the playback end, and to perform the processing operation corresponding to the target handling strategy for the live audio and video stream.

10. A computing device, characterized in that, This includes processing components and storage components; The storage component stores a computer program; the computer program is invoked and executed by the processing component to implement the audio and video stream processing method as described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by the processing component, implements the audio and video stream processing method as described in any one of claims 1-8.

12. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processing component, implement the method for processing audio and video streams as described in any one of claims 1-8.