Audio material auditing and storing method and system

By analyzing the metadata and waveform characteristics of audio materials, and combining historical review data with multimodal analysis, review reports and strategies are generated. This solves the problem that traditional audio material review methods cannot adapt to diverse dissemination scenarios, and enables accurate and compliant storage and effective dissemination of audio materials.

CN121765111BActive Publication Date: 2026-05-29HANGZHOU XIAOSHAN HUA NUMBER OF DIGITAL TV CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU XIAOSHAN HUA NUMBER OF DIGITAL TV CO LTD
Filing Date
2026-03-03
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional methods for reviewing and adding audio materials to the database are unable to adapt to diverse dissemination scenarios, resulting in inaccurate reviews, an inability to effectively screen out compliant audio materials, and an impact on the quality of materials added to the database and subsequent dissemination effects.

Method used

By analyzing the metadata information and audio waveform features of audio materials, the dissemination scenario tags are determined. Historical review data is used to calculate the feature correlation coefficient and the distribution trend of violation features, generate review reports and formulate adaptation or rectification strategies, and combine multimodal semantic analysis and smart contract technology to make compliance judgments.

Benefits of technology

It has improved the targeting and accuracy of audio material review, ensuring that the materials entering the database comply with regulations on various channels, providing clear review guidance and rectification directions, and guaranteeing the compliant entry and effective dissemination of audio materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765111B_ABST
    Figure CN121765111B_ABST
Patent Text Reader

Abstract

The application discloses an audio material auditing and warehousing method and system, relates to the technical field of audio auditing, and has the technical scheme as follows: obtaining a target audio material to be audited, and performing compliance screening on the target audio material; if the compliance of the target audio material cannot be determined, determining a propagation scene label after analyzing metadata information and audio waveform features of the target audio material; if the propagation scene label belongs to a cross-channel propagation type, extracting cross-channel audit records of audio of the same type from a historical cross-channel audit database, and calculating a feature correlation coefficient between the historical audio and the target audio material; separating a to-be-determined audio segment from the target audio material, generating an audit report after judging the multi-channel compliance of the to-be-determined audio segment according to the feature correlation coefficient and a real-time audit rule library of each channel, and formulating an adaptive strategy to complete warehousing determination; and the effect is to provide strong support for the standardized management and effective propagation of audio materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio review technology, and more specifically, to a method and system for reviewing and storing audio materials. Background Technology

[0002] In the field of audio material review and warehousing, traditional methods process audio materials according to fixed review rules. However, audio materials are distributed in diverse scenarios; some need to be distributed across multiple channels, while others need to be distributed for specific scenarios. Fixed rules are difficult to adapt to these complex and diverse scenarios. Audio materials distributed across multiple channels may have different review requirements and standards, and fixed rules cannot fully cover the differences between channels. For audio materials used in targeted scenarios, the violation characteristics of specific scenarios are unique and dynamically changing, and fixed rules cannot accurately capture the changing trends of these characteristics. This leads to inaccurate reviews and an inability to effectively screen out compliant audio materials, thus affecting the quality of audio materials entering the warehouse and the subsequent dissemination effect. Summary of the Invention

[0003] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a method and system for reviewing and storing audio materials.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A method for reviewing and adding audio materials to a database, comprising the following steps:

[0006] Obtain the target audio materials to be reviewed and conduct compliance screening on the target audio materials;

[0007] If the compliance of the target audio material cannot be determined, the dissemination scenario label is determined after parsing the metadata information and audio waveform characteristics of the target audio material.

[0008] If the dissemination scenario tag belongs to the cross-channel dissemination type, then extract the cross-channel review records of the same type of audio from the historical cross-channel review database, calculate the feature correlation coefficient between the historical audio and the target audio material, separate the audio segment to be judged from the target audio material, and generate a review report after judging the multi-channel compliance of the audio segment to be judged based on the feature correlation coefficient and the real-time review rule library of each channel, and formulate an adaptation strategy to complete the entry judgment.

[0009] If the dissemination scenario tag belongs to the targeted scenario application type, then audio materials of the same scenario are extracted from the historical review data of the targeted scenario, and the historical key feature distribution trend of the illegal content is statistically analyzed based on the audio materials of the same scenario; the current illegal feature correlation trend of the target audio material is judged, and the compliance is judged based on the matching overlap between the current illegal feature correlation trend and the historical key feature distribution trend. After that, an review report is generated, and a rectification strategy is formulated to complete the entry into the database.

[0010] Preferably, compliance screening of the target audio material is performed, specifically including the following steps:

[0011] The temporal signal of the target audio material is mapped to the semantic feature space through nonlinear tensor decomposition to generate a semantic differential manifold;

[0012] The text transcription content corresponding to the target audio material is projected onto the feature space to form a text semantic embedding layer;

[0013] The correspondence deviation between the semantic differential manifold and the text semantic embedding layer is calculated to obtain a multimodal semantic manifold containing both acoustic and textual semantic information;

[0014] A semantic normative field is generated based on a pre-set review rule base, and the semantic normative field is coupled to a multimodal semantic manifold to excite symmetry transformations;

[0015] Calculate the topological load and curvature tensor of the semantically differentiable manifold after symmetry transformation to identify topological defects in the semantically differentiable manifold;

[0016] Compliance screening of target audio materials is performed based on topological defects.

[0017] Preferably, the propagation scenario label is determined after parsing the metadata information and audio waveform features of the target audio material, specifically including the following steps:

[0018] The metadata information and audio waveform features of the target audio material are processed to obtain a multimodal semantic field;

[0019] Obtain the feature solution set of the multimodal semantic field; wherein the feature solution set contains multiple intrinsic modalities, and each intrinsic modality corresponds to a scene matching mode;

[0020] Based on a predefined scene classification system, multiple scene matching models corresponding to different propagation scenarios are constructed; the feature solution set is input into the scene matching model to obtain the degree of matching between the audio material and each scene;

[0021] The dissemination scene tags are determined based on the degree of matching between the audio material and each scene.

[0022] Preferably, the metadata information and audio waveform features of the target audio material are processed to obtain a multimodal semantic field, specifically including the following steps:

[0023] The metadata information of the target audio material is transformed and mapped into a contextual scalar field;

[0024] The audio waveform features are mapped into an acoustic vector field through spectral analysis;

[0025] A multimodal semantic field is generated by coupling a contextual scalar field with an acoustic vector field.

[0026] Preferably, cross-channel review records of the same type of audio are extracted from the historical cross-channel review database, and the feature correlation coefficient between the historical audio and the target audio material is calculated. This specifically includes the following steps:

[0027] Extract cross-channel review records of similar audio from the historical cross-channel review database; among them, the cross-channel review records of similar audio have the same dissemination scenario tags as the target audio material;

[0028] Each cross-channel review record is mapped to a corresponding historical review eigenstate, and all historical review eigenstates constitute a cluster of historical review eigenstates.

[0029] The feature information of the target audio material is mapped to the target review eigenstate; the feature correlation coefficient is obtained by performing operations on the target review eigenstate and the historical review eigenstate cluster.

[0030] Preferably, after determining the multi-channel compliance of the audio segment to be judged based on the feature correlation coefficient and the real-time review rule base of each channel, a review report is generated, which specifically includes the following steps:

[0031] After digital twin modeling the feature correlation coefficient with the voiceprint features and emotional tendency features of the audio segment to be judged, a corresponding virtual audio object is generated.

[0032] Deploy the real-time review rule base of each channel to the digital twin space in the form of smart contracts to build a feature recognition model;

[0033] The virtual audio object is input into the feature recognition model to obtain preliminary risk assessment results for violations;

[0034] The confidence level of the compliance decision-making path for the audio segment to be judged in various channels is obtained based on the preliminary risk assessment results.

[0035] The compliance status quantification indicators for each channel are obtained based on the confidence level of the compliance decision-making path, the dynamic weight of the authority of the channel review, and the real-time correlation characteristics of the audio segments to be judged.

[0036] The compliance status quantitative indicators of each channel are compared with the preset judgment threshold. If the compliance status quantitative indicators are greater than or equal to the threshold, the channel is judged to be compliant; otherwise, it is judged to be non-compliant.

[0037] Generate an audit report that includes an audio feature summary and a risk level identifier.

[0038] Preferably, audio materials from the same scene are extracted from the historical review data of the targeted scene, and the historical key feature distribution trend of the illegal content is statistically analyzed based on the audio materials from the same scene. This specifically includes the following steps:

[0039] Spatiotemporal annotations are performed on audio materials from the same scene extracted from historical review data of targeted scenes, and spatiotemporal data tags containing timestamps and geographical coordinates are constructed by combining the dissemination area and playback time period;

[0040] A generative adversarial network and variational autoencoder fusion model is used to enhance the features of audio materials in the same scene. Virtual violation audio features are generated through adversarial training, and the original features together constitute the enhanced feature dataset.

[0041] The enhanced feature dataset and spatiotemporal data labels are input into the feature association mining model to obtain the feature association matrix;

[0042] A reinforcement learning framework is adopted, with the feature association matrix as the state space, the feature selection operation as the action space, and the accuracy of violation feature recognition as the reward function, to train the optimal feature selection strategy.

[0043] Key violation features are selected from the enhanced feature dataset based on the optimal feature selection strategy, and time series decomposition is performed on the key violation features.

[0044] We model the key violation features after time series decomposition and capture the dependencies of key violation features in different time periods through a multi-head attention mechanism.

[0045] Calculate the dynamic fluctuation parameters and structural complexity measures of the time series of key violation features;

[0046] The historical key feature distribution trend of the illegal content is obtained based on dynamic fluctuation parameters, structural complexity metrics, and feature dependencies.

[0047] Preferably, after determining compliance based on the degree of overlap between the current trend of violation characteristics and the distribution trend of historical key characteristics, an audit report is generated, which specifically includes the following steps:

[0048] After mapping the current violation feature association trend and the historical key feature distribution trend to the topological space respectively, the feature topological structure is constructed, and the topological algebraic representation value of the feature topological structure is obtained.

[0049] Semantic enhancement processing is performed on the algebraic representation values ​​of topological structures, transforming topological features into feature vectors containing semantic information;

[0050] A dynamic weighted matching model based on attention mechanism is established. The semantic similarity, time span weight, and scene relevance weight of the feature vector are used as input. The weights of each dimension are adaptively adjusted through a multi-head attention mechanism, and the weighted matching degree between the current illegal feature association trend and the historical key feature distribution trend is calculated.

[0051] A fuzzy comprehensive evaluation model is introduced to perform a fuzzy mapping between the weighted matching degree and the preset multi-level compliance threshold range, and the membership degree function is combined to determine the degree of membership of the current audio material in different compliance levels.

[0052] A compliance probability distribution vector is constructed based on the degree of membership, and the probability distribution vector is processed to generate a compliance probability distribution map;

[0053] The final compliance status of the audio material is determined based on the compliance probability distribution map and the preset decision-making strategy;

[0054] Identify key violation segments that correlate trends with the violation characteristics of non-compliant audio materials;

[0055] Generate an audit report that includes basic information about the audio material and the compliance assessment results.

[0056] An audio material review and storage system includes:

[0057] Acquisition Module: Acquires target audio materials to be reviewed and performs compliance screening on the target audio materials;

[0058] Parsing module: If the compliance of the target audio material cannot be determined, the metadata information and audio waveform characteristics of the target audio material are parsed to determine the propagation scenario label;

[0059] First review module: If the dissemination scenario tag belongs to the cross-channel dissemination type, then extract the cross-channel review records of the same type of audio from the historical cross-channel review database, calculate the feature correlation coefficient between the historical audio and the target audio material; separate the audio segment to be judged from the target audio material, and generate a review report after judging the multi-channel compliance of the audio segment to be judged based on the feature correlation coefficient and the real-time review rule library of each channel, and formulate an adaptation strategy to complete the entry judgment.

[0060] The second review module: If the dissemination scenario tag belongs to the targeted scenario application type, then extract audio materials of the same scenario from the historical review data of the targeted scenario, and statistically analyze the historical key feature distribution trend of the illegal content based on the audio materials of the same scenario; determine the current illegal feature correlation trend of the target audio material, and determine compliance based on the matching overlap between the current illegal feature correlation trend and the historical key feature distribution trend, and generate a review report, and formulate rectification strategies to complete the entry into the database.

[0061] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for reviewing and storing audio materials.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] This invention improves review efficiency by acquiring target audio materials for review and performing compliance screening, enabling a preliminary judgment on obviously compliant or non-compliant audio materials. When the compliance of a target audio material cannot be directly determined, its metadata information and audio waveform characteristics are analyzed to determine the dissemination scenario label, clarifying its dissemination scenario type and providing a basis for subsequent targeted review, making the review more targeted and accurate. For audio materials whose dissemination scenario label belongs to the cross-channel dissemination type, cross-channel review records of the same type of audio are extracted from the historical cross-channel review database, the feature correlation coefficient between historical audio and target audio materials is calculated, and the multi-channel compliance of the audio segment to be judged is determined in combination with the real-time review rule base of each channel. It fully leverages the experiential value of historical review data, using historical data to assist current reviews and improve the reliability of the review process. The multi-channel compliance assessment comprehensively considers the performance of audio materials across different distribution channels, ensuring that the audio materials in the database comply with regulations on all channels and avoiding violations caused by channel differences. Simultaneously, the generated review reports and established adaptation strategies provide clear guidance for the database entry of audio materials and their subsequent cross-channel dissemination. For audio materials whose dissemination scenario tags belong to targeted scenario application types, it extracts audio materials from the same scenario from the historical review data of the targeted scenario, statistically analyzes the historical key feature distribution trends of the violation content, and matches and determines the overlap degree with the current violation features of the target audio material. It can deeply explore the violation patterns in targeted scenarios, judging the compliance of target audio materials by comparing historical trends with current trends. The generated review reports and established rectification strategies help audio materials in targeted scenarios to promptly correct violations, ensuring their compliant database entry, and also provide a historical basis and rectification direction for the subsequent review of audio materials in the same scenario, providing strong support for the standardized management and effective dissemination of audio materials. Attached Figure Description

[0064] Figure 1 This invention provides a schematic diagram illustrating the steps of an audio material review and storage method.

[0065] Figure 2 This invention provides a schematic diagram illustrating the steps involved in obtaining dissemination scenario tags in an audio material review and database entry method.

[0066] Figure 3 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention.

[0067] 610. Processor; 620. Communication interface; 630. Memory; 640. Communication bus. Detailed Implementation

[0068] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0069] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0070] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.

[0071] Reference Figures 1-3 As shown.

[0072] The embodiments further illustrate the audio material review and storage method and system proposed in this invention.

[0073] A method for reviewing and adding audio materials to a database, comprising the following steps:

[0074] Obtain the target audio materials to be reviewed and conduct compliance screening on the target audio materials;

[0075] If the compliance of the target audio material cannot be determined, the dissemination scenario label is determined after parsing the metadata information and audio waveform characteristics of the target audio material.

[0076] If the dissemination scenario tag belongs to the cross-channel dissemination type, then extract the cross-channel review records of the same type of audio from the historical cross-channel review database, calculate the feature correlation coefficient between the historical audio and the target audio material, separate the audio segment to be judged from the target audio material, and generate a review report after judging the multi-channel compliance of the audio segment to be judged based on the feature correlation coefficient and the real-time review rule library of each channel, and formulate an adaptation strategy to complete the entry judgment.

[0077] Develop an inbound strategy that adapts to the rules of multiple channels, such as making targeted modifications or restricting the distribution of non-compliant content on some channels, to ensure that the audio meets the requirements of each channel before being judged for inbound.

[0078] If the dissemination scenario tag belongs to the targeted scenario application type, then audio materials of the same scenario are extracted from the historical review data of the targeted scenario, and the historical key feature distribution trend of the illegal content is statistically analyzed based on the audio materials of the same scenario; the current illegal feature correlation trend of the target audio material is judged, and the compliance is judged based on the matching overlap between the current illegal feature correlation trend and the historical key feature distribution trend. After that, an review report is generated, and a rectification strategy is formulated to complete the entry into the database.

[0079] If non-compliance is determined, key violation segments are identified, an audit report containing basic audio information and compliance determination results is generated, and rectification strategies are formulated, such as adjusting sensitive parts of the content and optimizing audio expression. After rectification, the audit is re-approved and the entry into the database is completed. If compliance is determined, entry into the database is allowed directly.

[0080] The target audio materials will undergo compliance screening, specifically including the following steps:

[0081] The temporal signal of the target audio material is mapped to the semantic feature space through nonlinear tensor decomposition to generate a semantic differential manifold;

[0082] The text transcription content corresponding to the target audio material is projected onto the feature space to form a text semantic embedding layer;

[0083] The correspondence deviation between the semantic differential manifold and the text semantic embedding layer is calculated to obtain a multimodal semantic manifold containing both acoustic and textual semantic information;

[0084] A semantic normative field is generated based on a pre-set review rule base, and the semantic normative field is coupled to a multimodal semantic manifold to excite symmetry transformations;

[0085] Calculate the topological load and curvature tensor of the semantically differentiable manifold after symmetry transformation to identify topological defects in the semantically differentiable manifold;

[0086] Compliance screening of target audio materials is performed based on topological defects.

[0087] First, the mapping from time-domain signals to semantic features and the generation of semantic differential manifolds are achieved. In the initial stage of audio compliance screening, the time-domain signal of the target audio material is first acquired, that is, the raw signal data of the audio in the time dimension, such as the signal of the sound wave amplitude changing over time in a digital broadcast audio clip. Subsequently, the time-domain signal is processed using nonlinear tensor decomposition technology. This technology can overcome the limitations of linear analysis, uncover the complex semantic relationships hidden in the signal, and map the time-domain signal from the original time-dimensional space to a feature space that is more in line with semantic understanding. A semantic differential manifold is generated in the feature space. This manifold can be understood as a continuous geometric structure that can intuitively and accurately represent the semantic distribution pattern of the audio signal. For example, in an audio clip containing emergency warning content, its semantic differential manifold will form a specific geometric shape in the feature space that is related to the semantics of warning and emergency.

[0088] The feature space projection and semantic embedding layer construction of the transcribed text content are completed. For the target audio material, text transcription is performed, converting the speech content into text form. For example, the speech content of a digital broadcast stating "Please note that there will be heavy rainfall in the next two hours, please take precautions" is transcribed into corresponding text. The transcribed text content is projected into a feature space identical to the semantic differential manifold. Natural language processing techniques are then used to extract semantics from the text, forming a semantic embedding layer. This embedding layer locates key semantic information in the text within the feature space in the form of vectors, placing it on the same analytical dimension as the semantic representation of the audio signal, laying the foundation for subsequent multimodal semantic fusion.

[0089] Then, multimodal semantic fusion and multimodal semantic manifold construction are performed. The correspondence deviation between the previously generated semantic differential manifold and the text semantic embedding layer is calculated. This deviation essentially reflects the degree of matching between the audio acoustic semantics and the text semantics in the feature space. By analyzing and integrating this deviation, acoustic semantic information and text semantic information can be deeply fused, ultimately resulting in a multimodal semantic manifold containing dual semantic information. Taking an audio clip containing sensitive information as an example, if its acoustic signal contains specific emotional fluctuations, and the transcribed text also contains sensitive keywords, the multimodal semantic manifold will associate these two types of information to form a more comprehensive semantic representation, avoiding judgment bias caused by relying solely on acoustic or text information.

[0090] A semantic normative field is introduced to stimulate a symmetry transformation of the multimodal semantic manifold. The semantic normative field is generated based on a pre-built review rule base, which contains various compliance standards, such as rules prohibiting false warnings and sensitive words in digital broadcasting. The semantic normative field is the concrete representation of these rules in the feature space, essentially setting a compliance benchmark for semantic analysis. When this semantic normative field is coupled to the multimodal semantic manifold, a symmetry transformation is stimulated. That is, the manifold adjusts according to the benchmark of the semantic normative field. If the semantic representations in the multimodal semantic manifold conform to the rules, the transformed manifold maintains good symmetry; if there are illegal semantics, the symmetry of the manifold will be broken.

[0091] Topological defects in semantically differentiable manifolds are identified through topological charge and curvature tensor calculations. After symmetry transformation, two key topological indices—topological charge and curvature tensor—are calculated for the semantically differentiable manifold. Topological charge describes the overall morphological properties of the manifold in the feature space; different topological charges correspond to different semantic distribution patterns. The curvature tensor characterizes the local bending of the manifold, reflecting local anomalies in the semantic representation. Analysis of these two indices identifies topological defects in the semantically differentiable manifold, which typically correspond to non-compliant semantic information in audio. If an audio clip contains false earthquake warning content, its multimodal semantic manifold will show significant local bending in the semantic dimension corresponding to the authenticity of the warning, the curvature tensor will exhibit anomalous values, and the topological charge will also differ significantly from that of compliant warning audio, thus forming a topological defect.

[0092] The system performs compliance screening of target audio materials based on topological defects. Identified topological defects are matched against violation types in the review rule base to determine the specific violation content corresponding to the defects, such as false information, sensitive words, or inappropriate emotional guidance, thereby assessing the compliance of the target audio material. If no topological defects are detected, or the degree of topological defects does not reach the violation threshold, the audio is deemed compliant; if there are clear and serious topological defects, the audio is deemed non-compliant, and the specific non-compliant semantic segment can be located, providing precise evidence for subsequent rectification. Taking digital broadcast audio review as an example, if the system detects a topological defect related to a false flood warning in the multimodal semantic manifold of an audio segment, it will determine that the audio is non-compliant and point out the specific audio segment containing the false warning statement, facilitating modification by staff.

[0093] After analyzing the metadata and waveform features of the target audio material, the propagation scene label is determined, which includes the following steps:

[0094] The metadata information and audio waveform features of the target audio material are processed to obtain a multimodal semantic field;

[0095] Obtain the feature solution set of the multimodal semantic field; wherein, the feature solution set contains multiple intrinsic modalities, and each intrinsic modality corresponds to a scene matching mode;

[0096] Based on a predefined scene classification system, multiple scene matching models corresponding to different propagation scenarios are constructed; the feature solution set is input into the scene matching model to obtain the degree of matching between the audio material and each scene;

[0097] The dissemination scene tags are determined based on the degree of matching between the audio material and each scene.

[0098] Metadata and waveform features of the target audio material are extracted. The metadata includes the audio's source, associated region, event type, and background content encoded by the device. Taking the intelligent broadcast audio in the system as an example, its metadata includes the broadcast region being a level 4 emergency event and the associated terminal physical code being a specific value. This information collectively constitutes the audio's contextual background. The audio waveform features represent the physical characteristics of the audio in the time and frequency domains, such as the high-frequency waveform of warning sounds in intelligent broadcasts and the steady amplitude changes of the broadcast voice. These features directly reflect the acoustic nature of the audio.

[0099] These two types of information are processed and transformed separately. Metadata information is transformed and mapped into a contextual scalar field. This scalar field presents the scene-related attributes in the metadata in a quantified numerical form, such as mapping the type of emergency to a specific numerical value, making abstract contextual information computable and analyzable. Audio waveform features are mapped into an acoustic vector field using techniques such as spectral analysis. The direction of the vector corresponds to the type of acoustic feature, such as whether it is a warning sound or normal speech, and the magnitude of the vector represents the intensity of the feature, such as volume or frequency. The contextual scalar field and the acoustic vector field are coupled and fused to generate a multimodal semantic field that simultaneously contains contextual background and acoustic characteristics.

[0100] The generated multimodal semantic field is deeply analyzed, and feature sets representing the core attributes of the semantic field are extracted using eigenvalue decomposition techniques. Each feature set consists of multiple independent intrinsic modalities, representing a specific scene attribute within the multimodal semantic field and corresponding to a scene matching pattern. In the digital intelligent broadcasting scenario, there may be various intrinsic modalities, such as regionally oriented digital intelligent broadcasting, cross-regional daily broadcasting, and single-terminal early warning broadcasting. The intrinsic modalities of regionally oriented digital intelligent broadcasting emphasize features such as specific area coding, emergency event identification, and high-priority broadcasting, while the intrinsic modalities of cross-regional daily broadcasting focus on features such as multi-regional coverage identification and daily program identification. These intrinsic modalities provide accurate feature units for subsequent scene matching.

[0101] Based on business needs, a predefined scenario classification system is constructed. Digital intelligent broadcasting scenarios typically include core categories such as cross-channel dissemination, targeted scenario applications, regional digital intelligent broadcasting, and daily terminal broadcasting. Each category corresponds to an independent scenario matching model. These models are trained on a large amount of historical audio data and possess the ability to identify specific scenario feature combinations. For example, the regional digital intelligent broadcasting model focuses on identifying the feature combination of emergency event identifiers and specific regional codes, while the daily terminal broadcasting model focuses on the feature combination of daily program identifiers and single-terminal device codes.

[0102] The extracted feature set is input into all scene matching models. The models output the matching degree between the target audio and each scene category through feature similarity calculation methods. The feature set includes features such as specific area codes, emergency event identifiers, and outdoor multi-mode speaker terminal codes. After being input into the model, the matching degree of regional digital broadcasting scenes will be significantly higher than that of other scenes, forming a clear difference in matching degree.

[0103] The matching scores from each scene matching model are sorted, and the scene category with the highest matching score is selected as the propagation scene label for the target audio. If the highest matching score exceeds a preset threshold, the label is directly determined; if multiple scenes have similar matching scores but do not reach the threshold, a secondary verification is performed by combining additional information such as the audio's broadcast time and device type to ensure label accuracy.

[0104] The multimodal semantic field is obtained by processing the metadata information and audio waveform features of the target audio material. The specific steps include:

[0105] The metadata information of the target audio material is transformed and mapped into a contextual scalar field;

[0106] The audio waveform features are mapped into an acoustic vector field through spectral analysis;

[0107] A multimodal semantic field is generated by coupling a contextual scalar field with an acoustic vector field.

[0108] First, the metadata of the target audio material is processed and transformed into a contextual scalar field. The metadata encompasses a great deal of background information about the audio. For example, in this digital broadcast audio, the metadata includes the broadcast area being a specific town and village, the corresponding event type being a Level 4 emergency, and the associated terminal device code being a specific string of numbers. This metadata is then converted into a scalar field that can be quantified numerically—the contextual scalar field. For instance, a Level 4 emergency can correspond to a specific scalar value, and a specific town and village will also correspond to a corresponding scalar value. This presents the originally abstract metadata information in the form of a numerical scalar field, facilitating subsequent calculations and analysis.

[0109] Spectral analysis is performed on the audio waveform features, mapping them to an acoustic vector field. Audio waveform features refer to the physical characteristics of audio propagation. For example, in this digital broadcast audio, the waveform of the warning tone changes rapidly in amplitude and has a high frequency in the time domain, while the waveform of the normal speech broadcast is relatively stable and has a moderate frequency. Fourier transform is used to analyze these waveform features, converting them from time-domain waveforms to representations in the frequency domain or other domains, and then mapping them to an acoustic vector field. The vectors in the vector field can correspond to different types of acoustic features, such as the direction of the warning tone or the direction of normal speech; the magnitude of the vector corresponds to the intensity of the acoustic feature, such as the volume and frequency of the warning tone. This transforms the audio waveform features into an acoustic vector field that is easier to analyze.

[0110] A multimodal semantic field is generated by coupling a contextual scalar field with an acoustic vector field. This results in a multimodal semantic field that incorporates both contextual information represented by audio metadata and acoustic feature information represented by the audio waveform. Taking this digital broadcast audio as an example, the multimodal semantic field simultaneously includes contextual information such as the level four emergency event and the specific town and village, as well as acoustic feature information from the warning prompts and voice broadcasts. This allows for subsequent scene determination or compliance screening of the audio based on this multimodal semantic field, leading to more comprehensive and accurate analysis results.

[0111] Then, extract cross-channel review records of similar audio from the historical cross-channel review database, and calculate the feature correlation coefficient between the historical audio and the target audio material. This includes the following steps:

[0112] Extract cross-channel review records of similar audio from the historical cross-channel review database; among them, the cross-channel review records of similar audio have the same dissemination scenario tags as the target audio material;

[0113] Each cross-channel review record is mapped to a corresponding historical review eigenstate, and all historical review eigenstates constitute a cluster of historical review eigenstates.

[0114] The feature information of the target audio material is mapped to the target review eigenstate; the feature correlation coefficient is obtained by performing operations on the target review eigenstate and the historical review eigenstate cluster.

[0115] The first step is to construct virtual audio objects to achieve a precise mapping from physical audio to digital space. Core features of the segment to be judged in the target audio material are extracted, including feature correlation coefficients, the segment's own voiceprint features, and emotional tendency features. Taking an emergency warning audio message planned for cross-regional distribution as an example, the feature correlation coefficient reflects its correlation with the review features of similar historical cross-channel warning audio messages; the voiceprint features reflect the announcer's acoustic characteristics; and the emotional tendency features correspond to the seriousness, urgency, and other tonal characteristics of the broadcast. These feature data are integrated into a virtual audio object with attributes completely consistent with the physical audio segment, ensuring that the virtual object can accurately reproduce the propagation characteristics and semantic expression of real audio in different channel environments.

[0116] The real-time review rule bases of various channels are deployed to the digital twin space in the form of smart contracts. Smart contracts, with their immutability and automatic execution, can transform review rules scattered across different channels into unified and transparent execution logic. Taking the channels involved in digital broadcasting, such as short video platforms, local government platforms, and community broadcasting terminals, as examples, the rules of short video platforms include audio duration limits and sensitive word filtering; government platforms focus on verifying the authority of information releases; and community terminals focus on volume adaptation and content accessibility requirements. These rules are transformed into code logic in smart contracts, collectively forming the core detection basis of the feature recognition model.

[0117] The constructed virtual audio object is input into the feature recognition model, which then calls the channel rules encapsulated in the smart contract for automatic detection and matching. For emergency warning audio disseminated across multiple channels, the model verifies whether it complies with the rules of each platform: on short video platforms, it checks whether the duration is within the allowed range; on government platforms, it verifies whether the publishing label is authoritative; and on community terminals, it tests whether the volume parameters are compatible with the device. After the detection is completed, the model outputs preliminary violation risk assessment results, clearly indicating which channels the virtual audio object has potential violations on and what the violation type is, providing basic data for subsequent accurate judgment.

[0118] Next, the confidence level of the compliance decision path is calculated to quantify the reliability of the initial screening results. The confidence level reflects the degree of consistency between the preliminary assessment results and the decision-making logic of similar historical audit cases. The system will access the historical cross-channel audit database and compare the detection path of the current virtual audio object with the audit decision paths of similar audio in the past. For example, if a historical warning audio clip was flagged on a short video platform for exceeding the time limit, its decision path was: abnormal time detection, flagging violation, and suggesting editing. If the detection path of the current audio is highly similar to this, and the final judgment accuracy of historical cases is high, then the confidence level of the current audio's compliance decision path will increase accordingly; otherwise, it will decrease.

[0119] The compliance decision-making path confidence level is calculated by integrating the dynamic weight of channel review authority and real-time situation correlation characteristics. The dynamic weight of channel review authority is set according to the strictness of channel supervision and industry influence; for example, the weight of government platforms is usually higher than that of ordinary short video platforms. The real-time situation correlation characteristics focus on the current social commentary trends related to the audio topic. If the region involved in the emergency warning audio is in a sensitive period of the relevant event, the influence of the situation correlation characteristics will be correspondingly enhanced. Taking the rainstorm warning audio disseminated across channels as an example, if its compliance decision-making path confidence level is high on the government platform and the platform has a high weight, and the real-time situation urgently requires warning information, the quantitative indicator of compliance status obtained after integrating the three factors will be at a high level.

[0120] The compliance status quantification indicators of each channel are compared with preset judgment thresholds. If the indicator is greater than or equal to the threshold, the channel is judged to be compliant; otherwise, it is non-compliant. For compliant channels, compatibility is directly confirmed; for non-compliant channels, the reasons for the violation are clarified, such as exceeding the time limit on short video platforms or incompatible volume on community terminals. Subsequently, an audit report is generated, which includes an audio feature summary and a risk level label. The feature summary covers key information such as the audio's dissemination scenario and core characteristics, while the risk level label is divided into different levels such as high risk, medium risk, and low risk based on the compliance status quantification indicators. Taking emergency warning audio disseminated across channels as an example, the report will clearly list that it can be directly distributed on government platforms with low compliance risk levels, requires time editing on short video platforms with medium compliance risk levels, and is suitable for distribution on community terminals with low compliance risk levels, providing clear guidance for subsequent adaptation strategies.

[0121] After determining the multi-channel compliance of the audio segment to be judged based on the feature correlation coefficient and the real-time review rule base of each channel, a review report is generated, which includes the following steps:

[0122] After digital twin modeling the feature correlation coefficient with the voiceprint features and emotional tendency features of the audio segment to be judged, a corresponding virtual audio object is generated.

[0123] Deploy the real-time review rule base of each channel to the digital twin space in the form of smart contracts to build a feature recognition model;

[0124] The virtual audio object is input into the feature recognition model to obtain preliminary risk assessment results for violations;

[0125] The confidence level of the compliance decision-making path for the audio segment to be judged in various channels is obtained based on the preliminary risk assessment results.

[0126] The compliance status quantification indicators for each channel are obtained based on the confidence level of the compliance decision-making path, the dynamic weight of the authority of the channel review, and the real-time correlation characteristics of the audio segments to be judged.

[0127] The compliance status quantitative indicators of each channel are compared with the preset judgment threshold. If the compliance status quantitative indicators are greater than or equal to the threshold, the channel is judged to be compliant; otherwise, it is judged to be non-compliant.

[0128] Generate an audit report that includes an audio feature summary and a risk level identifier.

[0129] The first step is to construct virtual audio objects to achieve a precise mapping from physical audio to digital space. Core features of the segment to be judged in the target audio material are extracted, including the previously calculated feature correlation coefficient, the segment's own voiceprint features, and emotional tendency features. Taking an emergency warning audio message planned for cross-regional dissemination as an example, the feature correlation coefficient reflects its correlation with the review features of similar historical cross-channel warning audio messages; the voiceprint features reflect the announcer's acoustic characteristics; and the emotional tendency features correspond to the seriousness, urgency, and other tonal characteristics of the broadcast. These feature data are integrated into a virtual audio object with attributes completely consistent with the physical audio segment, ensuring that the virtual object can accurately reproduce the dissemination characteristics and semantic expression of real audio in different channel environments such as short video platforms, government broadcasting terminals, and community loudspeakers.

[0130] Secondly, a standardized multi-channel compliance detection environment is built by deploying feature recognition models. The real-time review rule bases of each channel are deployed to the digital twin space in the form of smart contracts. Smart contracts, with their immutability and automatic execution, can transform review rules scattered across different channels into unified and transparent execution logic. Taking the channels involved in digital broadcasting as an example, short video platforms may have rules including audio duration limits and sensitive word filtering; government platforms focus on verifying the authority of information releases; and community terminals focus on volume adaptation and content accessibility requirements. These rules will be transformed into code logic in smart contracts, collectively forming the core detection basis of the feature recognition model.

[0131] The constructed virtual audio object is input into the feature recognition model, which then calls the channel rules encapsulated in the smart contract for automatic detection and matching. For emergency warning audio transmitted across multiple channels, the model verifies whether it complies with the rules of each platform: on short video platforms, it checks whether the duration is within 3 minutes; on community terminals, it tests whether the volume parameters meet the 60-decibel limit. After the detection is completed, the model outputs preliminary violation risk assessment results, clearly indicating which channels the virtual audio object has potential violations on and what the violation type is, providing basic data for subsequent accurate judgment.

[0132] The confidence level of the compliance decision path reflects the degree of consistency between the preliminary assessment results and the decision-making logic of similar historical audit cases. The system will call the historical cross-channel audit database to compare the detection path of the current virtual audio object with the audit decision paths of similar audio in the past. For example, a rainstorm warning audio from Yunyang Street, Danyang City, Zhenjiang in 2019 was marked as non-compliant on the government affairs platform because it did not indicate the issuing unit. Its decision path was to identify the missing detection, mark it as non-compliant, and suggest supplementing the unit information. If the detection path of the current emergency warning audio is highly similar to that of the previous one, and the final judgment accuracy rate of historical cases is above 95%, then the confidence level of the compliance decision path of the current audio will increase to 92%. Conversely, if the detection path is significantly different, the confidence level will drop to below 60%.

[0133] The compliance decision-making path confidence level is calculated by integrating the dynamic weight of channel review authority and real-time situation correlation characteristics. The dynamic weight of channel review authority is set according to the strictness of channel supervision and industry influence; for example, the weight of government platforms is usually higher than that of ordinary short video platforms. The real-time situation correlation characteristics focus on the current social commentary trends related to the audio topic. If the area involved in the emergency warning audio is in the flood season, the attention to the warning information is high, and the weight of this characteristic will be higher than that in the non-flood season. Taking the rainstorm warning audio disseminated across channels as an example, if its compliance decision-making path confidence level on the government platform is 92%, the platform weight is 0.8, and the situation correlation characteristic weight is 0.3, the quantitative compliance status index obtained after integrating the three is 92%×0.8+85%×0.3=90.1%, which is at a relatively high level.

[0134] For channels identified as having a risk of violation in the initial assessment, feature vectors are constructed by extracting the corresponding violation segments from the virtual audio objects. Taking a short video platform as an example, if an emergency warning audio is marked as violating the rule due to exceeding the 4-minute limit, a 1-minute segment exceeding the 3-minute limit is extracted, and the timestamp features, semantic features, and acoustic features of this segment are extracted to form a feature vector with a dimension of 3. Simultaneously, feature vectors of similar violation cases are retrieved from the historical review database, such as the feature vector of an overdue warning audio from a test city in 2019, to prepare for subsequent similarity comparisons.

[0135] The cosine similarity algorithm is used to calculate the feature vector of the current violation segment and the feature vector of historical violation segments. If the feature vector of the current emergency warning audio timeout segment has a similarity of 88% with the feature vector of historical timeout cases, and the historical cases were ultimately determined to be minor violations, then the violation level of the current audio on the short video platform will also be initially determined to be a minor violation; if the similarity is less than 50%, the feature recognition model will be called again for secondary detection to eliminate the possibility of misjudgment.

[0136] The compliance status quantification indicators of each channel are compared with preset judgment thresholds. If the indicator is greater than or equal to the threshold, the channel is judged to be compliant; otherwise, it is non-compliant. For compliant channels, compatibility is directly confirmed; for non-compliant channels, the reasons for the violation and rectification suggestions are clarified. Subsequently, an audit report is generated, which includes an audio feature summary and a risk level label. The feature summary covers the audio's dissemination scenario and key information of core features. The risk level label is divided into low risk, medium risk, and high risk according to the compliance status quantification indicators. Short video platforms are marked as medium risk, while government platforms and community terminals are marked as low risk, providing clear guidance for the subsequent development of adaptation strategies.

[0137] Then, audio materials from the same scene are extracted from the historical review data of the targeted scene, and the historical key feature distribution trend of the illegal content is statistically analyzed based on the audio materials from the same scene. The specific steps include:

[0138] Spatiotemporal annotations are performed on audio materials from the same scene extracted from historical review data of targeted scenes, and spatiotemporal data tags containing timestamps and geographical coordinates are constructed by combining the dissemination area and playback time period;

[0139] A generative adversarial network and variational autoencoder fusion model is used to enhance the features of audio materials in the same scene. Virtual violation audio features are generated through adversarial training, and the original features together constitute the enhanced feature dataset.

[0140] The enhanced feature dataset and spatiotemporal data labels are input into the feature association mining model to obtain the feature association matrix;

[0141] A reinforcement learning framework is adopted, with the feature association matrix as the state space, the feature selection operation as the action space, and the accuracy of violation feature recognition as the reward function, to train the optimal feature selection strategy.

[0142] Key violation features are selected from the enhanced feature dataset based on the optimal feature selection strategy, and time series decomposition is performed on the key violation features.

[0143] We model the key violation features after time series decomposition and capture the dependencies of key violation features in different time periods through a multi-head attention mechanism.

[0144] Calculate the dynamic fluctuation parameters and structural complexity measures of the time series of key violation features;

[0145] The historical key feature distribution trend of the illegal content is obtained based on dynamic fluctuation parameters, structural complexity metrics, and feature dependencies.

[0146] Firstly, spatiotemporal annotation assigns geographic and temporal dimension feature anchors to audio materials. Taking the digital broadcasting scenario as an example, audio materials from the same scenario are extracted from historical review data of targeted scenarios. These audio materials are then spatiotemporally annotated, and spatiotemporal data tags containing timestamps and geographic coordinates are constructed by combining the dissemination region and playback time. In this way, each piece of digital broadcasting audio has a clear spatiotemporal attribute identifier, which can accurately associate regional and time-time characteristics during subsequent analysis.

[0147] Feature enhancement aims to broaden the sample diversity of violation features. A fusion model of generative adversarial networks (GANs) and variational autoencoders (VAEs) is used to process audio materials from the same scene. The GAN generates virtual violation audio features to simulate more similar but varied violation scenarios, such as generating different wording but fundamentally false warning audio features for false warnings in digital broadcasting. The VAE, on the other hand, encodes and decodes the original audio features to optimize the feature representation. The virtual violation audio features generated through adversarial training, together with the original features, constitute the enhanced feature dataset, improving the coverage of violation features.

[0148] Feature association mining establishes a correlation matrix between spatiotemporal and audio features. The enhanced feature dataset and spatiotemporal data labels are input into the feature association mining model. The model analyzes the correlation of audio features under different spatiotemporal conditions. For example, it identifies which geographical areas are more closely associated with features in intelligent broadcasting where warning levels are unclear, thus obtaining the feature association matrix. The feature association matrix clearly presents the correlation strength between spatiotemporal attributes and audio features, providing a basis for subsequent screening of key violation features.

[0149] The optimal feature selection strategy is obtained through reinforcement learning training. A reinforcement learning framework is employed, using the feature association matrix as the state space (representing currently available association information), the feature selection operation as the action space (determining which features to focus on), and the violation feature identification accuracy as the reward function; positive rewards are given if the selected features more accurately identify violations. Through continuous training, the system learns which features are more effective at identifying violations in digital broadcasting scenarios, ultimately leading to the optimal feature selection strategy.

[0150] Based on the optimal feature selection strategy, key violation features are selected from the enhanced feature dataset. For example, in the digital broadcasting scenario, key violation features include discrepancies between warning information and actual conditions, and broadcasting tone that does not conform to emergency regulations. Then, these key violation features are decomposed into time series, breaking them down into feature sequences for different time periods according to time order. For example, the feature of discrepancies between warning information and actual conditions is decomposed into hourly units to analyze the performance of the feature in each hour.

[0151] Key violation features after time series decomposition are modeled using a multi-head attention mechanism. This mechanism can simultaneously focus on key violation features across different time periods, capturing the dependencies between them. For example, in digital broadcasting, the features of warning-level errors in the previous hour can influence the features of public misunderstanding feedback in the following hour; this mechanism can clearly uncover such inter-time period correlations.

[0152] Next, dynamic fluctuation parameters and structural complexity metrics are calculated. For the time series of key violation features, dynamic fluctuation parameters are calculated to reflect the magnitude and frequency of feature changes over time, such as the number and magnitude of fluctuations in the false alarm feature within a day. At the same time, structural complexity metrics are calculated to measure the complexity of the feature sequence, such as whether the feature exhibits simple periodic changes or irregular complex changes.

[0153] Based on dynamic fluctuation parameters, structural complexity metrics, and feature dependencies captured through a multi-head attention mechanism, the historical key feature distribution trend of violating content is comprehensively derived. This provides a historical reference for subsequent audio review in targeted scenarios.

[0154] After assessing compliance based on the degree of overlap between the current trend of violation characteristics and the distribution trend of historical key characteristics, an audit report is generated, which includes the following steps:

[0155] After mapping the current violation feature association trend and the historical key feature distribution trend to the topological space respectively, the feature topological structure is constructed, and the topological algebraic representation value of the feature topological structure is obtained.

[0156] Semantic enhancement processing is performed on the algebraic representation values ​​of topological structures, transforming topological features into feature vectors containing semantic information;

[0157] A dynamic weighted matching model based on attention mechanism is established. The semantic similarity, time span weight, and scene relevance weight of the feature vector are used as input. The weights of each dimension are adaptively adjusted through a multi-head attention mechanism, and the weighted matching degree between the current illegal feature association trend and the historical key feature distribution trend is calculated.

[0158] A fuzzy comprehensive evaluation model is introduced to perform a fuzzy mapping between the weighted matching degree and the preset multi-level compliance threshold range, and the membership degree function is combined to determine the degree of membership of the current audio material in different compliance levels.

[0159] A compliance probability distribution vector is constructed based on the degree of membership, and the probability distribution vector is processed to generate a compliance probability distribution map;

[0160] The final compliance status of the audio material is determined based on the compliance probability distribution map and the preset decision-making strategy;

[0161] Identify key violation segments that correlate trends with the violation characteristics of non-compliant audio materials;

[0162] Generate an audit report that includes basic information about the audio material and the compliance assessment results.

[0163] First, topological space mapping and algebraic representation value acquisition are performed. Taking the digital broadcasting scenario as an example, the correlation trend of the violation characteristics of the audio to be reviewed and the historical key feature distribution trend of the violation content in this scenario, mined from historical data, are mapped into the topological space. In the topological space, these trends are combined into feature topological structures by topological elements such as points, lines, and surfaces. The algebraic representation values ​​of the topological structure that can represent the core attributes of the topological structure are extracted, and a set of values ​​is used to summarize the key characteristics of the topological structure.

[0164] The algebraic representation of the topological structure is semantically enhanced and transformed into semantically meaningful feature vectors. For example, a change in the value of a certain dimension in the digital broadcasting topology corresponds to a shift in warning information from accurate to vague in actual semantics. By combining the professional semantic knowledge of digital broadcasting, the originally mathematically-oriented topological features are transformed into feature vectors containing specific semantic information.

[0165] Then, a dynamic weighted matching model based on an attention mechanism is established to calculate the weighted matching degree. The semantic similarity, time span weight, and scene relevance weight of the feature vectors are used as inputs. With the help of a multi-head attention mechanism, the model can adaptively adjust the weights of these three dimensions. For example, when the semantic similarity is high, the weight of semantic similarity is increased, thereby calculating the weighted matching degree of the current trend of violation feature association and the historical trend of key feature distribution, and thus measuring the degree of matching between the two.

[0166] The calculated weighted matching degree is fuzzily mapped to a preset multi-level compliance threshold range. Then, a membership function is used to determine the degree of membership of the current audio material in different compliance levels. For example, if the weighted matching degree is in the slightly non-compliant range, the current audio material will have a higher degree of membership in the slightly non-compliant level and a lower degree of membership in the compliant level.

[0167] A compliance probability distribution vector is constructed based on the degree of membership to generate a compliance probability distribution map. The vector is created by constructing a compliance probability distribution vector according to the degree of membership of each compliance level, with each element corresponding to the probability of a compliance level. This probability distribution vector is then visualized to generate a compliance probability distribution map, which visually displays the probability of the current audio material belonging to different compliance levels. For example, the map shows that the probability of minor non-compliance is 60%, compliance is 30%, and serious non-compliance is 10%.

[0168] The final compliance status is determined based on the compliance probability distribution map and a preset decision-making strategy. For example, the preset decision-making strategy might select the level with the highest probability as the final status; if minor non-compliance has the highest probability, it would be classified as minor non-compliance.

[0169] Identify key violation segments in non-compliant audio materials. If an audio material is determined to be non-compliant, analyze the correlation trend of current violation characteristics to identify the key violation segments leading to the non-compliance. For example, in digital broadcast audio, if a segment in a broadcast contains an incorrect description of the warning time, that segment is the key violation segment.

[0170] Generate an audit report that includes basic information about the audio material and the compliance assessment results. The audit report integrates the source, duration, associated region, and compliance assessment results of the audio material.

[0171] An audio material review and storage system includes:

[0172] Acquisition Module: Acquires target audio materials to be reviewed and performs compliance screening on the target audio materials;

[0173] Parsing module: If the compliance of the target audio material cannot be determined, the metadata information and audio waveform characteristics of the target audio material are parsed to determine the propagation scenario label;

[0174] First review module: If the dissemination scenario tag belongs to the cross-channel dissemination type, then extract the cross-channel review records of the same type of audio from the historical cross-channel review database, calculate the feature correlation coefficient between the historical audio and the target audio material; separate the audio segment to be judged from the target audio material, and generate a review report after judging the multi-channel compliance of the audio segment to be judged based on the feature correlation coefficient and the real-time review rule library of each channel, and formulate an adaptation strategy to complete the entry judgment.

[0175] The second review module: If the dissemination scenario tag belongs to the targeted scenario application type, then extract audio materials of the same scenario from the historical review data of the targeted scenario, and statistically analyze the historical key feature distribution trend of the illegal content based on the audio materials of the same scenario; determine the current illegal feature correlation trend of the target audio material, and determine compliance based on the matching overlap between the current illegal feature correlation trend and the historical key feature distribution trend, and generate a review report, and formulate rectification strategies to complete the entry into the database.

[0176] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements a method for reviewing and storing audio materials.

[0177] like Figure 3As shown, the electronic device may include a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute an audio material review and storage method.

[0178] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0179] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute an audio material review and storage method.

[0180] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform an audio material review and storage method.

[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0182] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for reviewing and storing audio materials, characterized in that, The method includes the following steps: Obtain the target audio materials to be reviewed and conduct compliance screening on the target audio materials; If the compliance of the target audio material cannot be determined, the dissemination scenario label is determined after parsing the metadata information and audio waveform characteristics of the target audio material. If the dissemination scenario tag belongs to the cross-channel dissemination type, then the cross-channel review records of the same type of audio are extracted from the historical cross-channel review database, and the feature correlation coefficient between the historical audio and the target audio material is calculated; the audio segment to be judged is separated from the target audio material, and the multi-channel compliance of the audio segment to be judged is determined according to the feature correlation coefficient and the real-time review rule base of each channel before a review report is generated, which specifically includes the following steps: After digital twin modeling the feature correlation coefficient with the voiceprint features and emotional tendency features of the audio segment to be judged, a corresponding virtual audio object is generated. Deploy the real-time review rule base of each channel to the digital twin space in the form of smart contracts to build a feature recognition model; The virtual audio object is input into the feature recognition model to obtain preliminary risk assessment results for violations; The confidence level of the compliance decision-making path for the audio segment to be judged in various channels is obtained based on the preliminary risk assessment results. The compliance status quantification indicators for each channel are obtained based on the confidence level of the compliance decision path, the dynamic weight of the authority of the channel review, and the real-time situation correlation characteristics of the audio segment to be judged. Among them, the real-time situation correlation characteristics focus on the current social commentary trend related to the audio topic. If the region involved in the emergency warning audio is in a sensitive period of related events, the influence of the situation correlation characteristics will be correspondingly enhanced. The compliance status quantitative indicators of each channel are compared with the preset judgment threshold. If the compliance status quantitative indicators are greater than or equal to the threshold, the channel is judged to be compliant; otherwise, it is judged to be non-compliant. Generate an audit report that includes an audio feature summary and risk level identifier; And formulate an adaptation strategy to complete the data entry determination; If the dissemination scenario tag belongs to the targeted scenario application type, then audio materials of the same scenario are extracted from the historical review data of the targeted scenario, and the historical key feature distribution trend of the illegal content is statistically analyzed based on the audio materials of the same scenario; the current illegal feature correlation trend of the target audio material is judged, and the compliance is judged based on the matching overlap between the current illegal feature correlation trend and the historical key feature distribution trend. After that, an review report is generated, and a rectification strategy is formulated to complete the entry into the database.

2. The method for reviewing and storing audio materials according to claim 1, characterized in that, The target audio materials will undergo compliance screening, specifically including the following steps: After mapping the temporal signal of the target audio material to the semantic feature space through nonlinear tensor decomposition, a semantic differential manifold is generated. The semantic differential manifold is a continuous geometric structure that can intuitively and accurately represent the semantic distribution pattern of the audio signal. The text transcription content corresponding to the target audio material is projected onto the feature space to form a text semantic embedding layer; The correspondence deviation between the semantic differential manifold and the text semantic embedding layer is calculated to obtain a multimodal semantic manifold containing both acoustic and textual semantic information. The multimodal semantic manifold is used to associate the emotional fluctuations in the acoustic signals in the audio with the sensitive keywords in the text to form a comprehensive semantic representation. A semantic normative field is generated based on a pre-set review rule base, and the semantic normative field is coupled to a multimodal semantic manifold to excite symmetry transformations; Calculate the topological load and curvature tensor of the semantically differentiable manifold after symmetry transformation to identify topological defects in the semantically differentiable manifold; Compliance screening of target audio materials is performed based on topological defects.

3. The method for reviewing and storing audio materials according to claim 1, characterized in that, After analyzing the metadata and waveform features of the target audio material, the propagation scene label is determined, which includes the following steps: The metadata information and audio waveform features of the target audio material are processed to obtain a multimodal semantic field; Obtain the feature solution set of the multimodal semantic field; wherein the feature solution set contains multiple intrinsic modalities, and each intrinsic modality corresponds to a scene matching mode; Based on a predefined scene classification system, multiple scene matching models corresponding to different propagation scenarios are constructed; the feature solution set is input into the scene matching model to obtain the degree of matching between the audio material and each scene; The dissemination scene tags are determined based on the degree of matching between the audio material and each scene.

4. The method for reviewing and storing audio materials according to claim 3, characterized in that, The multimodal semantic field is obtained by processing the metadata information and audio waveform features of the target audio material. The specific steps include: The metadata information of the target audio material is transformed and mapped into a contextual scalar field, which presents the scene-related attributes in the metadata in a quantified numerical form. The audio waveform features are mapped into an acoustic vector field through spectral analysis; A multimodal semantic field is generated by coupling a contextual scalar field with an acoustic vector field.

5. The method for reviewing and storing audio materials according to claim 1, characterized in that, Then, extract cross-channel review records of similar audio from the historical cross-channel review database, and calculate the feature correlation coefficient between the historical audio and the target audio material. This includes the following steps: Extract cross-channel review records of similar audio from the historical cross-channel review database; among them, the cross-channel review records of similar audio have the same dissemination scenario tags as the target audio material; Each cross-channel review record is mapped to a corresponding historical review eigenstate, and all historical review eigenstates constitute a cluster of historical review eigenstates. The feature information of the target audio material is mapped to the target review eigenstate; the feature correlation coefficient is obtained by operating on the target review eigenstate and the historical review eigenstate cluster.

6. The method for reviewing and storing audio materials according to claim 5, characterized in that, Then, audio materials from the same scene are extracted from the historical review data of the targeted scene, and the historical key feature distribution trend of the illegal content is statistically analyzed based on the audio materials from the same scene. The specific steps include: Spatiotemporal annotations are performed on audio materials from the same scene extracted from historical review data of targeted scenes, and spatiotemporal data tags containing timestamps and geographical coordinates are constructed by combining the dissemination area and playback time period; A generative adversarial network and variational autoencoder fusion model is used to enhance the features of audio materials in the same scene. Virtual violation audio features are generated through adversarial training, and the original features together constitute the enhanced feature dataset. The enhanced feature dataset and spatiotemporal data labels are input into the feature association mining model to obtain the feature association matrix; A reinforcement learning framework is adopted, with the feature association matrix as the state space, the feature selection operation as the action space, and the accuracy of violation feature recognition as the reward function, to train the optimal feature selection strategy. Key violation features are selected from the enhanced feature dataset based on the optimal feature selection strategy, and time series decomposition is performed on the key violation features. We model the key violation features after time series decomposition and capture the dependencies of key violation features in different time periods through a multi-head attention mechanism. Calculate the dynamic fluctuation parameters and structural complexity measures of the time series of key violation features; The historical key feature distribution trend of the illegal content is obtained based on dynamic fluctuation parameters, structural complexity metrics, and feature dependencies.

7. The method for reviewing and storing audio materials according to claim 6, characterized in that, After assessing compliance based on the degree of overlap between the current trend of violation characteristics and the distribution trend of historical key characteristics, an audit report is generated, which includes the following steps: After mapping the current violation feature association trend and the historical key feature distribution trend to the topological space respectively, the feature topological structure is constructed, and the topological algebraic representation value of the feature topological structure is obtained. Semantic enhancement processing is performed on the algebraic representation values ​​of topological structures, transforming topological features into feature vectors containing semantic information; A dynamic weighted matching model based on attention mechanism is established. The semantic similarity, time span weight, and scene relevance weight of the feature vector are used as input. The weights of each dimension are adaptively adjusted through a multi-head attention mechanism, and the weighted matching degree between the current illegal feature association trend and the historical key feature distribution trend is calculated. A fuzzy comprehensive evaluation model is introduced to perform a fuzzy mapping between the weighted matching degree and the preset multi-level compliance threshold range, and the membership function is combined to determine the degree of membership of the current audio material in different compliance levels. A compliance probability distribution vector is constructed based on the degree of membership, and the probability distribution vector is processed to generate a compliance probability distribution map; The final compliance status of the audio material is determined based on the compliance probability distribution map and the preset decision-making strategy; Identify key violation segments that correlate with the violation characteristics of non-compliant audio materials. Generate an audit report that includes basic information about the audio material and the compliance assessment results.

8. An audio material review and warehousing system, characterized in that, include: The acquisition module acquires the target audio materials to be reviewed and performs compliance screening on the target audio materials; If the compliance of the target audio material cannot be determined by the parsing module, the dissemination scenario label is determined after parsing the metadata information and audio waveform characteristics of the target audio material. The first review module, if the dissemination scenario tag belongs to the cross-channel dissemination type, extracts cross-channel review records of the same type of audio from the historical cross-channel review database, calculates the feature correlation coefficient between the historical audio and the target audio material; separates the audio segment to be judged from the target audio material, and generates a review report after judging the multi-channel compliance of the audio segment to be judged based on the feature correlation coefficient and the real-time review rule base of each channel. Includes the following steps: After digital twin modeling the feature correlation coefficient with the voiceprint features and emotional tendency features of the audio segment to be judged, a corresponding virtual audio object is generated. Deploy the real-time review rule base of each channel to the digital twin space in the form of smart contracts to build a feature recognition model; The virtual audio object is input into the feature recognition model to obtain preliminary risk assessment results for violations; The confidence level of the compliance decision-making path for the audio segment to be judged in various channels is obtained based on the preliminary risk assessment results. The compliance status quantification indicators for each channel are obtained based on the confidence level of the compliance decision path, the dynamic weight of the authority of the channel review, and the real-time situation correlation characteristics of the audio segment to be judged. Among them, the real-time situation correlation characteristics focus on the current social commentary trend related to the audio topic. If the region involved in the emergency warning audio is in a sensitive period of related events, the influence of the situation correlation characteristics will be correspondingly enhanced. The compliance status quantitative indicators of each channel are compared with the preset judgment threshold. If the compliance status quantitative indicators are greater than or equal to the threshold, the channel is judged to be compliant; otherwise, it is judged to be non-compliant. Generate an audit report that includes an audio feature summary and risk level identifier; And formulate an adaptation strategy to complete the data entry determination; The second review module, if the dissemination scenario tag belongs to the targeted scenario application type, extracts audio materials of the same scenario from the historical review data of the targeted scenario, and statistically analyzes the historical key feature distribution trend of the illegal content based on the audio materials of the same scenario; judges the current illegal feature correlation trend of the target audio material, judges compliance based on the matching overlap between the current illegal feature correlation trend and the historical key feature distribution trend, generates a review report, and formulates rectification strategies to complete the entry into the database.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the audio material review and storage method as described in any one of claims 1 to 7.