ASR-based early media recognition method and system

By performing real-time noise reduction, restoration, and parameter evaluation on audio data, and combining it with the ASR model to identify keywords and sentiment, the problem of ASR technology's accuracy in complex environments has been solved, enabling real-time and personalized media content delivery.

WO2026157041A1PCT designated stage Publication Date: 2026-07-30BEIJING JIZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING JIZHI TECH CO LTD
Filing Date
2025-04-17
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing ASR technology is not accurate enough under the influence of environmental noise, accent differences and complex language structures, making it difficult to achieve real-time and accurate keyword and sentiment identification.

Method used

By collecting audio data in real time, noise reduction, popping removal, and volume normalization are performed. Combined with the evaluation of parameters such as signal-to-noise ratio, total harmonic distortion, and short-time energy, the time window is dynamically adjusted. The ASR model is used for audio restoration and keyword/phrase recognition. Combined with sentiment index, the system is used for identification and content push.

Benefits of technology

It enables real-time and efficient conversion of audio data into text data, accurately identifies keywords, phrases, and sentiment, supports personalized media content delivery, and is applicable to multiple fields such as news broadcasting and social media videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025089528_30072026_PF_FP_ABST
    Figure CN2025089528_30072026_PF_FP_ABST
Patent Text Reader

Abstract

An ASR-based early media recognition method and system. The ASR-based early media recognition method comprises: collecting in real time audio data corresponding to a target medium, and comparing the audio data for restoration to acquire restored audio data (S1); inputting the restored audio data into an ASR model to acquire text data corresponding to the audio data (S2); recognizing keywords and phrases from the text data to acquire target keywords and phrases (S3); comparing the target keywords and phrases with a preset keyword, phrase and emotional tendency table for recognition to acquire an emotional tendency result, the emotional tendency result being an early media recognition result (S4); and performing content pushing on media content on the basis of the emotional tendency result (S5). The system comprises modules corresponding to the steps of the method.
Need to check novelty before this filing date? Find Prior Art

Description

An Early Media Recognition Method and System Based on ASR Technical Field

[0001] This invention proposes an early media recognition method and system based on ASR, belonging to the field of audio recognition technology. Background Technology

[0002] In today's digital and information-explosive era, Automatic Speech Recognition (ASR) technology has become an important component of human-computer interaction. The core task of ASR technology is to convert human speech into text, a process involving several key components such as acoustic models, language models, and decoders. The acoustic model is responsible for converting speech signals into corresponding acoustic feature sequences, i.e., identifying phonemes or syllables in the speech; the language model, trained on large amounts of text data, is used to evaluate whether the generated text sequence conforms to the grammatical rules and idiomatic usage of the language; the decoder is responsible for finding the optimal speech-to-text conversion result based on these two models. ASR technology greatly simplifies the information input process. Compared to traditional manual input methods (such as keyboard input), voice input is more convenient and natural, saving significant time and effort. Especially on some mobile devices, voice input allows users to easily interact with the device even when their hands are occupied. Furthermore, ASR technology enables real-time interaction; users do not need to wait and can receive a system response the instant they speak, greatly improving user experience and work efficiency. However, ASR technology also faces some challenges in practical applications, such as environmental noise, accent differences, and complex language structures, all of which may affect the accuracy of recognition. Summary of the Invention

[0003] This invention provides an early media identification method and system based on ASR to solve the technical problems in the prior art. The technical solution adopted is as follows:

[0004] An ASR-based early media identification method, the ASR-based early media identification method comprising:

[0005] Real-time collection of audio data corresponding to the target media, comparison of the audio data for repair, and acquisition of repaired audio data;

[0006] The repaired audio data is input into the ASR model to obtain the text data corresponding to the audio data;

[0007] Keyword and phrase identification is performed on the text data to obtain target keywords and phrases;

[0008] The target keywords and phrases are compared and identified with a preset keyword and phrase sentiment index table to obtain sentiment index results, which are the early media identification results.

[0009] Media content is pushed based on the stated sentiment results.

[0010] Furthermore, the audio data corresponding to the target media is collected in real time, and the audio data is compared and repaired to obtain the repaired audio data, including:

[0011] Collect audio data corresponding to the target media in real time;

[0012] The audio data is enhanced to obtain the enhanced audio data, wherein the enhancement process includes noise reduction and popping removal and volume normalization.

[0013] The enhanced audio data is evaluated for audio quality to obtain audio quality results.

[0014] Audio data that does not meet the audio quality requirements is repaired, and the repaired audio data is obtained.

[0015] Furthermore, the enhanced audio data is subjected to audio quality evaluation to obtain audio quality results, including:

[0016] Extract audio parameters from the enhanced audio data, wherein the audio parameters include signal-to-noise ratio, total harmonic distortion, and short-time energy;

[0017] Audio quality assessment parameters are obtained using the signal-to-noise ratio, total harmonic distortion, and short-time energy contained in the audio parameters.

[0018] The audio quality evaluation parameters are obtained using the following formula:

[0019] Where P represents the audio quality assessment parameter; n represents the number of time windows corresponding to the short-time energy contained in the audio data; S i T represents the signal-to-noise ratio corresponding to the i-th time window; i S represents the total harmonic distortion corresponding to the i-th time window; max and S min W represents the maximum and minimum signal-to-noise ratios for n time windows corresponding to the audio data. max and W min W represents the maximum and minimum short-time energy values ​​for n time windows corresponding to the audio data. i W represents the short-time energy corresponding to the i-th time window; bi This represents the short-time energy standard deviation corresponding to the i-th time window;

[0020] The audio quality assessment parameters are compared with preset audio quality assessment parameter thresholds;

[0021] When the audio quality evaluation parameter is lower than the preset audio quality evaluation parameter threshold, the audio data is determined to be audio data that does not meet the audio quality requirements.

[0022] Furthermore, the time window corresponding to the short-time energy contained in the audio parameters of the audio data is set, including:

[0023] Extract the sampling frequency, bit depth, and number of channels corresponding to the audio data collection device;

[0024] The first window adjustment coefficient is obtained using the sampling frequency, bit depth, and number of channels corresponding to the audio data collection device; wherein, the first window adjustment coefficient is obtained by the following formula:

[0025] Among them, K 01 f represents the adjustment coefficient of the first window; f represents the sampling frequency corresponding to the audio data collection device; f c This indicates the preset sampling frequency reference value; N represents the number of channels; B represents the bit depth; int() indicates rounding up the value within the parentheses;

[0026] The first window adjustment coefficient is compared with a preset window adjustment coefficient threshold.

[0027] When the first window adjustment coefficient is lower than the preset window adjustment coefficient threshold, the preset initial time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy.

[0028] When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy.

[0029] Furthermore, when the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy, including:

[0030] When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length and initial window overlap ratio corresponding to the time window of the ASR model are extracted; wherein, the ASR model adopts a deep neural network model structure;

[0031] Extract the time length values ​​and corresponding window overlap ratio values ​​that occur during the dynamic adjustment of the time window in the ASR model.

[0032] The second window adjustment coefficient is obtained by combining the initial time length and initial window overlap ratio of the time window corresponding to the time window in the ASR model with the time length and corresponding window overlap ratio during the dynamic adjustment of the time window.

[0033] The second window adjustment coefficient is obtained using the following formula:

[0034] Among them, K 02 Indicates the second window adjustment coefficient; m represents the number of dynamic adjustments to the time window; T i P represents the time length corresponding to the i-th dynamically adjusted time window; i P represents the window overlap ratio after the i-th dynamic adjustment; ti T represents the rate of change in the length of the time window after the i-th dynamic adjustment compared to the time window before the dynamic adjustment; c and P c P represents the initial time length and the initial window overlap ratio. z P represents the median value of the window overlap ratio after m dynamic adjustments; b T represents the standard deviation of the window overlap ratio after m dynamic adjustments; b P represents the standard deviation of the time window length after m dynamic adjustments; max This represents the maximum value of the window overlap ratio after m dynamic adjustments.

[0035] The initial time length is adjusted using the first window adjustment coefficient and the second window adjustment coefficient to obtain the adjusted time length;

[0036] The adjusted time length is obtained using the following formula:

[0037] Among them, T x T0 represents the adjusted time length; T0 represents the original time length; K 01 K represents the first window adjustment coefficient; 02 represents the second window adjustment coefficient; t represents the adjustment factor, and the value range of the adjustment factor is 1.12-1.43;

[0038] The adjusted time length is used as the corresponding length of the time window to set the time window for short-term energy.

[0039] Furthermore, audio data whose audio quality results do not meet the audio quality requirements are subjected to audio repair to obtain repaired audio data, including:

[0040] Audio data whose audio quality results do not meet the audio quality requirements are extracted as audio data to be processed;

[0041] The audio data to be processed is subjected to noise removal processing using a filter to obtain the noise-removed audio data.

[0042] Distortion type analysis is performed on the noise-removed audio data to obtain the distortion type of the noise-removed audio data.

[0043] Retrieve the distortion processing method corresponding to the distortion type from the database;

[0044] The audio signal corresponding to each distortion type in the noise-removed audio data is compensated and repaired using the distortion processing method corresponding to the distortion type, and the distortion-repaired audio data is obtained.

[0045] The missing data in the distortion-repaired audio data is identified, and the index position corresponding to the missing data is obtained;

[0046] Extract the audio signal data values ​​of the data points before and after the index position;

[0047] The missing data is filled by using interpolation algorithms (such as linear interpolation and spline interpolation) and combining the audio signal data values ​​of the data points before and after the index position.

[0048] The missing data is filled with the corresponding padding values ​​to obtain the audio data after missing value repair.

[0049] Further, the repaired audio data is input into the ASR model to obtain the corresponding text data, including:

[0050] Extract the short-time energy average of the audio data after the current missing values ​​have been repaired;

[0051] The short-time energy average value is compared with a preset short-time energy threshold.

[0052] When the average short-time energy is lower than the preset short-time energy threshold, the model time window of the ASR model will not be adjusted.

[0053] When the average short-time energy is not lower than the preset short-time energy threshold, the model time window and window overlap ratio of the ASR model are adjusted to obtain the adjusted model time window length and window overlap ratio.

[0054] The adjusted model time window length is obtained using the following formula:

[0055] Among them, T s Indicates the adjusted model time window length; T s0 represents the adjusted model time window length; n represents the number of time windows corresponding to the short-time energy contained in the audio data; W i W represents the short-time energy corresponding to the i-th time window; max and W min W represents the maximum and minimum short-time energy values ​​for n time windows corresponding to the audio data. b E represents the short-time energy standard deviation corresponding to n time windows; m E represents the minimum short-time energy value of the time frame corresponding to the time frame containing the time window in which the maximum short-time energy value is located; n W represents the maximum short-time energy value of the time frame corresponding to the time frame containing the time window in which the short-time energy minimum occurs; p W represents the short-term energy average. y This indicates the preset short-term energy threshold;

[0056] Meanwhile, the adjusted window overlap ratio is obtained using the following formula:

[0057] Among them, B c Indicates the adjusted window overlap ratio; B c0 This indicates the window overlap ratio before adjustment; T s Indicates the adjusted model time window length; T s0 E represents the adjusted model time window length; m E represents the minimum short-time energy value of the time frame corresponding to the time frame containing the time window in which the maximum short-time energy value is located; n This represents the maximum short-time energy value of the time frame corresponding to the time frame containing the time window in which the short-time energy minimum occurs;

[0058] The ASR model was adjusted according to the adjusted model time window length and window overlap ratio.

[0059] The repaired audio data is input into the ASR model, and the ASR model processes the repaired audio data to obtain the text string corresponding to the repaired audio data.

[0060] The text string is processed to obtain the processed text data; wherein, the data processing includes, but is not limited to, removing extra spaces, processing punctuation marks, unifying font and font size, and correcting spelling errors.

[0061] Further, keyword and phrase recognition is performed from the text data to obtain target keywords and phrases, including:

[0062] Retrieve a preset list of keywords and phrases;

[0063] Iterate through the text data and use string matching methods to find whether keywords and phrases exist in the text data;

[0064] When keywords and phrases are found in text data, their positions in the text and their context are recorded and marked.

[0065] Use the found keywords and phrases as target keywords and phrases.

[0066] Furthermore, the target keywords and phrases are compared and identified with a preset keyword and phrase sentiment index table to obtain sentiment index results, including:

[0067] Retrieve a table comparing keywords and phrases with consistent structure and sentiment.

[0068] Remove empty and duplicate entries from the keyword and phrase versus sentiment comparison table to obtain the processed keyword and phrase versus sentiment comparison table.

[0069] The target keywords and phrases are compiled into a list of target keywords and phrases;

[0070] The list of target keywords and phrases is traversed.

[0071] During the traversal, the string matching method checks whether the keywords and phrases in the target keyword and phrase list appear in the processed keyword and phrase vs. sentiment table;

[0072] If the keywords and phrases in the target keyword and phrase list appear in the processed keyword and phrase sentiment comparison table, then the sentiment sentiment corresponding to the keywords and phrases is retrieved to obtain the sentiment sentiment result.

[0073] An ASR-based early media recognition system, the ASR-based early media recognition system comprising:

[0074] The audio data collection and processing module is used to collect audio data corresponding to the target media in real time, compare the audio data to repair it, and obtain the repaired audio data.

[0075] The text conversion module is used to input the repaired audio data into the ASR model to obtain the text data corresponding to the audio data.

[0076] The key information identification module is used to identify keywords and phrases from the text data and obtain target keywords and phrases;

[0077] The sentiment tendency recognition module is used to compare and identify the target keywords and phrases with a preset keyword and phrase sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are the early media recognition results.

[0078] The content push control module is used to push media content according to the sentiment trend results.

[0079] Beneficial effects of this invention:

[0080] This invention proposes an early media recognition method and system based on ASR (Acoustic Resonance Scale). By utilizing ASR technology, audio data is converted into text data in real time, significantly shortening recognition time and improving efficiency. Employing the acoustic and language models of the ASR model, along with natural language processing techniques, it can accurately identify keywords and phrases in audio data, as well as their corresponding sentiment tendencies. Based on the sentiment tendency results, media content can be personalized for filtering and classification, thereby meeting users' individual needs. This method is not only applicable to news broadcasts and social media videos, but can also be extended to smart homes, healthcare, finance, and other fields, enabling more intelligent and personalized human-computer interaction. Attached Figure Description

[0081] Figure 1 is a flowchart of the method described in this invention;

[0082] Figure 2 is a system block diagram of the system described in this invention. Detailed Implementation

[0083] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0084] An early media identification method based on ASR is proposed in this embodiment of the invention, as shown in Figure 1. The early media identification method based on ASR includes:

[0085] S1. Collect audio data corresponding to the target media in real time, compare the audio data to repair it, and obtain the repaired audio data;

[0086] S2. Input the repaired audio data into the ASR model to obtain the text data corresponding to the audio data;

[0087] S3. Recognize keywords and phrases from the text data to obtain target keywords and phrases;

[0088] S4. The target keywords and phrases are compared and identified with the preset keywords and phrases and sentiment tendency table to obtain the sentiment tendency result, and the sentiment tendency result is the early media identification result.

[0089] S5. Push media content according to the stated sentiment results.

[0090] The working principle of the above technical solution is as follows: Audio data corresponding to target media (such as news broadcasts, social media videos, etc.) is collected in real time. The collected audio data is preprocessed, including noise reduction and frame segmentation, to improve audio quality. The audio data is compared and necessary repairs are performed, such as removing background noise and filling in silent segments, to obtain repaired audio data. The repaired audio data is input into an ASR model, where its acoustic and language models are used to convert the audio data into corresponding text data. The acoustic model is responsible for converting speech signals into corresponding acoustic feature sequences, while the language model, trained on a large amount of text data, is used to evaluate whether the generated text sequences conform to the grammatical rules and idiomatic usage of the language. From the converted text data, natural language processing techniques are used to identify keywords and phrases. Representative and informational target keywords and phrases are extracted. The extracted target keywords and phrases are compared with a preset keyword and phrase sentiment index. Based on the comparison results, the sentiment index result corresponding to the audio data is obtained, such as positive, negative, or neutral. This sentiment index result is the early media identification result, used for subsequent content recommendation decisions. Media content is filtered and categorized based on sentiment analysis. Content that aligns with user preferences or needs is then pushed to users.

[0091] The above technical solution achieves the following results: By using ASR technology to convert audio data into text data in real time, it significantly shortens recognition time and improves recognition efficiency. Utilizing the acoustic and language models of the ASR model, along with natural language processing techniques, it can accurately identify keywords and phrases in audio data, as well as their corresponding sentiment tendencies. Based on the sentiment tendency results, media content can be personalized for filtering and categorization, thereby meeting users' individual needs. This method is not only applicable to news broadcasts and social media videos, but can also be extended to smart homes, healthcare, finance, and other fields, enabling more intelligent and personalized human-computer interaction.

[0092] In summary, the ASR-based early media identification method achieves rapid and accurate identification and personalized delivery of media content through steps such as real-time collection and repair of audio data, speech recognition, keyword and phrase recognition, sentiment analysis, and content delivery. It has broad application prospects and significant technical value.

[0093] One embodiment of the present invention involves collecting audio data corresponding to a target media in real time, comparing and repairing the audio data, and obtaining repaired audio data, including:

[0094] S101. Collect audio data corresponding to the target media in real time;

[0095] S102. Enhance the audio data to obtain enhanced audio data, wherein the enhancement process includes noise reduction and popping removal and volume normalization.

[0096] S103. Evaluate the audio quality of the enhanced audio data to obtain audio quality results;

[0097] S104. Perform audio repair on audio data whose audio quality results do not meet the audio quality requirements, and obtain the repaired audio data.

[0098] The working principle of the above technical solution is as follows: This step is the starting point of the entire process, involving real-time capture of audio data from the target media (such as microphones, recording equipment, network streaming media, etc.). This requires the system to have efficient audio data acquisition capabilities to ensure the real-time nature and integrity of the audio data. Noise reduction and popping removal: By analyzing the noise components and popping (i.e., sudden bursts of extremely high volume) characteristics in the audio signal, appropriate filtering or elimination algorithms are used to reduce or eliminate these adverse effects, thereby improving the clarity of the audio. The volume of the audio signal is adjusted to a relatively stable range to ensure consistent audio playback volume and avoid affecting the listening experience due to sudden volume changes. This step typically involves a series of objective and subjective audio quality evaluation indicators. Objective indicators, such as signal-to-noise ratio (SNR) and distortion, are used to quantify the purity and distortion level of the audio signal; subjective indicators assess the listening quality of the audio by listening to audio samples and giving subjective scores. By integrating these indicators, the system can comprehensively evaluate the audio quality and determine whether further repair processing is needed. Audio data whose audio quality results do not meet the requirements will undergo audio repair. For audio data whose audio quality assessment results do not meet the requirements, the system will perform targeted repair processing. This may include using more advanced noise reduction algorithms, audio synthesis techniques, or audio compensation algorithms to recover lost or damaged audio information, thereby improving the overall audio quality.

[0099] The above technical solution achieves the following effects: through noise reduction, popping, volume normalization, and targeted repair processing, it significantly improves audio clarity, purity, and listening experience. The ability to collect audio data in real-time ensures the real-time nature and integrity of the data, which is particularly important for audio applications requiring real-time processing (such as real-time communication and online conferencing). By integrating audio enhancement, quality assessment, and repair steps, this solution automates and automates audio processing, reducing the cost and time of manual intervention. Furthermore, by improving audio quality, it enhances the reliability and stability of audio applications, reducing communication interruptions or experience degradation caused by audio quality issues.

[0100] In summary, this technical solution achieves real-time collection, enhancement, quality assessment, and repair of audio data through a series of efficient audio processing steps, thereby significantly improving the overall audio quality and listening experience.

[0101] One embodiment of the present invention involves evaluating the audio quality of enhanced audio data to obtain an audio quality result, including:

[0102] S1031. Extract the audio parameters of the enhanced audio data, wherein the audio parameters include signal-to-noise ratio, total harmonic distortion, and short-time energy.

[0103] S1032. Obtain audio quality evaluation parameters using the signal-to-noise ratio, total harmonic distortion, and short-time energy contained in the audio parameters;

[0104] The audio quality evaluation parameters are obtained using the following formula:

[0105] Where P represents the audio quality assessment parameter; n represents the number of time windows corresponding to the short-time energy contained in the audio data; S i T represents the signal-to-noise ratio corresponding to the i-th time window; i S represents the total harmonic distortion corresponding to the i-th time window; max and S min W represents the maximum and minimum signal-to-noise ratios for n time windows corresponding to the audio data. max and W min W represents the maximum and minimum short-time energy values ​​for n time windows corresponding to the audio data. i W represents the short-time energy corresponding to the i-th time window; bi This represents the short-time energy standard deviation corresponding to the i-th time window;

[0106] S1033. Compare the audio quality evaluation parameters with preset audio quality evaluation parameter thresholds;

[0107] S1034. When the audio quality evaluation parameter is lower than the preset audio quality evaluation parameter threshold, the audio data is determined to be audio data whose audio quality result does not meet the audio quality requirements.

[0108] The working principle of the above technical solution is as follows: Key audio parameters, including signal-to-noise ratio (SNR), total harmonic distortion (THD), and short-time energy, are extracted from the enhanced audio data. These parameters reflect the quality characteristics of the audio signal, such as purity, distortion level, and energy distribution. Using the extracted audio parameters, the audio quality assessment parameter P is calculated through a specific formula. This formula considers multiple aspects of SNR, THD, and short-time energy, including their values, maximum values, minimum values, and standard deviations at different time windows. In the formula, n represents the number of time windows corresponding to the short-time energy contained in the audio data, and S... i T i W i S represents the signal-to-noise ratio, total harmonic distortion, and short-time energy corresponding to the i-th time window, respectively. max S min W max W min These represent the maximum and minimum signal-to-noise ratio (SNR) and the maximum and minimum short-time energy for the n time windows corresponding to the audio data, respectively. W bi Let P represent the short-time energy standard deviation corresponding to the i-th time window. Through this comprehensive calculation method, an evaluation parameter P that fully reflects audio quality can be obtained. The calculated audio quality evaluation parameter P is compared with a preset audio quality evaluation parameter threshold. If P is lower than the threshold, the audio quality result is determined to be unsatisfactory and further repair processing is required.

[0109] The above technical solution achieves the following results: By extracting multiple key audio parameters and comprehensively calculating the audio quality assessment parameter P, this solution can comprehensively and objectively evaluate the quality of audio signals. This helps to accurately identify problems in audio signals, such as excessive noise or severe distortion, providing strong support for subsequent repair processing. Through an automated evaluation process, this solution reduces the cost and time of manual intervention. Simultaneously, because the evaluation results are objective and accurate, it avoids inconsistencies in processing caused by subjective judgment, improving the efficiency and stability of audio processing. By comparing the audio quality assessment parameter P with a preset threshold, this solution can accurately determine which audio data requires repair processing. This helps to develop targeted repair strategies, optimize repair effects, and avoid unnecessary processing of audio data that does not require repair, saving resources. By improving the accuracy and efficiency of audio quality assessment, this solution can provide users with a higher quality audio experience. This is particularly important in applications such as audio communication, online conferencing, and music playback, helping to improve user satisfaction and loyalty.

[0110] On the other hand, this technical solution extracts multiple key audio parameters such as signal-to-noise ratio, total harmonic distortion, and short-time energy, and integrates these factors to calculate audio quality assessment parameters, thus reflecting the quality characteristics of the audio signal more comprehensively. This comprehensive assessment method is more accurate than single-parameter assessment and can reduce misjudgments caused by abnormalities in a single parameter. The calculation process of the assessment parameters is based on objective mathematical formulas and statistical methods, avoiding the uncertainty brought about by subjective judgment. This makes the assessment results more objective and accurate, providing a reliable basis for subsequent repair processing. This technical solution realizes the automation of the audio quality assessment process, reducing the cost and time of manual intervention. Automated assessment can significantly improve processing efficiency, especially in large-scale audio data processing scenarios, and can greatly shorten the assessment cycle. By extracting audio parameters and calculating assessment parameters in real time, this technical solution can achieve rapid response to audio quality issues. This helps to promptly detect and handle problems in audio signals, improving the real-time performance and effectiveness of audio processing. At the same time, this technical solution is applicable to various audio scenarios, such as music, speech, and noise, and has wide applicability. By adjusting the calculation formula and threshold settings of the assessment parameters, it can adapt to the quality assessment needs of different audio scenarios. This technical solution effectively suppresses the influence of noise and interference signals during the extraction and calculation of audio parameters. This makes the evaluation results more stable and reliable, maintaining high accuracy even in complex and variable audio environments. By comparing audio quality evaluation parameters with preset thresholds, this solution can accurately determine which audio data requires repair. This helps in developing targeted repair strategies, optimizing repair effects, and improving the overall quality of audio data. Accurate audio quality assessment and effective repair processing enhance the user's listening experience. In applications such as audio communication, online conferencing, and music playback, high-quality audio signals can enhance user engagement and satisfaction.

[0111] In summary, the technical benefits of this solution are mainly reflected in assessment accuracy, processing efficiency, robustness, and optimization effectiveness. These effects collectively enhance the reliability and effectiveness of audio quality assessment, providing strong support for subsequent audio processing and applications. Furthermore, by extracting key audio parameters, calculating audio quality assessment parameters, and comparing and judging them, this solution achieves comprehensive evaluation and optimization of the enhanced audio data. This helps improve the efficiency and stability of audio processing, optimize audio restoration strategies, and enhance the user experience.

[0112] In one embodiment of the present invention, setting a time window corresponding to the short-time energy contained in the audio parameters of the audio data includes:

[0113] Step 1: Extract the sampling frequency, bit depth, and number of channels corresponding to the audio data collection device;

[0114] Step 2: Obtain the first window adjustment coefficient using the sampling frequency, bit depth, and number of channels corresponding to the audio data collection device;

[0115] The first window adjustment coefficient is obtained by the following formula:

[0116] Among them, K 01 f represents the adjustment coefficient of the first window; f represents the sampling frequency corresponding to the audio data collection device; f c This indicates the preset sampling frequency reference value; N represents the number of channels; B represents the bit depth; int() indicates rounding up the value within the parentheses;

[0117] Step 3: Compare the first window adjustment coefficient with the preset window adjustment coefficient threshold;

[0118] Step 4: When the first window adjustment coefficient is lower than the preset window adjustment coefficient threshold, the preset initial time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy.

[0119] Step 5: When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy.

[0120] The working principle of the above technical solution is as follows: First, three key parameters—sampling frequency, bit depth, and number of channels—are extracted from the audio data collection device (such as a microphone or recorder). These parameters directly determine the quality and characteristics of the audio data. Using the extracted sampling frequency, bit depth, and number of channels, a first window adjustment coefficient is calculated using a specific formula. This coefficient reflects the impact of the audio data collection device's performance on the time window setting. The sampling frequency reference value in the formula is a preset value used to compare with the actual sampling frequency, thereby adjusting the size of the window adjustment coefficient. The number of channels and bit depth are also included in the formula as influencing factors. The calculated first window adjustment coefficient is compared with a preset window adjustment coefficient threshold. This threshold is used to determine whether the initial time length needs to be adjusted to adapt to different audio data collection devices and audio data quality. If the first window adjustment coefficient is lower than the threshold, it indicates that the current audio data collection device has low performance or poor audio data quality; in this case, the preset initial time length is used as the length of the time window. If the first window adjustment coefficient is not lower than the threshold, it indicates that the current audio data collection device has high performance or good audio data quality; in this case, the initial time length needs to be adjusted to better adapt to the characteristics of the audio data. The adjusted time length will be used as the length of the time window for the calculation and analysis of short-time energy.

[0121] The above technical solution achieves the following results: By setting the length of the time window according to the performance of the audio data collection device and the characteristics of the audio data, short-time energy can be extracted and analyzed more effectively. This helps to more accurately reflect the amplitude changes and characteristics of the audio signal. The solution automatically adjusts the time window length based on device parameters and audio data quality, avoiding the tediousness and uncertainty of manual settings. This improves the efficiency and automation of audio data processing. The solution is applicable to audio data collection devices with different performance and quality levels, as well as audio data with different characteristics. By adjusting the time window length, accurate results can be obtained in different scenarios. Accurate short-time energy analysis helps improve the processing effect and application performance of audio data. In fields such as audio communication, speech recognition, and audio editing, this can enhance the user's auditory experience and satisfaction.

[0122] On the other hand, by dynamically adjusting the length of the time window based on parameters such as the sampling frequency, bit depth, and number of channels of the audio data collection device, the short-time energy characteristics of the audio signal can be captured more effectively. Compared with a fixed time window setting, this dynamic adjustment method can more accurately reflect the amplitude changes and characteristics of the audio signal, thereby improving the accuracy of short-time energy analysis. By accurately calculating the first window adjustment coefficient and comparing it with a preset window adjustment coefficient threshold, the length of the time window can be automatically selected or adjusted to reduce errors and interference introduced by differences in device performance or fluctuations in audio data quality. This technical solution realizes an automated processing flow for time window setting, reducing the tediousness and uncertainty of manual setting. Automated processing can significantly improve the efficiency of audio data processing, especially in large-scale audio data processing scenarios, and can greatly shorten the processing cycle. By reasonably setting the length of the time window, unnecessary consumption of computing resources can be reduced while ensuring the accuracy of analysis. This helps to reduce the cost of audio data processing and improve the overall performance and efficiency of the system. This technical solution is applicable to audio data collection devices with different performance and quality, as well as audio data with different characteristics. By dynamically adjusting the length of the time window, accurate results can be obtained for short-time energy analysis in different scenarios, thereby enhancing the system's wide applicability. The formulas and thresholds in this technical solution can be adjusted and optimized according to actual needs to adapt to different application scenarios and changing requirements. This flexibility and scalability help maintain the system's advanced nature and competitiveness. Accurate short-time energy analysis helps improve the processing effect and application performance of audio data. In fields such as audio communication, speech recognition, and audio editing, this can improve the clarity and quality of audio signals, thereby improving the user's listening experience. Automated processing reduces the user's operational burden, making it more convenient for users to use the audio data processing system. This helps improve user satisfaction and loyalty.

[0123] In summary, the technical benefits of this solution in terms of performance indicators are mainly reflected in improved accuracy, optimized efficiency, enhanced adaptability, and improved user experience. These effects collectively improve the performance and quality of audio data processing, providing strong support for subsequent audio analysis and applications. Furthermore, by setting the time window length according to the performance of the audio data collection device and the characteristics of the audio data, this solution improves the accuracy and efficiency of short-time energy analysis, and enhances the adaptability and user experience of audio data processing.

[0124] In one embodiment of the present invention, when the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy, including:

[0125] Step 501: When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, extract the initial time length and initial window overlap ratio corresponding to the time window of the ASR model; wherein, the ASR model adopts a deep neural network model structure.

[0126] Step 502: Extract the time length values ​​and corresponding window overlap ratio values ​​that occur during the dynamic adjustment of the time window of the ASR model.

[0127] Step 503: Use the initial time length and initial window overlap ratio of the time window corresponding to the ASR model, combined with the time length and corresponding window overlap ratio that appear during the dynamic adjustment of the time window, to obtain the second window adjustment coefficient.

[0128] The second window adjustment coefficient is obtained using the following formula:

[0129] Among them, K 02 Indicates the second window adjustment coefficient; m represents the number of dynamic adjustments to the time window; T i P represents the time length corresponding to the i-th dynamically adjusted time window; i P represents the window overlap ratio after the i-th dynamic adjustment; ti T represents the rate of change in the length of the time window after the i-th dynamic adjustment compared to the time window before the dynamic adjustment; c and P c P represents the initial time length and the initial window overlap ratio. z P represents the median value of the window overlap ratio after m dynamic adjustments; b T represents the standard deviation of the window overlap ratio after m dynamic adjustments; b P represents the standard deviation of the time window length after m dynamic adjustments; max This represents the maximum value of the window overlap ratio after m dynamic adjustments.

[0130] Step 504: Adjust the initial time length using the first window adjustment coefficient and the second window adjustment coefficient to obtain the adjusted time length;

[0131] The adjusted time length is obtained using the following formula:

[0132] Among them, T x T0 represents the adjusted time length; T0 represents the original time length; K 01 K represents the first window adjustment coefficient; 02represents the second window adjustment coefficient; t represents the adjustment factor, and the value range of the adjustment factor is 1.12-1.43;

[0133] Step 505: Use the adjusted time length as the corresponding length of the time window to set the time window corresponding to the short-time energy.

[0134] The working principle of the above technical solution is as follows: When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the system first extracts the initial time length and initial window overlap ratio corresponding to the time window of the ASR model. These initial values ​​serve as the basis for subsequent dynamic adjustments. The system records the adjusted time length and corresponding window overlap ratio of the ASR model time window during the dynamic adjustment process. These records are used to calculate the second window adjustment coefficient. Using the dynamic adjustment data recorded above, combined with the initial time length and initial window overlap ratio, the second window adjustment coefficient is calculated using a specific formula. This coefficient reflects the overall change and trend during the dynamic adjustment of the time window. Combining the first window adjustment coefficient and the second window adjustment coefficient, as well as a preset adjustment factor, the initial time length is adjusted using a specific formula to obtain the adjusted time length. This adjustment process considers the performance of the audio data collection device, the characteristics of the ASR model, and historical data of the dynamic adjustment of the time window. Finally, the adjusted time length is used as the length of the time window to set the time window corresponding to the short-time energy. The time window set in this way can more accurately reflect the short-time energy characteristics of the audio signal, thereby improving the performance and accuracy of the ASR model.

[0135] The above technical solution achieves the following effects: By dynamically adjusting the length of the time window, it can more accurately capture the short-time energy characteristics of audio signals, thereby improving the recognition performance and accuracy of the ASR model. This solution can automatically adjust the time window length based on the performance of different audio data collection devices and the characteristics of the ASR model, thus enhancing the system's adaptability and flexibility. By rationally setting the time window length and overlap ratio, this solution can optimize the utilization of computing resources and reduce system power consumption and cost while ensuring ASR model performance. Accurate ASR model performance and optimized resource utilization can improve the user's auditory experience and satisfaction, especially in application scenarios requiring efficient and accurate speech recognition.

[0136] By dynamically adjusting the length of the time window, this technique can more accurately capture the short-time energy features of audio signals, thereby improving the recognition accuracy of the Automatic Speech Recognition (ASR) model. This helps reduce recognition errors and misjudgments, improving the overall performance of the system. Optimized time window settings reduce unnecessary computation and processing time, thus improving the response speed of the ASR model. This allows the system to complete recognition tasks faster when processing large amounts of audio data, improving processing efficiency. By reasonably setting the length and overlap ratio of the time window, this technique can optimize the utilization of computing resources while ensuring the performance of the ASR model. This helps reduce system power consumption and cost, improving resource utilization efficiency. Optimized time window settings reduce the storage requirements of audio data, thus saving storage space. This is particularly important for systems that need to process large amounts of audio data, reducing storage costs and improving system scalability. Dynamically adjusting the length of the time window helps enhance the ASR model's resistance to noise interference. In noisy environments, this technique can more effectively extract the short-time energy features of audio signals, thereby improving the stability and accuracy of recognition. By recording and analyzing the dynamic adjustment process of the time window, this technique can promptly detect and correct potential errors and deviations. This helps improve the system's fault tolerance and robustness, ensuring stable operation under various conditions. Optimized ASR model performance and response speed enhance user-system interactivity. Users can interact with the system more quickly and accurately, thereby improving the overall user experience. Accurate identification performance and optimized resource utilization increase user satisfaction and trust. This helps strengthen user confidence and reliance on the system, improving user stickiness and market competitiveness.

[0137] In summary, the technical benefits of this solution are mainly reflected in improved recognition performance, optimized resource utilization efficiency, enhanced system stability and reliability, and improved user experience. These effects collectively improve the overall system performance and user experience, providing strong support for subsequent audio analysis and applications. Furthermore, by dynamically adjusting the time window length, this solution improves the performance and accuracy of the ASR model, enhances the system's adaptability and flexibility, optimizes the utilization of computing resources, and improves the user experience.

[0138] One embodiment of the present invention involves audio repair of audio data whose audio quality results do not meet audio quality requirements, and obtaining repaired audio data, including:

[0139] S1041. Extract audio data whose audio quality results do not meet the audio quality requirements as audio data to be processed.

[0140] S1042. Use a filter to perform noise removal processing on the audio data to be processed, and obtain the audio data after noise removal processing.

[0141] S1043. Perform distortion type analysis on the audio data after noise removal to obtain the distortion type of the audio data after noise removal.

[0142] S1044. Retrieve the distortion processing method corresponding to the distortion type from the database;

[0143] S1045. Using the distortion processing method corresponding to the distortion type, the audio signal corresponding to each distortion type in the audio data after noise removal is compensated and repaired to obtain the distortion-repaired audio data.

[0144] S1046. Determine the missing data in the distortion-repaired audio data and obtain the index position corresponding to the missing data;

[0145] S1047. Extract the audio signal data values ​​of the data points before and after the index position;

[0146] S1048. Use interpolation algorithms (such as linear interpolation and spline interpolation) to combine the audio signal data values ​​of the data points before and after the index position to obtain the filling value corresponding to the missing data;

[0147] S1049. Fill in each missing data using the filling value corresponding to the missing data to obtain the audio data after missing value repair.

[0148] The working principle of the above technical solution is as follows: Audio data that does not meet the quality requirements is selected from the original audio data and designated as the audio data to be processed. A filter is used to process the audio data to remove noise components. The filter can be selected and set according to the frequency characteristics or statistical characteristics of the noise to effectively remove different types of noise. Distortion type analysis is performed on the noise-removed audio data. By analyzing the spectrum and time-domain characteristics of the audio signal, the type of distortion present in the audio data can be determined, such as noise distortion, compression distortion, and frequency response distortion. Based on the analyzed distortion type, the corresponding distortion processing method is retrieved from the database. The database can store multiple distortion types and their corresponding processing methods to quickly and accurately find suitable repair solutions. The retrieved distortion processing method is used to compensate and repair each type of distortion in the audio data. The repair methods may include filter adjustment, dynamic processing, equalizer adjustment, etc., to restore the original quality of the audio signal. In the distortion-repaired audio data, there may be missing data due to various reasons. By determining the index position of the missing data, the audio signal data values ​​of the data points before and after the corresponding index position are extracted. Then, interpolation algorithms (such as linear interpolation, spline interpolation, etc.) are used to calculate the filling value corresponding to the missing data. Finally, the filling value is used to fill each missing data to obtain the audio data with missing values ​​repaired.

[0149] The above technical solution achieves the following results: through noise removal and distortion compensation, it significantly reduces noise and distortion components in audio data, thereby improving the overall audio quality. The distortion compensation process can target different types of distortion, helping to restore the original details and features of the audio signal. By identifying and filling missing data, it can fill in blank or missing parts of the audio data, improving the integrity and continuity of the audio. The repaired audio data is of higher quality and more complete, and can be used in more application scenarios, such as music production, speech recognition, and audio analysis. This technical solution employs an automated processing workflow, enabling rapid and efficient processing of large amounts of audio data, improving processing efficiency and reducing labor costs.

[0150] In conclusion, this technical solution has high application value in audio restoration and can effectively improve the quality and usability of audio data.

[0151] In one embodiment of the present invention, the repaired audio data is input into an ASR model to obtain the text data corresponding to the audio data, including:

[0152] S201. Extract the short-time energy average of the audio data after the current missing values ​​have been repaired;

[0153] S202. Compare the short-time energy average value with a preset short-time energy threshold.

[0154] S203. When the average short-time energy is lower than the preset short-time energy threshold, the model time window of the ASR model will not be adjusted.

[0155] S204. When the average short-time energy value is not lower than the preset short-time energy threshold, the model time window and window overlap ratio of the ASR model are adjusted, and the adjusted model time window length and window overlap ratio are obtained.

[0156] The adjusted model time window length is obtained using the following formula:

[0157] Among them, T s Indicates the adjusted model time window length; T s0 represents the adjusted model time window length; n represents the number of time windows corresponding to the short-time energy contained in the audio data; W i W represents the short-time energy corresponding to the i-th time window; max and W min W represents the maximum and minimum short-time energy values ​​for n time windows corresponding to the audio data. b E represents the short-time energy standard deviation corresponding to n time windows; m E represents the minimum short-time energy value of the time frame corresponding to the time frame containing the time window in which the maximum short-time energy value is located; n W represents the maximum short-time energy value of the time frame corresponding to the time frame containing the time window in which the short-time energy minimum occurs; p W represents the short-term energy average. y This indicates the preset short-term energy threshold;

[0158] Meanwhile, the adjusted window overlap ratio is obtained using the following formula:

[0159] Among them, B c Indicates the adjusted window overlap ratio; B c0 This indicates the window overlap ratio before adjustment; T s Indicates the adjusted model time window length; T s0 E represents the adjusted model time window length; m E represents the minimum short-time energy value of the time frame corresponding to the time frame containing the time window in which the maximum short-time energy value is located; n This represents the maximum short-time energy value of the time frame corresponding to the time frame containing the time window in which the short-time energy minimum occurs;

[0160] S205. Adjust the ASR model according to the adjusted model time window length and window overlap ratio;

[0161] S206. Input the repaired audio data into the ASR model, process the repaired audio data through the ASR model, and obtain the text string corresponding to the repaired audio data.

[0162] S207. Perform data processing on the text string to obtain processed text data; wherein, the data processing includes, but is not limited to, removing redundant spaces, processing punctuation marks, unifying font and font size, and correcting spelling errors.

[0163] The working principle of the above technical solution is as follows: First, the short-time energy average of the audio data after the current missing value repair is extracted. Short-time energy is an important feature in audio signal analysis, reflecting the energy changes of the audio signal within a short period of time. Next, this short-time energy average is compared with a preset short-time energy threshold. This threshold is preset based on the performance of the ASR model and the characteristics of the audio data, and is used to determine whether the energy level of the audio data is high enough to support accurate speech recognition. If the short-time energy average is lower than the preset threshold, it indicates that the energy of the audio data is low and may not be sufficient to support accurate speech recognition; therefore, the model time window of the ASR model is not adjusted. If the short-time energy average is not lower than the preset threshold, it indicates that the energy of the audio data is high enough, and the model time window and window overlap ratio of the ASR model can be adjusted. The adjusted model time window length is calculated based on a series of complex formulas that consider factors such as the short-time energy distribution, maximum and minimum values, standard deviation, and minimum and maximum values ​​of the short-time energy of the time frame. Meanwhile, the adjusted window overlap ratio was calculated using a similar formula to ensure that appropriate window overlap was maintained while adjusting the time window length, in order to capture detailed information in the audio signal. Finally, the ASR model was adjusted according to the adjusted model time window length and window overlap ratio to adapt to the characteristics of the current audio data. The repaired audio data was input into the adjusted ASR model, which processed the audio data to obtain the corresponding text string. Data processing was performed on the obtained text string, including removing extra spaces, processing punctuation, standardizing font and font size, and correcting spelling errors, to obtain high-quality text data.

[0164] The above technical solution achieves the following results: By dynamically adjusting the model time window and window overlap ratio of the ASR model, optimization can be performed based on the characteristics of the audio data, thereby improving the accuracy of speech recognition. This technical solution can dynamically adjust according to different audio data characteristics, enhancing the adaptability and flexibility of the ASR model. A series of data processing steps on the acquired text string can remove redundant information and correct errors, thus optimizing the quality of the text data. High-quality text data can provide users with more accurate and clear information, improving user experience and satisfaction. This technical solution employs an automated processing flow, enabling rapid and efficient processing of large amounts of audio data, improving processing efficiency.

[0165] On the other hand, by dynamically adjusting the model time window and window overlap ratio of the ASR model, the model can better adapt to the characteristics of different audio data. This adjustment can significantly improve the accuracy of speech recognition, enabling the model to maintain a high recognition rate when processing audio data with different energy levels and speech characteristics. This technical solution employs an automated processing flow, including short-time energy analysis, ASR model adjustment, speech recognition, and text processing. This automated process significantly improves processing efficiency and reduces the time cost of manual intervention. Simultaneously, the dynamic adjustment of the model time window and window overlap ratio makes the model more efficient in processing audio data, thereby improving the overall processing speed. By accurately calculating the model time window length and window overlap ratio, it is ensured that the ASR model can fully utilize computing resources when processing audio data. The optimized model reduces unnecessary computational overhead and improves resource utilization when processing audio data. This not only helps reduce computational costs but also helps improve the stability and reliability of the system. This technical solution can dynamically adjust according to different audio data characteristics, enabling the ASR model to maintain high performance when processing different types of audio data. This dynamic adjustment capability enhances the system's adaptability, enabling it to better cope with various complex scenarios and changes. This helps improve the system's robustness and reliability, reducing the risk of performance degradation due to changes in audio data characteristics. High-quality text data is key to improving user experience. This technical solution provides users with more accurate and clear information by optimizing speech recognition accuracy and text processing quality. The optimized system can provide users with higher-quality text data, thereby improving user experience and satisfaction. This helps enhance user trust and reliance on the system, improving its market competitiveness.

[0166] In summary, the technical benefits of the above-mentioned solution in terms of performance indicators are mainly reflected in improved speech recognition accuracy, increased processing efficiency, optimized resource utilization, enhanced system adaptability, and improved user experience. These technical effects work together to make the system more efficient, accurate, and reliable in processing audio data. Furthermore, this technical solution has high application value in speech recognition and text processing of audio data, effectively improving the accuracy of speech recognition and the quality of text data.

[0167] One embodiment of the present invention involves identifying keywords and phrases from the text data to obtain target keywords and phrases, including:

[0168] S301. Retrieve the preset list of keywords and phrases;

[0169] S302. Traverse the text data and use string matching methods to find whether keywords and phrases exist in the text data;

[0170] S303. When keywords and phrases are found in text data, record and mark the position of the keywords and phrases in the text and their context;

[0171] S304. Use the found keywords and phrases as target keywords and phrases.

[0172] The working principle of the above technical solution is as follows: First, the system retrieves a list of keywords and phrases from a pre-set database. This list contains all the keywords and phrases that need to be searched in the text data, which may be related to a specific topic, domain, or task. Next, the system traverses the entire text data, using string matching methods (such as exact matching, fuzzy matching, regular expressions, etc.) to search for the existence of the pre-set keywords and phrases in the text. This process may involve preprocessing steps such as word segmentation and stop word removal of the text data to improve the accuracy and efficiency of the search. Recording and marking the position of keywords and phrases and their context (S303): When the system finds keywords and phrases in the text data, it records and marks the specific position of these keywords and phrases in the text (such as the start and end positions), as well as their context information. The context information may include the text content before and after the keywords and phrases, which helps in the subsequent understanding and analysis of the keywords and phrases. Using the found keywords and phrases as target keywords and phrases (S304): Finally, the system outputs or stores the found keywords and phrases as target keywords and phrases. These target keywords and phrases can be used for subsequent tasks such as information extraction, text classification, and sentiment analysis.

[0173] The above technical solution achieves the following results: by using a pre-defined list of keywords and phrases, combined with string matching methods, the system can accurately identify target keywords and phrases in text data. This helps reduce false positives and false negatives, improving accuracy. The solution employs a traversal approach to search the text data, incorporating possible preprocessing steps (such as word segmentation and stop word removal), enabling efficient processing of large volumes of text data. Furthermore, recording the location and context of keywords and phrases facilitates subsequent processing tasks, further enhancing efficiency. The pre-defined list of keywords and phrases can be flexibly configured and updated according to actual needs. This means the system can adapt to different application scenarios and task requirements, exhibiting strong flexibility and scalability. By identifying target keywords and phrases, the system can provide valuable information for subsequent tasks such as information extraction, text classification, and sentiment analysis. This information helps in a deeper understanding of the meaning and characteristics of text data, providing strong support for applications such as decision support and data mining.

[0174] In summary, this technical solution has high application value in keyword and phrase recognition, and can accurately and efficiently identify target keywords and phrases in text data, providing strong support for subsequent processing tasks.

[0175] In one embodiment of the present invention, the target keywords and phrases are compared and identified with a preset keyword and phrase sentiment tendency table to obtain sentiment tendency results, including:

[0176] S401. Retrieve a comparison table of keywords and phrases with a unified structure and their sentiment tendencies;

[0177] S402. Remove empty values ​​and duplicates from the comparison table of keywords and phrases with sentiment tendencies to obtain the processed comparison table of keywords and phrases with sentiment tendencies.

[0178] S403. Form a list of target keywords and phrases;

[0179] S404. Traverse the list of target keywords and phrases;

[0180] S405. During the traversal, the string matching method checks whether the keywords and phrases in the target keyword and phrase list appear in the processed keyword and phrase vs. sentiment table.

[0181] S406. When the keywords and phrases in the target keyword and phrase list appear in the processed keyword and phrase vs. sentiment trend comparison table, the sentiment trend corresponding to the keywords and phrases is retrieved to obtain the sentiment trend result.

[0182] The working principle of the above technical solution is as follows: A pre-defined table of keywords and phrases with corresponding sentiment tendencies is retrieved from a database or other storage location. This table contains multiple keywords and phrases and their corresponding sentiment tendencies (e.g., positive, negative, neutral). The table is preprocessed to remove null values ​​and duplicates. Null values ​​may indicate invalid data, while duplicates can lead to errors in subsequent matching processes; therefore, pre-processing is necessary. The target keywords and phrases to be analyzed are compiled into a list for subsequent traversal and matching. This list is traversed, meaning each keyword and phrase is checked individually. During traversal, string matching methods (e.g., forward maximum matching, backward maximum matching, bidirectional maximum matching, etc.) are used to check if the target keywords and phrases appear in the processed table. The matching process can be viewed as comparing characters one by one until a complete match is found or no match is determined. If a matching keyword or phrase is found in the table, its corresponding sentiment tendency is directly retrieved as the sentiment tendency result.

[0183] The above technical solution achieves the following results: Implemented through a computer program, it automatically matches keywords and phrases and acquires sentiment indicators, significantly improving work efficiency. By preprocessing the lookup table, empty values ​​and duplicates are removed, reducing errors in the matching process. Simultaneously, using string matching methods ensures accuracy. This solution can be applied to different fields and scenarios; simply adjust the keywords and phrases in the lookup table and their corresponding sentiment indicators according to actual needs. As business grows and new keywords emerge, the lookup table can be easily updated to adapt to new requirements. Furthermore, this solution can be combined with other technologies (such as natural language processing and machine learning) to further enhance the accuracy and efficiency of sentiment analysis.

[0184] In summary, this technical solution achieves sentiment analysis of target keywords and phrases in an automated, accurate, flexible, and scalable manner, providing strong support for subsequent decision-making and applications.

[0185] An early media identification system based on ASR is proposed in this embodiment of the invention, as shown in Figure 2. The early media identification system based on ASR includes:

[0186] The audio data collection and processing module is used to collect audio data corresponding to the target media in real time, compare the audio data to repair it, and obtain the repaired audio data.

[0187] The text conversion module is used to input the repaired audio data into the ASR model to obtain the text data corresponding to the audio data.

[0188] The key information identification module is used to identify keywords and phrases from the text data and obtain target keywords and phrases;

[0189] The sentiment tendency recognition module is used to compare and identify the target keywords and phrases with a preset keyword and phrase sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are the early media recognition results.

[0190] The content push control module is used to push media content according to the sentiment trend results.

[0191] The working principle of the above technical solution is as follows: Audio data corresponding to target media (such as news broadcasts, social media videos, etc.) is collected in real time. The collected audio data is preprocessed, including noise reduction and frame segmentation, to improve audio quality. The audio data is compared and necessary repairs are performed, such as removing background noise and filling in silent segments, to obtain repaired audio data. The repaired audio data is input into an ASR model, where its acoustic and language models are used to convert the audio data into corresponding text data. The acoustic model is responsible for converting speech signals into corresponding acoustic feature sequences, while the language model, trained on a large amount of text data, is used to evaluate whether the generated text sequences conform to the grammatical rules and idiomatic usage of the language. From the converted text data, natural language processing techniques are used to identify keywords and phrases. Representative and informational target keywords and phrases are extracted. The extracted target keywords and phrases are compared with a preset keyword and phrase sentiment index. Based on the comparison results, the sentiment index result corresponding to the audio data is obtained, such as positive, negative, or neutral. This sentiment index result is the early media identification result, used for subsequent content recommendation decisions. Media content is filtered and categorized based on sentiment analysis. Content that aligns with user preferences or needs is then pushed to users.

[0192] The above technical solution achieves the following results: By using ASR technology to convert audio data into text data in real time, it significantly shortens recognition time and improves recognition efficiency. Utilizing the acoustic and language models of the ASR model, along with natural language processing techniques, it can accurately identify keywords and phrases in audio data, as well as their corresponding sentiment tendencies. Based on the sentiment tendency results, media content can be personalized for filtering and categorization, thereby meeting users' individual needs. This method is not only applicable to news broadcasts and social media videos, but can also be extended to smart homes, healthcare, finance, and other fields, enabling more intelligent and personalized human-computer interaction.

[0193] In summary, the ASR-based early media identification method achieves rapid and accurate identification and personalized delivery of media content through steps such as real-time collection and repair of audio data, speech recognition, keyword and phrase recognition, sentiment analysis, and content delivery. It has broad application prospects and significant technical value.

[0194] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An early media identification method based on ASR, characterized in that, The ASR-based early media identification method includes: Real-time collection of audio data corresponding to the target media, comparison of the audio data for repair, and acquisition of repaired audio data; The repaired audio data is input into the ASR model to obtain the text data corresponding to the audio data; Keyword and phrase identification is performed on the text data to obtain target keywords and phrases; The target keywords and phrases are compared and identified with a preset keyword and phrase sentiment index table to obtain sentiment index results, which are the early media identification results. Media content is pushed based on the stated sentiment results.

2. The early media identification method based on ASR according to claim 1, characterized in that, Real-time collection of audio data corresponding to the target media, comparison and repair of the audio data, and acquisition of repaired audio data, including: Collect audio data corresponding to the target media in real time; The audio data is enhanced to obtain the enhanced audio data, wherein the enhancement process includes noise reduction and popping removal and volume normalization. The enhanced audio data is evaluated for audio quality to obtain audio quality results. Audio data that does not meet the audio quality requirements is repaired, and the repaired audio data is obtained.

3. The early media identification method based on ASR according to claim 2, characterized in that, The enhanced audio data is then evaluated for audio quality to obtain the following results: Extract audio parameters from the enhanced audio data, wherein the audio parameters include signal-to-noise ratio, total harmonic distortion, and short-time energy; Audio quality assessment parameters are obtained using the signal-to-noise ratio, total harmonic distortion, and short-time energy contained in the audio parameters. The audio quality evaluation parameters are obtained using the following formula: Where P represents the audio quality assessment parameter; n represents the number of time windows corresponding to the short-time energy contained in the audio data; S i T represents the signal-to-noise ratio corresponding to the i-th time window; i S represents the total harmonic distortion corresponding to the i-th time window; max and S min W represents the maximum and minimum signal-to-noise ratios for n time windows corresponding to the audio data. max and W min W represents the maximum and minimum short-time energy values ​​for n time windows corresponding to the audio data. i W represents the short-time energy corresponding to the i-th time window; bi This represents the short-time energy standard deviation corresponding to the i-th time window; The audio quality assessment parameters are compared with preset audio quality assessment parameter thresholds; When the audio quality evaluation parameter is lower than the preset audio quality evaluation parameter threshold, the audio data is determined to be audio data that does not meet the audio quality requirements.

4. The early media identification method based on ASR according to claim 1, characterized in that, Setting the time window corresponding to the short-time energy contained in the audio parameters of the audio data includes: Extract the sampling frequency, bit depth, and number of channels corresponding to the audio data collection device; The first window adjustment coefficient is obtained using the sampling frequency, bit depth, and number of channels corresponding to the audio data collection device; The first window adjustment coefficient is obtained by the following formula: Among them, K 01 f represents the adjustment coefficient of the first window; f represents the sampling frequency corresponding to the audio data collection device; f c This indicates the preset sampling frequency reference value; N represents the number of channels; B represents the bit depth; int() indicates rounding up the value within the parentheses; The first window adjustment coefficient is compared with a preset window adjustment coefficient threshold. When the first window adjustment coefficient is lower than the preset window adjustment coefficient threshold, the preset initial time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy. When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy.

5. The early media identification method based on ASR according to claim 4, characterized in that, When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy, including: When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length and initial window overlap ratio corresponding to the time window of the ASR model are extracted; wherein, the ASR model adopts a deep neural network model structure; Extract the time length values ​​and corresponding window overlap ratio values ​​that occur during the dynamic adjustment of the time window in the ASR model. The second window adjustment coefficient is obtained by combining the initial time length and initial window overlap ratio of the time window corresponding to the time window in the ASR model with the time length and corresponding window overlap ratio during the dynamic adjustment of the time window. The second window adjustment coefficient is obtained using the following formula: Among them, K 02 Indicates the second window adjustment coefficient; m represents the number of dynamic adjustments to the time window; T i P represents the time length corresponding to the i-th dynamically adjusted time window; i P represents the window overlap ratio after the i-th dynamic adjustment; ti T represents the rate of change in the length of the time window after the i-th dynamic adjustment compared to the time window before the dynamic adjustment; c and P c P represents the initial time length and the initial window overlap ratio. z P represents the median value of the window overlap ratio after m dynamic adjustments; b T represents the standard deviation of the window overlap ratio after m dynamic adjustments; b P represents the standard deviation of the time window length after m dynamic adjustments; max This represents the maximum value of the window overlap ratio after m dynamic adjustments. The initial time length is adjusted using the first window adjustment coefficient and the second window adjustment coefficient to obtain the adjusted time length; The adjusted time length is obtained using the following formula: Among them, T x T0 represents the adjusted time length; T0 represents the original time length; K 01 K represents the first window adjustment coefficient; 02 represents the second window adjustment coefficient; t represents the adjustment factor, and the value range of the adjustment factor is 1.12-1.43; The adjusted time length is used as the corresponding length of the time window to set the time window for short-term energy.

6. The early media identification method based on ASR according to claim 2, characterized in that, Audio data that does not meet audio quality requirements is repaired, and the repaired audio data is obtained, including: Audio data whose audio quality results do not meet the audio quality requirements are extracted as audio data to be processed; The audio data to be processed is subjected to noise removal processing using a filter to obtain the noise-removed audio data. Distortion type analysis is performed on the noise-removed audio data to obtain the distortion type of the noise-removed audio data. Retrieve the distortion processing method corresponding to the distortion type from the database; The audio signal corresponding to each distortion type in the noise-removed audio data is compensated and repaired using the distortion processing method corresponding to the distortion type, and the distortion-repaired audio data is obtained. The missing data in the distortion-repaired audio data is identified, and the index position corresponding to the missing data is obtained; Extract the audio signal data values ​​of the data points before and after the index position; The interpolation algorithm is used to combine the audio signal data values ​​of the data points before and after the index position to obtain the filling value corresponding to the missing data; The missing data is filled with the corresponding padding values ​​to obtain the audio data after missing value repair.

7. The early media identification method based on ASR according to claim 1, characterized in that, The repaired audio data is input into the ASR model to obtain the corresponding text data, including: Extract the short-time energy average of the audio data after the current missing values ​​have been repaired; The short-time energy average value is compared with a preset short-time energy threshold. When the average short-time energy is lower than the preset short-time energy threshold, the model time window of the ASR model will not be adjusted. When the average short-time energy is not lower than the preset short-time energy threshold, the model time window and window overlap ratio of the ASR model are adjusted to obtain the adjusted model time window length and window overlap ratio. The adjusted model time window length is obtained using the following formula: Among them, T s Indicates the adjusted model time window length; T s0 represents the adjusted model time window length; n represents the number of time windows corresponding to the short-time energy contained in the audio data; W i W represents the short-time energy corresponding to the i-th time window; max and W min W represents the maximum and minimum short-time energy values ​​for n time windows corresponding to the audio data. b E represents the short-time energy standard deviation corresponding to n time windows; m E represents the minimum short-time energy value of the time frame corresponding to the time frame containing the time window in which the maximum short-time energy value is located; n W represents the maximum short-time energy value of the time frame corresponding to the time frame containing the time window in which the short-time energy minimum occurs; p W represents the short-term energy average. y This indicates the preset short-term energy threshold; Meanwhile, the adjusted window overlap ratio is obtained using the following formula: Among them, B c Indicates the adjusted window overlap ratio; B c0 This indicates the window overlap ratio before adjustment; T s Indicates the adjusted model time window length; T s0 E represents the adjusted model time window length; m E represents the minimum short-time energy value of the time frame corresponding to the time frame containing the time window in which the maximum short-time energy value is located; n This represents the maximum short-time energy value of the time frame corresponding to the time frame containing the time window in which the short-time energy minimum occurs; The ASR model was adjusted according to the adjusted model time window length and window overlap ratio. The repaired audio data is input into the ASR model, and the ASR model processes the repaired audio data to obtain the text string corresponding to the repaired audio data. The text string is processed to obtain the processed text data.

8. The early media identification method based on ASR according to claim 1, characterized in that, Keyword and phrase identification is performed on the text data to obtain target keywords and phrases, including: Retrieve a preset list of keywords and phrases; Iterate through the text data and use string matching methods to find whether keywords and phrases exist in the text data; When keywords and phrases are found in text data, their positions in the text and their context are recorded and marked. Use the found keywords and phrases as target keywords and phrases.

9. The early media identification method based on ASR according to claim 1, characterized in that, The target keywords and phrases are compared and identified with a preset keyword and phrase sentiment index table to obtain sentiment index results, including: Retrieve a table comparing keywords and phrases with consistent structure and sentiment. Remove empty and duplicate entries from the keyword and phrase versus sentiment comparison table to obtain the processed keyword and phrase versus sentiment comparison table. The target keywords and phrases are compiled into a list of target keywords and phrases; The list of target keywords and phrases is traversed. During the traversal, the string matching method checks whether the keywords and phrases in the target keyword and phrase list appear in the processed keyword and phrase vs. sentiment table; If the keywords and phrases in the target keyword and phrase list appear in the processed keyword and phrase sentiment comparison table, then the sentiment sentiment corresponding to the keywords and phrases is retrieved to obtain the sentiment sentiment result.

10. An early media recognition system based on ASR, characterized in that, The ASR-based early media recognition system includes: The audio data collection and processing module is used to collect audio data corresponding to the target media in real time, compare the audio data to repair it, and obtain the repaired audio data. The text conversion module is used to input the repaired audio data into the ASR model to obtain the text data corresponding to the audio data. The key information identification module is used to identify keywords and phrases from the text data and obtain target keywords and phrases; The sentiment tendency recognition module is used to compare and identify the target keywords and phrases with a preset keyword and phrase sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are the early media recognition results. The content push control module is used to push media content according to the sentiment trend results.