A method and system for early media recognition based on ASR

Through the ASR-based early media recognition method, audio data is collected and repaired in real time, and voice-to-text conversion and keyword recognition are used to use the ASR model to perform speech-to-text conversion and keyword recognition, solving the problem that the recognition accuracy in the prior art is affected by environmental noise, accent differences and complex language structure, and achieving rapid and accurate identification and personalized push of media content.

CN119559936BActive Publication Date: 2025-06-06北京基智科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510115938.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The accuracy of the identification under factors such as environmental noise, accent differences and complex language structures is affected, making it difficult to achieve real-time and accurate media content recognition.

Method used

ASR-based early media recognition method is adopted to collect and repair audio data in real time, use the ASR model to perform speech-to-text conversion, identify keywords and phrases, and compare them with emotional tendency tables to obtain emotional tendency results for media content push.

Benefits of technology

It realizes fast and accurate identification and personalized push of media content, improves user experience and work efficiency, and is suitable for intelligent and personalized human-computer interaction in multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559936B_ABST
    Figure CN119559936B_ABST
Patent Text Reader

Abstract

The present invention proposes an early media identification method and system based on ASR. The ASR-based early media identification method includes: collecting audio data corresponding to the target media in real time, repairing the audio data by comparison, and obtaining the repaired audio data; inputting the repaired audio data into the ASR model to obtain text data corresponding to the audio data; performing keyword and phrase recognition from the text data to obtain target keywords and phrases; comparing and identifying the target keywords and phrases with preset keywords and phrases and sentiment tendency tables to obtain sentiment tendency results, and the sentiment tendency results are early media identification results; and pushing media content according to the sentiment tendency results. The system includes modules corresponding to the steps of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention proposes an early media recognition method and system based on ASR, belonging to the technical field of audio recognition. Background Art

[0002] In today's era of digitalization and information explosion, speech recognition (Automatic Speech Recognition, ASR) technology has become an important part of the field of human-computer interaction. The core task of ASR technology is to convert human speech into text. This process involves multiple key components such as acoustic models, language models, and decoders. The acoustic model is responsible for converting speech signals into corresponding acoustic feature sequences, that is, identifying phonemes or syllables in speech; the language model is trained based on a large amount of text data to evaluate whether the generated text sequence conforms to the grammatical rules and idioms of the language; the decoder is responsible for finding the optimal speech-to-text conversion result based on these two models. ASR technology greatly simplifies the process of information input. Compared with traditional manual input methods (such as keyboard input), voice input is more convenient and natural, which can greatly save time and energy. Especially on some mobile devices, voice input allows users to easily interact with the device even when their hands are occupied. In addition, ASR technology can achieve real-time interaction. Users do not need to wait and can get system responses at the moment of speaking. This immediacy greatly improves user experience and work efficiency. However, ASR technology also faces some challenges in practical applications. Factors such as environmental noise, accent differences, and complex language structures may affect the accuracy of recognition. Summary of the invention

[0003] The present invention provides an early media recognition method and system based on ASR to solve the technical problems in the above-mentioned prior art. The technical solutions adopted are as follows:

[0004] An early media identification method based on ASR, the early media identification method based on ASR comprising:

[0005] Collecting audio data corresponding to the target media in real time, repairing the audio data by comparing it, and obtaining the repaired audio data;

[0006] Inputting the repaired audio data into the ASR model to obtain text data corresponding to the audio data;

[0007] Perform keyword and phrase recognition from the text data to obtain target keywords and phrases;

[0008] Compare and identify the target keywords and phrases with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are early media recognition results;

[0009] The media content is pushed according to the emotional tendency result.

[0010] Furthermore, audio data corresponding to the target media is collected in real time, the audio data is repaired by comparison, and the repaired audio data is obtained, including:

[0011] Collect audio data corresponding to the target media in real time;

[0012] Performing enhancement processing on the audio data to obtain enhanced audio data, wherein the enhancement processing includes noise reduction and de-pop processing and volume normalization processing;

[0013] Performing audio quality evaluation on the enhanced audio data to obtain an audio quality result;

[0014] Perform audio repair on audio data whose audio quality result does not meet the audio quality requirement, and obtain repaired audio data.

[0015] Furthermore, the audio quality of the enhanced audio data is evaluated to obtain an audio quality result, including:

[0016] Extracting audio parameters of the enhanced audio data, wherein the audio parameters include signal-to-noise ratio, total harmonic distortion and short-time energy;

[0017] Acquire audio quality assessment parameters using the signal-to-noise ratio, total harmonic distortion and short-time energy included in the audio parameters;

[0018] The audio quality evaluation parameter is obtained by the following formula:

[0019]

[0020] Where P represents the audio quality assessment parameter; n represents the number of time windows corresponding to the short-time energy contained in the audio data; S i represents the signal-to-noise ratio corresponding to the i-th time window; T i represents the total harmonic distortion corresponding to the i-th time window; S max and S min W represents the maximum and minimum signal-to-noise ratio of n time windows corresponding to the audio data; max and W min W represents the maximum and minimum short-term energy of n time windows corresponding to the audio data; i represents the short-time energy corresponding to the i-th time window; W birepresents the short-time energy standard deviation corresponding to the i-th time window;

[0021] Comparing the audio quality assessment parameter with a preset audio quality assessment parameter threshold;

[0022] When the audio quality evaluation parameter is lower than a preset audio quality evaluation parameter threshold, it is determined that the audio quality result does not meet the audio quality requirement of the audio data.

[0023] Furthermore, setting a time window corresponding to the short-time energy contained in the audio parameters of the audio data includes:

[0024] Extract the sampling frequency, bit depth and number of channels corresponding to the audio data collection device;

[0025] Obtaining a first window adjustment coefficient using a sampling frequency, a bit depth, and a number of channels corresponding to the audio data collection device;

[0026] The first window adjustment coefficient is obtained by the following formula:

[0027]

[0028] Among them, K 01 represents the first window adjustment coefficient; f represents the sampling frequency corresponding to the audio data collection device; f c Indicates the preset sampling frequency reference value; N indicates the number of channels; B indicates the bit depth; int() indicates rounding up the value in the brackets;

[0029] Comparing the first window adjustment coefficient with a preset window adjustment coefficient threshold;

[0030] When the first window adjustment coefficient is lower than a preset window adjustment coefficient threshold, the preset initial time length is used as the time window corresponding length to set the time window corresponding to the short-time energy;

[0031] When the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, the initial time length is adjusted, and the time window corresponding to the short-time energy is set using the adjusted time length as the time window corresponding length.

[0032] Further, when the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the time window corresponding length to set the time window corresponding to the short-time energy, including:

[0033] When the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, extracting the initial time length and initial window overlap ratio values ​​corresponding to the time window of the ASR model; wherein the ASR model adopts a deep neural network model structure;

[0034] Extract the time length value and the corresponding window overlap ratio value that appear in the dynamic adjustment process of the time window of the ASR model;

[0035] The second window adjustment coefficient is obtained by using the initial time length and the initial window overlap ratio value corresponding to the time window of the ASR model in combination with the time length value and the corresponding window overlap ratio value appearing in the dynamic adjustment process of the time window;

[0036] The second window adjustment coefficient is obtained by the following formula:

[0037]

[0038] Among them, K 02 represents the second window adjustment coefficient; m represents the number of dynamic adjustments of the time window; T i represents the time length corresponding to the time window after the i-th dynamic adjustment; P i represents the window overlap ratio after the i-th dynamic adjustment; P ti represents the time length change rate between the time window after the i-th dynamic adjustment and the time window before the dynamic adjustment; T c and P c Indicates the initial time length and initial window overlap ratio; P z Represents the middle value of the window overlap ratio after m dynamic adjustments; P b represents the standard deviation of the window overlap ratio after m dynamic adjustments; T b represents the standard deviation of the time length corresponding to the time window after m dynamic adjustments; P max Indicates the maximum value of the window overlap ratio after m dynamic adjustments;

[0039] The initial time length is adjusted using the first window adjustment coefficient and the second window adjustment coefficient to obtain an adjusted time length;

[0040] The adjusted time length is obtained by the following formula:

[0041]

[0042] Among them, T x Indicates the adjusted time length; T 0 Indicates the length of time before adjustment; K 01 Represents the first window adjustment coefficient; K02 represents the second window adjustment coefficient; t represents the adjustment factor, and the value range of the adjustment factor is 1.12-1.43;

[0043] The time window corresponding to the short-time energy is set by using the adjusted time length as the corresponding length of the time window.

[0044] Further, performing audio repair on the audio data whose audio quality result does not meet the audio quality requirement to obtain the repaired audio data includes:

[0045] Extracting audio data whose audio quality results do not meet the audio quality requirements as audio data to be processed;

[0046] Using a filter to perform noise removal processing on the audio data to be processed, and obtaining the audio data after the noise removal processing;

[0047] Performing a distortion type analysis on the audio data after the noise removal process to obtain the distortion type of the audio data after the noise removal process;

[0048] Retrieving a distortion processing method corresponding to the distortion type from a database;

[0049] Using the distortion processing method corresponding to the distortion type, compensating and repairing the audio signal corresponding to each distortion type in the audio data after noise removal processing, to obtain the audio data after distortion repair;

[0050] Determining missing data in the distortion-repaired audio data, and obtaining an index position corresponding to the missing data;

[0051] Extract the audio signal data values ​​of the preceding and following data points corresponding to the index position;

[0052] Using an interpolation algorithm (such as linear interpolation or spline interpolation) to combine the audio signal data values ​​of the preceding and following data points corresponding to the index position to obtain the filling value corresponding to the missing data;

[0053] Each missing data is filled with a filling value corresponding to the missing data to obtain audio data after the missing value is repaired.

[0054] Further, the repaired audio data is input into an ASR model to obtain text data corresponding to the audio data, including:

[0055] Extract the short-time energy average value of the audio data after the current missing value is repaired;

[0056] Comparing the short-time energy average value with a preset short-time energy threshold;

[0057] When the short-time energy average value is lower than the preset short-time energy threshold, the model time window of the ASR model is not adjusted;

[0058] When the short-time energy average value is not lower than the preset short-time energy threshold, the model time window and the window overlap ratio value of the ASR model are adjusted to obtain the adjusted model time window length and the window overlap ratio value;

[0059] The adjusted model time window length is obtained by the following formula:

[0060]

[0061] Among them, T s represents the adjusted model time window length; T s0 represents the adjusted model time window length; n represents the number of time windows corresponding to the short-time energy contained in the audio data; W i represents the short-time energy corresponding to the i-th time window; W max and W min W represents the maximum and minimum short-term energy of n time windows corresponding to the audio data; b represents the short-time energy standard deviation corresponding to n time windows; E m Indicates the minimum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy maximum value is located; E n W represents the maximum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy minimum value is located; p Represents the short-time energy average; W y Indicates the preset short-time energy threshold;

[0062] At the same time, the adjusted window overlap ratio value is obtained by the following formula:

[0063]

[0064] Among them, B c Indicates the adjusted window overlap ratio value; B c0 Indicates the window overlap ratio value before adjustment; T s represents the adjusted model time window length; T s0 represents the adjusted model time window length; E m Indicates the minimum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy maximum value is located; E n Indicates the maximum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy minimum value is located;

[0065] Adjust the ASR model according to the adjusted model time window length and window overlap ratio values;

[0066] Inputting the repaired audio data into an ASR model, processing the repaired audio data through the ASR model, and obtaining a text string corresponding to the repaired audio data;

[0067] The text string is processed to obtain processed text data; wherein the data processing includes but is not limited to removing redundant spaces, processing punctuation marks, unifying fonts and font sizes, and correcting spelling errors.

[0068] Further, performing keyword and phrase recognition from the text data to obtain target keywords and phrases includes:

[0069] Recall a preset list of keywords and phrases;

[0070] Traversing the text data, and using a string matching method to find out whether the key words and phrases exist in the text data;

[0071] When keywords and phrases are found in text data, the location of the keywords and phrases in the text and their context are recorded and marked;

[0072] The found keywords and phrases are used as target keywords and phrases.

[0073] Furthermore, the target keywords and phrases are compared and identified with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, including:

[0074] Retrieve a comparison table of keywords and phrases with unified structure and sentiment tendency;

[0075] Remove the empty values ​​and duplicate items in the comparison table of the keywords and phrases and sentiment orientation, and obtain a processed comparison table of the keywords and phrases and sentiment orientation;

[0076] forming a target keyword and phrase list with the target keywords and phrases;

[0077] Traversing the target keyword and phrase list;

[0078] During the traversal process, the string matching method searches whether the keywords and phrases in the target keyword and phrase list appear in the comparison table of processed keywords and phrases and sentiment orientation;

[0079] When the keywords and phrases in the target keyword and phrase list appear in the comparison table of processed keywords and phrases and sentiment orientation, the sentiment orientation corresponding to the keywords and phrases is retrieved to obtain the sentiment orientation result.

[0080] An ASR-based early media recognition system, the ASR-based early media recognition system comprising:

[0081] The audio data collection and processing module is used to collect the audio data corresponding to the target media in real time, repair the audio data by comparing it, and obtain the repaired audio data;

[0082] A text conversion module, used for inputting the repaired audio data into an ASR model to obtain text data corresponding to the audio data;

[0083] A key information identification module, used to identify keywords and phrases from the text data to obtain target keywords and phrases;

[0084] The sentiment tendency recognition module is used to compare and recognize the target keywords and phrases with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are early media recognition results;

[0085] The content push control module is used to push the media content according to the emotional tendency result.

[0086] Beneficial effects of the present invention:

[0087] The present invention proposes an ASR-based early media recognition method and system that converts audio data into text data in real time through ASR technology, which greatly shortens the recognition time and improves the recognition efficiency. The acoustic model and language model of the ASR model, as well as natural language processing technology, can accurately identify keywords and phrases in the audio data, as well as the corresponding emotional tendencies. According to the emotional tendency results, the media content can be personalized screened and classified to meet the personalized needs of users. This method is not only applicable to media content such as news broadcasts and social media videos, but can also be expanded to multiple fields such as smart homes, medical care, and finance to achieve more intelligent and personalized human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 is a flow chart of the method of the present invention;

[0089] Figure 2 The system block diagram of the system of the present invention. DETAILED DESCRIPTION

[0090] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0091] An ASR-based early media recognition method is proposed in an embodiment of the present invention. Figure 1 As shown, the ASR-based early media recognition method includes:

[0092] S1. Collect audio data corresponding to the target media in real time, repair the audio data by comparing it, and obtain the repaired audio data;

[0093] S2, inputting the repaired audio data into the ASR model to obtain text data corresponding to the audio data;

[0094] S3, identifying keywords and phrases from the text data to obtain target keywords and phrases;

[0095] S4, comparing and identifying the target keywords and phrases with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are early media recognition results;

[0096] S5. Pushing media content according to the emotional tendency result.

[0097] The working principle of the above technical solution is as follows: collect audio data corresponding to the target media (such as news broadcasts, social media videos, etc.) in real time. Preprocess the collected audio data, including noise reduction, frame segmentation, etc., to improve the audio quality. Compare the audio data and perform necessary repairs, such as removing noise, filling silent segments, etc., to obtain the repaired audio data. Input the repaired audio data into the ASR model and use the acoustic model and language model of the ASR model to convert the audio data into corresponding text data. The acoustic model is responsible for converting the speech signal into the corresponding acoustic feature sequence, and the language model is trained based on a large amount of text data to evaluate whether the generated text sequence conforms to the grammatical rules and idiomatic usage of the language. From the converted text data, natural language processing technology is used to identify keywords and phrases. Extract target keywords and phrases with representative and information value. Compare the extracted target keywords and phrases with the preset keywords and phrases and sentiment tendency table. According to the comparison results, obtain the sentiment tendency results corresponding to the audio data, such as positive, negative or neutral. This sentiment tendency result is the early media recognition result, which is used for subsequent content push decisions. According to the results of emotional tendency, media content is screened and classified, and media content that meets user preferences or needs is pushed to users.

[0098] The effect of the above technical solution is: through ASR technology, audio data is converted into text data in real time, which greatly shortens the recognition time and improves the recognition efficiency. Using the acoustic model and language model of the ASR model, as well as natural language processing technology, it is possible to accurately identify keywords and phrases in the audio data, as well as the corresponding emotional tendencies. According to the emotional tendency results, media content can be personalized and classified to meet the personalized needs of users. This method is not only applicable to media content such as news broadcasts and social media videos, but can also be extended to smart homes, medical care, finance and other fields to achieve more intelligent and personalized human-computer interaction.

[0099] In summary, the ASR-based early media recognition method achieves fast, accurate recognition and personalized push of media content through real-time collection and repair of audio data, speech recognition, keyword and phrase recognition, sentiment tendency recognition, and content push. It has broad application prospects and important technical value.

[0100] In one embodiment of the present invention, audio data corresponding to target media is collected in real time, the audio data is repaired by comparison, and the repaired audio data is obtained, including:

[0101] S101, collecting audio data corresponding to the target media in real time;

[0102] S102, performing enhancement processing on the audio data to obtain enhanced audio data, wherein the enhancement processing includes noise reduction and de-pop processing and volume normalization processing;

[0103] S103, performing audio quality assessment on the enhanced audio data to obtain an audio quality result;

[0104] S104: Perform audio repair on the audio data whose audio quality result does not meet the audio quality requirement, and obtain repaired audio data.

[0105] The working principle of the above technical solution is as follows: This step is the starting point of the entire process, which involves real-time capture of audio data from the target media (such as microphones, recording devices, network streaming media, etc.). This requires the system to have efficient audio data acquisition capabilities to ensure the real-time and integrity of the audio data. Noise reduction and de-pop processing: By analyzing the noise components and pop (that is, the sudden appearance of extremely loud volume) characteristics in the audio signal, the corresponding filtering or elimination algorithm is adopted to reduce or eliminate these adverse effects, thereby improving the clarity of the audio. Adjust the volume of the audio signal to a relatively stable range to ensure that the audio playback volume is consistent and avoid affecting the auditory experience due to sudden changes in volume. This step usually involves a series of objective and subjective audio quality evaluation indicators. Objective indicators such as signal-to-noise ratio (SNR) and distortion are used to quantify the purity and distortion of the audio signal; subjective indicators evaluate the auditory quality of the audio by listening to audio samples and giving subjective scores. By combining these indicators, the system can comprehensively evaluate the audio quality and determine whether further repair processing is needed. Audio data whose audio quality results do not meet the audio quality requirements will be repaired. For audio data whose audio quality assessment results do not meet the requirements, the system will perform targeted repair processing. This may include using more advanced noise reduction algorithms, audio synthesis technology, or audio compensation algorithms to restore lost or damaged audio information, thereby improving the overall audio quality.

[0106] The effect of the above technical solution is: through enhanced processing such as noise reduction, de-pop, volume normalization and targeted repair processing, this technical solution can significantly improve the clarity, purity and auditory experience of the audio. The ability to collect audio data in real time ensures the real-time and integrity of the audio data, which is especially important for audio application scenarios that require real-time processing (such as real-time communications, online meetings, etc.). This technical solution realizes the automation and intelligence of audio processing by integrating processing steps such as audio enhancement, quality assessment and repair, reducing the cost and time of manual intervention. By improving audio quality, this technical solution can also enhance the reliability and stability of audio applications, and reduce problems such as communication interruptions or reduced experience due to audio quality problems.

[0107] In summary, this technical solution realizes real-time collection, enhancement, quality assessment and repair of audio data through a series of efficient audio processing steps, thereby significantly improving the overall audio quality and listening experience.

[0108] In one embodiment of the present invention, performing audio quality assessment on enhanced audio data to obtain an audio quality result includes:

[0109] S1031. Extracting audio parameters of the enhanced audio data, wherein the audio parameters include signal-to-noise ratio, total harmonic distortion, and short-time energy;

[0110] S1032, obtaining audio quality assessment parameters using the signal-to-noise ratio, total harmonic distortion and short-time energy included in the audio parameters;

[0111] The audio quality evaluation parameter is obtained by the following formula:

[0112]

[0113] Where P represents the audio quality assessment parameter; n represents the number of time windows corresponding to the short-time energy contained in the audio data; S i represents the signal-to-noise ratio corresponding to the i-th time window; T i represents the total harmonic distortion corresponding to the i-th time window; S max and S min W represents the maximum and minimum signal-to-noise ratio of n time windows corresponding to the audio data; max and W min W represents the maximum and minimum short-term energy of n time windows corresponding to the audio data; i represents the short-time energy corresponding to the i-th time window; W bi represents the short-time energy standard deviation corresponding to the i-th time window;

[0114] S1033. Compare the audio quality assessment parameter with a preset audio quality assessment parameter threshold;

[0115] S1034: When the audio quality evaluation parameter is lower than a preset audio quality evaluation parameter threshold, determine that the audio quality result does not meet the audio quality requirement of the audio data.

[0116] The working principle of the above technical solution is: extract key audio parameters from the enhanced audio data, including signal-to-noise ratio (SNR), total harmonic distortion (THD) and short-time energy. These parameters can reflect the quality characteristics of the audio signal, such as purity, distortion level and energy distribution. Using the extracted audio parameters, the audio quality evaluation parameter P is calculated through a specific formula. The formula takes into account multiple aspects of signal-to-noise ratio, total harmonic distortion and short-time energy, including their values ​​in different time windows, maximum value, minimum value, standard deviation, etc. The n in the formula represents the number of time windows corresponding to the short-time energy contained in the audio data, S i , T i , W i They represent the signal-to-noise ratio, total harmonic distortion and short-time energy corresponding to the i-th time window. max , S min , W max , W minThen they represent the maximum and minimum values ​​of the signal-to-noise ratio and the maximum and minimum values ​​of the short-time energy of the n time windows corresponding to the audio data. bi Represents the short-time energy standard deviation corresponding to the i-th time window. Through this comprehensive calculation method, an evaluation parameter P that can fully reflect the audio quality can be obtained. The calculated audio quality evaluation parameter P is compared with the preset audio quality evaluation parameter threshold. If P is lower than the threshold, it is determined that the audio quality result does not meet the audio quality requirements and further repair processing is required.

[0117] The effect of the above technical solution is: by extracting multiple key audio parameters and comprehensively calculating the audio quality evaluation parameter P, the technical solution can comprehensively and objectively evaluate the quality of the audio signal. This helps to accurately identify problems in the audio signal, such as excessive noise, severe distortion, etc., and provides strong support for subsequent repair processing. The technical solution reduces the cost and time of manual intervention through an automated evaluation process. At the same time, since the evaluation results are objective and accurate, the processing inconsistency caused by subjective judgment differences can be avoided, and the efficiency and stability of audio processing can be improved. By comparing the audio quality evaluation parameter P with the preset threshold, the technical solution can accurately determine which audio data needs to be repaired. This helps to formulate targeted repair strategies and optimize the repair effect, while avoiding unnecessary processing of audio data that does not need to be repaired, saving resources. The technical solution can provide users with a higher quality audio experience by improving the accuracy and efficiency of audio quality evaluation. This is particularly important in application scenarios such as audio communication, online conferencing, and music playback, and helps to improve user satisfaction and loyalty.

[0118] On the other hand, the technical solution can more comprehensively reflect the quality characteristics of the audio signal by extracting multiple key audio parameters such as signal-to-noise ratio, total harmonic distortion, short-time energy, etc., and calculating the audio quality evaluation parameters based on these factors. This comprehensive evaluation method is more accurate than single parameter evaluation, and can reduce the misjudgment caused by the abnormality of a single parameter. The calculation process of the evaluation parameters is based on objective mathematical formulas and statistical methods, avoiding the uncertainty caused by subjective judgment. This makes the evaluation results more objective and accurate, and can provide a reliable basis for subsequent repair processing. The technical solution realizes the automation process of audio quality evaluation and reduces the cost and time of manual intervention. Automated evaluation can significantly improve processing efficiency, especially in large-scale audio data processing scenarios, which can greatly shorten the evaluation cycle. By extracting audio parameters in real time and calculating evaluation parameters, the technical solution can achieve a rapid response to audio quality. This helps to timely discover and deal with problems in audio signals and improve the real-time and effectiveness of audio processing. At the same time, the technical solution can be applied to a variety of audio scenarios, such as music, speech, noise, etc., and has a wide range of applicability. By adjusting the calculation formula and threshold setting of the evaluation parameters, it can adapt to the quality evaluation needs in different audio scenarios. In the process of extracting audio parameters and calculating evaluation parameters, this technical solution can effectively suppress the influence of noise and interference signals. This makes the evaluation results more stable and reliable, and can maintain a high evaluation accuracy even in complex and changing audio environments. By comparing the audio quality evaluation parameters with the preset thresholds, the technical solution can accurately determine which audio data needs to be repaired. This helps to formulate targeted repair strategies, optimize the repair effect, and improve the overall quality of audio data. Accurate audio quality evaluation and effective repair processing can enhance the user's listening experience. In application scenarios such as audio communication, online conferencing, and music playback, high-quality audio signals can enhance user participation and satisfaction.

[0119] In summary, the technical effects of this technical solution in terms of performance indicators are mainly reflected in the aspects of evaluation accuracy, processing efficiency, robustness, and optimization effect. These effects jointly improve the reliability and effectiveness of audio quality evaluation, and provide strong support for subsequent audio processing and applications. At the same time, this technical solution realizes comprehensive evaluation and optimization processing of enhanced audio data by extracting key audio parameters, calculating audio quality evaluation parameters, and comparing and judging. This helps to improve the efficiency and stability of audio processing, optimize audio repair strategies, and enhance user experience.

[0120] In one embodiment of the present invention, setting a time window corresponding to the short-time energy contained in the audio parameters of the audio data includes:

[0121] Step 1, extracting the sampling frequency, bit depth and number of channels corresponding to the audio data collection device;

[0122] Step 2: Obtain a first window adjustment coefficient using the sampling frequency, bit depth and number of channels corresponding to the audio data collection device;

[0123] The first window adjustment coefficient is obtained by the following formula:

[0124]

[0125] Among them, K 01 represents the first window adjustment coefficient; f represents the sampling frequency corresponding to the audio data collection device; f c Indicates the preset sampling frequency reference value; N indicates the number of channels; B indicates the bit depth; int() indicates rounding up the value in the brackets;

[0126] Step 3: comparing the first window adjustment coefficient with a preset window adjustment coefficient threshold;

[0127] Step 4: When the first window adjustment coefficient is lower than a preset window adjustment coefficient threshold, the preset initial time length is used as the time window corresponding length to set the time window corresponding to the short-time energy;

[0128] Step 5: When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length is adjusted, and the adjusted time length is used as the corresponding length of the time window to set the time window corresponding to the short-time energy.

[0129] The working principle of the above technical solution is as follows: first, extract the three key parameters of sampling frequency, bit depth and number of channels from the audio data collection device (such as a microphone, a voice recorder, etc.). These parameters directly determine the quality and characteristics of the audio data. The first window adjustment coefficient is calculated by a specific formula using the extracted sampling frequency, bit depth and number of channels. This coefficient reflects the impact of the performance of the audio data collection device on the time window setting. The sampling frequency reference value in the formula is a preset value, which is used to compare with the actual sampling frequency to adjust the size of the window adjustment coefficient. The number of channels and bit depth are also included in the formula as influencing factors. The calculated first window adjustment coefficient is compared with the preset window adjustment coefficient threshold. This threshold is used to determine whether the initial time length needs to be adjusted to adapt to different audio data collection devices and audio data quality. If the first window adjustment coefficient is lower than the threshold, it means that the current audio data collection device has low performance or the audio data quality is poor. At this time, the preset initial time length is used as the length of the time window. If the first window adjustment coefficient is not lower than the threshold, it means that the current audio data collection device has high performance or the audio data quality is good. At this time, the initial time length needs to be adjusted to better adapt to the characteristics of the audio data. The adjusted time length will be used as the length of the time window for the calculation and analysis of short-time energy.

[0130] The effect of the above technical solution is: by setting the length of the time window according to the performance of the audio data collection device and the characteristics of the audio data, short-time energy can be extracted and analyzed more effectively. This helps to more accurately reflect the amplitude changes and characteristics of the audio signal. The technical solution can automatically adjust the length of the time window according to the device parameters and the quality of the audio data, thereby avoiding the tediousness and uncertainty of manual setting. This improves the efficiency and automation of audio data processing. The technical solution can be applied to audio data collection devices of different performance and quality and audio data of different characteristics. By adjusting the length of the time window, it can be ensured that short-time energy analysis can obtain accurate results in different scenarios. Accurate short-time energy analysis helps to improve the processing effect and application performance of audio data. In the fields of audio communication, speech recognition, audio editing, etc., this can enhance the user's auditory experience and satisfaction.

[0131] On the other hand, by dynamically adjusting the length of the time window according to parameters such as the sampling frequency, bit depth and number of channels of the audio data collection device, the short-time energy characteristics of the audio signal can be captured more effectively. Compared with the fixed time window setting, this dynamic adjustment method can more accurately reflect the amplitude changes and characteristics of the audio signal, thereby improving the accuracy of short-time energy analysis. By accurately calculating the first window adjustment coefficient and comparing it with the preset window adjustment coefficient threshold, the length of the time window can be automatically selected or adjusted to reduce the errors and interference introduced by device performance differences or audio data quality fluctuations. The technical solution realizes the automated processing flow of time window setting, reducing the tediousness and uncertainty of manual setting. Automated processing can significantly improve the efficiency of audio data processing, especially in large-scale audio data processing scenarios, which can greatly shorten the processing cycle. By reasonably setting the length of the time window, it is possible to reduce unnecessary computing resource consumption while ensuring the accuracy of the analysis. This helps to reduce the cost of audio data processing and improve the overall performance and efficiency of the system. The technical solution can be applied to audio data collection devices with different performance and quality and audio data with different characteristics. By dynamically adjusting the length of the time window, it can ensure that short-time energy analysis can obtain accurate results in different scenarios, thereby enhancing the wide applicability of the system. The formulas and thresholds in this technical solution can be adjusted and optimized according to actual needs to adapt to different application scenarios and changes in demand. This flexibility and scalability help maintain the advancement and competitiveness of the system. Accurate short-time energy analysis helps improve the processing effect and application performance of audio data. In the fields of audio communication, speech recognition, audio editing, etc., this can improve the clarity and quality of audio signals, thereby improving the user's auditory experience. The automated processing flow reduces the user's operating burden, allowing users to use the audio data processing system more conveniently. This helps to improve user satisfaction and loyalty.

[0132] In summary, the technical effects of this technical solution in terms of performance indicators are mainly reflected in improved accuracy, optimized efficiency, enhanced adaptability, and improved user experience. These effects jointly improve the performance and effect of audio data processing, and provide strong support for subsequent audio analysis and applications. At the same time, this technical solution improves the accuracy and efficiency of short-term energy analysis and enhances the adaptability and user experience of audio data processing by setting the length of the time window according to the performance of the audio data collection device and the characteristics of the audio data.

[0133] In one embodiment of the present invention, when the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, the initial time length is adjusted, and the time window corresponding to the short-time energy is set using the adjusted time length as the time window corresponding length, including:

[0134] Step 501: when the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, extracting an initial time length and an initial window overlap ratio value corresponding to a time window of an ASR model; wherein the ASR model adopts a deep neural network model structure;

[0135] Step 502: extracting the time length value and the corresponding window overlap ratio value appearing in the dynamic adjustment process of the time window of the ASR model;

[0136] Step 503: Obtain a second window adjustment coefficient by using the initial time length and initial window overlap ratio values ​​corresponding to the time window of the ASR model in combination with the time length values ​​and the corresponding window overlap ratio values ​​occurring during the dynamic adjustment of the time window;

[0137] The second window adjustment coefficient is obtained by the following formula:

[0138]

[0139] Among them, K 02 represents the second window adjustment coefficient; m represents the number of dynamic adjustments of the time window; T i represents the time length corresponding to the time window after the i-th dynamic adjustment; P i represents the window overlap ratio after the i-th dynamic adjustment; P ti represents the time length change rate between the time window after the i-th dynamic adjustment and the time window before the dynamic adjustment; T c and P c Indicates the initial time length and initial window overlap ratio; P z Represents the middle value of the window overlap ratio after m dynamic adjustments; P b represents the standard deviation of the window overlap ratio after m dynamic adjustments; T b represents the standard deviation of the time length corresponding to the time window after m dynamic adjustments; P max Indicates the maximum value of the window overlap ratio after m dynamic adjustments;

[0140] Step 504: Use the first window adjustment coefficient and the second window adjustment coefficient to adjust the initial time length to obtain an adjusted time length;

[0141] The adjusted time length is obtained by the following formula:

[0142]

[0143] Among them, T x Indicates the adjusted time length; T 0 Indicates the length of time before adjustment; K01 Represents the first window adjustment coefficient; K 02 represents the second window adjustment coefficient; t represents the adjustment factor, and the value range of the adjustment factor is 1.12-1.43;

[0144] Step 505: Use the adjusted time length as the corresponding length of the time window to set the time window corresponding to the short-time energy.

[0145] The working principle of the above technical solution is: when the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the system first extracts the initial time length and the initial window overlap ratio value corresponding to the time window of the ASR model. These initial values ​​are used as the basis for subsequent dynamic adjustment. The system records the time length and the corresponding window overlap ratio value of each adjustment of the ASR model time window during the dynamic adjustment process. These records are used to calculate the second window adjustment coefficient. Using the dynamic adjustment data recorded above, combined with the initial time length and the initial window overlap ratio value, the second window adjustment coefficient is calculated by a specific formula. This coefficient reflects the overall changes and trends in the dynamic adjustment process of the time window. Combined with the first window adjustment coefficient and the second window adjustment coefficient, as well as a preset adjustment factor, the initial time length is adjusted by a specific formula to obtain the adjusted time length. This adjustment process takes into account the performance of the audio data collection device, the characteristics of the ASR model, and the historical data of the dynamic adjustment of the time window. Finally, the time window corresponding to the short-time energy is set using the adjusted time length as the length of the time window. The time window set in this way can more accurately reflect the short-time energy characteristics of the audio signal, thereby improving the performance and accuracy of the ASR model.

[0146] The effect of the above technical solution is: by dynamically adjusting the length of the time window, the technical solution can more accurately capture the short-time energy characteristics of the audio signal, thereby improving the recognition performance and accuracy of the ASR model. The technical solution can automatically adjust the length of the time window according to the performance of different audio data collection devices and the characteristics of the ASR model, thereby enhancing the adaptability and flexibility of the system. By reasonably setting the length and overlap ratio of the time window, the technical solution can optimize the utilization of computing resources while ensuring the performance of the ASR model, reducing the power consumption and cost of the system. Accurate ASR model performance and optimized resource utilization can enhance the user's auditory experience and satisfaction, especially in application scenarios that require efficient and accurate speech recognition.

[0147] By dynamically adjusting the length of the time window, the technical solution can more accurately capture the short-term energy characteristics of the audio signal, thereby improving the recognition accuracy of the automatic speech recognition (ASR) model. This helps to reduce recognition errors and misjudgments and improve the overall performance of the system. The optimized time window setting can reduce unnecessary calculation and processing time, thereby improving the response speed of the ASR model. This enables the system to complete the recognition task faster and improve processing efficiency when processing a large amount of audio data. By reasonably setting the length and overlap ratio of the time window, the technical solution can optimize the utilization of computing resources while ensuring the performance of the ASR model. This helps to reduce the power consumption and cost of the system and improve the efficiency of resource utilization. The optimized time window setting can reduce the storage requirements of audio data, thereby saving storage space. This is particularly important for systems that need to process a large amount of audio data, which can reduce storage costs and improve the scalability of the system. Dynamically adjusting the length of the time window helps to enhance the ASR model's ability to resist noise. In a noisy environment, the technical solution can more effectively extract the short-term energy characteristics of the audio signal, thereby improving the stability and accuracy of recognition. By recording and analyzing the dynamic adjustment process of the time window, the technical solution can promptly discover and correct potential errors and deviations. This helps improve the system's fault tolerance and robustness, ensuring that the system can operate stably under various circumstances. Optimized ASR model performance and response speed can improve the interactivity between users and the system. Users can interact with the system more quickly and accurately, thereby improving the overall user experience. Accurate recognition performance and optimized resource utilization can improve user satisfaction and trust. This helps to enhance user confidence and dependence on the system, and improve the system's user stickiness and market competitiveness.

[0148] In summary, the technical effects of this technical solution in terms of performance indicators are mainly reflected in the improvement of recognition performance, optimization of resource utilization efficiency, enhancement of system stability and reliability, and improvement of user experience. These effects jointly improve the overall performance and user experience of the system, and provide strong support for subsequent audio analysis and applications. At the same time, this technical solution improves the performance and accuracy of the ASR model by dynamically adjusting the length of the time window, enhances the adaptability and flexibility of the system, optimizes the utilization of computing resources, and improves the user experience.

[0149] In one embodiment of the present invention, audio data whose audio quality result does not meet the audio quality requirement is repaired to obtain the repaired audio data, including:

[0150] S1041, extracting audio data whose audio quality results do not meet the audio quality requirements as audio data to be processed;

[0151] S1042, using a filter to perform noise removal processing on the audio data to be processed, to obtain audio data after the noise removal processing;

[0152] S1043, performing a distortion type analysis on the audio data after the noise removal process to obtain the distortion type of the audio data after the noise removal process;

[0153] S1044, retrieving a distortion processing method corresponding to the distortion type from a database;

[0154] S1045, using the distortion processing method corresponding to the distortion type to compensate and repair the audio signal corresponding to each distortion type in the audio data after noise removal processing, to obtain distortion-repaired audio data;

[0155] S1046, determining missing data in the distortion-repaired audio data, and obtaining an index position corresponding to the missing data;

[0156] S1047, extracting audio signal data values ​​of the preceding and following data points corresponding to the index position;

[0157] S1048, using an interpolation algorithm (such as linear interpolation, spline interpolation) in combination with the audio signal data values ​​of the preceding and following data points corresponding to the index position to obtain a filling value corresponding to the missing data;

[0158] S1049: Fill each missing data with the filling value corresponding to the missing data to obtain audio data after the missing value is repaired.

[0159] The working principle of the above technical solution is: audio data that does not meet the quality requirements are screened out from the original audio data as the audio data to be processed. The audio data to be processed is processed by a filter to remove the noise components therein. The filter can be selected and set according to the frequency characteristics or statistical characteristics of the noise to effectively remove different types of noise. The audio data after noise removal is subjected to distortion type analysis. By analyzing the spectrum and time domain characteristics of the audio signal, the distortion type existing in the audio data can be determined, such as noise distortion, compression distortion, frequency response distortion, etc. According to the distortion type obtained by the analysis, the corresponding distortion processing method is retrieved from the database. A variety of distortion types and their corresponding processing methods can be stored in the database so as to quickly and accurately find the applicable repair solution. Each distortion type in the audio data is compensated and repaired using the retrieved distortion processing method. The repair method may include filter adjustment, dynamic processing, equalizer adjustment, etc. to restore the original quality of the audio signal. In the audio data after distortion repair, there may be missing data due to various reasons. By determining the index position of the missing data, extracting the audio signal data values ​​of the preceding and following data points corresponding to the index position, and then using an interpolation algorithm (such as linear interpolation, spline interpolation, etc.) to calculate the filling value corresponding to the missing data. Finally, each missing data is filled with the filling value to obtain the audio data after the missing value is repaired.

[0160] The effect of the above technical solution is: through noise removal processing and distortion compensation repair, the noise and distortion components in the audio data can be significantly reduced, thereby improving the overall quality of the audio. The distortion compensation repair process can perform targeted processing on different types of distortion, which helps to restore the original details and characteristics of the audio signal. By determining and filling in missing data, the blanks or missing parts in the audio data can be filled to improve the integrity and continuity of the audio. The repaired audio data is of higher quality and more complete, and can be used in more application scenarios, such as music production, speech recognition, audio analysis, etc. This technical solution adopts an automated processing flow, which can quickly and efficiently process large amounts of audio data, improve processing efficiency and reduce labor costs.

[0161] In summary, this technical solution has high application value in audio restoration and can effectively improve the quality and availability of audio data.

[0162] In one embodiment of the present invention, the repaired audio data is input into an ASR model to obtain text data corresponding to the audio data, including:

[0163] S201, extracting the short-time energy average value of the audio data after the current missing value is repaired;

[0164] S202, comparing the short-time energy average value with a preset short-time energy threshold;

[0165] S203: when the short-time energy average value is lower than a preset short-time energy threshold, the model time window of the ASR model is not adjusted;

[0166] S204: When the short-time energy average value is not lower than the preset short-time energy threshold, the model time window and the window overlap ratio value of the ASR model are adjusted to obtain the adjusted model time window length and the window overlap ratio value;

[0167] The adjusted model time window length is obtained by the following formula:

[0168]

[0169] Among them, T s represents the adjusted model time window length; T s0 represents the adjusted model time window length; n represents the number of time windows corresponding to the short-time energy contained in the audio data; W i represents the short-time energy corresponding to the i-th time window; W max and W min W represents the maximum and minimum short-term energy of n time windows corresponding to the audio data; b represents the short-time energy standard deviation corresponding to n time windows; E m Indicates the minimum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy maximum value is located; E n W represents the maximum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy minimum value is located; p Represents the short-time energy average; W y Indicates the preset short-time energy threshold;

[0170] At the same time, the adjusted window overlap ratio value is obtained by the following formula:

[0171]

[0172] Among them, B c Indicates the adjusted window overlap ratio value; B c0 Indicates the window overlap ratio value before adjustment; T s represents the adjusted model time window length; T s0 represents the adjusted model time window length; E m Indicates the minimum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy maximum value is located; E n Indicates the maximum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy minimum value is located;

[0173] S205, adjusting the ASR model according to the adjusted model time window length and window overlap ratio values;

[0174] S206, inputting the repaired audio data into an ASR model, processing the repaired audio data through the ASR model, and obtaining a text string corresponding to the repaired audio data;

[0175] S207, performing data processing on the text string to obtain processed text data; wherein the data processing includes but is not limited to removing redundant spaces, processing punctuation marks, unifying fonts and font sizes, and correcting spelling errors.

[0176] The working principle of the above technical solution is as follows: first, extract the short-time energy average value of the audio data after the current missing value is repaired. Short-time energy is an important feature in audio signal analysis, which reflects the energy change of the audio signal in a short time. Then, compare this short-time energy average value with the preset short-time energy threshold. This threshold is preset according to the performance of the ASR model and the characteristics of the audio data, and is used to determine whether the energy level of the audio data is high enough to support accurate speech recognition. If the short-time energy average value is lower than the preset threshold, it means that the energy of the audio data is low and may not be enough to support accurate speech recognition, so the model time window of the ASR model is not adjusted. If the short-time energy average value is not lower than the preset threshold, it means that the energy of the audio data is high enough, and the model time window and window overlap ratio of the ASR model can be adjusted. The adjusted model time window length is calculated based on a series of complex formulas that take into account factors such as the short-time energy distribution, maximum and minimum values, standard deviation, and minimum and maximum values ​​of the short-time energy of the time frame of the audio data. At the same time, the adjusted window overlap ratio value is also calculated according to a similar formula to ensure that while adjusting the time window length, the appropriate window overlap can be maintained to capture the detailed information in the audio signal. Finally, the ASR model is adjusted according to the adjusted model time window length and window overlap ratio value to adapt to the characteristics of the current audio data. The repaired audio data is input into the adjusted ASR model, and the audio data is processed by the ASR model to obtain the text string corresponding to the audio data. The obtained text string is processed, including removing redundant spaces, processing punctuation marks, unifying fonts and font sizes, and correcting spelling errors, so as to obtain high-quality text data.

[0177] The effect of the above technical solution is: by dynamically adjusting the model time window and window overlap ratio values ​​of the ASR model, it can be optimized according to the characteristics of the audio data, thereby improving the accuracy of speech recognition. The technical solution can be dynamically adjusted according to different audio data characteristics, enhancing the adaptability and flexibility of the ASR model. A series of data processing is performed on the acquired text strings to remove redundant information and correct errors, thereby optimizing the quality of the text data. High-quality text data can provide users with more accurate and clear information, improving user experience and satisfaction. The technical solution adopts an automated processing flow, which can quickly and efficiently process large amounts of audio data and improve processing efficiency.

[0178] On the other hand, by dynamically adjusting the model time window and window overlap ratio values ​​of the ASR model, the model can better adapt to the characteristics of different audio data. This adjustment can significantly improve the accuracy of speech recognition, so that the model can maintain a high recognition rate when processing audio data with different energy levels and different speech characteristics. The technical solution adopts an automated processing flow, including short-time energy analysis, ASR model adjustment, speech recognition and text processing. The automated processing flow can significantly improve processing efficiency and reduce the time cost of manual intervention. At the same time, due to the dynamic adjustment of the model time window and window overlap ratio values, the model is more efficient in processing audio data, thereby improving the overall processing speed. By accurately calculating the model time window length and window overlap ratio values, it is ensured that the ASR model can make full use of computing resources when processing audio data. The optimized model can reduce unnecessary computing overhead and improve resource utilization when processing audio data. This not only helps to reduce computing costs, but also helps to improve the stability and reliability of the system. The technical solution can be dynamically adjusted according to different audio data characteristics, so that the ASR model can maintain high performance when processing different types of audio data. This dynamic adjustment capability enhances the adaptability of the system, allowing the system to better cope with various complex scenarios and changes. This helps improve the robustness and reliability of the system and reduce the risk of performance degradation due to changes in audio data characteristics. High-quality text data is the key to improving user experience. This technical solution provides users with more accurate and clear information by optimizing speech recognition accuracy and text processing quality. The optimized system can provide users with higher-quality text data, thereby improving user experience and satisfaction. This helps to enhance users' trust and reliance on the system and improve the market competitiveness of the system.

[0179] In summary, the technical effects of the above technical solutions in terms of performance indicators are mainly reflected in the improvement of speech recognition accuracy, processing efficiency, resource utilization optimization, system adaptability enhancement, and user experience improvement. These technical effects work together on the entire system, making the system more efficient, accurate, and reliable when processing audio data. At the same time, the technical solution has high application value in speech recognition and text processing of audio data, and can effectively improve the accuracy of speech recognition and the quality of text data.

[0180] In one embodiment of the present invention, keyword and phrase recognition is performed from the text data to obtain target keywords and phrases, including:

[0181] S301, retrieve a preset keyword and phrase list;

[0182] S302, traversing the text data, and using a string matching method to find out whether there are keywords and phrases in the text data;

[0183] S303, when keywords and phrases are found in the text data, the positions of the keywords and phrases in the text and their contexts are recorded and marked;

[0184] S304: The found keywords and phrases are used as target keywords and phrases.

[0185] The working principle of the above technical solution is as follows: First, the system retrieves a list of keywords and phrases from a preset database. This list contains all the keywords and phrases that need to be found in the text data, which may be related to a specific topic, field or task. Next, the system traverses the entire text data and uses string matching methods (such as exact matching, fuzzy matching, regular expressions, etc.) to find whether the preset keywords and phrases exist in the text. This process may involve pre-processing steps such as word segmentation and stop word removal of text data to improve the accuracy and efficiency of the search. Record and mark the location of keywords and phrases and their context (S303): When the system finds keywords and phrases in the text data, it records and marks the specific locations of these keywords and phrases in the text (such as the starting position and the ending position), as well as the context information in which they are located. The context information may include the text content before and after the keywords and phrases, which is helpful for the subsequent understanding and analysis of the keywords and phrases. The found keywords and phrases are used as target keywords and phrases (S304): Finally, the system outputs or stores the found keywords and phrases as target keywords and phrases. These target keywords and phrases can be used for subsequent tasks such as information extraction, text classification, and sentiment analysis.

[0186] The effect of the above technical solution is that by using a preset list of keywords and phrases, combined with a string matching method, the system can accurately identify the target keywords and phrases in the text data. This helps to reduce the situation of misidentification and missed identification, and improves the accuracy of recognition. The technical solution uses a method of traversing text data for searching, and combines possible preprocessing steps (such as word segmentation, stop word removal, etc.), which can efficiently process a large amount of text data. At the same time, by recording the location of keywords and phrases and their contextual information, it provides convenience for subsequent processing tasks and further improves processing efficiency. The preset keyword and phrase list can be flexibly configured and updated according to actual needs. This means that the system can adapt to different application scenarios and task requirements, and has strong flexibility and scalability. By identifying target keywords and phrases, the system can provide valuable information for subsequent tasks such as information extraction, text classification, and sentiment analysis. This information helps to deeply understand the meaning and characteristics of text data, and provides strong support for applications such as decision support and data mining.

[0187] In summary, this technical solution has high application value in keyword and phrase recognition. It can accurately and efficiently identify target keywords and phrases in text data, providing strong support for subsequent processing tasks.

[0188] In one embodiment of the present invention, the target keywords and phrases are compared and identified with preset keywords and phrases and sentiment tendency tables to obtain sentiment tendency results, including:

[0189] S401, retrieve a comparison table of keywords and phrases with a unified structure and sentiment tendency;

[0190] S402, removing the empty values ​​and duplicate items in the comparison table of the keywords and phrases and sentiment orientation, and obtaining a processed comparison table of the keywords and phrases and sentiment orientation;

[0191] S403, forming a target keyword and phrase list with the target keywords and phrases;

[0192] S404, traversing the target keyword and phrase list;

[0193] S405, during the traversal process, the string matching method searches whether the keywords and phrases in the target keyword and phrase list appear in the comparison table of processed keywords and phrases and sentiment orientation;

[0194] S406. When the keywords and phrases in the target keyword and phrase list appear in the comparison table of processed keywords and phrases and sentiment orientation, the sentiment orientation corresponding to the keywords and phrases is retrieved to obtain the sentiment orientation result.

[0195] The working principle of the above technical solution is: obtain a comparison table of keywords and phrases with unified structure and sentiment tendency from a database or other storage location. This table is pre-set and contains multiple keywords and phrases and their corresponding sentiment tendencies (such as positive, negative, neutral, etc.). Preprocess this table to remove null values ​​and duplicates. Null values ​​may be invalid data, and duplicates will cause errors in the subsequent matching process, so they need to be cleaned in advance. The target keywords and phrases to be analyzed are sorted into a list for subsequent traversal and matching. Traverse this list, that is, check the keywords and phrases one by one. During the traversal process, use a string matching method (such as forward maximum matching method, reverse maximum matching method, bidirectional maximum matching method, etc.) to find out whether the target keywords and phrases appear in the processed comparison table. The matching process can be regarded as a character-by-character comparison until a complete match is found or it is determined that there is no match. If a matching keyword or phrase is found in the comparison table, the sentiment tendency corresponding to the keyword or phrase is directly retrieved as the sentiment tendency result.

[0196] The effect of the above technical solution is: the technical solution is implemented through a computer program, which can automatically complete the matching of keywords and phrases and the acquisition of emotional tendencies, greatly improving work efficiency. By preprocessing the comparison table, null values ​​and duplicates are removed, reducing errors in the matching process. At the same time, the use of string matching methods for searching can ensure the accuracy of matching. The technical solution can be applied to different fields and scenarios. It only needs to adjust the keywords and phrases in the comparison table and their corresponding emotional tendencies according to actual needs. With the development of the business and the emergence of new keywords, the comparison table can be easily updated to meet new needs. At the same time, the technical solution can also be combined with other technologies (such as natural language processing, machine learning, etc.) to further improve the accuracy and efficiency of sentiment analysis.

[0197] In summary, this technical solution realizes the sentiment tendency analysis of target keywords and phrases in an automated, accurate, flexible and scalable manner, providing strong support for subsequent decision-making and application.

[0198] An ASR-based early media recognition system is proposed in an embodiment of the present invention. Figure 2 As shown, the ASR-based early media recognition system includes:

[0199] The audio data collection and processing module is used to collect the audio data corresponding to the target media in real time, repair the audio data by comparing it, and obtain the repaired audio data;

[0200] A text conversion module, used for inputting the repaired audio data into an ASR model to obtain text data corresponding to the audio data;

[0201] A key information identification module, used to identify keywords and phrases from the text data to obtain target keywords and phrases;

[0202] The sentiment tendency recognition module is used to compare and recognize the target keywords and phrases with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are early media recognition results;

[0203] The content push control module is used to push the media content according to the emotional tendency result.

[0204] The working principle of the above technical solution is as follows: collect audio data corresponding to the target media (such as news broadcasts, social media videos, etc.) in real time. Preprocess the collected audio data, including noise reduction, frame segmentation, etc., to improve the audio quality. Compare the audio data and perform necessary repairs, such as removing noise, filling silent segments, etc., to obtain the repaired audio data. Input the repaired audio data into the ASR model and use the acoustic model and language model of the ASR model to convert the audio data into corresponding text data. The acoustic model is responsible for converting the speech signal into the corresponding acoustic feature sequence, and the language model is trained based on a large amount of text data to evaluate whether the generated text sequence conforms to the grammatical rules and idiomatic usage of the language. From the converted text data, natural language processing technology is used to identify keywords and phrases. Extract target keywords and phrases with representative and information value. Compare the extracted target keywords and phrases with the preset keywords and phrases and sentiment tendency table. According to the comparison results, obtain the sentiment tendency results corresponding to the audio data, such as positive, negative or neutral. This sentiment tendency result is the early media recognition result, which is used for subsequent content push decisions. According to the results of emotional tendency, media content is screened and classified, and media content that meets user preferences or needs is pushed to users.

[0205] The effect of the above technical solution is: through ASR technology, audio data is converted into text data in real time, which greatly shortens the recognition time and improves the recognition efficiency. Using the acoustic model and language model of the ASR model, as well as natural language processing technology, it is possible to accurately identify keywords and phrases in the audio data, as well as the corresponding emotional tendencies. According to the emotional tendency results, media content can be personalized and classified to meet the personalized needs of users. This method is not only applicable to media content such as news broadcasts and social media videos, but can also be extended to smart homes, medical care, finance and other fields to achieve more intelligent and personalized human-computer interaction.

[0206] In summary, the ASR-based early media recognition method achieves fast, accurate recognition and personalized push of media content through real-time collection and repair of audio data, speech recognition, keyword and phrase recognition, sentiment tendency recognition, and content push. It has broad application prospects and important technical value.

[0207] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. An early media recognition method based on ASR, characterized in that: The ASR-based early media recognition method includes: Collecting audio data corresponding to the target media in real time, repairing the audio data by comparing it, and obtaining the repaired audio data; Inputting the repaired audio data into the ASR model to obtain text data corresponding to the audio data; Perform keyword and phrase recognition from the text data to obtain target keywords and phrases; Compare and identify the target keywords and phrases with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are early media recognition results; Pushing media content according to the sentiment tendency result; The method of collecting audio data corresponding to the target media in real time, repairing the audio data by comparing the audio data, and obtaining the repaired audio data includes: Collect audio data corresponding to the target media in real time; Performing enhancement processing on the audio data to obtain enhanced audio data, wherein the enhancement processing includes noise reduction and de-pop processing and volume normalization processing; Performing audio quality evaluation on the enhanced audio data to obtain an audio quality result; Performing audio repair on audio data whose audio quality result does not meet the audio quality requirement, and obtaining repaired audio data; The audio quality of the enhanced audio data is evaluated to obtain an audio quality result, including: Extracting audio parameters of the enhanced audio data, wherein the audio parameters include signal-to-noise ratio, total harmonic distortion and short-time energy; Acquiring audio quality assessment parameters using the signal-to-noise ratio, total harmonic distortion and short-time energy included in the audio parameters, Comparing the audio quality assessment parameter with a preset audio quality assessment parameter threshold; When the audio quality evaluation parameter is lower than a preset audio quality evaluation parameter threshold, it is determined that the audio quality result does not meet the audio quality requirement of the audio data.

2. The ASR-based early media recognition method according to claim 1, It is characterized in that The audio quality evaluation parameter is obtained by the following formula: in, P Represents audio quality assessment parameters; n Indicates the number of time windows corresponding to the short-time energy contained in the audio data; S i Indicates i The signal-to-noise ratio corresponding to the time window; T i Indicates i The total harmonic distortion corresponding to the time window; S max and S min Indicates the audio data corresponding to n The maximum and minimum signal-to-noise ratio of a time window; W max and W min Indicates the audio data corresponding to n The maximum and minimum short-term energy of a time window; W i Indicates i The short-term energy corresponding to the time window; W bi Indicates i The short-time energy standard deviation corresponding to the time window.

3. The ASR-based early media identification method according to claim 1, characterized in that: Setting a time window corresponding to the short-time energy contained in the audio parameters of the audio data includes: Extract the sampling frequency, bit depth and number of channels corresponding to the audio data collection device; Obtaining a first window adjustment coefficient using a sampling frequency, a bit depth, and a number of channels corresponding to the audio data collection device; The first window adjustment coefficient is obtained by the following formula: in, K 01 represents the first window adjustment coefficient; f Indicates the sampling frequency corresponding to the audio data collection device; f c Indicates the preset sampling frequency reference value; N Indicates the number of channels; B Indicates the bit depth; int() means rounding up the value in the brackets; comparing the first window adjustment coefficient with a preset window adjustment coefficient threshold; When the first window adjustment coefficient is lower than a preset window adjustment coefficient threshold, the preset initial time length is used as the time window corresponding length to set the time window corresponding to the short-time energy; When the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, the initial time length is adjusted, and the time window corresponding to the short-time energy is set using the adjusted time length as the time window corresponding length.

4. The ASR-based early media recognition method according to claim 3, characterized in that: When the first window adjustment coefficient is not lower than the preset window adjustment coefficient threshold, the initial time length is adjusted, and the time window corresponding to the short-time energy is set by using the adjusted time length as the time window corresponding length, including: When the first window adjustment coefficient is not lower than a preset window adjustment coefficient threshold, extracting the initial time length and initial window overlap ratio values ​​corresponding to the time window of the ASR model; wherein the ASR model adopts a deep neural network model structure; Extract the time length value and the corresponding window overlap ratio value that appear in the dynamic adjustment process of the time window of the ASR model; The second window adjustment coefficient is obtained by using the initial time length and the initial window overlap ratio value corresponding to the time window of the ASR model in combination with the time length value and the corresponding window overlap ratio value appearing in the dynamic adjustment process of the time window; The second window adjustment coefficient is obtained by the following formula: in, K 02 represents the second window adjustment coefficient; m Indicates the number of dynamic adjustments of the time window; T i Indicates i The time length corresponding to the dynamically adjusted time window; P i Indicates i The window overlap ratio value after the dynamic adjustment; P ti Indicates i The change rate of the time length between the time window after the dynamic adjustment and the time window before the dynamic adjustment; T c and P c Indicates the initial time length and initial window overlap ratio; P z express m The middle value of the window overlap ratio after the dynamic adjustment; P b express m The standard deviation of the window overlap ratio after dynamic adjustment; T b express m The standard deviation of the time length corresponding to the time window after the dynamic adjustment; P max express m The maximum value of the window overlap ratio after dynamic adjustment; The initial time length is adjusted using the first window adjustment coefficient and the second window adjustment coefficient to obtain an adjusted time length; The adjusted time length is obtained by the following formula: in, T x Indicates the adjusted length of time; T 0 indicates the length of time before adjustment; K 01 represents the first window adjustment coefficient; K 02 represents the second window adjustment coefficient; t represents the adjustment factor, and the value range of the adjustment factor is 1.12-1.43; The time window corresponding to the short-time energy is set by using the adjusted time length as the corresponding length of the time window.

5. The ASR-based early media identification method according to claim 1, characterized in that: Performing audio repair on audio data whose audio quality result does not meet the audio quality requirement, and obtaining repaired audio data, including: Extracting audio data whose audio quality results do not meet the audio quality requirements as audio data to be processed; Using a filter to perform noise removal processing on the audio data to be processed, and obtaining the audio data after the noise removal processing; Performing a distortion type analysis on the audio data after the noise removal process to obtain the distortion type of the audio data after the noise removal process; Retrieving a distortion processing method corresponding to the distortion type from a database; Using the distortion processing method corresponding to the distortion type, compensating and repairing the audio signal corresponding to each distortion type in the audio data after noise removal processing, to obtain the audio data after distortion repair; Determining missing data in the distortion-repaired audio data, and obtaining an index position corresponding to the missing data; Extract the audio signal data values ​​of the preceding and following data points corresponding to the index position; Using an interpolation algorithm to combine the audio signal data values ​​of the preceding and following data points corresponding to the index position to obtain the filling value corresponding to the missing data; Each missing data is filled with a filling value corresponding to the missing data to obtain audio data after the missing value is repaired.

6. The ASR-based early media identification method according to claim 1, characterized in that: Inputting the repaired audio data into the ASR model to obtain text data corresponding to the audio data includes: Extract the short-time energy average of the audio data after the current missing value is repaired; Comparing the short-time energy average value with a preset short-time energy threshold; When the short-time energy average value is lower than the preset short-time energy threshold, the model time window of the ASR model is not adjusted; When the short-time energy average value is not lower than the preset short-time energy threshold, the model time window and the window overlap ratio value of the ASR model are adjusted to obtain the adjusted model time window length and the window overlap ratio value; The adjusted model time window length is obtained by the following formula: in, T s Represents the adjusted model time window length; T s0 Indicates the model time window length before adjustment; n Indicates the number of time windows corresponding to the short-time energy contained in the audio data; W i Indicates i The short-term energy corresponding to the time window; W max and W min Indicates the audio data corresponding to n The maximum and minimum short-term energy of a time window; W b express n The short-time energy standard deviation corresponding to the time window; E m Indicates the minimum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy maximum value is located; E n Indicates the maximum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy minimum value is located; W p Represents the short-time energy average; W y Indicates the preset short-time energy threshold; At the same time, the adjusted window overlap ratio value is obtained by the following formula: in, B c Indicates the adjusted window overlap ratio value; B c0 Indicates the window overlap ratio value before adjustment; T s Represents the adjusted model time window length; T s0 Indicates the model time window length before adjustment; E m Indicates the minimum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy maximum value is located; E n Indicates the maximum short-time energy value of the time frame corresponding to the time frame contained in the time window where the short-time energy minimum value is located; Adjust the ASR model according to the adjusted model time window length and window overlap ratio values; Inputting the repaired audio data into an ASR model, processing the repaired audio data through the ASR model, and obtaining a text string corresponding to the repaired audio data; Perform data processing on the text character string to obtain processed text data.

7. The ASR-based early media identification method according to claim 1, characterized in that: Recognizing keywords and phrases from the text data to obtain target keywords and phrases includes: Recall a preset list of keywords and phrases; Traversing the text data, and using a string matching method to find out whether the key words and phrases exist in the text data; When keywords and phrases are found in text data, the location of the keywords and phrases in the text and their context are recorded and marked; The found keywords and phrases are used as target keywords and phrases.

8. The ASR-based early media recognition method according to claim 1, characterized in that: Compare and identify the target keywords and phrases with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, including: Retrieve a comparison table of keywords and phrases with unified structure and sentiment tendency; Remove the empty values ​​and duplicate items in the comparison table of the keywords and phrases and sentiment orientation, and obtain a processed comparison table of the keywords and phrases and sentiment orientation; forming a target keyword and phrase list with the target keywords and phrases; Traversing the target keyword and phrase list; During the traversal process, the string matching method searches whether the keywords and phrases in the target keyword and phrase list appear in the comparison table of processed keywords and phrases and sentiment orientation; When the keywords and phrases in the target keyword and phrase list appear in the comparison table of processed keywords and phrases and sentiment orientation, the sentiment orientation corresponding to the keywords and phrases is retrieved to obtain the sentiment orientation result.

9. An early media recognition system based on ASR, characterized in that: The ASR-based early media recognition system includes: The audio data collection and processing module is used to collect the audio data corresponding to the target media in real time, repair the audio data by comparing it, and obtain the repaired audio data; A text conversion module, used for inputting the repaired audio data into an ASR model to obtain text data corresponding to the audio data; A key information identification module, used to identify keywords and phrases from the text data to obtain target keywords and phrases; The sentiment tendency recognition module is used to compare and recognize the target keywords and phrases with the preset keywords and phrases and sentiment tendency table to obtain sentiment tendency results, and the sentiment tendency results are early media recognition results; A content push control module, used for pushing media content according to the emotional tendency result; The method of collecting audio data corresponding to the target media in real time, repairing the audio data by comparing the audio data, and obtaining the repaired audio data includes: Collect audio data corresponding to the target media in real time; Performing enhancement processing on the audio data to obtain enhanced audio data, wherein the enhancement processing includes noise reduction and de-pop processing and volume normalization processing; Performing audio quality evaluation on the enhanced audio data to obtain an audio quality result; Performing audio repair on audio data whose audio quality result does not meet the audio quality requirement, and obtaining repaired audio data; The audio quality of the enhanced audio data is evaluated to obtain an audio quality result, including: Extracting audio parameters of the enhanced audio data, wherein the audio parameters include signal-to-noise ratio, total harmonic distortion and short-time energy; Acquiring audio quality assessment parameters using the signal-to-noise ratio, total harmonic distortion and short-time energy included in the audio parameters, Comparing the audio quality assessment parameter with a preset audio quality assessment parameter threshold; When the audio quality evaluation parameter is lower than a preset audio quality evaluation parameter threshold, it is determined that the audio quality result does not meet the audio quality requirement of the audio data.

Citation Information

Patent Citations

  • Advertisement demand determination system and method based on voice analysis

    CN118261650A