Processing method for automatically optimizing audio effect of intelligent equipment
Through intelligent devices, they capture ambient audio data in real time, use the sound source-frequency domain feature table to identify the sound source and analyze the noise level, filter effective audio and compare it with historical standard characteristics, and optimize it according to the adjustment methods of industry standards, which solves the noise interference problem of intelligent devices when capturing ambient audio, and improves processing efficiency and accuracy.
Patent Information
- Application Number
- CN202510511293.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Existing smart devices are susceptible to background noise interference when capturing ambient audio data, and the audio processing standards in different fields are inconsistent, resulting in the audio adjustment method being unpredictable.
Ambient audio data is captured in real time through intelligent devices, the audio source-frequency domain feature table is used to identify the sound source, analyze the noise level and extract features, filter effective audio and compare it with historical standard features, and optimize according to the priority sorting of industry standards adjustment methods.
Improves the quality of audio data capture, reduces background noise interference, improves processing efficiency and accuracy, and enhances the applicability of audio.
Smart Images

Figure CN120356480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and particularly to a method for automatically optimizing the audio effect of intelligent devices. Background Art
[0002] The popularization of intelligent devices has made it easy to capture environmental audio data. The public's demand for the automation and intelligence of audio analysis has increased. The existing technologies mainly rely on traditional microphones and recording devices, with limited ability to capture environmental audio, being easily interfered by background noise, and inconsistent audio processing standards in different fields, resulting in the lack of universality in audio adjustment methods among industries.
[0003] Therefore, the present invention provides a method for automatically optimizing the audio effect of intelligent devices. Summary of the Invention
[0004] The method for automatically optimizing the audio effect of intelligent devices provided by the present invention captures environmental audio data in real time through an intelligent device, identifies the existing sound sources using a sound source-frequency domain feature table, analyzes the noise level and extracts features, based on the extracted noise features, screens out effective audio from the actual audio data and extracts its features, optimizes by comparing with historical standard features, and finally sorts the priorities of industry-standard adjustment methods to further improve the audio processing effect, realizing the improvement of the capture quality of audio data, reducing the interference of background noise, improving the processing efficiency and accuracy, and enhancing the applicability of the audio.
[0005] The present invention provides a method for automatically optimizing the audio effect of intelligent devices, including: Step 1: Use an intelligent device to capture environmental audio data of the surrounding environment, determine the existing sound sources in the environmental audio data based on a sound source-frequency domain feature table, perform noise level analysis based on the existing sound sources, and extract the noise features in the existing sound sources; Step 2: Use an intelligent device to capture actual audio data, extract effective audio data from the actual audio data based on the noise features, and extract the effective features of the effective audio data; Step 3: Determine the standard features corresponding to each effective feature according to a historical standard table, compare the effective features with the standard features, and perform a first optimization on the effective audio data based on the comparison results; Step 4: Obtain industry audio adjustment methods, sort the priorities of the industry audio adjustment methods to obtain an optimal adjustment method, and perform a second optimization on the first optimization result based on the optimal adjustment method.
[0006] The present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The intelligent device is used to capture the ambient audio data of the surrounding environment. Based on the sound source-frequency domain feature table, the existing sound sources in the ambient audio data are determined. Noise level analysis is performed based on the existing sound sources, and the noise characteristics in the existing sound sources are extracted, including: Perform frequency domain analysis on the ambient audio data to determine the ambient spectrum characteristics. Analyze the ambient spectrum characteristics, and perform sound source stripping on the ambient audio data according to the feature analysis results and the sound source-frequency domain feature table to obtain the existing sound sources; Divide the noise levels of the existing sound sources according to the noise duration and noise bands, and extract the noise characteristics of each existing sound source based on the noise levels.
[0007] The present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The level division is performed according to the noise duration and noise bands, including: Obtain the actual application scenario of the intelligent device, and divide the first level of the noise duration based on the actual application scenario; Determine the common bands according to the actual application scenario, and divide the second level of the noise bands based on the common bands; Comprehensively determine the noise levels of each existing sound source based on the first level and the second level, and then extract the noise characteristics of each existing sound source according to the noise fans.
[0008] The present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The intelligent device is used to capture the actual audio data. Based on the noise characteristics, the effective audio data in the actual audio data is extracted, and the effective characteristics of the effective audio data are extracted, including: Compare the noise characteristics with the actual audio data, determine the noise part and the remaining audio part according to the comparison result, and perform segmentation processing on the actual audio data according to the noise characteristics to obtain the segmentation processing result; Use the audio event detection algorithm to identify specific events in the segmentation processing result, determine the event category and confidence level of each time period, set an effective threshold according to the confidence level and event category, perform an effective judgment on the specific events, and comprehensively obtain all effective events to obtain the effective audio data; Extract the characteristics of each effective event, integrate the extracted characteristics to form an event vector, determine the task requirements based on the automatic optimization of the audio effect, and select the event vector according to the task requirements, and then obtain the effective characteristics of the effective audio data.
[0009] The present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. Use the audio event detection algorithm to identify specific events in the segmentation processing result, determine the event category and confidence level of each time period, and set an effective threshold according to the confidence level and event category, including: Obtain industry requirements, classify the types of the industry requirements, and determine the confidence threshold corresponding to a specific event in combination with the classification result. Collect historical audio data, match historical events of the same category as the specific event from the historical audio data, and set the time period threshold corresponding to the specific event based on the historical events. Derive an effective threshold by combining the confidence threshold and the time period threshold.
[0010] The present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. Determine the standard feature corresponding to each effective feature according to the historical standard table, compare the effective feature with the standard feature, and perform a first optimization on the effective audio data based on the comparison result, including: Comprehensively evaluate the effective feature and the standard feature, and determine the difference between the effective feature and the standard feature based on the comprehensive evaluation. Derive the feature to be optimized in the effective feature according to the difference situation, and propose a first optimization direction for the feature to be optimized according to the difference situation. Perform a first optimization on the effective audio data corresponding to the feature to be optimized according to the optimization direction.
[0011] The present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The comprehensive evaluation of the effective feature and the standard feature includes: , where P represents the comprehensive evaluation of the effective feature set; A represents the weighted overlap degree between the effective feature set and the standard feature set; B represents the clustering overlap degree of the clustering results of the effective feature set and the standard feature set; represents the weight coefficient of the weighted overlap degree in the comprehensive evaluation; represents the weight coefficient of the conditional mutual information in the comprehensive evaluation; represents the weight coefficient of the weighted Shannon diversity index in the comprehensive evaluation; represents the effective feature set; represents the standard feature set; Z represents the existence of a sound source variable; represents the existence of a sound source variable Z of the standard feature set under the condition entropy; represents the existence of a sound source variable Z of the effective feature set under the condition entropy; represents the existence of a sound source variable Z under the condition of the standard feature set and the effective feature set of the joint conditional entropy; Represents the standard feature set and the effective feature set The conditional mutual information under the condition of the existence of sound source variables Z ; Represents the weighted Shannon diversity index; Represents the weight of the effective feature; f Represents the effective feature; Represents the i th cluster in the effective feature set; Represents the effective feature f 's information gain; Represents the i th cluster's stability in the effective feature set; C represents the set of all clusters; D Represents the total number of resamplings; Represents the d th cluster obtained by the th resampling; d Represents the index of the number of resamplings; i Represents the index of the clustering of the effective feature set; Represents the weight coefficient of the clustering result in the comprehensive evaluation.
[0012] The present invention provides a processing method for automatically optimizing the audio effect of intelligent devices, performs a performance evaluation on the first optimization result, prioritizes the industry audio adjustment methods according to the optimization evaluation result, obtains the optimal adjustment method, and performs a second optimization on the first optimization result based on the optimal adjustment method, including: Performs a performance evaluation on the first optimization result, collects the user's feedback on the first optimization result, analyzes the advantages and disadvantages of the first optimization result based on the performance evaluation and the user's feedback, and determines the second optimization direction according to the advantages and disadvantages; Prioritizes the industry audio adjustment methods based on the second optimization direction, obtains the optimal adjustment method, and performs a second optimization on the first optimization result.
[0013] Compared with the prior art, the beneficial effects of the present application are as follows: By using intelligent devices to capture environmental audio data in real time, identifying the existing sound sources using the sound source-frequency domain feature table, analyzing the noise level and extracting features, based on the extracted noise features, screening effective audio from the actual audio data and extracting its features, performing optimization by comparing with historical standard features, and finally prioritizing the industry standard adjustment methods, further improving the audio processing effect, achieving an improvement in the capture quality of audio data, reducing the interference of background noise, improving the processing efficiency and accuracy, and enhancing the applicability of the audio.
[0014] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention can be realized and attained by the structure particularly pointed out in the written description and the drawings.
[0015] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0016] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings: Figure 1 is a schematic flowchart of a processing method for automatically optimizing the audio effect of an intelligent device provided by an embodiment of the present invention. Detailed Embodiments
[0017] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not used to limit the present invention.
[0018] Embodiment 1:
[0019] The embodiment of the present invention provides a processing method for automatically optimizing the audio effect of an intelligent device, as Figure 1 shown, including: Step 1: Use the intelligent device to capture the environmental audio data of the surrounding environment, determine the existing sound sources in the environmental audio data based on the sound source-frequency domain feature table, perform noise level analysis based on the existing sound sources, and extract the noise features in the existing sound sources; Step 2: Use the intelligent device to capture the actual audio data, extract the effective audio data in the actual audio data based on the noise features, and extract the effective features of the effective audio data; Step 3: Determine the standard features corresponding to each effective feature according to the historical standard table, compare the effective features with the standard features, and perform a first optimization on the effective audio data based on the comparison result; Step 4: Obtain the industry audio adjustment methods, rank the priorities of the industry audio adjustment methods to obtain the optimal adjustment method, and perform a second optimization on the first optimization result based on the optimal adjustment method.
[0020] In this embodiment, the environmental audio data refers to the audio signal recorded in a specific environment, including the mixture of various sound sources such as background noise, natural sounds, and human voices. For example, the recording of a city street contains car sounds, pedestrian conversations, and bird songs.
[0021] In this embodiment, the sound source - frequency domain feature table is a database that lists the features of various sound sources in the frequency domain, such as the energy distribution and duration in a specific frequency range, etc., for sound source recognition. For example, the frequency characteristics of bird calls listed in the table are 2000 - 4000 Hz, and those of car sounds are 100 - 500 Hz.
[0022] In this embodiment, the existence of a sound source refers to the specific sound source identified after sound source stripping, representing the actual sound existing in a specific environment. For example, after processing, "bird calls" and "car sounds" are identified as existing sound sources.
[0023] In this embodiment, an intelligent device refers to an electronic device with sensing, computing, and communication capabilities, capable of data exchange and processing through the Internet. For example, smart speakers, smart home monitoring systems, smartphones, etc.
[0024] In this embodiment, the noise characteristics refer to the specific attributes of each identified existing sound source, including its frequency, duration, intensity, etc. For example, the characteristics of traffic noise may include a frequency range of 100 - 500 Hz, a duration exceeding 10 minutes, and an intensity of 85 dB, etc.
[0025] In this embodiment, by analyzing the actual application scenarios of intelligent devices, the first level of noise duration is divided, and the common frequency bands are determined to divide the second level of noise frequency bands. Combining these two levels, the noise level of each existing sound source is determined, and the corresponding noise characteristics are extracted.
[0026] In this embodiment, the effective audio data refers to the audio data that contains effective events after processing, with the noise part removed. For example, in surveillance recordings, only the segment of "human conversation" is retained after removing the noise.
[0027] In this embodiment, the effective features refer to the important features extracted from effective events for subsequent analysis and applications. For example, the effective features may include information such as the frequency range, duration, and volume of human voices.
[0028] In this embodiment, by comparing the noise characteristics with the actual audio data, the noise part and the remaining audio part are identified and segmented. The audio event detection algorithm is used to identify specific events in the segmentation results, determine the event category and confidence level, judge the specific events according to the set effective threshold, synthesize the effective events to extract the effective audio data, finally form an event vector and optimize the feature selection according to the task requirements to obtain the effective features.
[0029] In this embodiment, standard features refer to features that are considered ideal or benchmark in audio analysis. These features are usually verified and can effectively represent the quality of a specific event. For example, in speech recognition, standard features may include spectral features of audio, pitch, and volume range, etc.
[0030] In this embodiment, the first optimization refers to the preliminary improvement or adjustment of the feature to be optimized according to the optimization direction, usually achieved through algorithms or manual adjustment. For example, frequency filtering is performed on the effective audio data to improve the clarity of the audio.
[0031] In this embodiment, a comprehensive evaluation is performed on the effective features and the standard features to determine the difference between the two. Based on the difference, the features to be optimized in the effective features are identified, and the corresponding optimization directions are proposed. According to these optimization directions, the first optimization is performed on the effective audio data corresponding to the features to be optimized.
[0032] In this embodiment, the historical standard table is a reference tool for comparing and evaluating effective features. Its content includes the audio features considered to be the best in a specific application or industry, containing historical data or cases to show the performance of these standard features under specific conditions. For example, it is assumed that the historical standard table stipulates that the standard signal-to-noise ratio of a certain environmental audio is 20 dB, while the extracted effective features show a signal-to-noise ratio of 15 dB. Through this comparison, the team can clearly identify the need to take measures to improve the signal-to-noise ratio to meet or exceed the industry standard.
[0033] In this embodiment, the priority ranking is to rank the industry audio adjustment methods according to the second optimization direction.
[0034] In this embodiment, the industry audio adjustment methods refer to various adjustment and optimization strategies that can be adopted in audio processing and analysis. For example, introducing a deep learning model for feature extraction, using data augmentation techniques to generate more training samples, and optimizing the audio signal processing algorithm to improve the processing speed.
[0035] In this embodiment, the optimal adjustment method is to select the most effective adjustment strategy based on the priority ranking and industry standards. For example, choosing to introduce a deep learning model as the optimal adjustment method because it can significantly improve the recognition accuracy.
[0036] In this embodiment, the second optimization is to implement further optimization according to the determined optimal adjustment method. For example, implementing the training and integration of the deep learning model, replacing the original traditional model, and conducting system testing to evaluate the effect.
[0037] The working principle and beneficial effects of the above technical solution are as follows: The intelligent device captures environmental audio data in real time, identifies the existing sound sources using the sound source-frequency domain feature table, analyzes the noise level and extracts features. Based on the extracted noise features, valid audio is screened from the actual audio data and its features are extracted. Through comparison with historical standard features, optimization is carried out. Finally, the priority order of the adjustment method is sorted according to industry standards, further improving the audio processing effect, achieving an improvement in the capture quality of audio data, reducing the interference of background noise, improving the processing efficiency and accuracy, and enhancing the applicability of the audio.
[0038] Embodiment 2: The embodiment of the present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The intelligent device is used to capture the environmental audio data of the surrounding environment, determine the existing sound sources in the environmental audio data based on the sound source-frequency domain feature table, perform noise level analysis based on the existing sound sources, and extract the noise features in the existing sound sources, including: Perform frequency domain analysis on the environmental audio data to determine the environmental spectrum features, analyze the environmental spectrum features, and perform sound source stripping on the environmental audio data according to the feature analysis results and the sound source-frequency domain feature table to obtain the existing sound sources; Divide the noise level of the existing sound sources according to the noise duration and noise band, and extract the noise features of each existing sound source based on the noise level.
[0039] In this embodiment, the frequency domain analysis is to perform a Fourier transform on the audio signal to convert the time domain signal into a frequency domain signal to analyze its frequency components and amplitude characteristics. For example, convert a piece of music signal into a spectrogram to observe the energy distribution of different frequencies.
[0040] In this embodiment, the environmental spectrum features refer to the spectrum information extracted from the environmental audio data, such as the energy of specific frequencies, the shape and distribution of the spectrum, etc., which reflect the characteristics of the environmental sound. For example, if it is found that the energy of a certain frequency band (such as 500 - 1000 Hz) is relatively high in the spectrum, it means that the sound source in this frequency band is more prominent in the environment.
[0041] In this embodiment, the feature analysis result is the conclusion obtained after analyzing the extracted spectrum features, such as the significance, correlation of the features, and their relationship with the sound source. For example, the analysis result shows that the feature near 300 Hz has a relatively high correlation with traffic noise.
[0042] In this embodiment, the sound source stripping is the process of extracting a specific sound source from a mixed audio signal. Usually, signal processing techniques are used to separate the target sound source from the background noise. For example, extract the human voice part from a recording containing human voice and music.
[0043] In this embodiment, the noise duration and noise band are used to classify the noise level of the existing sound source. According to the duration and frequency range of the noise, the noise of the existing sound source is classified to evaluate its impact on the environment. For example, traffic noise with a duration exceeding 5 minutes is classified as high-level noise, while brief bird chirping is classified as low-level noise.
[0044] The working principle and beneficial effects of the above technical solution are as follows: By performing frequency-domain analysis on environmental audio data, extracting environmental spectrum features, and comparing them with the sound source-frequency domain feature table, sound source separation is achieved. According to the duration and band of the noise, the noise level of the existing sound source is classified, and corresponding noise features are extracted based on the noise level, enabling more accurate identification and processing of audio data, reducing background noise interference, and enhancing the accuracy of sound source identification.
[0045] Embodiment 3: The embodiment of the present invention provides a method for automatically optimizing the audio effect of an intelligent device, which includes classifying according to the noise duration and noise band, and specifically includes: Obtain the actual application scenario of the intelligent device, and divide the first level of the noise duration based on the actual application scenario; Determine the common band according to the actual application scenario, and divide the second level of the noise band based on the common band; Comprehensively determine the noise level of each existing sound source based on the first level and the second level, and then determine the noise characteristics of each existing sound source according to the noise characteristics.
[0046] In this embodiment, the actual application scenario refers to the situation where the intelligent device is used in a specific environment, including the functional requirements of the device and the usage habits of the user. For example, environmental monitoring in an urban park, noise monitoring at home, audio analysis in an office, etc.
[0047] In this embodiment, the first level refers to the noise level divided according to the noise duration, usually divided into brief noise, continuous noise, intermittent noise, etc. For example, brief noise (such as a ringtone) is the first level, and continuous noise (such as air conditioner operation) is the second level.
[0048] In this embodiment, the common band refers to the common frequency range in a specific application scenario, which is used to analyze and classify noise. For example, in an urban environment, the common band may include low frequency (20 - 200 Hz) for traffic noise and medium frequency (200 - 2000 Hz) for human voices.
[0049] In this embodiment, the noise level is a level assigned to each existing sound source after comprehensively considering the noise duration and frequency band, reflecting its impact on the environment. For example, traffic noise with a long duration and high frequency is classified as a high noise level, while brief bird chirping is classified as a low noise level.
[0050] The working principle and beneficial effects of the above technical solution are as follows: By analyzing the actual application scenarios of intelligent devices, the first level of noise duration is divided, and the common frequency bands are determined to divide the second level of noise bands. Combining these two levels, the noise level of each existing sound source is determined, so as to extract the corresponding noise characteristics, in order to more effectively understand and manage environmental audio, improve the performance of intelligent devices in different scenarios, customize noise characteristics according to actual application scenarios, and enhance the pertinence and effectiveness of processing.
[0051] Embodiment 4: The embodiment of the present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The intelligent device is used to capture actual audio data, and based on the noise characteristics, the effective audio data in the actual audio data is extracted, and the effective characteristics of the effective audio data are extracted, including: Compare the noise characteristics with the actual audio data, determine the noise part and the remaining audio part according to the comparison result, and perform segmentation processing on the actual audio data according to the noise characteristics to obtain the segmentation processing result; Use an audio event detection algorithm to identify specific events in the segmentation processing result, determine the event category and confidence level of each time period, set an effective threshold according to the confidence level and event category, perform an effective judgment on the specific events, and integrate all effective events to obtain effective audio data; Extract the characteristics of each effective event, integrate the extracted characteristics to form an event vector, determine the task requirements based on the automatic optimization of the audio effect, and select the event vector according to the task requirements, thereby obtaining the effective characteristics of the effective audio data.
[0052] In this embodiment, the comparison process is to compare the extracted noise characteristics with the actual audio data to identify the noise part and other audio parts contained in the audio data. For example, match the background noise characteristics with the recording data to find out which parts are noise.
[0053] In this embodiment, the comparison result is the conclusion obtained after the comparison, usually including the time period of the noise part and the identification information of the remaining audio part. For example, the comparison result shows that the first 5 seconds of the recording is noise and the last 10 seconds is effective audio.
[0054] In this embodiment, the noise part refers to the area in the audio data that is recognized as noise, while the remaining audio part refers to the valid audio content remaining after removing the noise. For example, in a recording, traffic noise is recognized as the noise part, and human voices are the remaining audio part.
[0055] In this embodiment, a specific event refers to a specific sound or event that needs to be detected and recognized in the audio data, such as human voices, alarm sounds, music segments, etc. For example, in a surveillance recording, the specific events identified may be "fighting sounds" or "window-breaking sounds".
[0056] In this embodiment, the segmentation processing result refers to the audio data segments obtained after segmentation processing. Each segment may contain noise or valid audio. For example, the audio data is segmented into multiple segments, some of which contain valid human voices and others are noise.
[0057] In this embodiment, the event category and confidence level: The event category refers to the type of the recognized event, and the confidence level refers to the degree of confidence of the algorithm in recognizing this event, usually expressed as a percentage. For example, the recognized event category may be "human voice" with a confidence level of 85%.
[0058] In this embodiment, the valid threshold is a set standard used to determine whether an event is considered valid, usually set based on the confidence level. For example, if the valid threshold is set at 80%, events with a confidence level lower than 80% will be ignored.
[0059] In this embodiment, a valid event refers to a specific event confirmed after being judged by the valid threshold, usually with a relatively high confidence level. For example, an event of recognizing "human voice" with a confidence level of 90% is considered a valid event.
[0060] In this embodiment, the event vector is a vector formed by integrating the features of each extracted valid event, used for subsequent analysis and optimization. For example, an event vector may contain information such as the frequency characteristics, duration, and intensity of the event.
[0061] In this embodiment, determining the task requirements based on automatic optimization of audio effects is to automatically adjust and optimize the task requirements according to the characteristics and effects of the audio data to improve the processing effect. For example, if multiple "alarm sounds" are detected, the task requirements may be adjusted to strengthen the monitoring and response of the alarm sounds.
[0062] In this embodiment, the selection process is a process of evaluating and screening the event vectors to determine which features best meet the current task requirements. For example, event vectors containing the features of "human voice" and "alarm sound" are selected, while other irrelevant features are ignored.
[0063] The working principle and beneficial effects of the above technical solution are as follows: By comparing the noise characteristics with the actual audio data, the noise part and the remaining audio part are identified and segmented. The audio event detection algorithm is used to identify specific events in the segmentation results, determine the event categories and confidence levels, judge the specific events according to the set effective threshold, synthesize the effective events to extract the effective audio data, and finally form an event vector and optimize the feature selection according to the task requirements, effectively distinguishing noise from audio, improving the recognition accuracy of specific events, enhancing the audio effect and processing efficiency, and strengthening the adaptability of the system in different application scenarios.
[0064] Embodiment 5: The embodiment of the present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The audio event detection algorithm is used to identify specific events in the segmentation processing results, determine the event categories and confidence levels of each time period, and set an effective threshold according to the confidence level and event category, including: Obtain industry requirements, classify the industry requirements, and determine the confidence threshold corresponding to the specific event in combination with the classification result; Collect historical audio data, match historical events of the same category as the specific event from the historical audio data, and set a time period threshold corresponding to the specific event based on the historical events; Derive an effective threshold by synthesizing the confidence threshold and the time period threshold.
[0065] In this embodiment, the industry requirements refer to the specific requirements and goals of a specific industry in audio analysis or processing, usually involving aspects such as improving efficiency, accuracy, and user experience. For example, in the security industry, the requirements may include real-time monitoring and alarm systems to quickly respond to emergencies.
[0066] In this embodiment, the classification is to classify the industry requirements to better understand and meet the specific events and processing methods corresponding to different requirements. For example, the requirements are classified into categories such as "environmental monitoring", "security monitoring", and "customer service".
[0067] In this embodiment, the confidence threshold is the confidence standard for identifying specific events, usually set as a percentage, indicating the minimum confidence level at which the recognition result is considered valid. For example, for the "human voice" event, the set confidence threshold is 80%, that is, only the recognition results with a confidence level higher than 80% are considered valid.
[0068] In this embodiment, the historical audio data refers to the audio records collected in the past. These data can be used to train models or perform event matching to improve the recognition accuracy. For example, a surveillance recording containing various environmental noises and human voices records the audio events in the past few months.
[0069] In this embodiment, the time period threshold refers to the time range set for a specific event, usually determined based on the duration of historical events, in order to better identify and process events. For example, if the average duration of "alarm sound" in historical data is 5 seconds, the time period threshold may be set to 4 - 6 seconds.
[0070] In this embodiment, the effective threshold is a criterion obtained by combining the confidence threshold and the time period threshold, and is used to determine whether a specific event is effective. It ensures that an event is considered effective only when both of these conditions are met. For example, if the confidence threshold is 80% and the time period threshold is 5 seconds, then only when the confidence of the event exceeds 80% and the duration is between 4 - 6 seconds will it be considered an effective event.
[0071] The working principle and beneficial effects of the above technical solution are as follows: Obtain industry requirements and conduct category division, determine the confidence threshold of a specific event, collect historical audio data, match historical events of the same category as the specific event, set the time period threshold based on these historical events, combine the confidence threshold and the time period threshold to obtain the effective threshold, so as to improve the accuracy and reliability of event recognition and enable the system to flexibly adapt to different scenarios and requirements.
[0072] Embodiment 6: The embodiment of the present invention provides a method for automatically optimizing the audio effect of an intelligent device. Determine the standard feature corresponding to each effective feature according to the historical standard table, compare the effective feature with the standard feature, and perform a first optimization on the effective audio data based on the comparison result, including: Comprehensively evaluate the effective feature and the standard feature, and determine the difference situation between the effective feature and the standard feature based on the comprehensive evaluation; Obtain the feature to be optimized in the effective feature according to the difference situation, and propose a first optimization direction for the feature to be optimized according to the difference situation, and perform a first optimization on the effective audio data corresponding to the feature to be optimized according to the optimization direction.
[0073] In this embodiment, the comprehensive evaluation is a process of comprehensively analyzing the effective feature and the standard feature to judge the similarity and difference between the two. For example, by comparing the effective feature (such as the frequency feature of background noise) with the standard feature (such as the frequency feature in an ideal environment), evaluate their consistency and differences.
[0074] In this embodiment, the difference situation refers to the specific differences between the effective feature and the standard feature found in the comprehensive evaluation, including differences in terms of values, shapes, distributions, etc. For example, it is found that the frequency peak of the effective feature does not exist in the standard feature, or the frequency range of the effective feature is wider than that of the standard feature.
[0075] In this embodiment, the first optimization direction is an improvement suggestion proposed based on the difference situation, aiming to make the feature to be optimized closer to the standard feature and improve the quality of audio data. For example, if the energy of the effective feature is too high in a certain frequency band, the optimization direction may be to reduce the gain of that frequency band or perform filtering.
[0076] The working principle and beneficial effects of the above technical solution are as follows: comprehensively evaluate the effective feature and the standard feature to determine the difference situation between the two. According to the difference situation, identify the feature to be optimized in the effective feature and propose the corresponding optimization direction. Based on these optimization directions, perform the first optimization on the effective audio data corresponding to the feature to be optimized to improve the audio processing quality, enhance the overall performance of the effective audio data, and improve the efficiency and effect of audio processing.
[0077] Embodiment 7: The embodiment of the present invention provides a processing method for automatically optimizing the audio effect of an intelligent device, which comprehensively evaluates the effective feature and the standard feature, including: ,
[0078] Among them, P represents the comprehensive evaluation of the effective feature set; A represents the weighted overlap degree between the effective feature set and the standard feature set; B represents the clustering overlap degree of the clustering results of the effective feature set and the standard feature set; represents the weight coefficient of the weighted overlap degree in the comprehensive evaluation; represents the weight coefficient of the conditional mutual information in the comprehensive evaluation; represents the weight coefficient of the weighted Shannon diversity index in the comprehensive evaluation; represents the effective feature set; represents the standard feature set; Z represents the existence of a sound source variable; represents the existence of a sound source variable Z of the standard feature set under the condition of; represents the existence of a sound source variable Z of the effective feature set under the condition of; represents the existence of a sound source variable Z under the condition of the standard feature set and the effective feature set of the joint conditional entropy; represents the standard feature set and the effective feature set in the presence of a sound source variable Z under the condition of the conditional mutual information; represents the weighted Shannon diversity index; Represents the weight of the effective features; f Represents the effective features; Represents the i - th cluster in the set of effective features; Represents the information gain of the effective features f ; Represents the i - th cluster in the set of effective features; C represents the set of all clusters; D Represents the total number of resamplings; Represents the d - th resampling obtained cluster ; d Represents the index of the resampling times; i Represents the index of the clustering of the set of effective features; Represents the weight coefficient of the clustering result in the comprehensive evaluation.
[0079] In this embodiment, the effective features are standardized, and the clustering results corresponding to the standardized effective features are determined according to the clustering algorithm.
[0080] The working principle and beneficial effects of the above technical solution are: by comprehensively evaluating the weighted overlap degree, clustering results and other information - theoretic indicators of the set of effective features and the set of standard features, calculating the quality of the effective features, using conditional entropy and conditional mutual information to analyze the influence of the sound source variables on the features, and combining the stability and information gain of the clustering to optimize the set of effective features, so as to improve its matching degree with the standard features, ensure that the effective features are closer to the standard features, and improve the accuracy of audio processing.
[0081] Embodiment 8:
[0082] The embodiment of the present invention provides a processing method for automatically optimizing the audio effect of an intelligent device. The performance of the first optimization result is evaluated, the priority of the industry audio adjustment method is sorted according to the optimization evaluation result, and the optimal adjustment method is obtained. Based on the optimal adjustment method, the second optimization of the first optimization result is performed, including: Evaluating the performance of the first optimization result, collecting the feedback of the user on the first optimization result, analyzing the advantages and disadvantages of the first optimization result based on the performance evaluation and user feedback, and determining the second optimization direction according to the advantages and disadvantages; Based on the second optimization direction, the priority of the industry audio adjustment method is sorted to obtain the optimal adjustment method, and the second optimization of the first optimization result is performed.
[0083] In this embodiment, the performance evaluation is to quantitatively and qualitatively analyze the effect of the first optimization result, usually including indicators such as accuracy, recall rate, F1 - score, processing time, etc. For example, in an audio classification task, evaluating the accuracy of the model may be the proportion of correctly classified audio segments in the total segments.
[0084] In this embodiment, user feedback is to collect the opinions and suggestions of users after using the first optimization result, and understand their satisfaction and usage experience. For example, through questionnaire surveys or user interviews, users are asked about their satisfaction with the audio classification results, whether they think the classification is accurate, and feedback on the usability and functions of the audio processing tool is collected, such as the friendliness of the interface and the practicality of the functions.
[0085] In this embodiment, the analysis of advantages and disadvantages is based on performance evaluation and user feedback to analyze the advantages and disadvantages of the first optimization result. For example, advantages: high model accuracy and the ability to quickly process large-scale audio data; disadvantages: low recognition accuracy for specific audio types (such as speech in a noisy background), and the user interface is not intuitive enough.
[0086] In this embodiment, the second optimization direction is to determine the specific direction of improvement based on the analysis of advantages and disadvantages to improve performance and user experience. For example, for the problem of low recognition accuracy, consider introducing more training data or improving the feature extraction method; for the problem of an unfriendly user interface, plan to redesign the interface to improve usability.
[0087] The working principle and beneficial effects of the above technical solution are as follows: perform performance evaluation on the first optimization result, collect user feedback to analyze its advantages and disadvantages, determine the second optimization direction according to the analysis results, prioritize the industry audio adjustment methods, select the optimal adjustment method, and perform a second optimization on the first optimization result to further improve the audio processing effect, improve resource utilization efficiency, and improve the accuracy of audio processing.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A processing method for automatically optimizing the audio effect of an intelligent device, characterized in that, Including: Step 1: Use a smart device to capture environmental audio data of the surrounding environment, determine the existing sound sources in the environmental audio data based on the sound source-frequency domain feature table, perform noise level analysis based on the existing sound sources, and extract the noise characteristics in the existing sound sources; Step 2: Use a smart device to capture actual audio data, extract the valid audio data in the actual audio data based on the noise characteristics, and extract the valid characteristics of the valid audio data; Step 3: Determine the standard characteristics corresponding to each valid characteristic according to the historical standard table, compare the valid characteristics with the standard characteristics, and perform the first optimization on the valid audio data based on the comparison result; Step 4: Obtain the industry audio adjustment methods, rank the priorities of the industry audio adjustment methods to obtain the optimal adjustment method, and perform the second optimization on the first optimization result based on the optimal adjustment method.
2. The processing method for automatically optimizing the audio effect of the intelligent device according to claim 1, wherein Using a smart device to capture environmental audio data of the surrounding environment, determining the existing sound sources in the environmental audio data based on the sound source-frequency domain feature table, performing noise level analysis based on the existing sound sources, and extracting the noise characteristics in the existing sound sources, including: Perform frequency domain analysis on the environmental audio data to determine the environmental spectrum characteristics, analyze the environmental spectrum characteristics, and perform sound source stripping on the environmental audio data according to the feature analysis result and the sound source-frequency domain feature table to obtain the existing sound sources; Classify the noise levels of the existing sound sources according to the noise duration and noise band, and extract the noise characteristics of each existing sound source based on the noise level.
3. The processing method for automatically optimizing the audio effect of the intelligent device according to claim 2, wherein Classifying the levels according to the noise duration and noise band, including: Obtain the actual application scenario of the smart device, and classify the first level of the noise duration based on the actual application scenario; Determine the common bands according to the actual application scenario, and classify the second level of the noise band based on the common bands; Comprehensively determine the noise level of each existing sound source based on the first level and the second level, and then extract the noise characteristics of each existing sound source according to the noise characteristics.
4. The processing method for automatically optimizing the audio effect of the intelligent device according to claim 1, characterized in that, Using a smart device to capture actual audio data, extracting the valid audio data in the actual audio data based on the noise characteristics, and extracting the valid characteristics of the valid audio data, including: Compare the noise characteristics with the actual audio data, determine the noise part and the remaining audio part according to the comparison result, and perform segmentation processing on the actual audio data according to the noise characteristics to obtain the segmentation processing result; Use the audio event detection algorithm to identify specific events in the segmentation processing result, determine the event category and confidence level of each time period, set an effective threshold according to the confidence level and event category, perform an effective judgment on the specific events, and comprehensively obtain all valid events to obtain the valid audio data; Extract the characteristics of each valid event, integrate the extracted characteristics to form an event vector, determine the task requirements based on the automatic optimization of the audio effect, select the event vector according to the task requirements, and then obtain the valid characteristics of the valid audio data.
5. The processing method for automatically optimizing the audio effect of the intelligent device according to claim 4, wherein, Using the audio event detection algorithm to identify specific events in the segmentation processing result, determine the event category and confidence level of each time period, and setting an effective threshold according to the confidence level and event category, including: Obtain industry requirements, classify the types of the industry requirements, and determine the confidence threshold corresponding to a specific event in combination with the classification results. Collect historical audio data, match historical events of the same category as the specific event from the historical audio data, and set a time period threshold corresponding to the specific event based on the historical events. Derive an effective threshold by synthesizing the confidence threshold and the time period threshold.
6. The processing method for automatically optimizing the audio effect of the intelligent device according to claim 1, wherein Determine the standard feature corresponding to each effective feature according to the historical standard table, compare the effective feature with the standard feature, and perform a first optimization on the effective audio data based on the comparison result, including: Comprehensively evaluate the effective feature and the standard feature, and determine the difference between the effective feature and the standard feature based on the comprehensive evaluation. Derive the features to be optimized in the effective features according to the difference situation, propose a first optimization direction for the features to be optimized for the difference situation, and perform a first optimization on the effective audio data corresponding to the features to be optimized according to the optimization direction.
7. The processing method for automatically optimizing the audio effect of an intelligent device according to claim 6, characterized in that, The comprehensive evaluation of the effective feature and the standard feature includes: , where P represents the comprehensive evaluation of the effective feature set; A represents the weighted overlap degree between the effective feature set and the standard feature set; B represents the clustering overlap degree of the clustering results of the effective feature set and the standard feature set; represents the weight coefficient of the weighted overlap degree in the comprehensive evaluation; represents the weight coefficient of the conditional mutual information in the comprehensive evaluation; represents the weight coefficient of the weighted Shannon diversity index in the comprehensive evaluation; represents the effective feature set; represents the standard feature set; Z represents the existence of a sound source variable; represents the existence of a sound source variable Z under the standard feature set of the conditional entropy; represents the existence of a sound source variable Z under the effective feature set of the conditional entropy; represents the existence of a sound source variable Z under the condition of the standard feature set and the effective feature set of the joint conditional entropy; represents the standard feature set and the effective feature set in the presence of a sound source variable Z under the condition of the conditional mutual information; represents the weighted Shannon diversity index; represents the weight of the effective feature; f represents the effective feature; represents the i th clustering in the effective feature set; represents the effective feature f of the information gain; represents the i th clustering stability of the effective feature set; C represents the set of all clusterings; D represents the total number of resamplings; represents the d th clustering obtained by resampling ; d represents the index of the number of resamplings; i represents the index of the clustering of the effective feature set; represents the weight coefficient of the clustering result in the comprehensive evaluation.
8. The processing method for automatically optimizing the audio effect of the intelligent device according to claim 3, wherein, Perform a performance evaluation on the first optimization result, prioritize the industry audio adjustment methods according to the optimization evaluation result to obtain the optimal adjustment method, and perform a second optimization on the first optimization result based on the optimal adjustment method, including: Perform a performance evaluation on the first optimization result, collect the feedback of users on the first optimization result, analyze the advantages and disadvantages of the first optimization result based on the performance evaluation and user feedback, and determine the second optimization direction according to the advantages and disadvantages. Prioritize the industry audio adjustment methods based on the second optimization direction to obtain the optimal adjustment method, and perform a second optimization on the first optimization result.
Citation Information
Patent Citations
Audio optimization method and device
CN109087659A
Monitoring audio processing method and device, storage medium and electronic equipment
CN113409800A
Audio processing method and device, audio playing equipment and storage medium
CN118042348A
Sound environment adaptive USB audio optimization method and system
CN118230767A
Noise reduction optimization method for noise reduction type MEMS microphone for smart home
CN118314917A