Data identification method and system based on neural network model and application

By intercepting the target audio period in the audio recognition method and setting multiple sets of denoising indicators, an audio neural network model is created, which solves the problems of low audio denoising efficiency and low degree of automation in the existing technology, and achieves efficient audio denoising and target audio recognition.

CN120452432AActive Publication Date: 2025-08-08XIAN FULIYE MICROELECTRONICS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510826225.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-08-08
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The existing data recognition methods cannot effectively intercept the sound period of the target audio for targeted denoising, and cannot automatically identify and remove audio noise through the audio neural network model, resulting in low audio denoising efficiency and low degree of automation.

Method used

By obtaining the audio to be identified and intercepting the target audio period, setting multiple sets of denoising indicators for audio source marking, creating an audio neural network model, and using this model to identify and remove audio noise.

Benefits of technology

It improves the targeted and automated level of audio denoising, significantly improves the audio denoising efficiency and target audio recognition rate in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452432A_ABST
    Figure CN120452432A_ABST
Patent Text Reader

Abstract

The invention discloses a data recognition method and system based on a neural network model and application, relates to the field of audio processing, and solves the problem that an existing data recognition method is poor in audio data recognition processing, and the method comprises the steps: S1, obtaining a to-be-recognized audio, and intercepting a sound time period in which a target audio exists in the to-be-recognized audio, s2, acquiring multiple groups of historical target audios, performing sound source marking on the historical target audios by setting a denoising index to obtain audio sound source marking data, and creating an audio neural network model according to the audio sound source marking data, and S3, carrying out audio noise identification and removal on each target audio interception time period according to the audio neural network model. The audio denoising method can improve the audio denoising efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of audio processing and relates to artificial feedback reinforcement learning technology, specifically a data recognition method, system and application based on a neural network model. Background Art

[0002] Existing data recognition methods have the following specific defects when processing audio data: Existing data recognition methods are unable to intercept the sound period of the target audio in the audio to be recognized and perform targeted denoising on the intercepted sound period, resulting in low audio denoising efficiency; Existing data recognition methods are unable to create an audio neural network model by labeling the sound source of historical target audio, nor are they able to identify and remove audio noise for each target audio interception period based on the audio neural network model, resulting in a low degree of automation in the audio denoising process.

[0003] To this end, we propose a data recognition method, system and application based on a neural network model. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a data recognition method, system and application based on a neural network model. The present invention aims to improve the pertinence and accuracy of the audio denoising process.

[0005] In order to achieve the above object, the present invention adopts the following technical solution: a data recognition method based on a neural network model, comprising the following specific steps: Step S1: Acquire the audio to be recognized, and intercept the sound period of the target audio in the audio to be recognized to obtain the target audio period collection data; Step S2: Acquire multiple sets of historical target audio, perform audio source tagging on the historical target audio by setting multiple sets of denoising indicators to obtain audio source tag data, and create an audio neural network model based on the audio source tag data; Step S3: Identify and remove audio noise for each target audio capture period based on the audio neural network model.

[0006] Furthermore, the step S1 further includes the following specific steps: Step S11: During the process of real-time audio monitoring of the audio recognition location by the audio acquisition device, the time point at which the audio device starts working is marked as the audio start time point of the period, the time point corresponding to the current moment is marked as the audio end time point, and the ambient recordings generated by the audio recognition location at the audio start time point and the audio end time point are marked as audio to be recognized; Step S12: Segment the audio to be recognized into a plurality of time-continuous audio short-time frames, and select a sample audio short-time frame from the obtained plurality of audio short-time frames; Step S13: performing target audio energy analysis on the sample audio short-time frame, and obtaining the target audio energy ratio corresponding to the sample audio short-time frame according to the analysis result; Step S14: Obtaining a plurality of historical audio short-time frames in which target audio has been identified, obtaining the target audio energy ratio corresponding to each historical audio short-time frame, and comparing the values of the obtained multiple target audio energy ratios. The target audio energy ratio with the smallest value is marked as the target audio baseline energy ratio. Step S15: acquiring a target audio energy ratio for each audio short-time frame, comparing the obtained target audio energy ratio with the target audio reference energy ratio, and classifying the audio short-time frames into valid audio short-time frames and invalid audio short-time frames based on the comparison results; Step S16: intercepting the sound period corresponding to the valid audio short-time frame in the audio to be recognized to obtain multiple target audio interception periods; Step S17: defining the acquired target audio interception period as target audio period acquisition data.

[0007] Furthermore, the step S13 further includes the following specific steps: Divide the sample audio short time frame into a number of audio signal sampling points, and name the divided audio signal sampling points as X1 signal sampling point to Xa signal sampling point in the order of division time; The target audio signals corresponding to the target audio at the X1 signal sampling point to the Xa signal sampling point are respectively obtained to obtain the strength values of the target audio signals, and the strength values of the target signals X1 to Xa are obtained; At the X1 signal sampling point, obtain the strength values of any ambient audio signals other than the target audio signal, compare the obtained multiple ambient audio signal strength values, and mark the ambient audio signal strength value with the largest value as the X1 ambient signal strength value; Repeat the process of obtaining the X1 environmental signal strength value, and obtain the environmental signal strength values corresponding to the X2 signal sampling point to the Xa signal sampling point, and obtain the X2 environmental signal strength value to the Xa environmental signal strength value; Calculate the ratio of the X1 target signal strength value to the X1 ambient signal strength value to obtain the X1 target audio intensity ratio. Calculate the ratio of the X2 target signal strength value to the X2 ambient signal strength value to obtain the X2 target audio intensity ratio. Similarly, calculate the ratio of the Xa target signal strength value to the Xa ambient signal strength value to obtain the Xa target audio intensity ratio. The target audio intensity ratio X1 to the target audio intensity ratio Xa are averaged to obtain the target audio energy ratio corresponding to the sample audio short-time frame.

[0008] Furthermore, the step S15 further includes the following specific steps: If the target audio energy ratio is greater than or equal to the target audio reference energy ratio, the corresponding audio short-time frame is classified as a valid audio short-time frame; If the target audio energy ratio is less than the target audio reference energy ratio, the corresponding audio short-time frame is classified as an invalid audio short-time frame.

[0009] Furthermore, the step S2 further includes the following specific steps: Step S21: in the process of performing audio denoising on the target audio interception period, setting a plurality of different types of audio denoising indicators, and marking the set plurality of different types of audio denoising indicators as Z1 denoising indicators to Zb denoising indicators respectively; Step S22: Analyze the Z1 denoising index to the Zb denoising index based on the historical target audio to obtain the Z1 denoising index benchmark interval to the Zb denoising index benchmark interval; Step S23: Obtain several historical audio interception periods in the audio recognition field. In the historical audio interception periods, mark the sound sources whose Z1 denoising index to Zb denoising index are all in the Z1 denoising index benchmark interval to the Zb denoising index benchmark interval as the first type of sound source, and mark the sound source whose denoising index is not in the corresponding denoising index interval as the second type of sound source, to obtain audio source marking data.

[0010] Further, step S24: dividing the audio source labeled data into a denoised frequency training set and a denoised frequency test set according to the audio training test ratio; Step S25: Create an audio recognition model using an existing artificial intelligence platform, and train the audio recognition model using the denoised frequency training set until the audio recognition model is trained once for each historical audio interception period in the denoised frequency training set; Step S26: Use the denoised frequency test set to test the audio recognition model and obtain the recognition accuracy. When the recognition accuracy is greater than or equal to the target recognition accuracy, the audio recognition model training is completed and the audio neural network model is obtained. When the recognition accuracy is less than the target recognition accuracy, continue to use the denoised frequency training set to train the audio recognition model until the recognition accuracy is greater than or equal to the target recognition accuracy.

[0011] Furthermore, the step S22 further includes the following specific steps: Acquire multiple groups of historical target audios that have completed denoising, select a sample target audio from the multiple historical target audios, obtain the Z1 denoising index value corresponding to each time point in the sample target audio, obtain multiple Z1 denoising index values, compare the multiple Z1 denoising index values obtained, obtain a value range consisting of a maximum Z1 denoising index value and a minimum Z1 denoising index value, and obtain a Z1 denoising index value interval corresponding to the sample target audio; Repeat the process of obtaining the Z1 denoising index value interval corresponding to the sample target audio, obtain the Z1 denoising index value interval corresponding to each historical target audio, obtain multiple Z1 denoising index value intervals, and perform interval union on the multiple Z1 denoising index value intervals to obtain the Z1 denoising index benchmark interval; Repeat the process of obtaining the Z1 denoising index benchmark interval to obtain the Z2 denoising index benchmark interval to the Zb denoising index benchmark interval respectively.

[0012] Furthermore, the step S3 further includes the following specific steps: Obtain an audio neural network model, obtain target audio time period collection data, and obtain multiple target audio interception time periods based on the target audio time period collection data; An audio neural network model is used to identify the second type of sound source in the target audio interception period, and the second type of sound source in each target audio interception period is removed as audio noise.

[0013] A data recognition system based on a neural network model, comprising: Audio acquisition module: obtains the audio to be recognized, and intercepts the sound period of the target audio in the audio to be recognized to obtain the target audio period acquisition data; Model creation module: obtains multiple sets of historical target audio, labels the historical target audio by setting multiple sets of denoising indicators, obtains audio source labeling data, and creates an audio neural network model based on the audio source labeling data; Audio denoising module: identifies and removes audio noise for each target audio interception period based on the audio neural network model; Furthermore, the data recognition method can be applied in audio denoising.

[0014] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. The present invention intercepts the sound period where the target audio exists in the audio to be identified, and performs targeted denoising on the intercepted sound period, thereby improving the audio denoising efficiency.

[0015] 2. The present invention creates an audio neural network model by marking the sound source of historical target audio, and identifies and removes audio noise for each target audio interception period based on the audio neural network model, which can improve the degree of automation of the audio denoising process. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.

[0017] Figure 1 is a block diagram of the overall system of the present invention; Figure 2 It is a diagram of the implementation steps of the present invention. DETAILED DESCRIPTION

[0018] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] First, see Figure 1 The audio neural network model in the present invention involves artificial feedback reinforcement learning. The present invention provides a technical solution: a data recognition method based on a neural network model, comprising the following specific steps: Step S1: Acquire the audio to be recognized, and intercept the sound period of the target audio in the audio to be recognized to obtain the target audio period collection data; The step S1 further includes the following specific steps: Step S11: During the process of real-time audio monitoring of the audio recognition location by the audio acquisition device, the time point at which the audio device starts working is marked as the audio start time point of the period, the time point corresponding to the current moment is marked as the audio end time point, and the ambient recordings generated by the audio recognition location at the audio start time point and the audio end time point are marked as audio to be recognized; Step S12: Segment the audio to be recognized into a plurality of time-continuous audio short-time frames, and select a sample audio short-time frame from the obtained plurality of audio short-time frames; Step S13: performing target audio energy analysis on the sample audio short-time frame, and obtaining the target audio energy ratio corresponding to the sample audio short-time frame according to the analysis result; The step S13 further includes the following specific steps: Divide the sample audio short time frame into a number of audio signal sampling points, and name the divided audio signal sampling points as X1 signal sampling point to Xa signal sampling point in the order of division time; The target audio signals corresponding to the target audio at the X1 signal sampling point to the Xa signal sampling point are respectively obtained to obtain the strength values of the target audio signals, and the strength values of the target signals X1 to Xa are obtained; At the X1 signal sampling point, obtain the strength values of any ambient audio signals other than the target audio signal, compare the obtained multiple ambient audio signal strength values, and mark the ambient audio signal strength value with the largest value as the X1 ambient signal strength value; Repeat the process of obtaining the X1 environmental signal strength value, and obtain the environmental signal strength values corresponding to the X2 signal sampling point to the Xa signal sampling point, and obtain the X2 environmental signal strength value to the Xa environmental signal strength value; Calculate the ratio of the X1 target signal strength value to the X1 ambient signal strength value to obtain the X1 target audio intensity ratio. Calculate the ratio of the X2 target signal strength value to the X2 ambient signal strength value to obtain the X2 target audio intensity ratio. Similarly, calculate the ratio of the Xa target signal strength value to the Xa ambient signal strength value to obtain the Xa target audio intensity ratio. Calculate the average of the target audio intensity ratio X1 to the target audio intensity ratio Xa to obtain the target audio energy ratio corresponding to the sample audio short-time frame; Step S14: Obtaining a plurality of historical audio short-time frames in which target audio has been identified, obtaining the target audio energy ratio corresponding to each historical audio short-time frame, and comparing the values of the obtained multiple target audio energy ratios. The target audio energy ratio with the smallest value is marked as the target audio baseline energy ratio. Step S15: acquiring a target audio energy ratio for each audio short-time frame, comparing the obtained target audio energy ratio with the target audio reference energy ratio, and classifying the audio short-time frames into valid audio short-time frames and invalid audio short-time frames based on the comparison results; The step S15 further includes the following specific steps: If the target audio energy ratio is greater than or equal to the target audio reference energy ratio, the corresponding audio short-time frame is classified as a valid audio short-time frame; If the target audio energy ratio is less than the target audio reference energy ratio, the corresponding audio short-time frame is classified as an invalid audio short-time frame; Step S16: intercepting the sound period corresponding to the valid audio short-time frame in the audio to be recognized to obtain multiple target audio interception periods; Step S17: defining the acquired target audio interception period as target audio period collection data; The beneficial effects corresponding to the above step S1 are as follows: Real-time audio monitoring and a dynamic time period marking mechanism ensure the integrity and timeliness of audio acquisition, providing raw data for subsequent analysis. A short-frame segmentation strategy is used to break audio into local analysis units, preserving temporal continuity while facilitating transient feature capture, effectively balancing computational efficiency and feature resolution. The innovative introduction of the target audio energy ratio metric quantifies the intensity relationship between the target signal and ambient noise, constructing a characteristic parameter with strong anti-interference capabilities. Compared with the traditional energy threshold method, it significantly improves the target audio recognition rate in complex environments. The adaptive baseline energy ratio generated based on historical data enables the system to adapt to the environment and dynamically distinguish between valid signals and background noise, avoiding the misjudgment problem of fixed thresholds in changing scenarios. By comparing energy ratios, the system accurately captures effective time periods, removing silent or noise-dominated periods while preserving the complete temporal distribution of the target audio. This provides high-quality input data for subsequent audio recognition tasks while reducing the system resource consumption of invalid data. This solution, through the organic combination of signal processing and data screening, establishes a complete technical chain from raw audio acquisition to target time period extraction, significantly improving target detection accuracy and system efficiency in complex audio scenarios.

[0020] Step S2: Acquire multiple sets of historical target audio, perform audio source tagging on the historical target audio by setting multiple sets of denoising indicators to obtain audio source tag data, and create an audio neural network model based on the audio source tag data; The step S2 further includes the following specific steps: Step S21: in the process of performing audio denoising on the target audio interception period, setting a plurality of different types of audio denoising indicators, and marking the set plurality of different types of audio denoising indicators as Z1 denoising indicators to Zb denoising indicators respectively; Step S22: Analyze the Z1 denoising index to the Zb denoising index based on the historical target audio to obtain the Z1 denoising index benchmark interval to the Zb denoising index benchmark interval; The step S22 further includes the following specific steps: Acquire multiple groups of historical target audios that have completed denoising, select a sample target audio from the multiple historical target audios, obtain the Z1 denoising index value corresponding to each time point in the sample target audio, obtain multiple Z1 denoising index values, compare the multiple Z1 denoising index values obtained, obtain a value range consisting of a maximum Z1 denoising index value and a minimum Z1 denoising index value, and obtain a Z1 denoising index value interval corresponding to the sample target audio; Repeat the process of obtaining the Z1 denoising index value interval corresponding to the sample target audio, obtain the Z1 denoising index value interval corresponding to each historical target audio, obtain multiple Z1 denoising index value intervals, and perform interval union on the multiple Z1 denoising index value intervals to obtain the Z1 denoising index benchmark interval; Repeat the process of obtaining the Z1 denoising index benchmark interval to obtain the Z2 denoising index benchmark interval to the Zb denoising index benchmark interval respectively; Step S23: Acquire several historical audio interception periods in the audio recognition field. In these historical audio interception periods, mark the sound sources whose Z1 denoising index to Zb denoising index are all within the Z1 denoising index reference interval to the Zb denoising index reference interval as first-type sound sources, and mark the sound sources whose denoising index is not within the corresponding denoising index interval as second-type sound sources, thereby obtaining audio source label data. Step S24: dividing the audio source labeled data into a denoised frequency training set and a denoised frequency test set according to the audio training test ratio; Step S25: Create an audio recognition model using an existing artificial intelligence platform, and train the audio recognition model using the denoised frequency training set until the audio recognition model is trained once for each historical audio interception period in the denoised frequency training set; Step S26: Use the denoised frequency test set to test the audio recognition model and obtain the recognition accuracy. When the recognition accuracy is greater than or equal to the target recognition accuracy, the audio recognition model training is completed and the audio neural network model is obtained. When the recognition accuracy is less than the target recognition accuracy, continue to use the denoised frequency training set to train the audio recognition model until the recognition accuracy is greater than or equal to the target recognition accuracy.

[0021] The beneficial effects corresponding to step S2 are as follows: By constructing multiple sets of denoising index benchmark intervals, a quantitative assessment of audio quality is achieved. Compared with the single index threshold method, this solution can more comprehensively reflect the multi-dimensional characteristics of the audio signal. The dynamic benchmark intervals generated based on historical data enable the system to adapt to the environment and automatically distinguish between valid sound sources and noise interference, avoiding the misjudgment problem of fixed thresholds in complex scenarios. By creating training and test sets from labeled sound source data, a closed-loop optimization system was constructed, enabling the model to continuously optimize its feature extraction capabilities during iterative training. The audio neural network model, which adopts a deep learning architecture, can automatically learn the complex mapping relationship between sound source features and denoising indicators, and achieve a better balance between noise suppression and target sound source retention compared to traditional machine learning models. Finally, through multi-indicator collaborative analysis and neural network modeling, the target sound source recognition rate in complex audio scenarios was significantly improved, while the interference of environmental noise on the recognition results was reduced.

[0022] Step S3: Identify and remove audio noise for each target audio interception period according to the audio neural network model; The step S3 further includes the following specific steps: Obtain an audio neural network model, obtain target audio time period collection data, and obtain multiple target audio interception time periods based on the target audio time period collection data; An audio neural network model is used to identify the second type of sound source in the target audio interception period, and the second type of sound source in each target audio interception period is removed as audio noise.

[0023] In this application, if a corresponding calculation formula appears, the above calculation formula is dimensionless and its numerical calculation is performed. The weight coefficient, proportional coefficient and other coefficients in the formula are set to a result value obtained by quantifying each parameter. Regarding the size of the weight coefficient and the proportional coefficient, as long as it does not affect the proportional relationship between the parameter and the result value, it is acceptable.

[0024] Second, see Figure 2 Based on another concept of the same invention, a data recognition system based on a neural network model is proposed, comprising an audio acquisition module, a model creation module, an audio denoising module, and a server, wherein the audio acquisition module, the model creation module, and the audio denoising module are respectively connected to the server, and the server controls the audio acquisition module, the model creation module, and the audio denoising module respectively; The audio acquisition module obtains the audio to be recognized and intercepts the sound period of the target audio in the audio to be recognized to obtain the target audio period acquisition data; The details are as follows: During the process of real-time audio monitoring of the audio recognition location by the audio acquisition device, the time point when the audio device starts working is marked as the audio start time point of the period, the time point corresponding to the current moment is marked as the audio end time point, and the ambient recordings generated by the audio recognition location at the audio start time point and the audio end time point are marked as audio to be recognized; It should be noted here that: In this application, the audio to be recognized involved here can be updated in real time according to the change of the current time value; It should be noted here that: In this application, the audio recognition places involved here are public places that specifically comply with the audio signal collection rules and do not involve personal privacy issues.

[0025] Segmenting the audio to be recognized into a plurality of time-continuous audio short-time frames, and selecting a sample audio short-time frame from the obtained plurality of audio short-time frames; It should be noted here that: In this application, the frame duration corresponding to the sample audio short-time frame involved here is specifically 20ms.

[0026] Performing target audio energy analysis on the sample audio short-time frame, and obtaining a target audio energy ratio corresponding to the sample audio short-time frame according to the analysis result; The details are as follows: Divide the sample audio short time frame into a number of audio signal sampling points, and name the divided audio signal sampling points as X1 signal sampling point to Xa signal sampling point in the order of division time; It should be noted here that: In this application, X mentioned here is the sign symbol corresponding to the signal sampling point, and a mentioned here is the quantity value corresponding to the signal sampling point; The target audio signals corresponding to the target audio at the X1 signal sampling point to the Xa signal sampling point are respectively obtained to obtain the strength values of the target audio signals, and the strength values of the target signals X1 to Xa are obtained; It should be noted here that: In this application, the target audio signal referred to herein is specifically a human voice signal; At the X1 signal sampling point, obtain the strength values of any ambient audio signals other than the target audio signal, compare the obtained multiple ambient audio signal strength values, and mark the ambient audio signal strength value with the largest value as the X1 ambient signal strength value; Repeat the process of obtaining the X1 environmental signal strength value, and obtain the environmental signal strength values corresponding to the X2 signal sampling point to the Xa signal sampling point, and obtain the X2 environmental signal strength value to the Xa environmental signal strength value; It should be noted here that: In the present application, the environmental audio signals involved herein include but are not limited to natural environmental sounds, man-made environmental sounds, and biological noises generated by creatures other than humans.

[0027] Calculate the ratio of the X1 target signal strength value to the X1 ambient signal strength value to obtain the X1 target audio intensity ratio. Calculate the ratio of the X2 target signal strength value to the X2 ambient signal strength value to obtain the X2 target audio intensity ratio. Similarly, calculate the ratio of the Xa target signal strength value to the Xa ambient signal strength value to obtain the Xa target audio intensity ratio. Calculate the average of the target audio intensity ratio X1 to the target audio intensity ratio Xa to obtain the target audio energy ratio corresponding to the sample audio short-time frame; Obtaining several historical audio short-time frames in which target audio has been identified, obtaining the target audio energy ratio corresponding to each historical audio short-time frame, and comparing the numerical values of the obtained multiple target audio energy ratios, marking the target audio energy ratio with the smallest numerical value as the target audio baseline energy ratio; Repeat the process of obtaining the target audio energy ratio corresponding to the sample audio short-time frame, obtain the target audio energy ratio for each audio short-time frame, and compare the obtained target audio energy ratio with the target audio reference energy ratio. According to the comparison result, the audio short-time frame is divided into a valid audio short-time frame and an invalid audio short-time frame; If the target audio energy ratio is greater than or equal to the target audio reference energy ratio, the corresponding audio short-time frame is classified as a valid audio short-time frame; If the target audio energy ratio is less than the target audio reference energy ratio, the corresponding audio short-time frame is classified as an invalid audio short-time frame; In the audio to be recognized, the sound period corresponding to the valid audio short-time frame is intercepted to obtain multiple target audio interception periods; The obtained target audio interception period is defined as target audio period acquisition data; The model creation module obtains multiple sets of historical target audio, tags the historical target audio by setting multiple sets of denoising indicators, obtains audio source tag data, and creates an audio neural network model based on the audio source tag data; Obtain multiple sets of historical target audio that have completed denoising, and create an audio neural network model based on the historical target audio; It should be noted here that: In this application, the historical target audio referred to here is the target audio that has been denoised.

[0028] The details are as follows: In the process of performing audio denoising on the target audio interception period, several different types of audio denoising indicators are set, and the set multiple different types of audio denoising indicators are marked as Z1 denoising indicators to Zb denoising indicators respectively; Analyze the Z1 denoising index to the Zb denoising index based on the historical target audio to obtain the Z1 denoising index benchmark interval to the Zb denoising index benchmark interval; The details are as follows: It should be noted here that: In this application, Z mentioned here is the sign symbol corresponding to the denoising index, and a mentioned here is the numerical value corresponding to the denoising index; In the present application, the Z1 denoising index involved here may be the Mel-frequency cepstral coefficient, the Z2 denoising index involved here may be the spectrum centroid, and the Z3 denoising index involved here may be the spectrum zero-crossing rate.

[0029] A sample target audio is selected from the multiple historical target audios obtained, and a Z1 denoising index value corresponding to each time point in the sample target audio is obtained to obtain multiple Z1 denoising index values. The multiple Z1 denoising index values obtained are compared, and a value range consisting of a maximum Z1 denoising index value and a minimum Z1 denoising index value is obtained to obtain a Z1 denoising index value interval corresponding to the sample target audio; Repeat the process of obtaining the Z1 denoising index value interval corresponding to the sample target audio, obtain the Z1 denoising index value interval corresponding to each historical target audio, obtain multiple Z1 denoising index value intervals, and perform interval union on the multiple Z1 denoising index value intervals to obtain the Z1 denoising index benchmark interval; Repeat the process of obtaining the Z1 denoising index benchmark interval to obtain the Z2 denoising index benchmark interval to the Zb denoising index benchmark interval respectively; Acquire several historical audio interception periods in the audio recognition field. In the historical audio interception periods, mark the sound sources whose Z1 denoising index to Zb denoising index are all within the Z1 denoising index reference interval to the Zb denoising index reference interval as the first type of sound source, and mark the sound sources whose denoising index is not within the corresponding denoising index interval as the second type of sound source, thereby obtaining audio source labeling data; It should be noted here that: The first type of sound source involved here is specifically the sound source corresponding to the historical target audio, specifically a human voice source, and the second type of sound source involved here is specifically a non-human voice source.

[0030] The audio source labeled data is divided into a denoised frequency training set and a denoised frequency test set according to the audio training and testing ratio; It should be noted here that: In this application, the audio training test ratio is specifically set to 7:3, that is, the ratio of the number of medical denoised frequencies in the denoised frequency training set and the denoised frequency test set is 7:3; Create an audio recognition model using an existing artificial intelligence platform, and train the audio recognition model using a denoised frequency training set. Train the audio recognition model once for each historical audio interception period in the denoised frequency training set. The audio recognition model is tested using the denoised frequency test set, and the recognition accuracy is obtained. When the recognition accuracy is greater than or equal to the target recognition accuracy, the audio recognition model training is completed and the audio neural network model is obtained. When the recognition accuracy is less than the target recognition accuracy, the audio recognition model is continued to be trained using the denoised frequency training set until the recognition accuracy is greater than or equal to the target recognition accuracy.

[0031] It should be noted here that: The target recognition accuracy involved here is specifically set to 95% in this application.

[0032] The audio denoising module identifies and removes audio noise for each target audio interception period based on the audio neural network model; The details are as follows: Obtain an audio neural network model, obtain target audio time period collection data, and obtain multiple target audio interception time periods based on the target audio time period collection data; An audio neural network model is used to identify the second type of sound source in the target audio interception period, and the second type of sound source in each target audio interception period is removed as audio noise.

[0033] In a third aspect, the present invention provides an application of any of the aforementioned real-time data recognition methods in audio denoising.

[0034] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A data recognition method based on a neural network model, characterized in that: include: Step S1: Acquire the audio to be recognized, and intercept the sound period of the target audio in the audio to be recognized to obtain the target audio period collection data; Step S2: Acquire multiple sets of historical target audio, perform audio source tagging on the historical target audio by setting a denoising index to obtain audio source tag data, and create an audio neural network model based on the audio source tag data; Step S3: Identify and remove audio noise for each target audio capture period based on the audio neural network model.

2. The data recognition method based on a neural network model according to claim 1, characterized in that: The step S1 further includes the following specific steps: Step S11: During the process of real-time audio monitoring of the audio recognition location by the audio acquisition device, the time point at which the audio device starts working is marked as the audio start time point of the period, the time point corresponding to the current moment is marked as the audio end time point, and the ambient recordings generated by the audio recognition location at the audio start time point and the audio end time point are marked as audio to be recognized; Step S12: Segment the audio to be recognized into a plurality of time-continuous audio short-time frames, and select a sample audio short-time frame from the obtained plurality of audio short-time frames; Step S13: performing target audio energy analysis on the sample audio short-time frame, and obtaining the target audio energy ratio corresponding to the sample audio short-time frame according to the analysis result; Step S14: Obtaining a plurality of historical audio short-time frames in which target audio has been identified, obtaining the target audio energy ratio corresponding to each historical audio short-time frame, and comparing the values of the obtained multiple target audio energy ratios. The target audio energy ratio with the smallest value is marked as the target audio baseline energy ratio. Step S15: acquiring a target audio energy ratio for each audio short-time frame, comparing the obtained target audio energy ratio with the target audio reference energy ratio, and classifying the audio short-time frames into valid audio short-time frames and invalid audio short-time frames based on the comparison results; Step S16: intercepting the sound period corresponding to the valid audio short-time frame in the audio to be recognized to obtain multiple target audio interception periods; Step S17: defining the acquired target audio interception period as target audio period acquisition data.

3. The data recognition method based on a neural network model according to claim 2, characterized in that: The step S13 further includes the following specific steps: Divide the sample audio short time frame into a number of audio signal sampling points, and name the divided audio signal sampling points as X1 signal sampling point to Xa signal sampling point in the order of division time; The target audio signals corresponding to the target audio at the X1 signal sampling point to the Xa signal sampling point are respectively obtained to obtain the strength values of the target audio signals, and the strength values of the target signals X1 to Xa are obtained; At the X1 signal sampling point, obtain the strength values of any ambient audio signals other than the target audio signal, compare the obtained multiple ambient audio signal strength values, and mark the ambient audio signal strength value with the largest value as the X1 ambient signal strength value; Obtain the environmental signal strength values corresponding to the X2 signal sampling point to the Xa signal sampling point respectively, and obtain the X2 environmental signal strength value to the Xa environmental signal strength value; Calculate the ratio of the X1 target signal strength value to the X1 ambient signal strength value to obtain the X1 target audio strength ratio. Calculate the ratio of the Xa target signal strength value to the Xa ambient signal strength value to obtain the Xa target audio strength ratio. The target audio intensity ratio X1 to the target audio intensity ratio Xa are averaged to obtain the target audio energy ratio corresponding to the sample audio short-time frame.

4. The data recognition method based on a neural network model according to claim 2, characterized in that: The step S15 further includes the following specific steps: If the target audio energy ratio is greater than or equal to the target audio reference energy ratio, the corresponding audio short-time frame is classified as a valid audio short-time frame; If the target audio energy ratio is less than the target audio reference energy ratio, the corresponding audio short-time frame is classified as an invalid audio short-time frame.

5. The data recognition method based on a neural network model according to claim 1, characterized in that: The step S2 includes the following specific steps: Step S21: during the audio denoising process for the target audio interception period, setting the Z1 denoising index to the Zb denoising index; Step S22: Analyze the Z1 denoising index to the Zb denoising index based on the historical target audio to obtain the Z1 denoising index benchmark interval to the Zb denoising index benchmark interval; Step S23: Obtain several historical audio interception periods in the audio recognition field. In the historical audio interception periods, mark the sound sources whose denoising indicators are all in the corresponding denoising indicator benchmark interval as the first type of sound source, and mark the sound sources whose any denoising indicator is not in the corresponding denoising indicator interval as the second type of sound source to obtain audio source marking data.

6. The data recognition method based on a neural network model according to claim 1, characterized in that: The step S2 further includes the following specific steps: Step S24: dividing the audio source labeled data into a denoised frequency training set and a denoised frequency test set according to the audio training test ratio; Step S25: Create an audio recognition model using an existing artificial intelligence platform, and train the audio recognition model using the denoised frequency training set until the audio recognition model is trained once for each historical audio interception period in the denoised frequency training set; Step S26: Use the denoised frequency test set to test the audio recognition model and obtain the recognition accuracy. When the recognition accuracy is greater than or equal to the target recognition accuracy, the audio recognition model training is completed and the audio neural network model is obtained. When the recognition accuracy is less than the target recognition accuracy, continue to use the denoised frequency training set to train the audio recognition model until the recognition accuracy is greater than or equal to the target recognition accuracy.

7. The data recognition method based on a neural network model according to claim 5, characterized in that: The step S22 further includes the following specific steps: Acquire multiple groups of historical target audios that have completed denoising, select a sample target audio from the multiple historical target audios, obtain the Z1 denoising index value corresponding to each time point in the sample target audio, obtain multiple Z1 denoising index values, compare the multiple Z1 denoising index values obtained, obtain a value range consisting of the maximum Z1 denoising index value and the minimum Z1 denoising index value, and obtain the Z1 denoising index value range corresponding to the sample target audio; Obtaining the Z1 denoising index value interval corresponding to each historical target audio respectively to obtain multiple Z1 denoising index value intervals, and performing interval union on the multiple Z1 denoising index value intervals to obtain a Z1 denoising index benchmark interval; Repeat the process of obtaining the Z1 denoising index benchmark interval to obtain the Z2 denoising index benchmark interval to the Zb denoising index benchmark interval respectively.

8. The data recognition method based on a neural network model according to claim 1, characterized in that: The step S3 further includes the following specific steps: Obtain an audio neural network model, obtain target audio time period collection data, and obtain multiple target audio interception time periods based on the target audio time period collection data; An audio neural network model is used to identify the second type of sound source in the target audio interception period, and the second type of sound source in each target audio interception period is removed as audio noise.

9. A data recognition system based on a neural network model, applicable to a data recognition method based on a neural network model according to any one of claims 1 to 8, characterized in that: The data identification system comprises: Audio acquisition module: obtains the audio to be recognized, and intercepts the sound period of the target audio in the audio to be recognized to obtain the target audio period acquisition data; Model creation module: obtains multiple sets of historical target audio, labels the historical target audio by setting denoising indicators, obtains audio source labeling data, and creates an audio neural network model based on the audio source labeling data; Audio denoising module: identifies and removes audio noise for each target audio capture period based on the audio neural network model.

10. Application of the data recognition method according to any one of claims 1 to 8 in audio denoising.

Citation Information

Patent Citations

  • Method and system for screening noise in voice automatic annotation data

    CN115440238A

  • Method and system for intelligently identifying environmental noise

    CN115662464A

  • Vehicle operation noise data processing method and system

    CN116168714A

  • Acquisition method of voice recognition model in electric power emergency consultation environment, voice recognition method, device, equipment, storage medium and program product

    CN120048250A

  • 5G communication equipment intelligent voice interaction method and system based on deep learning

    CN120126463A