An audio data detection method, computer, and readable storage medium

By extracting and fusing the audio data to be detected in the audio output object, identifying the object quality results, the problems of high cost and low efficiency of speaker single-body quality detection in the prior art are solved, and efficient and accurate audio data detection is achieved.

CN115910107BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110968745.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-23
Publication Date
2025-05-27
Estimated Expiration
2041-08-23

AI Technical Summary

Technical Problem

In the prior art, the quality inspection cost of speaker singles is high, the testing environment is complex, and it is difficult to conduct large-scale batch inspections, resulting in low detection efficiency.

Method used

By obtaining the audio data to be detected generated by the audio output object, converting it into mono audio data, feature extraction and fusion are used for N target Fourier sizes to identify the object quality results of the audio output object.

Benefits of technology

It realizes efficient detection of audio data, improves detection accuracy and coverage, and reduces dependence on complex testing environments and personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910107B_ABST
    Figure CN115910107B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses an audio data detection method, a computer, and a readable storage medium. The method includes: obtaining the audio data to be detected generated by an audio output object, and obtaining the mono audio data corresponding to the audio data to be detected; using N target Fourier sizes to respectively extract features from the mono audio data to obtain N audio sub-features corresponding to the mono audio data, and performing feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data; N is a positive integer; identifying the object quality result of the audio output object according to the target audio feature. By using the present application, the accuracy of audio data detection and the test coverage rate of the audio output object can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to an audio data detection method, a computer, and a readable storage medium. Background Art

[0002] When many devices on the market are produced currently, there may be a situation where the speaker unit is defective, that is, problems caused by the quality of the speaker unit. Therefore, it is necessary to detect the quality of the speaker unit. Currently, generally, a certain number of speaker unit devices are randomly selected according to quantity, and the selected speaker unit devices are sent to an audio laboratory for acoustic single-unit verification. The cost of detection by this method is relatively high. On the one hand, the equipment in the audio laboratory is expensive, and on the other hand, the test environment, test cycle, test personnel, etc. are relatively complex, making it difficult to conduct large-scale batch detection. Only small-batch detection of speaker unit devices can be carried out, resulting in low detection efficiency for speaker unit devices. Summary of the Invention

[0003] The embodiments of this application provide an audio data detection method, a computer, and a readable storage medium, which can improve the accuracy of audio data detection, and improve the detection efficiency and test coverage rate of audio output objects.

[0004] On the one hand, the embodiments of this application provide an audio data detection method, which includes:

[0005] Obtain the audio data to be detected generated by an audio output object, and obtain the mono audio data corresponding to the audio data to be detected;

[0006] Adopt N target Fourier sizes to respectively extract features from the mono audio data, obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data; N is a positive integer;

[0007] Identify the object quality result of the audio output object according to the target audio feature.

[0008] On the one hand, the embodiments of this application provide an audio data detection device, which includes:

[0009] An audio acquisition module, configured to acquire the audio data to be detected generated by an audio output object;

[0010] A mono acquisition module, configured to acquire the mono audio data corresponding to the audio data to be detected;

[0011] A feature extraction module, which is used to perform feature extraction on the monophonic audio data respectively using N target Fourier sizes, obtain N audio sub-features corresponding to the monophonic audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the monophonic audio data; N is a positive integer;

[0012] A quality detection module, which is used to identify the object quality result of the audio output object according to the target audio feature.

[0013] Among them, the monophonic acquisition module includes:

[0014] A channel acquisition unit, which is used to acquire the number of audio channels of the audio data to be detected;

[0015] A monophonic determination unit, which is used to determine the audio data to be detected as monophonic audio data if the number of audio channels is the number of monophonic channels;

[0016] A channel fusion unit, which is used to split the audio data to be detected based on the number of audio channels if the number of audio channels is the number of multi-channels, obtain at least two split audio data, and perform audio fusion on the at least two split audio data to obtain monophonic audio data; the split audio data refers to the audio data generated by one channel.

[0017] Among them, the channel fusion unit includes:

[0018] An audio splicing sub-unit, which is used to splice the at least two split audio data to obtain monophonic audio data; or,

[0019] An audio equalization sub-unit, which is used to perform averaging processing on the at least two split audio data to obtain monophonic audio data.

[0020] Among them, the feature extraction module includes:

[0021] An audio division unit, which is used to obtain the detected audio duration and the audio processing duration threshold of the monophonic audio data, and divide the monophonic audio data based on the detected audio duration and the audio processing duration threshold to obtain M segmented audio data; M is a positive integer;

[0022] A segmented extraction unit, which is used to perform feature extraction on the i-th segmented audio data respectively using N target Fourier sizes to obtain N segmented audio sub-features corresponding to the i-th segmented audio data; i is a positive integer less than or equal to M;

[0023] A segmented fusion unit, which is used to perform feature fusion on the N segmented audio sub-features corresponding to the i-th segmented audio data to obtain the segmented audio feature corresponding to the i-th segmented audio data; the target audio feature includes the segmented audio features corresponding to the M segmented audio data respectively;

[0024] The quality detection module includes:

[0025] A segmented prediction unit for predicting the audio quality results corresponding to M segmented audio features respectively;

[0026] A segmented integration unit for determining the object quality result of the audio output object according to the audio quality results corresponding to M segmented audio features respectively.

[0027] Among them, the segmented integration unit includes:

[0028] A quantity statistics subunit for obtaining the normal quantity of the segmented quality normal results and the abnormal quantity of the segmented quality abnormal results in the audio quality results corresponding to M segmented audio features respectively;

[0029] A normal determination subunit for determining that the object quality result of the audio output object is the object normal result if the normal quantity is greater than the abnormal quantity;

[0030] An abnormal determination subunit for determining that the object quality result of the audio output object is the object abnormal result if the normal quantity is less than or equal to the abnormal quantity.

[0031] Among them, the segmented integration unit includes:

[0032] The quantity statistics subunit is further used for obtaining the normal quantity of the segmented quality normal results in the audio quality results corresponding to M segmented audio features respectively;

[0033] A ratio obtaining subunit for obtaining the quality normal rate between the normal quantity and the quantity of M audio quality results;

[0034] The normal determination subunit is further used for determining that the object quality result of the audio output object is the object normal result if the quality normal rate is greater than or equal to the object quality normal threshold;

[0035] The abnormal determination subunit is further used for determining that the object quality result of the audio output object is the object abnormal result if the quality normal rate is less than the object quality normal threshold.

[0036] Among them, the number of audio output objects is p, and p is a positive integer; the device further includes:

[0037] An abnormal object obtaining module for obtaining the object quality results corresponding to p audio output objects respectively, and determining the audio output objects with the object quality result being the object abnormal result as audio abnormal objects;

[0038] A communication obtaining module for obtaining the abnormal object information of the audio abnormal objects and obtaining the communication methods associated with p audio output objects;

[0039] A message sending module, configured to send an object exception message to a target terminal device based on a communication method, so that the target terminal device performs exception detection on an audio exception object based on the object exception message; the object exception message includes exception object information.

[0040] Wherein, the quality detection module includes:

[0041] A prediction acquisition unit, configured to input target audio features into an audio detection model, and output at least two candidate prediction results corresponding to the target audio features and prediction probability values respectively corresponding to each candidate prediction result through the audio detection model;

[0042] A quality determination unit, configured to determine the candidate prediction result with the largest prediction probability value as the object quality result of the audio output object.

[0043] Wherein, the device further includes:

[0044] An object sample acquisition module, configured to obtain a positive-negative sample ratio, obtain a normal audio output object and an abnormal audio output object based on the positive-negative sample ratio, obtain an object normal quality result label corresponding to the normal audio output object, and an object abnormal quality result label corresponding to the abnormal audio output object;

[0045] An audio acquisition module, configured to acquire normal audio data generated by a normal audio output object and abnormal audio data generated by an abnormal audio output object, and obtain normal mono audio data corresponding to the normal audio data and abnormal mono audio data corresponding to the abnormal audio data;

[0046] A normal feature extraction module, configured to extract features from the normal mono audio data by using N target Fourier sizes to obtain normal audio features corresponding to the normal mono audio data;

[0047] An abnormal feature extraction module, configured to extract features from the abnormal mono audio data by using N target Fourier sizes to obtain abnormal audio features corresponding to the abnormal mono audio data;

[0048] A model adjustment module, configured to adjust parameters of an initial detection model by using the normal audio features, the abnormal audio features, the object normal quality result label, and the object abnormal quality result label to obtain an audio detection model.

[0049] Wherein, the normal feature extraction module includes:

[0050] A first partitioning unit, configured to obtain the normal audio duration of the normal mono audio data and the audio processing duration threshold, and based on the normal audio duration and the audio processing duration threshold, partition the normal mono audio data to obtain f normal segmented audio data; f is a positive integer;

[0051] A first extraction unit, configured to perform feature extraction on each of the f normal segmented audio data by using N target Fourier sizes, to obtain the normal segmented audio features corresponding to each normal segmented audio data;

[0052] The abnormal feature extraction module includes:

[0053] A second partitioning unit, configured to obtain the abnormal audio duration of the abnormal mono audio data and the audio processing duration threshold, and based on the abnormal audio duration and the audio processing duration threshold, partition the abnormal mono audio data to obtain h abnormal segmented audio data; h is a positive integer;

[0054] A second extraction unit, configured to perform feature extraction on each of the h abnormal segmented audio data by using N target Fourier sizes, to obtain the abnormal segmented audio features corresponding to each abnormal segmented audio data.

[0055] Wherein, the apparatus further includes:

[0056] A segmented label adding module, configured to add segmented normal quality result labels to all of the f normal segmented audio features, and add segmented abnormal quality result labels to all of the h abnormal segmented audio features; the object normal quality result labels include f segmented normal quality result labels, and the object abnormal quality result labels include h segmented abnormal quality result labels;

[0057] The model adjustment module includes:

[0058] A first prediction unit, configured to respectively input the f normal segmented audio features into the initial detection model, and output the first quality prediction results respectively corresponding to the f normal segmented audio features through the initial detection model;

[0059] A first loss generation unit, configured to generate a first loss function according to the first quality prediction result corresponding to each normal segmented audio feature and the segmented normal quality result label; the number of the first loss functions is f;

[0060] A second prediction unit, configured to respectively input the h abnormal segmented audio features into the initial detection model, and output the second quality prediction results respectively corresponding to the h abnormal segmented audio features through the initial detection model;

[0061] A second loss generation unit, configured to generate a second loss function according to the second quality prediction result corresponding to each abnormal segmented audio feature and the segmented abnormal quality result label; the number of second loss functions is h;

[0062] A parameter adjustment unit, configured to adjust the parameters of the initial detection model based on f first loss functions and h second loss functions to obtain an audio detection model.

[0063] Wherein, the model adjustment module includes:

[0064] A forward adjustment unit, configured to perform forward parameter adjustment on the initial detection model by using normal audio features and object normal quality result labels;

[0065] A negative adjustment unit, configured to perform negative parameter adjustment on the initial detection model by using abnormal audio features and object abnormal quality result labels;

[0066] A model determination unit, configured to determine the initial detection model after normal parameter adjustment and negative parameter adjustment as the audio detection model.

[0067] Wherein, the audio acquisition module includes:

[0068] An initial acquisition unit, configured to obtain the audio test duration, acquire first initial audio data generated by a normal audio output object, and acquire second initial audio data generated by an abnormal audio output object;

[0069] A normal sample acquisition unit, configured to acquire first test audio data corresponding to the audio test duration from the first initial audio data, and determine the first initial audio data corresponding to the first test audio data in a qualified test state as normal audio data;

[0070] An abnormal sample acquisition unit, configured to acquire second test audio data corresponding to the audio test duration from the second initial audio data, and determine the second initial audio data corresponding to the second test audio data in a qualified test state as abnormal audio data.

[0071] One aspect of the embodiments of the present application provides a computer device, including a processor, a memory, and an input / output interface;

[0072] The processor is respectively connected to the memory and the input / output interface. Wherein, the input / output interface is used to receive and output data, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device including the processor executes the audio data detection method in one aspect of the embodiments of the present application.

[0073] One aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the audio data detection method in one aspect of the embodiments of the present application.

[0074] One aspect of the embodiments of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the methods provided in various alternative manners in one aspect of the embodiments of the present application.

[0075] Implementing the embodiments of the present application will have the following beneficial effects:

[0076] In the embodiments of the present application, the to-be-detected audio data generated by an audio output object is obtained, and the mono audio data corresponding to the to-be-detected audio data is obtained; N target Fourier sizes are used to respectively extract features from the mono audio data to obtain N audio sub-features corresponding to the mono audio data, and the N audio sub-features are feature-fused to obtain the target audio feature corresponding to the mono audio data; N is a positive integer; an object quality result of the audio output object is identified according to the target audio feature. Through the above process, the features of the audio data generated by the audio output object can be extracted, and the quality of the audio output object can be identified based on the extracted features, realizing the automatic detection of the audio output object, so that the audio output object can be detected without complex testers and test environments, etc., improving the efficiency of audio data detection and the detection coverage rate of the audio output object. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0078] Figure 1 FIG. is a network interaction architecture diagram of audio data detection provided by an embodiment of the present application;

[0079] Figure 2 FIG. is a schematic diagram of an audio data detection scenario provided by an embodiment of the present application;

[0080] Figure 3 FIG. is a flowchart of a method for audio data detection provided by an embodiment of the present application;

[0081] Figure 4 It is a schematic diagram of a monophonic audio data acquisition scenario provided by an embodiment of the present application;

[0082] Figure 5 It is a schematic diagram of an audio feature acquisition scenario provided by an embodiment of the present application;

[0083] Figure 6 It is a flowchart of a method for training an audio data detection model provided by an embodiment of the present application;

[0084] Figure 7 It is a schematic diagram of an audio data acquisition scenario provided by an embodiment of the present application;

[0085] Figure 8 It is a schematic diagram of a model training process provided by an embodiment of the present application;

[0086] Figure 9 It is a schematic diagram of an audio data detection device provided by an embodiment of the present application;

[0087] Figure 10 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Specific embodiments

[0088] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0089] In the embodiments of the present application, please refer to Figure 1 , Figure 1It is a network interaction architecture diagram for audio data detection provided by an embodiment of the present application, and the embodiment of the present application can be implemented by a computer device 101. Among them, the computer device 101 can obtain the audio data to be detected generated by the audio output object. The number of audio output objects is one or at least two, such as audio output object 102a, audio output object 102b, and audio output object 102c, etc. Among them, the audio output object refers to a device or object with a sound generating element (such as a speaker unit, etc.), such as a speaker, a microphone, or a phonograph head, etc., or an electronic device that can generate sound, such as a computer device like a laptop or a mobile phone. Among them, the speaker unit is an electroacoustic element that can convert an electrical signal into sound. The computer device 101 can extract features from the obtained audio data to be detected to obtain target audio features, and detect the quality of the audio output object based on the target audio features, realizing the automatic detection of the audio output object, so that complex human and material resources are not required, and the detection efficiency of the audio output object is improved. Moreover, by using different Fourier sizes, features can be extracted from the audio data generated by the audio output object, and relatively accurate and comprehensive features can be obtained, improving the accuracy of audio data detection.

[0090] Among them, the present application may involve machine learning technology in the field of artificial intelligence, and the quality of the audio output object is detected through machine learning technology.

[0091] Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machine to have the functions of perception, reasoning, and decision-making.

[0092] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0093] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0094] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0095] Specifically, please refer to Figure 2 , Figure 2 which is a schematic diagram of an audio data detection scenario provided by an embodiment of this application. As Figure 2As shown, the computer device can obtain the audio data to be detected 202 generated by the audio output object 201, and obtain the mono audio data corresponding to the audio data to be detected 202. Among them, the computer device can perform channel conversion on the audio data to be detected 202 based on the number of audio channels of the audio data to be detected 202 to obtain the mono audio data corresponding to the audio data to be detected 202. By obtaining the mono audio data, the interference between the audio data generated by different channels can be reduced, and at the same time, the information of the audio data to be detected 202 generated by the audio output object 201 can be relatively comprehensively retained, so that the information contained in the mono audio data is relatively comprehensive and the interference is less, which can improve the accuracy of audio data detection. Further, the computer device can use N target Fourier sizes to respectively extract features from the mono audio data to obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data. N is a positive integer. Optionally, if N is at least two, the N target Fourier sizes may not be exactly the same, that is, two target Fourier sizes may be the same or different. Through the N target Fourier sizes, the features of the mono audio data are extracted as comprehensively as possible to further improve the accuracy of audio data detection. Further, the computer device can identify the object quality result of the audio output object according to the target audio feature. The object quality result can be used to indicate whether the audio output object is an object normal result or an object abnormal result, that is, it is used to indicate the quality detection result for the audio output object, thus realizing the quality detection of the audio output object.

[0096] It is understandable that the computer devices mentioned in the embodiments of the present application include, but are not limited to, terminal devices or servers. In other words, the computer device can be a server or a terminal device, or a system composed of a server and a terminal device, etc. Among them, the above-mentioned terminal device can be an electronic device, including but not limited to mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, in-vehicle devices, augmented reality / virtual reality (AR / VR) devices, head-mounted displays, smart TVs, wearable devices, smart speakers, digital cameras, cameras, and other mobile internet devices (MIDs) with network access capabilities, or terminal devices in scenarios such as trains, ships, and flights. Among them, the above-mentioned server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-road collaboration, content delivery network (CDN), and big data and artificial intelligence platforms.

[0097] Optionally, the data involved in the embodiments of the present application can be stored in a computer device, or the data can be stored based on cloud storage technology or a blockchain network, and no limitation is made here.

[0098] Further, please refer to Figure 3 , Figure 3 which is a flowchart of a method for detecting audio data provided by an embodiment of the present application. As Figure 3 shown, in the method embodiment described in Figure 3 , the audio data detection process includes the following steps:

[0099] Step S301, obtain the audio data to be detected generated by the audio output object, and obtain the mono audio data corresponding to the audio data to be detected.

[0100] In an embodiment of the present application, a computer device may obtain the audio data to be detected generated by an audio output object, obtain the number of audio channels of the audio data to be detected, and convert the audio data to be detected into mono audio data based on the number of audio channels. Optionally, the computer device may obtain the target initial audio data generated by the audio output object and determine the target initial audio data as the audio data to be detected. The audio data to be detected may be timing information or a frequency domain signal; or, if the target initial audio data is a timing signal, obtain the window size and conversion step, and perform spectral conversion on the target initial audio data based on the window size and conversion step to obtain the audio data to be detected corresponding to the target initial audio data. The audio data to be detected may be a frequency domain signal. Wherein, the window size and conversion step refer to the parameters for performing spectral conversion on the target initial audio data. Specifically, the computer device may obtain the number of audio channels of the audio data to be detected. If the number of audio channels is a mono number, that is, the audio data to be detected is audio data generated by one channel, then determine the audio data to be detected as mono audio data; if the number of audio channels is a multi-channel number, that is, the audio data to be detected is audio data generated by at least two channels, then split the audio data to be detected based on the number of audio channels to obtain at least two split audio data, and perform audio fusion on the at least two split audio data to obtain mono audio data; the split audio data refers to audio data generated by one channel. Wherein, a channel refers to an independent audio signal collected or played back at different spatial positions during sound recording or playback. The number of channels is also the number of sound sources during sound recording or the number of corresponding speakers during playback. The number of audio channels includes mono numbers and multi-channel numbers. The multi-channel numbers include but are not limited to two-channel numbers (i.e., 2 channels), four-channel numbers, 5.1-channel numbers, and 7.1-channel numbers, etc. Optionally, the computer device may split the audio data to be detected by using a channel splitting method to obtain at least two split audio data. Wherein, the channel splitting method includes but is not limited to using a channel splitting application program or using a channel splitting algorithm (such as a multi-channel splitting function, etc.).

[0101] Further, when fusing at least two split audio data to obtain mono audio data, the computer device may splice the at least two split audio data to obtain mono audio data; or, may perform averaging processing on the at least two split audio data, for example, obtain the average value of the at least two split audio data, etc., to obtain mono audio data. Optionally, when the audio output object generates the audio data to be detected, different channels generally generate audio data synchronously. Therefore, based on the generation correspondence of the at least two split audio data, the at least two split audio data may be averaged to obtain the mono audio data corresponding to the at least two split audio data. Optionally, the present application may fuse the at least two split audio data to obtain mono audio data through, but not limited to, the above audio fusion methods, that is, other audio fusion methods may also be used to fuse the at least two split audio data.

[0102] For example, please refer to Figure 4 , Figure 4 which is a schematic diagram of a scenario for obtaining mono audio data provided by an embodiment of the present application. As Figure 4 shown, the computer device obtains the target initial audio data 401 generated by the audio output object. If the target initial audio data 401 is a time series signal, then perform spectrum conversion on the target initial audio data 401 to obtain the audio data to be detected 402 corresponding to the target initial audio data 401. Assume that the number of audio channels of the audio data to be detected 402 is d, and d is a positive integer. If the number of audio channels of the audio data to be detected 402 is a multi-channel number, that is, if d is greater than 1, then the audio data to be detected 402 is split, and the audio data to be detected 402 may be split into at least two split audio data, including split audio data 4031, split audio data 4032, and split audio data 403d, etc. Among them, an audio fusion method may be used to fuse the at least two split audio data to obtain mono audio data. For example, audio fusion method ①, splice the split audio data 4031, split audio data 4032, and split audio data 403d, etc., to obtain mono audio data; audio fusion method ②, perform averaging processing on the at least two split audio data to obtain mono audio data, etc., which is not limited here.

[0103] Step S302: Use N target Fourier sizes to respectively extract features from the mono audio data to obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data.

[0104] In an embodiment of the present application, the computer device may adopt N target Fourier sizes to respectively extract N audio sub-features of the mono audio data, where N is a positive integer. Optionally, if N is at least two, the N target Fourier sizes are not completely the same.

[0105] For example, the target Fourier sizes and the extracted audio sub-features involved in the present application include but are not limited to the following N cases.

[0106] Under the first type of target Fourier size, assuming that the target Fourier size is 4096, the computer device may adopt the target Fourier size "4096" to extract the audio sub-feature of the mono audio data. This audio sub-feature can be used to represent the mel128 feature of the mono audio data, which can be considered as the feature obtained by mapping linear frequency to the mel scale. Among them, the mel scale is a scale based on the perceptual judgment of pitch by listeners equidistant from each other. Optionally, the computer device may adopt the target Fourier size "4096" to extract the audio sub-feature of the mono audio data through a mel filter.

[0107] Under the second type of target Fourier size, assuming that the target Fourier size is 2048, the computer device may adopt the target Fourier size to extract the audio sub-feature of the mono audio data. This audio sub-feature can be used to represent the Mel-scale Frequency Cepstral Coefficients (mfcc) feature. The mfcc feature is the cepstral parameter extracted under the mel scale. Optionally, the computer device may adopt the target Fourier size to extract the mel spectrum feature of the mono audio data through a mel filter, perform logarithmic processing on the mel spectrum feature to obtain a logarithmic spectrum feature, and perform Discrete Cosine Transform (DCT) processing on the logarithmic spectrum feature to obtain the audio sub-feature corresponding to the mono audio data.

[0108] Under the third target Fourier size, assuming the target Fourier size is 1024, the computer device can use the target Fourier size "1024" to extract the audio sub-features of the mono audio data, and the audio sub-features can be used to represent the zero-crossing rate feature of the mono audio data. Among them, the zero-crossing rate feature is used to represent the number of zero-crossing points of the mono audio data, which reflects the frequency characteristics of the mono audio data. Optionally, the computer device can obtain the number of data zero-crossing points of the mono audio data based on the target Fourier size, determine the zero-crossing rate feature of the mono audio data according to the number of data zero-crossing points, and determine the zero-crossing rate feature as an audio sub-feature of the mono audio data; or, the computer device can perform frame splitting processing on the mono audio data based on the target Fourier size to obtain at least two audio data frames, obtain the data zero-crossing points corresponding to the at least two audio data frames respectively, and form the zero-crossing rate feature of the mono audio data by the data zero-crossing points corresponding to the at least two audio data frames respectively, and determine the zero-crossing rate feature as an audio sub-feature of the mono audio data. Among them, each audio data frame can record a voice unit, and the computer device can add a window signal to the mono audio data and perform truncation processing on the mono audio data based on the window signal to obtain at least two audio data frames. Optionally, in a method of frame splitting processing for mono audio data, the computer device can add a window signal to the mono audio data, obtain the frame shift amount, determine the frame overlap length based on the frame shift amount, and perform truncation processing on the mono audio data with the window signal added according to the window signal and the frame overlap length to obtain at least two audio data frames. Or, the computer device can directly obtain the single-frame length and perform truncation processing on the mono audio data based on the single-frame length to obtain at least two audio data frames.

[0109] Under the fourth target Fourier size, assuming the target Fourier size is 1024, the computer device can use the target Fourier size "1024" to extract the spectral flatness feature of the mono audio data, and determine the spectral flatness feature as an audio sub-feature of the mono audio data. Among them, the spectral flatness feature is used to represent the similarity between the mono audio data and noise. The greater the signal flatness represented by the spectral flatness feature, the greater the probability that the mono audio data is noise. By adding this audio sub-feature, the detection accuracy of the audio data to be detected can be improved. Optionally, the computer device can sample the mono audio data based on the target Fourier size to obtain the sampled audio data, and use the flatness acquisition algorithm to extract features from the sampled audio data to obtain the frequency flatness feature of the mono audio data.

[0110] Under the fifth target Fourier size, assuming that the target Fourier size is 1024, the computer device can use the target Fourier size "1024" to extract the spectral centroid feature of the mono audio data and determine the spectral centroid feature as an audio sub-feature of the mono audio data. Among them, the spectral centroid feature is used to represent the energy concentration point of the mono audio data in the spectrum and can represent the clarity of the mono audio data. Optionally, the computer device can use the spectral centroid acquisition algorithm and the target Fourier size to collect the spectral centroid feature of the mono audio data. Or, the computer device can obtain the signal frequency of the mono audio data, perform short-time Fourier transform on the mono audio data using the target Fourier size to obtain Fourier audio data; obtain the signal spectral energy corresponding to the signal frequency in the Fourier audio data; where the signal frequency includes d signal sub-frequencies, and the signal spectral energy includes the signal sub-spectral energy corresponding to each signal sub-frequency, and d is a positive integer; the computer device can obtain the total signal energy corresponding to the d signal sub-spectral energies, perform weighted summation of the d signal sub-frequencies on the d signal sub-spectral energies according to the correspondence between the d signal sub-frequencies and the d signal sub-spectral energies to obtain a weighted energy value; determine the ratio of the weighted energy value to the total signal energy as the spectral centroid feature of the mono audio data, etc., which are not limited here. Optionally, the computer device can also perform frame processing on the mono audio data to obtain at least two audio data frames, obtain the d signal sub-frequencies and d signal sub-spectral frequencies corresponding to each audio data frame, and obtain the spectral centroid sub-feature of each audio data frame through the above-mentioned determination method of the spectral centroid feature, and form the spectral centroid feature of the mono audio data by the spectral centroid sub-features of each audio data frame.

[0111] Optionally, the audio sub-feature may be, but is not limited to, spectral flux feature, pitch frequency feature, sharpness feature, etc. Among them, the spectral flux feature is used to represent the degree of change between two adjacent audio data frames in the mono audio data; the pitch frequency feature is used to represent the frequency corresponding to the pitch of the mono audio data; the sharpness feature is used to represent the energy of the high-frequency part of the mono audio data that is perceived by the human ear, and the greater the energy of the high-frequency part, the greater the sharpness corresponding to the sharpness feature. Among them, the computer device can obtain one or at least two of the above-mentioned audio sub-features, that is, N is 1 or at least two. Optionally, when the computer device obtains the N audio sub-features corresponding to the mono audio data, it can perform standardization processing on the N audio sub-features respectively, and perform feature fusion on the N standardized audio sub-features to obtain the target audio feature corresponding to the mono audio data. Among them, the standardization processing includes, but is not limited to, normalization processing and binary coding processing, etc. Among them, the computer device can obtain the standard processing algorithms corresponding to the N audio sub-features respectively, and use the standard processing algorithms to perform standardization processing on the N audio sub-features respectively; for example, the standard processing algorithm corresponding to the audio sub-feature "zero-crossing rate feature" can be a binary coding algorithm; the standard processing algorithm corresponding to the audio sub-feature "mfcc feature" can be a normalization algorithm, etc.

[0112] Optionally, taking the kth audio sub-feature as an example, the computer device can obtain the first signal type corresponding to the kth audio sub-feature parameter, and the second signal type of the mono audio data; among them, the kth audio sub-feature refers to the value of the kth audio sub-feature parameter obtained after feature acquisition of the mono audio data; the first signal type includes time-domain signal type and frequency-domain signal type, and the second signal type includes time-domain signal type and frequency-domain signal type; the first signal type is used to represent the type of signal that needs to be processed when collecting the kth audio sub-feature parameter, that is to say, after the computer device performs feature acquisition on the signal in the first signal type, it can obtain the kth audio sub-feature corresponding to the kth audio sub-feature parameter; the second signal type is used to represent the signal type of the mono audio data. If the first signal type is the same as the second signal type, the kth audio sub-feature of the mono audio data is obtained. If the first signal type is different from the second signal type, time-frequency conversion is performed on the mono audio data based on the first signal type, and the kth audio sub-feature of the time-frequency converted mono audio data is collected; for example, if the first signal type is the frequency-domain signal type and the second signal type is the time-domain signal type, the computer device can perform time-frequency conversion on the mono audio data based on the first signal type to convert the mono audio data from the time-domain signal to the frequency-domain signal, and collect the kth audio sub-feature of the time-frequency converted mono audio data.

[0113] Further, the computer device may obtain the detected audio duration and the audio processing duration threshold of the mono audio data, and based on the detected audio duration and the audio processing duration threshold, perform audio segmentation on the mono audio data to obtain M segmented audio data; M is a positive integer. The computer device uses N target Fourier sizes to respectively extract features from the i-th segmented audio data to obtain N segmented audio sub-features corresponding to the i-th segmented audio data; i is a positive integer less than or equal to M. Feature fusion is performed on the N segmented audio sub-features corresponding to the i-th segmented audio data to obtain the segmented audio feature corresponding to the i-th segmented audio data; the target audio feature includes the segmented audio features respectively corresponding to the M segmented audio data.

[0114] For example, please refer to Figure 5 , Figure 5 which is a schematic diagram of an audio feature acquisition scenario provided by an embodiment of the present application. As Figure 5 shown, the computer device may obtain the detected audio duration and the audio processing duration threshold of the mono audio data 501, and based on the detected audio duration and the audio processing duration threshold, perform audio segmentation on the mono audio data 501 to obtain M segmented audio data, including segmented audio data 5021, segmented audio data 5022, …, and segmented audio data 502M. The computer device may obtain N segmented audio sub-features of each segmented audio data. Specifically, the computer device may obtain N segmented audio sub-features 5031 corresponding to the segmented audio data 5021, N segmented audio sub-features 5032 corresponding to the segmented audio data 5022, …, and N segmented audio sub-features 503M corresponding to the segmented audio data 502M. The computer device may perform feature fusion on the N segmented audio sub-features 5031 to obtain the segmented audio feature 5041 corresponding to the segmented audio data 5021; perform feature fusion on the N segmented audio sub-features 5032 to obtain the segmented audio feature 5042 corresponding to the segmented audio data 5022; …; perform feature fusion on the N segmented audio sub-features 503M to obtain the segmented audio feature 504M corresponding to the segmented audio data 502M. The segmented audio features 5041, segmented audio features 5042, …, and segmented audio features 504M constitute the target audio feature corresponding to the mono audio data.

[0115] Step S303, identify the object quality result of the audio output object according to the target audio feature.

[0116] In an embodiment of the present application, a computer device may input target audio features into an audio detection model, and output at least two candidate prediction results corresponding to the target audio features and prediction probability values respectively corresponding to each candidate prediction result through the audio detection model; determine the candidate prediction result with the largest prediction probability value as the object quality result of the audio output object. Optionally, the at least two candidate prediction results refer to the results respectively corresponding to at least two candidate prediction labels of the audio detection model. The at least two candidate prediction labels may include a normal result label and an abnormal result label. At this time, if the candidate prediction result with the largest prediction probability value is the result corresponding to the normal result label, the object normal result is determined as the object quality result of the audio output object; if the candidate prediction result with the largest prediction probability value is the result corresponding to the abnormal result label, the object abnormal result is determined as the object quality result of the audio output object. The at least two candidate prediction labels may include, but are not limited to, a normal result label, a broken sound result label, a silent result label, a noise result label, etc. If the candidate prediction result with the largest prediction probability value is the result corresponding to the broken sound result label, the result corresponding to the silent result label, or the result corresponding to the noise result label, etc., the object abnormal result may be determined as the object quality result of the audio output object. Labels other than the normal result label may be used to represent the abnormal attributes of the audio output object, that is, to indicate what problems the audio output object specifically has.

[0117] Optionally, if the monophonic audio data is divided into M segmented audio data, the computer device may predict the audio quality results respectively corresponding to the M segmented audio features; determine the object quality result of the audio output object according to the audio quality results respectively corresponding to the M segmented audio features. Optionally, the computer device may input the M segmented audio features into the audio detection model respectively, and obtain the audio quality results respectively corresponding to the M segmented audio data through the audio detection model. The determination process of the audio quality result may refer to the above object quality result determination process.

[0118] Further, when determining the object quality result of the audio output object according to the audio quality results respectively corresponding to the M segmented audio features, the computer device may obtain the normal quantity of the normal result of the segmented quality in the audio quality results respectively corresponding to the M segmented audio features, and the abnormal quantity of the abnormal result of the segmented quality; if the normal quantity is greater than the abnormal quantity, determine that the object quality result of the audio output object is the object normal result; if the normal quantity is less than or equal to the abnormal quantity, determine that the object quality result of the audio output object is the object abnormal result. Alternatively, the computer device may obtain the normal quantity of the normal result of the segmented quality in the audio quality results respectively corresponding to the M segmented audio features; obtain the quality normal rate between the normal quantity and the quantity of the M audio quality results, if the quality normal rate is greater than or equal to the object quality normal threshold, determine that the object quality result of the audio output object is the object normal result; if the quality normal rate is less than the object quality normal threshold, determine that the object quality result of the audio output object is the object abnormal result. Among them, the determination method of the object quality result is not limited to the above normal quantity, abnormal quantity or quality normal rate, etc., and other object quality result determination methods can also be adopted based on needs to determine the object quality result of the audio output object.

[0119] Further, the number of audio output objects is p, and p is a positive integer. The computer device may obtain the object quality results respectively corresponding to the p audio output objects, and determine the audio output object with the object quality result of the object abnormal result as the audio abnormal object. Obtain the abnormal object information of the audio abnormal object, and obtain the communication methods associated with the p audio output objects. Send an object abnormal message to the target terminal device based on the communication method, so that the target terminal device performs abnormal detection on the audio abnormal object based on the object abnormal message; the object abnormal message includes the abnormal object information.

[0120] Among them, this application can be applied to the Design Verification Test (DVT) stage of the audio output object, or to the Mass Production (MP) stage of the audio output object, or to the Mass Verification Test (MVT) stage of the audio output object, etc., without limitation here. Specifically, for example, in the DVT stage, the computer device can obtain the abnormal elements of the audio abnormal object, and based on the abnormal elements, the object design method of the audio output object. Among them, the audio abnormal object refers to the audio output object whose object quality result is an object abnormal result, and the abnormal element refers to the reason for the object abnormal result. For example, if the abnormal element is the audio playback method of the audio abnormal object, then update the audio playback method in the object design method, etc. For example, in the MP stage, the computer device can obtain the audio abnormal object, which refers to the audio output object whose object quality result is an object abnormal result, that is, the audio output object with an abnormality. The computer device can, based on the abnormal object information of the audio abnormal object, instruct the business personnel to recycle the audio abnormal object or repair the audio abnormal object, etc.

[0121] Optionally, the computer device can generate a smart contract for quality detection of the audio output object based on the above Figure 3 shown steps. When receiving a quality detection request for the audio output object, the computer device can call the smart contract based on the quality detection request and obtain the object quality result of the audio output object based on the smart contract.

[0122] In the embodiment of this application, obtain the audio data to be detected generated by the audio output object, and obtain the mono audio data corresponding to the audio data to be detected; use N target Fourier sizes to perform feature extraction on the mono audio data respectively to obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data; N is a positive integer; identify the object quality result of the audio output object according to the target audio feature. Through the above process, the features of the audio data generated by the audio output object can be extracted, and the quality of the audio output object can be identified based on the extracted features, realizing the automatic detection of the audio output object, so that the audio output object can be detected without a complex tester and test environment, etc., improving the efficiency of audio data detection and the detection coverage rate of the audio output object.

[0123] Further, please refer to Figure 6 , Figure 6 which is a flowchart of a method for training an audio data detection model provided by an embodiment of this application. AsFigure 6 As shown in Figure 6 In the method embodiments described below, the model training process for audio data detection includes the following steps:

[0124] Step S601: Obtain the positive and negative sample ratio, obtain the normal audio output object and the abnormal audio output object based on the positive and negative sample ratio, obtain the object normal quality result label corresponding to the normal audio output object, and the object abnormal quality result label corresponding to the abnormal audio output object.

[0125] In the embodiments of the present application, the computer device can obtain an audio output object sample, where the audio output object sample can include an object positive sample and an object negative sample. Specifically, the computer device can obtain the positive and negative sample ratio, and obtain the normal audio output object and the abnormal audio output object based on the positive and negative sample ratio. Among them, the normal audio output object is the object positive sample, and the abnormal audio output object is the object negative sample. The ratio between the number of normal audio output objects and the number of abnormal audio output objects is the positive and negative sample ratio. Optionally, the computer device can obtain at least two audio physical attributes, including but not limited to the channel number attribute, the amplitude attribute, and the sampling precision attribute, etc. The computer device can obtain the attribute types corresponding to the at least two audio physical attributes respectively, and obtain the object positive sample and the object negative sample based on the attribute types corresponding to the at least two audio physical attributes respectively, so as to increase the diversity and comprehensiveness of the audio output object sample and improve the generalization of the model. Among them, the attribute type is used to represent the classification of the corresponding audio physical attribute, etc. For example, the attribute types corresponding to the channel number attribute include but not limited to the mono-channel type, the stereo-channel type, the 2.1-channel type, and the 5.1-channel type, etc.; the attribute types corresponding to the sampling precision attribute include but not limited to 48k, 32k, and 16k, etc. Optionally, the computer device can obtain the object normal quality result label corresponding to the normal audio output object, and the object abnormal quality result label corresponding to the abnormal audio output object. Through the object normal quality result label and the object abnormal quality result label, a binary classification model can be trained; optionally, the object abnormal quality result label can include but not limited to the normal result label, the crackling result label, the silent result label, and the noise result label, etc. Through the above labels, a multi-classification model can be trained, and this model can detect the specific abnormal conditions of the audio output object.

[0126] Step S602: Collect the normal audio data generated by the normal audio output object, and the abnormal audio data generated by the abnormal audio output object, and obtain the normal mono-channel audio data corresponding to the normal audio data, and the abnormal mono-channel audio data corresponding to the abnormal audio data.

[0127] In the embodiments of the present application, the computer device may collect the normal audio data generated by the normal audio output object and the abnormal audio data generated by the abnormal audio output object, and obtain the normal mono audio data corresponding to the normal audio data and the abnormal mono audio data corresponding to the abnormal audio data. For the relevant description of this process, reference may be made to Figure 3 the relevant description in step S301 in Figure 3 . Specifically, the computer device may obtain the normal audio channel number of the normal audio data. If the normal audio channel number is the mono number, the normal audio data is determined as the normal mono audio data; if the normal audio channel number is the multi-channel number, the normal audio data is split based on the normal audio channel number to obtain at least two normal split audio data, and the at least two normal split audio data are subjected to audio fusion to obtain the normal mono audio data. Similarly, the computer device may obtain the abnormal audio channel number of the abnormal audio data. If the abnormal audio channel number is the mono number, the abnormal audio data is determined as the abnormal mono audio data; if the abnormal audio channel number is the multi-channel number, the abnormal audio data is split based on the abnormal audio channel number to obtain at least two abnormal split audio data, and the at least two abnormal split audio data are subjected to audio fusion to obtain the abnormal mono audio data.

[0128] Optionally, the computer device may obtain the audio test duration, collect the first initial audio data generated by the normal audio output object, and obtain the second initial audio data generated by the abnormal audio output object. The first test audio data corresponding to the audio test duration is obtained from the first initial audio data, and the first initial audio data corresponding to the first test audio data in the qualified test state is determined as the normal audio data; wherein, the first test audio data being in the qualified test state means that the first test audio data is normal. The computer device may obtain the first initial audio data corresponding to the first test audio data in the qualified test state, denoted as the first qualified audio data, and delete the first test audio data in the first qualified audio data to obtain the normal audio data. The second test audio data corresponding to the audio test duration is obtained from the second initial audio data, and the second initial audio data corresponding to the second test audio data in the qualified test state is determined as the abnormal audio data. Wherein, the second test audio data being in the qualified test state means that the second test audio data is normal. The computer device may obtain the second initial audio data corresponding to the second test audio data in the qualified test state, denoted as the second qualified audio data, and delete the second test audio data in the second qualified audio data to obtain the abnormal audio data.

[0129] Among them, the computer device can collect the first test audio data generated by the auxiliary object during the audio test duration, collect the first object audio data generated by the normal audio output object, and combine the first test audio data and the first object audio data into the first initial audio data of the normal audio output object; during the audio test duration, collect the second test audio data generated by the auxiliary object, collect the second object audio data generated by the abnormal audio output object, and combine the second test audio data and the second object audio data into the second initial audio data of the abnormal audio output object. Among them, the auxiliary object refers to an audio output object whose output audio data is normal. Optionally, this process can also be implemented by an audio collection device, and the computer device can obtain the first initial audio data of the normal audio output object and the second initial audio data of the abnormal audio output object from the audio collection device. Through the first test audio data and the second test audio data, the first initial audio data and the second initial audio data can be screened. If the first test audio data is in a qualified test state, it means that the acquisition process of the first initial audio data where the first test audio data is located is normal, and the first initial audio data can be used; if the first test audio data is in an abnormal test state, it means that the acquisition process of the first initial audio data where the first test audio data is located is abnormal, that is, the obtained first initial audio data may be inaccurate, and the first initial audio data will affect the model training result. Similarly, if the second test audio data is in a qualified test state, it means that the acquisition process of the second initial audio data where the second test audio data is located is normal, and the second initial audio data can be used; if the second test audio data is in an abnormal test state, it means that the acquisition process of the second initial audio data where the second test audio data is located is abnormal, that is, the obtained second initial audio data may be inaccurate, and the second initial audio data will affect the model training result. Through the above process, the invalid data in the obtained samples can be reduced, that is, the obtained normal audio data and abnormal audio data can be made more accurate, and the impact of the acquisition process on the samples can be reduced, thereby improving the training accuracy of the model.

[0130] For example, please refer to Figure 7 , Figure 7 which is a schematic diagram of an audio data acquisition scenario provided by an embodiment of the present application. As Figure 7As shown, the computer device can obtain the first initial audio data 701, obtain the audio test duration, and obtain the first test audio data 702 corresponding to the audio test duration from the first initial audio data 701. If the first test audio data 702 is in an abnormal test state, the first initial audio data 701 is deleted; if the first test audio data 702 is in a qualified test state, the first test audio data 702 in the first initial audio data 701 is deleted to obtain the intercepted audio data 703, and the intercepted audio data 703 is determined as the normal audio data.

[0131] Step S603: Extract features from the normal mono audio data using N target Fourier sizes to obtain the normal audio features corresponding to the normal mono audio data.

[0132] In the embodiment of the present application, the computer device can extract features from the normal mono audio data using N target Fourier sizes to obtain the normal audio features corresponding to the normal mono audio data. Optionally, the computer device can obtain the normal audio duration of the normal mono audio data and the audio processing duration threshold, and based on the normal audio duration and the audio processing duration threshold, perform audio segmentation on the normal mono audio data to obtain f normal segmented audio data; f is a positive integer. Extract features from the f normal segmented audio data respectively using N target Fourier sizes to obtain the normal segmented audio features corresponding to each normal segmented audio data; add segmented normal quality result labels to the f normal segmented audio features; the object normal quality result label includes f segmented normal quality result labels. Among them, the normal audio features include f normal segmented audio features.

[0133] Step S604: Extract features from the abnormal mono audio data using N target Fourier sizes to obtain the abnormal audio features corresponding to the abnormal mono audio data.

[0134] In the embodiment of the present application, the computer device can extract features from the abnormal mono audio data using N target Fourier sizes to obtain the abnormal audio features corresponding to the abnormal mono audio data. Optionally, the computer device can obtain the abnormal audio duration of the abnormal mono audio data and the audio processing duration threshold, and based on the abnormal audio duration and the audio processing duration threshold, perform audio segmentation on the abnormal mono audio data to obtain h abnormal segmented audio data; h is a positive integer. Extract features from the h abnormal segmented audio data respectively using N target Fourier sizes to obtain the abnormal segmented audio features corresponding to each abnormal segmented audio data. Add segmented abnormal quality result labels to the h abnormal segmented audio features, and the object abnormal quality result label includes h segmented abnormal quality result labels. Among them, the abnormal audio features include h abnormal segmented audio features.

[0135] Among them, the execution order of step S603 and step S604 is not limited. That is, step S603 can be executed first, and then step S604; step S604 can be executed first, and then step S603; or step S603 and step S604 can be executed in parallel. The specific implementation processes of step S603 and step S604 can refer to Figure 3 the relevant description shown in step S302 in

[0136] Step S605: Use the normal audio feature, abnormal audio feature, object normal quality result label, and object abnormal quality result label to adjust the parameters of the initial detection model to obtain an audio detection model.

[0137] In the embodiment of the present application, the computer device can use the normal audio feature and the object normal quality result label to perform positive parameter adjustment on the initial detection model; use the abnormal audio feature and the object abnormal quality result label to perform negative parameter adjustment on the initial detection model; and determine the initial detection model after positive parameter adjustment and negative parameter adjustment as the audio detection model. Among them, the computer device can input the normal audio feature into the initial detection model, and the initial detection model outputs a first object prediction result corresponding to the normal audio feature. The first object prediction result includes at least two candidate prediction labels corresponding to the normal audio feature output by the initial detection model, and the probability of each candidate prediction label; perform positive parameter adjustment on the initial detection model through the first object prediction result and the object normal quality result label; input the abnormal audio feature into the initial detection model, and the initial detection model outputs a second object prediction result corresponding to the abnormal audio feature. The second object prediction result includes at least two candidate prediction labels corresponding to the abnormal audio feature output by the initial detection model, and the probability of each candidate prediction label; perform negative parameter adjustment on the initial detection model through the second object prediction result and the object abnormal quality result label.

[0138] Among them, any classification model can be obtained as the initial detection model. In this case, the selection of the model is relatively simple and convenient; or, a model including a residual network structure and a Batch Normalization network can be obtained as the initial detection model, which can improve the robustness and prediction effect of the model.

[0139] Optionally, if the normal audio feature includes f normal segmented audio features and the abnormal audio feature includes h abnormal segmented audio features, the f normal segmented audio features are respectively input into the initial detection model, and the first quality prediction results corresponding to the f normal segmented audio features are output through the initial detection model. The f first quality prediction results form the first object prediction result. According to the first quality prediction result corresponding to each normal segmented audio feature and the segmented normal quality result label, a first loss function is generated; the number of first loss functions is f; that is, each normal segmented audio feature corresponds to a first quality prediction result, and each normal segmented audio feature corresponds to a segmented normal quality result label. Based on this corresponding relationship, the first quality prediction result corresponding to each normal segmented audio feature and the segmented normal quality result label are obtained to generate a first loss function, and finally f first loss functions are obtained, where the first quality prediction result and the segmented normal quality result label for generating a first loss function correspond to the same normal segmented audio feature.

[0140] Further, the h abnormal segmented audio features are respectively input into the initial detection model, and the second quality prediction results corresponding to the h abnormal segmented audio features are output through the initial detection model; according to the second quality prediction result corresponding to each abnormal segmented audio feature and the segmented abnormal quality result label, a second loss function is generated; the number of second loss functions is h; based on the f first loss functions and the h second loss functions, the parameters of the initial detection model are adjusted to obtain the audio detection model.

[0141] Further, reference can be made to Figure 8 , Figure 8 which is a schematic diagram of a model training process provided by an embodiment of the present application. As Figure 8 shown, the model training includes processes such as collecting audio data -> audio feature engineering -> model design -> model training -> model evaluation. Among them, the collection of audio data refers to the process in which a computer device obtains audio output object samples and collects normal audio data and abnormal audio data, and reference can be made to Figure 6 the steps S601 and S602 shown therein; audio feature engineering refers to the process in which a computer device obtains normal audio features and abnormal audio features, and reference can be made to Figure 6 the steps S603 and S604 shown therein; model design refers to the process in which a computer device obtains an initial detection model; model training refers to the process of adjusting the parameters of the initial detection model based on normal audio features and abnormal audio features to obtain an audio detection model, and reference can be made to Figure 6 the step S605 shown therein; model evaluation refers to the process in which a computer device evaluates the quality of the trained audio detection model. Optionally, the computer device can further optimize and adjust the audio detection model based on the quality evaluation result.

[0142] Optionally, the computer device implementing Figure 3 the audio data detection process shown and the computer device implementing Figure 6 the audio data detection model training process shown can be the same device or different devices.

[0143] Furthermore, please refer to Figure 9 , Figure 9 FIG. is a schematic diagram of an audio data detection device provided in an embodiment of the present application. The audio data detection device can be a computer program (including program code, etc.) running in a computer device. For example, the audio data detection device can be an application software; the device can be used to execute the corresponding steps in the method provided in the embodiment of the present application. As Figure 9 shown, the audio data detection device 900 can be used for Figure 3 the computer device in the corresponding embodiment. Specifically, the device can include: an audio acquisition module 11, a mono acquisition module 12, a feature extraction module 13, and a quality detection module 14.

[0144] The audio acquisition module 11 is used to acquire the audio data to be detected generated by the audio output object;

[0145] The mono acquisition module 12 is used to acquire the mono audio data corresponding to the audio data to be detected;

[0146] The feature extraction module 13 is used to perform feature extraction on the mono audio data respectively using N target Fourier sizes to obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data; N is a positive integer;

[0147] The quality detection module 14 is used to identify the object quality result of the audio output object according to the target audio feature.

[0148] Among them, the mono acquisition module 12 includes:

[0149] A channel acquisition unit 121, which is used to acquire the number of audio channels of the audio data to be detected;

[0150] A mono determination unit 122, which is used to determine the audio data to be detected as mono audio data if the number of audio channels is the number of mono channels;

[0151] The channel fusion unit 123 is configured to, if the number of audio channels is a multi-channel number, split the audio data to be detected based on the number of audio channels to obtain at least two split audio data, and perform audio fusion on the at least two split audio data to obtain mono audio data; the split audio data refers to the audio data generated by one channel.

[0152] Among them, the channel fusion unit 123 includes:

[0153] The audio splicing subunit 1231 is configured to splice the at least two split audio data to obtain mono audio data; or,

[0154] The audio equalization subunit 1232 is configured to perform averaging processing on the at least two split audio data to obtain mono audio data.

[0155] Among them, the feature extraction module 13 includes:

[0156] The audio division unit 131 is configured to obtain the detection audio duration and the audio processing duration threshold of the mono audio data, and divide the mono audio data based on the detection audio duration and the audio processing duration threshold to obtain M segmented audio data; M is a positive integer;

[0157] The segmented extraction unit 132 is configured to use N target Fourier sizes to respectively extract features from the i-th segmented audio data to obtain N segmented audio sub-features corresponding to the i-th segmented audio data; i is a positive integer less than or equal to M;

[0158] The segmented fusion unit 133 is configured to perform feature fusion on the N segmented audio sub-features corresponding to the i-th segmented audio data to obtain the segmented audio feature corresponding to the i-th segmented audio data; the target audio feature includes the segmented audio features respectively corresponding to the M segmented audio data;

[0159] The quality detection module 14 includes:

[0160] The segmented prediction unit 141 is configured to predict the audio quality results respectively corresponding to the M segmented audio features;

[0161] The segmented synthesis unit 142 is configured to determine the object quality result of the audio output object according to the audio quality results respectively corresponding to the M segmented audio features.

[0162] Among them, the segmented synthesis unit 142 includes:

[0163] The quantity statistics subunit 1421 is configured to obtain the normal quantity of the segmented quality normal results and the abnormal quantity of the segmented quality abnormal results in the audio quality results respectively corresponding to the M segmented audio features;

[0164] A normal determination subunit 1422, configured to determine that the object quality result of the audio output object is an object normal result if the normal quantity is greater than the abnormal quantity;

[0165] An abnormal determination subunit 1423, configured to determine that the object quality result of the audio output object is an object abnormal result if the normal quantity is less than or equal to the abnormal quantity.

[0166] Wherein, the segmentation integration unit 142 includes:

[0167] The quantity statistics subunit 1421 is further configured to obtain the normal quantity of the normal segmentation results in the audio quality results respectively corresponding to the M segmented audio features;

[0168] A ratio obtaining subunit 1424, configured to obtain the normal quality rate between the normal quantity and the quantity of the M audio quality results;

[0169] The normal determination subunit 1422 is further configured to determine that the object quality result of the audio output object is an object normal result if the normal quality rate is greater than or equal to the object quality normal threshold;

[0170] The abnormal determination subunit 1423 is further configured to determine that the object quality result of the audio output object is an object abnormal result if the normal quality rate is less than the object quality normal threshold.

[0171] Wherein, the number of audio output objects is p, and p is a positive integer; the apparatus 900 further includes:

[0172] An abnormal object obtaining module 15, configured to obtain the object quality results respectively corresponding to the p audio output objects, and determine the audio output objects with the object quality result being the object abnormal result as audio abnormal objects;

[0173] A communication obtaining module 16, configured to obtain the abnormal object information of the audio abnormal objects, and obtain the communication methods associated with the p audio output objects;

[0174] A message sending module 17, configured to send an object abnormal message to a target terminal device based on the communication method, so that the target terminal device performs abnormal detection on the audio abnormal objects according to the object abnormal message; the object abnormal message includes the abnormal object information.

[0175] Wherein, the quality detection module 14 includes:

[0176] A prediction obtaining unit 143, configured to input a target audio feature into an audio detection model, and output at least two candidate prediction results corresponding to the target audio feature and the prediction probability values respectively corresponding to each candidate prediction result through the audio detection model;

[0177] A quality determination unit 144, configured to determine, as the object quality result of the audio output object, the candidate prediction result with the maximum predicted probability value.

[0178] Wherein, the apparatus 900 further includes:

[0179] An object sample acquisition module 18, configured to obtain the positive and negative sample ratio, obtain normal audio output objects and abnormal audio output objects based on the positive and negative sample ratio, obtain the object normal quality result labels corresponding to the normal audio output objects, and the object abnormal quality result labels corresponding to the abnormal audio output objects;

[0180] An audio acquisition module 19, configured to acquire the normal audio data generated by the normal audio output objects and the abnormal audio data generated by the abnormal audio output objects, and obtain the normal mono audio data corresponding to the normal audio data and the abnormal mono audio data corresponding to the abnormal audio data;

[0181] A normal feature extraction module 20, configured to perform feature extraction on the normal mono audio data by using N target Fourier sizes to obtain the normal audio features corresponding to the normal mono audio data;

[0182] An abnormal feature extraction module 21, configured to perform feature extraction on the abnormal mono audio data by using N target Fourier sizes to obtain the abnormal audio features corresponding to the abnormal mono audio data;

[0183] A model adjustment module 22, configured to adjust the parameters of the initial detection model by using the normal audio features, the abnormal audio features, the object normal quality result labels, and the object abnormal quality result labels to obtain an audio detection model.

[0184] Wherein, the normal feature extraction module 20 includes:

[0185] A first division unit 20a, configured to obtain the normal audio duration of the normal mono audio data and the audio processing duration threshold, and perform audio division on the normal mono audio data based on the normal audio duration and the audio processing duration threshold to obtain f normal segmented audio data; f is a positive integer;

[0186] A first extraction unit 20b, configured to perform feature extraction on each of the f normal segmented audio data by using N target Fourier sizes to obtain the normal segmented audio features corresponding to each normal segmented audio data;

[0187] The abnormal feature extraction module 21 includes:

[0188] A second division unit 211, configured to obtain the abnormal audio duration of the abnormal monophonic audio data and the audio processing duration threshold, and based on the abnormal audio duration and the audio processing duration threshold, divide the abnormal monophonic audio data to obtain h abnormal segmented audio data; h is a positive integer;

[0189] A second extraction unit 212, configured to perform feature extraction on the h abnormal segmented audio data respectively by using N target Fourier sizes, so as to obtain abnormal segmented audio features corresponding to each abnormal segmented audio data.

[0190] Wherein, the apparatus 900 further includes:

[0191] A segmented label adding module 23, configured to add segmented normal quality result labels to all f normal segmented audio features, and add segmented abnormal quality result labels to all h abnormal segmented audio features; the object normal quality result labels include f segmented normal quality result labels, and the object abnormal quality result labels include h segmented abnormal quality result labels;

[0192] The model adjustment module 22 includes:

[0193] A first prediction unit 221, configured to input the f normal segmented audio features into the initial detection model respectively, and output first quality prediction results corresponding to the f normal segmented audio features respectively through the initial detection model;

[0194] A first loss generation unit 222, configured to generate a first loss function according to the first quality prediction result corresponding to each normal segmented audio feature and the segmented normal quality result label; the number of the first loss functions is f;

[0195] A second prediction unit 223, configured to input the h abnormal segmented audio features into the initial detection model respectively, and output second quality prediction results corresponding to the h abnormal segmented audio features respectively through the initial detection model;

[0196] A second loss generation unit 224, configured to generate a second loss function according to the second quality prediction result corresponding to each abnormal segmented audio feature and the segmented abnormal quality result label; the number of the second loss functions is h;

[0197] A parameter adjustment unit 225, configured to adjust the parameters of the initial detection model based on the f first loss functions and the h second loss functions to obtain an audio detection model.

[0198] Wherein, the model adjustment module 22 includes:

[0199] A forward adjustment unit 226, configured to perform forward parameter adjustment on the initial detection model by using the normal audio features and the object normal quality result labels;

[0200] A negative adjustment unit 227, configured to perform negative parameter adjustment on the initial detection model by using abnormal audio features and object abnormal quality result labels;

[0201] A model determination unit 228, configured to determine the initial detection model after normal parameter adjustment and negative parameter adjustment as an audio detection model.

[0202] Wherein, the audio acquisition module 19 includes:

[0203] An initial acquisition unit 191, configured to obtain an audio test duration, acquire first initial audio data generated by a normal audio output object, and acquire second initial audio data generated by an abnormal audio output object;

[0204] A normal sample acquisition unit 192, configured to acquire first test audio data corresponding to the audio test duration from the first initial audio data, and determine the first initial audio data corresponding to the first test audio data in a qualified test state as normal audio data;

[0205] An abnormal sample acquisition unit 193, configured to acquire second test audio data corresponding to the audio test duration from the second initial audio data, and determine the second initial audio data corresponding to the second test audio data in a qualified test state as abnormal audio data.

[0206] The embodiment of the present application provides an audio data detection device. The device can acquire to-be-detected audio data generated by an audio output object, and acquire mono audio data corresponding to the to-be-detected audio data; perform feature extraction on the mono audio data respectively by using N target Fourier sizes to obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain target audio features corresponding to the mono audio data; N is a positive integer; identify an object quality result of the audio output object according to the target audio features. Through the above process, the features of the audio data generated by the audio output object can be extracted, and the quality of the audio output object can be identified based on the extracted features, realizing the automatic detection of the audio output object, so that the audio output object can be detected without a complex tester and test environment, etc., improving the efficiency of audio data detection and the detection coverage rate of the audio output object.

[0207] See Figure 10 , Figure 10 is a schematic structural diagram of a computer device provided by the embodiment of the present application. As Figure 10As shown in the figure, the computer device in the embodiment of the present application may include: one or more processors 1001, a memory 1002, and an input / output interface 1003. The processor 1001, the memory 1002, and the input / output interface 1003 are connected through a bus 1004. The memory 1002 is used to store a computer program, and the computer program includes program instructions. The input / output interface 1003 is used to receive and output data, such as for data interaction between the computer device and an audio output object; the processor 1001 is used to execute the program instructions stored in the memory 1002.

[0208] Among them, the processor 1001 may perform the following operations:

[0209] Obtain the audio data to be detected generated by the audio output object, and obtain the mono audio data corresponding to the audio data to be detected;

[0210] Adopt N target Fourier sizes to perform feature extraction on the mono audio data respectively, obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data; N is a positive integer;

[0211] Identify the object quality result of the audio output object according to the target audio feature.

[0212] In some feasible implementation manners, the processor 1001 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0213] The memory 1002 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1001 and the input / output interface 1003. A part of the memory 1002 may also include a non-volatile random access memory. For example, the memory 1002 may also store information about the device type.

[0214] In a specific implementation, the computer device may execute the following through its built-in various functional modules Figure 3 or Figure 6The implementation manners provided by each step in can be specifically referred to in this Figure 3 or Figure 6 The implementation manners provided by each step in will not be elaborated herein.

[0215] An embodiment of the present application provides a computer device, including: a processor, an input / output interface, and a memory. The processor obtains a computer program in the memory and executes each step of the method shown in this Figure 3 to perform an audio data detection operation. The embodiment of the present application realizes obtaining the to-be-detected audio data generated by an audio output object, and obtaining the corresponding mono audio data of the to-be-detected audio data; adopting N target Fourier sizes to respectively extract features from the mono audio data to obtain N audio sub-features corresponding to the mono audio data, and performing feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data; N is a positive integer; and identifying the object quality result of the audio output object according to the target audio feature. Through the above process, the features of the audio data generated by the audio output object can be extracted, and the quality of the audio output object can be identified based on the extracted features, realizing the automatic detection of the audio output object, so that the audio output object can be detected without complex testers and test environments, etc., improving the efficiency of audio data detection and the detection coverage rate of the audio output object.

[0216] The embodiment of the present application further provides a computer-readable storage medium storing a computer program, and the computer program is suitable for being loaded and executed by the processor Figure 3 or Figure 6 The audio data detection method provided by each step in, and the specific implementation manners can be specifically referred to in this Figure 3 or Figure 6 The implementation manners provided by each step in will not be elaborated herein. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiment of the present application. As an example, the computer program can be deployed to be executed on a computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network.

[0217] The computer-readable storage medium may be the audio data detection device provided in any of the foregoing embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0218] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 3 or Figure 6 the methods provided in the various alternative manners in, realizes extracting the features of the audio data generated by the audio output object, and performs quality identification on the audio output object based on the extracted features, realizes the automatic detection of the audio output object, so that the audio output object can be detected without a complex tester and test environment, etc., improves the efficiency of audio data detection and the detection coverage rate of the audio output object.

[0219] In the description, claims and drawings of the embodiments of the present application, the terms "first", "second", etc. are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products or equipment.

[0220] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in this description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0221] The methods and related devices provided in the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided in the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable audio data detection devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable audio data detection devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable audio data detection devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable audio data detection devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.

[0222] The steps in the methods of the embodiments of this application can be adjusted, combined, and deleted according to actual needs.

[0223] The modules in the devices of the embodiments of this application can be combined, divided, and deleted according to actual needs.

[0224] The above-disclosed are only the preferred embodiments of this application. Of course, the scope of the rights of this application cannot be limited thereby. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.

Claims

1. An audio data detection method, characterized in that, the method includes: Obtain the audio data to be detected generated by the audio output object, and obtain the mono audio data corresponding to the audio data to be detected; Adopt N target Fourier sizes to respectively extract features from the mono audio data, obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data; N is a positive integer; Identify the object quality result of the audio output object according to the target audio feature.

2. The method according to claim 1, characterized in that, the obtaining of the mono audio data corresponding to the audio data to be detected includes: Obtain the number of audio channels of the audio data to be detected; If the number of audio channels is the mono number, determine the audio data to be detected as the mono audio data; If the number of audio channels is the multi-channel number, split the audio data to be detected based on the number of audio channels to obtain at least two split audio data, and perform audio fusion on the at least two split audio data to obtain the mono audio data; the split audio data refers to the audio data generated by one channel.

3. The method according to claim 2, characterized in that, the performing of audio fusion on the at least two split audio data to obtain the mono audio data includes: Performing audio splicing on the at least two split audio data to obtain the mono audio data; or, Performing an averaging process on the at least two split audio data to obtain the mono audio data.

4. The method according to claim 1, characterized in that, the adopting of N target Fourier sizes to respectively extract features from the mono audio data, obtain N audio sub-features corresponding to the mono audio data, and perform feature fusion on the N audio sub-features to obtain the target audio feature corresponding to the mono audio data includes: Obtain the detection audio duration and the audio processing duration threshold of the mono audio data, and based on the detection audio duration and the audio processing duration threshold, perform audio partitioning on the mono audio data to obtain M segmented audio data; M is a positive integer; Adopt N target Fourier sizes to respectively extract features from the i-th segmented audio data to obtain N segmented audio sub-features corresponding to the i-th segmented audio data; i is a positive integer less than or equal to M; Perform feature fusion on the N segmented audio sub-features corresponding to the i-th segmented audio data to obtain the segmented audio feature corresponding to the i-th segmented audio data; the target audio feature includes the segmented audio features respectively corresponding to the M segmented audio data; the identifying of the object quality result of the audio output object according to the target audio feature includes: Predict the audio quality results respectively corresponding to the M segmented audio features; Determine the object quality result of the audio output object according to the audio quality results respectively corresponding to the M segmented audio features.

5. The method according to claim 4, wherein, determining the object quality result of the audio output object according to the audio quality results respectively corresponding to the M segmented audio features includes: obtaining the normal quantity of the normal result of the segmented quality in the audio quality results respectively corresponding to the M segmented audio features, and the abnormal quantity of the abnormal result of the segmented quality; if the normal quantity is greater than the abnormal quantity, determining the object quality result of the audio output object as the object normal result; if the normal quantity is less than or equal to the abnormal quantity, determining the object quality result of the audio output object as the object abnormal result.

6. The method according to claim 4, wherein, determining the object quality result of the audio output object according to the audio quality results respectively corresponding to the M segmented audio features includes: obtaining the normal quantity of the normal result of the segmented quality in the audio quality results respectively corresponding to the M segmented audio features; obtaining the quality normal rate between the normal quantity and the quantity of the M audio quality results, if the quality normal rate is greater than or equal to the object quality normal threshold, determining the object quality result of the audio output object as the object normal result; if the quality normal rate is less than the object quality normal threshold, determining the object quality result of the audio output object as the object abnormal result.

7. The method according to claim 1, wherein, the number of the audio output objects is p, and p is a positive integer; the method further includes: obtaining the object quality results respectively corresponding to the p audio output objects, and determining the audio output object with the object quality result of the object abnormal result as the audio abnormal object; obtaining the abnormal object information of the audio abnormal object, and obtaining the communication methods associated with the p audio output objects; sending an object abnormal message to the target terminal device based on the communication method, so that the target terminal device performs abnormal detection on the audio abnormal object based on the object abnormal message; the object abnormal message includes the abnormal object information.

8. The method according to claim 1, wherein, identifying the object quality result of the audio output object according to the target audio feature includes; inputting the target audio feature into an audio detection model, and outputting at least two candidate prediction results corresponding to the target audio feature and the prediction probability value respectively corresponding to each candidate prediction result through the audio detection model; determining the candidate prediction result with the maximum prediction probability value as the object quality result of the audio output object.

9. The method according to claim 8, wherein, the method further includes: obtaining the positive and negative sample ratio, obtaining the normal audio output object and the abnormal audio output object based on the positive and negative sample ratio, obtaining the object normal quality result label corresponding to the normal audio output object, and the object abnormal quality result label corresponding to the abnormal audio output object; Collect the normal audio data generated by the normal audio output object and the abnormal audio data generated by the abnormal audio output object, and obtain the normal mono audio data corresponding to the normal audio data and the abnormal mono audio data corresponding to the abnormal audio data; Use the N target Fourier sizes to extract features from the normal mono audio data to obtain the normal audio features corresponding to the normal mono audio data; Use the N target Fourier sizes to extract features from the abnormal mono audio data to obtain the abnormal audio features corresponding to the abnormal mono audio data; Use the normal audio features, the abnormal audio features, the object normal quality result label, and the object abnormal quality result label to adjust the parameters of the initial detection model to obtain the audio detection model.

10. The method according to claim 9, wherein, The step of using the N target Fourier sizes to extract features from the normal mono audio data to obtain the normal audio features corresponding to the normal mono audio data includes: Obtain the normal audio duration of the normal mono audio data and the audio processing duration threshold, and based on the normal audio duration and the audio processing duration threshold, divide the normal mono audio data to obtain f normal segmented audio data; f is a positive integer; Use the N target Fourier sizes to extract features from the f normal segmented audio data respectively to obtain the normal segmented audio features corresponding to each normal segmented audio data; The step of using the N target Fourier sizes to extract features from the abnormal mono audio data to obtain the abnormal audio features corresponding to the abnormal mono audio data includes: Obtain the abnormal audio duration of the abnormal mono audio data and the audio processing duration threshold, and based on the abnormal audio duration and the audio processing duration threshold, divide the abnormal mono audio data to obtain h abnormal segmented audio data; h is a positive integer; Use the N target Fourier sizes to extract features from the h abnormal segmented audio data respectively to obtain the abnormal segmented audio features corresponding to each abnormal segmented audio data.

11. The method according to claim 10, wherein, The method further includes: Add a segmented normal quality result label to each of the f normal segmented audio features, and add a segmented abnormal quality result label to each of the h abnormal segmented audio features; the object normal quality result label includes f segmented normal quality result labels, and the object abnormal quality result label includes h segmented abnormal quality result labels; The step of using the normal audio features, the abnormal audio features, the object normal quality result label, and the object abnormal quality result label to adjust the parameters of the initial detection model to obtain the audio detection model includes: Input the f normal segmented audio features into the initial detection model respectively, and output the first quality prediction results corresponding to the f normal segmented audio features through the initial detection model; Generate a first loss function based on the first quality prediction result corresponding to each normal segmented audio feature and the segmented normal quality result label; the number of the first loss functions is f; Input the h abnormal segmented audio features into the initial detection model respectively, and output the second quality prediction results corresponding to the h abnormal segmented audio features respectively through the initial detection model; Generate a second loss function based on the second quality prediction result corresponding to each abnormal segmented audio feature and the segmented abnormal quality result label; the number of the second loss functions is h; Based on the f first loss functions and the h second loss functions, adjust the parameters of the initial detection model to obtain the audio detection model.

12. The method according to claim 9, wherein, The step of using the normal audio features, the abnormal audio features, the object normal quality result label and the object abnormal quality result label to adjust the parameters of the initial detection model to obtain the audio detection model includes: Using the normal audio features and the object normal quality result label to perform positive parameter adjustment on the initial detection model; Using the abnormal audio features and the object abnormal quality result label to perform negative parameter adjustment on the initial detection model; Determine the initial detection model after the positive parameter adjustment and the negative parameter adjustment as the audio detection model.

13. The method according to claim 9, wherein, The step of collecting the normal audio data generated by the normal audio output object and the abnormal audio data generated by the abnormal audio output object includes: Obtain the audio test duration, collect the first initial audio data generated by the normal audio output object, and obtain the second initial audio data generated by the abnormal audio output object; Obtain the first test audio data corresponding to the audio test duration from the first initial audio data, and determine the first initial audio data corresponding to the first test audio data in the qualified test state as the normal audio data; Obtain the second test audio data corresponding to the audio test duration from the second initial audio data, and determine the second initial audio data corresponding to the second test audio data in the qualified test state as the abnormal audio data.

14. A computer device, wherein, It includes a processor, a memory, and an input / output interface; The processor is respectively connected to the memory and the input / output interface, wherein the input / output interface is used to receive and output data, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1-13.

15. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor so that a computer device with the processor executes the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Generative speech separation method and device introducing fundamental frequency clues

    CN115910091A

  • Noise perception time domain voice separation method based on joint constraint and shared encoder

    CN117524243A