Engine model detection method and device, electronic device, and storage medium
By acquiring engine audio for noise detection and feature extraction, and using voiceprint feature comparison technology to automatically identify engine models, the problem of low efficiency in manual detection is solved, and efficient engine model identification is achieved.
Patent Information
- Application Number
- CN202210974438.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-15
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-08-15
AI Technical Summary
In the current technology, automobile engine testing requires manual inspection by professional technicians, which is inefficient, time-consuming, and cannot accurately identify the engine model.
By acquiring engine audio, performing noise detection and feature extraction, automatically identifying the engine model using voiceprint feature comparison technology, and finding the engine model using a preset mapping relationship.
It enables automatic identification of engine models, improves testing efficiency, and avoids the time wasted on manual testing and model identification errors.
Smart Images

Figure CN115359800B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an engine model detection method and device, electronic equipment, and storage medium. Background Technology
[0002] In related technologies, car engine testing is carried out manually. However, manual testing requires professional engine testing skills, is time-consuming, has low testing efficiency, and cannot identify the engine model. Summary of the Invention
[0003] The main objective of this application is to provide an engine model detection method, device, electronic equipment, and storage medium that can automatically identify engine models and improve engine detection efficiency.
[0004] To achieve the above objectives, a first aspect of this application provides a method for detecting engine model numbers, comprising:
[0005] Acquire the initial audio of the engine to be tested;
[0006] The initial audio is subjected to silence detection to obtain the target audio;
[0007] If the target audio meets the preset conditions, then feature extraction is performed on the target audio to obtain the target voiceprint features corresponding to the target audio;
[0008] Obtain the voiceprint features of the sample;
[0009] The target voiceprint features are compared with the sample voiceprint features to obtain the comparison results;
[0010] If the comparison result shows that the target voiceprint feature matches the sample voiceprint feature, then the corresponding engine model is found from the preset mapping relationship based on the sample voiceprint feature, and the model of the engine to be detected is obtained based on the engine model.
[0011] In some embodiments, the step of performing silence detection on the initial audio to obtain the target audio includes:
[0012] Silence detection is performed on the initial audio to obtain speech segments and non-speech segments;
[0013] The target audio is obtained by removing the non-speech segments from the initial audio.
[0014] In some embodiments, if the target audio meets preset conditions, feature extraction is performed on the target audio to obtain the target voiceprint features corresponding to the target audio, including:
[0015] Calculate the signal-to-noise ratio, amplitude cutoff, and volume of the target audio;
[0016] If the signal-to-noise ratio is greater than a preset first threshold, the cutoff is greater than a preset second threshold, and the volume is greater than a preset third threshold, then feature extraction is performed on the target audio to obtain the target voiceprint features.
[0017] In some embodiments, if the signal-to-noise ratio is greater than a preset first threshold, the cutoff is greater than a preset second threshold, and the volume is greater than a preset third threshold, then feature extraction is performed on the target audio to obtain the target voiceprint features, including:
[0018] If the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, then the target audio is subjected to first feature extraction to obtain initial voiceprint features;
[0019] The initial voiceprint features are input into a preset feature extraction model, and a second feature extraction is performed on the initial voiceprint features to obtain the target voiceprint features.
[0020] In some embodiments, if the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, then a first feature extraction is performed on the target audio to obtain initial voiceprint features, including:
[0021] If the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, calculate the power spectrum of the target audio.
[0022] The power spectrum is subjected to Mel filtering based on the Mel filter bank to obtain the energy of the first filter bank.
[0023] The energy of the first filter bank is logarithmically transformed to obtain the energy of the second filter bank;
[0024] The initial acoustic signature features are obtained based on the energy of the second filter bank.
[0025] In some embodiments, obtaining the initial voiceprint features based on the energy of the second filter bank includes:
[0026] The static acoustic signature is obtained by performing a discrete cosine transform on the energy of the second filter bank.
[0027] Calculate the first-order and second-order difference parameters of the static voiceprint features;
[0028] Calculate the short-time energy of the target audio;
[0029] Calculate the first-order difference energy and the second-order difference energy based on the short-time energy;
[0030] Dynamic voiceprint features are obtained based on the first-order difference parameters, the second-order difference parameters, the first-order difference energy, and the second-order difference energy.
[0031] The initial voiceprint features are obtained based on the static voiceprint features and the dynamic voiceprint features.
[0032] In some embodiments, comparing the target voiceprint features with the sample voiceprint features to obtain a comparison result includes:
[0033] Calculate the cosine similarity between the target voiceprint features and the sample voiceprint features;
[0034] The cosine similarity is used as the score for comparing the target voiceprint feature with the sample voiceprint feature;
[0035] If the score value is greater than the preset score threshold, then the comparison result between the target voiceprint feature and the sample voiceprint feature is determined to be a match between the target voiceprint feature and the sample voiceprint feature.
[0036] A second aspect of this application provides an engine model detection device, comprising:
[0037] The first acquisition module is used to acquire the initial audio of the engine to be detected;
[0038] A silence detection module is used to perform silence detection on the initial audio to obtain the target audio;
[0039] The feature extraction module is used to extract features from the target audio if the target audio meets preset conditions, so as to obtain the target voiceprint features corresponding to the target audio.
[0040] The second acquisition module is used to acquire the voiceprint features of the samples;
[0041] The feature comparison module is used to compare the target voiceprint features with the sample voiceprint features to obtain the comparison result;
[0042] An engine model detection module is used to find the corresponding engine model from a preset mapping relationship based on the sample voiceprint features if the comparison result shows that the target voiceprint feature matches the sample voiceprint feature, and to obtain the model of the engine to be detected based on the engine model.
[0043] A third aspect of this application provides an electronic device comprising a memory and a processor, wherein the memory stores a program, and when the program is executed by the processor, the processor performs an engine model detection method as described in any one of the embodiments of the first aspect of this application.
[0044] A fourth aspect of this application provides a storage medium that is a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the engine model detection method as described in any one of the embodiments of the first aspect of this application.
[0045] The engine model detection method, device, electronic equipment, and storage medium proposed in this application acquire the initial audio of the engine to be detected, perform silence detection on the initial audio to obtain the target audio, and if the target audio meets preset conditions, perform feature extraction on the target audio to obtain the target voiceprint features corresponding to the target audio, obtain sample voiceprint features, compare the target voiceprint features with the sample voiceprint features to obtain the comparison result, and if the comparison result shows that the target voiceprint features match the sample voiceprint features, then find the corresponding engine model from the preset mapping relationship based on the sample voiceprint features, and obtain the model of the engine to be detected based on the engine model. This can automatically identify the model of the engine to be detected, improving the engine detection efficiency. Attached Figure Description
[0046] Figure 1 This is a first flowchart of the engine model detection method provided in the embodiments of this application;
[0047] Figure 2 yes Figure 1 The flowchart of step S120 in the middle;
[0048] Figure 3 yes Figure 1 The flowchart of step S130 in the process;
[0049] Figure 4 yes Figure 3 The flowchart of step S320 in the text;
[0050] Figure 5 yes Figure 4 The flowchart of step S410 in the middle;
[0051] Figure 6 yes Figure 5 The flowchart of step S540 in the text;
[0052] Figure 7 yes Figure 1 The flowchart of step S150 in the middle;
[0053] Figure 8 This is a second flowchart of the engine model detection method provided in the embodiments of this application;
[0054] Figure 9 This is a block diagram of the module structure of the engine model detection device provided in the embodiments of this application;
[0055] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.
[0057] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0059] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., may be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0060] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0061] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0062] First, let's analyze some of the terms used in this application:
[0063] Artificial Intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0064] Voice Activity Detector (VAD): Used to detect the state of speech to determine whether the speech is silent or active.
[0065] Short Time Energy (STE) refers to the energy of a single frame of speech signal.
[0066] Zero Crossing Counter (ZCC): refers to the number of times a frame of speech signal crosses zero in the time domain.
[0067] Signal-to-noise ratio (SNR): This refers to the ratio of signal to noise in an electronic system. It can be obtained from the ratio of average signal power to average noise power, or from the ratio of signal voltage to noise voltage.
[0068] Amplitude cutoff: This limits the amplitude of a signal to a certain maximum value.
[0069] Volume: also known as loudness, refers to the size or intensity of a sound. The volume depends on the amplitude; the larger the amplitude, the louder the volume; conversely, the smaller the amplitude, the quieter the volume.
[0070] Filter bank (Fbank) features are frequency domain features based on cepstral extraction. Because they better match auditory response characteristics, they have become a commonly used audio feature in speech recognition.
[0071] Mel-Frequency Cepstral Coefficient (MFCC): These are the coefficients that make up the Mel-Frequency Cepstral, which is a linear transformation of the logarithmic energy spectrum based on the nonlinear Mel-scale of sound frequencies.
[0072] In related technologies, car engine testing is carried out through manual inspection. However, this method not only requires professional engine testing skills, but also consumes a lot of time, has low testing efficiency, and may not be able to detect the engine model.
[0073] Based on this, embodiments of this application propose an engine model detection method, device, electronic device, and storage medium. By extracting the target voiceprint features of the audio of the engine to be detected, the target voiceprint features are compared with sample voiceprint features. When the comparison result shows that the target voiceprint features match the sample voiceprint features, the corresponding engine model is found from a preset mapping relationship based on the sample voiceprint features. The model of the engine to be detected is obtained based on the engine model. This method can automatically identify the model of the engine to be detected, avoid engine modification, and eliminate the need to spend a lot of time detecting the engine, thus improving the engine detection efficiency.
[0074] The engine model detection method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the engine model detection method in this application is described.
[0075] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0076] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0077] The engine model detection method provided in this application relates to the fields of artificial intelligence and voiceprint recognition. This engine model detection method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, or smartwatch, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the engine model detection method, but is not limited to the above forms.
[0078] The embodiments of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0079] Reference Figure 1 The engine model detection method according to the first aspect of the embodiments of this application includes, but is not limited to, steps S110 to S160.
[0080] Step S110: Obtain the initial audio of the engine to be tested;
[0081] Step S120: Perform silence detection on the initial audio to obtain the target audio;
[0082] Step S130: If the target audio meets the preset conditions, then the target audio is subjected to feature extraction to obtain the target voiceprint features corresponding to the target audio.
[0083] Step S140: Obtain the voiceprint features of the sample;
[0084] Step S150: Compare the target voiceprint features with the sample voiceprint features to obtain the comparison results;
[0085] Step S160: If the comparison result shows that the target voiceprint feature matches the sample voiceprint feature, then the corresponding engine model is found from the preset mapping relationship based on the sample voiceprint feature, and the model of the engine to be detected is obtained based on the engine model.
[0086] In step S110 of some embodiments, the initial audio of the engine to be tested is obtained, wherein the initial audio is the sound emitted by the engine to be tested when it is working.
[0087] In step S120 of some embodiments, in order to reduce the amount of data processing in the engine model detection process, the initial audio is subjected to silence detection to determine the silent segments and non-silent segments in the initial audio, and the silent segments are removed from the initial audio to obtain the target audio.
[0088] In step S130 of some embodiments, if the target audio meets preset conditions, it indicates that the speech quality of the target audio is high and the engine to be detected is working normally. In this case, feature extraction is performed on the target audio to obtain target voiceprint features. If the target audio does not meet the preset conditions, it indicates that the speech quality of the target audio is low, the engine to be detected is malfunctioning, or the engine to be detected has been modified and the current model does not match the initial model. In this case, no further operations are performed. By determining whether the speech quality of the target audio meets the preset conditions, the engine to be detected can be initially screened. If the state of the engine to be detected is abnormal, no further operations are performed, thus improving engine detection efficiency.
[0089] In step S140 of some embodiments, sample audio of the sample engine is acquired, and feature extraction is performed on the sample audio to obtain sample voiceprint features, wherein the sample voiceprint features are the voiceprint features corresponding to the sample engine in the voiceprint database. It is understood that the voiceprint database stores the mapping relationship between sample engine models and sample voiceprint features. It should be noted that the method for extracting sample voiceprint features is the same as the method for extracting target voiceprint features, and will not be repeated here.
[0090] In step S150 of some embodiments, the target voiceprint feature is compared with the sample voiceprint feature to obtain a comparison result. The comparison result is used to characterize the similarity between the target voiceprint feature and the sample voiceprint feature. If the similarity between the two is greater than a preset threshold, it means that the target voiceprint feature and the sample voiceprint feature match. If the similarity between the two is less than or equal to the preset threshold, it means that the target voiceprint feature and the sample voiceprint feature do not match.
[0091] In step S160 of some embodiments, if the comparison result is that the target voiceprint feature matches the sample voiceprint feature, then the sample engine model corresponding to the sample voiceprint feature is found from the mapping relationship of the voiceprint database, and the sample engine model is used as the model of the engine to be detected.
[0092] The engine model detection method of this application embodiment obtains the initial audio of the engine to be detected, performs silence detection on the initial audio to obtain the target audio, and if the target audio meets the preset conditions, performs feature extraction on the target audio to obtain the target voiceprint features corresponding to the target audio, obtains the sample voiceprint features, compares the target voiceprint features with the sample voiceprint features to obtain the comparison result, and if the comparison result shows that the target voiceprint features match the sample voiceprint features, then finds the corresponding engine model from the preset mapping relationship based on the sample voiceprint features, and obtains the model of the engine to be detected based on the engine model. This method can automatically identify the model of the engine to be detected, thereby improving the engine detection efficiency.
[0093] In some embodiments, such as Figure 2 As shown, step S120 specifically includes, but is not limited to, steps S210 to S220.
[0094] Step S210: Perform silence detection on the initial audio to obtain speech segments and non-speech segments;
[0095] Step S220: Remove non-speech segments from the initial audio to obtain the target audio.
[0096] In step S210 of some embodiments, silence detection can be performed on the initial audio using a method based on short-time energy and zero-crossing rate or a Gaussian mixture model to obtain speech segments and non-speech segments.
[0097] The steps for silence detection using short-time energy and zero-crossing rate are as follows: Obtain a preset signal sampling rate; sample the initial audio according to the signal sampling rate to obtain sampled audio; obtain a preset frame length; divide the sampled audio into frames according to the frame length to obtain framed audio; obtain the short-time energy based on the sum of squared amplitudes of the framed audio; calculate the zero-crossing rate of the framed audio; if the short-time energy is less than a preset short-time energy threshold and the zero-crossing rate is greater than a preset zero-crossing rate threshold, then the corresponding framed audio is a non-speech segment; otherwise, the corresponding framed audio is a speech segment. The signal sampling rate can be 8000Hz, and the frame length can be 20ms.
[0098] The steps for silence detection using a Gaussian mixture model are as follows: Obtain a preset signal sampling rate; sample the initial audio according to the sampling rate to obtain sampled audio; obtain a preset frame length; divide the sampled audio into frames according to the frame length to obtain framed audio; divide the framed audio into multiple sub-bands in the frequency domain; calculate the logarithmic energy feature of each sub-band; sum the logarithmic energy features of all sub-bands to obtain the global energy feature; if the global energy feature is greater than the energy threshold, input the logarithmic energy feature of the sub-bands into the Gaussian mixture model to calculate... The sub-frequency bands have a first probability and a second probability, where the first probability represents the probability that the audio segment corresponding to the sub-frequency band is a speech segment, and the second probability represents the probability that the audio segment corresponding to the sub-frequency band is a non-speech segment. The local log-likelihood ratio of the sub-frequency band is obtained based on the ratio of the first probability to the second probability. The global log-likelihood ratio is obtained by weighted summing the local log-likelihood ratios of all sub-frequency bands. If the local log-likelihood ratio of each sub-frequency band is greater than the first likelihood ratio threshold or the global log-likelihood ratio is greater than the second likelihood ratio threshold, then the framed audio is determined to be a speech segment; otherwise, it is a non-speech segment. The signal sampling rate can be 8kHz, the frame length can be 20ms, and the number of sub-frequency bands can be 6, with frequency ranges of [80Hz, 250Hz], [250Hz, 500Hz], [500Hz, 1kHz], [1kHz, 2kHz], [2kHz, 3kHz], and [3kHz, 4kHz].
[0099] In step S220 of some embodiments, non-speech segments are removed from the initial audio to obtain the target audio, wherein the non-speech segments can be noisy audio segments, silent segments, etc.
[0100] By executing steps S210 to S220, the target audio is made into a valid speech segment, avoiding the subsequent feature extraction of invalid speech segments, reducing the amount of data processing in the engine model detection process, and improving the efficiency of engine model detection.
[0101] In some embodiments, such as Figure 3 As shown, step S130 specifically includes, but is not limited to, steps S310 to S320.
[0102] Step S310: Calculate the signal-to-noise ratio, cutoff, and volume of the target audio.
[0103] Step S320: If the signal-to-noise ratio is greater than a preset first threshold, the cutoff is greater than a preset second threshold, and the volume is greater than a preset third threshold, then feature extraction is performed on the target audio to obtain the target voiceprint features.
[0104] In step S310 of some embodiments, the signal-to-noise ratio (SNR), cutoff, and volume of the target audio are calculated, and the speech quality of the target audio is measured by the magnitude of the SNR, cutoff, and volume. This can be achieved by adding a specific noise, such as Gaussian noise, to the target audio, calculating a first active power of the target audio and a second active power of the specific noise, and obtaining the SNR of the target audio by the ratio of the first active power to the second active power; or by calculating a first average amplitude of the target audio and a second average amplitude of the specific noise, and obtaining the SNR of the target audio by the ratio of the first average amplitude to the second average amplitude. Alternatively, the target audio can be separated into effective audio and noise audio using a preset filter, and the SNR can be obtained based on the first active power of the effective audio and the second active power of the noise audio, or based on the first average amplitude of the effective audio and the second average amplitude of the noise audio. The amplitude of the target audio is limited to a range less than a certain maximum value, and this maximum value is used as the calculated cutoff. The amplitude of the target audio is limited to a range greater than a certain minimum value, and this minimum value is used as the calculated volume.
[0105] In step S320 of some embodiments, if the signal-to-noise ratio is greater than a preset first threshold, the cutoff is greater than a preset second threshold, and the volume is greater than a preset third threshold, it indicates that the speech quality of the target audio is high and the engine to be detected can work normally. Feature extraction is performed on the target audio to obtain the target voiceprint features. Otherwise, it indicates that the speech quality of the target audio is low, the engine to be detected is malfunctioning, or the engine to be detected has been modified. In this case, no further operations are performed.
[0106] By executing steps S310 to S320, engines under test with abnormal operating conditions can be preliminarily screened, thereby improving engine testing efficiency.
[0107] In some embodiments, such as Figure 4 As shown, step S320 specifically includes, but is not limited to, steps S410 to S420.
[0108] Step S410: If the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, then the first feature is extracted from the target audio to obtain the initial voiceprint features.
[0109] Step S420: Input the initial voiceprint features into the preset feature extraction model, perform second feature extraction on the initial voiceprint features, and obtain the target voiceprint features.
[0110] In step S420 of some embodiments, the target voiceprint features are subjected to algorithmic feature extraction to generate a new voiceprint feature vector. It should be noted that algorithmic feature extraction involves reducing the dimensionality of the target voiceprint features or recombinating the target voiceprint features to generate a new voiceprint feature vector.
[0111] In steps S410 to S420 of some embodiments, if the signal-to-noise ratio of the target audio is greater than a first threshold, the cutoff is greater than a second threshold, and the volume is greater than a third threshold, the initial voiceprint features of the target audio are extracted. The initial voiceprint features can be Mel frequency cepstral coefficients, fundamental frequency, filter bank energy, etc. The initial voiceprint features are input into a convolutional neural network or a recurrent neural network for deep voiceprint feature extraction to obtain the target voiceprint features. The convolutional neural network can be a VGG network, a ResNet network, etc., and the recurrent neural network can be an LSTM network, a GRU network, a bidirectional LSTM network, etc.
[0112] In some embodiments, such as Figure 5 As shown, step S410 specifically includes, but is not limited to, steps S510 to S540.
[0113] Step S510: If the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, calculate the power spectrum of the target audio.
[0114] Step S520: Perform Mel filtering on the power spectrum based on the Mel filter bank to obtain the energy of the first filter bank;
[0115] Step S530: Perform a logarithmic transformation on the energy of the first filter bank to obtain the energy of the second filter bank;
[0116] Step S540: Obtain the initial voiceprint features based on the energy of the second filter bank.
[0117] In step S510 of some embodiments, if the signal-to-noise ratio is greater than a first threshold, the cutoff is greater than a second threshold, and the volume is greater than a third threshold, the target audio undergoes preprocessing such as pre-emphasis, framing, and windowing to obtain preprocessed audio. A short-time Fourier transform is then performed on the preprocessed audio to obtain its corresponding spectrum, and the power spectrum of the target audio is calculated based on the spectrum. To reduce the truncation effect, the window function in the windowing process can be a Hamming window or a Hanning window.
[0118] In step S520 of some embodiments, the power spectrum is subjected to Mel filtering based on the Mel filter bank to obtain the energy of the first filter bank, wherein the energy of the first filter bank is the FBank feature without logarithmic transformation.
[0119] In step S530 of some embodiments, the energy of the first filter bank is logarithmically transformed to obtain the energy of the second filter bank, wherein the energy of the second filter bank is an FBank feature.
[0120] In step S540 of some embodiments, the FBank features of each frame of audio can be used as the initial voiceprint features of the target audio.
[0121] In steps S510 to S540 of some embodiments, if the signal-to-noise ratio of the target audio is greater than a first threshold, the cutoff is greater than a second threshold, and the volume is greater than a third threshold, a high-pass filter is used to pre-emphasize the target audio to enhance the high-frequency part of the target audio. The enhanced target audio is then divided into frames to obtain multiple framed audios. A preset window function is multiplied with the framed audios and a short-time Fourier transform is performed to obtain the spectrum of the framed audios. The power spectrum is obtained by squared the spectrum value. The power spectrum is then subjected to Mel filtering using multiple Mel filters to obtain the energy of the first filter group. The energy of the first filter group is then subjected to logarithmic transformation to obtain the energy of the second filter group. The energy of the second filter group of all framed audios is used as the initial voiceprint feature of the target audio.
[0122] In some embodiments, such as Figure 6 As shown, step S540 specifically includes, but is not limited to, steps S610 to S660.
[0123] Step S610: Perform discrete cosine transform on the energy of the second filter bank to obtain static acoustic signature features;
[0124] Step S620: Calculate the first-order difference parameters and second-order difference parameters of the static voiceprint features;
[0125] Step S630: Calculate the short-time energy of the target audio.
[0126] Step S640: Calculate the first-order difference energy and the second-order difference energy based on the short-time energy;
[0127] Step S650: Obtain dynamic voiceprint features based on first-order difference parameters, second-order difference parameters, first-order difference energy, and second-order difference energy;
[0128] Step S660: Obtain initial voiceprint features based on static and dynamic voiceprint features.
[0129] In step S610 of some embodiments, the FBank features are subjected to discrete cosine transform to obtain 12-dimensional Mel frequency cepstral coefficients, which are then used as static voiceprint features.
[0130] In step S620 of some embodiments, the first-order difference parameters of the static voiceprint features are calculated, and the second-order difference parameters are obtained based on the first-order difference parameters. The calculation method of the first-order difference parameters is shown in formula (1).
[0131]
[0132] Wherein, d1(n) is the nth first-order difference parameter, c(n) is the nth Mel frequency cepstral coefficient, N1 is the order of the Mel frequency cepstral coefficient, N1 is 12 in this embodiment, and M1 is the time difference of the first-order difference parameter, which can be 1 or 2.
[0133] The calculation method for obtaining the second-order difference parameter from the first-order difference parameter is shown in formula (2).
[0134]
[0135] Where d2(n) is the nth first-order difference parameter, d1(n) is the nth first-order difference parameter, N2 is the order of the first-order difference parameter, and M2 is the time difference of the second-order difference parameter, which can take the value of 1 or 2.
[0136] In step S630 of some embodiments, the target audio is preprocessed by pre-emphasis, framing, windowing, etc., and the short-time energy corresponding to all frames in the target audio is calculated.
[0137] In step S640 of some embodiments, the first-order difference energy is calculated based on the short-time energy, and the second-order difference energy is calculated based on the first-order difference energy. If the short-time energy of the nth frame audio is e(n), and the first-order difference energy of the nth frame is Δe(n), the calculation method of the first-order difference energy is shown in formula (3).
[0138]
[0139] Where N is the total number of frames in the target audio.
[0140] The calculation method for the second-order difference energy is shown in formula (4).
[0141]
[0142] In step S650 of some embodiments, the first-order difference parameters, second-order difference parameters, first-order difference energy, and second-order difference energy of all frames in the target audio are used as dynamic voiceprint features.
[0143] In step S660 of some embodiments, the fundamental frequency of all frames in the target audio is extracted, and the static voiceprint features, dynamic voiceprint features and fundamental frequency of all frames are used as initial voiceprint features.
[0144] In steps S610 to S660 of some embodiments, a discrete cosine transform is performed on the energy of the second filter bank to obtain 12-dimensional Mel frequency cepstral coefficients. First-order difference parameters are obtained from the Mel frequency cepstral coefficients, and second-order difference parameters are obtained from the first-order difference parameters. The short-time energy of each frame audio is calculated. First-order difference energy is obtained from the short-time energy of two adjacent frame audios, and second-order difference energy is obtained from the two adjacent first-order difference energies. The short-time average amplitude difference of the frame audio is calculated. The pitch period is obtained from the short-time average amplitude difference, and the fundamental frequency is obtained from the reciprocal of the pitch period. The Mel frequency cepstral coefficients, first-order difference parameters, second-order difference parameters, first-order difference energy, second-order difference energy, and fundamental frequency are used as initial voiceprint features.
[0145] In some embodiments, such as Figure 7 As shown, step S150 specifically includes, but is not limited to, steps S710 to S730.
[0146] Step S710: Calculate the cosine similarity between the target voiceprint features and the sample voiceprint features;
[0147] Step S720: Use cosine similarity as the score for comparing the target voiceprint features with the sample voiceprint features.
[0148] Step S730: If the score value is greater than the preset score threshold, the comparison result between the target voiceprint feature and the sample voiceprint feature is determined to be a match between the target voiceprint feature and the sample voiceprint feature.
[0149] In step S710 of some embodiments, the cosine similarity between the target voiceprint feature and the sample voiceprint features in the voiceprint database is calculated. The cosine similarity characterizes the similarity between the target and sample voiceprint features; the more similar the target and sample voiceprint features are, the closer the cosine similarity value is to 1. A smaller cosine similarity indicates a greater difference between the target and sample voiceprint features, and vice versa.
[0150] In step S720 of some embodiments, in order to obtain the comparison result of comparing the target voiceprint features with the sample voiceprint features, the cosine similarity is used as the score value of the comparison.
[0151] In step S730 of some embodiments, if the score value is greater than a preset score threshold, it indicates that the similarity between the target voiceprint feature and the sample voiceprint feature is high, and the comparison result is determined to be a match between the target voiceprint feature and the sample voiceprint feature. If the score value is less than or equal to the score threshold, it indicates that the target voiceprint feature does not match the sample voiceprint feature, and the process continues to traverse the next sample voiceprint feature in the voiceprint library until the target voiceprint feature matches the sample voiceprint feature. If after traversing each sample voiceprint feature in the voiceprint library, the comparison result is that the target voiceprint feature does not match the sample voiceprint feature, then it is determined that the engine to be detected has been modified, and the current model does not match the original model.
[0152] By executing steps S710 to S730, the comparison results of the target voiceprint features and the sample voiceprint features can be obtained based on cosine similarity, which makes it easier to determine the model of the engine to be detected.
[0153] Reference Figure 8 Another embodiment of this application proposes an engine model detection method, including but not limited to steps S810 to S880.
[0154] Step S810: Obtain the first initial audio of the sample engine;
[0155] Step S820: Perform silence detection on the first initial audio, remove the first non-speech segment from the first initial audio, and obtain the first target audio;
[0156] Step S830: If the signal-to-noise ratio of the first target audio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, feature extraction is performed on the first target audio to obtain the first initial voiceprint feature of the first target audio. The first initial voiceprint feature is input into the neural network model, and feature extraction is performed on the first initial voiceprint feature to obtain the sample voiceprint feature. The sample voiceprint feature and its corresponding sample engine model are stored in the voiceprint database.
[0157] Step S840: Obtain the second initial audio of the engine to be detected;
[0158] Step S850: Perform silence detection on the second initial audio, remove the second non-speech segment from the second initial audio, and obtain the second target audio;
[0159] Step S860: If the signal-to-noise ratio of the second target audio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, feature extraction is performed on the second target audio to obtain the second initial voiceprint feature of the second target audio. The second initial voiceprint feature is input into the neural network model to extract features from the second initial voiceprint feature to obtain the target voiceprint feature.
[0160] Step S870: Read sample voiceprint features from the voiceprint database, calculate the cosine similarity between the target voiceprint features and the sample voiceprint features. If the cosine similarity is greater than a preset fourth threshold, the comparison result between the target voiceprint features and the sample voiceprint features is determined to be a match. If the cosine similarity is less than or equal to the fourth threshold, the comparison result between the target voiceprint features and the sample voiceprint features is determined to be a mismatch.
[0161] In step S880, if the comparison result shows that the target voiceprint feature matches the sample voiceprint feature, the sample engine model is retrieved from the voiceprint database based on the sample voiceprint feature, and the sample engine model is used as the model of the engine to be detected; if the comparison result shows that the target voiceprint feature does not match the sample voiceprint feature, the next sample voiceprint feature is read from the voiceprint database until all sample voiceprint features in the voiceprint database are completely traversed.
[0162] This application also provides an engine model detection device, such as... Figure 9 As shown, the above-mentioned engine model detection method can be implemented. The device includes a first acquisition module 910, a silence detection module 920, a feature extraction module 930, a second acquisition module 940, a feature comparison module 950, and an engine model detection module 960. The first acquisition module 910 is used to acquire the initial audio of the engine to be detected; the silence detection module 920 is used to perform silence detection on the initial audio to obtain the target audio; the feature extraction module 930 is used to extract features from the target audio if the target audio meets preset conditions to obtain the target voiceprint features corresponding to the target audio; the second acquisition module 940 is used to acquire sample voiceprint features; the feature comparison module 950 is used to compare the target voiceprint features with the sample voiceprint features to obtain the comparison result; the engine model detection module 960 is used to find the corresponding engine model from a preset mapping relationship based on the sample voiceprint features if the comparison result shows a match between the target voiceprint features and the sample voiceprint features, and obtain the model of the engine to be detected based on the engine model.
[0163] The engine model detection device of this application embodiment is used to execute the engine model detection method in the above embodiment. Its specific processing procedure is the same as that of the engine model detection method in the above embodiment, and will not be described in detail here.
[0164] The engine model detection device of this application embodiment acquires the initial audio of the engine to be detected through a first acquisition module, performs silence detection on the initial audio to obtain the target audio, and if the target audio meets the preset conditions, the feature extraction module extracts features from the target audio to obtain the target voiceprint features corresponding to the target audio, the second acquisition module acquires the sample voiceprint features, the feature comparison module compares the target voiceprint features with the sample voiceprint features to obtain the comparison result, and the engine model detection module judges the comparison result. If the comparison result shows that the target voiceprint features match the sample voiceprint features, the corresponding engine model is found from the preset mapping relationship based on the sample voiceprint features, and the model of the engine to be detected is obtained based on the engine model. This device can automatically identify the model of the engine to be detected, thereby improving the engine detection efficiency.
[0165] This application also provides an electronic device, including:
[0166] At least one processor, and,
[0167] A memory that is communicatively connected to at least one processor; wherein,
[0168] The memory stores instructions that are executed by at least one processor to enable the engine model detection method as described in any of the embodiments of the first aspect of this application when the at least one processor executes the instructions.
[0169] The following is combined with Figure 10 The hardware structure of the electronic device is described in detail. The electronic device includes: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050.
[0170] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0171] The memory 1020 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device, or RAM (Random Access Memory). The memory 1020 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and called and executed by the processor 1010 to execute the engine model detection method of the embodiments of this application.
[0172] The input / output interface 1030 is used to implement information input and output;
[0173] The communication interface 1040 is used to enable communication and interaction between this device and other devices. Communication can be achieved via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth).
[0174] Bus 1050 transmits information between various components of the device (e.g., processor 1010, memory 1020, input / output interface 1030, and communication interface 1040);
[0175] The processor 1010, memory 1020, input / output interface 1030 and communication interface 1040 are connected to each other within the device via bus 1050.
[0176] This application also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the engine model detection method of this application.
[0177] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0178] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0179] It will be understood by those skilled in the art that Figures 1 to 8 The technical solutions shown do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0181] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0182] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0183] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0184] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0185] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0186] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0187] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0188] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An engine model detection method, characterized in that, The method includes: Acquire the initial audio of the engine to be tested; wherein, the initial audio is the sound emitted by the engine to be tested when it is working; The initial audio is subjected to silence detection to obtain the target audio; Calculate the signal-to-noise ratio, cutoff, and volume of the target audio; wherein the cutoff is used to limit the amplitude of the target audio to a range less than a specific maximum value, the specific maximum value being the value of the cutoff; and the volume is used to limit the amplitude of the target audio to a range greater than a specific minimum value, the specific minimum value being the value of the volume. If the signal-to-noise ratio is greater than a preset first threshold, the cutoff is greater than a preset second threshold, and the volume is greater than a preset third threshold, then the target audio is subjected to first feature extraction to obtain initial voiceprint features, and the initial voiceprint features are input into a preset feature extraction model for second feature extraction to obtain target voiceprint features. Obtain the voiceprint features of the sample; The target voiceprint features are compared with the sample voiceprint features to obtain the comparison results; If the comparison result is that the target voiceprint feature matches the sample voiceprint feature, then the corresponding engine model is found from the preset mapping relationship based on the sample voiceprint feature, and the model of the engine to be detected is obtained based on the engine model. If the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, then the target audio is subjected to first feature extraction to obtain initial voiceprint features, including: If the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, calculate the power spectrum of the target audio; perform Mel filtering on the power spectrum according to the Mel filter bank to obtain the energy of the first filter bank; perform logarithmic transformation on the energy of the first filter bank to obtain the energy of the second filter bank; obtain the initial voiceprint features based on the energy of the second filter bank. The step of obtaining the initial acoustic signature features based on the energy of the second filter bank includes: The energy of the second filter bank is subjected to discrete cosine transform to obtain static voiceprint features; the first-order difference parameter and the second-order difference parameter of the static voiceprint features are calculated; the short-time energy of the target audio is calculated; the first-order difference energy and the second-order difference energy are calculated based on the short-time energy; the dynamic voiceprint features are obtained based on the first-order difference parameter, the second-order difference parameter, the first-order difference energy, and the second-order difference energy; the fundamental frequency of the target audio is extracted, and the static voiceprint features, the dynamic voiceprint features, and the fundamental frequency are used as initial voiceprint features.
2. The engine model detection method according to claim 1, characterized in that, The step of performing silence detection on the initial audio to obtain the target audio includes: Silence detection is performed on the initial audio to obtain speech segments and non-speech segments; The target audio is obtained by removing the non-speech segments from the initial audio.
3. The engine model detection method according to any one of claims 1 to 2, characterized in that, The step of comparing the target voiceprint features with the sample voiceprint features to obtain the comparison result includes: Calculate the cosine similarity between the target voiceprint features and the sample voiceprint features; The cosine similarity is used as the score for comparing the target voiceprint feature with the sample voiceprint feature; If the score value is greater than the preset score threshold, then the comparison result between the target voiceprint feature and the sample voiceprint feature is determined to be a match between the target voiceprint feature and the sample voiceprint feature.
4. An engine model detection device, characterized in that, The device includes: The first acquisition module is used to acquire the initial audio of the engine to be tested; wherein, the initial audio is the sound emitted by the engine to be tested when it is working; A silence detection module is used to perform silence detection on the initial audio to obtain the target audio; The feature extraction module is used to extract features from the target audio if the target audio meets preset conditions, so as to obtain the target voiceprint features corresponding to the target audio. The second acquisition module is used to acquire the voiceprint features of the samples; The feature comparison module is used to compare the target voiceprint features with the sample voiceprint features to obtain the comparison result; An engine model detection module is used to find the corresponding engine model from a preset mapping relationship based on the sample voiceprint features if the comparison result is that the target voiceprint feature matches the sample voiceprint feature, and to obtain the model of the engine to be detected based on the engine model. The device is also used for: Calculate the signal-to-noise ratio, cutoff, and volume of the target audio; wherein the cutoff is used to limit the amplitude of the target audio to a range less than a specific maximum value, the specific maximum value being the value of the cutoff; and the volume is used to limit the amplitude of the target audio to a range greater than a specific minimum value, the specific minimum value being the value of the volume. If the signal-to-noise ratio is greater than a preset first threshold, the cutoff is greater than a preset second threshold, and the volume is greater than a preset third threshold, then the target audio is subjected to first feature extraction to obtain initial voiceprint features, and the initial voiceprint features are input into a preset feature extraction model for second feature extraction to obtain target voiceprint features. If the signal-to-noise ratio is greater than the first threshold, the cutoff is greater than the second threshold, and the volume is greater than the third threshold, calculate the power spectrum of the target audio; perform Mel filtering on the power spectrum according to the Mel filter bank to obtain the energy of the first filter bank; perform logarithmic transformation on the energy of the first filter bank to obtain the energy of the second filter bank; obtain the initial voiceprint features based on the energy of the second filter bank. The energy of the second filter bank is subjected to discrete cosine transform to obtain static voiceprint features; the first-order difference parameter and the second-order difference parameter of the static voiceprint features are calculated; the short-time energy of the target audio is calculated; the first-order difference energy and the second-order difference energy are calculated based on the short-time energy; the dynamic voiceprint features are obtained based on the first-order difference parameter, the second-order difference parameter, the first-order difference energy, and the second-order difference energy; the fundamental frequency of the target audio is extracted, and the static voiceprint features, the dynamic voiceprint features, and the fundamental frequency are used as initial voiceprint features.
5. An electronic device, characterized in that, include: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes the at least one program to achieve the following: The engine model detection method as described in any one of claims 1 to 4.
6. A storage medium, wherein the storage medium is a computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform: The engine model detection method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Voiceprint recognition method and device
CN107886943A
Voiceprint recognition method based on Android
CN107993663A
Vehicle model identification method and device, computer equipment and storage medium
CN108847253A
Video generation method, video generation device, electronic equipment and storage medium
CN114786059A
Method and device for detecting abnormal condition
JP2000321176A