Sound detection method, device and equipment based on machine vision, and medium

By using a machine vision-based audio detection method, recorded audio is acquired and converted into a time-domain image to detect the acoustic parameters of the audio equipment. This solves the problems of low detection efficiency and high cost in existing technologies, and achieves efficient and accurate acoustic quality control of audio equipment.

CN117395587BActive Publication Date: 2026-08-25GOERTEK INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311526520.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2026-08-25
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

Existing technologies for acoustic performance testing in the development of new products and the NPI (New Product Introduction) stage of smart speakers suffer from low efficiency, high cost, and susceptibility to human factors, making it difficult to achieve efficient and accurate quality control.

Method used

A machine vision-based audio detection method is adopted. By acquiring the recorded audio and converting it into a time-domain image, the detection values ​​of preset detection items of the recorded audio are detected. If the preset threshold is met, the audio is determined to be qualified. The detection parameters include rise time, peak time, overshoot, adjustment time and tilt.

Benefits of technology

It enables low-cost, automated acoustic performance testing, improves testing efficiency and accuracy, reduces the impact of human factors, and ensures the acoustic quality of audio equipment, guaranteeing product quality from the R&D stage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117395587B_ABST
    Figure CN117395587B_ABST
Patent Text Reader

Abstract

The application discloses a sound detection method and device based on machine vision, equipment and medium, and belongs to the technical field of acoustic detection. In the application, the intelligent sound quality detection of new product research and development and NPI stage is realized through the automatic detection technology of machine vision, so as to ensure the product quality from the research and development end. First, the recorded audio of the to-be-detected sound when playing the preset test audio is converted into a time domain image. Then, the detection value of the preset detection item of the recorded audio is detected on the time domain image. If the detection values all meet the preset threshold of the preset detection item corresponding to the detection value, it is determined that the to-be-detected sound is qualified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of acoustic testing, and in particular to a machine vision-based sound testing method, a machine vision-based sound testing device, a machine vision-based sound testing equipment, and a computer-readable storage medium. Background Technology

[0002] Smart speakers combine traditional speakers with emerging technologies such as voice recognition and natural language processing, enabling them to simultaneously provide audio playback, intelligent voice interaction, and smart home control. Due to their rich functionality, smart speakers are considered the control hub of smart homes, thus becoming one of the fastest-growing electronic products today. Audio playback is the fundamental function of smart speakers; therefore, ensuring their acoustic performance is a basic quality requirement. This necessitates acoustic performance testing from the new product development and NPI (New Product Introduction) stages.

[0003] Currently, in traditional manufacturing scenarios, acoustic performance is often evaluated manually by listening during new product development and NPI (New Product Introduction) stages. This method is labor-intensive, inefficient, lacks robustness, is highly susceptible to subjective factors, and prolonged testing can lead to auditory fatigue for inspectors. In high-end manufacturing scenarios, acoustic performance is assessed using complex and precise acoustic instruments. This method is costly, requires stringent testing conditions, and its efficiency needs improvement. Summary of the Invention

[0004] The main objective of this application is to provide a machine vision-based sound detection method, a machine vision-based sound detection device, a machine vision-based sound detection equipment, and a computer-readable storage medium, which aim to accurately detect the acoustic quality of sound.

[0005] To achieve the above objectives, this application provides a machine vision-based sound detection method, the method comprising:

[0006] Acquire the audio recording of the speaker under test while playing a preset test audio, and convert the audio recording into a time-domain image;

[0007] Detect the detection value of a preset detection item for the recorded audio on the time-domain image;

[0008] If all the detected values ​​meet the preset threshold of the preset detection item corresponding to the detected value, then the audio device under test is determined to be qualified.

[0009] For example, the step of detecting the detection value of a preset detection item of the recorded audio on the time-domain image includes:

[0010] The rise time, peak time, and overshoot of the preset test audio are determined in the time-domain image.

[0011] When the preset test audio is a step signal audio, the adjustment time of the step signal audio is also determined in the time domain image;

[0012] When the preset test audio is a square wave signal audio, the tilt of the square wave signal audio is also determined in the time domain image.

[0013] For example, the step of determining the rise time, peak time, and overshoot of the preset test audio in the time-domain image includes:

[0014] The temporal image is binarized, and edge contour lines are extracted through edge detection.

[0015] By tracing the curve of the edge contour and performing local optimization, the starting point, the maximum value point of the vertical axis coordinate, and the steady-state value point of the edge contour are determined, and a two-dimensional coordinate system is established with the starting point as the origin.

[0016] The peak time of the preset test audio is determined as the abscissa value of the maximum value point of the vertical axis.

[0017] The overshoot of the preset test audio is determined as the ratio of the difference between the vertical coordinates of the maximum vertical coordinate point and the steady-state vertical coordinate point to the vertical coordinate of the steady-state vertical coordinate point.

[0018] The rise time of the preset test audio is determined as the abscissa value when the ordinate of the edge contour first reaches the reference ordinate, wherein the reference ordinate is determined by the ordinate of the steady-state point and the preset reference coefficient.

[0019] For example, the step of determining the adjustment time of the step signal audio in the time domain image when the preset test audio is a step signal audio includes:

[0020] The fluctuation range of the ordinate of the steady-state value point is determined based on the ordinate of the steady-state value point and the preset fluctuation coefficient.

[0021] Determine the extreme points of the monotonic variation of the edge contour line;

[0022] If the ordinate of the extreme point is within the ordinate fluctuation range, then before the extreme point, determine the target point where the edge contour line intersects with the maximum or minimum ordinate value of the ordinate fluctuation range.

[0023] The adjustment time of the step signal audio is determined as the x-coordinate value of the target point.

[0024] For example, the step of determining the tilt of the square wave signal audio in the time domain image when the preset test audio is a square wave signal audio includes:

[0025] Determine the tilt point, wherein the x-coordinate of the tilt point is the x-coordinate value when the y-coordinate of the edge contour line first reaches the reference y-coordinate, and the y-coordinate of the tilt point is the reference y-coordinate.

[0026] The tilt of the square wave signal audio is determined as the slope between the tilt point and the starting point.

[0027] For example, before the step of obtaining the audio recording of the speaker under test while playing a preset test audio, the following steps are included:

[0028] The free-field frequency response of the active loudspeaker is measured using a standard microphone, and the measurement result is used as the reference frequency response.

[0029] At the same measurement position of the standard microphone, the microphone under test is repeatedly measured, and the frequency response deviation between the frequency response of the microphone under test and the reference frequency response is recorded. The frequency response deviation is used as the compensation value of the microphone under test.

[0030] Using the standard microphone or the microphone under test compensated based on the compensation value, the audio recording of the speaker under test is collected when playing a preset test audio.

[0031] For example, the method further includes:

[0032] If all the speakers in the current batch are qualified, the speakers corresponding to the current batch will be sent to the next work station.

[0033] If any of the speakers in the current batch are found to be substandard, the number of speakers to be tested will be increased until all speakers in subsequent batches are found to be substandard.

[0034] This application also provides a machine vision-based sound detection device, the machine vision-based sound detection device comprising:

[0035] The acquisition module is used to acquire the audio recordings made by the speaker under test when playing a preset test audio, and convert the audio recordings into a time-domain image;

[0036] A detection module is used to detect the detection value of a preset detection item of the recorded audio on the time-domain image;

[0037] The determination module is used to determine that the audio device under test is qualified if all the detected values ​​meet the preset threshold of the preset detection item corresponding to the detected values.

[0038] The present application also provides an audio detection device based on machine vision. The audio detection device based on machine vision includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the audio detection method based on machine vision as described above are implemented.

[0039] The present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the audio detection method based on machine vision as described above are implemented.

[0040] An audio detection method based on machine vision, an audio detection device based on machine vision, an audio detection device based on machine vision, and a computer-readable storage medium provided in an embodiment of the present application. The recorded audio of a to-be-tested audio when playing a preset test audio is obtained, and the recorded audio is converted into a time-domain image; the detection value of a preset detection item of the recorded audio is detected on the time-domain image; if the detection values all meet the preset thresholds of the preset detection items corresponding to the detection values, it is determined that the to-be-tested audio is qualified.

[0041] In the present application, through the automated detection technology of machine vision, the acoustic quality detection of intelligent audio in the new product R & D and NPI stages is realized, so as to ensure product quality from the R & D end. First, the recorded audio of a to-be-tested audio when playing a preset test audio is converted into a time-domain image; then, the detection value of a preset detection item of the recorded audio is detected on the time-domain image; if the detection values all meet the preset thresholds of the preset detection items corresponding to the detection values, it is determined that the to-be-tested audio is qualified. To ensure the quality of intelligent audio, it is necessary to detect its acoustic performance in the new product R & D and NPI stages. If manual detection is used, it will consume a large amount of human resources, and long-term detection will cause auditory fatigue, and at the same time, missed detection and misdetection are likely to occur; if high-end acoustic instruments are used for detection, the cost is high and the detection environment requirements are relatively high. Through the quality inspection method of intelligent audio laboratory based on machine vision, low-cost automated acoustic performance detection can be realized, and the acoustic intelligence of the audio can be accurately detected. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic structural diagram of an operating device of a hardware operating environment involved in the solution of an embodiment of the present application;

[0043] Figure 2 It is a schematic flowchart of an embodiment of an audio detection method based on machine vision involved in the solution of an embodiment of the present application;

[0044] Figure 3 It is a schematic diagram of an audio detection algorithm of an embodiment of an audio detection method based on machine vision involved in the solution of an embodiment of the present application;

[0045] Figure 4 This is a schematic diagram of step signal detection in an embodiment of the machine vision-based audio detection method involved in the present application.

[0046] Figure 5 This is a schematic diagram of square wave signal detection in an embodiment of the machine vision-based audio detection method involved in the present application.

[0047] Figure 6 This is a schematic diagram of the audio quality inspection process of an embodiment of the machine vision-based audio detection method involved in the present application.

[0048] Figure 7 This is a schematic diagram illustrating an embodiment of the machine vision-based sound detection method involved in the embodiments of this application.

[0049] Figure 8 This is a schematic diagram of a machine vision-based audio detection device involved in an embodiment of this application.

[0050] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0051] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0052] Reference Figure 1 , Figure 1 This is a schematic diagram of the operating device structure of the hardware operating environment involved in the embodiments of this application.

[0053] like Figure 1As shown, the operating device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.

[0054] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the operating equipment and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0055] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and computer programs.

[0056] exist Figure 1 In the illustrated operating device, the network interface 1004 is mainly used for data communication with other devices; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and memory 1005 in the operating device of this application can be disposed in the operating device, and the operating device calls the computer program stored in the memory 1005 through the processor 1001 and performs the following operations:

[0057] Acquire the audio recording of the speaker under test while playing a preset test audio, and convert the audio recording into a time-domain image;

[0058] Detect the detection value of a preset detection item for the recorded audio on the time-domain image;

[0059] If all the detected values ​​meet the preset threshold of the preset detection item corresponding to the detected value, then the audio device under test is determined to be qualified.

[0060] In one embodiment, the processor 1001 may invoke a computer program stored in the memory 1005 and further perform the following operations:

[0061] The step of detecting the detection value of the preset detection item of the recorded audio on the time domain image includes:

[0062] The rise time, peak time, and overshoot of the preset test audio are determined in the time-domain image.

[0063] When the preset test audio is a step signal audio, the adjustment time of the step signal audio is also determined in the time domain image;

[0064] When the preset test audio is a square wave signal audio, the tilt of the square wave signal audio is also determined in the time domain image.

[0065] In one embodiment, the processor 1001 may invoke a computer program stored in the memory 1005 and further perform the following operations:

[0066] The step of determining the rise time, peak time, and overshoot of the preset test audio in the time-domain image includes:

[0067] The temporal image is binarized, and edge contour lines are extracted through edge detection.

[0068] By tracing the curve of the edge contour and performing local optimization, the starting point, the maximum value point of the vertical axis coordinate, and the steady-state value point of the edge contour are determined, and a two-dimensional coordinate system is established with the starting point as the origin.

[0069] The peak time of the preset test audio is determined as the abscissa value of the maximum value point of the vertical axis.

[0070] The overshoot of the preset test audio is determined as the ratio of the difference between the vertical coordinates of the maximum vertical coordinate point and the steady-state vertical coordinate point to the vertical coordinate of the steady-state vertical coordinate point.

[0071] The rise time of the preset test audio is determined as the abscissa value when the ordinate of the edge contour first reaches the reference ordinate, wherein the reference ordinate is determined by the ordinate of the steady-state point and the preset reference coefficient.

[0072] In one embodiment, the processor 1001 may invoke a computer program stored in the memory 1005 and further perform the following operations:

[0073] The step of determining the adjustment time of the step signal audio in the time domain image when the preset test audio is a step signal audio includes:

[0074] The fluctuation range of the ordinate of the steady-state value point is determined based on the ordinate of the steady-state value point and the preset fluctuation coefficient.

[0075] Determine the extreme points of the monotonic variation of the edge contour line;

[0076] If the ordinate of the extreme point is within the ordinate fluctuation range, then before the extreme point, determine the target point where the edge contour line intersects with the maximum or minimum ordinate value of the ordinate fluctuation range.

[0077] The adjustment time of the step signal audio is determined as the x-coordinate value of the target point.

[0078] In one embodiment, the processor 1001 may invoke a computer program stored in the memory 1005 and further perform the following operations:

[0079] The step of determining the tilt of the square wave signal audio in the time domain image when the preset test audio is a square wave signal audio includes:

[0080] Determine the tilt point, wherein the x-coordinate of the tilt point is the x-coordinate value when the y-coordinate of the edge contour line first reaches the reference y-coordinate, and the y-coordinate of the tilt point is the reference y-coordinate.

[0081] The tilt of the square wave signal audio is determined as the slope between the tilt point and the starting point.

[0082] In one embodiment, the processor 1001 may invoke a computer program stored in the memory 1005 and further perform the following operations:

[0083] Before the step of acquiring the audio recording of the speaker under test while playing a preset test audio, the following steps are included:

[0084] The free-field frequency response of the active loudspeaker is measured using a standard microphone, and the measurement result is used as the reference frequency response.

[0085] At the same measurement position of the standard microphone, the microphone under test is repeatedly measured, and the frequency response deviation between the frequency response of the microphone under test and the reference frequency response is recorded. The frequency response deviation is used as the compensation value of the microphone under test.

[0086] Using the standard microphone or the microphone under test compensated based on the compensation value, the audio recording of the speaker under test is collected when playing a preset test audio.

[0087] In one embodiment, the processor 1001 may invoke a computer program stored in the memory 1005 and further perform the following operations:

[0088] The method further includes:

[0089] If all the speakers in the current batch are qualified, the speakers corresponding to the current batch will be sent to the next work station.

[0090] If any of the speakers in the current batch are found to be substandard, the number of speakers to be tested will be increased until all speakers in subsequent batches are found to be substandard.

[0091] This application provides a machine vision-based sound detection method, referring to... Figure 2 In one embodiment of the machine vision-based sound detection method, the method includes:

[0092] Step S10: Obtain the audio recording of the speaker under test when playing a preset test audio, and convert the audio recording into a time-domain image;

[0093] To ensure the acoustic performance of smart speakers, quality testing must begin at the new product development and NPI (New Product Introduction) stages. Testing is conducted in a quality laboratory. First, the product under test (whole unit or module) plays test audio. A standard microphone collects the audio signal, which is then sent to a host computer via a serial port. The host computer converts the audio signal into a time-domain image, and then applies a machine vision-based detection algorithm to the time-domain image to evaluate its acoustic performance.

[0094] For example, before the step of obtaining the audio recording of the speaker under test while playing a preset test audio, the following steps are included:

[0095] The free-field frequency response of the active loudspeaker is measured using a standard microphone, and the measurement result is used as the reference frequency response.

[0096] At the same measurement position of the standard microphone, the microphone under test is repeatedly measured, and the frequency response deviation between the frequency response of the microphone under test and the reference frequency response is recorded. The frequency response deviation is used as the compensation value of the microphone under test.

[0097] Using the standard microphone or the microphone under test compensated based on the compensation value, the audio recording of the speaker under test is collected when playing a preset test audio.

[0098] Before acquiring the audio recording of the speaker under test while playing a preset test audio, it is necessary to calibrate the recording equipment, such as the microphone. A standard microphone must be used when recording audio, therefore the microphone must be calibrated before testing.

[0099] Calibration must be performed in a soundproof environment. Prepare a standard microphone and several microphones to be calibrated. First, fix the positions of the active loudspeaker and the standard microphone in the venue, and measure the free-field frequency response of the active loudspeaker. The measurement result is used as the reference frequency response. Replace the standard microphone with the microphone to be tested in the same position, repeat the measurement, and record its frequency response deviation. This deviation will be compensated by software in subsequent tests.

[0100] In addition, to improve the reliability of calibration, the signal-to-noise ratio of spectrum measurement should be greater than 20dB, and the consistency of calibration should be ensured by repeated calibration.

[0101] Step S20: Detect the detection value of a preset detection item for the recorded audio on the time-domain image;

[0102] After converting the audio recording of the speaker under test, which was recorded while playing a preset test audio, into a time-domain image, the detection values ​​of the preset detection items of the audio recording can be detected on the time-domain image. The detection values ​​are then used to determine whether the speaker under test is qualified.

[0103] In one embodiment, reference is made to Figure 3 First, the host computer receives the audio signal via serial port. The audio signal is a digital signal, which is then converted into a time-domain image based on information such as the total number of frames, sampling rate, and amplitude. A window function is used to identify image noise; if there is too much noise, the image is deemed unsuitable (NG); otherwise, subsequent detection steps are performed. Image segmentation is used to extract the main signal region (while simultaneously denoising the image), and then corresponding test algorithms are executed based on different test audio.

[0104] For example, the step of detecting the detection value of a preset detection item of the recorded audio on the time-domain image includes:

[0105] The rise time, peak time, and overshoot of the preset test audio are determined in the time-domain image.

[0106] When the preset test audio is a step signal audio, the adjustment time of the step signal audio is also determined in the time domain image;

[0107] When the preset test audio is a square wave signal audio, the tilt of the square wave signal audio is also determined in the time domain image.

[0108] When detecting preset detection items for recorded audio on a time-domain image, the test audio for non-detailed testing is a step signal, and the test content is to use machine vision algorithms to detect the rise time, peak time, overshoot, and settling time of the recorded step signal. The test audio for detailed testing is a square wave signal, and the test content is to use machine vision algorithms to detect the rise time, peak time, overshoot, and tilt of the recorded square wave signal.

[0109] For example, the step of determining the rise time, peak time, and overshoot of the preset test audio in the time-domain image includes:

[0110] The temporal image is binarized, and edge contour lines are extracted through edge detection.

[0111] By tracing the curve of the edge contour and performing local optimization, the starting point, the maximum value point of the vertical axis coordinate, and the steady-state value point of the edge contour are determined, and a two-dimensional coordinate system is established with the starting point as the origin.

[0112] The peak time of the preset test audio is determined as the abscissa value of the maximum value point of the vertical axis.

[0113] The overshoot of the preset test audio is determined as the ratio of the difference between the vertical coordinates of the maximum vertical coordinate point and the steady-state vertical coordinate point to the vertical coordinate of the steady-state vertical coordinate point.

[0114] The rise time of the preset test audio is determined as the abscissa value when the ordinate of the edge contour first reaches the reference ordinate, wherein the reference ordinate is determined by the ordinate of the steady-state point and the preset reference coefficient.

[0115] Taking a non-detailed test as an example, in one embodiment, refer to Figure 4 First, image preprocessing is performed. The temporal image is binarized, and edge detection (such as the Canny algorithm) is used to extract the signal edge contours. Figure 4 This is a magnified schematic diagram of a step signal after image preprocessing.

[0116] Then, the step signal features are labeled using image processing methods. The foreground and background of the binarized image are distinguished by 0 and 255, so key points can be obtained through curve tracing and local optimization. The starting point of the curve tracing, with absolute coordinates (x0, y0) in the image, is marked as the origin of the relative coordinates (0, 0).

[0117] Next, local optimization is performed, recording the absolute coordinates (x1, y1) of the maximum point (maximum point of the vertical axis) in the y-axis direction of the first segment of the curve. Its relative coordinates (t) are then obtained from x1-x0 and y1-y0. m y m ).

[0118] Curve tracking involves taking the ordinate y2 of any steady-state point in the latter part of the image, and obtaining its relative ordinate y from y2-y0. s .

[0119] In one embodiment, the preset reference coefficient is set to 0.9, and the reference ordinate is 0.9y. s +y0, from (x0, 0.9ys Starting from point (+y0), traverse along the positive x-axis, and record the x-coordinate of the first intersection point with the curve as x3. The relative x-coordinate t is obtained from x3 - x0. r .

[0120] Determine the rise time: t r Peak time: t m Overshoot: σ = (y m -y s ) / y s .

[0121] In one embodiment, reference is made to Figure 5 The rise time t of the square wave signal r Peak time t m With overshoot σ=(y m -y s ) / y s The calculation method is similar to that of the step signal, and will not be elaborated here.

[0122] For example, the step of determining the adjustment time of the step signal audio in the time domain image when the preset test audio is a step signal audio includes:

[0123] The fluctuation range of the ordinate of the steady-state value point is determined based on the ordinate of the steady-state value point and the preset fluctuation coefficient.

[0124] Determine the extreme points of the monotonic variation of the edge contour line;

[0125] If the ordinate of the extreme point is within the ordinate fluctuation range, then before the extreme point, determine the target point where the edge contour line intersects with the maximum or minimum ordinate value of the ordinate fluctuation range.

[0126] The adjustment time of the step signal audio is determined as the x-coordinate value of the target point.

[0127] In one embodiment, reference is made to Figure 4 Perform curve tracking from point (t) m y m Starting from the positive x-axis, traverse the curve along its relative ordinate y. x Compare y x With y s +Δ and y s -Δ, where the preset fluctuation coefficient is 0.02 and Δ is 0.02y. s y x The value of y will cycle from decreasing to increasing and then decreasing again until it stabilizes. If at some point y... x When the monotonic transformation of y satisfies s-Δ<y x <y s +Δ, then find the previous (y) x =y s -Δ)||(y x =y s The relative x-coordinate of the point (+Δ) is denoted as the settling time t of the step signal audio. s .

[0128] For example, the step of determining the tilt of the square wave signal audio in the time domain image when the preset test audio is a square wave signal audio includes:

[0129] Determine the tilt point, wherein the x-coordinate of the tilt point is the x-coordinate value when the y-coordinate of the edge contour line first reaches the reference y-coordinate, and the y-coordinate of the tilt point is the reference y-coordinate.

[0130] The tilt of the square wave signal audio is determined as the slope between the tilt point and the starting point.

[0131] In one embodiment, reference is made to Figure 5 The square wave signal needs to have its tilt detected additionally. The tilt is determined by the starting point (t0, 0) and the point (t... r 0.9y s The calculation yields: l = 0.9y s / (t r -t0).

[0132] Step S30: If all the detected values ​​meet the preset threshold of the preset detection item corresponding to the detected value, then the audio device under test is determined to be qualified.

[0133] When the audio signal being tested is a step signal, if the rise time, peak time, overshoot, and settling time all meet the corresponding preset thresholds, then the audio device under test is deemed qualified.

[0134] When the test audio is a square wave signal, if the rise time, peak time, overshoot, and tilt all meet the corresponding preset thresholds, then the audio device under test is deemed qualified.

[0135] For example, the method further includes:

[0136] If all the speakers in the current batch are qualified, the speakers corresponding to the current batch will be sent to the next work station.

[0137] If any of the speakers in the current batch are found to be substandard, the number of speakers to be tested will be increased until all speakers in subsequent batches are found to be substandard.

[0138] In one embodiment, if the test result is PASS, the product is deemed qualified and subsequent testing steps are performed; if the test result is NG, the test data is analyzed and uploaded to the R&D system, where retesting or other processing can be carried out (including increasing the sampling rate in the MP stage). Simultaneously, after laboratory testing is completed, the test results are transmitted to the R&D system after simple data analysis and report generation, facilitating targeted improvements to the product from the R&D perspective during new product development and NPI stages.

[0139] In one application scenario of the machine vision-based sound detection method of this application, referring to... Figure 6 The quality inspection of smart speakers in a machine vision-based laboratory is divided into two parts: the laboratory preparation stage and the laboratory testing stage. The laboratory preparation stage ensures a stable testing environment and calibrates a standard microphone, including setting up the test environment and calibrating the standard microphone. The test environment for smart speaker quality inspection must be stable, soundproof, free from audio interference, and the positions of the product under test and the standard microphone must be fixed.

[0140] During the laboratory testing phase, the standard microphone acquires the audio signal and transmits it to the host computer. The acoustic device or module under test is sampled and sent to the laboratory testing environment. Test audio (step signal or square wave signal) is played, and the standard microphone acquires the audio signal and transmits it to the host computer. Audio testing and subsequent processing are then performed.

[0141] Reference Figure 7 The process of quality inspection for smart speakers in the laboratory based on machine vision is detailed below:

[0142] (1) Set up the laboratory testing environment;

[0143] (2) Calibrate the test microphone to obtain a standard microphone;

[0144] (3) Select the product to be tested and play the corresponding audio (step signal or square wave signal) according to the test content;

[0145] (4) Record test audio with a standard microphone and transmit it to the host computer in the laboratory via serial port;

[0146] (5) Convert the audio signal into an image in the host computer so that it can be processed;

[0147] (6) Perform audio quality detection using a machine vision-based audio detection algorithm;

[0148] (7) If there are NG products in step (6), their analysis data is fed back to the R&D system, and then retesting, increasing the number of samples (MP stage), rework or other operations are performed;

[0149] (8) If all the test items in step (6) pass, the audio quality is deemed to be qualified, and subsequent testing steps are performed. Thus, the quality inspection of the smart speaker laboratory based on machine vision is completed.

[0150] This invention provides a low-cost, automated testing method for intelligent speakers in the laboratory stage, replacing manual audio quality testing, improving testing efficiency and accuracy, and ensuring product quality. Furthermore, the system integrates with the R&D system to optimize product design and processes, thereby enhancing the intelligence level of manufacturing and improving product development efficiency. This method features high integration, high reliability, high automation, and a certain degree of intelligence, reducing testing costs, improving testing efficiency, and can be extended to other products with audio playback functions.

[0151] Reference Figure 8 Furthermore, embodiments of this application also provide a machine vision-based sound detection device, which includes:

[0152] The acquisition module M1 is used to acquire the audio recording of the speaker under test when playing a preset test audio, and convert the audio recording into a time-domain image;

[0153] The detection module M2 is used to detect the detection value of a preset detection item of the recorded audio on the time domain image;

[0154] The determination module M3 is used to determine that the audio device under test is qualified if all the detected values ​​meet the preset threshold of the preset detection item corresponding to the detected values.

[0155] For example, the detection module is further configured to:

[0156] The rise time, peak time, and overshoot of the preset test audio are determined in the time-domain image.

[0157] When the preset test audio is a step signal audio, the adjustment time of the step signal audio is also determined in the time domain image;

[0158] When the preset test audio is a square wave signal audio, the tilt of the square wave signal audio is also determined in the time domain image.

[0159] For example, the detection module is further configured to:

[0160] The temporal image is binarized, and edge contour lines are extracted through edge detection.

[0161] By tracing the curve of the edge contour and performing local optimization, the starting point, the maximum value point of the vertical axis coordinate, and the steady-state value point of the edge contour are determined, and a two-dimensional coordinate system is established with the starting point as the origin.

[0162] The peak time of the preset test audio is determined as the abscissa value of the maximum value point of the vertical axis.

[0163] The overshoot of the preset test audio is determined to be the maximum value point of the vertical axis coordinate y. m and the steady-state point y s The difference between the ordinate and the steady-state value point y s The ratio of the ordinates;

[0164] The rise time of the preset test audio is determined as the abscissa value when the ordinate of the edge contour first reaches the reference ordinate, wherein the reference ordinate is determined by the ordinate of the steady-state point and the preset reference coefficient.

[0165] For example, the detection module is further configured to:

[0166] The fluctuation range of the ordinate of the steady-state value point is determined based on the ordinate of the steady-state value point and the preset fluctuation coefficient.

[0167] Determine the extreme points of the monotonic variation of the edge contour line;

[0168] If the ordinate of the extreme point is within the ordinate fluctuation range, then before the extreme point, determine the target point where the edge contour line intersects with the maximum or minimum ordinate value of the ordinate fluctuation range.

[0169] The adjustment time of the step signal audio is determined as the x-coordinate value of the target point.

[0170] For example, the detection module is further configured to:

[0171] Determine the tilt point, wherein the x-coordinate of the tilt point is the x-coordinate value when the y-coordinate of the edge contour line first reaches the reference y-coordinate, and the y-coordinate of the tilt point is the reference y-coordinate.

[0172] The tilt of the square wave signal audio is determined as the slope between the tilt point and the starting point.

[0173] For example, the acquisition module is further configured to:

[0174] Before the step of obtaining the audio recording made by the speaker under test while playing a preset test audio,

[0175] The free-field frequency response of the active loudspeaker is measured using a standard microphone, and the measurement result is used as the reference frequency response.

[0176] At the same measurement position of the standard microphone, the microphone under test is repeatedly measured, and the frequency response deviation between the frequency response of the microphone under test and the reference frequency response is recorded. The frequency response deviation is used as the compensation value of the microphone under test.

[0177] Using the standard microphone or the microphone under test compensated based on the compensation value, the audio recording of the speaker under test is collected when playing a preset test audio.

[0178] For example, the determining module is further configured to:

[0179] If all the speakers in the current batch are qualified, the speakers corresponding to the current batch will be sent to the next work station.

[0180] If any of the speakers in the current batch are found to be substandard, the number of speakers to be tested will be increased until all speakers in subsequent batches are found to be substandard.

[0181] The machine vision-based audio detection device provided in this application employs the machine vision-based audio detection method described in the above embodiments, aiming to accurately detect the acoustic quality of audio. Compared with conventional technologies, the beneficial effects of the machine vision-based audio detection device provided in this application are the same as those of the machine vision-based audio detection method described in the above embodiments, and other technical features in the machine vision-based audio detection device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0182] Furthermore, this application embodiment also provides a machine vision-based audio detection device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the machine vision-based audio detection method as described above.

[0183] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the machine vision-based sound detection method described above.

[0184] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0185] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to conventional technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0186] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A machine vision-based sound detection method, characterized in that, The method includes: Acquire the audio recording of the speaker under test while playing a preset test audio, and convert the audio recording into a time-domain image; Detect the detection value of a preset detection item for the recorded audio on the time-domain image; The step of detecting the detection value of the preset detection item of the recorded audio on the time domain image includes: The rise time, peak time, and overshoot of the preset test audio are determined in the time-domain image. When the preset test audio is a step signal audio, the adjustment time of the step signal audio is also determined in the time domain image; When the preset test audio is a square wave signal audio, the tilt of the square wave signal audio is also determined in the time domain image; The step of determining the rise time, peak time, and overshoot of the preset test audio in the time-domain image includes: The temporal image is binarized, and edge contour lines are extracted through edge detection. By tracing the curve of the edge contour and performing local optimization, the starting point, the maximum value point of the vertical axis coordinate, and the steady-state value point of the edge contour are determined, and a two-dimensional coordinate system is established with the starting point as the origin. The peak time of the preset test audio is determined as the abscissa value of the maximum value point of the vertical axis. The overshoot of the preset test audio is determined as the ratio of the difference between the vertical coordinates of the maximum vertical coordinate point and the steady-state vertical coordinate point to the vertical coordinate of the steady-state vertical coordinate point. The rise time of the preset test audio is determined as the abscissa value when the ordinate of the edge contour line first reaches the reference ordinate, wherein the reference ordinate is determined by the ordinate of the steady-state point and the preset reference coefficient. If all the detected values ​​meet the preset threshold of the preset detection item corresponding to the detected value, then the audio device under test is determined to be qualified.

2. The machine vision-based sound detection method as described in claim 1, characterized in that, The step of determining the adjustment time of the step signal audio in the time domain image when the preset test audio is a step signal audio includes: The fluctuation range of the ordinate of the steady-state value point is determined based on the ordinate of the steady-state value point and the preset fluctuation coefficient. Determine the extreme points of the monotonic variation of the edge contour line; If the ordinate of the extreme point is within the ordinate fluctuation range, then before the extreme point, determine the target point where the edge contour line intersects with the maximum or minimum ordinate value of the ordinate fluctuation range. The adjustment time of the step signal audio is determined as the x-coordinate value of the target point.

3. The machine vision-based sound detection method as described in claim 1, characterized in that, The step of determining the tilt of the square wave signal audio in the time domain image when the preset test audio is a square wave signal audio includes: Determine the tilt point, wherein the x-coordinate of the tilt point is the x-coordinate value when the y-coordinate of the edge contour line first reaches the reference y-coordinate, and the y-coordinate of the tilt point is the reference y-coordinate. The tilt of the square wave signal audio is determined as the slope between the tilt point and the starting point.

4. The machine vision-based sound detection method as described in claim 1, characterized in that, Before the step of acquiring the audio recording of the speaker under test while playing a preset test audio, the following steps are included: The free-field frequency response of the active loudspeaker is measured using a standard microphone, and the measurement result is used as the reference frequency response. At the same measurement position of the standard microphone, the microphone under test is repeatedly measured, and the frequency response deviation between the frequency response of the microphone under test and the reference frequency response is recorded. The frequency response deviation is used as the compensation value of the microphone under test. Using the standard microphone or the microphone under test compensated based on the compensation value, the audio recording of the speaker under test is collected when playing a preset test audio.

5. The machine vision-based sound detection method as described in claim 1, characterized in that, The method further includes: If all the speakers in the current batch are qualified, the speakers corresponding to the current batch will be sent to the next work station. If any of the speakers in the current batch are found to be substandard, the number of speakers to be tested will be increased until all speakers in subsequent batches are found to be substandard.

6. A machine vision-based sound detection device, characterized in that, The machine vision-based sound detection device includes: The acquisition module is used to acquire the audio recordings made by the speaker under test when playing a preset test audio, and convert the audio recordings into a time-domain image; A detection module is used to detect the detection value of a preset detection item of the recorded audio on the time-domain image; The detection module is also used to determine the rise time, peak time, and overshoot of the preset test audio in the time domain image; When the preset test audio is a step signal audio, the adjustment time of the step signal audio is also determined in the time domain image; When the preset test audio is a square wave signal audio, the tilt of the square wave signal audio is also determined in the time domain image; The detection module is also used to binarize the temporal image and extract edge contour lines through edge detection; By tracing the curve of the edge contour and performing local optimization, the starting point, the maximum value point of the vertical axis coordinate, and the steady-state value point of the edge contour are determined, and a two-dimensional coordinate system is established with the starting point as the origin. The peak time of the preset test audio is determined as the abscissa value of the maximum value point of the vertical axis. The overshoot of the preset test audio is determined as the ratio of the difference between the vertical coordinates of the maximum vertical coordinate point and the steady-state vertical coordinate point to the vertical coordinate of the steady-state vertical coordinate point. The rise time of the preset test audio is determined as the abscissa value when the ordinate of the edge contour line first reaches the reference ordinate, wherein the reference ordinate is determined by the ordinate of the steady-state point and the preset reference coefficient. The determination module is used to determine that the audio device under test is qualified if all the detected values ​​meet the preset threshold of the preset detection item corresponding to the detected values.

7. A machine vision-based sound detection device, characterized in that, The machine vision-based audio detection device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the machine vision-based audio detection method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the machine vision-based sound detection method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Loudspeaker aging detection method, device and system

    CN116866808A