Conference room audio system inspection method, medium and program product

By utilizing speakers and microphones for automated inspection in the conference room audio system, the inefficiency of existing technologies has been resolved, enabling efficient and accurate audio system monitoring and report generation, thus ensuring audio quality and system stability.

CN119629542BActive Publication Date: 2025-12-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411379781.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-12-30
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

In existing technologies, the inspection of conference room audio systems relies on manual on-site inspections, which is inefficient, time-consuming, and labor-intensive. It cannot achieve real-time status perception and rapid response to problems, making it difficult to meet the inspection needs of a large number of conference rooms.

Method used

By controlling the speakers in the target audio device to play test audio and having the corresponding microphone collect the audio, performance indicators are tested, and an inspection report is generated, thus achieving automatic inspection using existing audio equipment.

Benefits of technology

It improves the efficiency and accuracy of conference room audio system inspections, ensures audio quality and system stability, and saves hardware and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119629542B_ABST
    Figure CN119629542B_ABST
Patent Text Reader

Abstract

A method, medium and program product for inspecting a conference room audio system. The method comprises: determining at least one target audio device to be inspected in the conference room audio system; detecting performance indicators of the at least one target audio device by playing corresponding test audio through a loudspeaker in the at least one target audio device and collecting audio played by the loudspeaker through a corresponding microphone in the at least one target audio device; and generating an inspection report of the conference room audio system based on the performance indicator detection results of the at least one target audio device. In this way, the original audio devices in the conference room audio system can be efficiently utilized to automatically inspect the conference room audio system, the running health status of the conference room audio system can be monitored in real time, the inspection report can be quickly generated, and the inspection efficiency and detection accuracy of the conference room audio system are significantly improved, thereby ensuring the conference audio quality and system stability and saving hardware and manpower costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio equipment testing technology, specifically to a method, medium, and procedure product for inspecting a conference room audio system. Background Technology

[0002] With the increasing popularity of remote and virtual conferencing, conference rooms play a crucial role in modern communication, placing high demands on audio systems to ensure clear transmission of sound from both remote and local sources. Conference room audio systems typically include speakers, microphones, and all-in-one conference equipment. These devices often suffer from audio problems such as echo, reflections, environmental noise interference, audio distortion, signal interruptions, stuttering, and unclear speech. These issues can severely impact the audio quality and communication effectiveness of a meeting.

[0003] Currently, the audio systems in conference rooms typically rely on professionals who periodically conduct manual on-site inspections using additional equipment (such as sound level meters) to record the equipment's operating status and performance indicators. This inspection method is inefficient, cumbersome, time-consuming, and labor-intensive, with high hardware and manpower costs. It also fails to achieve real-time status awareness and rapid problem response, making it difficult to meet the inspection needs of a large number of conference rooms. Summary of the Invention

[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the subsequent detailed description section. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Firstly, this disclosure provides a method for inspecting a conference room audio system, including:

[0006] Identify at least one target audio device in the conference room audio system to be inspected;

[0007] The performance indicators of the at least one target audio device are tested by controlling the speaker in the at least one target audio device to play the corresponding test audio and the corresponding microphone in the at least one target audio device to collect the audio played by the speaker.

[0008] Based on the performance index test results of the at least one target audio device, an inspection report of the conference room audio system is generated.

[0009] Secondly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the inspection method for the conference room audio system provided in the first aspect of this disclosure.

[0010] Thirdly, this disclosure provides an electronic device, including:

[0011] A storage device on which computer programs are stored;

[0012] A processing device is configured to execute the computer program in the storage device to implement the steps of the inspection method for the conference room audio system provided in the first aspect of this disclosure.

[0013] Fourthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the inspection method for the conference room audio system provided in the first aspect of this disclosure.

[0014] In the above technical solution, when inspecting a conference room audio system, firstly, at least one target audio device to be inspected is identified. Then, by controlling the speakers in at least one target audio device to play corresponding test audio and the corresponding microphones in at least one target audio device to capture the audio played by the speakers, the performance indicators of at least one target audio device are tested. Finally, based on the performance indicator test results of at least one target audio device, an inspection report for the conference room audio system is generated. This allows for efficient use of the existing audio equipment in the conference room audio system, enabling automated inspection of the conference room audio system. It facilitates real-time monitoring of the operational health status of the conference room audio system, quickly generates inspection reports, significantly improves the inspection efficiency and accuracy of the conference room audio system, thereby ensuring conference audio quality and system stability, and saving hardware and labor costs. Attached Figure Description

[0015] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:

[0016] Figure 1A This is an architectural diagram of a conference room audio system according to an exemplary embodiment.

[0017] Figure 1B This is an architectural diagram of a conference room audio system according to another exemplary embodiment.

[0018] Figure 1C This is an architectural diagram of a conference room audio system according to yet another exemplary embodiment.

[0019] Figure 2 This is a flowchart illustrating an inspection method for a conference room audio system according to an exemplary embodiment.

[0020] Figure 3 This is a flowchart illustrating the inspection process of a conference room audio system according to an exemplary embodiment.

[0021] Figure 4 This is a schematic diagram illustrating the flow of inspection signals in a conference room audio system according to an exemplary embodiment.

[0022] Figure 5 This is a block diagram illustrating an inspection device for a conference room audio system according to an exemplary embodiment.

[0023] Figure 6 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment. Detailed Implementation

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0030] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0031] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0032] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0033] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0034] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0035] Conference room audio systems typically include: the physical acoustic environment, an audio control center (e.g., conference terminals, remote servers (e.g., main control console)), and audio equipment such as speakers, microphones, and all-in-one conference units. The physical acoustic environment includes sound-absorbing materials, the reflection characteristics of walls and ceilings, reverberation time, and ambient noise, all of which directly affect audio quality. The audio control center is used for overall control of the sound signal. Speakers are used for audio playback and amplification, while microphones are used for sound signal pickup. All-in-one conference units integrate hardware devices (such as microphones and speakers) and audio processing algorithms (such as Automatic Gain Control (AGC), Acoustic Echo Cancellation (AEC), and Automatic Noise Suppression (ANS)) to provide a comprehensive solution.

[0036] Based on actual application needs, conference room audio systems can be divided into three types: local deployment, remote deployment, and local + remote deployment, to meet different privacy and data processing requirements.

[0037] In one implementation, the conference room audio system can adopt... Figure 1A The remote deployment method is shown in the image. Figure 1A As shown, a conference room audio system may include audio devices and a remote server connected via a wide area network (WAN). The audio devices are located in the local physical conference room, and the remote server is a remote server, such as a cloud server. The conference room audio system includes audio devices such as microphone M1, speaker Y1, a conference all-in-one unit, microphone M2, microphone M3, and speaker Y2. The remote server is connected to these audio devices via a WAN. In this case, as... Figure 3 and Figure 4 As shown, the remote server directly controls the audio equipment via the network, performing data acquisition and control. All data analysis and computation (including the operation of artificial intelligence models) involved in the inspection of the conference room audio system—that is, artificial intelligence (AI) analysis—are performed on the remote server. This remote deployment method is suitable for scenarios with lower privacy requirements, higher computational performance requirements, and a need for a simplified system architecture. The AI ​​model includes various algorithm modules or sub-models for analyzing various performance indicators of audio signals, such as the deep noise reduction model, short-time objective intelligibility index algorithm, and recurrent neural networks mentioned below.

[0038] In another implementation, the conference room audio system can adopt... Figure 1B The local + remote deployment method is shown in the image. Figure 1B As shown, the conference room audio system includes audio equipment, conference terminals, and a remote server. The conference terminals are connected to the audio equipment via a local area network (LAN) (e.g., via Bluetooth, Ethernet cable, WiFi, etc.), and the conference terminals are connected to the remote server via a wide area network (WAN). The audio equipment and conference terminals are located in the local physical conference room, while the remote server is a remote server. Figure 3 and Figure 4 As shown, the remote server controls the conference terminal via a wide area network (WAN), and the conference terminal controls the audio equipment via a local area network (LAN). Figure 1BAs shown, the conference room audio system includes microphone M1, speaker Y1, a conference all-in-one machine, microphone M2, microphone M3, speaker Y2, and other audio devices. The conference terminal is connected to these audio devices via a local area network. Based on privacy and computing performance requirements, the data analysis and computation involved in the conference room audio system inspection (including the operation of artificial intelligence models) can be performed on a remote server or on the conference terminal. However, the performance indicator test results of the audio devices need to be transmitted back to the remote server (e.g., ...). Figure 4 (As shown).

[0039] In another implementation, the conference room audio system can adopt... Figure 1C The local deployment method is shown in the image. Figure 1C As shown, the conference room audio system includes audio equipment and conference terminals connected via a local area network, located in the local physical conference room. Figure 1C As shown, the conference room audio system includes audio devices such as microphone M1, speaker Y1, conference all-in-one machine, microphone M2, microphone M3, and speaker Y2. The conference terminal is connected to these audio devices via a local area network. All data analysis and calculations (including the operation of artificial intelligence models) involved in the inspection of the conference room audio system are performed on the conference terminal. This local deployment method is suitable for scenarios with high privacy requirements, as the data does not leave the local machine.

[0040] The remote server connects to the conference terminal or directly to the audio equipment via a wide area network (WAN). Its main functions include: running artificial intelligence models to process data collected from the conference terminal and audio equipment; providing remote control, management, maintenance, and upgrade functions for the conference terminal and audio equipment; storing all inspection data and analysis results (i.e., performance indicator test results of the audio equipment) for easy viewing and analysis later; and in the case of remote deployment, the remote server can bypass the conference terminal and directly connect to the audio equipment via the WAN for data collection and control, and send test audio to the audio equipment for playback.

[0041] The conference terminal connects to various audio devices (microphones, speakers, conference all-in-one machines) in the conference room via a local area network. Its main functions include: controlling the working status of audio devices; collecting data from audio devices and sending test audio to the audio devices for playback; and, in the case of local processing, the conference terminal can run an artificial intelligence model to analyze audio signals and, depending on the deployment method, decide whether to analyze and send the data back to a remote server.

[0042] Figure 2 This is a flowchart illustrating an inspection method for a conference room audio system according to an exemplary embodiment. Figure 2 As shown, the inspection method may include the following S101 to S103.

[0043] In S101, at least one target audio device in the conference room audio system to be inspected is identified.

[0044] In this disclosure, at least one target audio device to be inspected may include all or some of the audio devices in the conference room audio system, and the user may specify the audio devices to be inspected.

[0045] In S102, the performance indicators of at least one target audio device are tested by controlling the speaker in at least one target audio device to play the corresponding test audio and the corresponding microphone in at least one target audio device to collect the audio played by the speaker.

[0046] In S103, an inspection report of the conference room audio system is generated based on the performance index test results of at least one target audio device.

[0047] In this disclosure, the above-mentioned inspection method can be executed by a conference terminal, a remote server, or by a combination of both.

[0048] In one implementation, the conference room audio system is deployed locally. In this case, the aforementioned inspection method is executed by the conference terminal, such as... Figure 3 As shown.

[0049] In another implementation, the aforementioned conference room audio system is remotely deployed. In this case, the aforementioned inspection method is executed by a remote server, such as... Figure 3 As shown.

[0050] In another implementation, the aforementioned conference room audio system adopts a local + remote deployment method. In this case, the aforementioned inspection method can be executed collaboratively by the conference terminal and the remote server.

[0051] When the conference terminal and remote server collaboratively execute the above inspection method, the data analysis and calculations involved in the inspection of the conference room audio system can be performed on the conference terminal. After obtaining the performance index detection results of at least one target audio device, the conference terminal sends them back to the remote server, which then generates an inspection report for the conference room audio system based on the received performance index detection results of at least one target audio device. At this time, S101 may include: the conference terminal determining at least one target audio device to be inspected in the conference room audio system; S102 may include: the conference terminal performing performance index detection on at least one target audio device by controlling the speakers in at least one target audio device to play corresponding test audio and the corresponding microphones in at least one target audio device to collect the audio played by the speakers, and sending the performance index detection results of at least one target audio device to the remote server; S103 may include: the remote server generating an inspection report for the conference room audio system based on the performance index detection results of at least one target audio device sent by the conference terminal. That is, S101 and S102 can be executed by the conference terminal, and S103 is executed by the remote server.

[0052] When the conference terminal and remote server collaboratively execute the above inspection method, the data analysis and calculations involved in the conference room audio system inspection can be performed on the remote server. In this case, the remote server uses the conference terminal as the communication medium to execute the above inspection method, such as... Figure 3 As shown, the remote server controls the conference terminal via the network.

[0053] In the above technical solution, when inspecting a conference room audio system, firstly, at least one target audio device to be inspected is identified. Then, by controlling the speakers in at least one target audio device to play corresponding test audio and the corresponding microphones in at least one target audio device to capture the audio played by the speakers, the performance indicators of at least one target audio device are tested. Finally, based on the performance indicator test results of at least one target audio device, an inspection report for the conference room audio system is generated. This allows for efficient use of the existing audio equipment in the conference room audio system, enabling automated inspection of the conference room audio system. It facilitates real-time monitoring of the operational health status of the conference room audio system, quickly generates inspection reports, significantly improves the inspection efficiency and accuracy of the conference room audio system, thereby ensuring conference audio quality and system stability, and saving hardware and labor costs.

[0054] The following is a detailed description of the specific implementation method for determining at least one target audio device to be inspected in the conference room audio system in step S101 above. Specifically, it can be implemented through various methods. In one implementation method, the second state of each audio device in the conference room audio system can be obtained first, wherein the second state is used to characterize whether the audio device is online; then, the audio device with the second state being online is directly determined as the target audio device.

[0055] In another implementation, such as Figure 3 As shown, the second status of each audio device in the conference room audio system can be obtained first. If no audio device is offline in the second status (meaning all audio devices in the conference room audio system are online), the online audio device can be identified as the target audio device, and then an automatic inspection program can be started. If an audio device is offline in the second status, the offline device is processed. At this time, selection information can be sent to the user terminal, allowing the user to choose whether to skip the inspection of the offline audio device. After receiving the selection information, the user terminal displays the selection information, allowing the user to choose whether to skip the inspection of the offline audio device. If the user terminal receives a first selection instruction to skip the inspection of the offline audio device, it sends the first selection instruction to the conference terminal or remote server. After receiving the first selection instruction, the conference terminal or remote server identifies the online audio device as the target audio device. If the user terminal receives a second selection instruction indicating the inspection of offline audio devices, it sends the second selection instruction to the conference terminal or remote server. After receiving the second selection instruction, the conference terminal or remote server can control the audio devices in the second offline state to start, and then identify each audio device in the conference room audio system as the target audio device. Alternatively, after receiving the selection information through the user terminal, the user can manually start the offline audio devices and then input the second selection instruction into the user terminal. In this case, after receiving the second selection instruction, the conference terminal or remote server can identify the online audio devices as the target audio devices.

[0056] The following is a detailed description of the specific implementation method for performance index testing of at least one target audio device, specifically the method described in S102 above, which involves controlling the speakers in at least one target audio device to play corresponding test audio and the corresponding microphones in at least one target audio device to collect the audio played by the speakers. Specifically, the at least one target audio device may include m first speakers and n first microphones, where m ≥ 1 and n ≥ 1. In this case, as shown... Figure 3 As shown, a cyclic cross-validation strategy can be used to test the performance indicators of m first speakers and n first microphones, which can be achieved through the following steps (1) to (4):

[0057] Step (1): Combine m first speakers with n first microphones one by one to form m*n speaker-microphone pairs.

[0058] Step (2): For each speaker-microphone pair, control the first speaker in the speaker-microphone pair to play the first test audio, and control the first microphone in the speaker-microphone pair to collect the current audio to obtain the first re-collected audio.

[0059] Step (3): Based on the first audio sample, determine the first performance indicators of the first microphone and the first speaker in the speaker-microphone pair.

[0060] In this disclosure, the first test audio may include a continuously swept frequency signal from low to high frequency, and the first performance metric may include frequency response (FR) and / or total harmonic distortion (THD). The speaker-microphone pair is a device pair consisting of one first speaker from m first speakers and one first microphone from n first microphones.

[0061] For example, at least one target audio device includes speaker Y1, speaker Y2, microphone M1, microphone M2, and microphone M3, i.e., m=2, n=3. Two first speakers (i.e., speaker Y1 and speaker Y2) and three first microphones (microphone M1, microphone M2, and microphone M3) are combined one-to-one to obtain the following six speaker-microphone pairs: (speaker Y1, microphone M1), (speaker Y1, microphone M2), (speaker Y1, microphone M3), (speaker Y2, microphone M1), (speaker Y2, microphone M2), (speaker Y2, microphone M3). In this case, performance indicators can be tested for each speaker-microphone pair separately to detect the frequency response and / or total harmonic distortion of the microphone and speaker in each pair. Figure 3 The following explanation uses the microphone M1 and speaker Y1 as examples.

[0062] Specifically, based on the first recorded audio, the frequency responses of the first microphone and the first speaker in the speaker-microphone pair can be determined in the following way: First, the first recorded audio is converted from the time domain to the frequency domain using a short-time Fourier transform; then, the amplitude of each frequency component in the spectrum of the first recorded audio after the transformation is analyzed, and a frequency response curve is plotted based on the amplitude; finally, this frequency response curve is used as both the frequency response of the first microphone and the frequency response of the first speaker in the speaker-microphone pair.

[0063] Alternatively, the total harmonic distortion of the first microphone and the first speaker in the speaker-microphone pair can be determined based on the first recorded audio by the following method: First, the harmonic components and fundamental components of the first recorded audio are calculated; then, the total harmonic distortion rate is calculated by comparing the amplitudes of the harmonic components with those of the fundamental components, and this is used as the total harmonic distortion of the first microphone and the first speaker in the speaker-microphone pair.

[0064] Step (4): Summarize all the first performance indicators as the performance indicator test results of at least one target audio device.

[0065] In this disclosure, after detecting the first performance indicators of the first speaker and the first microphone in each speaker-microphone pair, the indicators can be summarized on a device-by-device basis. Specifically, for each of the m first speakers, if the first speaker includes multiple detected first performance indicators, the indicator with the best value among the multiple detected first performance indicators is determined as the final first performance indicator of the first speaker; if the first speaker includes a first performance indicator obtained in one detection, the first performance indicator obtained in that detection is determined as the final first performance indicator of the first speaker. Similarly, for each of the n first microphones, if the first microphone includes multiple detected first performance indicators, the indicator with the best value among the multiple detected first performance indicators is determined as the final first performance indicator of the first microphone; if the first microphone includes a first performance indicator obtained in one detection, the first performance indicator obtained in that detection is determined as the final first performance indicator of the first microphone. Finally, the final first performance indicator of each first speaker and the final first performance indicator of each first microphone are used as the performance indicator detection results of at least one target audio device.

[0066] To more comprehensively evaluate the performance of at least one target audio device, in addition to detecting the frequency response and / or total harmonic distortion of the first microphone and the first speaker through the first test audio, the speech intelligibility (STOI) and / or stuttering of the first speaker can also be detected through the second test audio. The second test audio can include speech segments covering low to high frequencies, with speech rate and pronunciation clarity conforming to preset standards. These preset standards can be, for example, a series of standards for voice communication quality assessment developed by the Telecommunication Standardization Sector of the International Telecommunication Union (ITU), such as ITU-T P.501. Specifically, in addition to steps (1) to (4) above, step S102 can also include steps (5) to (7).

[0067] Step (5): For each speaker-microphone pair, control the first speaker in the speaker-microphone pair to play the second test audio, and control the first microphone in the speaker-microphone pair to collect the current audio to obtain the second audio sample.

[0068] Step (6): Based on the second audio sample, determine the second performance indicators of the first microphone and the first speaker in the speaker-microphone pair.

[0069] In this disclosure, the second performance metric includes Short-Time Objective Intelligence (STOI) and / or stuttering.

[0070] Specifically, after acquiring the second re-acquisition audio, a short-time objective intelligibility index algorithm can be used to calculate the similarity between the second test audio (which belongs to a clean speech signal) and the second re-acquisition audio (which belongs to a lossy speech signal) to measure the speech intelligibility, which is then used as the speech intelligibility of the first microphone and the first speaker in the speaker-microphone pair.

[0071] In addition, after obtaining the second re-encapsulated audio, a recurrent neural network (RNN) can be used to capture the time dependence and changing trend in the second re-encapsulated audio in order to accurately identify the stuttering phenomenon in the second re-encapsulated audio.

[0072] Step (7): Summarize all the second performance indicators as the performance indicator test results of at least one target audio device.

[0073] In this disclosure, after detecting the second performance indicators of the first microphone and the first speaker in each speaker-microphone pair, the indicators can be summarized on a device-by-device basis. Specifically, for each of the m first speakers, if the first speaker includes multiple detected second performance indicators, the indicator with the best value among the multiple detected second performance indicators is determined as the final second performance indicator of the first speaker; if the first speaker includes a second performance indicator detected only once, the second performance indicator detected only once is determined as the final second performance indicator of the first speaker. Similarly, for each of the n first microphones, if the first microphone includes multiple detected second performance indicators, the indicator with the best value among the multiple detected second performance indicators is determined as the final second performance indicator of the first microphone; if the first microphone includes a second performance indicator detected only once, the second performance indicator detected only once is determined as the final second performance indicator of the first microphone. Finally, the final second performance indicators of each first speaker and each first microphone are used as the performance indicator detection results of at least one target audio device.

[0074] In addition to m first speakers and n first microphones, the aforementioned target audio device may also include p multi-function conferencing units, where p ≥ 1. In this case, besides performance testing of the m first speakers and n first microphones, performance testing of the p multi-function conferencing units is also required. Figure 3 As shown, when testing the conference all-in-one machine, the first state of m first speakers and n first microphones is first determined. The first state is used to characterize whether the corresponding device has malfunctioned. Then, different testing strategies are adopted for the conference all-in-one machine according to different first states. Specifically, the above S102 may also include the following steps (8) to (10).

[0075] Step (8): Determine the first state of the m first speakers and n first microphones based on the performance index detection results of the m first speakers and n first microphones.

[0076] In this disclosure, after obtaining the performance index detection results of m first speakers and n first microphones, for each of the m first speakers, each index (e.g., FR, THD, STOI, stuttering) in the performance index detection result (i.e., the final first performance index and the final second performance index) of the first speaker can be compared with its corresponding threshold to determine whether the corresponding index is normal; if all the indexes in the performance index detection result of the first speaker are normal, it is determined that the first speaker has not failed; if there are abnormal indexes in the performance index detection result of the first speaker, it is determined that the first speaker has failed.

[0077] Meanwhile, for each of the n first microphones, each indicator (e.g., FR, THD) in the performance indicator detection result of the first speaker (i.e., the final first performance indicator) is compared with its corresponding threshold to determine whether the corresponding indicator is normal; if all indicators in the performance indicator detection result of the first microphone are normal, it is determined that the first microphone is not faulty; if there are abnormal indicators in the performance indicator detection result of the first microphone, it is determined that the first microphone is faulty.

[0078] Step (9): Determine the target detection strategy for the conference all-in-one machine based on the first state.

[0079] Specifically, such as Figure 3 As shown, if there is a first speaker that is not faulty and a first microphone that is not faulty, the target detection strategy is determined to be to perform performance index detection on the conference all-in-one machine based on the target speaker and the target microphone; if there is a first speaker that is not faulty and each first microphone is faulty, the target detection strategy is determined to be to perform performance index detection on the conference all-in-one machine based on the target speaker; if each first speaker is faulty and a first microphone is not faulty, the target detection strategy is determined to be to perform performance index detection on the conference all-in-one machine based on the target microphone; if both the first speaker and the first microphone are faulty, the target detection strategy is determined to be to perform performance index detection on the conference all-in-one machine based on the second speaker and the second microphone of the conference all-in-one machine.

[0080] Wherein, the target speaker is any first speaker that has not malfunctioned, and the target microphone is any first microphone that has not malfunctioned. Figure 3 The example uses microphone M1 as the target microphone and speaker Y1 as the target speaker.

[0081] Step (10): Perform performance index testing on each conference all-in-one machine according to the target detection strategy.

[0082] In one implementation, when the target detection strategy is to detect the performance indicators of the conference all-in-one machine based on the target speaker and the target microphone, step (10) may include at least one of the following steps (a1), (a2) and (a3).

[0083] Step (a1): For each conference all-in-one machine, control the target speaker to play the first test audio, and control the second microphone of the conference all-in-one machine to collect the current audio to obtain the third re-collected audio; based on the third re-collected audio, determine the first performance index of the second microphone of the conference all-in-one machine.

[0084] In this disclosure, a similar method can be used to determine the first performance index of the first microphone and the first speaker in the speaker-microphone pair based on the first audio re-collection in step (3) above, and to determine the first performance index of the second microphone of the conference all-in-one machine based on the third audio re-collection. This disclosure will not elaborate further.

[0085] Step (a2): For each conference all-in-one machine, control the target speaker to play the third test audio, and control the second microphone of the conference all-in-one machine to collect the current audio to obtain the fourth re-collected audio; based on the fourth re-collected audio, determine the third performance index of the second microphone of the conference all-in-one machine and the automatic gain control index of the conference all-in-one machine.

[0086] In this disclosure, the third test audio includes speech segments that cover components from low to high frequencies, whose speech rate and pronunciation clarity meet preset standards, and which contain speech segments with volume varying from low to high.

[0087] The third performance metric may include speech intelligibility and / or stuttering. The third performance metric is a hardware performance metric of the conference all-in-one machine, while automatic gain control is a performance metric of the audio processing algorithm of the conference all-in-one machine.

[0088] In one possible implementation, the fourth re-acquired audio can be denoised based on a Deep Noise Suppression (DeepNS) model using artificial intelligence. Then, the difference between the total root mean square value of the denoised fourth re-acquired audio and the total root mean square value of the third test audio (which is a clean speech signal) is calculated to obtain the gain adjustment value and adjustment time, which are used as the automatic gain control index results of the conference all-in-one machine.

[0089] In this process, a method similar to that used in step (5) above to determine the second performance indicators of the first microphone and the first speaker in the speaker-microphone pair based on the second audio re-collection can be adopted to determine the speech intelligibility and stuttering indicators of the second microphone of the conference all-in-one machine based on the fourth audio re-collection. This disclosure will not elaborate further.

[0090] Step (a3): For each conference all-in-one machine, control the target speaker to play the fourth test audio, and control the second microphone of the conference all-in-one machine to collect the current audio to obtain the fifth re-collected audio; based on the fifth re-collected audio, determine the noise suppression index of the conference all-in-one machine.

[0091] In this disclosure, the fourth test audio may include speech segments covering low to high frequencies, with speech rate and pronunciation clarity meeting preset standards, and containing different noise types and signal-to-noise ratios. Noise suppression metrics are also performance indicators of the audio processing algorithm of the conference all-in-one machine.

[0092] In one possible implementation, after acquiring the fifth audio recording, the noise component in the fifth audio recording can be extracted based on the DeepNS model; then, the difference between the total root mean square value of the noise component and the total root mean square value of the fourth test audio (which belongs to the original noise signal) is calculated to obtain the overall noise residual level, which is used as the noise suppression index result of the conference all-in-one machine.

[0093] When the target detection strategy is to perform performance index testing on the conference all-in-one machine based on the target speaker and the target microphone, in addition to detecting the FR, THD, speech intelligibility, stuttering index of the second microphone of the conference all-in-one machine, as well as the automatic gain control and noise suppression index of the conference all-in-one machine, the echo cancellation index of the conference all-in-one machine can also be detected. Among them, the echo cancellation index is also a performance index of the audio processing algorithm of the conference all-in-one machine. Specifically, when the target detection strategy is to perform performance index testing on the conference all-in-one machine based on the target speaker and the target microphone, the above step (10) includes at least one of steps (a1), (a2) and (a3), and may also include the following step (a4).

[0094] Step (a4): For each conference all-in-one machine, control the target speaker and the second speaker of the conference all-in-one machine to play the second test audio simultaneously, and control the second microphone of the conference all-in-one machine to collect the current audio to obtain the sixth echo audio; determine the echo cancellation index of the conference all-in-one machine based on the sixth echo audio.

[0095] In one possible implementation, the sixth audio recording can be denoised based on the DeepNS model. The difference between the total root mean square value of the denoised sixth audio recording and the total root mean square value of the second test audio (which is a clean speech signal) is calculated to obtain the overall level of speech clipping and echo residue, which is used as the echo cancellation index result of the conference all-in-one machine.

[0096] When the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target speaker and the target microphone, in addition to performing performance index detection on the second microphone and the audio processing algorithm of the conference all-in-one machine, the performance index detection on the second speaker of the conference all-in-one machine can also be performed. Specifically, when the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target speaker and the target microphone, the above step (10) includes at least one of steps (a1), (a2) and (a3), and may also include at least one of the following steps (a5) and (a6).

[0097] Step (a5): For each conference all-in-one machine, control the second speaker of the conference all-in-one machine to play the first test audio, and control the target microphone to collect the current audio to obtain the seventh audio recording; based on the seventh audio recording, determine the first performance index of the second speaker of the conference all-in-one machine.

[0098] In this disclosure, a similar method can be used to determine the first performance index of the first microphone and the first speaker in the speaker-microphone pair based on the first audio re-collection in step (3) above, and the first performance index of the second speaker of the conference all-in-one machine can be determined based on the seventh audio re-collection. This disclosure will not elaborate further.

[0099] Step (a6): For each conference all-in-one machine, control the second speaker of the conference all-in-one machine to play the second test audio, and control the target microphone to collect the current audio to obtain the eighth audio recording; based on the eighth audio recording, determine the second performance index of the second speaker of the conference all-in-one machine.

[0100] In this disclosure, a similar method as that used in step (5) above to determine the second performance index of the first microphone and the first speaker in the speaker-microphone pair based on the second re-acquired audio can be adopted to determine the second performance index of the second speaker of the conference all-in-one machine based on the eighth re-acquired audio. This disclosure will not elaborate further.

[0101] In another implementation, such as Figure 3 As shown, when the target detection strategy is to detect the performance indicators of the conference all-in-one machine based on the target speaker, the above step (10) may include at least one of the above steps (a1), (a2) and (a3).

[0102] like Figure 3 As shown, when the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target speaker, in addition to detecting the FR, THD, speech intelligibility, stuttering index of the second microphone of the conference all-in-one machine, as well as the automatic gain control and noise suppression index of the conference all-in-one machine, the echo cancellation index of the conference all-in-one machine can also be detected. Specifically, when the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target speaker, the above step (10) may include at least one of steps (a1), (a2), and (a3), as well as step (a4).

[0103] In another implementation, when the target detection strategy is to detect the performance indicators of the conference all-in-one machine based on the target microphone, the above step (10) may include at least one of the following steps (b1) and (b2).

[0104] Step (b1): For each conference all-in-one machine, control the second speaker of the conference all-in-one machine to play the first test audio, and control the target microphone to collect the current audio to obtain the seventh audio recording; based on the seventh audio recording, determine the first performance index of the second speaker of the conference all-in-one machine.

[0105] In this disclosure, a similar method can be used to determine the first performance index of the first microphone and the first speaker in the speaker-microphone pair based on the first audio re-collection in step (3) above, and the first performance index of the second speaker of the conference all-in-one machine can be determined based on the seventh audio re-collection. This disclosure will not elaborate further.

[0106] Step (b2): For each conference all-in-one machine, control the second speaker of the conference all-in-one machine to play the second test audio, and control the target microphone to collect the current audio to obtain the eighth audio recording; based on the eighth audio recording, determine the second performance index of the second speaker of the conference all-in-one machine.

[0107] In this disclosure, a similar method as that used in step (5) above to determine the second performance index of the first microphone and the first speaker in the speaker-microphone pair based on the second re-acquired audio can be adopted to determine the second performance index of the second speaker of the conference all-in-one machine based on the eighth re-acquired audio. This disclosure will not elaborate further.

[0108] like Figure 3 As shown, when the target detection strategy is to perform performance index testing on the conference all-in-one machine based on the target microphone, in addition to testing the conference all-in-one machine's FR, THD, speech intelligibility, automatic gain control, stuttering, and noise suppression indicators, the conference all-in-one machine's echo cancellation indicator can also be tested. Specifically, when the target detection strategy is to perform performance index testing on the conference all-in-one machine based on the target microphone, the above step (10) may include at least one of the above steps (b1) and (b2), and may also include the following step (b3).

[0109] Step (b3): ​​For each conference all-in-one machine, control the second speaker of the conference all-in-one machine to play the second test audio, and control the second microphone of the conference all-in-one machine to collect the current audio to obtain the ninth echo audio; based on the ninth echo audio, determine the echo cancellation index of the conference all-in-one machine.

[0110] In this disclosure, a similar method to that used in step (a4) above to determine the echo cancellation index of the conference all-in-one machine based on the sixth audio recording can be adopted to determine the echo cancellation index of the conference all-in-one machine based on the ninth audio recording. This disclosure will not elaborate further.

[0111] In another embodiment, when the target detection strategy is to detect the performance indicators of the conference all-in-one machine based on the second speaker and the second microphone of the conference all-in-one machine, the above step (10) may include the above step (b3).

[0112] The following is a detailed description of the specific implementation method for generating an inspection report for the conference room audio system based on the performance index test results of at least one target audio device in S103 above.

[0113] Specifically, based on the performance indicator test results of at least one target audio device, an overall health status of at least one target audio device can be generated, and / or a first state of at least one target audio device can be determined, resulting in an evaluation result. Then, based on the overall health status and / or the evaluation result, an inspection report of the conference room audio system can be generated.

[0114] In one implementation, an overall health status of at least one target audio device can be generated based on the performance index test results of at least one target audio device. Then, based on the at least one overall health status, an inspection report of the conference room audio system is generated. At this time, the inspection report may include the overall health status of the conference room audio system and the performance index test results of at least one target audio device.

[0115] The inspection report includes the overall health status of the conference room audio system, which can help users intuitively understand the overall health status of the conference room audio system.

[0116] In another implementation, the first state of at least one target audio device can be determined based on the performance index detection results of at least one target audio device to obtain an evaluation result; then, based on the evaluation result, an inspection report of the conference room audio system can be generated. At this time, the inspection report may include the evaluation result (i.e., the first state of at least one target audio device) and the performance index detection results of at least one target audio device.

[0117] The inspection report includes both the initial status of at least one target audio device as to whether it is malfunctioning and the performance indicator test results of at least one target audio device. In this way, users can quickly understand the status of the corresponding audio device through the inspection report.

[0118] In another embodiment, the overall health of at least one target audio device can be generated based on the performance index detection results of at least one target audio device, and the first state of at least one target audio device can be determined based on the performance index detection results of at least one target audio device to obtain an evaluation result; then, an inspection report of the conference room audio system can be generated based on the overall health and the evaluation result. At this time, the inspection report may include the overall health, the evaluation result (i.e., the first state of at least one target audio device), and the performance index detection results of at least one target audio device.

[0119] In addition, the performance indicators of at least one target audio device can be displayed in the form of charts, making it easier for users to understand the performance indicators of the corresponding audio device more intuitively.

[0120] In addition, alarm messages can be generated for audio devices whose first state is faulty, to notify users (e.g., maintenance personnel) to troubleshoot and repair the fault. To facilitate troubleshooting, the target fault cause corresponding to the performance index test results of the faulty audio device can be determined based on the pre-established correspondence between reference performance indicators and fault causes, and this fault cause can be added to the inspection report.

[0121] To allow users to intuitively understand the status of the conference room audio system from different perspectives, a radar chart of the conference room audio system can be generated based on the performance indicators of at least one target audio device, and this radar chart can be displayed in the inspection report. The radar chart can be constructed from dimensions such as sound pickup performance, amplification performance, conference all-in-one machine algorithm performance, and overall performance.

[0122] The following is a detailed description of the specific implementation method for generating the overall health score of at least one target audio device based on the performance index detection results of at least one target audio device. Specifically, for each target audio device, a sub-health score for each performance index of the target audio device can be calculated. Then, the weighted sum of the sub-health scores of each performance index of the target audio device is used to determine the overall health score of the target audio device. Finally, the weighted sum of the health scores of all target audio devices is used to determine the overall health score of at least one target audio device.

[0123] Specifically, for each performance metric of the audio device, a correspondence between the performance metric range and the reference sub-health can be pre-built. In this way, the reference sub-health corresponding to the range to which the performance metric of the audio device belongs can be determined based on the correspondence, and used as the sub-health of the performance metric of the audio device.

[0124] To more comprehensively evaluate the performance of the conference room audio system, in addition to performance testing of the hardware audio equipment, such as... Figure 3 and Figure 4 As shown, the performance of the acoustic environment in the conference room can also be tested. Specifically, the above-mentioned inspection method for the conference room audio system may also include the following steps (C1) to (C3).

[0125] Step (C1): Determine the first state of the m first speakers and n first microphones based on the performance index test results of the m first speakers and n first microphones.

[0126] Step (C2): If there is a first speaker that is not malfunctioning and a first microphone that is not malfunctioning, then use the target speaker and the target microphone to detect the spatial acoustic environment of the physical conference room where at least one target audio device is located.

[0127] Step (C3): If each first speaker and each first microphone fails, abandon the testing of the acoustic environment of the physical conference room (e.g., Figure 4 (As shown).

[0128] In one implementation, step (C2) may include the following steps (C11) to (C14):

[0129] Step (C11): Control the target speaker to play the fifth test audio and control the target microphone to collect the current audio to obtain the tenth audio recording.

[0130] In this disclosure, the fifth test audio includes a pulse signal.

[0131] Step (C12): Noise reduction is performed on the tenth audio recording.

[0132] Step (C13): Generate the energy attenuation curve of the tenth audio sample obtained after noise reduction.

[0133] Step (C14): Determine the reverberation time (Reverb Time 60 dB Decay, RT60) of the spatial acoustic environment based on the energy decay curve.

[0134] After obtaining the energy decay curve of the tenth audio sample after noise reduction, linear regression can be used to fit the linear part of the energy decay curve, and the time required for the initial energy to decay by 60 dB can be calculated to obtain the reverberation time index of the spatial acoustic environment.

[0135] In addition to using pulse signals to detect the reverberation time of spatial acoustic environments, other signals suitable for measuring reverberation time can also be used, and this disclosure does not make specific limitations on this.

[0136] To more comprehensively evaluate the performance of the conference room's acoustic environment, in addition to detecting the reverberation time, the ambient noise level (BNL) can also be detected. Specifically, in addition to steps (C11) to (C14), step (C2) may also include steps (C15) to (C17).

[0137] Step (C15): With the physical conference room in a quiet state, control the target microphone to collect the ambient sound signal of the physical conference room.

[0138] Step (C16): Extract pure background noise information from the ambient sound signal.

[0139] Pure background noise information usually refers to the background noise that still exists in the environment when there is no obvious sound source.

[0140] Step (C17): Determine the ambient noise floor of the spatial acoustic environment based on the pure noise floor information.

[0141] In this disclosure, an A-weighted filter can be used to process the pure noise floor information and calculate the total root mean square value of the weighted signal to obtain the overall level of environmental noise, which can be used as the environmental noise floor index result of the spatial acoustic environment.

[0142] In addition, if each first speaker and / or each first microphone fails, to more comprehensively evaluate the performance of the conference room audio system, additional microphones or speakers can be added to the physical conference room as needed to test the spatial acoustic environment. Specifically, the above-mentioned inspection method for the conference room audio system may also include the following steps:

[0143] If there is a first speaker that is not faulty and every first microphone is faulty, then the target speaker and the third microphone are used to detect the spatial acoustic environment. The third microphone is a device temporarily added to the physical conference room.

[0144] If each of the first speakers fails and there is a first microphone that does not fail, then the third speaker and the target microphone are used to detect the spatial acoustic environment. The third speaker is a device that is temporarily added to the physical conference room.

[0145] If each of the first loudspeakers and each of the first microphones fails, the third loudspeaker and the third microphone are used to detect the spatial acoustic environment.

[0146] The method can be similar to that used in step (C2) above to detect the spatial acoustic environment of the physical conference room where at least one target audio device is located. The spatial acoustic environment can be detected using a third speaker and a target microphone (in which case the third speaker acts as the target speaker), the spatial acoustic environment can be detected using a target speaker and a third microphone (in which case the third microphone acts as the target microphone), and the spatial acoustic environment can be detected using a third speaker and a third microphone (in which case the third microphone acts as the target microphone and the third speaker acts as the target speaker). This disclosure will not elaborate further.

[0147] In addition, when all first speakers and / or all first microphones fail, to more comprehensively evaluate the performance of the conference all-in-one machine, additional microphones or speakers can be added to the physical conference room as needed to test the all-in-one machine. Specifically, if there are unfailed first speakers and all first microphones fail, the performance indicators of the conference all-in-one machine are tested based on the target speaker and the third microphone; if all first speakers fail and there are unfailed first microphones, the performance indicators of the conference all-in-one machine are tested based on the third speaker and the target microphone; if both first speakers and first microphones fail, the performance indicators of the conference all-in-one machine are tested based on the target speaker and the target microphone. This can be done in a manner similar to that used when the target detection strategy is to test the performance indicators of the conference all-in-one machine based on the target speaker and the target microphone, where performance indicators are tested separately for each conference all-in-one machine according to the target detection strategy. This disclosure will not elaborate further on these methods.

[0148] Figure 5 This is a block diagram illustrating an inspection device for a conference room audio system according to an exemplary embodiment. Figure 5 As shown, the inspection device 500 for the conference room audio system includes:

[0149] The first determining module 501 is used to determine at least one target audio device to be inspected in the conference room audio system.

[0150] The first detection module 502 is used to detect the performance indicators of the at least one target audio device by controlling the speaker in the at least one target audio device to play corresponding test audio and the corresponding microphone in the at least one target audio device to collect the audio played by the speaker.

[0151] The generation module 503 is used to generate an inspection report of the conference room audio system based on the performance index test results of the at least one target audio device.

[0152] In the above technical solution, when inspecting a conference room audio system, firstly, at least one target audio device to be inspected is identified. Then, by controlling the speakers in at least one target audio device to play corresponding test audio and the corresponding microphones in at least one target audio device to capture the audio played by the speakers, the performance indicators of at least one target audio device are tested. Finally, based on the performance indicator test results of at least one target audio device, an inspection report for the conference room audio system is generated. This allows for efficient use of the existing audio equipment in the conference room audio system, enabling automated inspection of the conference room audio system. It facilitates real-time monitoring of the operational health status of the conference room audio system, quickly generates inspection reports, significantly improves the inspection efficiency and accuracy of the conference room audio system, thereby ensuring conference audio quality and system stability, and saving hardware and labor costs.

[0153] Optionally, the at least one target audio device includes m first speakers and n first microphones, where m ≥ 1 and n ≥ 1;

[0154] The first detection module 502 includes:

[0155] The combination submodule is used to combine the m first speakers with the n first microphones one by one to form m*n speaker-microphone pairs;

[0156] The first control submodule is configured to control the first speaker in each speaker-microphone pair to play a first test audio, and control the first microphone in each speaker-microphone pair to collect the current audio to obtain a first back-collected audio, wherein the first test audio includes a continuous sweep frequency signal from low frequency to high frequency;

[0157] A first determining submodule is configured to determine a first performance index of the first microphone and the first speaker in the speaker-microphone pair based on the first retrieved audio, wherein the first performance index includes frequency response and / or total harmonic distortion;

[0158] The first aggregation submodule is used to aggregate all the first performance indicators as the performance indicator detection results of the at least one target audio device.

[0159] Optionally, the first detection module 502 further includes:

[0160] The second control submodule is used to control the first speaker in each speaker-microphone pair to play the second test audio, and to control the first microphone in each speaker-microphone pair to collect the current audio to obtain the second re-collected audio, wherein the second test audio includes speech segments that cover low-frequency to high-frequency components and whose speech rate and pronunciation clarity meet preset standards.

[0161] The second determining submodule is used to determine a second performance index of the first microphone and the first speaker in the speaker-microphone pair based on the second re-collected audio, wherein the second performance index includes speech intelligibility and / or stuttering;

[0162] The second aggregation submodule is used to aggregate all the second performance indicators as the performance indicator detection results of the at least one target audio device.

[0163] Optionally, the at least one target audio device further includes p multi-function conferencing machines, where p ≥ 1;

[0164] The first detection module 502 further includes:

[0165] The third determining submodule is used to determine the first state of the m first speakers and the n first microphones based on the performance index detection results of the m first speakers and the n first microphones, wherein the first state is used to characterize whether the corresponding device has malfunctioned;

[0166] The fourth determining submodule is used to determine the target detection strategy of the conference all-in-one machine based on the first state;

[0167] The all-in-one machine detection submodule is used to perform performance index detection on each of the conference all-in-one machines according to the target detection strategy.

[0168] Optionally, the fourth determining submodule includes:

[0169] The fifth determining submodule is used to determine the target detection strategy as performing performance index detection on the conference all-in-one machine based on the target speaker and the target microphone if there is a first speaker that has not failed and a first microphone that has not failed.

[0170] The sixth determining submodule is used to determine the target detection strategy as performing performance index detection on the conference all-in-one machine based on the target speaker if there is a first speaker that has not failed and each of the first microphones has failed.

[0171] The seventh determining submodule is used to determine that if each of the first speakers fails and there is a first microphone that does not fail, the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target microphone.

[0172] The eighth determination submodule is used to determine the target detection strategy as performing performance index detection on the conference all-in-one machine based on the second speaker and the second microphone of the conference all-in-one machine if each of the first speakers and each of the first microphones fails.

[0173] Optionally, the target detection strategy may be to detect the performance indicators of the conference all-in-one machine based on the target speaker and the target microphone, or to detect the performance indicators of the conference all-in-one machine based on the target speaker;

[0174] The integrated machine detection submodule includes at least one of the following:

[0175] The third control submodule is used to control the target speaker to play the first test audio for each of the conference all-in-one machines, and to control the second microphone of the conference all-in-one machine to collect the current audio to obtain the third re-collected audio; the ninth determination submodule is used to determine the first performance index of the second microphone of the conference all-in-one machine based on the third re-collected audio.

[0176] The fourth control submodule is used to control the target speaker to play a third test audio for each of the conference all-in-one machines, and to control the second microphone of the conference all-in-one machine to collect the current audio to obtain a fourth re-collected audio. The third test audio includes components covering low to high frequencies, speech rate and pronunciation clarity that meet preset standards, and speech segments with volume changes from low to high. The tenth determination submodule is used to determine the third performance index of the second microphone of the conference all-in-one machine and the automatic gain control index of the conference all-in-one machine based on the fourth re-collected audio. The third performance index includes speech intelligibility and / or stuttering.

[0177] The fifth control submodule is used to control the target speaker to play the fourth test audio for each of the conference all-in-one machines, and to control the second microphone of the conference all-in-one machine to collect the current audio to obtain the fifth re-collected audio. The fourth test audio includes speech segments that cover low-frequency to high-frequency components, whose speech rate and pronunciation clarity meet the preset standards, and which contain different noise types and signal-to-noise ratios. The eleventh determination submodule is used to determine the noise suppression index of the conference all-in-one machine based on the fifth re-collected audio.

[0178] Optionally, the integrated machine detection submodule further includes:

[0179] The sixth control submodule is used to control the target speaker and the second speaker of the conference all-in-one machine to play the second test audio simultaneously for each of the aforementioned conference all-in-one machines, and to control the second microphone of the conference all-in-one machine to collect the current audio to obtain the sixth re-collected audio, wherein the second test audio includes speech segments that cover components from low frequency to high frequency, and whose speech rate and pronunciation clarity all meet preset standards;

[0180] The twelfth determination submodule is used to determine the echo cancellation index of the conference all-in-one machine based on the sixth echo audio.

[0181] Optionally, the target detection strategy involves detecting the performance indicators of the conference all-in-one machine based on the target speaker and the target microphone;

[0182] The integrated machine detection submodule also includes:

[0183] The seventh control submodule is used to control the second speaker of each conference all-in-one machine to play the first test audio and control the target microphone to collect the current audio to obtain the seventh re-collected audio; the thirteenth determination submodule is used to determine the first performance index of the second speaker of the conference all-in-one machine based on the seventh re-collected audio; and / or

[0184] The eighth control submodule is used to control the second speaker of each conference all-in-one machine to play a second test audio, and to control the target microphone to collect the current audio to obtain an eighth re-collected audio, wherein the second test audio includes speech segments covering low-frequency to high-frequency components and whose speech rate and pronunciation clarity meet preset standards; the fourteenth determination submodule is used to determine a second performance index of the second speaker of the conference all-in-one machine based on the eighth re-collected audio, wherein the second performance index includes speech intelligibility and / or stuttering.

[0185] Optionally, the target detection strategy involves detecting the performance indicators of the conference all-in-one machine based on the target microphone;

[0186] The integrated machine detection submodule includes:

[0187] The ninth control submodule is used to control the second speaker of each of the aforementioned conference all-in-one machines to play the first test audio, and to control the target microphone to collect the current audio to obtain the seventh re-collected audio; the fifteenth determination submodule is used to determine the first performance index of the second speaker of the conference all-in-one machine based on the seventh re-collected audio; and / or

[0188] The tenth control submodule is used to control the second speaker of each conference all-in-one machine to play a second test audio, and to control the target microphone to collect the current audio to obtain an eighth re-collected audio, wherein the second test audio includes speech segments covering low-frequency to high-frequency components and whose speech rate and pronunciation clarity meet preset standards; the sixteenth determination submodule is used to determine a second performance index of the second speaker of the conference all-in-one machine based on the eighth re-collected audio, wherein the second performance index includes speech intelligibility and / or stuttering.

[0189] Optionally, the integrated machine detection submodule further includes:

[0190] The eleventh control submodule is used to control the second speaker of each conference all-in-one machine to play the second test audio, and to control the second microphone of the conference all-in-one machine to collect the current audio to obtain the ninth audio sample. The second test audio includes speech segments that cover low-frequency to high-frequency components, and whose speech rate and pronunciation clarity all meet preset standards.

[0191] The seventeenth determination submodule is used to determine the echo cancellation index of the conference all-in-one machine based on the ninth re-acquired audio.

[0192] Optionally, the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the second speaker and the second microphone of the conference all-in-one machine;

[0193] The step of performing performance index testing on each of the conference all-in-one machines according to the target detection strategy includes:

[0194] The eleventh control submodule is used to control the second speaker of each conference all-in-one machine to play the second test audio, and to control the second microphone of the conference all-in-one machine to collect the current audio to obtain the ninth audio sample. The second test audio includes speech segments that cover low-frequency to high-frequency components, and whose speech rate and pronunciation clarity all meet preset standards.

[0195] The seventeenth determination submodule is used to determine the echo cancellation index of the conference all-in-one machine based on the ninth re-acquired audio.

[0196] Optionally, the first determining module 501 includes:

[0197] The acquisition submodule is used to acquire the second status of each audio device in the conference room audio system, wherein the second status is used to indicate whether the audio device is online;

[0198] The eighteenth determination submodule is used to determine the audio device whose second state is online as the target audio device.

[0199] Optionally, the inspection device 500 further includes:

[0200] The sending module is used to send selection information to the user terminal before the eighteenth determining submodule determines the audio device in the second state as the target audio device. If there is an audio device in the second state as offline, the user can choose whether to skip the inspection of the offline audio device.

[0201] The triggering submodule is configured to, in response to receiving a first selection instruction sent by the user terminal to indicate skipping the inspection of offline audio devices, or if there is no audio device in the second state of being offline, trigger the eighteenth determining submodule to determine the audio device in the second state of being online as the target audio device;

[0202] The twelfth control submodule is configured to respond to receiving a second selection instruction sent by the user terminal to indicate the inspection of offline audio devices, control the audio devices in the second offline state to start, and identify each audio device in the conference room audio system as the target audio device.

[0203] Optionally, the inspection device 500 further includes:

[0204] The second determining module is used to determine the first state of the m first speakers and the n first microphones based on the performance index detection results of the m first speakers and the n first microphones, wherein the first state is used to characterize whether the corresponding device has malfunctioned;

[0205] The second detection module is used to detect the spatial acoustic environment of the physical conference room where the at least one target audio device is located, using the target speaker and the target microphone, if there is a first speaker that is not faulty and a first microphone that is not faulty.

[0206] Optionally, the second detection module includes:

[0207] The thirteenth control submodule is used to control the target speaker to play the fifth test audio and control the target microphone to collect the current audio to obtain the tenth audio sample, wherein the fifth test audio includes a pulse signal;

[0208] A noise reduction submodule is used to reduce noise in the tenth audio sample.

[0209] The first generation submodule is used to generate the energy attenuation curve of the tenth audio sample obtained after noise reduction.

[0210] The nineteenth determination submodule is used to determine the reverberation time of the spatial acoustic environment based on the energy decay curve.

[0211] Optionally, the second detection module further includes:

[0212] The fourteenth control submodule is used to control the target microphone to collect ambient sound signals of the physical conference room when the conference room is in a quiet state.

[0213] An extraction submodule is used to extract pure background noise information from the ambient sound signal;

[0214] The twentieth determination submodule is used to determine the ambient noise of the spatial acoustic environment based on the pure noise floor information.

[0215] Optionally, the inspection device 500 further includes:

[0216] The third detection module is used to detect the spatial acoustic environment using the target speaker and the third microphone if there is a first speaker that is not faulty and each of the first microphones is faulty. The third microphone is a device temporarily added to the physical conference room.

[0217] The fourth detection module is used to detect the spatial acoustic environment using a third speaker and the target microphone if each of the first speakers fails and there is a first microphone that does not fail. The third speaker is a device temporarily added to the physical conference room.

[0218] The fourth detection module is used to detect the spatial acoustic environment using the third speaker and the third microphone if each of the first speakers and each of the first microphones fails.

[0219] Optionally, the conference room audio system includes audio equipment and a conference terminal connected via a local area network, and the inspection device 500 is applied to the conference terminal.

[0220] Optionally, the conference room audio system includes audio equipment and a remote server connected via a wide area network, and the inspection device 500 is applied to the remote server.

[0221] Optionally, the conference room audio system includes audio equipment, a conference terminal, and a remote server, wherein the conference terminal is connected to the audio equipment via a local area network (LAN), and the conference terminal is connected to the remote server via a wide area network (WAN).

[0222] The first determining module 501 is used to determine at least one target audio device to be inspected in the conference room audio system through the conference terminal;

[0223] The first detection module 502 is used to perform performance index detection on the at least one target audio device by controlling the speaker in the at least one target audio device to play corresponding test audio and the corresponding microphone in the at least one target audio device to collect the audio played by the speaker through the conference terminal, and send the performance index detection results of the at least one target audio device to the remote server;

[0224] The generation module 503 is used to generate an inspection report of the conference room audio system through the remote server based on the performance index detection results of the at least one target audio device sent by the conference terminal.

[0225] Optionally, the conference room audio system includes audio equipment, a conference terminal, and a remote server, wherein the conference terminal is connected to the audio equipment via a local area network (LAN), and the conference terminal is connected to the remote server via a wide area network (WAN).

[0226] The remote server includes the inspection device 500, and the remote server controls the operation of the inspection device 500 using the conference terminal as a communication medium.

[0227] Optionally, the generation module 503 includes:

[0228] The second generation submodule is used to generate the overall health of the at least one target audio device based on the performance index detection results of the at least one target audio device; and / or the twenty-first determination submodule is used to determine the first state of the at least one target audio device based on the performance index detection results of the at least one target audio device, and obtain an evaluation result, wherein the first state is used to characterize whether the corresponding device has malfunctioned;

[0229] The third generation submodule is used to generate an inspection report of the conference room audio system based on the overall health status and / or the assessment results.

[0230] In addition, this disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described inspection method for the conference room audio system provided in this disclosure.

[0231] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described inspection method for a conference room audio system provided in this disclosure.

[0232] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device (conference terminal or remote server) 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0233] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0234] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, first microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), first speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0235] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0236] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0237] In some implementations, the conferencing terminal and remote server can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0238] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0239] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine at least one target audio device to be inspected in the conference room audio system; perform performance index testing on the at least one target audio device by controlling the speakers in the at least one target audio device to play corresponding test audio and the corresponding microphones in the at least one target audio device to collect the audio played by the speakers; and generate an inspection report for the conference room audio system based on the performance index testing results of the at least one target audio device.

[0240] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0241] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0242] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, the first determining module can also be described as "a module for determining at least one target audio device to be inspected in a conference room audio system".

[0243] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0244] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0245] According to one or more embodiments of this disclosure, Example 1 provides a method for inspecting a conference room audio system, comprising: identifying at least one target audio device to be inspected in the conference room audio system; performing performance index testing on the at least one target audio device by controlling a speaker in the at least one target audio device to play corresponding test audio and a corresponding microphone in the at least one target audio device to collect the audio played by the speaker; and generating an inspection report of the conference room audio system based on the performance index testing results of the at least one target audio device.

[0246] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the at least one target audio device includes m first speakers and n first microphones, where m ≥ 1 and n ≥ 1; the method of performing performance index detection on the at least one target audio device by controlling the speakers in the at least one target audio device to play corresponding test audio and the corresponding microphones in the at least one target audio device to collect the audio played by the speakers includes: combining the m first speakers and the n first microphones one by one to form m*n speaker-microphone pairs; for each speaker-microphone pair, controlling the first speaker in the speaker-microphone pair to play the first test audio and controlling the first microphone in the speaker-microphone pair to collect the current audio to obtain a first sampled audio, wherein the first test audio includes a continuous sweep frequency signal from low frequency to high frequency; determining a first performance index of the first microphone and the first speaker in the speaker-microphone pair based on the first sampled audio, wherein the first performance index includes frequency response and / or total harmonic distortion; and summing all the first performance indexes as the performance index detection result of the at least one target audio device.

[0247] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein the method of performing performance index detection on the at least one target audio device by controlling a speaker in the at least one target audio device to play a corresponding test audio and a corresponding microphone in the at least one target audio device to collect the audio played by the speaker, further includes: for each speaker-microphone pair, controlling the first speaker in the speaker-microphone pair to play a second test audio, and controlling the first microphone in the speaker-microphone pair to collect the current audio to obtain a second re-collected audio, wherein the second test audio includes speech segments covering low-frequency to high-frequency components and whose speech rate and pronunciation clarity both meet preset standards; determining a second performance index of the first microphone and the first speaker in the speaker-microphone pair based on the second re-collected audio, wherein the second performance index includes speech intelligibility and / or stuttering; and summarizing all the second performance indexes as the performance index detection result of the at least one target audio device.

[0248] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2, wherein the at least one target audio device further includes p conference all-in-one machines, where p ≥ 1; the method of performing performance index detection on the at least one target audio device by controlling the speakers in the at least one target audio device to play corresponding test audio and the corresponding microphones in the at least one target audio device to collect the audio played by the speakers further includes: determining a first state of the m first speakers and the n first microphones based on the performance index detection results of the m first speakers and the n first microphones, wherein the first state is used to characterize whether the corresponding device has malfunctioned; determining a target detection strategy for the conference all-in-one machines based on the first state; and performing performance index detection on each of the conference all-in-one machines according to the target detection strategy.

[0249] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, wherein determining the target detection strategy of the conference all-in-one machine based on the first state includes: if there is a first speaker that is not faulty and a first microphone that is not faulty, then the target detection strategy is determined to be to perform performance index detection on the conference all-in-one machine based on the target speaker and the target microphone, wherein the target speaker is any first speaker that is not faulty and the target microphone is any first microphone that is not faulty; if there is a first speaker that is not faulty and each first microphone is faulty, then the target detection strategy is determined to be to perform performance index detection on the conference all-in-one machine based on the target speaker; if each first speaker is faulty and there is a first microphone that is not faulty, then the target detection strategy is determined to perform performance index detection on the conference all-in-one machine based on the target microphone; if each first speaker and each first microphone are faulty, then the target detection strategy is determined to perform performance index detection on the conference all-in-one machine based on the second speaker and the second microphone of the conference all-in-one machine.

[0250] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target speaker and the target microphone, or to perform performance index detection on the conference all-in-one machine based on the target speaker; the performance index detection of each conference all-in-one machine according to the target detection strategy includes at least one of the following: for each conference all-in-one machine, controlling the target speaker to play the first test audio, and controlling the second microphone of the conference all-in-one machine to collect the current audio to obtain a third re-collected audio; determining the first performance index of the second microphone of the conference all-in-one machine based on the third re-collected audio; for each conference all-in-one machine, controlling the target speaker to play the third test audio, and controlling the second microphone of the conference all-in-one machine to collect the current audio to obtain a fourth re-collected audio, wherein the third test audio includes components covering low to high frequencies, speech rate and pronunciation clarity meeting preset standards, and includes speech segments with volume changing from low to high; and determining the first performance index of the second microphone of the conference all-in-one machine based on the fourth re-collected audio. The following steps are taken to determine the third performance index of the second microphone of the conference all-in-one machine and the automatic gain control index of the conference all-in-one machine, wherein the third performance index includes speech intelligibility and / or stuttering; for each conference all-in-one machine, the target speaker is controlled to play a fourth test audio, and the second microphone of the conference all-in-one machine is controlled to collect the current audio to obtain a fifth re-collected audio, wherein the fourth test audio includes speech segments covering low-frequency to high-frequency components, speech rate and pronunciation clarity all meeting the preset standards, and containing different noise types and signal-to-noise ratios; based on the fifth re-collected audio, the noise suppression index of the conference all-in-one machine is determined; for each conference all-in-one machine, the target speaker and the second speaker of the conference all-in-one machine are controlled to simultaneously play a second test audio, and the second microphone of the conference all-in-one machine is controlled to collect the current audio to obtain a sixth re-collected audio, wherein the second test audio includes speech segments covering low-frequency to high-frequency components, speech rate and pronunciation clarity all meeting the preset standards; based on the sixth re-collected audio, the echo cancellation index of the conference all-in-one machine is determined.

[0251] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 6, wherein the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target speaker and the target microphone; the step of performing performance index detection on each of the conference all-in-one machines according to the target detection strategy further includes: for each conference all-in-one machine, controlling the second speaker of the conference all-in-one machine to play the first test audio, and controlling the target microphone to collect the current audio to obtain a seventh re-collected audio; determining the first performance index of the second speaker of the conference all-in-one machine based on the seventh re-collected audio; and / or for each conference all-in-one machine, controlling the second speaker of the conference all-in-one machine to play the second test audio, and controlling the target microphone to collect the current audio to obtain an eighth re-collected audio, wherein the second test audio includes speech segments covering low-frequency to high-frequency components and whose speech rate and pronunciation clarity both meet preset standards; determining the second performance index of the second speaker of the conference all-in-one machine based on the eighth re-collected audio, wherein the second performance index includes speech intelligibility and / or stuttering.

[0252] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 5, wherein the target detection strategy is to perform performance index detection on the conference all-in-one machine based on the target microphone; the performance index detection of each conference all-in-one machine according to the target detection strategy includes at least one of the following: for each conference all-in-one machine, controlling the second speaker of the conference all-in-one machine to play the first test audio, and controlling the target microphone to collect the current audio to obtain a seventh re-collected audio; determining the first performance index of the second speaker of the conference all-in-one machine based on the seventh re-collected audio; for each conference all-in-one machine, controlling the second speaker of the conference all-in-one machine to play a second test audio, and controlling the target microphone to collect the current audio to obtain a seventh re-collected audio; determining the first performance index of the second speaker of the conference all-in-one machine based on the seventh re-collected audio; and controlling the second speaker of the conference all-in-one machine to play a second test audio, and controlling the target microphone to collect the current audio to obtain a seventh re-collected audio. The first audio is used to obtain the eighth audio recording, wherein the second test audio includes speech segments covering low to high frequencies and whose speech rate and pronunciation clarity meet preset standards; based on the eighth audio recording, the second performance index of the second speaker of the conference all-in-one machine is determined, wherein the second performance index includes speech intelligibility and / or stuttering; for each conference all-in-one machine, the second speaker of the conference all-in-one machine is controlled to play the second test audio, and the second microphone of the conference all-in-one machine is controlled to collect the current audio to obtain the ninth audio recording, wherein the second test audio includes speech segments covering low to high frequencies and whose speech rate and pronunciation clarity meet preset standards; based on the ninth audio recording, the echo cancellation index of the conference all-in-one machine is determined.

[0253] According to one or more embodiments of this disclosure, Example 9 provides the method of Example 1, wherein determining at least one target audio device to be inspected in a conference room audio system includes: obtaining a second state of each audio device in the conference room audio system, wherein the second state is used to characterize whether the audio device is online; and determining the audio device whose second state is online as the target audio device.

[0254] According to one or more embodiments of this disclosure, Example 10 provides the method of Example 9, which, prior to the step of determining the audio device in the second state as online as the target audio device, further includes: if there is an audio device in the second state as offline, sending selection information to a user terminal for the user to choose whether to skip the inspection of offline audio devices; in response to receiving a first selection instruction sent by the user terminal to indicate skipping the inspection of offline audio devices, or if there is no audio device in the second state as offline, performing the step of determining the audio device in the second state as online as the target audio device; in response to receiving a second selection instruction sent by the user terminal to indicate that the inspection includes offline audio devices, controlling the audio device in the second state as offline to start, and determining each of the audio devices in the conference room audio system as the target audio device.

[0255] According to one or more embodiments of this disclosure, Example 11 provides a method of any one of Examples 2-8, the method further comprising: determining a first state of the m first speakers and the n first microphones based on performance index detection results of the m first speakers and the n first microphones, wherein the first state is used to characterize whether the corresponding device has malfunctioned; if there is a first speaker that has not malfunctioned and there is a first microphone that has not malfunctioned, then using a target speaker and a target microphone, detecting the spatial acoustic environment of the physical conference room where the at least one target audio device is located, wherein the target speaker is any first speaker that has not malfunctioned and the target microphone is any first microphone that has not malfunctioned.

[0256] According to one or more embodiments of this disclosure, Example 12 provides a method of Example 11, wherein detecting the spatial acoustic environment of a physical conference room where at least one target audio device is located using a target loudspeaker and a target microphone includes: controlling the target loudspeaker to play a fifth test audio and controlling the target microphone to collect the current audio to obtain a tenth sampled audio, wherein the fifth test audio includes a pulse signal; denoising the tenth sampled audio; generating an energy decay curve of the denoised tenth sampled audio; and determining the reverberation time of the spatial acoustic environment based on the energy decay curve.

[0257] According to one or more embodiments of this disclosure, Example 13 provides the method of Example 12, which further includes: using a target loudspeaker and a target microphone to detect the spatial acoustic environment of a physical conference room where the at least one target audio device is located, and controlling the target microphone to collect the ambient sound signal of the physical conference room when the physical conference room is in a quiet state; extracting pure background noise information from the ambient sound signal; and determining the ambient noise floor of the spatial acoustic environment based on the pure background noise information.

[0258] According to one or more embodiments of this disclosure, Example 14 provides the method of Example 11, the method further comprising: if there is a first speaker that is not faulty and each of the first microphones is faulty, then using the target speaker and a third microphone to detect the spatial acoustic environment, wherein the third microphone is a device temporarily added to the physical conference room; if each of the first speakers is faulty and there is a first microphone that is not faulty, then using the third speaker and the target microphone to detect the spatial acoustic environment, wherein the third speaker is a device temporarily added to the physical conference room; if each of the first speakers and each of the first microphones is faulty, then using the third speaker and the third microphone to detect the spatial acoustic environment.

[0259] According to one or more embodiments of this disclosure, Example 15 provides a method of any one of Examples 1-8, wherein the conference room audio system includes audio devices and a conference terminal connected via a local area network, and the method is applied to the conference terminal; or the conference room audio system includes audio devices and a remote server connected via a wide area network, and the method is applied to the remote server.

[0260] According to one or more embodiments of this disclosure, Example 16 provides a method of any one of Examples 1-8, wherein the conference room audio system includes audio equipment, a conference terminal, and a remote server, wherein the conference terminal is connected to the audio equipment via a local area network (LAN), and the conference terminal is connected to the remote server via a wide area network (WAN); the remote server executes the inspection method using the conference terminal as a communication medium; or, determining at least one target audio device to be inspected in the conference room audio system includes: the conference terminal determining at least one target audio device to be inspected in the conference room audio system; and controlling the speakers in the at least one target audio device to play corresponding test audio and the corresponding microphones in the at least one target audio device to collect audio from the speakers. The method of playing audio and performing performance index testing on the at least one target audio device includes: the conference terminal controlling the speakers in the at least one target audio device to play corresponding test audio and the corresponding microphones in the at least one target audio device to collect the audio played by the speakers, thereby performing performance index testing on the at least one target audio device, and sending the performance index testing results of the at least one target audio device to the remote server; the step of generating an inspection report for the conference room audio system based on the performance index testing results of the at least one target audio device includes: the remote server generating an inspection report for the conference room audio system based on the performance index testing results of the at least one target audio device sent by the conference terminal.

[0261] According to one or more embodiments of this disclosure, Example 17 provides a method of any one of Examples 1-8, wherein the conference room audio system includes an audio device, a conference terminal, and a remote server, wherein the conference terminal is connected to the audio device via a local area network, and the conference terminal is connected to the remote server via a wide area network; the remote server executes the inspection method using the conference terminal as a communication medium.

[0262] According to one or more embodiments of this disclosure, Example 18 provides the method of Example 1, wherein generating an inspection report of the conference room audio system based on the performance index detection results of the at least one target audio device includes: generating an overall health status of the at least one target audio device based on the performance index detection results of the at least one target audio device, and / or determining a first state of the at least one target audio device to obtain an evaluation result, wherein the first state is used to characterize whether the corresponding device has malfunctioned; and generating an inspection report of the conference room audio system based on the overall health status and / or the evaluation result.

[0263] According to one or more embodiments of the present disclosure, Example 19 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any one of Examples 1-18.

[0264] According to one or more embodiments of the present disclosure, Example 20 provides a computer program product including a computer program that, when executed by a processor, implements the steps of the method described in any one of Examples 1-18.

[0265] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0266] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0267] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A method of auditing a conference room audio system, the method comprising: The method comprises: determining at least one target audio device to be inspected in a conference room audio system; detecting performance indicators of the at least one target audio device by controlling a loudspeaker in the at least one target audio device to play corresponding test audio and a corresponding microphone in the at least one target audio device to collect audio played by the loudspeaker; generating an inspection report of the conference room audio system according to the performance indicator detection results of the at least one target audio device; wherein the at least one target audio device comprises m first loudspeakers and n first microphones, wherein m≥1 and n≥1; the performance indicator detection of the at least one target audio device by controlling the loudspeaker in the at least one target audio device to play corresponding test audio and the corresponding microphone in the at least one target audio device to collect audio played by the loudspeaker comprises: combining the m first speakers one-to-one with the n first microphones to form speaker-microphone pairs; for each loudspeaker-microphone pair, controlling the first loudspeaker in the loudspeaker-microphone pair to play first test audio and controlling the first microphone in the loudspeaker-microphone pair to collect current audio to obtain first collected audio, wherein the first test audio comprises a continuous sweep frequency signal from low frequency to high frequency; determining first performance indicators of the first microphone and the first loudspeaker in the loudspeaker-microphone pair according to the first collected audio, wherein the first performance indicators comprise frequency response and / or total harmonic distortion; summarizing all the first performance indicators as the performance indicator detection results of the at least one target audio device.

2. The method of claim 1, wherein, the performance indicator detection of the at least one target audio device by controlling the loudspeaker in the at least one target audio device to play corresponding test audio and the corresponding microphone in the at least one target audio device to collect audio played by the loudspeaker further comprises: for each loudspeaker-microphone pair, controlling the first loudspeaker in the loudspeaker-microphone pair to play second test audio and controlling the first microphone in the loudspeaker-microphone pair to collect current audio to obtain second collected audio, wherein the second test audio comprises a voice segment covering components from low frequency to high frequency and meeting preset standards in terms of speech speed and pronunciation clarity; determining second performance indicators of the first microphone and the first loudspeaker in the loudspeaker-microphone pair according to the second collected audio, wherein the second performance indicators comprise voice intelligibility and / or stuttering; summarizing all the second performance indicators as the performance indicator detection results of the at least one target audio device.

3. The method of claim 1, wherein, the at least one target audio device further comprises p conference all-in-one machines, wherein p≥1; the performance indicator detection of the at least one target audio device by controlling the loudspeaker in the at least one target audio device to play corresponding test audio and the corresponding microphone in the at least one target audio device to collect audio played by the loudspeaker further comprises: According to the performance index detection results of the m first loudspeakers and the n first microphones, a first state of the m first loudspeakers and the n first microphones is determined, wherein the first state is used to represent whether a corresponding device fails; According to the first state, a target detection strategy of the conference integrated machine is determined; According to the target detection strategy, performance index detection is performed on each conference integrated machine.

4. The method of claim 3, wherein, The step of determining the target detection strategy of the conference integrated machine according to the first state comprises: If there is a first loudspeaker that does not fail and there is a first microphone that does not fail, the target detection strategy is determined to be performance index detection on the conference integrated machine based on a target loudspeaker and a target microphone, wherein the target loudspeaker is any first loudspeaker that does not fail and the target microphone is any first microphone that does not fail; If there is a first loudspeaker that does not fail and each first microphone fails, the target detection strategy is determined to be performance index detection on the conference integrated machine based on the target loudspeaker; If each first loudspeaker fails and there is a first microphone that does not fail, the target detection strategy is determined to be performance index detection on the conference integrated machine based on the target microphone; If each first loudspeaker and each first microphone fail, the target detection strategy is determined to be performance index detection on the conference integrated machine based on a second loudspeaker and a second microphone of the conference integrated machine.

5. The method of claim 4, wherein, The target detection strategy is performance index detection on the conference integrated machine based on a target loudspeaker and a target microphone, or performance index detection on the conference integrated machine based on the target loudspeaker; The step of performing performance index detection on each conference integrated machine according to the target detection strategy comprises at least one of the following: For each conference integrated machine, the target loudspeaker is controlled to play the first test audio, and the second microphone of the conference integrated machine is controlled to collect current audio to obtain third collected audio; and the first performance index of the second microphone of the conference integrated machine is determined according to the third collected audio; For each conference integrated machine, the target loudspeaker is controlled to play third test audio, and the second microphone of the conference integrated machine is controlled to collect current audio to obtain fourth collected audio, wherein the third test audio comprises components covering low frequencies to high frequencies, speech speed and pronunciation clarity that meet preset standards, and speech fragments with volume changing from low to high; and a third performance index of the second microphone of the conference integrated machine and an automatic gain control index of the conference integrated machine are determined according to the fourth collected audio, wherein the third performance index comprises speech intelligibility and / or lag. For each of the conference integrated machines, the target speaker is controlled to play fourth test audio, and the second microphone of the conference integrated machine is controlled to collect current audio to obtain seventh collected audio, wherein the fourth test audio includes speech segments covering low frequency to high frequency components, speech rate and pronunciation clarity conforming to the preset standard, and containing different noise types and signal-to-noise ratios; and the noise suppression index of the conference integrated machine is determined according to the seventh collected audio. For each of the conference integrated machines, the target speaker and the second speaker of the conference integrated machine are controlled to simultaneously play second test audio, and the second microphone of the conference integrated machine is controlled to collect current audio to obtain sixth collected audio, wherein the second test audio includes speech segments covering low frequency to high frequency components, speech rate and pronunciation clarity conforming to the preset standard; and the echo cancellation index of the conference integrated machine is determined according to the sixth collected audio.

6. The method of claim 5, wherein, The target detection strategy is to detect performance indexes of the conference integrated machine based on the target speaker and the target microphone; The performance index detection of each of the conference integrated machines according to the target detection strategy further includes: For each of the conference integrated machines, the second speaker of the conference integrated machine is controlled to play the first test audio, and the target microphone is controlled to collect current audio to obtain seventh collected audio; and the first performance index of the second speaker of the conference integrated machine is determined according to the seventh collected audio; and / or For each of the conference integrated machines, the second speaker of the conference integrated machine is controlled to play second test audio, and the target microphone is controlled to collect current audio to obtain eighth collected audio, wherein the second test audio includes speech segments covering low frequency to high frequency components, and speech rate and pronunciation clarity conforming to the preset standard; and the second performance index of the second speaker of the conference integrated machine is determined according to the eighth collected audio, wherein the second performance index includes speech intelligibility and / or lag.

7. The method of claim 4, wherein, The target detection strategy is to detect performance indexes of the conference integrated machine based on the target microphone; The performance index detection of each of the conference integrated machines according to the target detection strategy includes at least one of the following: For each of the conference integrated machines, the second speaker of the conference integrated machine is controlled to play the first test audio, and the target microphone is controlled to collect current audio to obtain seventh collected audio; and the first performance index of the second speaker of the conference integrated machine is determined according to the seventh collected audio; and / or For each of the conference integrated machines, the second speaker of the conference integrated machine is controlled to play a second test audio, and the target microphone is controlled to collect current audio to obtain an eighth collected audio, wherein the second test audio includes a speech segment covering low frequency to high frequency and having a speech speed and pronunciation clarity meeting preset standards; and a second performance index of the second speaker of the conference integrated machine is determined according to the eighth collected audio, wherein the second performance index includes speech intelligibility and / or stuttering. For each of the conference integrated machines, the second speaker of the conference integrated machine is controlled to play a second test audio, and the second microphone of the conference integrated machine is controlled to collect current audio to obtain a ninth collected audio, wherein the second test audio includes a speech segment covering low frequency to high frequency and having a speech speed and pronunciation clarity meeting preset standards; and an echo cancellation index of the conference integrated machine is determined according to the ninth collected audio.

8. The method of claim 1, wherein, The method further includes: obtaining a second state of each audio device in the conference room audio system, wherein the second state is used to represent whether the audio device is online; determining the audio device with the second state being online as the target audio device.

9. The method of claim 8, wherein, Before the step of determining the audio device with the second state being online as the target audio device, the method further includes: if there is the audio device with the second state being offline, sending selection information to a user terminal for the user to select whether to skip the patrol of the offline audio device; in response to receiving a first selection instruction sent by the user terminal for indicating the skip of the patrol of the offline audio device, or there is no audio device with the second state being offline, performing the step of determining the audio device with the second state being online as the target audio device; in response to receiving a second selection instruction sent by the user terminal for indicating the patrol of the offline audio device, controlling the audio device with the second state being offline to start, and determining each of the audio devices in the conference room audio system as the target audio device.

10. The method according to any one of claims 1-7, characterized in that, The method further includes: determining a first state of the m first speakers and the n first microphones according to the performance index detection results of the m first speakers and the n first microphones, wherein the first state is used to represent whether a corresponding device has failed; if there is the first speaker that has not failed and there is the first microphone that has not failed, detecting a spatial acoustic environment of an entity conference room where the at least one target audio device is located by using a target speaker and a target microphone, wherein the target speaker is any first speaker that has not failed, and the target microphone is any first microphone that has not failed.

11. The method of claim 10, wherein, The detecting the spatial acoustic environment of the entity conference room where the at least one target audio device is located by using the target speaker and the target microphone includes: control the target speaker to play a fifth test audio, and control the target microphone to collect a current audio to obtain a tenth collected audio, wherein the fifth test audio comprises a pulse signal; de-noise the tenth collected audio; generate an energy decay curve of the de-noised tenth collected audio; determine a reverberation time of the spatial acoustic environment according to the energy decay curve.

12. The method of claim 11, wherein, The method of detecting the spatial acoustic environment of the physical conference room where the at least one target audio device is located by using the target speaker and the target microphone further comprises: in a case that the physical conference room is in a quiet state, control the target microphone to collect an environmental sound signal of the physical conference room; extract pure background noise information from the environmental sound signal; determine an environmental background noise of the spatial acoustic environment according to the pure background noise information.

13. The method of claim 10, wherein, The method further comprises: if there is the first speaker that does not fail and each of the first microphone fails, detecting the spatial acoustic environment by using the target speaker and a third microphone, wherein the third microphone is a device temporarily added in the physical conference room; if each of the first speaker fails and there is the first microphone that does not fail, detecting the spatial acoustic environment by using a third speaker and the target microphone, wherein the third speaker is a device temporarily added in the physical conference room; if each of the first speaker and each of the first microphone fails, detecting the spatial acoustic environment by using the third speaker and the third microphone.

14. The method of any one of claims 1-7, wherein, The conference room audio system comprises audio devices and a conference terminal connected through a local area network, and the method is applied to the conference terminal; or The conference room audio system comprises audio devices and a remote server connected through a wide area network, and the method is applied to the remote server.

15. The method of any one of claims 1-7, wherein, The conference room audio system comprises audio devices, a conference terminal and a remote server, wherein the conference terminal is connected with the audio devices through a local area network, and the conference terminal is connected with the remote server through a wide area network; The remote server executes the inspection method by taking the conference terminal as a communication medium; or The remote server executes the inspection method by taking the conference terminal as a communication medium; or The determining of the at least one target audio device to be inspected in the conference room audio system comprises: the conference terminal determining the at least one target audio device to be inspected in the conference room audio system; the conference terminal performing the steps of detecting the performance indicators of the at least one target audio device by controlling the loudspeaker in the at least one target audio device to play corresponding test audio and the corresponding microphone in the at least one target audio device to collect the audio played by the loudspeaker; and the conference terminal further sending the performance indicator detection results of the at least one target audio device to the remote server; and the conference terminal generating the inspection report of the conference room audio system according to the performance indicator detection results of the at least one target audio device.

16. The method of any one of claims 1-7, wherein, The conference room audio system comprises audio devices, a conference terminal, and a remote server, wherein the conference terminal is connected to the audio devices through a local area network, and the conference terminal is connected to the remote server through a wide area network. The remote server executes the inspection method by taking the conference terminal as a communication medium.

17. The method of claim 1, wherein, The generating of the inspection report of the conference room audio system according to the performance indicator detection results of the at least one target audio device comprises: generating an overall health degree of the at least one target audio device and / or determining a first state of the at least one target audio device according to the performance indicator detection results of the at least one target audio device, to obtain an evaluation result, wherein the first state is used to represent whether the corresponding device has failed; generating the inspection report of the conference room audio system according to the overall health degree and / or the evaluation result.

18. A computer readable medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processing device, implements the steps of the method of any one of claims 1-17.

19. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-17.

Citation Information

Patent Citations

  • Automatic inspection method and device for meeting place, medium and electronic equipment

    CN114760432A

  • Automatic detection method for audio equipment in conference room

    CN118075676A