Audio network, audio network establishing method and recording method

By having master and slave cameras work together in an audio network to compare and overlay audio data quality in real time, and by updating audio equalizer parameters through self-learning, the problem of poor recording effect of security cameras in different environments is solved, achieving better recording effect and noise reduction capability.

CN120980407APending Publication Date: 2025-11-18SHENZHEN OCEANWING SMART INNOVATIONS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410620469.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing security cameras cannot update audio equalizer parameters according to actual conditions in different installation environments, resulting in poor audio data quality.

Method used

By working collaboratively between the master and slave cameras in an audio network, audio data quality is compared and superimposed in real time to optimize recording quality, and audio equalizer parameters are updated through self-learning and self-calibration.

Benefits of technology

It improves recording quality, achieves optimal audio data recording in different installation environments, reduces noise, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980407A_ABST
    Figure CN120980407A_ABST
Patent Text Reader

Abstract

The invention discloses an audio network, an audio network establishment method and a recording method, and the recording method comprises the steps: obtaining first audio data when collecting video data, and obtaining second audio data sent from a camera device; comparing the audio quality information of the first audio data with the audio quality information of the second audio data; and in response to the fact that the audio quality of the second audio data is higher than that of the first audio data, superposing the first audio data and the second audio data to obtain target audio data. According to the application, the first audio data is collected through the main camera device in the audio network, the second audio data sent by the slave camera device in the audio network is received at the same time, the first audio data and the second audio data are further compared, and if the audio quality of the second audio data is higher than the audio quality of the first audio data, the first audio data is sent to the slave camera device. Optimization of the first audio data is completed through superposition of the second audio data and the first audio data, and the recording effect of the main camera device can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of recording technology, and in particular to audio networks, methods for establishing audio networks, and recording methods. Background Technology

[0002] With technological advancements, acquiring audio and video data from remote camera devices has become a key focus in various technological fields, including security and education. Currently, security cameras record audio using their built-in microphones. The camera's System-on-Chip (SoC) chip and processing system convert the recorded analog sound signals into audio electrical signals, which are then integrated and encoded with the video signal data stream. Finally, the data stream is transmitted to a remote location via wireless or wired networks. The remote device receives the data stream, decodes it, and then plays back the recorded audio content.

[0003] However, current security cameras are manufactured with pre-set audio equalizer parameters based on standard installation conditions. This means that if users use the same audio equalizer parameters to record in different installation environments, the audio equalizer parameters cannot be updated according to the actual installation environment, and the recorded audio data cannot achieve the best results. Summary of the Invention

[0004] This application provides an audio network, a method for establishing the audio network, and a recording method, which can solve the problem that existing security cameras cannot provide optimal recording based on the installation environment.

[0005] The first aspect of this application provides a recording method applied to a main camera device in an audio network, the audio network further including a slave camera device, the main camera device being used to acquire first audio data, and the slave camera device being used to acquire second audio data, the recording method comprising:

[0006] When acquiring video data, first audio data is acquired, and second audio data sent from the camera device is acquired; the acquisition time of the second audio data is the same as the acquisition time of the first audio data.

[0007] Compare the audio quality information of the first audio data and the audio quality information of the second audio data, wherein the audio quality information is used to represent the audio quality of the first audio data or the second audio data;

[0008] Since the audio quality of the second audio data is stronger than that of the first audio data, the first audio data and the second audio data are superimposed to obtain the target audio data.

[0009] Further, the step of comparing the audio quality information of the first audio data and the audio quality information of the second audio data includes:

[0010] Obtain the first audio index of the first audio data and the second audio index of the second audio data;

[0011] The first audio metric and the second audio metric are compared to obtain a comparison result. The comparison result is that the first audio metric is greater than or equal to the second audio metric, or the first audio metric is less than the second audio metric.

[0012] Furthermore, in response to the fact that the audio quality of the second audio data is stronger than that of the first audio data, the target audio data is obtained by superimposing the first audio data and the second audio data, including:

[0013] In response to the comparison result that the first audio index is less than the second audio index, a first weight of the first audio data and a second weight of the second audio data are determined.

[0014] The target audio data is obtained by superimposing the first audio data and the second audio data based on the first weight and the second weight.

[0015] Further, in response to the comparison result that the first audio metric is less than the second audio metric, the step of determining the first weight of the first audio data and the second weight of the second audio data includes:

[0016] Calculate the difference between the first audio metric and the second audio metric;

[0017] If the difference in response indicators is greater than or equal to the preset parameter, the first weight is determined to be greater than the second weight.

[0018] If the difference in response metrics is less than the preset parameter, the first weight is determined to be less than the second weight.

[0019] Furthermore, the step of the main camera device acquiring the first audio data includes:

[0020] Read audio equalizer parameters;

[0021] Collect audio data and obtain the first audio data based on the audio equalizer parameters.

[0022] Furthermore, the recording methods also include:

[0023] In response to the second weight being greater than zero, obtain the initial audio equalizer parameters;

[0024] The initial audio equalizer parameters are calibrated based on the target audio data to obtain the target audio equalizer parameters;

[0025] Replace the initial audio equalizer parameters with the target audio equalizer parameters.

[0026] Furthermore, the recording methods also include:

[0027] Send the target audio data to the camera device;

[0028] The system receives noise-reduced frequency data fed back from the camera device. The noise-reduced frequency data is feedback data generated by the camera device based on the noise component of the target audio data; wherein the noise-reduced frequency data and the noise component are out of phase.

[0029] A second aspect of this application provides a method for establishing an audio network, the audio network including multiple camera devices that are communicatively connected to each other, the method for establishing the audio network including:

[0030] Acquire image data of the target object and video footage from multiple camera devices;

[0031] Compare the image data of the target object with multiple camera frames, and in response to the presence of the target object in any camera frame, determine one of the camera frames in which the target object exists as the target frame;

[0032] The camera device corresponding to the target image is identified as the main camera device, and the other camera devices are identified as slave camera devices; the main camera device is used to perform the recording method described above.

[0033] Furthermore, methods for establishing audio networks also include:

[0034] Acquire the characteristic sound wave and timestamp output by the main camera device; the timestamp is used to characterize the output time of the characteristic sound wave.

[0035] Obtain the current time of recording the characteristic sound waves from the camera device;

[0036] Get the delay time based on the current time and timestamp;

[0037] The distance between the main camera and the slave camera is determined based on the time delay.

[0038] An audio network is constructed based on distance, with the main camera device at the center.

[0039] A third aspect of this application provides an audio network, the audio network comprising:

[0040] The main camera device is used to capture the first audio data;

[0041] At least one camera device is used to acquire second audio data and send the second audio data to a main camera device; the acquisition time of the second audio data is the same as the acquisition time of the first audio data;

[0042] The main camera device is also used to receive second audio data, compare the audio quality information of the first audio data and the audio quality information of the second audio data, wherein the audio quality information is used to represent the audio quality of the first audio data or the second audio data; in response to the audio quality of the second audio data being stronger than the audio quality of the first audio data, the main camera device superimposes the first audio data and the second audio data to obtain target audio data.

[0043] Unlike existing technologies, the recording method of this application is applied to a main camera device in an audio network. The audio network also includes a slave camera device. The main camera device acquires first audio data and simultaneously receives second audio data acquired and transmitted by the slave camera device. The acquisition time of the second audio data is the same as that of the first audio data. The first and second audio data are further compared. When it is determined that the audio quality of the second audio data is stronger than that of the first audio data, the main camera device optimizes the first audio data by superimposing the second and first audio data, enabling the main camera device to acquire target audio data with better recording effect. Even if the main camera device of this application cannot acquire audio data with better recording effect on its own, it can still work with the slave camera device to optimize the acquired first audio data, effectively improving the recording effect of the main camera device.

[0044] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating the first embodiment of the recording method of this application;

[0047] Figure 2 yes Figure 1 A detailed flowchart of step S11 is shown below;

[0048] Figure 3 yes Figure 1 A detailed flowchart of step S12 is shown below;

[0049] Figure 4 yes Figure 1 A detailed flowchart of step S13 is shown below;

[0050] Figure 5 yes Figure 4A detailed flowchart of step S131 is shown below;

[0051] Figure 6 yes Figure 4 A detailed flowchart of step S132 is shown below;

[0052] Figure 7 This is a flowchart illustrating the second embodiment of the recording method of this application;

[0053] Figure 8 This is a flowchart illustrating the third embodiment of the recording method of this application;

[0054] Figure 9 This is a flowchart illustrating the first embodiment of the method for establishing an audio network according to this application;

[0055] Figure 10 This is a flowchart illustrating the second embodiment of the method for establishing an audio network according to this application;

[0056] Figure 11 This is a schematic diagram of the audio network structure of this application;

[0057] Figure 12 This is a schematic diagram of the framework of an embodiment of the electronic device of this application;

[0058] Figure 13 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0059] To enable those skilled in the art to better understand the technical solutions of this application, the audio network, audio network establishment method, and recording method provided in this application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It is understood that the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0060] The terms "first," "second," etc., used in this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0061] To address the problem that existing security cameras cannot update audio equalizer parameters according to the actual installation environment, resulting in suboptimal audio recording quality, this application provides a recording method that uses different audio data recorded simultaneously by multiple cameras for self-learning and self-calibration. This method can update the audio equalizer parameters stored in the camera devices in real time, thereby improving the quality of the recorded audio data.

[0062] The recording method of this application can be implemented by a camera device, specifically a master camera device in an audio network, which also includes several slave camera devices. Specifically, the audio network includes multiple camera devices, and these multiple camera devices are communicatively connected. Optionally, in one embodiment, all multiple camera devices can achieve wireless communication via Wi-Fi (Wireless Fidelity); or, all multiple camera devices can perform data communication via a wired Ethernet network; or, some of the multiple camera devices can perform data communication via Wi-Fi, while others can perform data communication via a wired Ethernet network.

[0063] Optionally, the camera device in this audio network may specifically refer to a security camera, which includes a microphone and a speaker. The microphone is used for recording, and the speaker is used for audio playback. Furthermore, the security camera device stores the typical frequency response curves for recording and playback, as well as the typical audio equalizer parameters required to achieve the optimal frequency response, in the device's storage medium at the time of manufacture.

[0064] Alternatively, in some possible implementations, the recording method of this application can also be implemented by the processor calling computer-readable instructions stored in memory.

[0065] Please see Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the recording method of this application. Specifically, the recording method of this disclosure may include the following steps:

[0066] Step S11: When acquiring video data, acquire first audio data and second audio data sent from the camera device.

[0067] In this embodiment, the multiple camera devices in the audio network can be divided into main camera devices and slave camera devices. There is one main camera device, which is used to collect audio and video data of the target object, and there is at least one slave camera device.

[0068] The main camera device collects the first audio data and the secondary camera device collects the second audio data at the same time. The main camera device and the secondary camera device record audio and video of the target object from different directions.

[0069] Optionally, the audio network of this embodiment may include multiple camera devices. Based on different installation environments or different locations of the monitored objects within the installation environment, one camera device is selected as the master camera device, and the remaining camera devices serve as slave camera devices. Optionally, the master camera device may be a camera device capable of capturing a complete image of the monitored object, or a camera device that is closest to the monitored object and can record the clearest audio data. Therefore, in different installation environments or usage scenarios, the same camera device can be used as either a master camera device or a slave camera device.

[0070] Specifically, in this embodiment, the main camera device and the slave camera device are communicatively connected. The main camera device can collect audio data on its own, and can also receive audio data sent by the slave camera device which is communicatively connected to the main camera device.

[0071] For further details on the specific process by which the main camera device acquires the first audio data, please refer to [link / reference needed]. Figure 2 , Figure 2 yes Figure 1 A detailed flowchart of step S11 is provided. Specifically, it includes the following steps:

[0072] Step S111: Read the audio equalizer parameters.

[0073] Each camera device stores the frequency response curve of the recording and playback, as well as the audio equalizer parameters required to achieve the frequency response in its storage medium. In this embodiment, the main camera device reads the audio equalizer parameters stored in its own storage medium.

[0074] Step S112: Collect audio data and obtain the first audio data based on the audio equalizer parameters.

[0075] In this embodiment, the main camera device continuously collects audio data and performs filtering, amplification, and other operations on the collected audio data based on the audio equalizer parameters read in step S111 to obtain the first audio data.

[0076] Step S12: Compare the audio quality information of the first audio data and the audio quality information of the second audio data.

[0077] In this embodiment, the audio quality information is used to represent the audio quality of the first audio data or the second audio data. The audio quality of the audio data can be specifically reflected in whether the recording effect is clear, and the signal-to-noise ratio and / or frequency response fitting degree of the speech stream can be used as specific characterization values.

[0078] Furthermore, please refer to the following for the specific process of the main camera device comparing the audio quality information of the first audio data and the audio quality information of the second audio data. Figure 3 , Figure 3 yes Figure 1 A detailed flowchart of step S12 is provided. Specifically, it includes the following steps:

[0079] Step S121: Obtain the first audio index of the first audio data and the second audio index of the second audio data.

[0080] The main camera device processes the first audio data collected by its own microphone through a built-in processing chip, and at the same time, the processing chip processes the second audio data collected from the camera device through wireless communication, thereby obtaining the first audio index of the first audio data and the second audio index of the second audio data.

[0081] Step S122: Compare the first audio index and the second audio index to obtain a comparison result. The comparison result is that the first audio index is greater than or equal to the second audio index, or the first audio index is less than the second audio index.

[0082] In this embodiment, the audio metrics include signal-to-noise ratio and standard frequency response fit. That is, both the first and second audio metrics include the corresponding signal-to-noise ratio and standard frequency response fit. The main camera device in this embodiment can obtain the comparison result by judging the audio metrics of different audio data.

[0083] Specifically, the comparison result can be that the first audio index is greater than or equal to the second audio index. In this case, based on the comparison result, it can be known that the audio quality of the first audio data is stronger than that of the second audio data, or that the two are equal. Therefore, no further processing is required on the first audio data.

[0084] Alternatively, the comparison result may show that the first audio index is less than the second audio index. In this case, based on the comparison result, it can be known that the audio quality of the second audio data is stronger than that of the first audio data. Therefore, further processing of the first audio data is needed to improve its audio quality.

[0085] Step S13: In response to the fact that the audio quality of the second audio data is stronger than that of the first audio data, the first audio data and the second audio data are superimposed to obtain the target audio data.

[0086] When the comparison result shows that the first audio index is less than the second audio index, it is determined that the audio quality of the second audio data is stronger than that of the first audio data. Data processing is required on the first audio data to obtain target audio data with better audio quality. Specifically, the target audio data is obtained by superimposing the first audio data and the second audio data.

[0087] Specifically, in this embodiment, the main camera device can determine the corresponding ratio based on the audio quality of the second audio data and the audio quality of the first audio data, and then superimpose the first audio data and the second audio data according to different ratios to obtain the target audio data. Alternatively, the main camera device can obtain a preset superposition ratio stored in the storage medium and superimpose the first audio data and the second audio data to obtain the target audio data.

[0088] Furthermore, please refer to the following for the specific process of how the main camera device superimposes the first and second audio data to obtain the target audio data. Figure 4 , Figure 4 yes Figure 1 A detailed flowchart of step S13 is provided. Specifically, it includes the following steps:

[0089] Step S131: In response to the comparison result that the first audio index is less than the second audio index, determine the first weight of the first audio data and the second weight of the second audio data.

[0090] In this embodiment, the main camera device, based on the comparison result obtained in step S12, can determine which of the first audio data and the second audio data has better recording quality, and can increase its weight ratio to improve the audio quality of the obtained target audio data. Specifically, the main camera device needs to determine the first weight of the first audio data and the second weight of the second audio data.

[0091] Furthermore, please refer to the following for the specific process by which the main camera device determines the first weight of the first audio data and the second weight of the second audio data. Figure 5 , Figure 5 yes Figure 4 A detailed flowchart of step S131 is provided. Specifically, it includes the following steps:

[0092] Step S1311: Calculate the difference between the first audio index and the second audio index.

[0093] In this embodiment, the main camera device subtracts the audio index of the second audio data from the audio index of the first audio data to obtain the index difference. Specifically, the main camera device can subtract the signal-to-noise ratio of the second audio data from the signal-to-noise ratio of the first audio data.

[0094] In this process, after determining the difference between the audio metrics of the first audio data and the audio metrics of the second audio data, it can be determined that the weight of the first audio data or the weight of the second audio data is larger.

[0095] Step S1312: If the difference in response indicators is greater than or equal to the preset parameter, determine that the first weight is greater than the second weight.

[0096] If the index difference calculated in step S1311 is greater than or equal to a preset parameter, where the preset parameter can be zero, meaning the signal-to-noise ratio of the first audio data is greater than that of the second audio data, it indicates that the recording quality of the first audio data is stronger than that of the second audio data, or that their recording qualities are equal. In this case, it can be determined that the first weight of the first audio data is greater than the second weight of the second audio data. Optionally, the second weight can be a value infinitely close to zero, such as 0.1, etc.

[0097] Step S1313: If the difference in response indicators is less than the preset parameter, determine that the first weight is less than the second weight.

[0098] If the index difference calculated in step S1311 is less than a preset parameter (which can be zero), meaning the signal-to-noise ratio of the first audio data is less than that of the second audio data, it indicates that the recording quality of the first audio data is weaker than that of the second audio data. In this case, it can be determined that the first weight of the first audio data is less than the second weight of the second audio data. Optionally, the second weight in this embodiment can be calculated based on the index difference.

[0099] Optionally, if there are multiple second audio data received, they can be compared one by one with the first audio data to obtain the second weights corresponding to the multiple second audio data.

[0100] Step S132: Based on the first weight and the second weight, the first audio data and the second audio data are superimposed to obtain the target audio data.

[0101] In this embodiment, the main camera device obtains target audio data by superimposing the first audio data and the second audio data based on the first weight of the first audio data and the second weight of the second audio data obtained in step S131.

[0102] Furthermore, the main camera device obtains the target audio data by superimposing the first audio data and the second audio data based on the first weight and the second weight. Please refer to the following for the specific process. Figure 6 , Figure 6 yes Figure 4 A detailed flowchart of step S132 is provided. Specifically, it includes the following steps:

[0103] Step S1321: Calculate the product of the first audio data and the first weight to obtain the first sub-audio data.

[0104] In this embodiment, the main camera device multiplies the first audio data and the first weight to obtain the first sub-audio data.

[0105] Step S1322: Calculate the product of the second audio data and the second weight to obtain the second sub-audio data.

[0106] In this embodiment, the main camera device multiplies the second audio data with the corresponding second weight to obtain the second sub-audio data. When there are multiple second audio data sets, the number of calculated second sub-audio data sets is also multiple.

[0107] Step S1323: Superimpose the first sub-audio data and the second sub-audio data to obtain the target audio data.

[0108] In this embodiment, the main camera device superimposes the first sub-audio data and all the second sub-audio data to obtain the target audio data, which is the audio data with the best recording effect that the main camera device can acquire under the current environment.

[0109] The recording method of this application is applied to a main camera device in an audio network, and the audio network also includes a slave camera device. The main camera device acquires first audio data and simultaneously receives second audio data acquired and transmitted by the slave camera device, further comparing the first audio data and the second audio data. When it is determined that the audio quality of the second audio data is stronger than that of the first audio data, the main camera device optimizes the first audio data by superimposing the second audio data and the first audio data, enabling the main camera device to acquire target audio data with better recording effect. Even if the main camera device of this application cannot acquire audio data with better recording effect on its own, it can still work with the slave camera device to optimize the acquired first audio data, effectively improving the recording effect of the main camera device.

[0110] Furthermore, this application also provides a recording method capable of self-learning and self-calibrating the main camera device, thereby further improving the recording effect of the main camera device. Please refer to [link to relevant documentation]. Figure 7 , Figure 7 This is a flowchart illustrating the second embodiment of the recording method of this application. Specifically, it includes the following steps:

[0111] Step S21: In response to the second weight being greater than zero, obtain the initial audio equalizer parameters.

[0112] When the main camera device determines, based on the comparison result of step S12, that the audio quality of the second audio data is stronger than that of the first audio data, it can be determined that the audio equalizer parameters stored in the main camera device are not the optimal audio equalizer parameters in the current installation environment, and need to be calibrated.

[0113] Specifically, the main camera device in this embodiment can read the typical frequency response curves of the recording and playback stored in the storage medium, as well as the typical audio equalizer parameters required to achieve the most ideal frequency response. The typical audio equalizer parameters are the initial audio equalizer parameters.

[0114] Step S22: Calibrate the initial audio equalizer parameters based on the target audio data to obtain the target audio equalizer parameters.

[0115] In this embodiment, the storage medium of the main camera device can store an audio parameter model. The audio parameter model can be any existing audio parameter model. The target audio data obtained in step S13 is input into the audio parameter model, and the target audio equalizer parameters can be obtained through model training.

[0116] Step S23: Replace the initial audio equalizer parameters with the target audio equalizer parameters.

[0117] In this embodiment, the main camera device replaces the initial audio equalizer parameters with the target audio equalizer parameters obtained in step S22, and stores the target audio equalizer parameters in the storage medium of the current camera device.

[0118] After the audio equalizer parameters are changed, the main camera device further processes the acquired audio data based on the target audio equalizer parameters to obtain the first audio data. The audio quality of the first audio data is stronger than that of the first audio data obtained by the main camera device based on the initial target audio equalizer parameters.

[0119] Specifically, in this embodiment, the main camera device can receive second audio data sent from other camera devices multiple times. Based on the multiple second audio data acquired from the camera devices, multiple target audio data are obtained. Furthermore, based on these multiple target audio data, the audio equalizer parameters are updated and calibrated multiple times, ultimately obtaining the optimal audio equalizer parameters. These parameters then replace and update the preset typical audio equalizer parameter settings initially considered to be required for the ideal frequency response, achieving self-learning and self-calibration. In this embodiment, the main camera device uses deep learning to update and calibrate the audio equalizer parameters, achieving self-learning and self-calibration, thereby enabling it to obtain the optimal audio equalizer parameters under different installation environments.

[0120] Meanwhile, by selecting the main camera device and the slave camera device, this embodiment can customize the best audio equalizer for the sound source or audience in the current installation environment based on the characteristic position of the sound source or audience in the audio space.

[0121] This application also provides another recording method that can effectively reduce noise in the audio data recorded by the main camera device; please refer to further details. Figure 8 , Figure 8 This is a flowchart illustrating the third embodiment of the recording method of this application. Specifically, it includes the following steps:

[0122] Step S31: Send the target audio data to the camera device.

[0123] The main camera device can wirelessly transmit the target audio data to the slave camera device. The processing chip built into the slave camera device processes the target audio data to obtain the noise component in the target audio data.

[0124] Step S32: Receive noise-reduced frequency data fed back from the camera device.

[0125] In this embodiment, the noise reduction frequency data is feedback data generated by the camera device based on the noise component of the target audio data. In particular, the noise reduction frequency data and the noise component are out of phase.

[0126] Specifically, the camera device outputs noise-reduced frequency data using its own speaker based on the noise components in the obtained target audio data. The camera device can read its distance from the main camera device and generate noise-reduced frequency data with an out-of-phase phase with the noise components based on this distance, further outputting this noise-reduced frequency data through its speaker. When the main camera device receives this noise-reduced frequency data, it can reduce noise, thereby achieving active noise reduction.

[0127] The main camera device of this application uses the speaker of the camera device as the sound source for active noise reduction, which can acquire audio data with extremely low background noise, extremely low distortion, high fidelity and high definition, thereby improving the user experience.

[0128] Furthermore, before performing step S11, it is necessary to construct an audio network and identify the master and slave camera devices among the multiple camera devices. For details on the process of identifying the master and slave camera devices among the multiple camera devices, please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a flowchart illustrating the first embodiment of the method for establishing an audio network according to this application. Specifically, it includes the following steps:

[0129] Step S41: Acquire image data of the target object and video footage from multiple camera devices.

[0130] Since the camera device in this embodiment is a security camera device, the purpose of which is to conduct security monitoring of a specific target, each camera device takes pictures at a specific angle, and the video images of each camera device are different, that is, the remote display device can display multiple video images at the same time.

[0131] Step S42: Compare the image data of the target object with multiple camera frames. In response to the presence of the target object in any camera frame, determine one of the camera frames containing the target object as the target frame.

[0132] In this embodiment, the main camera device is determined by selecting a target image from multiple camera feeds. The target image contains a target object, which is the object to be monitored in the current installation environment of the camera device. Specifically, it can be a person or an object. In particular, the target image can also be a frontal image of the target object.

[0133] Specifically, in this embodiment, a target object can be preset, and the image data of the target object can be stored in the storage medium of multiple camera devices. The camera devices read the stored image data of the target object and compare it with the captured video footage. When it is determined that the video footage contains the target object, the video footage is the target footage captured by monitoring the target object.

[0134] Step S43: Determine the camera device corresponding to the target image as the master camera device, and the other camera devices as slave camera devices.

[0135] In step S42, the target image can be determined, and the camera device used to capture the target image can be defined as the main camera device. Multiple camera devices can be defined as slave camera devices, and the slave camera devices can all be in a waiting state or standby mode, waiting to provide technical support for self-learning and self-calibration to the main camera device.

[0136] Specifically, in this embodiment, multiple camera devices are interconnected to form an audio network, and each camera device is a node in this audio network. For a detailed description of the process of constructing the audio network, please refer to [link to documentation]. Figure 10 , Figure 10 This is a flowchart illustrating the second embodiment of the method for establishing an audio network according to this application.

[0137] Specifically, it includes the following steps:

[0138] Step S44: Obtain the characteristic sound wave and timestamp output by the main camera device.

[0139] In this embodiment, the timestamp is used to characterize the output time of the characteristic sound wave. The main camera device sends the timestamp to the slave camera device via wireless communication and transmits the characteristic sound wave through a speaker.

[0140] Step S45: Obtain the current time of recording the characteristic sound wave from the camera device.

[0141] The camera device receives a timestamp sent by the main camera device via wireless communication, and generates a second timestamp upon receiving the first timestamp. The difference between the second timestamp and the first timestamp is the delay time of the wireless transmission.

[0142] At the same time, the characteristic sound wave is received from the camera device through the microphone, and the time when the characteristic sound wave is recorded is the current time.

[0143] Optionally, when there are multiple camera devices, the distance between each pair of camera devices can also be calculated using the above-described acoustic pairing method.

[0144] Step S46: Obtain the delay time based on the current time and timestamp.

[0145] In this embodiment, the delay time of the characteristic sound wave can be obtained by obtaining the difference between the current time and the timestamp.

[0146] Step S47: Determine the distance between the main camera device and the slave camera device based on the delay time.

[0147] The delay time of the characteristic sound wave is also the transmission time of the sound wave. The distance between the main camera device and the slave camera device can be calculated by the speed of sound propagation in the air and the transmission time of the sound wave.

[0148] Step S48: Construct an audio network based on distance, centered on the main camera device.

[0149] The audio network can be constructed based on the distance between the main camera device and the slave camera devices calculated in step S47, and based on preset distance parameters. The main camera device is the center of the audio network, and the slave camera devices are multiple nodes. The node coordinates are calculated based on the distance and preset distance parameters.

[0150] Optionally, when there are multiple camera devices, the distance between each pair of camera devices can also be calculated using the above-described acoustic pairing method.

[0151] Optionally, when switching the main camera device, i.e., when several secondary camera devices are switched to a new main camera device, a new audio network is constructed using the new main camera device. Specifically, after the relative positions of the multiple camera devices are initially calibrated, the corresponding position data is stored in memory. At preset intervals, the stored position data can be directly read to construct the corresponding audio network.

[0152] This embodiment, by pre-constructing an audio network, can quickly locate the sound source or audience when switching any camera device as the main camera device, specifically including the specific location of the person or other sound source or audience in space.

[0153] This application also provides an audio network for use as an application subject in any of the above-described recording method embodiments, and the audio network of this application can be constructed based on any of the above-described audio network establishment method embodiments. Please refer to... Figure 11 , Figure 11 This is a schematic diagram of the audio network structure of this application. (See diagram below.) Figure 11As shown, the audio network 50 in this embodiment includes a main camera device 51 and a slave camera device 52. The main camera device 51 is used to perform the steps of any of the above-described recording method embodiments.

[0154] In this embodiment, there is at least one slave camera device 52. The slave camera device 52 and the main camera device 51 are communicatively connected. Optionally, the communication connection can be made through a wired network or a wireless network, thereby realizing the transmission of control commands and audio data.

[0155] For example, in one embodiment, the main camera device 51 and all the slave camera devices 52 are connected via a wireless network; or, in one embodiment, the main camera device 51 and all the slave camera devices 52 are connected via a wired network; or, in one embodiment, the main camera device 51 and some of the slave camera devices 52 are connected via a wired network, and the other part of the slave camera devices 52 are connected via a wireless network.

[0156] Optionally, the audio network 50 in this embodiment may also include a data relay, such as a router. The main camera device 51 and all slave camera devices 52 can communicate with the router through a wired network or a wireless network to realize the transmission of control commands and audio data.

[0157] For example, in one embodiment, the main camera device 51 and all the slave camera devices 52 communicate with the router via a wireless network; or, in one embodiment, the main camera device 51 and all the slave camera devices 52 communicate with the router via a wired network; or, in one embodiment, the main camera device 51 and some of the slave camera devices 52 communicate with the router via a wired network, while the other part of the slave camera devices 52 communicate with the router via a wireless network; or, in one embodiment, the main camera device 51 and some of the slave camera devices 52 communicate with the router via a wireless network, while the other part of the slave camera devices 52 communicate with the router via a wired network.

[0158] Optionally, the communication methods between the main camera device 51 and the slave camera device 52, the communication methods between the main camera device 51 and the router, and the communication methods between the slave camera device 52 and the router can be combined arbitrarily, and this embodiment does not impose any restrictions.

[0159] Specifically, the main camera device 51 is used to acquire first audio data, and the secondary camera device 52 is used to acquire second audio data. The main camera device 51 can also receive the second audio data sent from the secondary camera device 52. The acquisition time of the second audio data is the same as the acquisition time of the first audio data. The main camera device 51 is also used to compare the audio quality information of the first audio data and the audio quality information of the second audio data, whereby the audio quality information represents the audio quality of either the first or second audio data. In response to the second audio data having a higher audio quality than the first audio data, the main camera device 51 is also used to superimpose the first and second audio data to obtain target audio data.

[0160] Optionally, the main camera device 51 can determine the corresponding weight ratio between the first audio data and the second audio data by comparing their audio quality information, and then perform data superposition by combining the corresponding weight ratio. The main camera device 51 is also used to read the audio equalizer parameters of the main camera device, calibrate the audio equalizer parameters based on the target audio data, and obtain the target audio equalizer parameters.

[0161] Optionally, in this embodiment, the secondary camera device 52 can receive target audio data sent by the primary camera device 51, and can also use its own speaker as a source to output noise-reduced frequency data with an opposite phase to the noise component in the target audio data, so as to achieve the effect of active noise reduction.

[0162] The audio network 50 of this application can update and calibrate the audio equalizer parameters in real time based on different installation environments, obtain the best audio equalizer parameters in the current installation environment, and at the same time, it can use the speaker of the camera device 52 as the sound source for active noise reduction, so that the main camera device 51 can acquire audio data with extremely low background noise, extremely low distortion, high fidelity and high definition, thereby improving the user experience.

[0163] This application also provides an electronic device, please refer to... Figure 12 , Figure 12 This is a schematic diagram of a framework of an embodiment of the electronic device of this application. The electronic device 60 includes a memory 61 and a processor 62 coupled to each other. The processor 62 is used to execute program instructions stored in the memory 61 to implement the steps in any of the above-described recording method embodiments. In a specific implementation scenario, the electronic device 60 may include, but is not limited to, a microcomputer, a server, etc. In addition, the electronic device 60 may also include mobile devices such as laptops and tablets, which are not limited here.

[0164] Specifically, processor 62 controls itself and memory 61 to implement the steps in any of the above-described recording method embodiments. Processor 62 can also be referred to as a CPU (Central Processing Unit). Processor 62 may be an integrated circuit chip with signal processing capabilities. Processor 62 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 62 can be implemented using integrated circuit chips.

[0165] This application also provides a computer-readable storage medium; please refer to [link to relevant documentation]. Figure 13 , Figure 13 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 70 stores a computer program 71 that can be executed by a processor. The computer program 71 is used to implement the steps in any of the above-described recording method embodiments.

[0166] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0167] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0168] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0169] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0170] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0171] The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

[0172] The above are merely embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A recording method applied to a master camera device in an audio network, the audio network further comprising slave camera devices, the master camera device for acquiring first audio data, and the slave camera devices for acquiring second audio data, characterized in that, The recording method includes: When acquiring video data, the first audio data is acquired, and the second audio data sent from the camera device is acquired; the acquisition time of the second audio data is the same as the acquisition time of the first audio data. Compare the audio quality information of the first audio data and the audio quality information of the second audio data, wherein the audio quality information is used to represent the audio quality of the first audio data or the second audio data; In response to the fact that the audio quality of the second audio data is stronger than that of the first audio data, the first audio data and the second audio data are superimposed to obtain the target audio data.

2. The recording method according to claim 1, characterized in that, The step of comparing the audio quality information of the first audio data and the audio quality information of the second audio data includes: Obtain the first audio index of the first audio data and the second audio index of the second audio data; The first audio metric and the second audio metric are compared to obtain a comparison result, wherein the first audio metric is greater than or equal to the second audio metric, or the first audio metric is less than the second audio metric.

3. The recording method according to claim 2, characterized in that, The step of responding to the second audio data having a higher audio quality than the first audio data, and superimposing the first audio data and the second audio data to obtain the target audio data includes: In response to the comparison result that the first audio index is less than the second audio index, a first weight of the first audio data and a second weight of the second audio data are determined; The target audio data is obtained by superimposing the first audio data and the second audio data based on the first weight and the second weight.

4. The recording method according to claim 3, characterized in that, The step of determining a first weight for the first audio data and a second weight for the second audio data in response to the comparison result that the first audio metric is less than the second audio metric includes: Calculate the difference between the first audio metric and the second audio metric; If the difference in the indicators is greater than or equal to a preset parameter, it is determined that the first weight is greater than the second weight. If the difference in the index is less than a preset parameter, it is determined that the first weight is less than the second weight.

5. The recording method according to claim 3, characterized in that, The step of acquiring the first audio data includes: Read audio equalizer parameters; Audio data is collected, and the first audio data is obtained based on the audio equalizer parameters.

6. The recording method according to claim 5, characterized in that, The recording method further includes: In response to the second weight being greater than zero, obtain the initial audio equalizer parameters; The initial audio equalizer parameters are calibrated based on the target audio data to obtain the target audio equalizer parameters; Replace the initial audio equalizer parameters with the target audio equalizer parameters.

7. The recording method according to claim 1, characterized in that, The recording method further includes: Send the target audio data to the camera device; The system receives noise-reduced frequency data fed back from the camera device, wherein the noise-reduced frequency data is feedback data generated by the camera device based on the noise component of the target audio data; wherein the noise-reduced frequency data and the noise component are out of phase.

8. A method for establishing an audio network, characterized in that, The audio network includes multiple camera devices, which are communicatively connected to each other. The method for establishing the audio network includes: Acquire image data of the target object and video footage from the multiple camera devices; Compare the image data of the target object with the multiple camera frames, and in response to the presence of the target object in any of the camera frames, determine one of the camera frames in which the target object is present as the target frame; The camera device corresponding to the target image is identified as the main camera device, and the other camera devices are identified as slave camera devices; wherein the main camera device is used to perform the recording method as described in any one of claims 1-8.

9. The method for establishing an audio network according to claim 8, characterized in that, The method for establishing the audio network also includes: Acquire the characteristic sound wave and timestamp output by the main camera device; wherein the timestamp is used to characterize the output time of the characteristic sound wave. Obtain the current time at which the characteristic sound wave is recorded from the camera device; The delay time is obtained based on the current time and the timestamp; The distance between the main camera device and the slave camera device is determined based on the delay time; The audio network is constructed based on the distance, with the main camera device as the center.

10. An audio network, characterized in that, include: The main camera device is used to capture the first audio data; At least one camera device is used to acquire second audio data and send the second audio data to the main camera device; the acquisition time of the second audio data is the same as the acquisition time of the first audio data; The main camera device is further configured to receive the second audio data, compare the audio quality information of the first audio data and the audio quality information of the second audio data, wherein the audio quality information is used to represent the audio quality of the first audio data or the second audio data; in response to the audio quality of the second audio data being stronger than the audio quality of the first audio data, the main camera device superimposes the first audio data and the second audio data to obtain target audio data.