Audio processing method and device for monitoring video, electronic equipment and medium

By dividing the surveillance video into segments and processing the audio, and using array microphones to determine the weighted sum of the image frame and the distance to the sound source, the problem of unstable audio signals in surveillance videos was solved, thus improving audio auxiliary functions.

CN116132879BActive Publication Date: 2026-04-14CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing surveillance video has problems such as unstable audio signals, unclear human voices at a distance, and excessive environmental noise when processing audio signals, which makes it impossible to effectively assist in judgment due to audio blind spots and video blind spots.

Method used

The surveillance video frame is divided by the arrangement of array microphones to determine the key video frame. The frame size and the distance to the sound source are combined and weighted to output the audio data of the actual distance.

Benefits of technology

The audio assistance function has been improved, which can change the distribution and number of microphones according to user needs, collect and link audio data in the full-frame monitoring video space, and output the real distance audio data of objects in the specified frame video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116132879B_ABST
    Figure CN116132879B_ABST
Patent Text Reader

Abstract

The application provides an audio processing method and device for monitoring video, electronic equipment and medium, which belong to the technical field of audio and video. The method comprises the following steps: when a full-frame monitoring video is acquired, dividing the picture of the full-frame monitoring video according to the arrangement order of an array microphone to obtain a plurality of picture videos; in response to a received key selection instruction, determining a key picture video indicated by the key selection instruction from the plurality of picture videos; determining the picture distance of an object in the key picture video based on the picture proportion of the key picture video in the full-frame monitoring video; determining the sound source distance of the object in the key picture video based on the sound source decibel proportion of a sound source object in the array microphone; weighting and summing the picture distance and the sound source distance to obtain the real distance of the object; and outputting the audio data of the real distance collected by the array microphone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio and video technology, and specifically relates to an audio processing method, apparatus, electronic device and medium for surveillance video. Background Technology

[0002] In scenarios such as smart community construction and criminal investigation, surveillance video is often an important monitoring method. It allows for real-time monitoring of security and illegal issues arising from a range of uncertainties within the visible environment.

[0003] However, in many scenarios where video footage is required to play back the surveillance footage and address the above issues, current surveillance videos often suffer from problems such as unstable audio signals, unclear human voices at a distance from the monitoring source, excessive environmental noise making it impossible to extract useful audio signals, numerous audio blind spots, and the inability to extract audio signals in video blind spots to assist in judgment. Summary of the Invention

[0004] This application provides an audio processing method, apparatus, electronic device, and medium for surveillance video.

[0005] Some embodiments of this application provide an audio processing method for surveillance video, the method comprising:

[0006] When the full-frame surveillance video is acquired, the image of the full-frame surveillance video is divided according to the arrangement order of the array microphones to obtain multiple frame videos;

[0007] In response to a received emphasis selection instruction, the emphasis frame video indicated by the emphasis selection instruction is determined from a plurality of said frame videos;

[0008] Based on the proportion of the key frame video in the full-frame surveillance video, the frame distance of objects in the key frame video is determined;

[0009] Based on the decibel ratio of the sound source object in the key frame video in the array microphone, the sound source distance of the object in the key frame video is determined;

[0010] The true distance of the object is obtained by weighted summation of the image distance and the sound source distance;

[0011] Output the audio data of the actual distance collected by the array microphones.

[0012] Optionally, determining the frame distance of objects in the key frame video based on the proportion of the key frame video in the full-frame surveillance video includes:

[0013] Calculate the aspect ratio of the key frame video to the full-frame surveillance video;

[0014] The frame distance is obtained by dividing the video adjustment parameters by the frame ratio.

[0015] Optionally, determining the sound source distance of the object in the key frame video based on the proportion of the sound source decibels of the object in the array microphones includes:

[0016] Obtain the decibel percentage of the sound source of the microphone corresponding to the key video frame in the array microphone;

[0017] The distance to the sound source is obtained by using the percentage of decibels at which the audio adjustment parameters are applied to the sound source.

[0018] Optionally, before obtaining the decibel percentage of the sound source corresponding to the microphone in the array microphones for the key video frame, the method further includes:

[0019] Determine whether the timing of the video of the key frame is the same as the recording audio of the corresponding microphone;

[0020] If the recording audio timing of the key frame video is different from that of the corresponding microphone, a microphone with the same recording audio timing as the key frame video should be selected.

[0021] Optionally, before determining the key frame video indicated by the key selection instruction from the plurality of frame videos in response to the received key selection instruction, the method further includes:

[0022] Display multiple video frames as described;

[0023] In response to the selection operation of the key frame video in the frame video, a key selection instruction is generated.

[0024] Optionally, the step of dividing the full-frame surveillance video according to the arrangement order of the array microphones to obtain multiple microphone-corresponding video frames includes:

[0025] Obtain the number and arrangement order of the microphones in the array microphone;

[0026] The full-frame surveillance video is divided into the number of video frames corresponding to the number of microphones, according to the arrangement order, with each video frame corresponding to one microphone.

[0027] Optionally, the step of weighted summing of the image distance and the sound source distance to obtain the true distance of the object includes:

[0028] The true distance is obtained by combining the product of the sound source distance and the sound source weight, and the product of the frame distance and the frame weight.

[0029] Some embodiments of this application provide an audio processing apparatus for surveillance video, the apparatus comprising:

[0030] The audio and video processing module is used to divide the full-frame monitoring video into multiple frame videos according to the arrangement order of the array microphones when the full-frame monitoring video is acquired.

[0031] The focus selection module is used to determine the focus frame video indicated by the focus selection instruction from a plurality of said frame videos in response to a received focus selection instruction;

[0032] The distance detection module is used to determine the frame distance of an object in the key frame video based on the frame proportion of the key frame video in the full-frame monitoring video; to determine the sound source distance of the object in the key frame video based on the sound source decibel proportion of the sound source object in the key frame video in the array microphone; and to perform a weighted summation of the frame distance and the sound source distance to obtain the true distance of the object.

[0033] The result output module is used to output the audio data of the actual distance collected by the array microphone.

[0034] Optionally, the distance detection module is further configured to:

[0035] Calculate the aspect ratio of the key frame video to the full-frame surveillance video;

[0036] The frame distance is obtained by dividing the video adjustment parameters by the frame ratio.

[0037] Optionally, the distance detection module is further configured to:

[0038] Obtain the decibel percentage of the sound source of the microphone corresponding to the key video frame in the array microphone;

[0039] The distance to the sound source is obtained by using the percentage of decibels at which the audio adjustment parameters are applied to the sound source.

[0040] Optionally, the distance detection module is further configured to:

[0041] Determine whether the timing of the video of the key frame is the same as the recording audio of the corresponding microphone;

[0042] If the recording audio timing of the key frame video is different from that of the corresponding microphone, a microphone with the same recording audio timing as the key frame video should be selected.

[0043] Optionally, the selected module is also used for

[0044] Display multiple video frames as described;

[0045] In response to the selection operation of the key frame video in the frame video, a key selection instruction is generated.

[0046] Optionally, the audio and video processing module is further configured to:

[0047] Obtain the number and arrangement order of the microphones in the array microphone;

[0048] The full-frame surveillance video is divided into the number of video frames corresponding to the number of microphones, according to the arrangement order, with each video frame corresponding to one microphone.

[0049] Optionally, the distance detection module is further configured to:

[0050] The true distance is obtained by combining the product of the sound source distance and the sound source weight, and the product of the frame distance and the frame weight.

[0051] Some embodiments of this application provide a computing processing device, including:

[0052] Memory containing computer-readable code;

[0053] One or more processors, when the computer-readable code is executed by the one or more processors, the computing processing device performs the audio processing method for the surveillance video as described above.

[0054] Some embodiments of this application provide a non-transient computer-readable medium that stores computer-readable code, which, when run on a computing processing device, causes the computing processing device to perform the aforementioned audio processing method for surveillance video.

[0055] This application provides an audio processing method, apparatus, electronic device, and medium for surveillance video. By adding a focus frame selection function, the distribution and number of microphones can be changed according to user needs, and all audio in the space where the full-frame surveillance video is located can be collected. The audio is linked with different frame videos to output audio data of the actual distance of objects in the focus frame video specified by the user, thereby improving the audio auxiliary function.

[0056] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 The schematic diagram illustrates a flowchart of an audio processing method for surveillance video provided in some embodiments of this application;

[0059] Figure 2 The diagram illustrates a device schematically representing an audio processing method for surveillance video provided in some embodiments of this application.

[0060] Figure 3 The illustration schematically shows a logical diagram of an audio processing method for surveillance video provided in some embodiments of this application;

[0061] Figure 4 This illustration shows one of the schematic diagrams illustrating the principle of an audio processing method for surveillance video provided in some embodiments of this application;

[0062] Figure 5 The second schematic diagram illustrates the principle of an audio processing method for surveillance video provided in some embodiments of this application.

[0063] Figure 6 The third schematic diagram illustrates the principle of an audio processing method for surveillance video provided in some embodiments of this application;

[0064] Figure 7 The schematic diagram illustrates the structure of an audio processing device for surveillance video provided in some embodiments of this application;

[0065] Figure 8 A block diagram schematically illustrates a computing processing apparatus for performing methods according to some embodiments of this application;

[0066] Figure 9 A storage unit for holding or carrying program code implementing methods according to some embodiments of this application is illustrated schematically. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0068] Figure 1 The schematic diagram illustrates a flowchart of an audio processing method for surveillance video provided in this application, the method comprising:

[0069] Step 101: When the full-frame surveillance video is acquired, the image of the full-frame surveillance video is divided according to the arrangement order of the array microphones to obtain multiple frame videos.

[0070] Step 102: In response to the received emphasis selection instruction, determine the emphasis frame video indicated by the emphasis selection instruction from the plurality of frame videos.

[0071] Step 103: Based on the proportion of the key frame video in the full-frame monitoring video, determine the frame distance of the object in the key frame video.

[0072] Step 104: Determine the sound source distance of the object in the key frame video based on the decibel ratio of the sound source object in the array microphone.

[0073] Step 105: The distance between the image frame and the distance between the sound source are weighted and summed to obtain the true distance of the object.

[0074] Step 106: Output the audio data of the actual distance collected by the array microphone.

[0075] Reference Figure 2 The execution device in some embodiments of this application includes: a video acquisition and transmission module, an array microphone and transmission module, a video processing module, an audio processing module, a focus selection module, a distance detection module, and a result display module.

[0076] The video acquisition and transmission module is used to acquire video footage of an event area within a certain time period and spatial segment to be detected, and transmit the entire video frame to the video processing module.

[0077] The video processing module has pre-set data, primarily consisting of array microphone arrangement data. The video images captured by the video acquisition module are divided according to the arrangement and specifications of the array microphones. Each frame in the microphone array arrangement is named frame λ. There is a one-to-one correspondence between each frame and a single microphone in the pre-set microphone array of the audio processing module.

[0078] An array microphone and transmission module collects sound signals from different directions in space and transmits them to an audio processing module. The array microphone in the transmission module consists of a certain number of microphones that sample and filter the spatial characteristics of the sound field. The microphones have consistent frequency responses, and their sampling clocks are synchronized. Microphone arrays are generally used in sound source localization technology, including angle and distance measurement, suppression of background noise, interference, reverberation, and echo, signal extraction, and signal separation. This sound source localization technology uses the microphone array to calculate the angle and distance between the sound source and the array, enabling the tracking of the target sound source.

[0079] The audio processing module has a preset microphone array, in which each microphone is named microphone β. Each microphone in the array has a unique mapping relationship with the entire frame of the video processing module, i.e., β corresponds to λ. Each microphone β operates independently without affecting others.

[0080] The audio processing module also has a built-in audio processing algorithm. According to the user's needs for the audio signal, the audio signal is processed through a specific composite filter. The filter includes human voice analysis (distinguishing between voiced and unvoiced voices), noise removal (Gaussian white noise, environmental noise, etc.), amplitude enhancement (signal attenuation and restoration of small-scale and large-scale noise), and other composite filter methods for audio signal processing.

[0081] The key detection module is for users to select specific image frames to be processed and transmit the user's intentions to the audio processing module via the video processing module, where the audio processing module will further process the data.

[0082] The distance detection module works simultaneously with the video processing module and the audio processing module to detect audio distance and frame distance.

[0083] Reference Figure 3 The following is a logical schematic diagram illustrating an audio processing method for surveillance video provided in some embodiments of this application:

[0084] S1, Initialize each module. Divide and arrange the video images captured by the video acquisition module according to the arrangement order and specifications of the array microphones. The arrangement data is preset to the video processing module.

[0085] S2, the video surveillance enters recording mode and captures full-frame video;

[0086] S3, transmit the full-frame video to the video processing module, at which point the video processing module divides the full-frame according to the built-in frame size of the video processing module;

[0087] S4, the video processing module determines whether the important selection module should select the key frame λ. If there is no key selection module, the frame selection is performed.

[0088] S5, the distance detection module acts on the video processing module to perform distance processing on the selected frame of the focus selection module, and obtain the processed distance d. 画幅 And stored in the video processing module;

[0089] The processed frame distance is obtained by dividing the video adjustment parameter e by the frame ratio;

[0090] S6, the audio processing module selects microphone β from the corresponding array microphones in the selected frame λ;

[0091] S7, determine whether the audio recording timing of the image frame λ and the microphone β is the same. If they are not the same, return to S6 to reselect and adjust the audio. If they are the same, execute S8.

[0092] S8, the distance detection module acts on the audio processing module, selecting the corresponding microphone in the audio processing module based on the selected frame size by the aforementioned key selection module, and determining the frame size ratio H of the sound source object in frame λ. λ Perform distance processing to obtain the processed distance d. 声源 And it is stored in the audio processing module;

[0093] The method for handling the sound source distance is: dividing the audio adjustment parameter f by the decibel ratio.

[0094] S9, the audio processing module extracts the d from the video processing module. 画幅 Then, based on the sound source object's decibel ratio V at the array microphone β, β Perform distance processing, and based on the protection d 画幅 With d 声源 The final true distance d is determined by the weighting coefficients A and B in the true distance d, and the frame and audio adjustment parameters e and f:

[0095] S10, the audio processing module retains the audio signal of microphone 1 at the actual distance d, and the other array microphones temporarily store the audio signal in Ω;

[0096] S11: The audio processing module only processes audio signals at the actual distance d.

[0097] This application embodiment adds a focus frame selection function, which can change the distribution and number of microphones according to user needs, and collect all audio in the space where the full-frame monitoring video is located, and link the audio with different frame videos to output audio data of the actual distance of objects in the focus frame video specified by the user, thereby improving the audio auxiliary function.

[0098] Optionally, step 103 includes:

[0099] A1, Calculate the aspect ratio of the key frame video to the full-frame monitoring video;

[0100] A2. Divide the video adjustment parameters by the frame ratio to obtain the frame distance.

[0101] In the embodiments of this application, reference is made to Figure 4 In the video processing module, the distance of the object to the video monitoring source is determined based on its size and proportion within the full frame. A larger object occupies more space, indicating a closer proximity to the monitoring source. In the audio processing module, based on the video processing module's determination and the audio amplitude transmitted from the array microphones, the distance of the object to the monitoring source is assessed. Then, based on the input from the key detection module, audio signals outside the detection range are removed, and the remaining audio signals are processed.

[0102] Optionally, step 104 includes:

[0103] B1, determine whether the timing of the video of the key frame is the same as the recording audio of the corresponding microphone;

[0104] B2. If the recording audio timing of the key frame video and its corresponding microphone is different, a microphone with the same recording audio timing as the key frame video shall be selected.

[0105] B3, obtain the decibel ratio of the sound source of the microphone corresponding to the key video frame in the array microphone;

[0106] B4. The distance to the sound source is obtained by using the audio adjustment parameters at the decibel level of the sound source.

[0107] In the embodiments of this application, reference is made to Figure 5 The closer the sound source is to the microphone, the higher the decibel ratio of the sound source. Therefore, based on the sound source object's position relative to the array microphone β, the decibel ratio V is determined. β Perform distance processing, V β It follows an exponential distribution with a unit mean between (0, 1). Then, according to d... 画幅 With d 声源The weighting coefficients A and B in the true distance d are commonly A = 0.4 and B = 0.6; the frame and audio adjustment parameters e and f determine the final true distance d; commonly e = 5 and f = 4.

[0108] Optionally, prior to step 102, the method further includes:

[0109] C1 displays multiple video frames as described;

[0110] C2, in response to the selection operation of the key frame video in the frame video, generates a key selection instruction.

[0111] In the embodiments of this application, reference is made to Figure 6 It provides full-frame surveillance video containing multiple frame videos for users to select the key frame video, thereby enabling the output of audio data on the target distance of the object in the key frame video.

[0112] Optionally, step 101 includes:

[0113] D1, obtain the number and arrangement order of the microphones in the array microphones;

[0114] D2, the full-frame surveillance video is divided into the number of video frames corresponding to the number of microphones according to the arrangement order, and each video frame corresponds to one microphone.

[0115] In this embodiment of the application, when the array microphones are arranged in an M*N order, the full-frame monitoring video is divided into M*N proportionally, with each frame of video being the same size and each frame of video corresponding to a microphone.

[0116] Optionally, step 105 includes:

[0117] The true distance is obtained by combining the product of the sound source distance and the sound source weight, and the product of the frame distance and the frame weight.

[0118] In this embodiment of the application, the true distance can be calculated using the following formula (1):

[0119]

[0120] Where A is the weighting parameter for frame distance; B is the weighting parameter for sound source distance; H λ V represents the proportion of the frame occupied by the sound source object within the frame λ; β is the decibel ratio of the sound source when the object is in the array microphone β; e is the frame adjustment parameter, and f is the audio frequency adjustment parameter. Output d is the distance between the actual sound source and the monitored source.

[0121] Examples of audio processing methods for surveillance videos provided in some embodiments of this application are as follows:

[0122] Step 1: Initialize each module and set the array microphone to N*M. For common segmentation ratios, N can be set to 4 and M to 5. Users can adjust these parameters later according to different application needs. The video processing module has a built-in segmentation frame of n*m, where n = N = 4 and m = M = 5. Each microphone unit in the array microphone assembly can work independently and collect audio from different directions.

[0123] Step 2: The video surveillance enters recording mode and shoots full-frame video. The specific parameters of the video surveillance can be the commonly used parameters: resolution 1440P, lens pixels 4 million, 110° wide angle, shooting angle horizontal 360°, vertical 360°, and the commonly used full-frame aspect ratio 16:9.

[0124] Step 3: Transmit the full-frame video to the video processing module;

[0125] Preferably, the video processing module divides the full-frame image into 20 parts according to the built-in 4*5 aspect ratio;

[0126] Step 4: The video processing module determines whether the important selection module selects the key frame λ. In this video processing module, all frames will be sorted in order from left to right and from top to bottom, that is, frame 1 to frame 20. At this time, λ=1, so frame 1 is taken as an example. If there is no important selection module, the frame selection will be performed again.

[0127] Step 5: Determine if the distance detection module is applied to the video processing module, and select the frame for distance processing in the focus selection module to obtain the processed distance d. 画幅 And stored in the video processing module;

[0128] Step 6: The audio processing module selects microphone 1 from the corresponding array microphones for the selected frame 1. The arrangement order of the array microphones is consistent with the frame segmentation order, that is, microphones 1 to 20 are arranged from left to right and from top to bottom.

[0129] Step 7: Determine if the audio recording timing of frame 1 and microphone 1 is the same. If they are not the same, return to Step 6 to select again and adjust the audio.

[0130] Step 8: Determine whether the distance detection module is applied to the audio processing module, select the frame for the focus selection module, and determine the frame ratio H based on the proportion of the sound source object in frame 1. λ Distance processing, taking a common application scenario as an example, can obtain H. λ The only constant, H, after testing λIt follows an exponential distribution with a unit mean between (0, 1). The processed distance d 声源 And it is stored in the audio processing module;

[0131] Step 9: The audio processing module extracts the d from the video processing module. 画幅 Then, based on the sound source object's decibel ratio V at the array microphone β, β Perform distance processing, V β It follows an exponential distribution with a unit mean between (0, 1). Then, according to d... 画幅 With d 声源 The weighting coefficients A and B in the true distance d are commonly A = 0.4 and B = 0.6; the frame and audio adjustment parameters e and f determine the final true distance d; commonly e = 5 and f = 4.

[0132] Step 10: The audio processing module retains the audio signal from microphone 1 at the actual distance d, while the audio signals from the other array microphones are temporarily stored in Ω;

[0133] Step 11: The audio processing module only processes audio signals at the actual distance d.

[0134] The values ​​of some parameters of environmental noise are shown in Table 1 below:

[0135]

[0136] Table 1

[0137] Figure 7 The schematic diagram illustrates the structure of an audio processing device 20 for surveillance video provided in this application, the device comprising:

[0138] The audio and video processing module 201 is used to divide the full-frame monitoring video into multiple frame videos according to the arrangement order of the array microphones when the full-frame monitoring video is acquired.

[0139] The focus selection module 202 is used to determine the focus frame video indicated by the focus selection instruction from a plurality of frame videos in response to a received focus selection instruction;

[0140] The distance detection module 203 is used to determine the frame distance of an object in the key frame video based on the frame proportion of the key frame video in the full-frame monitoring video; to determine the sound source distance of the object in the key frame video based on the sound source decibel proportion of the sound source object in the key frame video in the array microphone; and to perform a weighted summation of the frame distance and the sound source distance to obtain the true distance of the object.

[0141] The result output module 204 is used to output the audio data of the actual distance collected by the array microphone.

[0142] Optionally, the distance detection module 203 is further configured to:

[0143] Calculate the aspect ratio of the key frame video to the full-frame surveillance video;

[0144] The frame distance is obtained by dividing the video adjustment parameters by the frame ratio.

[0145] Optionally, the distance detection module 203 is further configured to:

[0146] Obtain the decibel percentage of the sound source of the microphone corresponding to the key video frame in the array microphone;

[0147] The distance to the sound source is obtained by using the percentage of decibels at which the audio adjustment parameters are applied to the sound source.

[0148] Optionally, the distance detection module 203 is further configured to:

[0149] Determine whether the timing of the video of the key frame is the same as the recording audio of the corresponding microphone;

[0150] If the recording audio timing of the key frame video is different from that of the corresponding microphone, a microphone with the same recording audio timing as the key frame video should be selected.

[0151] Optionally, the focus selection module 202 is also used for

[0152] Display multiple video frames as described;

[0153] In response to the selection operation of the key frame video in the frame video, a key selection instruction is generated.

[0154] Optionally, the audio and video processing module 201 is further configured to:

[0155] Obtain the number and arrangement order of the microphones in the array microphone;

[0156] The full-frame surveillance video is divided into the number of video frames corresponding to the number of microphones, according to the arrangement order, with each video frame corresponding to one microphone.

[0157] Optionally, the distance detection module 203 is further configured to:

[0158] The true distance is obtained by combining the product of the sound source distance and the sound source weight, and the product of the frame distance and the frame weight.

[0159] This application embodiment adds a focus frame selection function, which can change the distribution and number of microphones according to user needs, and collect all audio in the space where the full-frame monitoring video is located, and link the audio with different frame videos to output audio data of the actual distance of objects in the focus frame video specified by the user, thereby improving the audio auxiliary function.

[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0161] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the computing processing device according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a non-transient computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0162] For example, Figure 4 A computing processing apparatus is shown that can implement the methods according to this application. This computing processing apparatus conventionally includes a processor 310 and a computer program product or non-transitory computer-readable medium in the form of a memory 320. The memory 320 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory 320 has a storage space 330 for program code 331 for performing any of the method steps described above. For example, the storage space 330 for program code may include various program codes 331 respectively for implementing the various steps in the methods described above. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. Such computer program products are typically as shown in the reference. Figure 5The portable or fixed storage unit is described above. This storage unit may have the same characteristics as... Figure 4 The memory 320 in the computing processing device is arranged similarly to storage segments, storage spaces, etc. Program code can be compressed, for example, in an appropriate form. Typically, the storage unit includes computer-readable code 331', that is, code that can be read by a processor such as 310, which, when run by the computing processing device, causes the computing processing device to perform the various steps in the methods described above.

[0163] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0164] The terms "an embodiment," "embodiment," or "one or more embodiments" as used herein mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this application. Furthermore, please note that the examples of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0165] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0166] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An audio processing method for surveillance video, characterized in that, The method includes: When the full-frame surveillance video is acquired, the image of the full-frame surveillance video is divided according to the arrangement order of the array microphones to obtain multiple frame videos; In response to a received emphasis selection instruction, the emphasis frame video indicated by the emphasis selection instruction is determined from a plurality of said frame videos; Based on the proportion of the key frame video in the full-frame surveillance video, the frame distance of objects in the key frame video is determined; Based on the decibel ratio of the sound source object in the key frame video in the array microphone, the sound source distance of the object in the key frame video is determined; The true distance of the object is obtained by weighted summation of the image distance and the sound source distance; Output the audio data of the actual distance collected by the array microphones.

2. The method according to claim 1, characterized in that, Determining the frame distance of objects in the key frame video based on the proportion of the key frame video in the full-frame surveillance video includes: Calculate the aspect ratio of the key frame video to the full-frame surveillance video; The frame distance is obtained by dividing the video adjustment parameters by the frame ratio.

3. The method according to claim 1, characterized in that, Determining the sound source distance of an object in the key frame video based on the decibel ratio of the sound source object in the array microphones includes: Obtain the decibel percentage of the sound source of the microphone corresponding to the key frame video in the array microphone; The distance to the sound source is obtained by dividing the audio adjustment parameters by the decibel ratio of the sound source.

4. The method according to claim 3, characterized in that, Before obtaining the decibel percentage of the sound source corresponding to the microphone in the array microphone for the key frame video, the method further includes: Determine whether the timing of the video of the key frame is the same as the recording audio of the corresponding microphone; If the recording audio timing of the key frame video is different from that of the corresponding microphone, a microphone with the same recording audio timing as the key frame video should be selected.

5. The method according to claim 1, characterized in that, Before determining the key frame video indicated by the key selection instruction from the plurality of frame videos in response to the received key selection instruction, the method further includes: Display multiple video frames as described; In response to the selection operation of the key frame video in the frame video, a key selection instruction is generated.

6. The method according to claim 1, characterized in that, The process of dividing the full-frame surveillance video according to the arrangement order of the array microphones to obtain multiple video frames corresponding to different microphones includes: Obtain the number and arrangement order of the microphones in the array microphone; The full-frame surveillance video is divided into the number of video frames corresponding to the number of microphones, according to the arrangement order, with each video frame corresponding to one microphone.

7. The method according to claim 1, characterized in that, The step of weighted summing of the image distance and the sound source distance to obtain the true distance of the object includes: The true distance is obtained by combining the product of the sound source distance and the sound source weight, and the product of the frame distance and the frame weight.

8. An audio processing device for surveillance video, characterized in that, The device includes: The audio and video processing module is used to divide the full-frame monitoring video into multiple frame videos according to the arrangement order of the array microphones when the full-frame monitoring video is acquired. The focus selection module is used to determine the focus frame video indicated by the focus selection instruction from a plurality of said frame videos in response to a received focus selection instruction; The distance detection module is used to determine the frame distance of an object in the key frame video based on the frame proportion of the key frame video in the full-frame monitoring video; to determine the sound source distance of the object in the key frame video based on the sound source decibel proportion of the sound source object in the key frame video in the array microphone; and to perform a weighted summation of the frame distance and the sound source distance to obtain the true distance of the object. The result output module is used to output the audio data of the actual distance collected by the array microphone.

9. A computing processing device, characterized in that, include: Memory containing computer-readable code; One or more processors, when the computer-readable code is executed by the one or more processors, the computing processing device performs the audio processing method for surveillance video as described in any one of claims 1-7.

10. A non-transient computer-readable medium, characterized in that, The computer-readable code is stored, and when the computer-readable code is run on a computing processing device, the computing processing device causes the computing processing device to perform the audio processing method for surveillance video as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for positioning abnormal sound source in video image

    CN105554443A

  • Method and device for switching video images

    CN109905616A