Data processing method and computer equipment
By adjusting the shooting and sound pickup parameters in the monitoring system, focusing on the target area and performing precise directional sound pickup, the problem of audio interference by noise in public open spaces is solved, and high-quality audio and video evidence collection and identification are achieved.
Patent Information
- Application Number
- CN202411280412.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2026-03-13
AI Technical Summary
In public open spaces, the audio data of monitoring systems is easily interfered with by environmental noise, resulting in poor audio quality and affecting the effectiveness of evidence collection for abnormal events. Existing technologies are unable to achieve effective audio and video recognition in open spaces.
The location information of abnormal events is determined by the imaging device, and the parameters of the imaging and sound pickup devices are adjusted to focus on the target area, shield external noise, and perform precise directional sound pickup when recognizing human voices to enhance the quality of human voices.
It improves the quality of audio and video evidence collection, enhances the identification of abnormal events, and is applicable to any public open space.
Smart Images

Figure CN121665103A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio and video processing, and more particularly to a data processing method and a computer device. Background Technology
[0002] In public open spaces, areas prone to fights and brawls typically require the installation of surveillance systems. When an incident occurs, the audio and video data collected by the system can be used to determine liability. During such incidents, human voices are of particular interest and have a significant impact on the determination of responsibility.
[0003] Although cameras today generally have audio and video capture capabilities, they are severely affected by environmental noise. Audio is easily submerged by environmental noise or has a very low signal-to-noise ratio, which makes audio data unable to play an effective role in scenarios such as evidence collection. Usually, only video data is used for evidence collection.
[0004] With the development of artificial intelligence technology, some monitoring systems have the ability to identify abnormal events in real time. However, due to poor audio quality, intelligent algorithms are mostly based on video data. The few audio and video abnormal behavior recognition systems exist only in enclosed spaces such as ATM rooms and vertical elevators, and are not suitable for public open spaces with severe environmental noise interference. Summary of the Invention
[0005] This application provides a data processing method for obtaining spatial location information (i.e., first location information) of an abnormal event through a shooting device, and adjusting the parameters of the shooting device and the sound pickup device accordingly (coarse adjustment process) to focus on the area where the abnormal event occurs (i.e., the target area), thus shielding noise outside the target area. When a human voice is detected within the target area, the parameters of the sound pickup device are further fine-tuned to achieve precise directional sound pickup of the target object in the first sub-area (fine adjustment process). At this point, the sound pickup device, after further parameter adjustment, accurately enhances the human voice of the target object, making the picked-up human voice clear and effectively improving the quality of audio and video evidence collection. This application is not limited to the location of the abnormal event and can be applied to any public open space.
[0006] Based on this, the embodiments of this application provide the following technical solutions:
[0007] Firstly, this application provides a method that specifically includes: First, determining the first location information of an abnormal event actually occurring in physical space using a shooting device. This physical space, also known as geographic space, refers to the three-dimensional space of biological activity, and its location information can be represented by a physical coordinate system (also known as a world coordinate system). In this application's embodiment, the abnormal event refers to events such as fights, brawls, and arguments that may cause social security problems. The specific scope of the abnormal event can be defined based on the type of data to be acquired and the purpose of acquisition (e.g., whether it is for evidence collection), which will not be elaborated upon in this application. Then, based on the acquired first location information of the abnormal event in physical space, adjusting the shooting parameters of the shooting device and the sound pickup parameters of the sound pickup device so that the shooting device and the sound pickup device focus on the target area where the abnormal event occurred. After initial adjustments to the shooting and sound pickup parameters, when a sound source is detected in the target area, one or more sub-regions can be determined based on this sound source. These sub-regions can be referred to as the first sub-region, which is the area where one or more target objects in the abnormal event are located. These target objects refer to the people in the sound-emitting state. After determining the first sub-region from the target area, the sound pickup parameters can be adjusted again to limit the sound pickup range to the first sub-region.
[0008] In the above embodiments of this application, the parameters of the audio-visual related devices undergo two adjustment processes: coarse adjustment and fine adjustment. First, the first location information of the abnormal event is obtained through the shooting device, and the parameters of the shooting device and the sound pickup device are adjusted accordingly to focus on the area where the abnormal event occurred (i.e., the target area). This can shield noise outside the target area; this parameter adjustment process is the coarse adjustment process. When a human voice is detected in the target area, the parameters of the sound pickup device are further fine-tuned to achieve accurate directional sound pickup of the target object in the first sub-area; this parameter adjustment process is the fine adjustment process. In this application, the sound pickup device, after further parameter adjustment, accurately enhances the human voice of the target object, making the picked-up human voice clear and effectively improving the quality of audio-visual evidence collection. This application is not limited to the location where the abnormal event occurred and can be applied to any public open space.
[0009] In one possible implementation of the first aspect, the method of determining the first sub-region from the target region based on the sound source can be as follows: First, when a sound source is emitting sound, it is first determined whether the sound source is a human voice. If so, it means that the sound source is the desired target sound source (which can be one or more). The location information of the sound source can then be determined by the sound pickup device (one sound source corresponds to one location information, so the location information can also be one or more). This location information can be called the second location information. After that, the first sub-region can be determined from the target region based on the second location information.
[0010] In the above embodiments of this application, a specific method for determining the first sub-region based on the sound source is described, so that noise outside the abnormal event region is shielded and only the sound within the region is picked up, thereby improving the recognition effect of abnormal events.
[0011] In one possible implementation of the first aspect, the method of determining the first sub-region from the target region based on the second location information can be as follows: First, the shooting device takes a picture of the object in the target region. The shooting device can determine the location information of the person (which may be one or more) in the abnormal event in the physical space. This location information can be called the third location information. Then, the target object (which may be one or more, and the target object refers to the person in the vocal state) is determined based on the second location information and the third location information. Finally, the first sub-region is determined from the target region based on the target object.
[0012] In the above embodiments of this application, the implementation method of determining the first sub-region based on the second directional information is specifically described, thereby achieving accurate voice positioning enhancement.
[0013] In one possible implementation of the first aspect, when the number of target objects is 1, the implementation of determining the first sub-region from the target region based on the target object can be: determining the target detection box of the target object and using the target detection box as the first sub-region, wherein the target detection box is contained in the target region.
[0014] In the above embodiments of this application, when the number of target objects is 1, the first sub-region is the target detection box region of the target object, which is feasible.
[0015] In one possible implementation of the first aspect, when there are multiple target objects, determining the first sub-region from the target region based on the target objects can be achieved by: determining the outer bounding region (e.g., the minimum outer bounding region) of these multiple target objects, and using this outer bounding region as the first sub-region, which is contained within the target region. This outer bounding region can be obtained from the target detection boxes of these multiple target objects; for example, it can be the union of the regions occupied by the respective target detection boxes of these multiple target objects. The outer bounding region can also be obtained through other methods, which are not limited in this application.
[0016] In the above embodiments of this application, when there are multiple target objects, the first sub-region is the outer region of these multiple target objects, which has universal applicability.
[0017] In one possible implementation of the first aspect, if the pickup parameters of the sound pickup device are adjusted to a first value based on the first directional information, and then adjusted again to a second value, the method may further include, after the sound source (human voice) stops emitting sound, adjusting the sound pickup device back to the coarsely adjusted pickup parameters, that is, adjusting the pickup parameters from the second value back to the first value. Alternatively, after the sound source (human voice) stops emitting sound, the steps of determining the actual directional information of the abnormal event in physical space through the shooting device, adjusting the shooting parameters of the shooting device, and adjusting the pickup parameters of the sound pickup device based on the first directional information can be repeated, so that the shooting device and the sound pickup device refocus on the new target area.
[0018] In the above embodiments of this application, it is specifically described that when there is voice in the target area, the first sub-region is focused for more accurate voice pickup, and when there is no voice in the target area, the audio in the target area is recorded, which has flexibility.
[0019] In one possible implementation of the first aspect, adjusting the shooting parameters of the shooting device can be achieved by adjusting the rotation angle of the shooting device so that the abnormal event is located in the center area of the imaging screen of the shooting device.
[0020] In the above embodiments of this application, one method of adjusting the shooting parameters of the shooting device is described, which can make the shooting of abnormal events more comprehensive.
[0021] In one possible implementation of the first aspect, another way to adjust the shooting parameters of the shooting device is to adjust the focal length of the shooting device according to the ratio between the number of pixels of the abnormal event in the target frame image (which may be referred to as the first number of pixels) and the total number of pixels in the target frame image, so that the proportion of the abnormal event in the image of the shooting device is appropriate (i.e., meets the preset requirements), for example, the first number of pixels / the total number of pixels = 0.7. Here, the target frame image is the frame image in which the abnormal event is detected.
[0022] In the above embodiments of this application, another way of adjusting the shooting parameters of the shooting device is described, which can make the shooting of abnormal events clearer.
[0023] In one possible implementation of the first aspect, adjusting the pickup parameters of the pickup device can be achieved by: pre-establishing a sound field coordinate system (a coordinate system stationary relative to the pickup device); then, based on the first azimuth information, determining the azimuth information of the abnormal event in the sound field coordinate system according to the mapping relationship between the sound field coordinate system and the physical coordinate system; this azimuth information of the abnormal event in the sound field coordinate system can be called fourth azimuth information. Subsequently, based on this fourth azimuth information, adjusting the rotation angle of the pickup device so that the line connecting the center of the abnormal event (which can be called the first center) and the center of the microphone array within the pickup device (which can be called the second center) is perpendicular to the microphone array.
[0024] In the above embodiments of this application, an implementation method for adjusting the pickup parameters of the pickup device is described, so that the audio and video in the abnormal event area are clear and the environmental noise is significantly suppressed.
[0025] In one possible implementation of the first aspect, determining the first location information of an abnormal event in physical space using an imaging device could involve detecting the abnormal event from captured video data or multiple consecutive frame images, where the frame image where the abnormal event was detected is called the target frame image. Then, the first location information of the abnormal event in physical space is determined based on the pixel position of the abnormal event in the target frame image. For example, the first location information of the abnormal event in physical space can be determined based on the mapping relationship between the image coordinate system and the physical coordinate system.
[0026] In the above embodiments of this application, a specific implementation method for determining the first orientation information is described, which is feasible.
[0027] A second aspect of this application provides a computer device having the function of implementing the method of the first aspect or any possible implementation thereof. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described function.
[0028] A third aspect of this application provides a computer device that may include a memory, a processor, and a bus system. The memory is used to store a computer program (also referred to as a program or computer-readable instructions), and the processor is used to invoke the program stored in the memory to execute the method of the first aspect of the embodiments of this application or any possible implementation of the first aspect.
[0029] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0030] The fifth aspect of this application provides a computer program or a computer program product containing instructions that, when the computer program or computer program product is run on a computer, causes the computer to perform the method described in the first aspect or any possible implementation of the first aspect.
[0031] A sixth aspect of this application provides a chip including at least one processor and at least one interface circuit coupled to the processor. The interface circuit performs transceiver functions and sends instructions to the at least one processor. The at least one processor runs a computer program or instructions, having the functionality to implement the methods described in the first aspect or any possible implementation of the first aspect. This functionality can be implemented in hardware, software, or a combination of hardware and software, including one or more modules corresponding to the described functions. Furthermore, the interface circuit is used to communicate with other modules outside the chip.
[0032] In some implementations of this application, some of the one or more processors may implement some steps of the above method through dedicated hardware. For example, the processing involving neural network models may be implemented by a dedicated neural network processor or graphics processor.
[0033] The method provided in this application embodiment can be implemented by a single chip or by multiple chips working together. Attached Figure Description
[0034] Figure 1 A schematic diagram of the system architecture provided in the embodiments of this application;
[0035] Figure 2 A flowchart illustrating the data processing method provided in this application embodiment;
[0036] Figure 3 A schematic diagram illustrating an example of an embodiment of this application;
[0037] Figure 4 This is another example diagram provided for an embodiment of this application;
[0038] Figure 5 A schematic diagram of a computer device provided in an embodiment of this application;
[0039] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0040] This application provides a data processing method and computer equipment for two adjustment processes: coarse and fine adjustment of the parameters of audio-visual related devices. First, the location information of the abnormal event is obtained through a shooting device, and the parameters of the shooting device and the audio pickup device are adjusted accordingly to focus on the area where the abnormal event occurred (i.e., the target area). This effectively blocks noise outside the target area; this parameter adjustment process is the coarse adjustment process. When a human voice is detected within the target area, the parameters of the audio pickup device are further fine-tuned to achieve precise directional audio pickup of the target object in the first sub-area; this parameter adjustment process is the fine adjustment process. In this application, the audio pickup device, after further parameter adjustment, accurately enhances the human voice of the target object, making the picked-up voice clear and effectively improving the quality of audio-visual evidence collection. This application is not limited to the location of the abnormal event and can be applied to any public open space.
[0041] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0042] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0043] First, the system architecture and overall process of the method in the embodiments of this application will be described. Please refer to [link / reference needed] for details. Figure 1 , Figure 1This is a schematic diagram of the system architecture provided in an embodiment of this application. The system architecture includes a camera device 100 and a microphone device 200. The camera device 100 is used to acquire video / image data, and the microphone device 200 is used to acquire ambient audio data. The camera device 100 performs abnormal event detection based on real-time imaging, while the microphone device 200 picks up ambient audio in real time. After the camera device 100 detects an abnormal event, the microphone device can coarsely adjust its microphone array parameters based on the physical spatial location information of the abnormal event to establish a regional sound barrier. The sound barrier is used to define a specific area of sound, within which the microphone device can pick up sound, while sound outside the specific area is blocked. Simultaneously, after detecting an abnormal event, the camera device 100 also adjusts its parameters to focus on the target area where the abnormal event occurred. When a human voice is detected within the target area, the microphone device 200 further fine-tunes its parameters to achieve precise directional sound pickup from the target object.
[0044] In this embodiment, the shooting device 100 and the audio pickup device 200 track the dynamic changes of the target area and adaptively adjust their respective parameters to maintain the video / image data and audio data captured from different angles at an optimal level. For ease of explanation, the video / image data and audio data can be collectively referred to as audio-visual data. Finally, the enhanced audio-visual data output by the shooting device 100 and the audio pickup device 200 after parameter adjustment can be applied to various practical application scenarios. As an example, the enhanced audio-visual data can be saved to a disk for subsequent evidence collection; as another example, the enhanced audio-visual data can also be input into an abnormal event detection algorithm to improve the detection effect of the shooting device 100 on abnormal events and enhance the capabilities of the entire system architecture.
[0045] It should be noted that, in the embodiments of this application, Figure 1 The system architecture described herein is for illustrative purposes only and does not limit the deployment method of each device. For example, the shooting device 100 and the sound pickup device 200 can be deployed on the same computer device, or the shooting device 100 and the sound pickup device 200 can be deployed independently. This application does not limit this.
[0046] Based on the above system architecture, the data processing method provided in the embodiments of this application will be described below. This method is applied to a shooting device and a sound pickup device, wherein there can be one or more shooting devices, and one or more sound pickup devices. The shooting device and the sound pickup device can be deployed on the same computer device or deployed independently; this application does not limit this. Please refer to [link / reference] for details. Figure 2 , Figure 2 A flowchart illustrating the data processing method provided in this application embodiment specifically includes the following steps:
[0047] 201. Determine the first location information of the actual occurrence of the abnormal event in physical space through the imaging device.
[0048] First, the location information of the abnormal event in physical space is determined by the imaging device; this location information can be called the first location information. This physical space can also be called geographic space, referring to the three-dimensional space of biological activity. The location information of this physical space can also be called physical location, which can be represented by a physical coordinate system (also called the world coordinate system).
[0049] It should be noted that, in the embodiments of this application, the abnormal events refer to events such as fighting, brawling, and quarreling that may cause social security problems. The specific scope of abnormal events can be defined based on the data type to be obtained and the purpose of obtaining the data (e.g., whether it is used for evidence collection), which will not be elaborated in this application.
[0050] In some embodiments of this application, the capturing device (or a device on which the capturing device is deployed) may activate an abnormal event detection algorithm to detect abnormal events from captured video data or multiple consecutive frame images. The frame image in which the abnormal event is detected is called the target frame image. Then, based on the pixel position of the abnormal event in the target frame image, the first location information of the actual occurrence of the abnormal event in physical space is determined. For example, the first location information of the actual occurrence of the abnormal event in physical space can be determined based on the mapping relationship between the image coordinate system and the physical coordinate system.
[0051] 202. Adjust the shooting parameters of the shooting device and adjust the sound pickup parameters of the sound pickup device according to the first azimuth information so that the shooting device and the sound pickup device focus on the target area when the abnormal event occurs.
[0052] Subsequently, based on the first location information of the abnormal event in physical space, the shooting parameters of the shooting device and the sound pickup parameters of the sound pickup device can be adjusted so that the shooting device and the sound pickup device focus on the target area when the abnormal event occurs.
[0053] The following sections explain how to adjust the shooting parameters and sound pickup parameters:
[0054] (1) Adjustment of shooting parameters of the shooting device
[0055] In one implementation, the rotation angle of the imaging device can be adjusted so that the detected abnormal event is located in the center area of the imaging device's image.
[0056] For example, suppose the pixel position of the center of the area where the abnormal event is captured by the shooting device in the image is P0. When the shooting device has a built-in gimbal, the rotation angle of the shooting device can be determined based on the offset between P0 and the center position of the image, so as to keep the abnormal event in the center of the image.
[0057] In another implementation, the focal length of the shooting device can be adjusted according to the ratio between the number of pixels of the abnormal event in the target frame image (which can be called the first number of pixels) and the total number of pixels in the target frame image, so that the proportion of the abnormal event in the imaging screen of the shooting device is appropriate (i.e., meets the preset requirements). For example, the first number of pixels / the total number of pixels = 0.7.
[0058] It should be noted that the above only illustrates the adjustment process of two of the shooting parameters of the shooting device. Other shooting parameters can be adjusted in a similar way, and will not be elaborated here.
[0059] (2) Adjustment of pickup parameters of the pickup device
[0060] In the embodiments of this application, the initial adjustment process of the pickup parameters of the pickup device is also the process of establishing a sound barrier. The so-called sound barrier is used to define a specific area of sound, in which the sound can be picked up by the pickup device, while the sound outside the specific area is blocked.
[0061] Specifically, a sound field coordinate system (a coordinate system stationary relative to the pickup device) can be pre-established. Then, based on the first directional information, the directional information of the abnormal event in the sound field coordinate system is determined according to the mapping relationship between the sound field coordinate system and the physical coordinate system. This directional information of the abnormal event in the sound field coordinate system can be called the fourth directional information. Subsequently, based on this fourth directional information, the rotation angle of the pickup device is adjusted so that the line connecting the center of the abnormal event (which can be called the first center) and the center of the microphone array within the pickup device (which can be called the second center) is perpendicular to the microphone array.
[0062] In this embodiment, the orientation information of the target area boundary in the sound field coordinate system can be determined based on the mapping relationship between the target area boundary and the physical space, and the mapping relationship between the sound field coordinate system and the physical space. Based on this, the area can be divided into a sound pickup area (i.e., the area where the abnormal event occurs) and a shielding area. Finally, beamforming sound pickup technology is used to enhance the sound in the sound pickup area and suppress the sound in the shielding area when the microphone array picks up the sound, thus forming a sound curtain for the target area.
[0063] 203. If a sound source is present in the target area, determine the first sub-region from the target area based on the sound source. This first sub-region is the area where the target object in the abnormal event is located, and the target object is the person in the sound-emitting state.
[0064] After initially adjusting the shooting and sound pickup parameters, when a sound source is detected in the target area, one or more sub-regions can be determined from the target area based on the sound source. These one or more sub-regions can be called the first sub-region. The first sub-region is the area where one or more target objects in the abnormal event are located. Here, these one or more target objects refer to the people in the sound-emitting state.
[0065] It should be noted that in some embodiments of this application, the method for determining the first sub-region from the target area based on the sound source can be as follows: First, when a sound source is emitted, it is first determined whether the sound source is a human voice. If so, it indicates that the sound source is the desired target sound source (which can be one or more). The location information of the sound source can then be determined by the sound pickup device (one sound source corresponds to one location information, therefore, the location information can also be one or more). This location information can be called the second location information. Then, the first sub-region can be determined from the target area based on the second location information. Specifically, the sound pickup device can pick up all the sounds in the target area, and then determine the human voice (e.g., a person's voice) as the target sound source based on the audio feature information of the sound source in space. The audio feature can be a sound frequency or other features. For example, the frequency of a human voice is generally between 200-600Hz. Afterward, the sound pickup device locates the target sound source based on the microphone array directional sound pickup technology, that is, it determines the location information of the target sound source (i.e., the second location information mentioned above) based on the time difference between the target sound source and each microphone in the microphone array and the distance between each microphone. Finally, the first sub-region can be determined based on the second location information of the target sound source.
[0066] It should be noted that in some other embodiments of this application, the method of determining the first sub-region from the target region based on the second location information may be as follows: First, the shooting device takes a picture of the object in the target region. The shooting device can determine the location information of the person (which may be one or more) in the abnormal event in the physical space. This location information may be called the third location information. Then, the target object (which may be one or more, and the target object refers to the person in the voice state) is determined based on the second location information and the third location information. Finally, the first sub-region is determined from the target region based on the target object.
[0067] For example, assuming the imaging device determines that the abnormal event includes four people: A, B, C, and D, we can obtain the physical location information of each of these four people, let's say directions a, b, c, and d. These four locations belong to the third location information. Now, suppose there are two target sound sources, P and Q, and their physical location information is p and q, respectively. These two locations belong to the second location information. Then, based on directions a, b, c, and d, and directions p and q, we can determine the target objects. Assuming directions a and p are considered consistent within the error range, and directions c and q are considered consistent within the error range, we can consider person A as target object 1, and person C as target object 2, who is also making a sound. Finally, based on the area occupied by these two target objects, we can determine the first sub-region containing these two target objects from the target region.
[0068] It should be noted that in some embodiments of this application, when the number of target objects is one, the first sub-region is the target detection box region of that target object. This target detection box is determined by the imaging device and is included in the target region when the abnormal event occurs. For ease of understanding, please refer to [reference needed]. Figure 3 , Figure 3 This is a schematic diagram illustrating an example of an embodiment of this application. In this example, an abnormal event occurs between two people, and the area where the abnormal event occurs is as follows: Figure 3 Assuming that the person speaking at the current moment is the person on the right (i.e., the target object), the camera device determines the target area as follows: Figure 3 The target detection box shown is the first sub-region.
[0069] It should also be noted that in some other embodiments of this application, when there are multiple target objects, the first sub-region is the outer region (e.g., the minimum outer region) of these multiple target objects. This outer region can be obtained from the target detection boxes of these multiple target objects. For example, the outer region can be the union of the regions occupied by the target detection boxes of each of these multiple target objects. This outer region can also be obtained in other ways, and this application does not limit this. However, this outer region is also included in the target region when the abnormal event occurs. It should be noted that in the embodiments of this application, when all objects in the abnormal event are in a audible state, it means that all objects are target objects, and at this time, the target region is the first sub-region.
[0070] 204. Adjust the pickup parameters of the pickup device again so that the pickup range of the pickup device is limited to the first sub-region.
[0071] After determining the first sub-region from the target area, the pickup parameters of the pickup device can be adjusted again. Based on this, the beamforming algorithm parameters can be further adjusted to limit the pickup range of the microphone array of the pickup device to the first sub-region and block the sound outside the first sub-region, so that the pickup is more focused on the target object in the first sub-region.
[0072] It should be noted that in some embodiments of this application, when the target sound source (i.e., the sound source is human voice) stops emitting sound, the pickup device can be adjusted back to the coarsely adjusted pickup parameters. Assuming that in step 202, the pickup parameters of the pickup device are adjusted to a first value based on the first directional information (if there are multiple pickup parameters, the first value is the value of each pickup parameter), and in step 204, the pickup parameters are adjusted again to a second value (if there are multiple pickup parameters, the second value is the value of each pickup parameter), then when the target sound source stops emitting sound, the pickup parameters of the pickup device can be adjusted from the second value back to the first value.
[0073] In some other embodiments of this application, after the target sound source stops emitting sound, steps 201 and 202 can be re-executed so that the sound pickup device and the shooting device can readjust their respective parameters to focus on a new target area.
[0074] It should also be noted that, in this embodiment, if human voice is picked up again, steps 203 and 204 can be repeated. The purpose is to focus on the first sub-region for more accurate voice pickup when there is human voice in the target area; and to record audio within the target area when there is no human voice.
[0075] The method in this application embodiment can also track dynamic changes in the target area and adaptively adjust the shooting parameters of the shooting device and the sound pickup parameters of the sound pickup device. For example, steps 201 to 204 can be repeated periodically to achieve adaptive adjustment of the shooting and sound pickup parameters. Alternatively, steps 201 to 204 can be repeated when the shooting device detects a change in the position of the target area to achieve adaptive adjustment of the shooting and sound pickup parameters. This application does not limit the method of adaptively adjusting the shooting and sound pickup parameters. As an example, please refer to [reference needed]. Figure 4 , Figure 4 This is another schematic diagram illustrating an embodiment of the present application. It is assumed that both the shooting device and the sound pickup device are installed in the target device, which can be installed as follows: Figure 4 In the public scenario shown, when an abnormal event occurs, the target device can pick up ambient noise around the device, as well as the echo of the abnormal event. Figure 4It can be seen that when the target area becomes smaller, it means that the distance between the abnormal event and the shooting device is greater; when the target area becomes larger, it means that the distance between the abnormal event and the shooting device is closer. By tracking the dynamic changes of the target area, the shooting parameters of the shooting device and the sound pickup parameters of the sound pickup device can be adaptively adjusted.
[0076] It is important to note that when the target area changes, and the target object in the first sub-region emits sound, the sound pickup parameters of the sound pickup device are adjusted first in the fine-tuning process (step 6); when the target area changes, and the target object in the first sub-region stops emitting sound, the sound pickup parameters of the sound pickup device are adjusted first in the coarse-tuning process. However, regardless of the situation, the shooting parameters of the shooting device only need to undergo the coarse-tuning process in step 202.
[0077] It should also be noted that in some embodiments of this application, the enhanced audio and video data output can be saved to a disk for subsequent evidence collection; or the enhanced audio and video data can be input into the above-mentioned abnormal event detection algorithm to improve the detection effect of the shooting device on abnormal events and enhance the capabilities of the entire system architecture.
[0078] Based on the above embodiments, in order to better implement the above solutions of this application, related equipment for implementing the above solutions is also provided below. See details. Figure 5 , Figure 5 This is a schematic diagram of a computer device provided in an embodiment of this application. The computer device 500 may specifically include: a first determining module 501, a coarse adjustment module 502, a second determining module 503, and a fine adjustment module 504. The first determining module 501 is used to determine the first location information of the abnormal event actually occurring in physical space using the shooting device. The coarse adjustment module 502 is used to adjust the shooting parameters of the shooting device and, based on the first location information, adjust the pickup parameters of the sound pickup device so that the shooting device and the sound pickup device focus on the target area where the abnormal event occurs. The second determining module 503 is used to determine a first sub-region from the target area based on the sound source when there is a sound source in the target area. The first sub-region is the area where the target object in the abnormal event is located, and the target object is a person in a sound-emitting state. The fine adjustment module 504 is used to further adjust the pickup parameters of the sound pickup device so that the pickup range of the sound pickup device is limited to the first sub-region.
[0079] In one possible design, the second determining module 503 is specifically used to: determine the second directional information of the sound source by means of the sound pickup device when the sound source is determined to be a human voice; and determine the first sub-region from the target region based on the second directional information.
[0080] In one possible design, the second determining module 503 is further configured to: determine the third location information of the person in the abnormal event in the physical space through the shooting device; determine the target object based on the second location information and the third location information; and determine a first sub-region from the target region based on the target object.
[0081] In one possible design, the number of target objects is one, and the second determining module 503 is further used to: determine the target detection box of the target object, and use the target detection box as the first sub-region, the target detection box being contained in the target region.
[0082] In one possible design, the number of target objects is multiple, and the second determining module 503 is further used to: determine the outer area of the target object and use the outer area as the first sub-region, and the minimum outer area is contained in the target area.
[0083] In one possible design, the coarse adjustment module 502 is specifically used to adjust the pickup parameters of the pickup device to a first value based on the first directional information, and the fine adjustment module 504 is specifically used to adjust the pickup parameters of the pickup device to a second value again. The fine adjustment module 504 is specifically used to: after adjusting the pickup parameters of the pickup device again, when the sound source is a human voice and the sound source stops emitting sound, adjust the pickup parameters of the pickup device from the second value back to the first value, or trigger the first determination module 501 to re-execute the steps of determining the first directional information of the abnormal event actually occurring in the physical space through the shooting device, and trigger the coarse adjustment module 503 to adjust the shooting parameters of the shooting device, and adjust the pickup parameters of the pickup device according to the first directional information.
[0084] In one possible design, the coarse adjustment module 503 is specifically used to adjust the rotation angle of the shooting device so that the abnormal event is located in the center area of the imaging screen of the shooting device.
[0085] In one possible design, the coarse adjustment module 503 is specifically used to: adjust the focal length of the shooting device according to the ratio of the first number of pixels of the abnormal event in the target frame image to the total number of pixels of the target frame image, so that the proportion of the abnormal event in the imaging screen of the shooting device meets the preset requirements, and the target frame image is the frame image in which the abnormal event is detected.
[0086] In one possible design, the coarse adjustment module 503 is further configured to: determine the fourth orientation information of the abnormal event in a pre-established sound field coordinate system based on the first orientation information; and adjust the rotation angle of the pickup device based on the fourth orientation information so that the line connecting the first center of the abnormal event and the second center of the microphone array in the pickup device is perpendicular to the microphone array.
[0087] In one possible design, the first determining module 501 is specifically used to: determine a target frame image from a plurality of consecutive frame images captured by the capturing device, the target frame image being a frame image in which an abnormal event is detected; and determine the first location information of the abnormal event in physical space based on the pixel position of the abnormal event in the target frame image.
[0088] It should be noted that the information interaction and execution process between the modules / units in the computer device 500 are based on the same concept as the method embodiments described above in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0089] Next, we will introduce another computer device provided in the embodiments of this application. Please refer to [link / reference]. Figure 6 , Figure 6 This is a schematic diagram of a computer device provided in an embodiment of this application. The computer device 600 may be equipped with... Figure 5 The computer device 500 described in the corresponding embodiment is used to implement Figure 5 Corresponding to the functionality of computer device 500 in the corresponding embodiment, specifically, computer device 600 is implemented by one or more servers. Computer device 600 can vary significantly due to differences in configuration or performance, and may include one or more central processing units (CPUs) 622 and memory 632, and one or more storage media 630 (e.g., one or more mass storage devices) for storing application programs 642 or data 644. The memory 632 and storage media 630 can be temporary or persistent storage. The program stored in storage media 630 may include one or more modules (not shown in the figure), each module including a series of instruction operations on computer device 600. Furthermore, the CPU 622 may be configured to communicate with storage media 630 and execute the series of instruction operations in storage media 630 on computer device 600.
[0090] The computer device 600 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0091] In this embodiment, the central processing unit 622 is used to execute... Figure 2 The steps performed by the computer device in the corresponding embodiment are as follows. For example, the central processing unit 622 can be used to: firstly, determine the first location information of the abnormal event occurring in physical space through the shooting device, and guide the shooting device and the sound pickup device to adjust parameters in conjunction with the first location information to focus on the target area where the abnormal event occurs, with the purpose of focusing only to collect audio and video within the target area. When there is a sound source emitting sound in the target area, then determine a first sub-region (the area where the person in the abnormal event is in the sounding state) from the target area, and adjust the sound pickup parameters again to limit the sound pickup range to the first sub-region.
[0092] It should be noted that the specific manner in which the central processing unit 622 executes the above steps is different from that in this application. Figure 2 The corresponding method embodiments are based on the same concept, and the technical effects they bring are the same as those in the above embodiments of this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0093] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0095] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0096] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A data processing method, characterized in that, include: The imaging device is used to determine the primary location information of the actual occurrence of the abnormal event in the physical space. Adjust the shooting parameters of the shooting device and adjust the sound pickup parameters of the sound pickup device according to the first azimuth information so that the shooting device and the sound pickup device focus on the target area when the abnormal event occurs. If a sound source is present in the target area, a first sub-region is determined from the target area based on the sound source. The first sub-region is the area where the target object in the abnormal event is located, and the target object is a person in a sound-emitting state. The pickup parameters of the pickup device are adjusted again so that the pickup range of the pickup device is limited to the first sub-region.
2. The method according to claim 1, characterized in that, Determining the first sub-region from the target region based on the sound source includes: If the sound source is determined to be a human voice, the second directional information of the sound source is determined by the sound pickup device; The first sub-region is determined from the target region based on the second location information.
3. The method according to claim 2, characterized in that, Determining the first sub-region from the target region based on the second location information includes: The imaging device is used to determine the third-party location information of the person in the abnormal event in physical space. The target object is determined based on the second location information and the third location information; A first sub-region is determined from the target region based on the target object.
4. The method according to claim 3, characterized in that, The number of target objects is one, and the step of determining the first sub-region from the target region based on the target object includes: The target detection box of the target object is determined, and the target detection box is used as the first sub-region, wherein the target detection box is contained within the target region.
5. The method according to claim 3, characterized in that, The number of target objects is multiple, and the step of determining the first sub-region from the target region based on the target objects includes: The outer region of the target object is determined, and the outer region is used as the first sub-region, wherein the outer region is contained within the target region.
6. The method according to any one of claims 1-5, characterized in that, The method further includes adjusting the pickup parameters of the pickup device to a first value based on the first directional information, adjusting the pickup parameters of the pickup device to a second value, and after adjusting the pickup parameters of the pickup device again, the method further includes: When the sound source is a human voice and the sound source stops emitting sound, the pickup parameters of the pickup device are adjusted from the second value back to the first value, or the steps of determining the first location information of the abnormal event in the physical space through the shooting device, adjusting the shooting parameters of the shooting device, and adjusting the pickup parameters of the pickup device according to the first location information are re-executed.
7. The method according to any one of claims 1-6, characterized in that, The adjustment of the shooting parameters of the shooting device includes: Adjust the rotation angle of the shooting device so that the abnormal event is located in the center area of the image captured by the shooting device.
8. The method according to any one of claims 1-6, characterized in that, The adjustment of the shooting parameters of the shooting device includes: Based on the ratio of the number of pixels of the abnormal event in the target frame image to the total number of pixels in the target frame image, the focal length of the shooting device is adjusted so that the proportion of the abnormal event in the imaging screen of the shooting device meets a preset requirement, and the target frame image is the frame image in which the abnormal event is detected.
9. The method according to any one of claims 1-8, characterized in that, The step of adjusting the pickup parameters of the pickup device according to the first azimuth information includes: Based on the first azimuth information, determine the fourth azimuth information of the abnormal event in the pre-established sound field coordinate system; Based on the fourth position information, the rotation angle of the pickup device is adjusted so that the line connecting the first center of the abnormal event and the second center of the microphone array within the pickup device is perpendicular to the microphone array.
10. The method according to any one of claims 1-9, characterized in that, The first location information for determining the actual occurrence of the abnormal event in physical space through the imaging device includes: A target frame image is determined from a plurality of consecutive frame images captured by the imaging device, wherein the target frame image is a frame image in which an abnormal event is detected; The first location information of the actual occurrence of the abnormal event in physical space is determined based on the pixel position of the abnormal event in the target frame image.
11. A computer device, characterized in that, include: The first determining module is used to determine the first location information of the actual occurrence of the abnormal event in physical space through the imaging device; The coarse adjustment module is used to adjust the shooting parameters of the shooting device and adjust the pickup parameters of the pickup device according to the first azimuth information, so that the shooting device and the pickup device focus on the target area when the abnormal event occurs. The second determining module is used to determine a first sub-region from the target region based on the sound source when there is a sound source in the target region. The first sub-region is the region where the target object in the abnormal event is located, and the target object is a person in a sound-emitting state. The fine-tuning module is used to readjust the pickup parameters of the pickup device so that the pickup range of the pickup device is limited to the first sub-region.
12. The device according to claim 11, characterized in that, The second determining module is specifically used for: If the sound source is determined to be a human voice, the second directional information of the sound source is determined by the sound pickup device; The first sub-region is determined from the target region based on the second location information.
13. The device according to claim 12, characterized in that, The second determining module is further used for: The imaging device is used to determine the third-party location information of the person in the abnormal event in physical space. The target object is determined based on the second location information and the third location information; A first sub-region is determined from the target region based on the target object.
14. The device according to claim 13, characterized in that, The number of target objects is one, and the second determining module is further configured to: The target detection box of the target object is determined, and the target detection box is used as the first sub-region, wherein the target detection box is contained within the target region.
15. The device according to claim 13, characterized in that, The number of target objects is multiple, and the second determining module is further used for: The outer region of the target object is determined, and the outer region is used as the first sub-region, wherein the outer region is contained within the target region.
16. The device according to any one of claims 11-15, characterized in that, The coarse adjustment module is specifically used to adjust the pickup parameters of the pickup device to a first value based on the first azimuth information. The fine adjustment module is specifically used to further adjust the pickup parameters of the pickup device to a second value. The fine adjustment module is specifically used for: After readjusting the pickup parameters of the pickup device, if the sound source is a human voice and the sound source stops emitting sound, the pickup parameters of the pickup device are adjusted from the second value back to the first value, or the first determining module is triggered to re-execute the steps of determining the first location information of the abnormal event in the physical space through the shooting device, and triggering the coarse adjustment module to adjust the shooting parameters of the shooting device, and adjusting the pickup parameters of the pickup device according to the first location information.
17. The device according to any one of claims 11-16, characterized in that, The coarse adjustment module is specifically used for: Adjust the rotation angle of the shooting device so that the abnormal event is located in the center area of the image captured by the shooting device.
18. The device according to any one of claims 11-16, characterized in that, The coarse adjustment module is specifically used for: Based on the ratio of the number of pixels of the abnormal event in the target frame image to the total number of pixels in the target frame image, the focal length of the shooting device is adjusted so that the proportion of the abnormal event in the imaging screen of the shooting device meets a preset requirement, and the target frame image is the frame image in which the abnormal event is detected.
19. The device according to any one of claims 11-18, characterized in that, The coarse adjustment module is further used for: Based on the first azimuth information, determine the fourth azimuth information of the abnormal event in the pre-established sound field coordinate system; Based on the fourth position information, the rotation angle of the pickup device is adjusted so that the line connecting the first center of the abnormal event and the second center of the microphone array within the pickup device is perpendicular to the microphone array.
20. The device according to any one of claims 11-19, characterized in that, The first determining module is specifically used for: A target frame image is determined from a plurality of consecutive frame images captured by the imaging device, wherein the target frame image is a frame image in which an abnormal event is detected; The first location information of the actual occurrence of the abnormal event in physical space is determined based on the pixel position of the abnormal event in the target frame image.
21. A computer device comprising a processor and a memory, the processor being coupled to the memory, characterized in that, The memory is used to store programs; The processor is configured to execute a program in the memory, causing the computer device to perform the method as described in any one of claims 1-10.
22. A computer storage medium, characterized in that, The device stores computer-readable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-10.
23. A computer program product, characterized in that, The computer program product includes computer-readable instructions that, when executed by a processor, implement the method as described in any one of claims 1-10.
24. A chip, the chip comprising a processor and a data interface, characterized in that, The processor reads instructions stored in the memory through the data interface and executes the method as described in any one of claims 1-10.