Sound source positioning method and device, equipment and storage medium
By detecting the target sound source and positioning it in the camera's camera component, the problem that the webcam cannot locate when the object to be located is not within the field of view is solved, and positioning capability and accuracy are improved.
Patent Information
- Application Number
- CN202311587212.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing webcams cannot trigger positioning tracking when the object to be located is not within the camera's field of view, resulting in limited positioning capability and accuracy.
By detecting the target sound source in each camera component, acquiring audio evaluation data, determining the target image component, and controlling it to locate the target sound source, so that the target sound source is within the field of view of the camera of the target image component.
The triggering conditions of the camera component are expanded, positioning capabilities and positioning accuracy are improved, and the object to be positioned can be positioned within the field of view of the camera.
Smart Images

Figure CN120050524A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, device, equipment and storage medium for sound source localization. Background Art
[0002] An Internet Protocol Camera (IPC) is a digital camera that uses a network to transmit data and communicate. Currently, the recognition and localization of objects by an IPC mainly rely on the camera itself for localization. How to improve the localization accuracy of the IPC for objects has become a hot topic in this field. Summary of the Invention
[0003] The present disclosure provides a method, device, equipment and storage medium for sound source localization to solve or alleviate one or more technical problems in the prior art.
[0004] In a first aspect, the present disclosure provides a method for sound source localization, including:
[0005] When at least one of the camera components detects a target sound source, obtaining audio evaluation data of each camera component, where the camera components in each camera component include a camera, a first microphone and a second microphone located on both sides of the camera in the horizontal direction, the orientation of the first microphone is opposite to that of the second microphone, and the orientation of the first microphone is perpendicular to that of the camera, and the audio evaluation data of the camera component is obtained based on the audio data of the target sound source detected by the first microphone and the audio data of the target sound source detected by the second microphone;
[0006] Determining a target camera component from each camera component based on the audio evaluation data of each camera component;
[0007] Controlling the target camera component to localize the target sound source so that the target sound source is within the field of view of the camera of the target camera component.
[0008] In a second aspect, the present disclosure provides a sound source localization device, including:
[0009] An obtaining unit, configured to obtain audio evaluation data of each camera component when at least one of the camera components detects a target sound source, where the camera components in each camera component include a camera, a first microphone and a second microphone located on both sides of the camera in the horizontal direction, the orientation of the first microphone is opposite to that of the second microphone, and the orientation of the first microphone is perpendicular to that of the camera, and the audio evaluation data of the camera component is obtained based on the audio data of the target sound source detected by the first microphone and the audio data of the target sound source detected by the second microphone;
[0010] A determination unit, configured to determine a target imaging component from each imaging component based on the audio evaluation data of each imaging component;
[0011] A positioning unit, configured to control the target imaging component to locate a target sound source so that the target sound source is within the field of view of the camera of the target imaging component.
[0012] In a third aspect, an electronic device is provided, including:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method in the embodiments of the present disclosure.
[0016] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0017] According to the sound source localization method, device, equipment and storage medium provided by the embodiments of the present disclosure, it is possible to trigger the localization of the object to be localized by detecting the target sound source, so that the object to be localized can be within the field of view of the camera, thereby expanding the trigger condition of the imaging component and improving the localization ability and localization accuracy of the imaging component.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided by the present disclosure and should not be regarded as limiting the scope of the present disclosure.
[0020] Figure 1 is a structural diagram of a system applying the sound source localization method provided by an embodiment of the present disclosure;
[0021] Figure 2 is a flowchart of the sound source localization method provided by an embodiment of the present disclosure;
[0022] Figure 3 is a schematic diagram of a scenario of the sound source localization method provided by an embodiment of the present disclosure;
[0023] Figure 4 is Figure 3 a schematic structural diagram of the camera assembly in
[0024] Figure 5 a flowchart of a sound source localization method provided by an embodiment of the present disclosure;
[0025] Figure 6 a schematic block diagram of a sound source localization device provided by an embodiment of the present disclosure;
[0026] Figure 7 a block diagram of an electronic device for implementing the sound source localization method of the embodiment of the present disclosure. Specific Embodiments
[0027] The present disclosure will be described in further detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0028] The disclosed embodiments provide a sound source localization method, apparatus, electronic device, and storage medium. Specifically, the sound source localization method of the embodiments of the present disclosure can be executed by an electronic device, where the electronic device can be a device such as a terminal or a server. The terminal can be a device such as a smart phone, a tablet computer, a notebook computer, a smart voice interaction device, a smart home appliance, a wearable smart device, an aircraft, a smart vehicle terminal, etc. The terminal can also include a client, and the client can be an audio client, a video client, a browser client, an instant messaging client, or a small program, etc. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0029] In the related art, a network camera can be used for the recognition and tracking positioning of an object, and the object can be a person, an animal, or a moving object, etc. For example, a user can locate the position of a pet, a baby, or an elderly person at home through the network camera to timely discover whether they are in a dangerous state. Or, the user can also conduct live broadcasts through the network camera, and the network camera can recognize and locate the position of the user so that the user can always be in the picture.
[0030] However, the positioning and tracking of webcams usually rely on the shooting range of the camera. When the object to be positioned is within its shooting range, the camera can position and track the object to be positioned. When the camera fails to capture the object to be positioned, for example, when the object to be positioned is behind the camera, the positioning of the object cannot be triggered, that is, the positioning of the webcam is limited by the position of the object to be positioned, and the positioning ability is poor.
[0031] To solve at least one of the above problems, an embodiment of the present disclosure provides a sound source positioning method, including: when at least one of the camera components detects a target sound source, obtaining audio evaluation data of each camera component, where the camera components in each camera component include a camera, a first microphone and a second microphone located on both sides of the camera in the horizontal direction, the orientation of the first microphone is opposite to that of the second microphone, and the orientation of the first microphone is perpendicular to the orientation of the camera, and the audio evaluation data of the camera component is obtained based on the audio data of the target sound source detected by the first microphone and the audio data of the target sound source detected by the second microphone; determining a target camera component from each camera component based on the audio evaluation data of each camera component; controlling the target camera component to position the target sound source so that the target sound source is within the field of view of the camera of the target camera component. The embodiment of the present disclosure can trigger the positioning of the object to be positioned by detecting the target sound source, so that the object to be positioned can be within the field of view of the camera, thereby expanding the trigger condition of the camera component and improving the positioning ability and positioning accuracy of the camera component.
[0032] The solution of the present disclosure will be described below with reference to the accompanying drawings.
[0033] Figure 1 is a structural diagram of a system applying the sound source positioning method provided by an embodiment of the present disclosure; please refer to Figure 1 , the system includes a terminal 110 and a server 120, etc.; the terminal 110 and the server 120 are connected through a network, for example, through a wired or wireless network connection, etc.
[0034] Among them, the terminal 110 can be used to display a graphical user interface. The terminal is used to interact with the user through the graphical user interface. For example, the corresponding client is downloaded and installed through the terminal and run, for example, by calling the corresponding applet and running, for example, by logging in to a website to present the corresponding graphical user interface, etc. In the embodiments of the present disclosure, the server 120 is used to obtain the audio evaluation data of each camera component when at least one of the camera components detects a target sound source. Among them, the camera components in each camera component include a camera, a first microphone and a second microphone located on both sides of the camera in the horizontal direction. The orientation of the first microphone is opposite to that of the second microphone, and the orientation of the first microphone is perpendicular to the orientation of the camera. The audio evaluation data of the camera component is obtained based on the audio data of the target sound source detected by the first microphone and the audio data of the target sound source detected by the second microphone; based on the audio evaluation data of each camera component, a target camera component is determined from each camera component; the target camera component is controlled to locate the target sound source so that the target sound source is within the field of view of the camera of the target camera component. In addition, the terminal 110 can receive and display the picture within the field of view of the camera transmitted by the server.
[0035] It should be noted that although the display interface is taken as an example of the page of the application program, the display interface can also be other pages such as a web page. In addition, the application program can be an application program installed on a desktop computer, an application program installed on a mobile terminal, or a small program embedded in the application program, etc.
[0036] It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0037] The following will be described in detail. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.
[0038] Figure 2 is a flowchart of a sound source localization method provided by an embodiment of the present disclosure; Figure 3 is a scenario diagram of a sound source localization method provided by an embodiment of the present disclosure; Figure 4 is Figure 3 a schematic structural diagram of the camera component in Figures 2 to 4 Referring to
[0039] Step S201, when at least one of the camera components detects the target sound source 310, obtain the audio evaluation data of each camera component. Among them, the camera components in each camera component include a camera 321, a first microphone 322 and a second microphone 323 located on both sides of the camera 321 in the horizontal direction. The orientation of the first microphone 322 is opposite to that of the second microphone 323, and the orientation of the first microphone 322 is perpendicular to the orientation of the camera 321. The audio evaluation data of the camera component is obtained based on the audio data of the target sound source detected by the first microphone 322 and the audio data of the target sound source detected by the second microphone 323;
[0040] Step S202, based on the audio evaluation data of each camera component, determine the target camera component from each camera component;
[0041] Step S203, control the target camera component to locate the target sound source so that the target sound source 310 is within the field of view of the camera of the target camera component.
[0042] In this embodiment, the method 200 can be used in a family or a relatively large area, and multiple camera components can be set in this area. Each camera component is a camera component installed in the same environment to be located. For example Figure 3 In a house area, there can be three rooms, and each room can be provided with a camera component, namely camera component 320a, camera component 320b and camera component 320c.
[0043] As Figure 4 shown, each camera component 320 can include a camera 321 and two microphones (the first microphone 322 and the second microphone 323). The camera 321 can be a network camera, and the microphone can be a device capable of receiving sound, such as a microphone. Figure 4 The plane shown can be a horizontal plane. The first microphone 322 can be set horizontally to the left, the second microphone 323 can be set horizontally to the right, and the camera 321 is horizontally forward ( Figure 4 in the upward direction in Figure 4 ), so that the orientation of the first microphone is opposite to that of the second microphone, and the orientation of the first microphone is perpendicular to the orientation of the camera. It can be understood that the orientation of the camera can be the optical axis direction of the camera. In addition, the camera component 320 can rotate horizontally along the arc arrow L direction in
[0044] The target sound source can be an object to be located. For example, the sound to be detected and recognized can be processed in advance to obtain the waveform corresponding to the sound. By setting a preset sound waveform (such as a crying sound, a sound of falling, etc.), when a sound with a certain preset waveform is detected, it can be determined that the target sound source is detected.
[0045] In some embodiments, the target sound source is a sound source with a sound waveform identical to a preset sound waveform; the preset sound waveform may be the waveform of a sound that needs to be focused on, such as a crying sound, a sound of falling, etc. The Mel (Mel) spectral features of these sounds can be extracted in advance through MFCC (Mel Frequency Cepstrum Coefficient) or PLP (Perceptual Linear Predictive) technology to convert the timbre of these sounds into a sound waveform and obtain the preset sound waveform. By comparing the sound waveform of the sound detected by the radio with the preset sound waveform, if they are the same, the camera component is triggered to locate the target sound source.
[0046] And / or, the target sound source is a sound source with a sound intensity greater than or equal to a preset intensity threshold. The preset intensity threshold can also be specified, such as 100 decibels, etc. When the sound intensity of the target sound source is greater than the preset intensity threshold, the location of the target sound source can be triggered.
[0047] Of course, the triggering condition can also include that the sound source is the sound source of the target object. For example, by learning the timbre of the target object to be located, when the sound intensity is greater than or equal to the preset intensity threshold and it is the sound of the target object, it is determined that the target sound source is detected. Or, when the sound waveform is the same as the preset sound waveform and it is the sound of the target object, it is determined that the target sound source is detected.
[0048] By setting the triggering condition, it is not necessary for all camera components to be in a cruising state all the time. Only the states that need to be focused on are located and tracked, saving resources.
[0049] In step S201, when one or more of the camera components detect the target sound source, the audio evaluation data of the target sound source detected by all camera components can be obtained. It can be understood that for a certain camera component, its audio evaluation data can be obtained based on the audio data collected by the first radio and the second radio therein respectively.
[0050] Step S202, according to the audio evaluation data of each camera component, a target camera component can be determined from each camera component. The target camera component can be the one closest to the target sound source or the one that is most likely to capture the target sound source.
[0051] Step S203, the camera component can be controlled to rotate so as to locate the position of the target sound source, so that the target sound source can be within the camera field of view of the target camera component.
[0052] It can be understood that for the method provided in this embodiment, when the object to be located is not within the field of view of the camera component, the sound of the object to be located can be detected in real time by multiple camera components. If a target sound source is detected, audio evaluation data for the target sound source is obtained through the multiple camera components (during the time period when the target sound source is detected, the audio data detected by the microphones of each camera component during this time period is obtained, and based on this, the audio evaluation data of each camera component is obtained), and based on the audio evaluation data, a target camera component for locating the object to be located is determined, and then the target camera component is controlled to rotate, so that the object to be located can be within the range of the camera. Then, the object to be located can be located and tracked through the tracking algorithm of the camera.
[0053] This embodiment overcomes the situation in the related art where it is impossible to achieve positioning and tracking if the object to be located is not within the field of view of the camera. It can trigger the positioning of the object to be located by detecting the target sound source, so that the object to be located can be within the field of view of the camera, thereby expanding the triggering conditions of the camera component and improving the positioning ability and positioning accuracy of the camera component.
[0054] In some embodiments, obtaining the audio evaluation data of each camera component in step S201 includes:
[0055] For the first camera component among each camera component, obtain the first audio data detected by the first microphone in the first camera component and the second audio data detected by the second receiver in the first camera component;
[0056] Determine the average audio data of the first audio data and the second audio data, and use the average audio data as the audio evaluation data of the first camera component;
[0057] Obtain the audio evaluation data of each camera component based at least on the audio evaluation data of the first camera component.
[0058] The audio evaluation data of each camera component can be obtained by traversing each camera component, and the method for obtaining the audio evaluation data of each camera component is the same. Taking the first camera component among each camera component as an example, the first audio data detected by the first microphone and the second audio data detected by the second microphone can be obtained. Among them, the audio data can be audio intensity data, such as the magnitude of decibels. If the first audio data is Aa and the second audio data is Ab, the average audio data (audio evaluation data) of the first camera component can be (Aa + Ab) / 2. Similarly, the audio evaluation data of all camera components can be obtained according to this method.
[0059] By calculating the average value of the first audio data and the second audio data in each camera component, audio evaluation data can be obtained, which can comprehensively evaluate the audio data of each camera component, and the calculation method is simple and fast.
[0060] In some embodiments, step S202, determining a target camera component from each camera component based on the audio evaluation data of each camera component may include:
[0061] Taking the camera component corresponding to the target audio evaluation data with the largest value among the audio evaluation data of each camera component as the target camera component.
[0062] As Figure 3 shown, if the first audio data of camera component 320a is Aa, the second audio data is Ab, and its audio evaluation data is (Aa + Ab) / 2, the first audio data of camera component 320b is Ba, the second audio data is Bb, and its audio evaluation data is (Ba + Bb) / 2, the first audio data of camera component 320c is Ca, the second audio data is Cb, and its audio evaluation data is (Ca + Cb) / 2. Assuming (Ca + Cb) / 2 > (Ba + Bb) / 2 > (Aa + Ab) / 2, then the camera component 320c corresponding to the audio evaluation data (Ca + Cb) / 2 can be taken as the target camera component.
[0063] It can be understood that the largest value of the audio evaluation data can be considered that the target camera component is closest to the target sound source, it is easiest to clearly locate the target sound source, and the positioning effect is the best.
[0064] In some embodiments, step S202, determining a target camera component from each camera component based on the audio evaluation data of each camera component further includes:
[0065] Determining a first position of the target sound source based on the third audio data detected by the first microphone in the target camera component and the fourth audio data detected by the second microphone in the target camera component;
[0066] In the case where the rotation angle of the target camera component rotating to the first position does not meet the preset angle condition of the target camera component, arranging the audio evaluation data of each camera component in descending order of value;
[0067] Taking the camera component ranked next to the target camera component among each camera component as the new target camera component;
[0068] Step S203, controlling the target camera component to locate the target sound source includes: controlling the new target camera component to locate the target sound source.
[0069] It can be understood that when each camera component is installed in a home environment such as a residence, since each room is isolated by walls, in order to achieve a more comprehensive positioning, for example, to locate a solitary elderly person to prevent accidents from occurring without being discovered in time, there can be at least one camera component in each room.
[0070] After the camera component is installed, for each camera component, the camera component can be first controlled to rotate horizontally (in a cruising state), and the image within the field of view of the camera in the camera component is obtained. In the case where the image is a large-area color block, the rotation angle ∠θ of the camera component is recorded (the angle rotated from the initially calibrated 0° position to the large-area color block). The original rotation angle of the camera component can be freely rotated by 360°, then the effective rotation angle of the camera component is ∠α = 360° - ∠θ. This effective rotation angle can be the preset angle condition corresponding to the camera component. It can be understood that for camera components at different positions, their effective rotation angles can be different and can be specifically set according to the actual situation.
[0071] When obtaining the target camera component relying on the maximum audio evaluation data, the first position of the target sound source can be roughly determined first through the magnitudes of the third audio data and the fourth audio data and the distance between the first sound receiver and the second sound receiver. This first position can be determined by the difference between the third audio data and the fourth audio data and the distance between the first sound receiver and the second sound receiver. It can be understood that the first position is an approximate position and the positioning may not be accurate. In addition, the position of the target sound source may change, so the first position can be used as a rough reference.
[0072] Next, the rotation angle of the camera component rotated from the initially calibrated 0° position to the camera facing the first position can be determined through the first position, that is, the rotation angle of the target camera component rotated to the first position. Then, this rotation angle is compared with the preset angle condition of the target camera component. If this rotation angle is not within the range of the preset angle condition (the rotation angle is greater than ∠α), it means that the straight-line direction between the target camera component and the target sound source will pass through the wall and the target camera component cannot capture the target object. For example, as Figure 3 shown, assuming (Aa + Ab) / 2 > (Ca + Cb) / 2 > (Ba + Bb) / 2, then the camera component 320a is first confirmed as the target camera component. However, due to the wall obstruction, it cannot directly capture the target sound source 310. At this time, the camera component 320c with the second-largest audio evaluation data can be used as the new target camera component. Then, the target sound source 310 is located through the new target camera component.
[0073] When the method provided in this embodiment is applied to a scenario of multiple rooms, it can accurately locate the camera components that effectively capture, quickly capture the position of the sound source, avoid all camera components from starting cruise capture, and save resources.
[0074] In some embodiments, step S203, controlling the target camera component to locate the target sound source, may include:
[0075] Based on the third audio data detected by the first microphone in the target camera component and the fourth audio data detected by the second microphone in the target camera component, determine the initial horizontal rotation direction of the target camera component, where the initial horizontal rotation direction is the orientation of the microphone corresponding to the audio data with a larger value among the third audio data and the fourth audio data;
[0076] Control the target camera component to rotate horizontally in the initial horizontal rotation direction;
[0077] Obtain the difference between the third audio data and the fourth audio data in real time;
[0078] Based on the difference, control the rotation angle of the target camera component.
[0079] After determining the target camera component, according to the third audio data detected by the first microphone and the fourth audio data detected by the second microphone in the target camera component, a larger-valued audio data can be determined therefrom.
[0080] Assume that the third audio data is greater than the fourth audio data. The orientation of the first microphone can be used as the initial horizontal rotation direction of the target camera component. It can be understood that when the orientation of the first microphone is horizontally to the right, the horizontal right direction can be used as the initial horizontal rotation direction of the target camera component. Additionally, if the third audio data is less than the fourth audio data, the orientation of the second microphone can be used as the initial horizontal rotation direction of the target camera component.
[0081] Next, control the target camera component to rotate along the initial horizontal rotation direction, and simultaneously calculate the difference between the third audio data and the fourth audio data, because as the target camera component rotates, the third audio data and the fourth audio data will change.
[0082] Based on the difference, control the rotation angle of the target camera component, that is, control the target camera component to rotate towards the direction where the difference tends to 0 until the target sound source is within the camera's field of view of the target camera component, thereby completing the positioning process relying on sound, that is, the target sound source can also be located when it is outside the camera's field of view.
[0083] In some embodiments, based on the difference, controlling the rotation angle of the target camera component may include: when the difference is zero, controlling the target camera component to stop rotating.
[0084] When the difference becomes 0, it can be understood that the target sound source is equidistant from the first microphone and the second microphone, and the target camera component can be controlled to stop horizontal rotation. That is, the target sound source is located within the field of view of the camera, thus completing the process of positioning relying on sound.
[0085] In some embodiments, controlling the rotation angle of the target camera component based on the difference may include: when the difference is zero and the target sound source does not appear within the field of view of the camera of the target camera component, controlling the camera component to rotate longitudinally until the target sound source is within the field of view of the camera of the target camera component, and the plane where the longitudinal direction is located is perpendicular to the horizontal plane.
[0086] It can be understood that the camera component can be installed on a pan-tilt head, and the pan-tilt head can control the camera component to rotate horizontally. Of course, it can also control the camera component to rotate longitudinally. The longitudinal direction can be perpendicular to the horizontal direction. For example, it can be Figure 4 the direction perpendicular to the paper surface in []. If the horizontal direction is left-right rotation, then the longitudinal direction can be up-down pitching.
[0087] When the difference is zero, but the target sound source still does not appear within the test range of the camera of the target camera component, it can be considered that the position of the target sound source is relatively high or low. Controlling the target camera component to rotate longitudinally can use the camera in the target camera component to locate the target sound source, so that it is within the field of view.
[0088] This embodiment can perform initial positioning through microphones and finally rely on the camera for precise positioning, effectively realizing the positioning of the target object.
[0089] In some embodiments, controlling the rotation angle of the target camera component based on the difference may include: when the difference is not zero and the target sound source appears within the field of view of the camera of the target camera component, obtaining the field of view image of the camera in the target camera component; based on the field of view image, controlling the rotation angle of the target camera component.
[0090] In addition, when the target camera component rotates along the initial horizontal rotation direction, the camera can obtain the images within the field of view in real time. Even if the difference is not 0 and the target sound source appears within the field of view of the camera, the positioning using the sound source can be stopped and the positioning can be switched to rely on the image recognition processing of the camera. Since the positioning accuracy of the camera is relatively high, after the target object appears within the field of view, the camera can be directly used for positioning detection, and the positioning effect is better.
[0091] In some embodiments, method 200 further includes: when the target sound source appears within the field of view of the camera of the target camera component, using the camera of the target camera component to locate the target sound source.
[0092] It can be understood that when controlling the target camera component to locate the target sound source, the rotation of the target camera component can be controlled first by relying on the third audio data and the fourth audio data. During the rotation process, if the target sound source (such as a human figure) appears in the picture, the calculation of the difference between the third audio data and the fourth audio data can be stopped, and instead, the picture within the field of view of the camera can be used to identify the picture through image recognition to achieve the positioning of the target sound source. For example, it can follow the movement of the target sound source so that it is always located in the center of the picture, thereby improving the positioning accuracy and effect.
[0093] Of course, in other embodiments, if the target sound source is interrupted and the interruption time is within a preset time range, such as more than 2 s, it can be considered that the target sound source is lost. At this time, the target camera component can be controlled to rotate to the initial position.
[0094] Figure 5 is a flowchart of a sound source localization method provided by an embodiment of the present disclosure; please refer to Figure 5 In one embodiment, a sound source localization method is provided, which can be used in devices such as security, etc., and supplements the related technology that only locates and tracks when the picture of the camera changes. That is, in the case where the picture does not change, the approximate azimuth of the object can be identified by relying on the sound source decibel difference detected by the dual microphones (mic, microphone), and then the camera component is rotated for sound tracking.
[0095] The specific implementation solution can be to set multiple camera components, each camera component includes a camera and a dual mic (a first microphone and a second microphone). Two mics are installed horizontally on the camera, and the two mics form a 90° angle with the lens direction (the orientation of the camera), as Figure 4 shown.
[0096] Each camera component can share its own mic (microphone) and upload the sound source (sound source) to the cloud for processing. In a user's home, the mic gain (audio data) of each camera component is judged to determine the sound source azimuth, and a rotation instruction is sent to a specific camera component (the target camera component) for sound tracking. For example Figure 3In this case, a camera component 320a, a camera component 320b, and a camera component 320c are respectively provided in the bedroom, the study, and the living room. When a specific sound source is emitted in the bedroom (the target sound source is detected), the camera components 320a - 320c can simultaneously detect the sound source (the target sound source). The camera components 320a - 320c respectively upload the audio data of the sound source detected within the same time period to the cloud server. If the server recognizes that the sound source is similar to the learned sound wave and is the same sound wave (the preset sound waveform), it preferentially retrieves the decibels (audio data) picked up by each camera component, determines the position of the camera component with the highest decibels, and calls this camera component for sound source tracking. Considering that each room in the home environment is isolated by walls, if omnidirectional positioning detection is required, at least 1 camera component is placed in each room. The camera in each camera component can perform cruising. During the continuous rotation process, if a large - area color block is recognized in the image as a wall, the rotation angle ∠θ, and the effective rotation angle of the camera (camera component) is ∠α = 360° - ∠θ.
[0097] It can be understood that each camera component has 2 oppositely - placed mics. The audio gains (audio data) picked up by these two mics (a, b) are different, and the resulting gain difference is the object's bias. For example, if the gain of a is greater than that of b, the sound source is closer to a.
[0098] In the case where multiple camera components are set, each camera component needs to upload the audio gains picked up by its two oppositely - placed mics to the background server. If the audio gains picked up by the dual mics of the camera component 320a are Aa and Ab respectively, the audio gains picked up by the dual mics of the camera component 320b are Ba and Bb respectively, and the audio gains picked up by the dual mics of the camera component 320c are Ca and Cb respectively, if (Aa + Ab) / 2 > (Ba + Bb) / 2 > (Ca + Cb) / 2, then the approximate rotation angle of the camera component 320a is calculated preferentially. If the rotation angle is greater than ∠α, it is determined that the camera component 320a does not start sound source tracking, and the camera component 320b is started to calculate the approximate rotation angle.
[0099] This function can accurately locate the effective capture camera component, quickly capture the sound source position, and avoid all camera components starting cruising to capture.
[0100] In addition, as Figure 4 , for each camera component in the camera components, it can be understood that it is achieved through the following steps:
[0101] 1. Install two mics in the horizontal direction of the camera; the two mics are at a 90° angle to the lens direction.
[0102] 2. The camera component can learn a specific sound source (preset sound waveform), learn the timbre of the specific sound source, and convert it into a sound waveform; this can be achieved by using MFCC or PLP technology to extract Mel spectrum features.
[0103] 3. When it is recognized that the sound source (sound source) is similar to the learned audio waveform (preset sound waveform), the camera component rotation step is triggered. Specifically, refer to Figure 5 , for the two mics (the first microphone and the second microphone), the audio sampling data of the sound source can be collected respectively, the respective sound waveforms can be obtained through audio decoding, and then the sound waveforms are compared with the preset sound waveform. If they are the same, the decibels (audio data) obtained by copying the two audio decoding data are uploaded to the cloud server. If they are different, the sound source continues to be monitored and detected.
[0104] 4. The cloud server determines the target camera component according to the decibels uploaded by each camera component, and then determines the approximate direction of the sound source (initial horizontal rotation direction) according to the decibels picked up by the two mics in the target camera component; the approximate direction is the side with the larger decibel.
[0105] 5. The system issues a rotation command to the pan-tilt motor, so that the target camera component turns to the approximate direction of the sound source.
[0106] 6. During the rotation process, if the sound continues, the sound source coordinates are continuously updated according to the decibel difference (difference) between the two mics; during the turning process, the decibel difference will gradually approach 0. When the decibel difference is equal to 0, it means that the horizontal (lateral) rotation is in place.
[0107] 7. When the sound is interrupted, generally set to more than 2 seconds, then continue to rotate to the home position according to the previous rotation direction, that is, return to the initial position.
[0108] In some embodiments, during the rotation process, when the rotation reaches the decibel difference approaching 0, but no human form is detected within the field of view of the camera, the camera is triggered to rotate longitudinally.
[0109] The method provided in this embodiment solves the deficiency in the related art that relies on the camera to implement object tracking through the sound tracking algorithm, that is, the problem that the tracking and positioning cannot be triggered when the field of view of the camera does not change.
[0110] Figure 6 It is a schematic block diagram of a sound source localization device provided in another embodiment of the present disclosure. Please refer to Figure 6 , the present disclosure provides a sound source localization device 600, including the following units.
[0111] An acquisition unit 601, configured to acquire audio evaluation data of each camera component when at least one camera component among the camera components detects a target sound source. The camera components among the camera components include a camera, a first microphone, and a second microphone located on both sides of the camera in the horizontal direction. The orientation of the first microphone is opposite to that of the second microphone, and the orientation of the first microphone is perpendicular to that of the camera. The audio evaluation data of the camera component is obtained based on the audio data of the target sound source detected by the first microphone and the audio data of the target sound source detected by the second microphone.
[0112] A determination unit 602, configured to determine a target camera component from each camera component based on the audio evaluation data of each camera component.
[0113] A positioning unit 603, configured to control the target camera component to locate the target sound source so that the target sound source is within the field of view of the camera of the target camera component.
[0114] In some embodiments, the acquisition unit 601 is further configured to:
[0115] For a first camera component among each camera component, acquire first audio data detected by the first microphone in the first camera component and second audio data detected by the second receiver in the first camera component;
[0116] Determine the average audio data of the first audio data and the second audio data, and use the average audio data as the audio evaluation data of the first camera component;
[0117] Obtain the audio evaluation data of each camera component based at least on the audio evaluation data of the first camera component.
[0118] In some embodiments, the determination unit 602 is further configured to:
[0119] Use the camera component corresponding to the target audio evaluation data with the largest value among the audio evaluation data of each camera component as the target camera component.
[0120] In some embodiments, the determination unit 602 is further configured to:
[0121] Based on the third audio data detected by the first microphone in the target camera component and the fourth audio data detected by the second microphone in the target camera component, determine the first position of the target sound source;
[0122] In a case where the rotation angle of the target camera component rotating to the first position does not meet the preset angle condition of the target camera component, arrange the audio evaluation data of each camera component in descending order of value;
[0123] Use the camera component whose arrangement order in each camera component is the next one after the target camera component as the new target camera component;
[0124] The positioning unit 603 is further configured to:
[0125] Control the new target camera component to locate the target sound source.
[0126] In some embodiments, the target sound source is a sound source whose sound waveform is the same as the preset sound waveform; and / or, the target sound source is a sound source whose sound intensity is greater than or equal to the preset intensity threshold.
[0127] In some embodiments, the positioning unit 603 is further configured to:
[0128] Based on the third audio data detected by the first microphone in the target camera component and the fourth audio data detected by the second microphone in the target camera component, determine the initial horizontal rotation direction of the target camera component, where the initial horizontal rotation direction is the orientation of the microphone corresponding to the audio data with a larger value in the third audio data and the fourth audio data;
[0129] Control the target camera component to horizontally rotate in the initial horizontal rotation direction;
[0130] Real-time obtain the difference between the third audio data and the fourth audio data;
[0131] Based on the difference, control the rotation angle of the target camera component.
[0132] In some embodiments, the positioning unit 603 is further configured to: when the difference is zero, control the target camera component to stop rotating.
[0133] In some embodiments, the positioning unit 603 is further configured to: when the difference is zero and the target sound source does not appear within the field of view of the camera of the target camera component, control the camera component to rotate longitudinally until the target sound source is within the field of view of the camera of the target camera component, and the plane where the longitudinal direction is located is perpendicular to the horizontal plane.
[0134] In some embodiments, the positioning unit 603 is further configured to: when the difference is not zero and the target sound source appears within the field of view of the camera of the target camera component, obtain the field of view image of the camera in the target camera component; based on the field of view image, control the rotation angle of the target camera component.
[0135] In some embodiments, the device 600 further includes: a shooting unit, configured to, when the target sound source appears within the field of view of the camera of the target camera component, use the camera of the target camera component to locate the target sound source.
[0136] For the specific functions and examples of each module and sub-module of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the foregoing method embodiments, which will not be elaborated herein.
[0137] Embodiments of the present disclosure provide a spinneret inspection device, including:
[0138] at least one processor; and
[0139] a memory communicatively connected to the at least one processor; wherein,
[0140] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method of any of the foregoing embodiments.
[0141] Embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to any of the foregoing embodiments.
[0142] Figure 7 is a block diagram of an electronic device for implementing the sound source localization method according to the embodiments of the present disclosure. As Figure 7 shown, the electronic device includes: a memory 710 and a processor 720. The memory 710 stores a computer program that can run on the processor 720. The number of the memory 710 and the processor 720 may be one or more. The memory 710 may store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device is caused to execute the method provided by the foregoing method embodiments. The electronic device may further include: a communication interface 730 for communicating with external devices and performing data interaction and transmission.
[0143] If the memory 710, the processor 720, and the communication interface 730 are implemented independently, the memory 710, the processor 720, and the communication interface 730 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is shown in, but it does not mean that there is only one bus or one type of bus.
[0144] Optionally, in a specific implementation, if the memory 710, the processor 720, and the communication interface 730 are integrated on a single chip, the memory 710, the processor 720, and the communication interface 730 can communicate with each other through an internal interface.
[0145] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced reduced instruction set machine (ARM) architecture.
[0146] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0147] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, Bluetooth, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Versatile Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the present disclosure can be a non-volatile storage medium, in other words, a non-transitory storage medium.
[0148] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, or the like.
[0149] In the description of the embodiments of the present disclosure, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0150] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. "And / or" herein is merely a description of the relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0151] In the description of the embodiments of the present disclosure, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, unless otherwise specified, "a plurality of" means two or more.
[0152] The above are only exemplary embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A sound source localization method, include: In the case where at least one of the camera assemblies detects a target sound source, audio evaluation data of the camera assemblies are obtained, wherein the camera assemblies in the camera assemblies include a camera and a first microphone and a second microphone located on both sides of the camera in a horizontal direction, the orientation of the first microphone is opposite to the orientation of the second microphone, and the orientation of the first microphone is perpendicular to the orientation of the camera, and the audio evaluation data of the camera assemblies is obtained based on the audio data of the target sound source detected by the first microphone and the audio data of the target sound source detected by the second microphone; determining a target camera assembly from among the camera assemblies based on the audio evaluation data of the camera assemblies; The target camera assembly is controlled to locate the target sound source so that the target sound source is located within the field of view of the camera of the target camera assembly.
2. The method according to claim 1, in, Acquiring audio evaluation data of each camera assembly, including: For a first camera assembly among the camera assemblies, acquiring first audio data detected by a first receiver in the first camera assembly and second audio data detected by a second receiver in the first camera assembly; Determine average audio data of the first audio data and the second audio data, and use the average audio data as audio evaluation data of the first camera assembly; The audio evaluation data of each camera component is obtained based on at least the audio evaluation data of the first camera component.
3. The method according to claim 1, in, Determining a target camera assembly from among the camera assemblies based on the audio evaluation data of the camera assemblies includes: The camera component corresponding to the target audio evaluation data with the largest value among the audio evaluation data of each camera component is used as the target camera component.
4. The method according to claim 3, in, Determining a target camera assembly from among the camera assemblies based on the audio evaluation data of the camera assemblies further includes: Determining a first position of the target sound source based on third audio data detected by the first microphone in the target camera assembly and fourth audio data detected by the second microphone in the target camera assembly; When the rotation angle of the target camera assembly to the first position does not satisfy the preset angle condition of the target camera assembly, arranging the audio evaluation data of each camera assembly in descending order of numerical value; The camera assembly that is next to the target camera assembly in the arrangement order among the camera assemblies is used as a new target camera assembly; Controlling the target camera assembly to locate the target sound source includes: The new target camera assembly is controlled to locate the target sound source.
5. The method according to claim 1, in, The target sound source is a sound source with a sound waveform identical to a preset sound waveform; and / or, The target sound source is a sound source whose sound intensity is greater than or equal to a preset intensity threshold.
6. The method according to any one of claims 1 to 5, in, Controlling the target camera assembly to locate the target sound source includes: Determine an initial horizontal rotation direction of the target camera assembly based on third audio data detected by a first microphone in the target camera assembly and fourth audio data detected by a second microphone in the target camera assembly, wherein the initial horizontal rotation direction is the direction of the microphone corresponding to the audio data having a larger value between the third audio data and the fourth audio data; Controlling the target camera assembly to rotate horizontally toward the initial horizontal rotation direction; Acquire a difference between the third audio data and the fourth audio data in real time; Based on the difference, the rotation angle of the target camera assembly is controlled.
7. The method according to claim 6, in, Based on the difference, controlling the rotation angle of the target camera assembly includes: When the difference is zero, the target camera assembly is controlled to stop rotating.
8. The method according to claim 6, in, Based on the difference, controlling the rotation angle of the target camera assembly includes: When the difference is zero and the target sound source does not appear within the field of view of the camera of the target camera assembly, the camera assembly is controlled to rotate longitudinally until the target sound source is within the field of view of the camera of the target camera assembly, and the plane in which the longitudinal direction is located is perpendicular to the horizontal plane.
9. The method according to claim 6, in, Based on the difference, controlling the rotation angle of the target camera assembly includes: When the difference is not zero and the target sound source appears within the field of view of the camera of the target camera assembly, acquiring a field of view image of the camera in the target camera assembly; Based on the field of view, the rotation angle of the target camera assembly is controlled.
10. The method according to any one of claims 1 to 5, further comprising: include: When the target sound source appears within the field of view of the camera of the target camera assembly, the target sound source is located using the camera of the target camera assembly.
11. A sound source localization device, include: an acquisition unit, configured to acquire audio evaluation data of each camera assembly when at least one camera assembly among the camera assemblies detects a target sound source, wherein the camera assembly among the camera assemblies comprises a camera and a first microphone and a second microphone located on both sides of the camera in a horizontal direction, the orientation of the first microphone is opposite to that of the second microphone, and the orientation of the first microphone is perpendicular to that of the camera, and the audio evaluation data of the camera assembly is obtained based on the audio data of the target sound source detected by the first microphone and the audio data of the target sound source detected by the second microphone; a determination unit, configured to determine a target camera assembly from among the camera assemblies based on the audio evaluation data of the camera assemblies; The positioning unit is used to control the target camera assembly to position the target sound source so that the target sound source is located within the field of view of the camera of the target camera assembly.
12. The device according to claim 11, in, The acquisition unit is also used for: For a first camera assembly among the camera assemblies, acquiring first audio data detected by a first receiver in the first camera assembly and second audio data detected by a second receiver in the first camera assembly; Determine average audio data of the first audio data and the second audio data, and use the average audio data as audio evaluation data of the first camera assembly; The audio evaluation data of each camera component is obtained based on at least the audio evaluation data of the first camera component.
13. The device according to claim 11, in, The determining unit is further configured to: The camera component corresponding to the target audio evaluation data with the largest value among the audio evaluation data of each camera component is used as the target camera component.
14. The device according to claim 13, in, The determining unit is further configured to: Determining a first position of the target sound source based on third audio data detected by the first microphone in the target camera assembly and fourth audio data detected by the second microphone in the target camera assembly; When the rotation angle of the target camera assembly to the first position does not satisfy the preset angle condition of the target camera assembly, arranging the audio evaluation data of each camera assembly in descending order of numerical value; The camera assembly that is next to the target camera assembly in the arrangement order among the camera assemblies is used as a new target camera assembly; The positioning unit is also used for: The new target camera assembly is controlled to locate the target sound source.
15. The device according to claim 11, in, The target sound source is a sound source with a sound waveform identical to a preset sound waveform; and / or, The target sound source is a sound source whose sound intensity is greater than or equal to a preset intensity threshold.
16. The device according to any one of claims 11 to 15, in, The positioning unit is also used for: Determine an initial horizontal rotation direction of the target camera assembly based on third audio data detected by a first microphone in the target camera assembly and fourth audio data detected by a second microphone in the target camera assembly, wherein the initial horizontal rotation direction is the direction of the microphone corresponding to the audio data having a larger value between the third audio data and the fourth audio data; Controlling the target camera assembly to rotate horizontally toward the initial horizontal rotation direction; Acquire a difference between the third audio data and the fourth audio data in real time; Based on the difference, the rotation angle of the target camera assembly is controlled.
17. The device according to claim 16, in, The positioning unit is also used for: When the difference is zero, the target camera assembly is controlled to stop rotating.
18. The device according to claim 16, in, The positioning unit is also used for: When the difference is zero and the target sound source does not appear within the field of view of the camera of the target camera assembly, the camera assembly is controlled to rotate longitudinally until the target sound source is within the field of view of the camera of the target camera assembly, and the plane in which the longitudinal direction is located is perpendicular to the horizontal plane.
19. The device according to claim 16, in, The positioning unit is also used for: When the difference is not zero and the target sound source appears within the field of view of the camera of the target camera assembly, acquiring a field of view image of the camera in the target camera assembly; Based on the field of view, the rotation angle of the target camera assembly is controlled.
20. The device according to any one of claims 11 to 15, further comprising: include: The shooting unit is used to locate the target sound source by using the camera of the target camera assembly when the target sound source appears within the field of view of the camera of the target camera assembly.
21. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
22. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.