Positioning method and apparatus, video conference system, electronic device, and storage medium

By detecting motion events through a radar array and combining the collaborative work of image acquisition equipment and microphone arrays, fast portrait focusing and audio beam positioning are achieved in video conferencing, solving the problem of slow response speed in existing technologies and improving user experience.

CN115019337BActive Publication Date: 2025-10-24ZHEJIANG HUACHUANG VISION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210445353.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-10-24
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

The response speed of portrait focus and audio beam positioning in existing video conferencing is slow, affecting the user experience.

Method used

A radar array is used to detect motion events, control the image acquisition equipment to focus on the portrait, and use the microphone array to perform audio beam positioning. Combined with zoom and fixed-focus camera equipment, fast focusing and audio beam positioning can be achieved.

Benefits of technology

It improves the efficiency of speaker positioning in video conferences, shortens the response time of portrait focus and audio beam positioning, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019337B_ABST
    Figure CN115019337B_ABST
Patent Text Reader

Abstract

The application relates to a positioning method and device, a video conference system, an electronic device and a storage medium. The positioning method comprises the following steps: in the case that a radar array detects a motion event in a video conference process, controlling an image acquisition device to perform portrait focusing on a detected target object based on a motion event detection result of the radar array, obtaining a portrait focusing result; based on the motion event detection result, controlling a microphone array to position an audio beam of the target object, obtaining an audio beam positioning result; and based on the portrait focusing result and the audio beam positioning result, outputting close-up video data of the target object. The positioning efficiency of a speaker in the video conference is improved, and the response speed of the portrait focusing and the audio beam positioning is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video conferencing, and in particular to a positioning method and device, a video conferencing system, an electronic device and a storage medium. BACKGROUND

[0002] With the development of science and technology, in a video conference, a speaker can be focused on and an audio beam is positioned to output a close-up picture of the speaker in the video conference process, thereby improving the efficiency of the speaker conveying information. In the current video conference, the speaker is often focused on by a face detection algorithm combined with a face focusing algorithm, and the audio beam of the speaker is positioned based on sound collection and an audio algorithm. The response speed of the above-mentioned portrait focusing process and audio beam positioning process is slow, thereby reducing the user experience.

[0003] In view of the problem in the related art that the portrait focusing and audio beam positioning in a video conference are slow, there is no effective solution at present. SUMMARY

[0004] A positioning method, device, video conferencing system, electronic device and storage medium are provided in the present embodiment to solve the problem in the related art that the portrait focusing and audio beam positioning are slow.

[0005] In a first aspect, a positioning method is provided in the present embodiment, for a video conferencing system, the video conferencing system comprising a radar array, an image collection device, and a microphone array, the method comprising:

[0006] In the case where the radar array detects a motion event in a video conference process, based on a motion event detection result of the radar array, the image collection device is controlled to focus on a target object detected, to obtain a portrait focusing result;

[0007] Based on the motion event detection result, the microphone array is controlled to position an audio beam of the target object, to obtain an audio beam positioning result;

[0008] Based on the portrait focusing result and the audio beam positioning result, close-up video data of the target object is output.

[0009] In some embodiments, based on the motion event detection result of the radar array, the image collection device is controlled to focus on a target object detected, to obtain a portrait focusing result, comprising:

[0010] In a case where the radar array detects a single motion event, a motion object in the single motion event is taken as a target object, and first positioning information of the target object detected by the radar array is acquired;

[0011] The first positioning information is sent to the image acquisition device, and the image acquisition device is controlled to perform portrait focusing on the target object based on the first positioning information, to obtain a portrait focusing result.

[0012] In some embodiments, the controlling, based on the motion event detection result of the radar array, the image acquisition device to perform portrait focusing on the detected target object to obtain a portrait focusing result further includes:

[0013] In a case where the radar array detects two or more motion events, position information of motion objects in the two or more motion events is sent to the microphone array, so that the microphone array performs voice positioning based on the position information to determine a positioning range of a target object in the two or more motion events;

[0014] The positioning range of the target object is sent to the radar array, and the radar array is controlled to determine second positioning information of the target object according to the positioning range of the target object;

[0015] The second positioning information is sent to the image acquisition device, so that the image acquisition device performs portrait focusing on the target object based on the second positioning information, to obtain the portrait focusing result.

[0016] In some embodiments, the image acquisition device includes a zoom camera device and a fixed-focus camera device, and the method further includes:

[0017] Based on the motion event detection result of the radar array, a gimbal motor carried by the zoom camera device is controlled to rotate, so that the zoom camera device performs portrait zooming on the target object to obtain a portrait zooming result;

[0018] Based on the motion event detection result of the radar array, the fixed-focus camera device is controlled to perform portrait framing on the target object to obtain a portrait framing result;

[0019] Based on the portrait zooming result and the portrait framing result, an enlarged portrait of the target object is output.

[0020] In some embodiments, the controlling, based on the motion event detection result, the microphone array to perform positioning on an audio beam of the target object to obtain an audio beam positioning result includes:

[0021] According to the motion event detection result, the microphone array is controlled to position an audio beam of the target object according to a preset voice beam algorithm, and an audio beam positioning result is obtained.

[0022] In some embodiments, the method further comprises the following steps:

[0023] The video conference system is controlled to start in a case where the radar array detects a motion event in real time.

[0024] In some embodiments, the method further comprises the following steps:

[0025] The video conference system is controlled to enter a preset power saving state in a case where the radar array does not detect a motion event within a preset time period.

[0026] In a second aspect, the embodiment provides a positioning device for a video conference system, the video conference system comprising a radar array, an image acquisition device, and a microphone array, and the positioning device comprising:

[0027] A focusing module is configured to, in a case where the radar array detects a motion event during a video conference, control the image acquisition device to perform portrait focusing on a detected target object based on a motion event detection result of the radar array, and obtain a portrait focusing result.

[0028] A beam positioning module is configured to, based on the motion event detection result, control the microphone array to position an audio beam of the target object, and obtain an audio beam positioning result.

[0029] An output module is configured to output close-up video data of the target object based on the portrait focusing result and the audio beam positioning result.

[0030] In a third aspect, the embodiment provides a video conference system, comprising a radar array, an image acquisition device, a microphone array, and a processor, wherein:

[0031] The radar array is configured to, in a case where a motion event is detected in real time, send a signal indicating that a motion event is detected to the processor.

[0032] The image acquisition device is configured to perform face focusing on a target object during a video conference.

[0033] The microphone array is configured to position an audio beam of a target object during a video conference.

[0034] The processor is configured to execute the positioning method of the first aspect.

[0035] In a fourth aspect, an electronic device is provided in the present embodiment, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the positioning method of the first aspect when executing the computer program.

[0036] In a fourth aspect, a storage medium is provided in the present embodiment, which stores a computer program executable by a processor to implement the positioning method of the first aspect.

[0037] Compared with the related art, in the present embodiment, the positioning method, device, video conference system, electronic device and storage medium are provided, in the case that the radar array detects a motion event in the video conference process, the image acquisition device is controlled to perform portrait focusing on the detected target object based on the motion event detection result of the radar array, to obtain a portrait focusing result; the audio beam of the target object is positioned by the microphone array based on the motion event detection result, to obtain an audio beam positioning result; and the close-up video data of the target object is output based on the portrait focusing result and the audio beam positioning result. The efficiency of positioning the speaker in the video conference is improved, and the response speed of portrait focusing and audio beam positioning is accelerated.

[0038] The details of one or more embodiments of the present application are presented in the following drawings and description to make other features, objects and advantages of the present application more apparent. BRIEF DESCRIPTION OF DRAWINGS

[0039] The drawings described herein are intended to provide further understanding of the present application, form a part of the present application, and the illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0040] Figure 1 is a hardware structure block diagram of a terminal of the positioning method of the related art;

[0041] Figure 2 is a flowchart of the positioning method of the present embodiment;

[0042] Figure 3 is a flowchart of the video conference positioning method of the preferred embodiment;

[0043] Figure 4 is a structure block diagram of the positioning device of the present embodiment;

[0044] Figure 5 is a structure block diagram of the video conference system of the present embodiment. DETAILED DESCRIPTION

[0045] In order to clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and explained below in conjunction with the accompanying drawings and embodiments.

[0046] Unless otherwise defined, technical terms or scientific terms used in the present application shall have the general meaning understood by a person skilled in the art to which the present application belongs. In the present application, "one", "a", "an", "the", "these" and similar words do not represent a quantitative limitation, and they can be singular or plural. In the present application, the terms "include", "contain", "have" and any variants thereof are intended to cover non-exclusive inclusion; for example, a process, method and system, product or device containing a series of steps or modules (units) are not limited to the listed steps or modules (units), but can include steps or modules (units) not listed, or can include other steps or modules (units) inherent to the process, method, product or device. In the present application, the terms "connected", "connected", "coupled" and the like do not limit to physical or mechanical connection, but can include electrical connection, whether direct or indirect. In the present application, "multiple" means two or more. The association between the associated objects is described by "and / or", which means that there can be three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. In general, the character " / " represents an "or" relationship between the objects before and after. In the present application, the terms "first", "second", "third" and the like are only used to distinguish similar objects, and do not represent a specific order of the objects.

[0047] The method embodiments provided in the present embodiment can be executed in a terminal, a computer or a similar computing device. For example, the method embodiments are executed on a terminal, Figure 1 is a hardware structure diagram of the terminal of the positioning method of the present embodiment. As shown in Figure 1 , the terminal can include one or more (only one is shown in Figure 1 ) processor 102 and memory 104 for storing data, wherein the processor 102 can include but not limited to processing devices such as microprocessor MCU or programmable logic device FPGA. The above terminal can also include a transmission device 106 for communication function and an input / output device 108. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above terminal. For example, the terminal can include more or less components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0048] The memory 104 can be used to store computer programs, such as software programs of application software and modules, such as a computer program corresponding to the positioning method in the embodiment, and the processor 102 can execute various functional applications and data processing, i.e., implement the method described above, by running the computer programs stored in the memory 104. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include memories remotely arranged with respect to the processor 102, which can be connected to the terminal through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0049] The transmission device 106 is configured to receive or send data via a network. The network includes a wireless network provided by a communication provider of the terminal. In an example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In an example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.

[0050] In the embodiment, a positioning method is provided for a video conference system, wherein the video conference system includes a radar array, an image acquisition device, and a microphone array. Figure 2 A flowchart of the positioning method in the embodiment is shown in FIG. 2, which includes the following steps: Figure 2

[0051] In step S210, in a case where the radar array detects a motion event during a video conference, the image acquisition device is controlled to perform portrait focusing on the detected target object based on a motion event detection result of the radar array, to obtain a portrait focusing result.

[0052] The radar array can be set according to an actual application scenario. For example, the radar array can be a linear radar array. When the radar array is integrated into the video conference system, the radar array can be arranged at a front part of the video conference system, so that a signal detection range of the radar array covers the entire video conference area. In addition, in order to reduce the influence of interference signals in the video conference site on the radar array, a radar array with a working frequency band higher than 24 GHz is preferably used. Further, the sensitivity of the radar array can be adjusted in advance to reduce the interference of motion events in the video conference site. It can be understood that the target object can be a speaker in the video conference site.

[0053] ​Specifically, the radar array can achieve detection of a motion event in a video conference site based on the Doppler effect of electromagnetic waves. If the radar array detects a motion event within the signal coverage range of the radar array, the detection result of the motion event can be reported to a processor of the video conference system. The detection result of the motion event can include the position coordinates of a motion target and the number of motion targets. After the processor receives the motion event reported by the radar array, the processor controls an image acquisition device to perform portrait focusing on the detected target object based on the motion event detection result, thereby obtaining a portrait focusing result.

[0054] Further, the radar array can achieve detection of a single motion event or multiple motion events. In the case of detection of a single motion event by the radar array, the motion object of the single motion event can be taken as a target object. The processor controls the image acquisition device to perform portrait focusing on the target object based on the first positioning information of the target object reported by the radar array. In the case of detection of multiple motion events by the radar array, a speaker in a video conference needs to be determined from the multiple motion events. That is, a target object needs to be determined from the multiple motion events. Therefore, the processor can send the position information of the multiple motion events reported by the radar array to a microphone array, and control the microphone array to combine the positioning information of the target object detected by the radar array. For example, the position information of the multiple motion events can be sent to the microphone array, and after the microphone array detects positioning information with a relatively large target object range, the radar array further determines the final positioning information of the target object based on the positioning information with the relatively large range.

[0055] In addition, the image acquisition device includes a zoom camera device and a fixed-focus camera device. After the positioning information of the target object is determined, the processor can send the positioning information of the target object to the zoom camera device, so that the zoom camera device performs portrait zooming on the target object based on the positioning information, and performs portrait focusing on the target object according to a preset portrait focusing algorithm. The portrait focusing algorithm can be any general portrait focusing algorithm in the related art, which is not limited in the present embodiment. The processor can also send the positioning information of the target object to the fixed-focus camera device, and control the fixed-focus camera device to automatically frame the target object based on the positioning information. Through the portrait zooming and portrait focusing of the zoom camera device and the automatic framing of the fixed-focus camera device, a relatively clear close-up picture of the target object can be output in the video conference site.

[0056] In step S220, the microphone array is controlled to locate the audio beam of the target object based on the motion event detection result, and an audio beam positioning result is obtained.

[0057] The processor can send the positioning information of the target object to the microphone array, and the microphone array positions the audio beam of the target object according to the positioning information of the target object. Since the radar array reports the positioning information of the target object at a relatively fast speed, the preset audio algorithm in the microphone array can be accelerated to converge, and the time of echo cancellation can be shortened, so as to further realize the in-beam enhancement and out-of-beam suppression of the target object audio, and improve the experience of the user receiving the audio. In addition, the video conference system can be started in the case that the radar array detects a motion event in real time, and the video conference system can be controlled to be on standby or turned off in the case that the radar array does not detect a motion event within a preset time period, thereby improving the experience of the participants, improving the operation convenience of the video conference, and avoiding waste of electric resources.

[0058] In step S230, the close-up video data of the target object is output based on the portrait focusing result and the audio beam positioning result.

[0059] The radar array is used to realize real-time detection of the motion event. Compared with the method of positioning the target object based on the face detection algorithm and the audio algorithm in the related art, the positioning method of the embodiment improves the efficiency of positioning the target object, and further improves the response speed of the portrait focusing of the target object by the image acquisition device and the response speed of the audio beam positioning of the target object by the microphone array, so that the clear portrait video and audio of the speaker can be output, the efficiency and accuracy of information transmission of the speaker in the video conference are improved, and the experience of the participants is improved.

[0060] The steps S210 to S230 are as follows: in the case that the radar array detects a motion event during the video conference, the image acquisition device is controlled to focus on the detected target object based on the motion event detection result of the radar array, to obtain a portrait focusing result; the microphone array is controlled to position the audio beam of the target object based on the motion event detection result, to obtain an audio beam positioning result; and the close-up video data of the target object is output based on the portrait focusing result and the audio beam positioning result. The efficiency of positioning the speaker in the video conference is improved, and the response speed of the portrait focusing and the audio beam positioning is further improved.

[0061] In one embodiment, based on the step S210, the image acquisition device is controlled to focus on the detected target object based on the motion event detection result of the radar array, to obtain a portrait focusing result, which can include the following steps:

[0062] In step S211, in the case that the radar array detects a single motion event, the motion object in the single motion event is taken as the target object, and the first positioning information of the target object detected by the radar array is obtained.

[0063] Step S212, the first positioning information is sent to the image acquisition device, and the image acquisition device is controlled to perform portrait focusing on the target object based on the first positioning information, to obtain a portrait focusing result.

[0064] Specifically, the first positioning information can be sent to the zoom camera device, and the zoom camera device is controlled to perform portrait focusing on the target object, i.e., the speaker, based on its built-in portrait focusing algorithm, so as to obtain a video picture after portrait focusing processing. In the case where the radar array detects a single motion event, the steps S211 to S212 send the first positioning information of the motion object of the single motion event to the image acquisition device as the target object, and control the image acquisition device to perform portrait focusing on the target object, which can improve the response speed of portrait focusing in the video conference, thereby improving the participation experience of the participants.

[0065] In addition, in an embodiment, based on the motion event detection result of the radar array based on the step S210, the image acquisition device is controlled to perform portrait focusing on the detected target object to obtain a portrait focusing result, which can further include the following steps:

[0066] Step S213, in the case where the radar array detects two or more motion events, the position information of the motion objects in the two or more motion events is sent to the microphone array, so that the microphone array performs voice positioning based on the position information to determine the positioning range of the target object in the two or more motion events.

[0067] In the case where the radar array detects multiple motion events in the video conference site, it is necessary to filter the interference of other motion targets from the multiple motion events to determine the speaker of the video conference, i.e., to determine a target object from the multiple motion events. Therefore, the position information of the motion objects in the multiple motion events can be processed in combination with the microphone array to determine the positioning range where the speaker is located.

[0068] Step S214, the positioning range of the target object is sent to the radar array, and the radar array is controlled to determine the second positioning information of the target object according to the positioning range of the target object.

[0069] After the microphone array determines the positioning range of the speaker, the positioning information of the target object can be further determined from the positioning range by the radar array, so as to exclude the interference of other motion objects in the video conference site on the positioning of the target object, and improve the accuracy of positioning the speaker in the video conference site.

[0070] Step S215, the second positioning information is sent to the image acquisition device, so that the image acquisition device focuses on the portrait of the target object based on the second positioning information, and a portrait focusing result is obtained.

[0071] Additionally, in an embodiment, the image acquisition device includes a zoom camera device and a fixed-focus camera device, and the positioning method can further include the following steps:

[0072] Step S241, based on the motion event detection result of the radar array, a pan-tilt motor of the zoom camera device is controlled to rotate, so that the zoom camera device performs portrait zooming on the target object, and a portrait zooming result is obtained.

[0073] Step S242, based on the motion event detection result of the radar array, the fixed-focus camera device is controlled to frame the target object, and a portrait framing result is obtained.

[0074] Step S243, based on the portrait zooming result and the portrait framing result, an enlarged portrait of the target object is output.

[0075] Through the above steps S241 to S243, after the radar array reports the positioning information of the target object, the zoom camera device is driven by the pan-tilt to realize portrait zooming on the target object, and the fixed-focus camera device realizes portrait framing on the target object based on the positioning information, so that the portrait focusing function can be linked to improve the participation experience of the participants.

[0076] Additionally, in an embodiment, based on the above step S220, based on the motion event detection result, the audio beam of the target object is positioned by the microphone array, and an audio beam positioning result is obtained, which can include the following steps:

[0077] Step S221, according to the motion event detection result, the audio beam of the target object is positioned by the microphone array according to a preset voice beam algorithm, and an audio beam positioning result is obtained.

[0078] The preset voice beam algorithm is an algorithm built in the microphone array, which is used to realize the positioning of the audio beam of the target object, and the in-beam enhancement and out-of-beam suppression of the audio, so as to improve the participation experience of the participants. It should be noted that the voice beam algorithm can be any common voice beam algorithm in related technologies, and the specific voice beam algorithm can be determined according to the actual application scene and requirements, which is not limited herein.

[0079] Additionally, in an embodiment, the positioning method can further include the following steps:

[0080] Step S250, control the video conference system to start in the case that the radar array determines to detect a motion event in real time. Specifically, the video conference system starts in the case that the motion event reported by the radar array is received, thereby reducing the operation of the user, improving the convenience of the video conference starting operation and the experience of the participants.

[0081] Additionally, in one embodiment, the positioning method described above can further include the following steps:

[0082] Step S260, control the video conference system to enter a preset power saving state in the case that the radar array does not detect a motion event within a preset time period. The power saving state can be a standby state and a shutdown state. If the radar array does not detect any motion event within the preset time period, in order to reduce the waste of power resources of the video conference system, other devices in the video conference system can be controlled to be in a standby state or a shutdown state.

[0083] The preferred embodiments will be described and explained below.

[0084] Figure 3 The flowchart of the video conference positioning method of the preferred embodiments is shown in FIG. 3, which includes the following steps: Figure 3

[0085] Step S301, in the case that the video conference system is closed, the radar array reports a motion event to the processor in the case that a motion event is detected;

[0086] Step S302, start the video conference system based on the motion event reported by the radar array;

[0087] Step S303, in the video conference process, the radar array judges whether it is a single motion event in the case that a motion event is detected, if yes, steps S304, S305, S306, and S307 are executed respectively, otherwise, step S308 is executed;

[0088] Step S304, use a portrait focusing algorithm to focus on the target object according to the positioning information of the target object;

[0089] Step S305, control the pan-tilt motor of the zoom camera device to rotate based on the positioning information of the target object, so as to zoom the target object by the zoom camera device;

[0090] Step S306, control the fixed focus camera device to automatically frame the target object for local enlargement according to the positioning information of the target object;

[0091] ​Step S307, controlling the microphone array to realize audio beam positioning of the target object based on the positioning information of the target object by using a voice beam algorithm;

[0092] Step S308, filtering the moving object from the multiple motion events by using the microphone array to determine the positioning range of the target object;

[0093] Step S309, determining the positioning information of the target object according to the positioning range of the target object by using the radar array;

[0094] Step S310, outputting the close-up picture and audio data of the target object;

[0095] Step S311, in the case where the radar array does not detect a motion event within a preset time period, controlling the video conference system to enter a standby or shutdown state.

[0096] It should be noted that the steps shown in the above flow or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here. For example, S305 and S306.

[0097] In this embodiment, a positioning device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and the description of which has been made above. The terms "module", "unit", "sub-unit" and the like used below can be a combination of software and / or hardware that realizes a predetermined function. Although the device described in the following embodiments is preferably realized in software, hardware or a combination of software and hardware is also possible and is conceived.

[0098] Figure 4 is a structural block diagram of the positioning device 40 of this embodiment, as Figure 4 shown, the positioning device 40 comprises:

[0099] The focusing module 42 is configured to, in the case where the radar array detects a motion event during the video conference, control the image acquisition device to perform portrait focusing on the detected target object based on the motion event detection result of the radar array, to obtain a portrait focusing result.

[0100] The beam positioning module 44 is configured to control the microphone array to perform positioning on the audio beam of the target object based on the motion event detection result, to obtain an audio beam positioning result; and

[0101] The output module 46 is configured to output close-up video data of the target object based on the portrait focusing result and the audio beam positioning result.

[0102] It should be noted that the above various modules can be functional modules or program modules, which can be implemented by software or hardware. For the modules implemented by hardware, the above various modules can be located in the same processor; or the above various modules can also be located in different processors in any combination.

[0103] In this embodiment, a video conference system 50 is also provided, Figure 5 The structural block diagram of the video conference system 50 of this embodiment is shown in Figure 5 The video conference system 50 includes a radar array 52, an image acquisition device 54, a microphone array 56, and a processor 58, wherein:

[0104] The radar array 52 is configured to send a signal of detecting a motion event to the processor 58 when detecting the motion event in real time;

[0105] The image acquisition device 54 is configured to perform face focusing on a target object during a video conference;

[0106] The microphone array 56 is configured to position an audio beam of the target object during the video conference;

[0107] The processor 58 is configured to perform the positioning method provided in any of the above embodiments

[0108] In this embodiment, an electronic device is also provided, which includes a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0109] Optionally, the above electronic device can further include a transmission device and an input-output device, wherein the transmission device is connected with the processor, and the input-output device is connected with the processor.

[0110] Optionally, in this embodiment, the processor can be configured to perform the following steps through the computer program:

[0111] In the case that the radar array detects a motion event during a video conference, based on the motion event detection result of the radar array, the image acquisition device is controlled to perform portrait focusing on the detected target object, to obtain a portrait focusing result;

[0112] Based on the motion event detection result, the microphone array is controlled to position an audio beam of the target object, to obtain an audio beam positioning result;

[0113] Based on the portrait focusing result and the audio beam positioning result, close-up video data of the target object is output.

[0114] It should be noted that the specific examples in the present embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be described herein again.

[0115] In addition, in combination with the positioning method provided in the above embodiments, a storage medium can also be provided in the present embodiment to implement. The storage medium has a computer program stored thereon; the computer program is executed by a processor to implement any one of the positioning methods in the above embodiments.

[0116] It should be understood that the specific embodiments described herein are only used to explain this application, but not to limit it. According to the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.

[0117] Obviously, the drawings are only some examples or embodiments of the present application, and those of ordinary skill in the art can also apply the present application to other similar situations without creative labor. In addition, it can be understood that although the work done in the development process may be complex and long, some design, manufacture or production changes made by those of ordinary skill in the art according to the technical content disclosed in the present application are only routine technical means and should not be regarded as insufficient disclosure of the present application.

[0118] The word "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiments can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean independence or alternative to other embodiments. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0119] The above-described embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A positioning method for a video conferencing system, characterized by, The video conference system comprises a radar array, an image acquisition device, and a microphone array, and the method comprises: In the case that the radar array detects a motion event during a video conference, the image acquisition device is controlled to perform portrait focusing on a detected target object based on a motion event detection result of the radar array, to obtain a portrait focusing result; The microphone array is controlled to perform audio beam positioning on the target object based on the motion event detection result, to obtain an audio beam positioning result; Based on the portrait focusing result and the audio beam positioning result, close-up video data of the target object is outputted; The image acquisition device is controlled to perform portrait focusing on a detected target object based on a motion event detection result of the radar array, to obtain a portrait focusing result, comprising: In the case that the radar array detects two or more motion events, position information of a moving object in the two or more motion events is sent to the microphone array, so that the microphone array performs voice positioning based on the position information to determine a positioning range of a target object in the two or more motion events; The positioning range of the target object is sent to the radar array, and the radar array is controlled to determine second positioning information of the target object according to the positioning range of the target object; The second positioning information is sent to the image acquisition device, so that the image acquisition device performs portrait focusing on the target object based on the second positioning information to obtain the portrait focusing result.

2. The positioning method according to claim 1, characterized in that, The image acquisition device comprises a zoom camera device and a fixed-focus camera device, and the method further comprises: The zoom camera device is controlled to rotate a gimbal motor carried thereby based on a motion event detection result of the radar array, so that the zoom camera device performs portrait zooming on the target object to obtain a portrait zooming result; The fixed-focus camera device is controlled to perform portrait framing on the target object based on a motion event detection result of the radar array, to obtain a portrait framing result; Based on the portrait zooming result and the portrait framing result, an enlarged portrait of the target object is outputted.

3. The positioning method of claim 1, wherein, The microphone array is controlled to perform audio beam positioning on the target object based on a motion event detection result, to obtain an audio beam positioning result, comprising: According to the motion event detection result, the microphone array is controlled to perform audio beam positioning on the target object according to a preset voice beam algorithm, to obtain the audio beam positioning result.

4. The positioning method of claim 1, wherein, Further comprising the following steps: The video conference system is controlled to start in the case that it is determined that the radar array detects a motion event in real time.

5. The positioning method of claim 1, wherein, Further comprising the following steps: The video conference system is controlled to enter a preset power saving state in the case that the radar array does not detect a motion event within a preset time period.

6. A positioning device for use in a video conferencing system, characterized by The video conference system comprises a radar array, an image acquisition device, and a microphone array, and the positioning device comprises: The focusing module is configured to, in a case where the radar array detects a motion event during the video conference, control the image acquisition device to perform portrait focusing on a detected target object based on a motion event detection result of the radar array, to obtain a portrait focusing result. The beam positioning module is configured to control the microphone array to position an audio beam of the target object based on the motion event detection result, to obtain an audio beam positioning result; and The output module is configured to output close-up video data of the target object based on the portrait focusing result and the audio beam positioning result. The focusing module is configured to, in a case where the radar array detects a motion event during the video conference, control the image acquisition device to perform portrait focusing on a detected target object based on a motion event detection result of the radar array, to obtain a portrait focusing result. In a case where the radar array detects two or more motion events, the position information of a moving object in the two or more motion events is sent to the microphone array, so that the microphone array performs voice positioning based on the position information to determine a positioning range of a target object in the two or more motion events. The positioning range of the target object is sent to the radar array, and the radar array is controlled to determine second positioning information of the target object according to the positioning range of the target object. The second positioning information is sent to the image acquisition device, so that the image acquisition device performs portrait focusing on the target object based on the second positioning information to obtain the portrait focusing result.

7. A video conferencing system characterized by The video conference system comprises a radar array, an image acquisition device, a microphone array, and a processor, wherein: The radar array is configured to, in a case where a motion event is detected in real time, send a signal indicating that a motion event is detected to the processor; The image acquisition device is configured to perform face focusing on a target object during a video conference; The microphone array is configured to position an audio beam of a target object during a video conference; The processor is configured to perform the positioning method of any one of claims 1 to 5. 8.An electronic device comprising a memory and a processor, the electronic device comprising: The memory stores a computer program, and the processor is configured to run the computer program to perform the positioning method of any one of claims 1 to 5.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the positioning method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Position directed acoustic array and beamforming methods

    CN104244143A

  • Method and device for identifying speakers in multi-person video

    CN112487246A

  • Video conference method, system and device based on microphone array and storage medium

    CN113099160A

  • Modified vehicle radar intelligent monitoring trigger equipment

    CN205545687U