A video conference method, apparatus, electronic device and storage medium

By establishing a 3D coordinate system and correcting camera height and target data in video conferencing, and combining this with voice signals to determine the speaker's location, the problem of inaccurate location recognition in multi-person video conferencing is solved, improving recognition accuracy and system robustness.

CN119893025BActive Publication Date: 2026-01-02SHENZHEN ZHIWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411157138.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-01-02
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

In existing video conferencing, the uncertainty of the speaker's voiceprint characteristics due to differences in scene environment and camera hardware affects the accuracy of participant identification, especially in multi-person or multi-terminal video conferencing, where the problem of inaccurate location identification is more prominent.

Method used

By acquiring scene images and target images at different locations using a camera at the same fixed position, a 3D coordinate system is established, and the camera height and target length/width data are corrected. Combined with voice signals, the location and identity of the speaker are determined.

Benefits of technology

It improves the recognition accuracy in multi-person or multi-terminal video conferences, reduces recognition errors caused by hardware differences, enhances the system's robustness and the ability to quickly identify participants, and optimizes the meeting process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119893025B_ABST
    Figure CN119893025B_ABST
Patent Text Reader

Abstract

The application discloses a kind of video conference method, device, electronic equipment and storage medium, belong to video conference technical field, the method includes the same fixed position camera, obtains scene image, and the same target in the same scene at least different position image of interval preset distance;According to scene image, establish 3D coordinate system, and determine scene and target parameter;According to the height value of camera and the parameter data of target, and horizon plane, the scene and its corresponding 3D coordinate system are corrected;Video image and the voice signal that is synchronous with video image are collected;Based on 3D coordinate system, and according to video image and voice signal, determine the position of current speaker in video image;According to the position of current speaker in video image and corresponding voice information, the identity of current speaker is marked.Solve the problem that there can be inaccurate position recognition in video conference, improve the overall performance and user experience of system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of video conferencing, and particularly relates to a video conferencing method and device, electronic equipment and a storage medium. BACKGROUND

[0002] The video conferencing method is a conference in which people in two or more places have face-to-face conversations through communication equipment and a network. According to the number of participating places, the video conference can be divided into point-to-point conference and multi-point conference.

[0003] Using the video conference system, the participants can hear the voices of other conference sites, see the images, actions and expressions of the on-site participants of other conference sites, and send electronic demonstration contents, so that the participants have a sense of being on the scene.

[0004] At present, due to the influence of the scene environment and the uncertainty of the voice signal, the voiceprint features of the speaker also have various uncertainties, which further affects the identification accuracy of the participants, especially in multi-person participation or multi-terminal video conference. For example, due to the influence of the scene environment and the camera hardware itself, there may be problems of inaccurate position recognition and camera selection, that is, for different cameras, due to the differences in their imaging components and other hardware, the identification accuracy may differ greatly. Therefore, the existing video conference method and device still have a large room for improvement. SUMMARY

[0005] The technical problem to be solved by the application is to overcome the shortcomings of the prior art and provide a video conference method and device.

[0006] The technical solution adopted to solve the above technical problem is: a video conference method, comprising:

[0007] acquiring a scene image by a camera at a same fixed position, and at least different position images of a same target in a same scene at an interval of a preset distance;

[0008] establishing a 3D coordinate system according to the scene image, determining a horizon plane of the scene, measuring a camera height value, and a length and / or width data of the target;

[0009] correcting the scene and the corresponding 3D coordinate system according to the camera height value and the at least two length or width data of the target, and the horizon plane;

[0010] when a participant appears in the scene, collecting a video image and a voice signal synchronized with the video image;

[0011] determining a position of a current speaker in the video image according to the video image and the voice signal based on the 3D coordinate system;

[0012] annotating the identity of the current speaker according to the position of the current speaker in the video image and the corresponding voice information.

[0013] Further, the method of acquiring the scene image and the at least different position images of the same target in the same scene at an interval of a preset distance through the camera at the same fixed position comprises:

[0014] fixing the camera at a certain height and acquiring the scene image through the camera;

[0015] sequentially setting the target at the first position and the second position and respectively acquiring the images containing the complete length and / or width of the target.

[0016] Further, the method of sequentially setting the target at the first position and the second position and respectively acquiring the images containing the complete length and / or width of the target further comprises:

[0017] sequentially setting the target at the third position and the fourth position and respectively acquiring the images containing the complete length and / or width of the target; wherein the first position, the second position, the third position and the fourth position are located at different directions and at least two positions are located at the edge of the scene.

[0018] Further, the method of establishing the 3D coordinate system according to the scene image, determining the horizon plane of the scene, and measuring the height value of the camera and the length and / or width data of the target comprises:

[0019] establishing the internal 3D coordinate system according to the preset AR system and the scene image and mapping the ground of the scene as the horizon plane in the image;

[0020] measuring the height value of the camera relative to the ground of the scene and the length and / or width data of the target.

[0021] Further, the method of correcting the scene and the corresponding 3D coordinate system according to the height value of the camera and the at least two length or width data of the target and the horizon plane comprises:

[0022] determining the field of view angle of the camera according to the height value of the camera;

[0023] determining the proportional relationship between the real object and the image of the target according to the length or width data of the target and the imaging of the target in the scene;

[0024] determining the mapping relationship of each orthographic plane of the scene relative to the 3D coordinate system according to the field of view angle and the horizon plane;

[0025] According to the proportional relationship, the mapping correspondence between the scene position and the 3D coordinate position is corrected.

[0026] Further, the position of the current speaker in the video image is determined based on the 3D coordinate system and according to the video image and the voice signal, including:

[0027] According to the voice signal determination, the voiceprint feature of the speaker is determined;

[0028] According to the voiceprint feature, the identity information of the speaker is determined;

[0029] In a conference scene, the position of the speaker in the video image is determined by using a directional microphone and a camera.

[0030] Further, it also includes:

[0031] According to the position of the current speaker in the video image and according to the corresponding voiceprint information, the identity of the current speaker is labeled, the voice information is converted into text information, and the identity information of the speaker is added to the text in each piece of text information.

[0032] A video conference device, comprising:

[0033] An image acquisition module is configured to acquire a scene image by a camera at a same fixed position, and at least different position images of a same target at a same scene with an interval of a preset distance;

[0034] A coordinate system generation module is configured to establish a 3D coordinate system according to the scene image, determine a horizon plane of the scene, measure a camera height value, and measure length and / or width data of the target;

[0035] A data correction module is configured to correct the scene and the corresponding 3D coordinate system according to the camera height value, the at least two length or width data of the target, and the horizon plane;

[0036] A synchronous acquisition module is configured to acquire a video image and a voice signal synchronized with the video image when a participant appears in the scene;

[0037] A position determination module is configured to determine the position of a current speaker in the video image based on the 3D coordinate system and according to the video image and the voice signal;

[0038] An identity labeling module is configured to label the identity of the current speaker according to the position of the current speaker in the video image and the corresponding voice information.

[0039] An electronic device, comprising:

[0040] One or more processors; and one or more machine-readable media having stored thereon instructions that, when executed by the one or more processors, cause the electronic device to perform a method of video conferencing.

[0041] A computer-readable storage medium stores a computer program that causes a processor to perform a method of video conferencing.

[0042] The beneficial effects of the present application are as follows:

[0043] By the same fixed position camera, the scene image is obtained, and at least different position images of the same target in the same scene at an interval preset distance; a 3D coordinate system is established according to the scene image, and a horizon plane of the scene is determined, and a camera height value and length and / or width data of the target are measured; the scene and the corresponding 3D coordinate system are corrected according to the camera height value and at least two length or width data of the target, and the horizon plane; when a participant appears in the scene, a video image and a voice signal synchronized with the video image are collected; the position of the current speaker in the video image is determined based on the 3D coordinate system and according to the video image and the voice signal; and the current speaker is identity labeled according to the position of the current speaker in the video image and the corresponding voice information. Through the 3D space modeling and multi-modal information fusion method, the problem of possible inaccurate position recognition in video conference and the problem of camera selection are effectively solved, that is, for different cameras, due to the difference in imaging components and other hardware, the recognition accuracy may differ greatly, thereby improving the overall performance and user experience of the system. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 is a step flow chart of a video conference method provided by an embodiment of the present application;

[0045] Figure 2 is a structural block diagram of a video conference device provided by an embodiment of the present application;

[0046] Figure 3 is a structural schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0048] Referring to Figure 1 , a step flow chart of a video conference method according to an embodiment of the present application is shown, which includes:

[0049] Step S100, acquiring a scene image and at least different position images of the same target in the same scene by a camera at a same fixed position;

[0050] Step S200, establishing a 3D coordinate system according to the scene image, determining a horizon plane of the scene, measuring a camera height value, and length and / or width data of the target;

[0051] Step S300, correcting the scene and the corresponding 3D coordinate system according to the camera height value, at least two length or width data of the target, and the horizon plane;

[0052] Step S400, collecting a video image and a synchronous voice signal when a participant appears in the scene;

[0053] Step S500, determining a position of a current speaker in the video image based on the 3D coordinate system and according to the video image and the voice signal;

[0054] Step S600, identity labeling the current speaker according to the position of the current speaker in the video image and corresponding voice information.

[0055] In the embodiment, the scene image and the images of the target at different positions are acquired by a camera at a same fixed position, a 3D coordinate system is established, a horizon plane of the scene is determined, and a camera height value and length and / or width data of the target are measured. This step helps to establish an accurate three-dimensional space model for subsequent positioning and correction; the scene and the corresponding 3D coordinate system are corrected according to the camera height value, length or width data of the target, and the horizon plane; even under different camera perspectives, consistent spatial relationships and proportions are ensured, and recognition errors caused by hardware differences are reduced; when a participant appears in the scene, a video image and a synchronous voice signal are collected. The position of the current speaker in the video image is determined by the 3D coordinate system in combination with the video image and the voice signal. The multi-modal information fusion can improve the positioning accuracy, especially when the voiceprint features are unclear or the environmental noise is large. The current speaker is identity labeled according to the position of the speaker in the video image and corresponding voice information. This helps participants to quickly identify and improve the interactive efficiency of the conference.

[0056] By 3D space modeling and correction, the recognition error caused by scene environment and hardware differences can be reduced, and the recognition accuracy in multi-person participation or multi-terminal video conference can be improved. Combining visual and voice information for positioning can provide supplement when single modal information is not enough for accurate recognition, and enhance the robustness of the system; accurate voiceprint and location recognition helps participants to identify faster, thereby optimizing the conference process and participant experience.

[0057] In an embodiment of the present application, the scene image is obtained by a camera at a fixed position, and at least different position images of the same target at a preset distance apart in the same scene, comprising:

[0058] The camera is fixed at a certain height, and the scene image is obtained by the camera; for example, the camera is fixed at a height of 1.5 to 1.8 meters, specifically, the camera can be fixed to the upper end of a large screen or the position of a whiteboard; the scene image obtained by the camera at the fixed position is used as the background image of the scene.

[0059] A measuring target can be set, which can be a ruler or a target with a fixed length, or a person; the target is sequentially set at a first position and a second position, and images containing the complete length and / or width of the target are obtained, for example, a fixed length measuring ruler or a certain length rod is respectively placed at different positions in the field of view of the camera, and at least two images are obtained at the two positions; the horizontal distance between the two positions from the camera is different, for example, 3 meters apart; since the same target is closer to the camera than farther away, its imaging size in the same scene is different. Two images of the same target are used as the measuring target for subsequent comparison.

[0060] Further, the target is sequentially set at a first position and a second position, and images containing the complete length and / or width of the target are obtained, and then further comprising:

[0061] The target is sequentially set at a third position and a fourth position, and images containing the complete length and / or width of the target are obtained; wherein the first position, the second position, the third position and the fourth position are located at different positions, and at least two positions are located at the edge of the scene. By setting the edge position, the scene can be divided into local areas, and only the images and audio content within the rendering division area are responsive, and the system resources can be saved, and the processing efficiency of the system for data can be improved. Only the sound in the local area is responded to, when the distance of the audio signal is not in the scene, that is, in the video image, if the sound is outside the divided scene, the system will shield the audio as noise.

[0062] It should be noted that the scale of the image can be non-linear, so more positions can be set, the target is located in these positions, the corresponding images are acquired, and the scene and the 3D coordinates are corrected according to the difference between the images and the actual measurement, which can make the recognized position more accurate.

[0063] It should be further noted that in a video conference, video and audio are different signals, and the recognition methods are different, so the video and the audio can be better corresponded in subsequent processing and recognition. Although the current face recognition technology is quite mature, in a video conference, the participants can only show a side face or be blocked by others in the video, and due to the limitation of the conference room, especially in a small conference room, the participants can be close to each other, which increases the difficulty of recognition. Through the above correction of the scene, the position of the participant can be accurately recognized, and the voiceprint and identity information thereof can be corresponded. When it is detected that the sound source is not in the scene, the voice information can be shielded, and only the audio in the scene is effective audio, which can make the sound not affected by external noise, especially when the environment is not ideal and unstable, which can cause a poor conference effect. Thus, the audio effect in the video conference is improved.

[0064] In an embodiment of the present application, the 3D coordinate system is established according to the scene image, the horizon plane of the scene is determined, and the height value of the camera and the length and / or width data of the target are measured, comprising:

[0065] The internal 3D coordinate system is established according to the preset AR system and the scene image, and the ground of the scene is mapped as the horizon plane in the image; the height value of the camera relative to the ground of the scene and the length and / or width data of the target are measured. Since the video or the system itself does not know the actual ground position, by establishing the 3D coordinate, the nearest ground in the video that the vision can reach is taken as the horizon, and the scene ground is formed with the nearest ground in the imaging.

[0066] In an embodiment of the present application, the scene and the corresponding 3D coordinate system are corrected according to the height value of the camera and at least two length or width data of the target, and the horizon plane, comprising:

[0067] According to the height value of the camera, the field of view angle of the camera is determined; the field of view angle (FOV) is divided into a horizontal field of view angle (HFOV) and a vertical field of view angle (VFOV), which respectively determine the field of view range of the camera in the horizontal and vertical directions. The focal length and the size of the imaging area are the main factors affecting the field of view angle. The longer the focal length, the smaller the field of view angle, and the narrower the field of view range; the shorter the focal length, the larger the field of view angle, and the wider the field of view range. The determining factors of the horizontal field of view angle are mainly the focal length and the width of the imaging area, and the main determining factors of the vertical field of view angle are the focal length and the height of the imaging area. The calculation formula is horizontal field of view angle (HFOV) = 2arctan (w / 2f), vertical field of view angle (VFOV) = 2arctan (h / 2f), wherein w is the field of view width, h is the field of view height, and f is the focal length of the lens.

[0068] According to the length or width data of the target and the imaging of the target in the scene, the proportional relationship between the real object and the imaging is determined; wherein the length or width of the imaging of the target in the system can be measured by using the tools provided by the system, and the size corresponding relationship between the real object and the imaging can be obtained by measuring the length or width of the target;

[0069] According to the view angle and the horizon plane, the mapping relationship of each orthographic plane of the scene relative to the 3D coordinate system is determined; according to the proportional relationship, the mapping corresponding relationship of the scene position and the 3D coordinate position is corrected; after the corresponding relationship is determined according to the field of view angle and the horizon plane, the ground of the scene can be determined as the intersection of the X axis and the Y axis of the 3D coordinate system, and the actual scene and the 3D coordinate are mapped according to the size corresponding relationship. Through the scene mapping relationship and the position mapping relationship, when the participant speaks, the participant who speaks can be more accurately positioned. The problem of "misattribution" in the target AR video conference is solved, that is, the position of the participant who originally speaks is identified as the position of another participant due to the coordinate problem. Therefore, through the above 3D coordinate and scene correction, the identified position can be more accurate.

[0070] In an embodiment of the present application, the position of the current speaker in the video image is determined based on the 3D coordinate system and according to the video image and the voice signal, comprising:

[0071] According to the voice signal, the voiceprint feature of the speaker is determined; according to the voiceprint feature, the identity information of the speaker is determined; the position of the speaker in the video image is determined in the conference scene by using the directional microphone and the camera.

[0072] In this embodiment, after the multi-person video conference starts, the video image of the conference site and the language signal synchronized with the video image can be collected in real time, and then the speaker is determined based on the collected video image and voice signal, and the identity of the determined speaker is labeled in the video image.

[0073] In an embodiment of the present application, the current speaker is also labeled with identity based on the position of the current speaker in the video image and the corresponding voiceprint information, the voice information is converted into text information, and identity information of the speaker is added to the text in each piece of text information.

[0074] For example, when the identity of the determined speaker is labeled in the video image, the relevant information of the corresponding speaker can be labeled at the position of the determined speaker in the video image, wherein the relevant information of the speaker includes adding a label to the participants, such as the name, gender, position, contact information, etc. of the participants, or a code, such as A001, A002, etc.

[0075] After the identity of the determined speaker is labeled in the video image, the video image can be sent to the conference equipment of each participant for use by the multiple parties participating in the conference, such as for generating a conference summary; or the video image can also be sent to a third party; through the generated summary, the sound is converted into corresponding text, and the corresponding identity label is added to the text according to the voiceprint information, and the text can be sorted in chronological order, so that when viewed, the speaker's speech content at a certain time can be known.

[0076] It should be noted that the approximate position of the speaker can be well obtained directly by the directional microphone, and the recognition is further improved in accuracy by combining voiceprint analysis and image analysis.

[0077] In this embodiment, the voice signal can be processed by using a voice model, for example, a deep neural network DNN model, a convolutional neural network CNN model, etc. The voiceprint feature is extracted from the obtained voice signal by the voice model to obtain a voiceprint vector, and the voiceprint vector is input into a classification model to obtain the identity information of the speaker.

[0078] In an embodiment of the present application, the image recognition model can be used to distinguish the speaker, that is, in this embodiment, the collected video image is only recognized for the lips in the image, without recognizing the face, for distinguishing whether the person is speaking, and in combination with the voiceprint recognition, the speaker in the image is located and the identity information is labeled, wherein for the lip analysis, the CNN-LSTM model or the 3D-ConvNet model can be used. The CNN-LSTM model is an integrated model of the convolutional neural network CNN and the long short-term memory network LSTM, the CNN part of the model processes the data, and the one-dimensional result is input into the LSTM model. The 3D-ConvNet model is a convolutional neural network model with a 3D convolution kernel, and the 3D convolution module in the 3D-ConvNet model can be used to extract the time and space features of the video frames in the first video image, for recognizing whether the participant is speaking.

[0079] For the device embodiment, it is basically similar to the method embodiment, so the description is relatively simple, and the related parts are described in the part of the method embodiment.

[0080] As shown in Figure 2 In an embodiment of the present application, a video conference device is also disclosed, comprising:

[0081] The image acquisition module 100 is configured to acquire scene images and images of the same target at different positions in the same scene at intervals of a preset distance through the same fixed position camera.

[0082] The coordinate system generation module 200 is configured to establish a 3D coordinate system according to the scene images, determine a horizon plane of the scene, measure a camera height value, and measure length and / or width data of the target.

[0083] The data correction module 300 is configured to correct the scene and the corresponding 3D coordinate system according to the camera height value, the at least two length or width data of the target, and the horizon plane.

[0084] The synchronous acquisition module 400 is configured to acquire video images and voice signals synchronized with the video images when a participant appears in the scene.

[0085] The position determination module 500 is configured to determine the position of the current speaker in the video images based on the 3D coordinate system and according to the video images and the voice signals.

[0086] The identity labeling module 600 is configured to label the identity of the current speaker according to the position of the current speaker in the video images and the corresponding voice information.

[0087] Referring to Figure 3In embodiments of the present application, the present application also provides a computer device in the form of a general purpose computing device 12, the components of which can include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including the system memory 28 to the processing unit 16.

[0088] The bus 18 represents one or more of several types of bus structures, including a memory bus 18 or memory controller, a peripheral bus 18, a graphics acceleration bus, a processor or local bus using any of a variety of bus architectures including an Industry Standard Architecture (ISA), Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0089] The computer device 12 typically includes a variety of computer system readable media. Such media can be any available media that is located either in or out of the computer device 12, such as volatile and non-volatile media, removable and non-removable media.

[0090] The system memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 31 and / or cache memory 32. The computer device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 can be provided for reading from and writing to non-removable, non-volatile magnetic media (typically called a "hard drive"). Although not specifically shown, the computer device 12 typically includes a removable mass storage medium such as, for example, a floppy disk device, a magnetic tape Figure 3 Although not specifically shown, the computer device 12 typically includes a removable mass storage medium such as, for example, a floppy disk device, a magnetic tape

[0091] Program / utility 41, having a set of programs / modules 42, can be stored in, for example, memory by way of example, such programs includes, but is not limited to, an operating system, one or more application programs, other program modules 42, and program data, each of which or a combination can include implementation of a networking environment. Program modules 42 generally carry out the functions and / or methodologies of embodiments of the present application as described herein.

[0092] Computer device 12 can also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, a camera, etc.; can communicate with one or more devices that enable a user to interact with computer device 12; and / or can communicate with any devices (such as a network card, a modem, etc.) that enable computer device 12 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface 22. Still yet, computer device 12 can communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or the Internet) through network adapter 20. As depicted, network adapter 21 communicates with the other components of computer device 12 via bus 18. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with computer device 12. Examples, include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems 34, etc.

[0093] Processing unit 16 performs various function applications and data processing by running programs stored in system memory 28, such as implementing the video conference method provided by embodiments of the present application.

[0094] When the above processing unit 16 executes the above program, it realizes: acquiring scene images by the same fixed position camera, and at least different position images of the same target in the same scene at an interval of a preset distance; establishing a 3D coordinate system according to the scene images, and determining a horizon plane of the scene, and measuring a camera height value, and length and / or width data of the target; correcting the scene and the corresponding 3D coordinate system according to the camera height value and at least two length or width data of the target, and the horizon plane; when a participant appears in the scene, collecting a video image and a voice signal synchronized with the video image; determining a position of a current speaker in the video image based on the 3D coordinate system, and according to the video image and the voice signal; and marking the identity of the current speaker according to the position of the current speaker in the video image and corresponding voice information.

[0095] In embodiments of the present application, the present application also provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to realize the video conference method provided by all embodiments of the present application.

[0096] That is, the program is executed by the processor to achieve: through the same fixed position camera, obtain scene images, and at least different position images of the same target in the same scene at an interval preset distance; establish a 3D coordinate system according to the scene images, and determine a horizon plane of the scene, and measure a camera height value, and length and / or width data of the target; correct the scene and the corresponding 3D coordinate system according to the camera height value and at least two length or width data of the target, and the horizon plane; when a participant appears in the scene, collect a video image and a voice signal synchronized with the video image; determine a position of a current speaker in the video image based on the 3D coordinate system and according to the video image and the voice signal; and perform identity labeling on the current speaker according to the position of the current speaker in the video image and corresponding voice information.

[0097] Any combination of one or more computer readable medium can be employed. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this document, the computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0098] A computer readable signal medium can include a propagated data signal with computer readable program code embodied therein, for use by or in connection with an instruction execution system, apparatus, or device. The computer readable signal medium can be, for example, but not limited to, a computer readable program code that can be transmitted over a computer readable medium. The computer readable signal medium can also be a computer readable storage medium, which can be any tangible medium that can be used to store the computer readable program code for use by or in connection with an instruction execution system, apparatus, or device.

[0099] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++, Python, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0100] The various embodiments in the specification are described in progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts among the various embodiments can be mutually referred to.

[0101] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device, or computer program product. Therefore, the embodiments of the present application can be in the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0102] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal devices produce the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The apparatus that implements the functions specified in a flow or multiple flows and / or blocks

[0103] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or steps.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or steps.

[0105] Although preferred embodiments of the present application have been described, those skilled in the art will be able to make additional modifications and variations to these embodiments without departing from the scope of the present application. Accordingly, the appended claims are intended to encompass all such modifications and variations as falling within the scope of the present application.

[0106] Finally, it should be noted that, in this document, the terms "first", "second", etc. are used merely to identify similar entities or operations from one another, and do not necessarily require or imply any actual relationship or order between such entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0107] The above provides a video conference method, device, electronic equipment and storage medium. The principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application. For those skilled in the art, according to the idea of the present application, the specific implementation manners and application scope can be changed. The above description of the present application should not be understood as a limitation.

Claims

1. A method of video conferencing, characterized by, The application relates to a method for establishing a 3D coordinate system in a conference scene, and a system thereof. The method comprises the following steps: acquiring a scene image and at least two images of a same target at different positions in the scene by using a camera fixed at a same position; establishing a 3D coordinate system according to the scene image; determining a horizon plane of the scene, a height value of the camera and length and / or width data of the target; correcting the scene and the corresponding 3D coordinate system according to the height value of the camera, the length and / or width data of the target and the horizon plane; acquiring a video image and a voice signal synchronized with the video image when a participant appears in the scene; determining a position of a current speaker in the video image according to the 3D coordinate system, the video image and the voice signal; and marking the identity of the current speaker according to the position of the current speaker in the video image and corresponding voice information. The system comprises a camera, a microphone, a processor and a storage device. The camera is used for acquiring a scene image and at least two images of a same target at different positions in the scene. The microphone is used for acquiring a voice signal synchronized with the video image. The processor is used for establishing a 3D coordinate system according to the scene image, determining a horizon plane of the scene, a height value of the camera and length and / or width data of the target, correcting the scene and the corresponding 3D coordinate system according to the height value of the camera, the length and / or width data of the target and the horizon plane, determining a position of a current speaker in the video image according to the 3D coordinate system, the video image and the voice signal, and marking the identity of the current speaker according to the position of the current speaker in the video image and corresponding voice information. The storage device is used for storing the 3D coordinate system, the video image, the voice signal, the position of the current speaker in the video image and the identity of the current speaker.

2. The method of claim 1, wherein, The method further comprises the following steps: The target is sequentially arranged at a third position and a fourth position, and images containing complete length and / or width of the target are acquired.

3. The method of claim 1, wherein, The first position, the second position, the third position and the fourth position are located at different positions, and at least two positions are located at edges of the scene. The method further comprises the following steps: The voice signal is used for determining a voiceprint feature of the speaker. The voiceprint feature is used for determining identity information of the speaker.

4. The method of claim 1, wherein, The camera and the microphone are used for determining the position of the speaker in the video image in the conference scene. The method further comprises the following steps:

5. A video conferencing apparatus, characterized by The position of the current speaker in the video image and the corresponding voice information are used for marking the identity of the current speaker, converting the voice information into text information, and adding identity information of the speaker to text in each piece of text information. The system further comprises a display device. The display device is used for displaying the text information and the identity information of the speaker. An image acquisition module is configured to acquire a scene image and at least two images of a same target at different positions in the same scene at a preset interval distance through a camera at a same fixed position. The camera is fixed at a certain height, and the scene image is acquired through the camera. The target is sequentially arranged at a first position and a second position, and images containing the complete length and / or width of the target are acquired respectively. A coordinate system generation module is configured to establish a 3D coordinate system according to the scene image, determine a horizon plane of the scene, measure a camera height value, and measure length and / or width data of the target. The 3D coordinate system is established according to a preset AR system and the scene image, and the ground of the scene is mapped as the horizon plane in the image. A data correction module is configured to correct the scene and a corresponding 3D coordinate system according to the camera height value, at least two length or width data of the target, and the horizon plane. The camera field of view angle is determined according to the camera height value. The proportional relationship between the real object and the image of the target is determined according to the length or width data of the target and the imaging of the target in the scene. The mapping relationship of each orthographic plane of the scene relative to the 3D coordinate system is determined according to the field of view angle and the horizon plane. The mapping corresponding relationship between the scene position and the 3D coordinate position is corrected according to the proportional relationship. A synchronous acquisition module is configured to acquire a video image and a voice signal synchronized with the video image when a participant appears in the scene. A position determination module is configured to determine the position of a current speaker in the video image based on the 3D coordinate system and according to the video image and the voice signal. An identity labeling module is configured to label the identity of the current speaker according to the position of the current speaker in the video image and the corresponding voice information.

6. An electronic device, comprising: One or more processors; One or more machine-readable media having instructions stored thereon that, when executed by the one or more processors, cause the electronic device to perform the method of any one of claims 1-4. The stored computer program causes the processor to perform the method of any one of claims 1-4.

7. A computer readable storage medium characterized in that, ​

Citation Information

Patent Citations

  • Video conference method and device and readable storage medium

    CN114125365A

  • Sound source positioning method and system for virtual video conference

    CN116405633A