Video calling method and device
By adjusting the gaze angle of the person in the display screen during a video call to align it with the eye area of the target person, the problem of not looking directly during a video call is solved, and the experience of eye contact and user experience is improved.
Patent Information
- Application Number
- CN202210151740.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-18
AI Technical Summary
In video calls, due to the deviation between the display screen and the camera, the person's eyes in the video screen do not look straight ahead, making it difficult to experience the same sight.
By obtaining the images taken by multiple cameras, determine the coordinates of the target's eye area, and adjust the angle of the view of the person in the display to align it with the target's eye area.
It realizes the experience of making both parties look at each other in video calls and improves the user experience.
Smart Images

Figure CN114531560B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of human-computer interaction technology, and in particular to a video calling method and device. Background Art
[0002] Video calls, also known as video telephones, usually refer to a communication method based on the Internet, which transmits people's voices and images in real time between display devices such as mobile phones or computers. When making a video call, due to the deviation between the display screen and the camera, when people look at the display screen, their line of sight will be at a certain angle to the camera's optical axis. Therefore, in the presented video screen, people's eyes are often not looking straight ahead, making it difficult for the two parties in the video call to have a sense of eye contact.
[0003] In order to solve the above problem, it is possible to store images of people at different angles, and generate an image of the person at a fixed sight angle through a neural network. The image of the person at the fixed sight angle can be synthesized with the image of the person recorded in real time to obtain an image of the person with a fixed sight angle to replace the image of the person presented during the video call.
[0004] However, the above-mentioned character picture used to generate a fixed sight angle is recorded in advance, and there is a difference between the recorded scene and the scene during the current video call. Therefore, the character picture with a fixed sight angle and the real-time recorded character picture cannot be perfectly synthesized together, resulting in abnormal pixel frames in the character picture presented during the video call, which is not conducive to user experience. Summary of the invention
[0005] In order to solve the technical problems caused by the above-mentioned video calls, the present disclosure is proposed. The embodiments of the present disclosure provide a video call method and device, which are used to solve the problem that it is difficult for the two parties in a video call to have a sense of eye contact. Specifically, the embodiments of the present disclosure provide the following technical solutions:
[0006] According to a first aspect of the present disclosure, a video calling method is provided, comprising:
[0007] Acquire a plurality of pictures taken by a plurality of cameras, wherein the plurality of pictures are pictures of a target area taken by the plurality of cameras at the same time, and the plurality of pictures at least include a first picture taken by a first camera and a second picture taken by a second camera;
[0008] Determine a first target person, where the first target person is a person in the target area;
[0009] Acquire first eye region coordinates according to the first picture and the second picture, where the first eye region coordinates are three-dimensional coordinates of the eye region of the first target person in a reference coordinate system of the first camera;
[0010] Determine a second target person image, where the second target person image is an image of a second target person displayed on the display screen, and the second target person is a person with whom the first target person makes a video call;
[0011] According to the second target person image, obtaining second eye region coordinates, where the second eye region coordinates are three-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the first camera;
[0012] According to the first eye region coordinates and the second eye region coordinates, the angle of the eye region image in the second target person image is adjusted to be oriented toward the eye region of the first target person.
[0013] According to a second aspect of the present disclosure, there is provided a video call device, comprising:
[0014] A plurality of cameras, for capturing a plurality of images, wherein the plurality of images are images of a target area captured by the plurality of cameras at the same time, and the plurality of images include a first image captured by a first camera and a second image captured by a second camera;
[0015] A display screen, used to present the images of both parties in the video call;
[0016] A first person acquisition module, used to determine a first target person, where the first target person is a person in the target area photographed by the multiple cameras;
[0017] a first parsing module, configured to obtain first eye region coordinates according to the first picture and the second picture taken by the first camera and the second camera, wherein the first eye region coordinates are three-dimensional coordinates of the eye region of the first target person in a reference coordinate system of the first camera;
[0018] A second person acquisition module, used to determine a second target person image, where the second target person image is an image of the second target person displayed on the display screen, and the second target person is the object of the video call with the first target person;
[0019] a second parsing module, configured to obtain second eye region coordinates according to the second target person image determined by the second person acquisition module, wherein the second eye region coordinates are three-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the first camera;
[0020] A sight redirection module is used to adjust the angle of the eye area image in the second target person image toward the eye area of the first target person according to the first eye area coordinates obtained by the first parsing module and the second eye area coordinates obtained by the second parsing module.
[0021] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the above-mentioned video call method.
[0022] According to a fourth aspect of the present disclosure, an electronic device is provided, the electronic device comprising
[0023] processor;
[0024] a memory for storing instructions executable by the processor;
[0025] The processor is used to read the executable instructions from the memory and execute the instructions to implement the above-mentioned video call method.
[0026] The present disclosure provides a video call method, device, computer-readable storage medium and electronic device, which adjust the sight angle of a person in a video screen having a video call with a video caller according to the sight angle of the video caller, so that both parties in the video have a sense of looking at each other, which is beneficial to the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and other purposes, features and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0028] Figure 1 is a schematic diagram of a scene structure provided by an embodiment of the present disclosure;
[0029] Figure 2 is a schematic diagram of a video call scenario provided by an embodiment of the present disclosure;
[0030] Figure 3 is a flowchart of a video calling method provided by an exemplary embodiment of the present disclosure;
[0031] Figure 4 is a flowchart of a step of determining a first target person provided by an exemplary embodiment of the present disclosure;
[0032] Figure 5is a flowchart of a step of obtaining the coordinates of the second mouth region provided by an exemplary embodiment of the present disclosure;
[0033] Figure 6 is a flowchart of another step of determining a first target person provided by an exemplary embodiment of the present disclosure;
[0034] Figure 7 is a flowchart of another step of determining a first target person provided by an exemplary embodiment of the present disclosure;
[0035] Figure 8 is a flowchart of a step of obtaining first eye region coordinates provided by an exemplary embodiment of the present disclosure;
[0036] Fig. 9 is a flowchart of a step of obtaining the coordinates of a second eye region provided by an exemplary embodiment of the present disclosure;
[0037] Fig.10 is a schematic diagram of a flow chart of adjusting the angle of the eyeball image in the second target person image provided by an exemplary embodiment of the present disclosure;
[0038] Fig.11 is a comparison diagram before and after adjusting the sight angle of the second target person image provided by the present disclosure;
[0039] Fig.12 is a structural schematic diagram of a video call device provided by an embodiment of the present disclosure;
[0040] Fig.13 is a structural diagram of a first character acquisition module provided by an exemplary embodiment of the present disclosure;
[0041] Fig.14 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0042] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described here.
[0043] Application Overview
[0044] When making a video call, due to the deviation between the display screen and the camera, the line of sight of a person looking at the display screen will be at a certain angle to the optical axis of the camera. Therefore, in the presented video, the person's eyes are often not looking straight ahead, making it difficult for the two parties in the video call to have a sense of eye contact. For example, when a driver is driving in a car, for driving safety, he usually looks straight ahead. Although his eyes occasionally look left and right, his eyes generally do not shift up and down. When the driver is making a video call with others through the video call system in the car (the video call system includes a display screen for displaying human images and a camera for capturing human images), because the driver is driving, his line of sight is often not towards the camera of the video call system. On the one hand, the line of sight of the human image captured by the camera is at a certain angle to the optical axis of the camera, so that the eyes of the person in the video screen received by the person who is having a video call with the driver are often not looking straight ahead, making it difficult to have a sense of eye contact. On the other hand, when the driver looks at the video call system, if the person who is having a video call with the driver does not look directly at the camera for capturing the image of the person having a video call with the driver, the line of sight of the human image received by the driver is also not looking straight ahead, making it difficult for the driver to have a sense of eye contact.
[0045] Based on the above technical problems, the present disclosure provides a video call system, method and device, which can adjust the sight angle of the person on the display screen according to the eye position of the person in front of the display screen, so that both parties in the video have a sense of looking at each other, which is beneficial to the user experience.
[0046] Exemplary Systems
[0047] See also Figure 1 , is a schematic diagram of a scenario structure provided by an embodiment of the present disclosure. The system includes: a display screen 100, a detector 200 and a server 300. The display screen 100 and the detector 200 are respectively connected to the server via a wireless network, for example, communicating with the server 300 via a gateway device.
[0048] The server 300 may be a network device. Optionally, the server 300 may also be a controller, a data center, a cloud platform, etc.
[0049] The wireless network can be any wireless communication system, such as a Long Term Evolution (LTE) system or a fifth generation mobile communication system (5G). In addition, it can also be applied to later communication systems, such as sixth generation and seventh generation mobile communication systems.
[0050] See also Figure 2 , which is a schematic diagram of a video call scenario provided in an embodiment of the present disclosure.
[0051] like Figure 2 As shown, the display screen 100 is used to receive and display the picture of the second user's environment transmitted by the server 300 when the first user has a video call with the second user. Specifically, the display screen 100 can be a display screen 100 with a physical screen such as a car display screen, a TV display screen, a wearable device display screen, and a mobile phone display screen, and the video picture can be directly displayed on such a display screen 100 with a physical screen; the display screen 100 can also be a virtual display screen without a physical screen, such as a head-up display (HUD) projected onto the windshield.
[0052] Further Figure 1 and 2 As shown, the detector 200 is used to collect signals of the external environment or external interaction. For example, the detector 200 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 200 includes an image collector, such as a camera 200a, which can be used to collect external environment scenes, user attributes or user interaction gestures, or, the detector 200 includes a sound collector, such as a microphone 200b, for receiving external sounds.
[0053] In an exemplary embodiment, the camera 200a-1 is used to capture the image of the first user's environment, and when the first user is having a video call with the second user, the captured image of the first user's environment is sent to the server 300, so that the second user's display screen 100b receives and displays the image of the first user's environment. Further, the camera 200a-2 is used to capture the image of the second user's environment, and when the first user is having a video call with the second user, the captured image of the second user's environment is sent to the server 300, so that the first user's display screen 100a receives and displays the image of the second user's environment.
[0054] The microphone 200b-1 is used to collect the voice of the first user, and when the first user has a video call with the second user, the collected voice of the first user is sent to the server 300 in real time, so that the second user can receive the voice of the first user. Furthermore, the microphone 200b-2 is used to collect the voice of the second user, and when the second user has a video call with the first user, the collected voice of the second user is sent to the server 300 in real time, so that the first user can receive the voice of the second user, thereby realizing a voice conversation between the first user and the second user.
[0055] The technical solution provided in this embodiment can be implemented by any method of software, hardware, or a combination of software and hardware. Among them, the hardware can provide sound and image input, the software can be implemented by C++ programming language, Java, etc., the video call function can be developed and implemented by Python-based programming voice, or can also be implemented by other software and hardware. This disclosure does not limit the specific hardware, software structure, and function of the implementation.
[0056] Exemplary Methods
[0057] Figure 3 It is a flowchart of a video calling method provided by an exemplary embodiment of the present disclosure.
[0058] This embodiment can be applied to electronic devices, and specifically to various electronic devices with video call functions. Figure 3 As shown, a video calling method provided by an exemplary embodiment of the present disclosure includes at least the following steps:
[0059] Step 101: Acquire multiple images captured by multiple cameras.
[0060] The multiple images are images of the target area captured by the multiple cameras at the same time, and the multiple images at least include a first image captured by the first camera and a second image captured by the second camera.
[0061] In one embodiment, the target area is the area where several people who are having a video call in front of the display screen are located. These people who are having a video call in front of the display screen can be called speakers. The display screen can be a display screen with a physical screen such as a car display screen, a TV display screen, a wearable device display screen, and a mobile phone display screen. The display screen can also be a virtual display screen without a physical screen that is projected onto the windshield such as a head-up display (HUD). The display screen is used to display the image of the person who is having a conversation with the speaker. The person who is having a conversation with the speaker displayed on the display screen is called the interlocutor. The speaker can obtain his / her video image through any camera, and transmit the video image to the interlocutor's calling device so that the interlocutor can see the speaker's video image on his / her display screen. In addition, the speaker and the interlocutor can communicate with each other through their respective microphones.
[0062] In one embodiment, multiple pictures can be obtained by shooting with multiple cameras, and two of the multiple pictures are selected, one of which is determined as the first picture, and the camera that shoots the first picture is determined as the first camera, and the other is determined as the second picture, and the camera that shoots the second picture is determined as the second camera, wherein both the first picture and the second picture at least include a facial image of a speaker, and since the first camera and the second camera are located at different positions, the eyeballs of the same speaker in the first picture and the second picture are directed in different angles.
[0063] Step 102: Determine the first target person.
[0064] The first target person is a person in the target area.
[0065] In one embodiment, the first target person is a speaker who is speaking in front of a display screen, and both the first screen and the second screen include a facial image of the first target person. The facial image of the first target person may include the eye area, mouth area, nose area, ear area, etc. of the first target person.
[0066] Step 103: Obtain first eye region coordinates according to the first picture and the second picture.
[0067] The first eye region coordinates are three-dimensional coordinates of the eye region of the first target person in the reference coordinate system of the first camera.
[0068] In one embodiment, the reference coordinate system of the first camera is a three-dimensional coordinate system established with the first camera as the coordinate origin. The plane formed by the x-axis and y-axis of the three-dimensional coordinate system is the horizontal plane, and the z-axis is perpendicular to the horizontal plane. The first eye area coordinates are the coordinates of the eye area of the first target person in the coordinate system with the first camera as the coordinate origin.
[0069] The coordinates of the eye region of the first target person may be the coordinates of key points of the entire eye region of the first target person obtained from the first picture and the second picture in the reference coordinate system of the first camera.
[0070] Step 104: Determine a second target person image.
[0071] The second target person image is an image of a second target person displayed on a display screen, and the second target person is the object of the video call with the first target person.
[0072] In some embodiments, when the display screen is a virtual display screen without a physical screen, such as a head-up display screen, the picture displayed on the virtual display screen can be scanned to obtain the positional relationship of each pixel in the picture displayed on the display screen and the filled color information, and based on the obtained positional relationship of each pixel and the filled color information, the second target person image can be determined.
[0073] In some embodiments, when the display screen is a display screen with a physical screen, the second target person image can be directly obtained based on a program for instructing the display screen to display an image.
[0074] In one embodiment, the speaker can see character images of several interlocutors on the display screen, and the second target character image is the character image of one or several interlocutors seen by the speaker on the display screen.
[0075] Step 105: Acquire the second eye region coordinates according to the second target person image.
[0076] The second eye region coordinates are three-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the first camera.
[0077] Specifically, the second target person image at least includes the facial image of the second target person, and the facial image of the second target person may include the eye area, mouth area, nose area, ear area, etc. of the second target person. The second eye area coordinates are the coordinates of the eye area of the second target person in a coordinate system with the first camera as the coordinate origin.
[0078] In some embodiments, when the display screen is a display screen with a physical screen, the picture displayed on the display screen can be directly obtained through a program for instructing the display screen to display a picture, and the obtained picture displayed on the display screen can be controlled to be input into a deep neural network model for face detection (face key point detection), and the eye key points of the entire eye area of the second target person image in the picture displayed on the display screen are detected to obtain the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the display screen, and combined with the positional relationship between the display screen and the first camera, the three-dimensional coordinates of the eye key points of the entire eye area in the second target person image in the reference coordinate system of the first camera are obtained, that is, the second eye area coordinates.
[0079] In some embodiments, when the display screen is a virtual display screen without a physical screen, the picture displayed on the virtual display screen can be scanned by a scanner. The scanner can include a scan line algorithm application. When the scanner scans the picture displayed on the virtual display screen, the scan line algorithm application is triggered to start, and the scanned picture is analyzed by the scan line algorithm to obtain the positional relationship of each pixel in the picture displayed on the display screen and the filled color information, and control the input of the obtained positional relationship of each pixel and the filled color information into a deep neural network model for face detection (face key point detection). The deep neural network model can restore the picture displayed on the virtual display screen based on the input positional relationship of each pixel and the filled color information, and detect the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the restored picture in the reference coordinate system of the display screen, and combine the positional relationship between the display screen and the first camera to obtain the three-dimensional coordinates of the eye key points of the entire eye area in the second target person image in the reference coordinate system of the first camera, that is, the second eye area coordinates.
[0080] Step 106: According to the first eye region coordinates and the second eye region coordinates, adjust the angle of the eye region image in the second target person image to be oriented toward the eye region of the first target person.
[0081] In one embodiment, the angle of the eye area image in the second target person image displayed on the display screen is adjusted according to the first eye area coordinates and the second eye area coordinates, so that the speaker can have a visual experience with the second target person image displayed on the screen.
[0082] In this embodiment, the first target person is determined based on the images of the target area captured at different angles at the same time, and the first eye area coordinates corresponding to the first target person are obtained; according to each second target person image displayed on the display screen, the second eye area coordinates corresponding to each second target person image are obtained, and according to the first eye area coordinates and the second eye area coordinates, the angle of the eye area image in each second target person image is adjusted to be toward the eye area of the first target person, so that the first target person has a sense of eye contact when looking at the second target person image displayed on the display screen, which is beneficial to user experience.
[0083] Figure 4 Shown as Figure 3 The flowchart of the step of determining the first target person in the embodiment shown is as follows.
[0084] like Figure 4 As shown in the above Figure 3Based on the illustrated embodiment, in an exemplary embodiment of the present disclosure, the step of determining the first target person shown in step 102 may specifically include the following steps:
[0085] Step 201: Acquire the first mouth region coordinates of each person in the target region according to the first picture and the second picture.
[0086] The first mouth area coordinates are three-dimensional coordinates of the mouth area of the person in the target area in the reference coordinate system of the first camera.
[0087] In some embodiments, the third mouth area coordinates of each person in the target area can be obtained based on the person image of each person in the target area in the first picture, wherein the third mouth area coordinates are the two-dimensional coordinates of the mouth area image of the person in the target area in the first picture in the reference coordinate system of the first picture, and the fourth mouth area coordinates of each person in the target area can be obtained based on the person image of each person in the target area in the second picture, wherein the fourth mouth area coordinates are the two-dimensional coordinates of the mouth area image of the person in the target area in the second picture in the reference coordinate system of the second picture. Based on the obtained third mouth area coordinates and fourth mouth area coordinates, the first mouth area coordinates of each person in the target area in the reference coordinate system of the first camera can be obtained based on the binocular positioning principle.
[0088] Step 202: localize the sound source according to the sound of the person in the target area collected by the microphone array, and obtain the coordinates of the second mouth area.
[0089] The second mouth area coordinates are the three-dimensional coordinates of the person corresponding to the sound source position of the sound collected by the microphone array in the reference coordinate system of the first camera.
[0090] In some embodiments, the intensities of the sounds collected by each microphone and emitted by each person in the target area can be compared, and the sound with the maximum intensity can be matched with the frequency of the sound source to locate the distance between the sound source and each microphone, thereby locating the position of the sound source and obtaining the corresponding second mouth area coordinates.
[0091] Specifically, taking the number of microphones as 2 as an example, namely microphone A and microphone B, if there are 3 people in the target area making sounds at the same time, microphone A and microphone B respectively collect the sounds made by the 3 people, which can be recorded as (A1, B1), (A2, B2) and (A3, B3). If, after comparison, it is found that (A2, B2) has the largest sound intensity, the sound corresponding to the collected (A2, B2) is matched with the frequency of the sound source. For example, different sound sources such as male voice, female voice and child's voice have different frequencies, so as to locate the distance between the sound source and each microphone, thereby accurately locating the position of the sound source, and based on the positional relationship between microphone A and microphone B and the first camera, the coordinates of the second mouth area in the reference coordinate system with the first camera as the coordinate origin are obtained.
[0092] Step 203: Determine the person corresponding to the first mouth area coordinate that is closest in straight-line distance to the second mouth area coordinate as the first target person.
[0093] Specifically, since the first mouth area coordinate and the second mouth area coordinates are all in the reference coordinate system of the first camera, the distance between the second mouth area coordinate and each of the first mouth area coordinates can be obtained based on the distance formula between two points in three-dimensional space, and the person corresponding to the first mouth area coordinate that is closest in straight-line distance to the second mouth area coordinate is determined as the first target person, wherein the distance formula between two points in three-dimensional space is:
[0094]
[0095] Figure 5 It shows that Figure 4 A schematic diagram of the flow of the step of obtaining the coordinates of the second mouth area in the embodiment shown.
[0096] like Figure 5 As shown in the above Figure 4 Based on the illustrated embodiment, in an exemplary embodiment of the present disclosure, the step of obtaining the coordinates of the second mouth region shown in step 202 may specifically include the following steps:
[0097] Step 301: Acquire first conversion parameters, where the first conversion parameters are used to convert coordinates in the reference coordinate system of the microphone into coordinates in the reference coordinate system of the first camera.
[0098] Specifically, any microphone in the microphone array is selected as the main microphone, and the conversion relationship between the reference coordinate system of the main microphone and the reference coordinate system of the first camera is calibrated to obtain the first conversion parameter. For example, microphone M is selected as the main microphone, and the conversion formula P A =R MA ·P M, obtain the first conversion parameter R from the reference coordinate system of the microphone M to the reference coordinate system of the first camera MA , where P A and P M represents the coordinates of any point P in the world coordinate system in the reference coordinate system of the first camera and the reference coordinate system of the microphone M, such as represents the coordinates of point P in the reference coordinate system of the first camera, represents the coordinates of the point P in the reference coordinate system of the microphone M. The reference coordinate system of the microphone M is a three-dimensional coordinate system established with the position of the microphone M as the coordinate origin.
[0099] Step 302: Acquire the coordinates of the second mouth region according to the coordinates of the sound collected by the microphone in the reference coordinate system of the microphone and the first conversion parameters.
[0100] Specifically, according to the distance between the person corresponding to the sound source position of the sound with the maximum sound intensity and the microphone M, the coordinates of the person corresponding to the sound source position of the sound with the maximum sound intensity in the reference coordinate system of the microphone M are obtained, and according to the first conversion parameter R MA , convert the coordinates of the person corresponding to the sound source position of the sound with the maximum sound intensity in the reference coordinate system of the microphone M into coordinates in the reference coordinate system of the first camera, wherein the coordinates of the person corresponding to the sound source position of the sound with the maximum sound intensity in the reference coordinate system of the first camera are the second mouth area coordinates.
[0101] Figure 6 It shows that Figure 4 The illustrated embodiment is a flow chart of the process when the step of obtaining the coordinates of the second mouth region cannot be implemented.
[0102] like Figure 6 As shown in the above Figure 4 Based on the illustrated embodiment, in an exemplary embodiment of the present disclosure, if the step of obtaining the coordinates of the second mouth region shown in step 202 cannot be implemented, the following steps may also be included:
[0103] Step 401: if the coordinates of the second mouth area cannot be obtained based on the sound collected by the microphone array, the facial contours of all the person images in the first picture are identified, and face detection frames corresponding to the number of the identified facial contours are generated.
[0104] Specifically, when no person in the target area makes any sound, or when there is noise in the target area so that the microphone cannot clearly collect the sound made by the person in the target area, the second mouth area coordinates cannot be obtained.
[0105] When the coordinates of the second mouth area cannot be obtained, the facial contours of all the character images in the first picture can be identified, and based on the highest point, the lowest point, the leftmost point and the rightmost point of the facial contour of each character image, a face detection frame corresponding to the number of identified facial contours is generated. The face detection frame can be a rectangle, and the lengths of two adjacent sides of the rectangle correspond to the distance between the highest point to the lowest point and the distance between the leftmost point to the rightmost point of the facial contour, respectively.
[0106] Step 402: determining the face contour corresponding to the face detection frame having the largest area as the target face contour;
[0107] Specifically, the area of each generated face detection frame is calculated according to the length parameters of the two adjacent sides of each face detection frame, and the face contour corresponding to the face detection frame with the largest area is determined as the target face contour. For example, if there are three people in the target area, when the coordinates of the second mouth area cannot be obtained, three face detection frames corresponding to the three people will be generated, namely face detection frame A, face detection frame B and face detection frame C. The length parameters of any two adjacent sides of each face detection frame are obtained respectively, which are recorded as (A 11 , A 12 )、(B 11 , B 12 ) and (C 11 , C 12 ), calculate A 11 ×A 12 , B 11 ×B 12 and C 11 ×C 12 The value of 11 ×A 12 >B 11 ×B 12 >C 11 ×C 12 , then the face detection frame A is determined as the target face contour.
[0108] Step 403: Determine the person corresponding to the target facial contour as the first target person.
[0109] Specifically, a person image corresponding to the target facial contour in the first picture is obtained, and the person corresponding to the person image is determined as the first target person.
[0110] Optional, Figure 7 It shows that Figure 3 The flowchart of the step of determining the first target person in the embodiment shown is as follows.
[0111] like Figure 7 As shown in the above Figure 3 On the basis of the illustrated embodiment, in an exemplary embodiment of the present disclosure, if there are two or more first mouth region coordinates that are closest in straight-line distance to the second mouth region coordinates, the following steps may also be included:
[0112] Step 501: Identify the facial contours of all persons corresponding to the first mouth area coordinates that are closest in linear distance to the second mouth area coordinates in the first picture, and generate face detection frames corresponding to the number of the identified facial contours.
[0113] Step 502: Determine the facial contour corresponding to the face detection frame having the largest area as the target facial contour.
[0114] Step 503: Determine the person corresponding to the target facial contour as the first target person.
[0115] In this embodiment, the generation of the face detection frame, the method of determining the target face contour, and the method of determining the first target person according to the target face contour can refer to the above steps 401 to 403, which will not be repeated here.
[0116] Figure 8 Shown as Figure 3 A schematic diagram of a flow chart of the step of obtaining the first eye region coordinates in the embodiment shown.
[0117] like Figure 8 As shown in the above Figure 3 Based on the illustrated embodiment, in an exemplary embodiment of the present disclosure, the step of obtaining the first eye region coordinates shown in step 103 may specifically include the following steps:
[0118] Step 601: Acquire second conversion parameters, where the second conversion parameters are used to convert coordinates in the reference coordinate system of the first camera into coordinates in the reference coordinate system of the second camera.
[0119] Specifically, the conversion relationship between the reference coordinate system of the first camera and the reference coordinate system of the second camera can be calibrated to obtain the second conversion parameter, and the second conversion parameter includes the rotation parameter R BA and the translation parameter T BA , for example, according to the conversion formula P A =R BA ·P B +T BA , get the rotation parameter R from the reference coordinate system of the first camera to the reference coordinate system of the second camera BA and the translation parameter T BA , where P A and P BRepresents the coordinate values of any point P in the world coordinate system in the reference coordinate system of the first camera and the reference coordinate system of the second camera, such as represents the coordinates of point P in the reference coordinate system of the first camera, represents the coordinates of point P in the reference coordinate system of the second camera.
[0120] Step 602: Obtain third eye region coordinates, where the third eye region coordinates are two-dimensional coordinates of the eye region image of the first target person in the first picture in the reference coordinate system of the first picture.
[0121] Specifically, a point in the first picture is used as the coordinate origin to establish a two-dimensional reference coordinate system in the first picture. For example, a certain endpoint in the first picture is used as the coordinate origin, and the boundaries of two adjacent target pictures passing through the coordinate origin are respectively determined as the X axis and the Y axis to establish a two-dimensional reference coordinate system in the first picture, wherein the coordinate origin in the first picture is defined as P10 = [0, 0, 0] T , based on the eye area image of the first target person in the first picture and the coordinate origin P10 = [0,0,0] T The distance can be used to obtain the two-dimensional coordinates of the eye area image of the first target person in the first picture in the reference coordinate system of the first picture, and the two-dimensional coordinates of the eye area image of the first target person in the first picture in the reference coordinate system of the first picture are determined as the third eye area coordinates.
[0122] Step 603: Acquire fourth eye region coordinates, where the fourth eye region coordinates are two-dimensional coordinates of the eye region image of the first target person in the second picture in the reference coordinate system of the second picture.
[0123] Specifically, a point in the second picture is used as the coordinate origin to establish a two-dimensional reference coordinate system in the second picture. For example, a certain endpoint in the second picture is used as the coordinate origin, and the boundaries of two adjacent target pictures passing through the coordinate origin are respectively determined as the X axis and the Y axis to establish a two-dimensional reference coordinate system in the second picture, wherein the coordinate origin in the second picture is defined as P20 = [0, 0, 0] T , based on the eye area image of the first target person in the second picture and the coordinate origin P20 = [0,0,0] T The distance can be used to obtain the two-dimensional coordinates of the eye area image of the first target person in the second picture in the reference coordinate system of the second picture, and the two-dimensional coordinates of the eye area image of the first target person in the second picture in the reference coordinate system of the second picture are determined as the fourth eye area coordinates.
[0124] Step 604: Acquire the first eye area coordinates according to the third eye area coordinates, the fourth eye area coordinates and the second conversion parameters.
[0125] Specifically, the third eye region coordinate [x1, y1] is converted to the corresponding coordinate in the reference coordinate system of the first camera: The fourth eye region coordinate [x2, y2] is converted to the corresponding coordinate in the reference coordinate system of the second camera: Among them, K1 represents the camera intrinsic parameter of the first camera, and K2 represents the camera intrinsic parameter of the second camera. It should be noted that the camera intrinsic parameter is the essential parameter of the camera, which is fixed after the camera leaves the factory and can be directly obtained.
[0126] After the third eye region coordinate [x1, y1] is converted into the corresponding coordinate P11 in the reference coordinate system of the first camera and the fourth eye region coordinate [x2, y2] is converted into the corresponding coordinate P21 in the reference coordinate system of the second camera, based on the rotation parameter R BA and the translation parameter T BA , the coordinates in the reference coordinate system of the second camera can be converted to the coordinates in the reference coordinate system of the first camera through rotation and translation. The conversion formula is as follows:
[0127] P120=R BA *P20+T BA ,
[0128] P121=R BA *P21+T BA ;
[0129] Among them, P120 is the origin coordinate P20 in the reference coordinate system of the second camera converted to the coordinate P120 in the reference coordinate system of the first camera through rotation and translation, and P121 is the coordinate P21 converted to the coordinate P120 in the reference coordinate system of the first camera through rotation and translation.
[0130] Based on P120 and P121 calculated above, two straight line equations in the reference coordinate system of the first camera can be defined to obtain the equation group:
[0131]
[0132] Among them, d1 and d2 are two different one-dimensional variables, and P1 p =P2 p , solve the values of d1 and d2 in the above equations, and substitute the obtained values of d1 and d2 into the above equations to obtain P1 p (or P2 p ) is expressed in the form of coordinates, P1 p (or P2p ) is the first eye area coordinate.
[0133] Fig. 9 It shows that Figure 3 A schematic diagram of the flow of the step of obtaining the second eye region coordinates in the embodiment shown.
[0134] like Fig. 9 As shown in the above Figure 3 Based on the illustrated embodiment, in an exemplary embodiment of the present disclosure, the step of obtaining the second eye region coordinates shown in step 105 may specifically include the following steps:
[0135] Step 701: Acquire a third conversion parameter, where the third conversion parameter is used to convert coordinates in the reference coordinate system of the display screen into coordinates in the reference coordinate system of the first camera;
[0136] Specifically, the conversion relationship between the reference coordinate system of the display screen and the reference coordinate system of the first camera can be calibrated to obtain the third conversion parameter, which includes the rotation parameter R SA and translation parameter T SA For example, according to the conversion formula P A =R SA ·P S +T SA , obtain the rotation parameter R from the reference coordinate system of the display screen to the reference coordinate system of the first camera SA and translation parameter T SA , where P A and P S represents the coordinate values of any point P in the world coordinate system in the reference coordinate system of the first camera and the reference coordinate system of the display screen, such as represents the coordinates of point P in the reference coordinate system of the first camera, It represents the coordinates of point P in the reference coordinate system of the display screen. The reference coordinate system of the display screen is a three-dimensional coordinate system established with the position of the display screen as the coordinate origin.
[0137] Step 702: Acquire fifth eye region coordinates, where the fifth eye region coordinates are two-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the display screen.
[0138] In some embodiments, when the display screen is a virtual display screen without a physical screen, such as a head-up display screen, the virtual display screen can be scanned by a scanner, and the scanner includes a scan line algorithm application. When the scanner scans the virtual display screen, the scan line algorithm application is triggered to start, and the scanned image is analyzed by the scan line algorithm to obtain the positional relationship of each pixel in the image displayed by the scanned virtual display screen and the filled color information, and control the input of the obtained positional relationship of each pixel and the filled color information into a deep neural network model for face detection (face key point detection). The deep neural network model can restore the image displayed by the virtual display screen based on the input positional relationship of each pixel and the filled color information, and establish a two-dimensional reference coordinate system in the restored image, and obtain the position of the second target person image in the two-dimensional reference coordinate system through detection, and obtain the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system, wherein the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system are the fifth eye area coordinates. Specifically, the pixel point corresponding to the first pixel point information input into the deep neural network model can be used as the coordinate origin, and any two mutually perpendicular directions passing through the coordinate origin are respectively determined as the X-axis and the Y-axis, so as to establish a two-dimensional reference coordinate system in the picture displayed on the virtual display screen restored in the deep neural network model, wherein the coordinate origin in the restored picture can be defined as P30 = [0,0,0] T , based on the eye area image of the second target person in the restored image and the coordinate origin P30 = [0,0,0] T The distance can be used to obtain the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the restored image, that is, the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the virtual display screen can be obtained, and the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the virtual display screen are determined as the fifth eye area coordinates.
[0139] In some embodiments, when the display screen is a display screen with a physical screen, the screen displayed on the display screen can be directly obtained through a program for indicating the screen display screen to display a screen, and the obtained screen displayed on the display screen can be controlled to be input into a deep neural network model for face detection (face key point detection), and the eye key points of the entire eye area of the second target person image in the screen displayed on the display screen are detected to obtain the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the display screen, and the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the display screen are determined as the fifth eye area coordinates. Specifically, a certain point (or pixel point) in the screen displayed on the display screen can be used as the coordinate origin to establish a two-dimensional reference coordinate system in the display screen. For example, a certain point on the display screen is used as the coordinate origin, and two adjacent boundaries passing through the coordinate origin are respectively determined as the X axis and the Y axis to establish a two-dimensional reference coordinate system in the display screen, wherein the coordinate origin in the display screen is defined as P30 = [0,0,0] T , based on the eye area image of the second target person on the display screen and the coordinate origin P30 = [0,0,0] T The distance can be used to obtain the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the physical display screen, and the two-dimensional coordinates of the eye key points of the entire eye area of the second target person image in the reference coordinate system of the physical display screen are determined as the fifth eye area coordinates.
[0140] Step 703: Acquire the second eye region coordinates according to the fifth eye region coordinates and the third conversion parameter.
[0141] Specifically, according to the fifth eye region coordinates, a three-dimensional representation P of the fifth eye region coordinates in the reference coordinate system of the display screen can be obtained. S , and obtain the second eye area coordinates according to the three-dimensional representation of the fifth eye area coordinates in the reference coordinate system of the display screen and the third conversion parameter. For example, the coordinates of the fifth eye area coordinates in the reference coordinate system of the display screen are [x3, y3], then the three-dimensional representation of the fifth eye area coordinates in the reference coordinate system of the display screen can be P S =[x3,y3,z3], where z3=0. And according to P A =R SA ·P S +T SA , converting the fifth eye area coordinates in the reference coordinate system of the display screen into coordinates in the reference coordinate system of the first camera, and converting the fifth eye area coordinates in the reference coordinate system of the display screen into coordinates in the reference coordinate system of the first camera to be determined as second eye area coordinates.
[0142] Fig.10 Shown as Figure 3 The illustrated embodiment is a schematic flow chart of the step of adjusting the angle of the eye region image in the second target person image to be oriented toward the eye region of the first target person.
[0143] like Fig.10 As shown in the above Figure 3 On the basis of the illustrated embodiment, in an exemplary embodiment of the present disclosure, the first eye region coordinates include first left eye region coordinates and first right eye region coordinates, the second eye region coordinates include second left eye region coordinates and second right eye region coordinates, and the step of adjusting the angle of the eye region image in the second target person image toward the eye region of the first target person according to the first eye region coordinates and the second eye region coordinates shown in step 106 may specifically include the following steps:
[0144] Step 801: adjusting the angle of the left eye area image in the second target person image according to the difference between the first left eye area coordinates and the second left eye area coordinates.
[0145] Specifically, based on the gaze redirection algorithm, the angle of the left eye area image in the second target person image is adjusted to face the left eye area of the first target person, so as to achieve the effect of "eye contact". Second left eye area coordinates The Euler angle that needs to be adjusted for the left eye area image in the second target person's image can be calculated. The Euler angle includes the pitch angle pitch and the yaw angle yaw. According to the obtained pitch angle pitch and yaw, the angle of the left eye area image in the second target person's image is adjusted so that the angle of the left eye area image in the second target person's image is toward the left eye area of the first target person, so as to achieve a "eye-to-eye" effect.
[0146] The calculation formulas for pitch angle pitch and yaw angle yaw are:
[0147] pitch=arcsin(y),
[0148]
[0149] in,
[0150]
[0151] Step 802: adjusting the angle of the right eye area image in the second target person image according to the difference between the first right eye area coordinates and the second right eye area coordinates.
[0152] Specifically, based on the gaze redirection algorithm, the angle of the right eyeball area image in the second target person image is adjusted to face the right eyeball area of the first target person, so as to achieve the effect of "eye contact". Second right eye area coordinates Then the Euler angle that needs to be adjusted for the right eye area image in the second target person's image can be calculated. The Euler angle includes the pitch angle pitch and the yaw angle yaw. According to the obtained pitch angle pitch and yaw, the angle of the right eye area image in the second target person's image is adjusted so that the angle of the right eye area image in the second target person's image is toward the right eye area of the first target person, so as to achieve the effect of "looking into each other's eyes".
[0153] The calculation formulas for pitch angle pitch and yaw angle yaw are:
[0154] pitch=arcsin(y),
[0155]
[0156] in,
[0157]
[0158] For example, see Fig.11 , which is a comparison diagram before and after adjusting the sight angle of the second target person image provided by the present disclosure. Fig.11 As shown, Figure a is the picture before adjusting the sight line angle of the second target person image. In Figure a, each person image displayed on the display screen can be the second target person image. The second eye area coordinates of each second target person image are calculated, and according to the first eye area coordinates of the first target person, the pitch angle and yaw angle that need to be adjusted for the left eye area image and the right eye area image in the second target person image are calculated respectively, and according to the calculated pitch angle pitch and yaw angle yaw, the eye area angles of the left eye area image and the right eye area image in the second target person image are adjusted. The sight line angles of the second target person images after adjustment are shown in Figure b, and the sight line angles of each second target person image are toward the first target person.
[0159] Exemplary Devices
[0160] See also Fig.12, is a schematic diagram of the structure of a video call device provided in an embodiment of the present disclosure, and the device is used to implement all or part of the functions of the aforementioned method embodiment. Specifically, the video call device includes a camera 111, a display screen 112, a first person acquisition module 113, a first parsing module 114, a second person acquisition module 115, a second parsing module 116, and a sight redirection module 117, etc. In addition, the device may also include other more modules, such as a storage module, a sending module, etc., which are not limited in this embodiment.
[0161] Specifically, the number of cameras 111 in the embodiment of the present disclosure is at least 2, which are used to capture images in the target area. The multiple images are images of the target area captured by the multiple cameras 111 at the same time, and the multiple images include a first image captured by the first camera and a second image captured by the second camera.
[0162] The display screen 112 is used to present the images of both parties in the video call.
[0163] The first person acquisition module 113 is used to determine a first target person.
[0164] The first target person is a person in the target area photographed by the multiple cameras 111.
[0165] The first parsing module 114 is configured to obtain first eye region coordinates according to the first image and the second image captured by the first camera and the second camera.
[0166] The first eye region coordinates are three-dimensional coordinates of the eye region of the first target person in the reference coordinate system of the first camera.
[0167] The second person acquisition module 115 is used to determine a second target person image.
[0168] The second target person image is an image of a second target person displayed on the display screen 112 , and the second target person is the object of the video call with the first target person.
[0169] The second parsing module 116 is used to obtain the second eye region coordinates according to the second target person image determined by the second person obtaining module 115 .
[0170] The second eye region coordinates are three-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the first camera.
[0171] The sight redirection module 117 is used to adjust the angle of the eye area image in the second target person image displayed on the display screen 112 toward the eye area of the first target person according to the first eye area coordinates obtained by the first analysis module 116 and the second eye area coordinates obtained by the second analysis module.
[0172] Among them, optional, such as Fig.13 As shown, the first person acquisition module 113 also includes a sound collection module 118 .
[0173] Optionally, in an implementation of this embodiment, when the first person acquisition module 113 determines the first target person, the first target person may be obtained through the sound collection module 118 of the microphone / microphone array.
[0174] Furthermore, the sound collection module 118 is used to collect the sound emitted by each person in the target area, and locate the sound source according to the collected sounds of the people in the target area, and obtain the three-dimensional coordinates of the person corresponding to the sound source position in the reference coordinate system of the first camera.
[0175] Optionally, in another implementation of the present embodiment, determining the first target person includes: obtaining first mouth area coordinates of each person in the target area according to the first picture and the second picture, the first mouth area coordinates being the three-dimensional coordinates of the mouth of the person in the target area in the reference coordinate system of the first camera; performing sound source positioning according to the sound of the person in the target area collected by the microphone array, and obtaining second mouth area coordinates, the second mouth area coordinates being the three-dimensional coordinates of the person corresponding to the sound source position of the sound collected by the microphone array in the reference coordinate system of the first camera;
[0176] A person corresponding to a first mouth area coordinate that is closest in straight-line distance to the second mouth area coordinate is determined as a first target person.
[0177] Optionally, in another implementation of the present embodiment, obtaining the second mouth area coordinates based on the sound collected by the microphone array also includes: obtaining a first conversion parameter, wherein the first conversion parameter is used to convert the coordinates in the reference coordinate system of the microphone array into the coordinates in the reference coordinate system of the first camera; obtaining the second mouth area coordinates based on the coordinates of the sound collected by the microphone array in the reference coordinate system of the microphone array and the first conversion parameter.
[0178] Alternatively, in another implementation of the present embodiment, based on the sound collected by the microphone array, if the coordinates of the second mouth area cannot be obtained, the facial contours of all the character images in the first picture are identified, and a number of face detection frames corresponding to the identified facial contours are generated; the facial contour corresponding to the face detection frame with the largest area is determined as the target facial contour; and the person corresponding to the target facial contour is determined as the first target person.
[0179] Optionally, in another implementation of the present embodiment, obtaining first eye area coordinates according to the first picture and the second picture includes: obtaining second conversion parameters, the second conversion parameters being used to convert coordinates in a reference coordinate system of the first camera into coordinates in a reference coordinate system of the second camera; obtaining third eye area coordinates, the third eye area coordinates being two-dimensional coordinates of an eye area image of the first target person in the first picture in the reference coordinate system of the first picture; obtaining fourth eye area coordinates, the fourth eye area coordinates being two-dimensional coordinates of an eye area image of the first target person in the second picture in the reference coordinate system of the second picture; and obtaining the first eye area coordinates according to the third eye area coordinates, the fourth eye area coordinates and the second conversion parameters.
[0180] Optionally, in another implementation of the present embodiment, obtaining second eye area coordinates according to the second target person image includes: obtaining third conversion parameters, the third conversion parameters being used to convert coordinates in the reference coordinate system of the display screen into coordinates in the reference coordinate system of the first camera; obtaining fifth eye area coordinates, the fifth eye area coordinates being two-dimensional coordinates of the eye area image in the second target person image in the reference coordinate system of the display screen; and obtaining the second eye area coordinates according to the fifth eye area coordinates and the third conversion parameters.
[0181] Optionally, in another implementation of the present embodiment, the first eye area coordinates include first left eye area coordinates and first right eye area coordinates, the second eye area coordinates include second left eye area coordinates and second right eye area coordinates, and adjusting the angle of the eye area image in the second target person image toward the eye area of the first target person according to the first eye area coordinates and the second eye area coordinates includes: adjusting the angle of the left eye area image in the second target person image according to the difference between the first left eye area coordinates and the second left eye area coordinates; adjusting the angle of the right eye area image in the second target person image according to the difference between the first right eye area coordinates and the second right eye area coordinates.
[0182] In addition, in the embodiment of the present device, Fig.12 The functions of the modules shown are similar to those described above. Figure 3 The method embodiments shown correspond to, for example, multiple cameras being used to execute the aforementioned method step 101, or the first character acquisition module being used to execute the aforementioned method step 102, the first analysis module being used to execute the aforementioned method step 103, the second character acquisition module and the display screen being used to execute the aforementioned method step 104, the second analysis module being used to execute the aforementioned method step 105, and the line of sight redirection module being used to execute the aforementioned method step 106, etc.
[0183] Exemplary Electronic Devices
[0184] Below, reference Fig.12 The electronic device 10 may be one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device may communicate with the first device and the second device to receive the collected input signals from them.
[0185] Fig.14 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0186] like Fig.14 As shown, the electronic device 10 includes one or more processors 11 and a memory 12 .
[0187] The processor 11 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0188] The memory 12 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute the program instructions to implement the steps in the video call method of each embodiment of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage medium.
[0189] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0190] For example, when the electronic device is the first device or the second device, the input device 13 may be the microphone or microphone array described above, for capturing input signals from a sound source. When the electronic device is a stand-alone device, the input device 13 may be a communication network connector, for receiving collected input signals from the first device and the second device.
[0191] In addition, the input device 13 may also include, for example, a keyboard, a mouse, and the like.
[0192] The output device 14 can output various information to the outside, including determined distance information, direction information, etc. The output device 14 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.
[0193] Of course, to simplify, Fig.13 Only some of the components related to the present disclosure in the electronic device 10 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application situations, the electronic device 10 may also include any other appropriate components.
[0194] Exemplary computer program products and computer-readable storage media
[0195] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the video call method according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.
[0196] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0197] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the video call method according to various embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.
[0198] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0199] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0200] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0201] It should also be noted that in the apparatus, device and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0202] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0203] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A video calling method, comprising: Acquire a plurality of pictures taken by a plurality of cameras, wherein the plurality of pictures are pictures of a target area taken by the plurality of cameras at the same time, and the plurality of pictures at least include a first picture taken by a first camera and a second picture taken by a second camera; Determine a first target person, where the first target person is a person in the target area; Acquire first eye region coordinates according to the first picture and the second picture, where the first eye region coordinates are three-dimensional coordinates of the eye region of the first target person in a reference coordinate system of the first camera; Determine a second target person image, where the second target person image is an image of a second target person displayed on the display screen, and the second target person is a person with whom the first target person makes a video call; According to the second target person image, obtaining second eye region coordinates, where the second eye region coordinates are three-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the first camera; According to the first eye region coordinates and the second eye region coordinates, adjusting the angle of the eye region image in the second target person image to be oriented toward the eye region of the first target person; Wherein, determining the first target person includes: Acquire first mouth region coordinates of each person in the target area according to the first picture and the second picture, where the first mouth region coordinates are three-dimensional coordinates of the mouth region of the person in the target area in the reference coordinate system of the first camera; Performing sound source positioning according to the sound of the person in the target area collected by the microphone array, and obtaining second mouth area coordinates, where the second mouth area coordinates are three-dimensional coordinates of the person corresponding to the sound source position of the sound collected by the microphone array in the reference coordinate system of the first camera; A person corresponding to a first mouth area coordinate that is closest in straight-line distance to the second mouth area coordinate is determined as a first target person.
2. The method according to claim 1, wherein: The step of obtaining the coordinates of the second mouth area according to the sound collected by the microphone array further includes: Acquire a first conversion parameter, where the first conversion parameter is used to convert coordinates in a reference coordinate system of the microphone array into coordinates in a reference coordinate system of the first camera; The second mouth area coordinates are acquired according to the coordinates of the sound collected by the microphone array in the reference coordinate system of the microphone array and the first conversion parameters.
3. The method according to claim 1, wherein: Also includes: If the coordinates of the second mouth area cannot be obtained according to the sound collected by the microphone array, the facial contours of all the person images in the first picture are identified, and a number of face detection frames corresponding to the identified facial contours are generated; Determining the facial contour corresponding to the face detection frame having the largest area as the target facial contour; The person corresponding to the target facial contour is determined as the first target person.
4. The method according to claim 1, wherein: The acquiring first eye region coordinates according to the first picture and the second picture includes: Acquire a second conversion parameter, where the second conversion parameter is used to convert the coordinates in the reference coordinate system of the first camera into the coordinates in the reference coordinate system of the second camera; Acquire third eye region coordinates, where the third eye region coordinates are two-dimensional coordinates of the eye region image of the first target person in the first picture in a reference coordinate system of the first picture; Acquire fourth eye region coordinates, where the fourth eye region coordinates are two-dimensional coordinates of the eye region image of the first target person in the second picture in a reference coordinate system of the second picture; The first eye area coordinates are acquired according to the third eye area coordinates, the fourth eye area coordinates and the second conversion parameters.
5. The method according to claim 1, wherein: The acquiring the second eye region coordinates according to the second target person image includes: Acquire a third conversion parameter, where the third conversion parameter is used to convert the coordinates in the reference coordinate system of the display screen into the coordinates in the reference coordinate system of the first camera; Acquire fifth eye region coordinates, where the fifth eye region coordinates are two-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the display screen; The second eye area coordinates are acquired according to the fifth eye area coordinates and the third conversion parameters.
6. The method according to claim 1, wherein: The first eye region coordinates include first left eye region coordinates and first right eye region coordinates, the second eye region coordinates include second left eye region coordinates and second right eye region coordinates, and adjusting the angle of the eye region image in the second target person image toward the eye region of the first target person according to the first eye region coordinates and the second eye region coordinates, comprises: adjusting the angle of the left eye area image in the second target person image according to the difference between the first left eye area coordinates and the second left eye area coordinates; According to the difference between the first right eye area coordinates and the second right eye area coordinates, the angle of the right eye area image in the second target person image is adjusted.
7. A video calling device, comprising: A plurality of cameras, for capturing a plurality of images, wherein the plurality of images are images of a target area captured by the plurality of cameras at the same time, and the plurality of images include a first image captured by a first camera and a second image captured by a second camera; A display screen, used to present the images of both parties in the video call; A first person acquisition module, used to determine a first target person, where the first target person is a person in the target area photographed by the multiple cameras; a first parsing module, configured to obtain first eye region coordinates according to the first picture and the second picture taken by the first camera and the second camera, wherein the first eye region coordinates are three-dimensional coordinates of the eye region of the first target person in a reference coordinate system of the first camera; A second person acquisition module, used to determine a second target person image, where the second target person image is an image of the second target person displayed on the display screen, and the second target person is the object of the video call with the first target person; a second parsing module, configured to obtain second eye region coordinates according to the second target person image determined by the second person acquisition module, wherein the second eye region coordinates are three-dimensional coordinates of the eye region image in the second target person image in the reference coordinate system of the first camera; a sight redirection module, configured to adjust the angle of the eye region image in the second target person image toward the eye region of the first target person according to the first eye region coordinates obtained by the first parsing module and the second eye region coordinates obtained by the second parsing module; The first person acquisition module includes a sound acquisition module; the sound acquisition module is used to collect the sound emitted by each person in the target area, and perform sound source positioning according to the collected sounds of the people in the target area, and obtain the three-dimensional coordinates of the person corresponding to the sound source position in the reference coordinate system of the first camera; The first person acquisition module is also used to acquire first mouth area coordinates of each person in the target area according to the first picture and the second picture, the first mouth area coordinates being the three-dimensional coordinates of the mouth of the person in the target area in the reference coordinate system of the first camera; perform sound source positioning according to the sound of the person in the target area collected by the microphone array to acquire second mouth area coordinates, the second mouth area coordinates being the three-dimensional coordinates of the person corresponding to the sound source position of the sound collected by the microphone array in the reference coordinate system of the first camera; and determine the person corresponding to the first mouth area coordinate that is closest in straight-line distance to the second mouth area coordinate as the first target person.
8. A computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the computer program is used to implement the video calling method described in any one of claims 1 to 6.
9. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the video calling method described in any one of claims 1-6.
Citation Information
Patent Citations
Method for correcting user's gaze direction in image, machine-readable storage medium and communication terminal
CN103310186A
audiovisual MULTI-CHANNEL VOICE DETECTOR
RU174044U1
System and method for determining directionality of imagery using head tracking
WO2021258201A1