Live broadcast interaction method, live broadcast interaction system and live broadcast device
Patent Information
- Application Number
- CN202110982604.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-25
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2041-08-25
AI Technical Summary
[0047]基于本申请实施例的上述内容,相对于现有技术而言,本申请实施例提供的直播互动方法、直播互动系统及直播设备,可以从主播视频图像中提取包括眼球关键点以及多个眼皮关键点的多个人眼关键点,并根据所述眼球关键点以及所述眼皮关键点的位置信息获得主播的眼睛动作信息用于控制直播互动画面中展示的虚拟形象执行相应的眼睛互动操作。如此,可以根据主播视频图像实现虚拟形象的眼睛动作与主播眼部动作的随动控制,将主播的眼部动作复制到虚拟形象上,进而对主播行为进行生动的表达,可以增强直播的互动效果并提升用户体验。
Smart Images

Figure CN115720274B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of online live streaming technology, and more specifically, to a live streaming interactive method, a live streaming interactive system, and live streaming equipment. Background Technology
[0002] With the continuous development of mobile internet and network communication technologies, live streaming has rapidly developed and been applied in people's daily work and life. For example, users can watch live streams provided by various hosts on live streaming platforms online using smartphones, computers, tablets, and other devices. Alternatively, users can also provide live stream content anytime, anywhere on relevant live streaming platforms using smartphones, computers, tablets, and other devices for others to watch. In some specific live streaming scenarios, to provide a diverse live streaming experience, live streaming based on virtual avatars has also been widely used. Compared to live streaming by real hosts, live streaming with virtual avatars eliminates the need for real-person interaction; the host can control the virtual avatar in the background to simulate the host's behavior and interact with the audience. However, how to make the virtual avatar vividly express the host's behavior is a major technical problem that needs to be solved. Summary of the Invention
[0003] Based on the above, in a first aspect, embodiments of this application provide a live streaming interaction method applied to a live streaming device, the method comprising:
[0004] Acquire a video image of the anchor and extract multiple key points of the human eye from the video image of the anchor. The multiple key points of the human eye include key points of the eyeball and multiple key points of the eyelid.
[0005] The anchor's eye movement information is obtained based on the positional information of the key points of the eyeball and the key points of the eyelid;
[0006] Based on the anchor's eye movement information, the virtual avatar displayed in the live interactive screen is controlled to perform corresponding eye interaction operations.
[0007] Based on one possible implementation of the first aspect, obtaining the anchor's eye movement information according to the positional information of the eyeball key points and the eyelid key points includes:
[0008] Obtain at least one reference key point from the plurality of eyelid key points;
[0009] The distance between the eyeball key point and each of the reference key points is calculated based on the location information of the eyeball key points and the location information of the at least one reference key point.
[0010] The anchor's eye movement information is obtained based on the distance between the eyeball key points and each of the reference key points.
[0011] Based on one possible implementation of the first aspect, the method further includes:
[0012] Obtain a predetermined reference image including the anchor's eye image, and analyze the reference image to obtain the first standard distance between the anchor's left and right eye corners and the second standard distance between the upper and lower eyelids;
[0013] The step of obtaining the anchor's eye movement information based on the positional information of the key points of the eyeball and the key points of the eyelid includes:
[0014] One of the upper eyelid point and the lower eyelid point among the plurality of eyelid key points is taken as the first reference key point, and one of the left corner point and the right corner point among the plurality of eyelid key points is taken as the second reference key point.
[0015] Calculate the first distance and the second distance between the eyeball key points and the first reference key point and the second reference key point, respectively;
[0016] The first distance and the second distance are compared with the first standard distance and the second standard distance, respectively, to obtain the anchor's eye movement information.
[0017] Based on one possible implementation of the first aspect, the eye movement information of the broadcaster is obtained by comparing the first distance and the second distance with the first standard distance and the second standard distance, respectively, including:
[0018] Calculate the first ratio between the first distance and the first standard distance, and the second ratio between the second distance and the second standard distance;
[0019] Based on the analysis of the first ratio and the second ratio, the rotation information of the anchor's eyes moving up, down, left, and right is obtained, and thus the eye movement information is obtained.
[0020] Based on one possible implementation of the first aspect, obtaining the anchor's eye movement information according to the positional information of the eyeball key points and the eyelid key points includes:
[0021] According to a preset frame interval, the previous video frame image of the broadcaster before the current video frame image of the broadcaster is obtained as a reference frame image;
[0022] Obtain the distance between the eyeball key points extracted from the reference frame image and each eyelid key point to obtain a first distance sequence;
[0023] Calculate the distance between the eyeball key points extracted from the current anchor video frame image and each eyelid key point to obtain a second distance sequence;
[0024] The first distance sequence and the second distance sequence are compared and analyzed to obtain the anchor's eye movement information.
[0025] Based on one possible implementation of the first aspect, the method further includes:
[0026] Obtain a predetermined reference image including the anchor's eye image, extract key points of the upper eyelid and the lower eyelid from the reference image, and calculate the distance between the key points of the upper eyelid and the key points of the lower eyelid as the first calibration distance.
[0027] The step of obtaining the anchor's eye movement information based on the positional information of the key points of the eyeball and the key points of the eyelid includes:
[0028] Based on the lower eyelid point, upper eyelid point and the first calibration distance among the plurality of eyelid key points, a virtual upper eyelid point is determined in the anchor video image as a first reference key point, wherein the virtual upper eyelid point, the lower eyelid point and the upper eyelid point are collinear, and the line distance between the virtual upper eyelid point and the lower eyelid point is equal to the first calibration distance;
[0029] Calculate the first distance between the eyeball key point and the left corner of the multiple eyelid key points, the second distance between the eyeball key point and the right corner of the multiple eyelid key points, the third distance between the eyeball key point and the lower eyelid key point, and the fourth distance between the eyeball key point and the virtual upper eyelid point.
[0030] The anchor's eye movement coefficients are calculated based on the first distance, the second distance, the third distance, and the fourth distance. The eye movement coefficients include a first movement coefficient negatively correlated with the first distance, a second movement coefficient negatively correlated with the second distance, a third movement coefficient negatively correlated with the third distance, and a fourth movement coefficient negatively correlated with the fourth distance. The eye movement information includes the eye movement coefficients.
[0031] Based on one possible implementation of the first aspect, the calculation of the anchor's eye movement coefficient according to the first distance, the second distance, the third distance, and the fourth distance includes:
[0032] Extract the key points at the left and right corners of the eyes from the reference image, and calculate the distance between the key points at the left and right corners of the eyes as the second calibration distance;
[0033] The eye movement coefficients were calculated using the following formulas:
[0034] C1 = (Pd2 - D1) / Pd2;
[0035] C2 = (Pd2 - D2) / Pd2;
[0036] C3 = (Pd1 - D3) / Pd1;
[0037] C4 = (Pd1 - D4) / Pd1;
[0038] Wherein, C1, C2, C3, and C4 are the first action coefficient, the second action coefficient, the third action coefficient, and the fourth action coefficient, respectively, and their values are in the range of [0-1]; D1, D2, D3, and D4 are the first distance, the second distance, the third distance, and the fourth distance, respectively; Pd1 and Pd2 are the first calibration distance and the second calibration distance, respectively.
[0039] Based on one possible implementation of the first aspect, controlling the virtual image displayed in the live interactive screen to perform corresponding eye interaction operations according to the anchor's eye movement information includes:
[0040] The live streaming device, acting as a live streaming terminal, sends the eye movement information to the live streaming service platform. The platform then renders the virtual character's eye movements based on this information, generating a corresponding live stream which is sent to the receiving terminal for display.
[0041] The live streaming device, acting as a live streaming service platform, sends the eye movement information to the live streaming receiving terminal, which then renders and displays the eye movements of the virtual character.
[0042] Secondly, embodiments of this application also provide a live interactive system, running on a live streaming device, the live interactive system comprising:
[0043] The key point extraction module is used to acquire the anchor video image and extract multiple human eye key points from the anchor video image. The multiple human eye key points include eyeball key points and multiple eyelid key points.
[0044] The motion information acquisition module is used to obtain the anchor's eye motion information based on the positional information of the eyeball key points and the eyelid key points;
[0045] The interactive control module is used to control the virtual image displayed in the live interactive screen to perform corresponding eye interaction operations based on the anchor's eye movement information.
[0046] Thirdly, embodiments of this application also provide a live streaming device, including a machine-readable storage medium and one or more processors, wherein the machine-readable storage medium stores machine-executable instructions, which, when executed by the one or more processors, implement the method described in any one of claims 1-8.
[0047] Based on the above content of the embodiments of this application, compared with the prior art, the live interactive method, live interactive system, and live streaming equipment provided in the embodiments of this application can extract multiple human eye key points, including eyeball key points and multiple eyelid key points, from the anchor's video image. Based on the positional information of the eyeball key points and the eyelid key points, the anchor's eye movement information is obtained to control the virtual image displayed in the live interactive screen to perform corresponding eye interactive operations. In this way, the eye movements of the virtual image and the anchor's eye movements can be dynamically controlled based on the anchor's video image, copying the anchor's eye movements onto the virtual image, thereby vividly expressing the anchor's behavior, enhancing the interactive effect of the live stream, and improving the user experience. Attached Figure Description
[0048] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a schematic diagram of a live streaming system provided in an embodiment of this application.
[0050] Figure 2 This is a flowchart illustrating the live interactive method provided in the embodiments of this application.
[0051] Figure 3 yes Figure 2 One of the specific implementation sub-process diagrams of step S200 in the process.
[0052] Figure 4 This is one of the exemplary schematic diagrams provided in the embodiments of this application for the extraction of multiple key points of the human eye.
[0053] Figure 5 yes Figure 2 The second schematic diagram of the specific implementation of step S200 in the process.
[0054] Figure 6 This is a second exemplary schematic diagram provided in the embodiments of this application for the extraction of multiple key points of the human eye.
[0055] Figure 7 yes Figure 2 The third schematic diagram of the specific implementation of step S200 in the process.
[0056] Figure 8 This is the third exemplary schematic diagram provided in the embodiments of this application for the extraction of multiple key points of the human eye.
[0057] Figure 9 yes Figure 2 The fourth schematic diagram of the specific implementation of step S200 in the process.
[0058] Figure 10 This is the fourth exemplary schematic diagram provided in the embodiments of this application for the extraction of multiple key points of the human eye.
[0059] Figure 11 This is a schematic diagram illustrating the effect of controlling the virtual image in the anchor's screen using the method provided in the embodiments of this application.
[0060] Figure 12 This is a schematic diagram of a live streaming device provided in an embodiment of this application for implementing the above-described live streaming interaction method.
[0061] Figure 13 yes Figure 12 A schematic diagram of the functional modules of the live interactive system. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0063] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0064] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0065] In the description of this application, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0066] In the description of this application, it should also be noted that, unless otherwise expressly specified and limited, the terms "set up," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0067] After researching and analyzing current common methods of live streaming using virtual avatars, the inventors discovered that most suffer from the problem of virtual avatars failing to vividly express the streamer's behavior. In particular, to make virtual live streaming more vivid and natural, capturing the streamer's eye movements (such as eyeball movements) is extremely important. By accurately capturing the streamer's eye movements and controlling the virtual avatar to follow those movements, the streamer's eye movements can be "copied" onto the virtual avatar, making the live streaming process more vivid and lifelike.
[0068] Currently, in order to capture the anchor's eye movements, professional eye trackers capable of capturing eyeballs or pupils can be used. However, professional eye trackers are very expensive. If virtual live streaming is to be popularized, using such equipment would be too costly and would make it difficult to promote virtual live streaming to the general public.
[0069] Based on the aforementioned problems, and in order to achieve accurate capture of the anchor's eye movements while avoiding the high costs associated with using specialized equipment such as eye trackers, this application provides a real-time, low-cost eye movement capture solution. This solution uses conventional camera devices to accurately identify and capture the anchor's eye movements in video images, thereby controlling the virtual avatar to interact with the anchor's eyes. This enhances the interactivity and entertainment value of virtual live streaming, expanding the market prospects and user experience of virtual live streaming applications. The embodiments of this application will be described below by way of example.
[0070] First, the application scenarios of the embodiments of this application will be introduced. For example... Figure 1The diagram shown is a schematic of a live streaming system provided in an embodiment of this application. In this embodiment, the live streaming system includes a live streaming providing terminal 100, a live streaming service platform 200, and a live streaming receiving terminal 300. Exemplarily, the live streaming providing terminal 100 and the live streaming receiving terminal 300 can access the live streaming service platform 200 via a network to use the live streaming services provided by the live streaming service platform 200. For example, as an example, the live streaming providing terminal 100 can download a broadcaster application (APP) through the live streaming service platform 200, register through the broadcaster application, and then broadcast content through the live streaming service platform 200. Correspondingly, the live streaming receiving terminal 300 can also download a viewer application through the live streaming service platform 200, and access the live streaming service platform 200 through the viewer application to watch the live streaming content provided by the live streaming providing terminal 100. In some possible implementations, the broadcaster application and the viewer application can also be an integrated application.
[0071] For example, a live streaming terminal 100 can send live streaming content (such as a live video stream) to a live streaming service platform 200, and viewers can access the live streaming service platform 200 through a live streaming receiving terminal 300 to watch the live streaming content. The live streaming content pushed by the live streaming service platform 200 can be real-time content currently being streamed on the platform, or it can be historical live streaming content stored after the stream has finished. This is understandable. Figure 1 The live streaming system shown is merely an alternative example; in other possible embodiments, the live streaming system may also include only... Figure 1 The components shown may include one or more other components.
[0072] Furthermore, it should be noted that in specific application scenarios, the roles of the live streaming provider terminal 100 and the live streaming receiver terminal 300 can be interchanged. For example, the host of the live streaming provider terminal 100 can use it to provide live streaming services or watch live streaming content provided by other hosts as a viewer. Similarly, a user of the live streaming receiver terminal 300 can use it to watch live streaming content provided by hosts they follow, or they can also broadcast live through the live streaming receiver terminal 300 as a host.
[0073] In this embodiment, the live streaming providing terminal 100 and the live streaming receiving terminal 300 can be, but are not limited to, smartphones, personal digital assistants, tablets, personal computers, laptops, virtual reality terminal devices, augmented reality terminal devices, etc. The live streaming providing terminal 100 and the live streaming receiving terminal 300 can be equipped with relevant applications or program components for enabling live streaming interaction, such as applications (APPs), web pages, live streaming mini-programs, live streaming plugins or components, etc., but not limited to these. The live streaming service platform 200 can be a backend device providing live streaming services, such as, but not limited to, servers, server clusters, cloud service centers, etc.
[0074] In this embodiment, the live streaming terminal 100 may include an image acquisition device for capturing images of the broadcaster. Additionally, it may include an audio acquisition device for capturing the broadcaster's voice and input / output devices for the broadcaster to input information, such as, but not limited to, a keyboard, mouse, touchscreen, microphone, and speaker. The image acquisition device, audio acquisition device, and input / output devices may be directly installed on or integrated into the live streaming terminal 100, or they may be independent of the live streaming terminal 100 but communicatively connected to it for data communication and interaction.
[0075] like Figure 2 The diagram shown is a flowchart illustrating the live streaming interaction method provided in an embodiment of this application. In this embodiment, the live streaming interaction method is executed and implemented by a live streaming device. The live streaming device may be... Figure 1 The live streaming terminal 100 shown can also be the live streaming service platform 200, and there is no specific limitation. It should be understood that the order of some steps in the live streaming interaction method provided in this embodiment can be interchanged according to actual needs during actual implementation, or some steps can be omitted or deleted, and this embodiment does not specifically limit this.
[0076] The following is combined with Figure 2 The steps of the live interactive method in this embodiment are described in detail by way of examples, as follows: Figure 2 As shown, the method may include the live interactive content provided in the embodiments of this application, as described in steps S100-S300 below.
[0077] Step S100: Obtain the anchor video image and extract multiple human eye key points from the anchor video image. The multiple human eye key points include eyeball key points and multiple eyelid key points.
[0078] In this embodiment, the anchor video image can be obtained by the live streaming provider 100 through a built-in or externally connected graphics acquisition device capturing the anchor. The multiple eye key points can be obtained by extracting eye key points from the current video image frame in the anchor video image. In one possible implementation, in step S100, eye key points can be extracted based on the anchor video image using a key point SDK (Software Development Kit).
[0079] Step S200: Obtain the anchor's eye movement information based on the position information of the key points of the eyeball and the key points of the eyelid.
[0080] In this embodiment, the eye movement information may include, but is not limited to, eye movements such as looking inward (eyeLookIn), looking outward (eyeLookOut), looking up (eyeLookUp), and looking down (eyeLookDown), as well as the degree of eye movement corresponding to each eye movement. This embodiment does not specifically limit this.
[0081] Step S300: Control the virtual image displayed in the live interactive screen to perform corresponding eye interaction operations based on the anchor's eye movement information.
[0082] For example, in one possible scenario, the virtual character displayed on the live streaming receiving terminal 300 can be controlled to move according to the aforementioned eye movement information, thereby copying the anchor's eye movements onto the virtual character's eye movements. For instance, the live streaming providing terminal 100 can send the eye movement information to the live streaming service platform 200, which then renders the virtual character's eye movements based on the eye movement information and generates a corresponding live stream, which is then sent to the live streaming receiving terminal 300 for display. Alternatively, the live streaming service platform 100 can extract and analyze key points of the anchor's eyes from the video image to obtain the eye movement information, and then send this information to the live streaming receiving terminal 300, which renders and displays the virtual character's eye movements. The specific implementation depends on the actual scenario, and this embodiment does not specifically limit this approach.
[0083] Furthermore, in a first possible implementation of the embodiments of this application, such as Figure 3 As shown, step S200 can be specifically implemented through the following steps S201-S203, which are illustrated below.
[0084] Step S201: Obtain at least one reference key point from the plurality of eyelid key points.
[0085] As an example, in this embodiment, for example Figure 4 As shown, Figure 4 This is one of the schematic diagrams of multiple key points extracted from the human eye. Taking the eye image corresponding to one eye (such as the left or right eye) of the anchor as an example, at least one of the key points, namely the left corner of the eye (L), the right corner of the eye (R), the upper eyelid point (U), and the lower eyelid point (D), can be used as the reference key point based on the key point description information carried in the multiple eyelid key points obtained above. Preferably, in this embodiment, the left corner of the eye can be the leftmost eyelid key point among the multiple eyelid key points; the right corner of the eye can be the rightmost eyelid key point among the multiple eyelid key points; the upper eyelid point can be the uppermost eyelid key point among the multiple eyelid key points; and the lower eyelid point can be the lowermost eyelid key point among the multiple eyelid key points.
[0086] Step S202: Calculate the distance between the eyeball key point and each of the reference key points based on the location information of the eyeball key points and the location information of the at least one reference key point.
[0087] As an example, in this instance, the distance between the eyeball keypoint and each of the reference keypoints can be calculated using the Euclidean distance calculation formula based on the position coordinates of the eyeball keypoint and the reference keypoint. The specific calculation method is not limited in this embodiment.
[0088] Step S203: Obtain the anchor's eye movement information based on the distance between the eyeball key points and each of the reference key points.
[0089] For example, in the above content, if an eyelid key point is selected as the reference key point, taking the left corner of the eye as the reference key point, the left and right eye movement of the broadcaster can be determined based on the distance between the eyeball key point and the left corner of the eye. For example, before the broadcast, an image of the broadcaster's eyes in a naturally open state can be obtained as a reference image, and the standard distance between the left and right corners of the broadcaster's eyes (such as the distance between the left and right corners of the eye) can be analyzed. Based on this, the distance between the eyeball key point and the left corner of the eye can be compared with the standard distance to determine whether the broadcaster's eyes are turning to the left or right, thereby obtaining the eye movement information. For example, assuming the distance between the eyeball key point and the left corner of the eye is d1, and the standard distance is d0, the ratio of d1 to d0 can be calculated, and the degree of eye movement can be determined based on the ratio of d1 to d0. For example, if the ratio of d1 to d0 is less than 50%, it indicates that the anchor's eye movement is a leftward eye movement, the degree of which can be determined based on the specific ratio. Conversely, if the ratio of d1 to d0 is greater than 50%, it indicates that the anchor's eye movement is a rightward eye movement, the degree of which can be determined based on the specific ratio. Alternatively, the right corner of the eye, the upper eyelid, or the lower eyelid can be chosen as the reference key point. In this case, the method for determining the anchor's movement information is similar to that using the left corner of the eye as the reference key point, and will not be elaborated further. In this implementation, using a single reference key point allows the virtual image to be dynamically controlled in one eye movement direction (e.g., horizontal and vertical) in sync with the anchor's eye movements.
[0090] Furthermore, in a second possible implementation of this application, given a pre-obtained image of the anchor's eyes in a naturally open state as a reference image, a first standard distance between the anchor's left and right eye corners and a second standard distance between the upper and lower eyelids can be obtained based on the reference image. The first standard distance between the left and right eye corners can be the distance between the left and right eye corner points extracted from the reference image, and the second standard distance between the upper and lower eyelids can be the distance between the upper and lower eyelid points extracted from the reference image.
[0091] Based on this, for step S200, another alternative implementation can be achieved through, for example... Figure 5 The specific implementation of steps S211-S213 shown is illustrated below.
[0092] Step S211: Take one of the upper eyelid point and the lower eyelid point as the first reference key point, and take one of the left eye corner point and the right eye corner point as the second reference key point.
[0093] Preferably, in this embodiment, considering the physiological characteristics of humans, the lower eyelid basically does not move, so the lower eyelid point is preferred as the second reference key point.
[0094] Step S212: Calculate the first distance and the second distance between the eyeball key point and the first reference key point and the second reference key point, respectively.
[0095] Step S213: Compare the first distance and the second distance with the first standard distance and the second standard distance respectively to obtain the anchor's eye movement information.
[0096] In this embodiment, a first ratio between the first distance and the first standard distance, and a second ratio between the second distance and the second standard distance, can be calculated. Based on the first and second ratios, the rotation information of the eyeballs moving up, down, left, and right can be analyzed, thereby obtaining the eye movement information. For example, taking the first reference key point as the left corner of the eye and the second reference key point as the lower eyelid point, when the first ratio is less than 50%, it indicates that the broadcaster's eye movement is a leftward eye movement, the degree of which can be determined based on the specific ratio; while when the first ratio is greater than 50%, it indicates that the broadcaster's eye movement is a rightward eye movement, the degree of which can be determined based on the specific ratio. Similarly, when the second ratio is less than 50%, it indicates that the broadcaster's eye movement is a downward eye movement, the degree of which can be determined based on the specific ratio; while when the second ratio is greater than 50%, it indicates that the broadcaster's eye movement is an upward eye movement, the degree of which can be determined based on the specific ratio.
[0097] For example, Figure 6As shown, assuming the left corner of the eye L is taken as the first reference key point and the lower eyelid point D is taken as the second reference key point, when the eyeball key point B is closer to the left corner of the eye L, the ratio between the first distance and the first standard distance (L1) is smaller (gradually approaching 0), indicating that the anchor's eyes are gradually turning to the left (looking outward); when the eyeball key point B is farther from the left corner of the eye L, the ratio between the first distance and the first standard distance (L1) is larger (gradually approaching 1), indicating that the anchor's eyes are gradually turning to the right (looking inward). Similarly, when the eyeball key point B is closer to the lower eyelid point D, the ratio between the second distance and the second standard distance is smaller (gradually approaching 0), indicating that the anchor's eyes are gradually turning downward (looking downward); when the eyeball key point B is farther from the lower eyelid point D, the ratio between the second distance and the second standard distance is larger (gradually approaching 1), indicating that the anchor's eyes are gradually turning upward (looking upward). In this way, the above method can be used to obtain the anchor's eye movement information (such as eye movement information in both horizontal and vertical directions), which can be used to drive the virtual image to perform follow-up control of eye movements, thereby replicating the anchor's eye movements onto the virtual image to achieve live interaction with the live viewers.
[0098] The second implementation method described above analyzes the anchor's eye movements by selecting two key eyelid points in different directional dimensions as the first and second key points. Compared to using only one key point, this method provides more comprehensive eye movement information, allowing for more accurate capture of eye movements. Therefore, the eye movement information obtained through this method enables virtual avatars to more vividly express the anchor's current eye movements during interactive operations, enhancing the interactive effect of the live stream.
[0099] Furthermore, to further enhance the accuracy of capturing the anchor's eye movements, in a third possible implementation of this application's embodiments, step S200 can also be achieved through methods such as... Figure 7 The specific implementation of steps S221-S224 shown is illustrated below.
[0100] Step S221: Obtain the previous video frame image of the broadcaster before the current video frame image of the broadcaster according to the preset frame interval as the reference frame image.
[0101] For example, as an example, the frame interval can be set according to the actual control needs and the computing power of the terminal device. For instance, the frame interval can be set to 5, 10, 15, etc. In this way, eye movements can be captured every 5, 10, or 15 frames to drive the virtual character to perform eye movement interactions. The specific interval can be determined according to the actual situation, and this embodiment does not limit it. For example, taking 5 frames as a frame interval, the reference frame image corresponding to the current anchor video frame image is the current anchor video frame image corresponding to the last time eye movement capture was performed.
[0102] Step S222: Obtain the distance between the eyeball keypoints extracted from the reference frame image and each eyelid keypoint to obtain a first distance sequence. In this embodiment, the first distance sequence can be calculated from the previous corresponding current anchor video frame image during the previous eye motion capture, and can be directly obtained from the cache of the live streaming device without recalculation, thereby reducing the computational load of the device.
[0103] Step S223: Calculate the distance between the eyeball key points extracted from the current anchor video frame image and each eyelid key point to obtain the second distance sequence.
[0104] Preferably, in this embodiment, the currently calculated second distance sequence can be used as the first distance sequence corresponding to the next eye motion capture.
[0105] Step S224: Compare and analyze the first distance sequence and the second distance sequence to obtain the anchor's eye movement information. For example, the movement trajectory of the anchor's key eye points can be obtained by comparing and analyzing the first distance sequence and the second distance sequence, and used as the eye movement information.
[0106] In this way, by analyzing the changes in distance between key points of the eyeball and key points of each eyelid between two consecutive frames of the anchor's video image, the anchor's eye movement information can be obtained more accurately.
[0107] Based on this, for step S300, virtual eyelid key points corresponding to the key points of each eyelid of the anchor and virtual eyeball key points corresponding to the key points of the anchor's eyes can be pre-marked in the eye image of the virtual image. In this way, when controlling the virtual image displayed in the live interactive screen of the live receiving terminal to perform the corresponding eye interaction operation, the virtual eyeball of the virtual image can be controlled to move accordingly based on the change in distance between the key points of the eyeball and each key point of each eyelid between two consecutive frames of anchor video images, thereby matching the anchor's eye movements and making the live broadcast process more vivid and lifelike.
[0108] Furthermore, generally speaking, while the corner of the eye is fixed, the eyelids (especially the upper eyelid) move with the opening and closing of the eyes. Therefore, calculating the degree of upward / downward looking based solely on the distance from the key point of the eyeball to the upper or lower eyelid point can lead to errors in certain special cases. For example, if the eyeball remains still and the eyes are slightly closed, the upper eyelid will move downward, reducing the distance between the key point of the eyeball and the upper eyelid point, thus increasing the calculated degree of upward looking. However, in reality, the eyeball may not be moving, so there is no issue of the degree of upward looking increasing. Therefore, to solve this problem and further enhance the accuracy of capturing the anchor's eye movements, in the fourth possible implementation of this application, a predetermined reference image including the anchor's eye image can be obtained. Key points of the upper and lower eyelids can be extracted from the reference image, and the distance between the key points of the upper and lower eyelids can be calculated as the first calibration distance.
[0109] For example Figure 8 The image shown is a schematic diagram of key points of the human eye extracted from a pre-acquired reference image of the anchor in a naturally open-eye state. The distance Pd1 between the key point U on the upper eyelid and the key point on the lower eyelid is used as the first calibration distance. Based on this, for step S200, further steps can be taken... Figure 9 The specific implementation of steps S231-S233 shown is illustrated below.
[0110] Step S231: Determine a virtual upper eyelid point in the anchor video image based on the lower eyelid point, upper eyelid point and the first calibration distance among the plurality of eyelid key points. The virtual upper eyelid point, the lower eyelid point and the upper eyelid point are collinear, and the line distance between the virtual upper eyelid point and the lower eyelid point is equal to the first calibration distance.
[0111] For example, in Figure 10 In the state shown, when the anchor's upper eyelids are slightly closed downwards, in order to accurately calculate the degree of vertical movement of the eyeballs, in this embodiment, the lower eyelid point D is taken as the starting point, and the endpoint of the first calibration distance Pd1 extended along the direction of the real upper eyelid point is taken as the virtual upper eyelid point F. In this way, the vertical movement information of the eyeballs (such as the degree of vertical movement) can be analyzed based on the change in distance between the key points of the eyeballs and the virtual upper eyelid point F.
[0112] Step S232: Calculate the first distance between the eyeball key point and the left corner of the multiple eyelid key points, the second distance between the eyeball key point and the right corner of the multiple eyelid key points, the third distance between the eyeball key point and the lower eyelid key point, and the fourth distance between the eyeball key point and the virtual upper eyelid point.
[0113] Step S233: Calculate the anchor's eye movement coefficients based on the first distance, second distance, third distance, and fourth distance. The eye movement coefficients include a first movement coefficient negatively correlated with the first distance, a second movement coefficient negatively correlated with the second distance, a third movement coefficient negatively correlated with the third distance, and a fourth movement coefficient negatively correlated with the fourth distance. The eye movement information includes the eye movement coefficients.
[0114] For more details, please refer to further reading. Figure 8 As shown, in this embodiment, key points at the left and right corners of the eyes can also be extracted from the reference image, and the distance between the key points at the left and right corners of the eyes can be calculated (e.g., ...). Figure 8 The Pd2 shown is used as the second calibration distance.
[0115] Based on this, regarding step S233, in this embodiment, the eye movement coefficient can be calculated according to the following formulas:
[0116] C1 = (Pd2 - D1) / Pd2;
[0117] C2 = (Pd2 - D2) / Pd2;
[0118] C3 = (Pd1 - D3) / Pd1;
[0119] C4 = (Pd1 - D4) / Pd1;
[0120] Wherein, C1, C2, C3, and C4 are the first action coefficient, the second action coefficient, the third action coefficient, and the fourth action coefficient, respectively, and their values are in the range of [0-1]; D1, D2, D3, and D4 are the first distance, the second distance, the third distance, and the fourth distance, respectively; Pd1 and Pd2 are the first calibration distance and the second calibration distance, respectively.
[0121] Based on the above method, it can be seen that the closer the eyeball key point is to the left corner of the eye, the closer the first coefficient is to 1; the farther the distance is from the left corner of the eye, the closer the first coefficient is to 0. Therefore, the magnitude of the first coefficient can characterize the change in distance between the eyeball key point and the left corner of the eye, thereby expressing the degree to which the anchor's eyes turn to the left (look outward / look to the left).
[0122] Correspondingly, the closer the eyeball key point is to the right corner of the eye, the closer the second coefficient is to 1; the farther the distance is from the right corner of the eye, the closer the second coefficient is to 0. Therefore, the magnitude of the second coefficient can characterize the change in distance between the eyeball key point and the right corner of the eye, thereby expressing the degree to which the anchor's eyes turn to the right (looking inward / to the right).
[0123] Correspondingly, the closer the eyeball key point is to the lower eyelid point, the closer the third coefficient is to 1; the farther the eyeball key point is from the lower eyelid point, the closer the third coefficient is to 0. Therefore, the magnitude of the third coefficient can characterize the change in distance between the eyeball key point and the lower eyelid, thereby expressing the degree to which the anchor's eyes turn downward (look down).
[0124] Correspondingly, the closer the eyeball key point is to the virtual upper eyelid point, the closer the fourth coefficient is to 1; the farther the eyeball key point is from the virtual upper eyelid point, the closer the fourth coefficient is to 0. Therefore, the magnitude of the fourth coefficient can characterize the change in distance between the eyeball key point and the virtual upper eyelid point, thereby expressing the degree to which the anchor's eyes turn upward (look upward).
[0125] After obtaining the aforementioned eye movement coefficients, these coefficients can be included in the eye movement information. Thus, in step S300, based on the eye movement coefficients obtained from the analysis of key points of the anchor's eyes in each video image, the eyes of the virtual character in the live stream can be controlled to rotate according to the corresponding eye movement coefficients. This allows the anchor's eye movements to be copied onto the virtual character as vividly and realistically as possible, thereby enhancing the interactive effect of the live stream.
[0126] For example, see Figure 11As shown, when the first action coefficient C1 in the eye movement coefficients obtained by the above method is close to 1 (as shown in box M1), it indicates that the anchor's eyes are looking to the left, and correspondingly, the virtual image's eyes can be controlled to turn to the left; when the second action coefficient C2 in the eye movement coefficients obtained by the above method is close to 1 (as shown in box M2), it indicates that the anchor's eyes are looking to the right, and correspondingly, the virtual image's eyes can be controlled to turn to the right; when the third action coefficient C3 in the eye movement coefficients obtained by the above method is close to 1 (as shown in box M4), it indicates that the anchor's eyes are looking down, and correspondingly, the virtual image's eyes can be controlled to turn down; when the fourth action coefficient C4 in the eye movement coefficients obtained by the above method is close to 1 (as shown in box M3), it indicates that the anchor's eyes are looking up, and correspondingly, the virtual image's eyes can be controlled to turn up.
[0127] Please refer to Figure 12 , Figure 12 This is a schematic diagram of a live streaming device for implementing the above-described live streaming interaction method, provided in an embodiment of this application. In this embodiment, the live streaming device may be the live streaming providing device 100 or the live streaming service platform 200 described above. Specifically, the live streaming device may include one or more processors 110, a machine-readable storage medium 120, and a live streaming interaction system 130. The processor 110 and the machine-readable storage medium 120 are communicatively connected via a system bus. The machine-readable storage medium 120 stores machine-executable instructions, and the processor 110 implements the live streaming interaction method described above by reading and executing the machine-executable instructions in the machine-readable storage medium 120.
[0128] The machine-readable storage medium 120 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The machine-readable storage medium 120 is used to store a program, which the processor 110 executes upon receiving an execution instruction.
[0129] The processor 110 may be an integrated circuit chip with signal processing capabilities. The processor described above can be, but is not limited to, a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.
[0130] Please refer to Figure 13 This is a functional module diagram of the live interactive system 130. In this embodiment, the live interactive system 130 may include one or more software functional modules running on the live streaming device. These software functional modules may be stored in the machine-readable storage medium 120 in the form of computer programs, so that when these software functional modules are called and executed by the processor 130, they can implement the live interactive method described in this application embodiment.
[0131] In detail, the live interactive system 130 includes a key point extraction module 131, an action information acquisition module 132, and an interactive control module 133.
[0132] The key point extraction module 131 is used to acquire the anchor video image and extract multiple human eye key points from the anchor video image. The multiple human eye key points include eyeball key points and multiple eyelid key points.
[0133] In this embodiment, the anchor video image can be obtained by the live streaming provider 100 through its built-in or connected graphics acquisition device by capturing images of the anchor. The multiple eye key points can be obtained by extracting eye key points from the current video image frame in the anchor video image. In one possible implementation, in step S100, eye key points can be extracted based on the anchor video image using a key point SDK (Software Development Kit).
[0134] The motion information acquisition module 132 is used to obtain the anchor's eye motion information based on the position information of the eyeball key points and the eyelid key points.
[0135] In this embodiment, the eye movement information may include, but is not limited to, eye movements such as looking inward (eyeLookIn), looking outward (eyeLookOut), looking up (eyeLookUp), and looking down (eyeLookDown), as well as the degree of eye movement corresponding to each eye movement.
[0136] The interactive control module 133 is used to control the virtual image displayed in the live interactive screen to perform corresponding eye interaction operations based on the anchor's eye movement information.
[0137] For example, in one possible scenario, the virtual character displayed on the live streaming receiving terminal 300 can be controlled to move according to the aforementioned eye movement information, thereby copying the anchor's eye movements onto the virtual character's eye movements. For instance, the live streaming providing terminal 100 can send the eye movement information to the live streaming service platform 200, which then renders the virtual character's eye movements based on the eye movement information and generates a corresponding live stream, which is then sent to the live streaming receiving terminal 300 for display. Alternatively, the live streaming service platform 100 can extract and analyze key points of the anchor's eyes in the video image to obtain the eye movement information, and then send this information to the live streaming receiving terminal 300, which renders and displays the virtual character's eye movements. The specific implementation depends on the actual scenario, and this embodiment does not specifically limit this approach.
[0138] Furthermore, in the embodiments of this application, the key point extraction module 131, the action information acquisition module 132, and the interaction control module 133 can respectively execute steps S100-S300 in the live interaction method of the embodiments of this application. The specific implementation methods and contents of these modules can be referred to the detailed description of the corresponding steps, which will not be repeated in this embodiment.
[0139] In summary, the live streaming interaction method, system, and equipment provided in this application can extract multiple key points of the human eye, including key points of the eyeball and multiple key points of the eyelid, from the anchor's video image. Based on the positional information of the key points of the eyeball and the eyelids, the anchor's eye movement information is obtained to control the virtual avatar displayed in the live streaming interaction screen to perform corresponding eye interaction operations. Thus, the eye movements of the virtual avatar in the virtual live stream can be dynamically controlled to mirror the anchor's eye movements, replicating the anchor's eye movements onto the virtual avatar, thereby vividly expressing the anchor's behavior, enhancing the interactive effect of the live stream, and improving the user experience.
[0140] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0141] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0142] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0143] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0144] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A live streaming interactive method, applied to a live streaming device, characterized in that, The method includes: Acquire a video image of the anchor and extract multiple key points of the human eye from the video image of the anchor. The multiple key points of the human eye include key points of the eyeball and multiple key points of the eyelid. Obtain a predetermined reference image including the anchor's eye image, extract key points of the upper eyelid and the lower eyelid from the reference image, and calculate the distance between the key points of the upper eyelid and the key points of the lower eyelid as the first calibration distance. Based on the lower eyelid point, upper eyelid point, and the first calibration distance among the plurality of eyelid key points, a virtual upper eyelid point is determined as a first reference key point in the anchor video image. The virtual upper eyelid point, the lower eyelid point, and the upper eyelid point are collinear, and the distance between the virtual upper eyelid point and the lower eyelid point is equal to the first calibration distance. The first distance between the eyeball key point and the left corner of the plurality of eyelid key points, the second distance between the eyeball key point and the right corner of the plurality of eyelid key points, and the distance between the eyeball key point and the lower eyelid point are calculated respectively. The third distance between the lower eyelid point and the eyeball key point, and the fourth distance between the eyeball key point and the virtual upper eyelid point; the eyeball movement coefficient of the anchor is calculated based on the first distance, the second distance, the third distance, and the fourth distance, wherein the eyeball movement coefficient includes a first movement coefficient negatively correlated with the first distance, a second movement coefficient negatively correlated with the second distance, a third movement coefficient negatively correlated with the third distance, and a fourth movement coefficient negatively correlated with the fourth distance, wherein eye movement information includes the eyeball movement coefficient; Based on the anchor's eye movement information, the virtual avatar displayed in the live interactive screen is controlled to perform corresponding eye interaction operations.
2. The method according to claim 1, characterized in that, The calculation of the anchor's eye movement coefficient based on the first distance, second distance, third distance, and fourth distance includes: Extract the key points at the left and right corners of the eyes from the reference image, and calculate the distance between the key points at the left and right corners of the eyes as the second calibration distance; The eye movement coefficients were calculated using the following formulas: C1 = (Pd2 - D1) / Pd2; C2 = (Pd2 - D2) / Pd2; C3 = (Pd1 - D3) / Pd1; C4 = (Pd1 - D4) / Pd1; Wherein, C1, C2, C3, and C4 are the first action coefficient, the second action coefficient, the third action coefficient, and the fourth action coefficient, respectively, and their values are in the range of [0-1]; D1, D2, D3, and D4 are the first distance, the second distance, the third distance, and the fourth distance, respectively; Pd1 and Pd2 are the first calibration distance and the second calibration distance, respectively.
3. The method according to any one of claims 1-2, characterized in that, The step of controlling the virtual avatar displayed in the live interactive screen to perform corresponding eye interaction operations based on the anchor's eye movement information includes: The live streaming device, acting as a live streaming terminal, sends the eye movement information to the live streaming service platform. The platform then renders the virtual character's eye movements based on this information, generating a corresponding live stream which is sent to the receiving terminal for display. The live streaming device, acting as a live streaming service platform, sends the eye movement information to the live streaming receiving terminal, which then renders and displays the eye movements of the virtual character.
4. A live interactive system, running on a live streaming device, characterized in that, The live interactive system includes: The key point extraction module is used to acquire the anchor video image and extract multiple human eye key points from the anchor video image. The multiple human eye key points include eyeball key points and multiple eyelid key points. The calibration distance extraction module is used to acquire a predetermined reference image including the anchor's eye image, extract the upper eyelid key points and the lower eyelid key points from the reference image, and calculate the distance between the upper eyelid key points and the lower eyelid key points as the first calibration distance. The motion information acquisition module is used to determine a virtual upper eyelid point as a first reference key point in the anchor video image based on the lower eyelid point, upper eyelid point, and the first calibration distance among the plurality of eyelid key points. The virtual upper eyelid point, the lower eyelid point, and the upper eyelid point are collinear, and the distance between the virtual upper eyelid point and the lower eyelid point is equal to the first calibration distance. The module also calculates the first distance between the eyeball key point and the left corner of the plurality of eyelid key points, the second distance between the eyeball key point and the right corner of the plurality of eyelid key points, and the third distance between the eyeball key point and the right corner of the plurality of eyelid key points. The third distance between the eyeball key point and the lower eyelid point among the plurality of eyelid key points, and the fourth distance between the eyeball key point and the virtual upper eyelid point; the eyeball movement coefficient of the anchor is calculated based on the first distance, the second distance, the third distance, and the fourth distance, wherein the eyeball movement coefficient includes a first movement coefficient negatively correlated with the first distance, a second movement coefficient negatively correlated with the second distance, a third movement coefficient negatively correlated with the third distance, and a fourth movement coefficient negatively correlated with the fourth distance, wherein eye movement information includes the eyeball movement coefficient; The interactive control module is used to control the virtual image displayed in the live interactive screen to perform corresponding eye interaction operations based on the anchor's eye movement information.
5. A live streaming device, characterized in that, The method includes a machine-readable storage medium and one or more processors, the machine-readable storage medium storing machine-executable instructions that, when executed by the one or more processors, implement the method of any one of claims 1-3.
Citation Information
Patent Citations
Intelligent home controller based on eye-movement tracking and intelligent home control method based on eye-movement tracking
CN105159460A
Image display method and device in live broadcast and storage medium
CN109120985A
Interaction method and device based on eye movement, medium and electronic equipment
CN110941333A