Live broadcast interaction method and device, equipment and storage medium
By processing viewers' touch operation data and skeletal key points on the broadcaster's end, special effects video frames are generated, solving the problems of low accuracy and high computing power in existing live streaming interactions, and realizing efficient live streaming special effects interactions.
Patent Information
- Application Number
- CN202511636322.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-10
AI Technical Summary
Among existing live streaming interaction methods, the accuracy of interactive effects is relatively low, the requirements for device computing power are high, and the interactive effect of live streaming effects is poor.
By acquiring the original video frames and skeletal keypoint set sent by the broadcaster, and combining them with the audience's touch operation data, special effects rendering is performed on the broadcaster's end to generate special effects video frames, allowing the audience to interact directly with the broadcaster.
It improved the accuracy of interactive effects, reduced the computing power requirements of devices, enhanced the interactive effect of live streaming effects, and increased user participation and immersion.
Smart Images

Figure CN121509741A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of image processing, and particularly relate to a live broadcast interaction method and device, equipment and a storage medium. BACKGROUND
[0002] With the development of computer technology and image processing technology, the interaction mode of the live broadcast platform is also more and more. In order to improve the activity of the host and the audience, various forms of interaction can be carried out between the host and the audience.
[0003] The live broadcast interaction mode in the related art is generally text barrage, static expression, likes and gift giving, and such a live broadcast interaction mode is one-way interaction, and the special effect of the live broadcast is usually initiated by the host or triggered randomly by the system, and the interactivity between the host and the audience is poor. The interaction scheme based on pure software image recognition (such as skin color detection and template matching) has a significant decrease in recognition accuracy when dealing with complex background, fast motion and multi-angle human body. If the interaction scheme based on machine learning is processed in the cloud, the network delay is difficult to meet the real-time interaction requirement; if it is processed in the audience end, the algorithm requirement of the audience end is high, and it is difficult to popularize in low-end devices, and the interaction effect of the live broadcast special effect is poor. SUMMARY
[0004] Embodiments of the present application provide a live broadcast interaction method, device, equipment and storage medium, by obtaining the original video frame and the skeleton key point set sent by the host end, obtaining the touch data of the touch operation on the screen, determining the touch hit area according to the touch data and the skeleton key point set, performing special effect rendering processing on the original video frame according to the touch hit area, obtaining the first special effect video frame, and playing the first special effect video frame, the touch operation of the audience on the screen can realize the special effect interaction with the host combined with the skeleton key point set of the host, and the skeleton key point set does not need to be recognized and processed in the audience end, but is obtained by unified recognition and processing in the host end, effectively solving the technical problems that the interaction special effect accuracy of the live broadcast is low, the algorithm requirement of the device is high, and the interaction effect of the live broadcast special effect is poor in the related art. The live broadcast special effect can be triggered according to the touch operation of the audience, the algorithm requirement of the device is reduced while improving the interaction special effect accuracy, and the interaction effect of the live broadcast special effect is improved.
[0005] In a first aspect, embodiments of the present application provide a live broadcast interaction method, comprising: obtaining the original video frame and the skeleton key point set sent by the host end, the skeleton key point set being obtained by the host end performing skeleton recognition processing on the original video frame; obtaining the touch data of the touch operation on the screen, and determining the touch hit area according to the touch data and the skeleton key point set; According to the touch hit region, the original video frame is processed for special effect rendering to obtain a first special effect video frame, and the first special effect video frame is played.
[0006] In a second aspect, the embodiments of the present application provide a live interaction device, comprising a key point acquisition module, a region determination module and a special effect processing module, wherein: The key point acquisition module is configured to acquire an original video frame and a skeleton key point set sent by a host end, wherein the skeleton key point set is obtained by performing skeleton recognition processing on the original video frame by the host end. The region determination module is configured to acquire touch data of a touch operation on a screen, and determine a touch hit region according to the touch data and the skeleton key point set. The special effect processing module is configured to perform special effect rendering processing on the original video frame according to the touch hit region to obtain a first special effect video frame, and play the first special effect video frame.
[0007] In a third aspect, the embodiments of the present application provide a live interaction device, comprising a memory and one or more processors. The memory is configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the live interaction method of the first aspect.
[0008] In a fourth aspect, the embodiments of the present application provide a non-volatile storage medium storing computer executable instructions, which, when executed by a computer processor, are used to execute the live interaction method of the first aspect.
[0009] In a fifth aspect, the embodiments of the present application provide a computer program product, which comprises a computer program stored in a computer readable storage medium, and at least one processor of a device reads and executes the computer program from the computer readable storage medium, so that the device executes the live interaction method of the first aspect.
[0010] The embodiment of the application obtains the original video frame and the skeleton key point set sent by the anchor end, obtains the touch data for performing a touch operation on the screen, determines the touch hit area according to the touch data and the skeleton key point set, performs special effect rendering processing on the original video frame according to the touch hit area, obtains the first special effect video frame, and plays the first special effect video frame. The audience can perform a touch operation on the screen to realize special effect interaction with the anchor in combination with the skeleton key point set of the anchor. The skeleton key point set does not need to be identified on the audience end, but is uniformly identified on the anchor end, thereby improving the accuracy of the interactive special effect, reducing the requirement for the computing power of the device, and improving the interactive effect of the live broadcast special effect. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is a flowchart of a live broadcast interaction method provided by the embodiment of the application; Figure 2 is a flowchart of another live broadcast interaction method provided by the embodiment of the application; Figure 3 is a structural schematic diagram of a live broadcast interaction device provided by the embodiment of the application; Figure 4 is a structural schematic diagram of a live broadcast interaction device provided by the embodiment of the application. DETAILED DESCRIPTION
[0012] In order to make the objectives, technical solutions and advantages of the application clearer, the following further describes the specific embodiments of the application in conjunction with the drawings. It can be understood that the specific embodiments described herein are only used to explain the application, rather than limit the application. In addition, it should be noted that, for the convenience of description, only parts related to the application are shown in the drawings, rather than all contents. Before discussing the example embodiments in more detail, it should be mentioned that some example embodiments are described as processes or methods depicted by flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The above processes can be terminated when the operations are completed, but can also have additional steps not included in the drawings. The above processes can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0013] The live broadcast interaction method provided by the application can be applied to a video live broadcast scene (for example, a show live broadcast, a game live broadcast, an outdoor live broadcast, etc.), and aims to introduce an intuitive and natural interaction mode of directly interacting with the visual image of the anchor, associate the interactive operation of the user with the visual image of the anchor, and enhance the participation of the user.
[0014] The live interaction method of the related technology is generally text barrage, static expression, likes and gift giving, etc. Such live interaction methods are one-way interaction, and the special effects of the live broadcast are usually initiated by the host or triggered randomly by the system, and the interactivity between the host and the audience is poor. The interaction in the live broadcast room is generally performed by generating special effects in the video frame, and the special effect generation can be performed by a pure software image recognition interaction scheme or a machine learning-based interaction scheme. Among them, the pure software image recognition interaction scheme has a significant decrease in recognition accuracy when processing complex backgrounds, fast motion and multi-angle human bodies. If the machine learning-based interaction scheme is processed in the cloud, the network delay is difficult to meet the real-time interaction demand; if it is processed in the audience end, the algorithm requirement of the audience end is high, and it is difficult to popularize in low-end devices, and the interaction effect of the live broadcast special effect is poor. Therefore, the live interaction method provided in the embodiments of the present application is provided to solve the technical problems of low accuracy of interactive special effects, high requirement of device algorithm, and poor interaction effect of live broadcast special effects in the existing live interaction scheme.
[0015] Figure 1 A flowchart of a live interaction method provided in the embodiments of the present application is given, and the live interaction method provided in the embodiments of the present application can be executed by a live interaction device, which can be realized by hardware and / or software and integrated in a live interaction device.
[0016] The live interaction method executed by the live interaction device is described below. Referring to Figure 1 , the live interaction method comprises: S110: acquiring an original video frame sent by a host end and a skeleton key point set, the skeleton key point set being obtained by the host end performing skeleton recognition processing on the original video frame.
[0017] Exemplarily, the original video frame sent by the host end and the skeleton key point set are acquired, wherein the skeleton key point set records the coordinates of a plurality of preset body parts of a human body in the original video frame. After receiving the original video frame, the live interaction device can play the original video frame, or when detecting a touch operation, perform special effect rendering processing on the original video frame based on the touch operation to obtain a first special effect video frame, and play the first special effect video frame for a user to watch a live broadcast picture. Optionally, the human body in the original video frame can be a real person or a virtual image. The skeleton key point set of the virtual image can be determined according to the skeleton parameters configured for the virtual image, or can be obtained by performing skeleton recognition processing on the virtual image.
[0018] The live interaction device provided in the application can be a viewer end in a video live room, and the live room also has a host end and other viewer ends. The host end collects images through a camera and generates original video frames, and sends the original video frames to the viewer ends in the live room. Optionally, the original video frames received by each viewer end can be original video frames after transcoding by a transcoding server, or original video frames without transcoding processing.
[0019] In one embodiment, after collecting the original video frames, the host end also performs skeleton recognition processing on the original video frames based on a preset skeleton recognition algorithm or a preset skeleton recognition model to obtain a set of skeleton key points. The skeleton recognition processing can be understood as real-time detection of a human body in the original video frames, positioning of the positions (two-dimensional or three-dimensional coordinates) and postures of key body parts (such as head, nose, shoulder, palm, wrist, eye, waist, knee, etc.), and formation of human skeleton key point coordinate data.
[0020] After obtaining the set of skeleton key points, the host end sends the original video frames and the set of skeleton key points to each viewer end in the live room. For example, the host end sends the original video frames and the set of skeleton key points to a server, and the server sends the original video frames and the set of skeleton key points to each viewer end in the live room. The host end can also send the original video frames and the set of skeleton key points to each viewer end in the live room in a peer-to-peer file sharing (P2P) manner.
[0021] Optionally, in addition to obtaining the set of skeleton key points by performing skeleton recognition processing on the original video frames, the host end can also obtain the confidence of each key point, and can filter out skeleton key points with a confidence less than a preset confidence threshold, thereby improving the skeleton recognition accuracy. The current set of skeleton key points can also be subjected to time sequence filtering processing (such as Kalman filtering processing) according to a plurality of previously obtained sets of skeleton key points, thereby reducing the jitter of the skeleton key points and improving the stability of special effect generation.
[0022] S120: Obtain touch data of a touch operation performed on the screen, and determine a touch hit region according to the touch data and the set of skeleton key points.
[0023] In one embodiment, when a user watches a live broadcast through the live interaction device, the user can perform a touch operation on the screen of the live interaction device to interact with the live broadcast picture. The user can perform a touch (such as a single click, a continuous click, a long press, a swipe, etc.) on the screen according to the position and posture of the human body in the live broadcast picture, for example, clicking the head, eye, shoulder, hand, etc. of the host in the live broadcast picture.
[0024] Optionally, a touch interaction prompt can be displayed on the screen to guide the user to interact with the host by tapping the screen when the user enters the live room or at a preset time. After the user triggers the touch operation, visual (e.g., playing an animation effect) and / or auditory (e.g., playing a sound effect) touch feedback can be provided to the user to provide immediate operation feedback.
[0025] Illustratively, touch data of a touch operation performed on the screen is obtained, and a touch hit region is determined according to the touch data and the set of skeletal key points. The effective region of different body parts can be determined according to the positions of different skeletal key points in the original video frame in the set of skeletal key points. According to the positional relationship between the touch position reflected by the touch data and the effective region, it can be determined whether the user hits the effective region, and the hit effective region is determined as the touch hit region.
[0026] Optionally, since the original video frame can be scaled and cropped on the live interaction device, the mapping relationship between the screen pixel points and the image pixel points of the original video frame can be determined before and after the scaling and cropping. The live interaction device can listen to the touch operation, obtain the screen coordinates of the touch point corresponding to the touch operation under the screen coordinates when the touch operation is detected, convert the screen coordinates to image coordinates of the image coordinate system corresponding to the original video frame according to the mapping relationship, and generate touch data according to the image coordinates to ensure that the interaction special effect is correctly applied to the original video frame.
[0027] S130: performing special effect rendering processing on the original video frame according to the touch hit region to obtain a first special effect video frame, and playing the first special effect video frame.
[0028] Illustratively, after the touch hit region is determined, special effect rendering processing can be performed on the original video frame according to the touch hit region to obtain a first special effect video frame. The special effect rendering processing on the original video frame can be understood as adding a preset special effect in the touch hit region of the original video frame. Optionally, different touch hit regions corresponding to different body parts can correspond to different special effects, for example, the head corresponds to a head-touch special effect, the eyes correspond to a panda-eye special effect, the shoulders correspond to a shoulder-pat special effect, etc.
[0029] After obtaining the first special effect video frame, the first special effect video frame can be played, and the live interaction device displays the image after adding the special effect to the corresponding body part of the host. The live special effect is no longer triggered by the host or the system, but can also be triggered by the user. The user changes from a passive viewer and text sender to an active interactive participant, who can directly interact with the host image, greatly enhancing the live immersion, interest, and user stickiness. At the same time, by performing skeletal recognition processing on the original video frame at the host end, pixel-level touch positioning accuracy is achieved, which makes the interactive feedback closely combined with the host posture, and the special effect generation more realistic and natural.
[0030] The above, by acquiring the original video frame and the skeleton key point set sent by the anchor end, acquiring the touch data of the touch operation on the screen, determining the touch hit area according to the touch data and the skeleton key point set, performing special effect rendering processing on the original video frame according to the touch hit area, obtaining the first special effect video frame, and playing the first special effect video frame, the audience can realize the special effect interaction with the anchor by performing the touch operation on the screen combined with the skeleton key point set of the anchor. The skeleton key point set does not need to be recognized and processed at the audience end, but is uniformly recognized and processed at the anchor end, which improves the accuracy of the interactive special effect, reduces the requirement for the computing power of the device, and improves the interactive effect of the live special effect.
[0031] On the basis of the above embodiment, Figure 2 A flowchart of another live interaction method provided by the embodiment of the application is given, which is a specific embodiment of the above live interaction method. Referring to Figure 2 The live interaction method comprises: S210: acquiring the original video frame and the skeleton key point set sent by the anchor end, the skeleton key point set being obtained by performing skeleton recognition processing on the original video frame by the anchor end.
[0032] S220: acquiring touch data of a touch operation on a screen, and determining a touch hit area according to the touch data and the skeleton key point set.
[0033] In one possible embodiment, in the live interaction method provided by the application, determining the touch hit area according to the touch data and the skeleton key point set comprises: determining an effective area of one or more preset body parts according to the skeleton key point set; determining whether the touch data hits the effective area, and determining the hit effective area as the touch hit area.
[0034] Illustratively, the effective area of one or more preset body parts is determined according to the skeleton key point set, wherein different preset body parts correspond to different skeleton key points, and different preset body parts can correspond to different numbers of skeleton key points, and the effective areas of different preset body parts can also correspond to different shapes. For example, the arm effective area can be a belt-shaped area connected by the skeleton key points corresponding to the shoulder and the wrist, the shoulder effective area can be a belt-shaped area connected by the skeleton key points corresponding to the left shoulder and the right shoulder, and the head effective area can be a circular area of a preset radius with the skeleton key point corresponding to the nose as the center.
[0035] In an embodiment, it is determined whether the touch position corresponding to the touch data is in the determined valid region, and if the touch data is in one of the valid regions, it is determined that the touch data hits the valid region, and the hit valid region is determined as the touch hit region. Alternatively, the touch hit region can also be determined based on distance nearest neighbor matching, for example, the distance (such as pixel distance, Euclidean distance, etc.) between the touch data and each skeletal key point is calculated, and the valid region corresponding to the skeletal key point with the smallest distance and less than a preset distance threshold is determined as the hit valid region. The touch data and the skeletal key point set can also be input into a trained machine model, and the touch data and the skeletal key point set are analyzed and processed by the machine model to obtain the touch hit region.
[0036] Alternatively, the touch data of one touch operation can hit one or more valid regions, and one or more touch hit regions can be determined, and then special effects will be generated for the one or more touch hit regions. According to the present application, the valid regions of one or more preset body parts are determined according to the skeletal key point set, and the touch hit region is accurately determined according to whether the touch data hits the valid region, which can more accurately generate live special effects. The identification of the skeletal key point set does not need to be performed on the live interactive device, and even devices with weak computing power can accurately and quickly identify the touch hit region and generate live special effects, effectively improving the interactive effect of live special effects.
[0037] S230: generating interactive event information according to the gesture type corresponding to the touch operation and the touch hit region.
[0038] S240: sending the interactive event information to the anchor end and / or other audience ends in the live room, for the anchor end and / or other audience ends in the live room to perform special effect rendering processing on the original video frame based on the interactive event information to obtain second special effect video frames, and playing the second special effect video frames.
[0039] For example, after determining the touch hit region according to the touch data and the skeletal key point set, the interactive event information can be generated according to the gesture type corresponding to the touch operation and the touch hit region. Alternatively, the interactive event information can also include the user identifier corresponding to the live interactive device, the timestamp, and the gesture type of the touch operation.
[0040] In an embodiment, the interactive event information is sent to the anchor end and / or other spectator ends in the live broadcast room. After receiving the interactive event information, the anchor end and / or other spectator ends in the live broadcast room can perform special effect rendering processing on the original video frame based on the interactive event information to obtain a second special effect video frame, and play the second special effect video frame. The special effect rendering processing of the anchor end and / or other spectator ends in the live broadcast room on the original video frame can refer to the special effect rendering processing of the live broadcast interactive device on the original video frame, and achieve corresponding technical effects, which will not be described herein again. According to the gesture type and the touch hit area corresponding to the touch operation, the application generates interactive event information and sends the interactive event information to the anchor end and / or other spectator ends in the live broadcast room, so that the anchor end and / or other spectator ends in the live broadcast room perform special effect rendering processing on the original video frame to obtain a second special effect video frame, and play the second special effect video frame. By synchronizing the interactive event information to each device in the live broadcast room, each device in the live broadcast room can synchronously display live broadcast special effects, so as to ensure that the interactive effects seen by each user in the live broadcast room with a large number of spectators are timely and consistent, avoid visual confusion, and improve the live broadcast interactive effect.
[0041] In a possible embodiment, the application provides a live broadcast interactive method, wherein the interactive event information is sent to the anchor end, including: sending the interactive event information to the server, for the server to perform screening processing on the received interactive event information, and sending the interactive event information after the screening processing to the anchor end and / or other spectator ends in the live broadcast room.
[0042] For example, the interactive event information is sent to the server. At the same time, the server can also receive the interactive event information uploaded by other spectator ends in the live broadcast room. After receiving the interactive event information, the server performs screening processing on the interactive event information to screen out interactive event information with conflicts, and sends the interactive event information after the screening processing to the anchor end and / or other spectator ends in the live broadcast room.
[0043] Among them, the interactive event information with conflicts can be interactive event information corresponding to the same touch hit area within a preset time window (for example, the window length can be 50ms-200ms), or can be interactive event information corresponding to touch hit areas with a preset conflict relationship. Optionally, screening out the interactive event information with conflicts can be merging the conflicting interactive event information, or can be selecting one of the conflicting interactive event information to be retained.
[0044] By performing screening processing on the received interactive event information, and then sending the interactive event information after the screening processing to the anchor end and / or other spectator ends in the live broadcast room, the application can reduce the situation that the anchor end and / or spectator end are affected by the special effects with conflicts to live broadcast experience, and ensure the live broadcast interactive effect.
[0045] Exemplarily, in the live interaction method provided in the application, the server performs screening processing on the received interaction event information, which can be: performing time sorting processing on the received interaction event information, and performing merging processing or selection processing on the interaction event information subjected to the time sorting processing based on a preset time window.
[0046] Exemplarily, the server performs sorting processing on the interaction event information according to the time stamp recorded in the received interaction event information. Through a sliding window manner, the interaction event information subjected to the time sorting processing is subjected to merging processing or selection processing based on a preset time window (for example, a time window with a window length of 100 ms). After completing the screening processing on the interaction event information, the interaction event information subjected to the screening processing can be put into a message queue, and the interaction event information subjected to the screening processing is pushed to the anchor end and the audience end (including the live interaction device and other audience ends) in the live room through a broadcast or multicast manner.
[0047] Optionally, the merging processing on the interaction event information can be to merge the interaction event information corresponding to the same touch hit region into one interaction event message, so as to reduce repeated sending and processing of the interaction event information. The selection processing on the interaction event information can be to select and retain one interaction event information (for example, the interaction event information with the highest ranking) from the multiple interaction event information in the same touch hit region. Through the merging processing or selection processing on the interaction event information subjected to the time sorting processing based on the preset time window, the application can reduce the situation that the same touch hit region is repeatedly rendered with special effects in a short time, so as to cause abnormal or lagging screen, and improve the live interaction effect.
[0048] S250: performing special effect rendering processing on the original video frame according to the touch hit region to obtain a first special effect video frame, and playing the first special effect video frame.
[0049] In one possible embodiment, in the live interaction method provided in the application, before the original video frame is subjected to special effect rendering processing according to the touch hit region to obtain a first special effect video frame, a special effect resource can also be determined according to the gesture type corresponding to the touch operation and / or the touch hit region. Different special effect resources contain anchor points corresponding to different skeletal key points, and the alignment and synchronization of the special effect and the human body action can be realized through the binding of the anchor points and the skeletal key points. Different special effect resources can correspond to different gesture types and / or touch hit regions, and the correspondence between the special effect resource and the gesture type and / or the touch hit region can be recorded through a mapping rule table (for example, a JSON configuration file).
[0050] Accordingly, in the live interaction method provided in the application, the original video frame is rendered with special effects according to the touch hit area to obtain a first special effect video frame. For example, the anchor point of the special effect resource is bound to the corresponding skeletal key point in the original video frame, and the original video frame is rendered based on the special effect resource to obtain the first special effect video frame. In the application, the special effect resource is determined according to the gesture type corresponding to the touch operation and / or the touch hit area, and the original video frame is rendered with special effects according to the special effect resource and the touch hit area to obtain the first special effect video frame. The special effect corresponding to the user touch operation is accurately generated, the user touch operation is accurately mapped to a specific body part of the host, the triggered special effect is accurately matched with the posture and action of the host, and the authenticity and interest of the live interaction are improved.
[0051] Optionally, the host end and / or the other audience end can render the original video frame with special effects based on the interactive event information. The special effect resource can be determined according to the gesture type and / or the touch hit area recorded in the interactive event information, and the original video frame can be rendered with special effects according to the special effect resource and the touch hit area to obtain a second special effect video frame.
[0052] Optionally, the original video frame used by the live interaction device, the host end and the other audience end for special effect rendering processing can be the original video frame corresponding to the touch operation, the latest original video frame to be rendered, or the latest original video frame received or generated. That is, the original video frame for rendering special effects can be after the original video frame corresponding to the touch operation. Optionally, the server can also specify the original video frame for rendering special effects, and the live interaction device, the host end and the other audience end can perform special effect rendering processing on the specified original video frame.
[0053] In one possible embodiment, the live interaction method provided in the application further includes receiving the interactive event information filtered by the server after the interactive event information is sent to the host end and / or the other audience end in the live room. For example, the server sends the filtered interactive event information to the live interaction device after filtering the received interactive event information. Accordingly, the live interaction device renders the original video frame with special effects according to the filtered interactive event information to obtain a first special effect video frame. In the application, the original video frame is rendered with special effects according to the filtered interactive event information, which can ensure that the interactive effect seen by the live interaction device, the host end and / or the other audience end is consistent, avoid visual confusion, and improve the live interaction effect.
[0054] According to the above, by acquiring the original video frame and the skeleton key point set sent by the anchor end, the touch data of the touch operation on the screen is acquired, the touch hit region is determined according to the touch data and the skeleton key point set, the original video frame is rendered and processed according to the touch hit region, the first special effect video frame is obtained, and the first special effect video frame is played. The audience can perform touch operation on the screen to realize special effect interaction with the anchor by combining the skeleton key point set of the anchor. The skeleton key point set does not need to be recognized and processed at the audience end, but is uniformly recognized and processed at the anchor end. The accuracy of the interactive special effect is improved, the requirement for the computing power of the device is reduced, and the interactive effect of the live special effect is improved. At the same time, the interactive event information is generated according to the gesture type corresponding to the touch operation and the touch hit region, and the interactive event information is sent to the anchor end and / or other audience ends in the live room, so that the original video frame is rendered and processed to obtain the second special effect video frame at the anchor end and / or other audience ends in the live room, and the second special effect video frame is played. By synchronizing the interactive event information to each device in the live room, each device in the live room can synchronously display the live special effect, so that in a live room with a large number of audiences, the interactive effect seen by each user is timely and consistent, and the situation of visual confusion is avoided, and the live interactive effect is improved.
[0055] Figure 3 A structure schematic diagram of a live interactive device provided by an embodiment of the application is given. Referring to Figure 3 The live interactive device includes a key point acquisition module 31, a region determination module 32, and a special effect processing module 33.
[0056] The key point acquisition module 31 is configured to acquire the original video frame and the skeleton key point set sent by the anchor end, and the skeleton key point set is obtained by the anchor end performing skeleton recognition processing on the original video frame. The region determination module 32 is configured to acquire the touch data of the touch operation on the screen, and determine the touch hit region according to the touch data and the skeleton key point set. The special effect processing module 33 is configured to render and process the original video frame according to the touch hit region, obtain the first special effect video frame, and play the first special effect video frame.
[0057] According to the above, by acquiring the original video frame and the set of skeleton key points sent by the anchor end, the touch data of the touch operation on the screen is acquired, the touch hit region is determined according to the touch data and the set of skeleton key points, the original video frame is rendered and processed according to the touch hit region, the first special effect video frame is obtained, and the first special effect video frame is played. The audience can perform touch operation on the screen to realize special effect interaction with the anchor by combining the set of skeleton key points of the anchor. The set of skeleton key points does not need to be recognized and processed at the audience end, but is uniformly recognized and processed at the anchor end. The accuracy of the interactive special effect is improved, the requirement for the computing power of the device is reduced, and the interactive effect of the live special effect is improved.
[0058] In one possible embodiment, the region determination module 32 determines the touch hit region according to the touch data and the set of skeleton key points, and is configured to: determine the effective region of one or more preset body parts according to the set of skeleton key points; determine whether the touch data hits the effective region, and determine the hit effective region as the touch hit region.
[0059] In one possible embodiment, the live interactive device further includes a resource determination module, which is configured to: determine the special effect resource according to the gesture type corresponding to the touch operation and / or the touch hit region; Correspondingly, the special effect processing module 33 renders and processes the original video frame according to the touch hit region to obtain the first special effect video frame, and is configured to: render and process the original video frame according to the special effect resource and the touch hit region to obtain the first special effect video frame.
[0060] In one possible embodiment, the live interactive device further includes a message synchronization module, which is configured to: generate interactive event information according to the gesture type corresponding to the touch operation and the touch hit region; send the interactive event information to the anchor end and / or other audience ends in the live room, so that the anchor end and / or other audience ends in the live room render and process the original video frame according to the interactive event information to obtain the second special effect video frame, and play the second special effect video frame.
[0061] In one possible embodiment, the message synchronization module sends the interactive event information to the anchor end, and is configured to: send the interactive event information to the server, so that the server performs screening processing on the received interactive event information, and sends the screened interactive event information to the anchor end and / or other audience ends in the live room.
[0062] In a possible embodiment, the server performs the screening processing on the received interaction event information, and is configured to: perform time sorting processing on the received interaction event information, and perform merging processing or selection processing on the interaction event information after the time sorting processing based on a preset time window.
[0063] In a possible embodiment, the live interaction apparatus further includes an information receiving module, which is configured to: receive the screened interaction event information sent by the server; Accordingly, the special effect processing module 33 performs special effect rendering processing on the original video frame according to the touch hitting area to obtain a first special effect video frame, and is configured to: perform special effect rendering processing on the original video frame according to the screened interaction event information to obtain a first special effect video frame.
[0064] It should be noted that, in the embodiments of the above live interaction apparatus, each unit and module included is only logically divided according to function, but is not limited to the above division, as long as the corresponding function can be implemented; in addition, the specific names of each functional unit are only for convenient mutual distinction, and do not serve to limit the protection scope of the embodiments of the present application.
[0065] The embodiments of the present application also provide a live interaction device, which can integrate the live interaction apparatus provided by the embodiments of the present application. Figure 4 FIG. 1 is a structural schematic diagram of a live interaction device provided by an embodiment of the present application. Referring to FIG. 1, Figure 4 The live interaction device includes an input apparatus 43, an output apparatus 44, a memory 42, and one or more processors 41; the memory 42 is used to store one or more programs; when the one or more programs are executed by the one or more processors 41, the one or more processors 41 implement the live interaction method provided by the above embodiments. The live interaction apparatus, device, and computer provided above can be used to execute the live interaction method provided by any of the above embodiments, and have the corresponding functions and beneficial effects.
[0066] The embodiment of the present application further provides a nonvolatile storage medium storing computer executable instructions, which, when executed by a computer processor, are used to execute the live interaction method provided by the above embodiment. Of course, the nonvolatile storage medium storing computer executable instructions provided by the embodiment of the present application is not limited to the live interaction method provided above, and can also execute the related operations in the live interaction method provided by any embodiment of the present application. The live interaction apparatus, device and storage medium provided in the above embodiment can execute the live interaction method provided by any embodiment of the present application, and the technical details not described in detail in the above embodiment can be referred to the live interaction method provided by any embodiment of the present application.
[0067] On the basis of the above embodiment, the embodiment of the present application further provides a computer program product, the technical solution of the present application or the part of the contribution to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes a plurality of instructions for causing a computer device, a mobile terminal or a processor therein to execute all or part of the steps of the live interaction method provided by each embodiment of the present application.
Claims
1. A live streaming interactive method, characterized in that, include: The original video frames and the set of skeletal key points sent by the broadcaster are obtained. The set of skeletal key points is obtained by the broadcaster performing skeletal recognition processing on the original video frames. Acquire touch data of touch operations performed on the screen, and determine the touch hit area based on the touch data and the set of skeletal key points; The original video frame is processed with special effects rendering based on the touch hit area to obtain a first special effects video frame, and the first special effects video frame is played.
2. The live streaming interactive method according to claim 1, characterized in that, The step of determining the touch hit area based on the touch data and the set of skeletal key points includes: Based on the set of skeletal key points, determine the effective area of one or more preset body parts; Determine whether the touch data hits the valid area, and define the hit valid area as the touch hit area.
3. The live streaming interaction method according to claim 1, characterized in that, Before performing special effects rendering on the original video frame based on the touch hit area to obtain the first special effects video frame, the method further includes: Special effects resources are determined based on the gesture type corresponding to the touch operation and / or the touch hit area; Accordingly, the step of performing special effects rendering processing on the original video frame based on the touch hit area to obtain the first special effects video frame includes: Based on the special effects resources and the touch hit area, the original video frame is processed to perform special effects rendering to obtain the first special effects video frame.
4. The live streaming interactive method according to claim 1, characterized in that, After determining the touch hit area based on the touch data and the set of skeletal key points, the method further includes: Interactive event information is generated based on the gesture type corresponding to the touch operation and the touch hit area; The interactive event information is sent to the broadcaster's terminal and / or other viewers' terminals in the live broadcast room, so that the broadcaster's terminal and / or other viewers' terminals in the live broadcast room can perform special effects rendering processing on the original video frame based on the interactive event information to obtain a second special effects video frame, and play the second special effects video frame.
5. The live interactive method according to claim 4, characterized in that, Sending the interactive event information to the broadcaster includes: The interactive event information is sent to the server, whereby the server filters and processes the received interactive event information and sends the filtered interactive event information to the broadcaster's terminal and / or other viewers' terminals in the live broadcast room.
6. The live streaming interactive method according to claim 5, characterized in that, The server performs filtering processing on the received interactive event information, including: The received interactive event information is sorted by time, and the sorted interactive event information is merged or selected based on a preset time window.
7. The live streaming interactive method according to claim 5, characterized in that, After sending the interactive event information to the broadcaster's terminal and / or other viewers' terminals in the live broadcast room, the method further includes: Receive the filtered interactive event information sent by the server; Accordingly, the step of performing special effects rendering processing on the original video frame based on the touch hit area to obtain the first special effects video frame includes: Based on the filtered interactive event information, the original video frame is subjected to special effects rendering to obtain the first special effects video frame.
8. A live interactive device, characterized in that, It includes a key point acquisition module, a region determination module, and a special effects processing module, among which: The key point acquisition module is configured to acquire the original video frame and the set of skeletal key points sent by the broadcaster, wherein the set of skeletal key points is obtained by the broadcaster performing skeletal recognition processing on the original video frame. The region determination module is configured to acquire touch data of touch operations performed on the screen, and determine the touch hit region based on the touch data and the set of skeletal key points; The special effects processing module is configured to perform special effects rendering processing on the original video frame according to the touch hit area to obtain a first special effects video frame, and play the first special effects video frame.
9. A live streaming interactive device, characterized in that, include: Memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the live interactive method as described in any one of claims 1-7.
10. A non-volatile storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the live interactive method as described in any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the live interactive method according to any one of claims 1-7.