Video display method and device, electronic equipment, storage medium and product

By recognizing and scoring user action videos and combining attribute information to filter target videos, the problem of strong subjectivity in scoring in live streaming scenarios is solved, and more interactive and interesting video displays are achieved.

CN121815008APending Publication Date: 2026-04-07MIGU INTERACTIVE ENTERTAINMENT CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing video display solutions cannot effectively quantify the relevance of user selfies to the played videos in live streaming scenarios, resulting in highly subjective ratings and failing to showcase outstanding audience videos to other users in the virtual space.

Method used

By acquiring action videos from multiple users, action recognition is performed to obtain user action data. This data is then compared with standard action data to analyze the degree of matching and generate an objective action score. Based on the action score and user attribute information, target videos are selected and pushed to the client associated with the virtual space for display.

Benefits of technology

It reduces the subjectivity of ratings, improves user experience, enhances the interactivity and fun of live streaming, showcases excellent action videos in the virtual space, and increases the interactivity and appeal of live streaming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815008A_ABST
    Figure CN121815008A_ABST
Patent Text Reader

Abstract

The invention provides a video display method and device, electronic equipment, a storage medium and a product. The method comprises the following steps: acquiring action videos of a plurality of users; performing action recognition on the action videos to obtain user action data corresponding to each action video; performing matching degree analysis on the user action data and standard action data to obtain an action score of each user; determining a target video from the plurality of action videos based on the action score and the attribute information of each user; and pushing the target video to a client associated with a virtual space for display.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of video display, and in particular to a video display method and device, an electronic device, a storage medium and a product. BACKGROUND

[0002] With the rapid development of immersive interactive technologies such as virtual reality, augmented reality and metaverse, real-time interactive live streaming scenarios based on virtual space are increasingly popular. In the video display scheme of related live streaming scenarios, feedback is given by comparing the consistency of the content of the played video and the user's selfie content, but the score determined based on the consistency of the content of the played video and the user's selfie content is highly subjective, and the related video display scheme cannot display excellent action videos to other users in the live streaming room. SUMMARY

[0003] The present disclosure provides a video display method, device, electronic device, storage medium and product to solve the problems in the related art.

[0004] A first aspect embodiment of the present disclosure provides a video display method, which comprises: obtaining action videos of multiple users; performing action recognition on the action videos to obtain user action data corresponding to each action video; performing matching degree analysis on the user action data and standard action data to obtain an action score of each user; determining a target video from the multiple action videos based on the action score and attribute information of each user; pushing the target video to a client associated with a virtual space for display.

[0005] In an embodiment, the matching degree analysis on the user action data and the standard action data to obtain the action score of each user comprises: based on human body skeleton key point information, analyzing the similarity in spatial form of the user action data and the standard action data to obtain a spatial similarity score; aligning the time sequence of the user action data and the standard action data to obtain a time synchronization score corresponding to each action video; comparing the action stages in the user action data and the standard action data to determine an action completeness score corresponding to each action video; based on the spatial similarity score, the synchronization score and the completeness score, determining the action score of each user.

[0006] In an embodiment, the determination of the target video from the multiple action videos based on the action score and the attribute information of each user comprises: Filtering, from the plurality of action videos, a candidate video meeting a target condition based on the action scores; Prioritizing, from high to low, users in the candidate video based on attribute information of each user; Determining, from the candidate video, a target video according to the prioritization result.

[0007] In an embodiment, the target video is pushed to a client associated with the virtual space for display, including: Obtaining an initial display duration of the target video; Determining a target display duration and a display frequency based on the action scores, the interaction data in the attribute information, and the initial display duration; Pushing the target video to the client associated with the virtual space and displaying the target video according to the target display duration and the display frequency.

[0008] In an embodiment, the action videos of the plurality of users are obtained, including: Sending a follow-up practice participation request to the plurality of users; In response to receiving request confirmation information of the plurality of users, sending a recording authorization request to a user sending the request confirmation information; In response to receiving recording authorization confirmation information of the plurality of users, obtaining action videos of the plurality of users sending the recording authorization confirmation information.

[0009] In an embodiment, after the target video is pushed to the client associated with the virtual space for display, the method provided by the present disclosure includes: Obtaining evaluation information for the target video and pushing the evaluation information to the client associated with the virtual space; Issuing a virtual reward to a user account corresponding to the target video.

[0010] A second aspect embodiment of the present disclosure provides a video display device, including: An obtaining unit configured to obtain action videos of a plurality of users; An identifying unit configured to perform action recognition on the action videos to obtain user action data corresponding to each action video; An analyzing unit configured to perform matching degree analysis on the user action data and standard action data to obtain an action score of each user; A determining unit configured to determine a target video from the plurality of action videos based on the action scores and attribute information of each user; A display unit configured to push the target video to a client associated with a virtual space for display.

[0011] A third aspect embodiment of the present disclosure provides an electronic device, including: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described in the first aspect of this disclosure.

[0012] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described in the first aspect of this disclosure.

[0013] A fifth aspect of this disclosure provides a computer program product including a computer program that, when executed by a processor, implements the methods described in the first aspect of this disclosure.

[0014] In summary, this disclosure proposes a video display method, which includes: acquiring action videos of multiple users; performing action recognition on the action videos to obtain user action data corresponding to each action video; performing matching degree analysis on the user action data and standard action data to obtain an action score for each user; determining a target video from the multiple action videos based on the action score and the attribute information of each user; and pushing the target video to a client associated with the virtual space for display.

[0015] According to the solution provided in this disclosure, by acquiring action videos from multiple users; performing action recognition on the action videos to obtain user action data corresponding to each action video; analyzing the matching degree between the user action data and standard action data to obtain an action score for each user, the subjectivity of the score can be reduced and the user experience improved; based on the action score and the attribute information of each user, the target video is determined from the multiple action videos; and the target video is pushed to the client associated with the virtual space for display, which can show excellent action videos to users in the virtual space and improve the interactivity and fun of the live broadcast.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0018] Figure 1 A flowchart illustrating the video display method provided in this embodiment of the disclosure; Figure 2 A flowchart illustrating a method for determining an action score for each user according to an embodiment of this disclosure; Figure 3 A flowchart illustrating the method for determining a target video provided in an embodiment of this disclosure; Figure 4 This is a flowchart illustrating a method for pushing a target video to a client associated with a virtual space for display, as provided in an embodiment of this disclosure. Figure 5 A flowchart illustrating a method for acquiring motion videos of multiple users according to an embodiment of this disclosure; Figure 6 This is a schematic diagram of the structure of the video display device provided in the embodiments of this disclosure; Figure 7 This is a schematic diagram of the hardware composition structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0019] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.

[0020] With increasing health awareness and rapid technological development, people's demand for fitness is growing, and they also have higher expectations for innovative and fun fitness methods.

[0021] The following is a brief introduction to one video display method in related technologies: The solution proposes a video interaction method that displays live video and viewer video on the same screen, identifies viewer posture data, determines whether the viewer's posture matches that in the live video, and generates evaluation information, thereby improving viewer participation and immersion, and effectively enhancing the user experience of online video viewers through video interaction.

[0022] The above solution has the following drawbacks: This method uses open-source human pose recognition to compare the content of the played video with the viewer's selfie to provide feedback. However, it lacks clear quantitative indicators for the match between the viewer's selfie and the played video content, leading to subjectivity and uncertainty in the scoring. Furthermore, it cannot showcase outstanding viewer videos to other viewers in the virtual space.

[0023] To address the shortcomings of related technologies, this disclosure involves acquiring action videos from multiple users; performing action recognition on the action videos to obtain user action data corresponding to each action video; analyzing the matching degree between the user action data and standard action data to obtain an action score for each user, which reduces the subjectivity of the score and improves the user experience; determining the target video from multiple action videos based on the action score and each user's attribute information; and pushing the target video to a client associated with the virtual space for display, which can showcase excellent action videos to users in the virtual space and enhance the interactivity and fun of live streaming.

[0024] The present disclosure will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0025] The video display method provided in this disclosure can be applied to real-time audio and video interactive scenarios where users need to follow standard movements to imitate, practice, and achieve interactive feedback. For example, it can be applied to online fitness / yoga / dance live streaming, rehabilitation training guidance live streaming, and imitation of virtual character movements, such as hand gesture dance. The executing entity of the method can be the server of the live streaming platform.

[0026] like Figure 1 As shown, Figure 1 This is a flowchart illustrating the video display method provided in this embodiment of the disclosure. The video display method provided in this embodiment includes the following steps: Step 101: Obtain action videos from multiple users; In one embodiment, in a fitness live-streaming scenario, the user refers to the terminal operator participating in the live stream, which can be an individual with active follow-up behavior, such as a viewer or student. The user submits a video of the movements through their terminal device. This video of the movements is video data recorded and uploaded by the user during the live stream, containing the process of their physical movements; it is typically real-time captured video.

[0027] In one embodiment, after a user clicks to participate in the follow-up exercise on the live streaming interface, the client receives video captured by the camera via Real-Time Messaging Protocol (RTMP) or Secure Reliable Transport (SRT), thereby obtaining video of the actions of multiple users.

[0028] In one embodiment, for the same user, there may be only one action video or multiple action videos from different perspectives. An action video may contain only one user or multiple users.

[0029] Step 102: Perform motion recognition on the motion videos to obtain user motion data corresponding to each motion video; In one embodiment, computer vision technology, such as human pose estimation, is used to extract human key points from motion videos, thereby obtaining user motion data corresponding to each motion video. The user motion data is used to indicate the sequence of motion postures performed by the user, which can usually be represented as a set of key point coordinates in a time series, such as the coordinates of 17 skeletal points in each frame.

[0030] In one embodiment, models such as MoveNet, HRNet, and OpenPose can be used to detect human key points in each motion video frame by frame to obtain user motion data corresponding to each motion video.

[0031] In one embodiment, the motion videos can also be sent to a server cluster of graphics processing units (GPUs) for batch processing to obtain user motion data corresponding to each motion video.

[0032] In one embodiment, if a user has multiple action videos from different perspectives, action recognition can be performed on each action video, thereby improving the accuracy of action recognition.

[0033] In one embodiment, by performing motion recognition on motion videos to obtain user motion data corresponding to each motion video, visual information can be converted into a computable numerical sequence.

[0034] Step 103: Analyze the matching degree between user action data and standard action data to obtain the action score for each user; In one embodiment, the standard action data is a standardized action template that is predefined or provided in real time by the content provider, such as the broadcaster or course designer. It is also represented by a key point time series and serves as a scoring benchmark.

[0035] In one embodiment, the matching degree analysis of user action data and standard action data refers to comparing user action data and standard action data in terms of spatial form, temporal rhythm, and completeness of action stages, and quantifying the matching degree between the two through action scoring. The action score is used to measure the matching degree between user actions and standard actions, and is usually a normalized comprehensive score, such as 0 to 100 points.

[0036] In one embodiment, by analyzing the degree of matching between user action data and standard action data, an action score for each user can be obtained, thereby generating an objective evaluation index.

[0037] Step 104: Based on the action score and each user's attribute information, determine the target video from multiple action videos; In one embodiment, user attribute information refers to data related to user identity or behavior, including but not limited to age, gender, geographical location, whether it is a new user, user level, and frequency of historical interactions such as likes, comments, and gifts.

[0038] In one embodiment, the target video is a user action video that has been selected after screening for public display in virtual space.

[0039] In one embodiment, candidate action videos for departments can be filtered based on action scores, and then action videos for virtual space display, i.e., target videos, can be determined based on user attribute information.

[0040] Step 105: Push the target video to the client associated with the virtual space for display.

[0041] In one embodiment, the virtual space is a live broadcast room, and the client associated with the virtual space is all user terminal devices that are watching the same live broadcast. The user terminal devices can be mobile phones, web browsers, smart TVs, etc.

[0042] In one embodiment, by pushing the target video to a client associated with the virtual space for display, social incentives and demonstration effects can be achieved, enhancing the interactivity and fun of the live stream.

[0043] In one embodiment, motion videos from multiple users are acquired, the video streams are parsed, and the user motion data contained within is extracted. By analyzing the matching degree between the user motion data and standard motion data, an objective motion score is generated for each user, eliminating the interference of subjective judgment. Based on this, the highest quality target video is determined from multiple motion videos, taking into account both the motion score and associated user attribute information. The selected target video is then pushed to clients associated with the virtual space via a content distribution interface. For example, in a live-streamed practice application scenario, the standard motion data is the coach's demonstration motions. Motion scores are generated by comparing the motion videos of users (live-stream viewers), and selected practice videos from some viewers are displayed in the virtual space.

[0044] By acquiring action videos from multiple users, performing action recognition on these videos to obtain user action data for each video, and analyzing the matching degree between user action data and standard action data to obtain an action score for each user, the subjectivity of the scoring can be reduced, improving the user experience. Based on the action score and each user's attribute information, a target video is determined from multiple action videos. The target video is then pushed to a client associated with the virtual space for display, showcasing excellent action videos to users in the virtual space, enhancing the interactivity and fun of the live stream, thereby increasing the attractiveness of the live stream interaction and the user experience.

[0045] In one embodiment, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating a method for determining the action score of each user according to an embodiment of the present disclosure. The method involves analyzing the matching degree between user action data and standard action data to obtain the action score for each user, including: Step 201: Based on the key point information of the human skeleton, analyze the similarity between the user's action data and the standard action data in terms of spatial morphology to obtain a spatial similarity score. In one embodiment, human skeletal key point information refers to the human joint coordinate data extracted from the video by the pose estimation algorithm, which typically includes 14 to 17 key points such as head, shoulder, elbow, wrist, hip, knee, and ankle, used to characterize human posture.

[0046] In one embodiment, spatial morphology refers to the relative positions and structural relationships of key points of the human body in two-dimensional or three-dimensional space in a certain frame of the motion video, which is used to reflect whether the shape of the motion is standard.

[0047] In one embodiment, the spatial similarity score is a quantitative indicator that measures the degree of consistency between user action data and standard action data in spatial configuration. The higher the spatial similarity score, the closer the user action data and standard action data are in spatial configuration.

[0048] In one embodiment, before analyzing the spatial similarity between user action data and standard action data to obtain a spatial similarity score, key points in the user action data can be normalized and scaled to a unit scale to eliminate differences in shooting distance and human body size.

[0049] In one embodiment, different weights can be assigned to different key points in user action data, such as hands and legs being more important in dance. The weighted average distance is calculated and then mapped to a spatial similarity score.

[0050] In one embodiment, a deep learning model can also be used, such as directly deriving a spatial similarity score between user action data and standard action data.

[0051] Step 202: Align the time series of user action data with standard action data to obtain the time synchronization score for each action video; In one embodiment, a time series refers to a sequence of motion key point data arranged in chronological order, such as an array of key point coordinates for each frame, used to reflect the dynamic evolution of the motion.

[0052] In one embodiment, the time synchronization score is an indicator that measures whether the rhythm, start and end times of a user's actions are synchronized with standard actions, such as determining whether the user's actions are correct but too fast or too slow.

[0053] In one embodiment, the Dynamic Time Warping (DTW) algorithm can be used to align the time series of user action data with standard action data to obtain a time synchronization score for each action video.

[0054] In one embodiment, the time series of user action data and standard action data can also be aligned based on key event alignment to obtain a time synchronization score for each action video. Specifically, the iconic events in the action, such as raising both hands to the highest point, are identified, these time points are forcibly aligned, and the deviation of the remaining parts is calculated to obtain a time synchronization score for each action video.

[0055] Step 203: Compare the user action data with the action stages in the standard action data to determine the action integrity score corresponding to each action video; In one embodiment, the action phase is a complete action divided into several logical sub-units, such as start-transition-apex-fall, with each phase corresponding to a specific posture or motion feature.

[0056] In one embodiment, the action completeness score is an indicator that measures whether the user has fully performed all the necessary stages included in the standard action data, thus avoiding situations where the action is only half performed but the score is high.

[0057] In one embodiment, keyframe detection technology is used to identify the action stages in user action data, check the completion status of key action stages in user action data and standard action data, and determine the action integrity score corresponding to each action video.

[0058] Step 204: Determine the action score for each user based on spatial similarity score, synchronization score, and integrity score.

[0059] In one embodiment, spatial similarity score, synchronization score and integrity score can be accumulated to obtain the action score for each user.

[0060] In one embodiment, the spatial similarity score, synchronization score, and integrity score can be weighted and summed to obtain each user's action score. The weights for the spatial similarity score, synchronization score, and integrity score can be pre-set fixed values, with the sum of the weights being 1, or they can be dynamically adjusted based on the action category. For example, if the action video is a dance video, emphasizing spatiality, a larger weight is assigned to the spatial similarity score; similarly, if the action video is a Tai Chi video, emphasizing tempo, a larger weight is assigned to the synchronization score.

[0061] In one embodiment, a total score is calculated using a weighted average method, taking into account spatial similarity, synchronicity, and integrity. Each dimension can be assigned a different weight; for example, spatial similarity accounts for 50%, synchronicity for 30%, and integrity for 20%. If user A has a spatial similarity score of 80, a synchronicity score of 70, and an integrity score of 90, then according to the weighted average method, user A's action score is: 80 × 50% + 70 × 30% + 90 × 0.2 = 79 points.

[0062] In one embodiment, spatial morphological similarity analysis is performed based on human skeletal keypoint information extracted from the video stream. Specifically, this involves calculating the Euclidean distance or cosine similarity between the user's skeletal keypoints and the corresponding keypoints of the standard action in spatial coordinates, followed by coordinate normalization to eliminate the influence of individual body shape differences, thereby obtaining a spatial similarity score. Next, temporal synchronization analysis is performed. A dynamic time warping algorithm is used to non-linearly align the user's action sequence with the standard action sequence to address the issue of inconsistent action rhythms. After alignment, a synchronization score is calculated based on the time offset of keyframes. Finally, action integrity analysis is performed. Keyframe detection technology is used to identify key stages such as the start, middle, and end of the user's action sequence and matches them with preset stages of the standard action. An integrity score is calculated based on the proportion of completed stages. The final action score is obtained by weighted fusion of the above three scores according to preset weights.

[0063] By introducing refined quantitative analysis based on three dimensions—spatial similarity, temporal synchronization, and completeness—the one-sidedness of single-dimensional evaluation is overcome, and the objectivity and accuracy of action scoring are improved.

[0064] In one embodiment, such as Figure 3 As shown, Figure 3 This is a flowchart illustrating a method for determining a target video provided in an embodiment of the present disclosure. Based on action scores and attribute information of each user, the method determines a target video from multiple action videos, including: Step 301: Based on the action score, select candidate videos that meet the target conditions from multiple action videos; In one embodiment, the target condition is a preset filtering condition used to filter low-quality videos. The target condition is usually expressed in the form of an action score threshold, such as an action score of not less than 80. The candidate videos are a set of user action videos that meet the target condition.

[0065] In one embodiment, the action scoring threshold can be adaptively adjusted based on the current number of participants, such as selecting the top 30% of users, to avoid having too many action videos during peak periods or no action videos available during off-peak periods.

[0066] In one embodiment, different difficulty levels of an action can correspond to different action scoring thresholds. For example, a score of 70 is sufficient for a high-difficulty action, while a score of 90 is required for a simple action.

[0067] In one embodiment, by filtering candidate videos that meet the target conditions from multiple action videos based on action scores, the size of the candidate videos can be controlled to ensure the quality of the action videos displayed.

[0068] Step 302: Based on the attribute information of each user, sort the users in the candidate videos from high to low priority; In one embodiment, if multiple users have the same action score in the candidate video, the priority ranking can be determined based on whether the user is a new user. Specifically, in order to retain new users, new users can be given higher priority.

[0069] In one embodiment, if multiple users in a candidate video have the same action score, the priority ranking can be determined based on the user's historical interaction frequency in the live broadcast room. Specifically, in order to ensure fan loyalty, users with more interaction frequency can be given higher priority.

[0070] In one embodiment, if multiple users in a candidate video have the same action score, the priority ranking can be determined based on the user's age and geographical location.

[0071] Step 303: Determine the target video from the candidate videos based on the priority ranking results.

[0072] In one embodiment, the target video is one or more action videos that have been sorted and ultimately selected for public display in the live broadcast room.

[0073] In one embodiment, the action video of the user with the highest action score can be determined as the target video based on the known number of display slots.

[0074] In one embodiment, after determining the action video of the user with the highest action score as the target video, if there are still display slots, the action videos of new users or users with high interaction frequency can be determined as the target videos.

[0075] In one embodiment, a minimum action score threshold, such as 80 points, is set, and only user videos with action scores higher than this threshold are included in the candidate video set. The attribute information of users in the candidate set is obtained. First, within the same score range, users with high interactivity are prioritized; if interactivity is similar, new users are prioritized for display; further, users can be grouped by age group or region to ensure that each user group has a display opportunity. Finally, a comprehensive priority ranking list is generated, and the target video for final display is determined from the candidate videos based on this list.

[0076] By sorting users by multiple dimensions of user attributes, not only is the basic quality of the displayed videos guaranteed, but the diversity and fairness of the display strategy are also achieved.

[0077] In one embodiment, such as Figure 4 As shown, Figure 4 This is a flowchart illustrating a method for pushing a target video to a client associated with a virtual space for display, as provided in an embodiment of this disclosure. Pushing the target video to a client associated with a virtual space for display includes: Step 401: Obtain the initial display duration of the target video; In one embodiment, the initial display duration is a preset baseline display time, such as 10 seconds, which is usually determined by the platform's operation strategy or the type of live stream. For example, for a fitness live stream, the initial display duration of the target video is set to 12 seconds; for a dance teaching live stream, the initial display duration of the target video is set to 8 seconds; and for a rehabilitation training live stream, the initial display duration of the target video is set to 15 seconds.

[0078] In one embodiment, the initial display duration of the target video can also be dynamically determined according to time periods. For example, during peak periods when there are many users, the initial display duration can be set to 8 seconds to speed up the rotation, while during off-peak periods when there are few users, the initial display duration can be set to 15 seconds to enhance the sense of presence.

[0079] Step 402: Based on the action score, interaction data in the attribute information, and the initial display duration, determine the target display duration and display frequency; In one embodiment, the interaction data in the attribute information refers to the user's behavioral data during the live broadcast, such as the number of likes, the frequency of comments, the gift-giving record, the dwell time, and the sharing behavior, which are used to measure the user's activity and participation in the virtual space.

[0080] In one embodiment, the target display duration is the actual single playback duration determined dynamically for a specific target video, and is usually longer than the initial display duration.

[0081] In one embodiment, the display frequency refers to the number of times a target video is displayed in a carousel per unit of time, such as per minute. It determines the density at which the video is repeatedly displayed in the carousel queue. The higher the display frequency, the more popular the target video is with users.

[0082] In one embodiment, the initial display duration T is 10 seconds. If the action score is higher than a preset benchmark score... This will increase the additional display time. If the audience's interaction data is higher than the preset baseline interaction data This will increase the additional display time. Based on action ratings, interaction data in attribute information, and initial display duration, the mathematical expressions for display duration D and display frequency F are as follows:

[0083]

[0084] Among them, the rating adjustment factor and the interaction adjustment factor are preset parameters used to adjust the display duration based on the audience's rating and interactivity. The audience rating is the aforementioned action rating, and the audience interactivity is a value determined based on the user's activity level.

[0085] Step 403: Push the target video to the client associated with the virtual space and display it according to the target display duration and display frequency.

[0086] In one embodiment, a preset base display duration is used as the initial display duration, such as 10 seconds. Based on this initial display duration, dynamic adjustments are made according to the audience's action rating for the video. For example, if the rating is higher than a preset baseline score, additional display time is added proportionally. Simultaneously, if the audience's real-time interactivity data, such as the number of gifts received per unit time, exceeds a preset interaction threshold, the display time is also increased accordingly. The display frequency is determined based on a function positively correlated with rating and interactivity, such as a frequency coefficient that is the product of (audience rating / rating threshold) and (audience interactivity / interaction threshold). According to the calculated target display duration and frequency, the target video is pushed and displayed in a carousel interface in the virtual space.

[0087] By dynamically linking display parameters with user action ratings and interaction data, personalized and motivating content display is achieved, which can give high-quality, highly interactive users longer exposure time and higher frequency of appearance, and can continuously motivate users' enthusiasm for participation.

[0088] In one embodiment, such as Figure 5 As shown, Figure 5 This is a flowchart illustrating a method for acquiring motion videos of multiple users according to an embodiment of the present disclosure. The method for acquiring motion videos of multiple users includes: Step 501: Send follow-up participation requests to multiple users; In one embodiment, during the live stream, requests to participate in the practice session are sent to multiple users. The live stream refers to the online audio and video session during which the host is streaming in real time and viewers can watch and interact synchronously.

[0089] In one embodiment, the follow-up participation request is an interactive prompt sent by the live streaming platform server to the audience (user), inviting them to participate in practicing by following the host's actions. It is usually presented in the form of a pop-up window, button or voice prompt.

[0090] In one embodiment, a follow-up request can be sent to all currently online users, or it can be sent only to users who are currently online and have been in the live stream for more than five minutes.

[0091] Step 502: In response to receiving confirmation requests from multiple users, a recording authorization request is sent to the user who sent the confirmation request. In one embodiment, the request confirmation information is a response signal returned by the user through client operation, such as clicking "I want to practice", indicating that the user is willing to participate in this practice activity.

[0092] In one embodiment, the recording authorization request is a privacy compliance prompt from the live streaming platform server after the user confirms participation, requesting authorization to use the camera / microphone to collect audio and video data.

[0093] Step 503: In response to receiving recording authorization confirmation information from multiple users, obtain the action videos of the multiple users who sent the recording authorization confirmation information.

[0094] In one embodiment, the recording authorization confirmation information is a response in which the user explicitly agrees to authorize the collection of audio and video, such as clicking "Allow" or "Accept".

[0095] In one embodiment, during the live stream, the application pushes an invitation to all online viewers to participate in the practice session via a graphical interface. Once a user confirms participation, the application further triggers a system-level authorization request to access the device's camera and microphone. Only after the user explicitly authorizes the request does the application initiate video recording to capture the user's practice movements. This ensures that user data collection strictly adheres to privacy protection principles.

[0096] In one embodiment, video data can be locally optimized before uploading, such as noise reduction and enhancement, to ensure the accuracy of subsequent motion analysis.

[0097] In one embodiment, after pushing the target video to a client associated with the virtual space for display, the video display method includes: Obtain evaluation information for the target video and push the evaluation information to the client associated with the virtual space; Distribute virtual rewards to the user accounts corresponding to the target video.

[0098] In one embodiment, the evaluation information refers to the feedback content input by the anchor for a target video, which can be in the form of text, voice, emoticons, preset tags, standards, encouragement, etc., and is used to express affirmation, suggestions or interaction.

[0099] In one embodiment, the user account corresponding to the target video refers to the unique identity identifier registered by the user who submitted the target video on the platform, such as a user identifier (UID).

[0100] In one embodiment, virtual rewards are non-physical incentives issued by the live streaming platform, such as points, badges, virtual gifts, level experience, vouchers, exclusive titles, etc., which can be used to enhance users' sense of honor or redeem benefits.

[0101] In one embodiment, evaluation information can be pushed to the client associated with the virtual space via WebSocket, Message Queuing Telemetry Transport (MQTT), or a private long connection.

[0102] In one embodiment, when the target video begins to be displayed in a loop in the virtual space, the live streaming platform provides a dedicated interactive interface for the broadcaster, allowing the broadcaster to generate evaluation information for the currently displayed video, such as action standards, by clicking buttons or using voice input. This evaluation information is synchronized to all virtual space clients in real time in the form of bullet comments or floating layers. Simultaneously, the live streaming platform's backend automatically triggers reward distribution logic, distributing a certain number of virtual gifts or points to the user account corresponding to the displayed video according to preset rules, such as the first time the video is displayed or its rating level.

[0103] By introducing a dual mechanism of real-time commentary from the host and automatic reward distribution, a simple video presentation is transformed into a positive and interactive event. This not only enhances the sense of honor and participation of the user being showcased but also creates a more attractive and engaging viewing experience for other viewers in the virtual space, thereby effectively increasing user engagement and the overall activity level of the live stream.

[0104] For example, the video display method provided in this disclosure uses human skeletal point recognition technology to detect key points of the audience's movements and scores them comprehensively. It combines collected basic information and real-time interaction data of the audience to set video display screening criteria, comprehensively considering audience scores and information to select qualified audience videos. Then, the qualified audience videos are displayed in a dynamic loop, and the display duration and frequency are adjusted based on real-time feedback. Finally, the host provides real-time comments and interactions on the displayed videos and offers certain rewards to enhance the interactivity and fun of the live stream. The specific steps are as follows: Step 1: Acquiring Audience Videos: 1) Sending Participation Invitations: To attract viewers to participate in the live practice session, a dynamic prompting method is implemented through the virtual space interface. Participation invitations are sent to all online viewers in a timely manner, and the interface clearly explains that outstanding participants have the opportunity to broadcast their practice videos in the virtual space and receive corresponding rewards, thus attracting audience attention and encouraging participation. 2) Video Authorization and Recording: After accepting the invitation, viewers are first guided to complete the necessary authorization operations to enable their device's camera and microphone. The authorization process adheres to privacy protection to ensure the security of viewer data. After authorization, the live streaming application will begin recording the viewer's practice video. During recording, the application optimizes video quality through image processing technology to ensure clear capture of the practice movements, providing accurate video data for subsequent motion analysis. 3) Video Upload: The recorded practice video is uploaded to the live streaming platform's server. Utilizing the efficient transmission characteristics of protocols such as RTMP, stable video data upload is ensured. Simultaneously, Content Delivery Network (CDN) distribution and load balancing technologies further improve upload speed and stability.

[0105] Step 2 (1) Extract action key point information from the captured audience video and give a comprehensive score: Use the MoveNet human pose estimation model to extract features from the video, extract human key point information, and infer the human pose based on this key point information. Based on the extracted human key point information, give a comprehensive score from three aspects: spatial similarity, temporal synchronization and action integrity.

[0106] 1) Calculating Similarity Scores: An adaptive keypoint-based human motion similarity algorithm is employed. This algorithm utilizes keypoint spatial coordinate normalization and adaptive dynamic time window sequence alignment technology to mask errors caused by differences in camera viewpoint, human form, and speed. By adaptively adjusting keypoint weights, a similarity score is calculated based on key features of the motion.

[0107] 2) Calculate the synchronicity score: Use the DTW algorithm to align the time series of the audience and the standard movements to find the optimal matching path. Based on the Euclidean distance of the aligned keypoints, evaluate the time synchronicity and normalize it into a percentage score.

[0108] 3) Calculate the motion integrity score: Use keyframe detection technology to identify different stages of the audience's motion. Match the audience's motion stages with standard motions to check the completion of necessary actions. Determine the motion integrity score based on the proportion of completed motion stages.

[0109] A total score is calculated using a weighted average method, taking into account accuracy, synchronicity, and completeness. Different weights can be assigned to each dimension, for example, similarity 50%, synchronicity 30%, and completeness 20%.

[0110] For example, if a viewer's similarity score is 80, synchronicity score is 70, and integrity score is 90, then according to the weighted average method, this viewer's total score is: 80×50%+70×30%+90×0.2=79 points.

[0111] Step 2 (2): Collect basic information and real-time interaction data of the audience: Based on the registration data of the audience on the platform, collect the basic information of the audience, such as age, gender, and geographical location. Use real-time data analysis tools to monitor the interactive behavior of the audience during the live broadcast, including likes, comments, and gift-giving.

[0112] Step 3: Set the selection criteria for practice videos: Combine audience ratings with their basic information, consider the characteristics and interaction of different audience groups, and determine the priority of display.

[0113] 1) Set a minimum scoring threshold based on the preset action scoring criteria. For example, only viewers who score above 80 points in the follow-up action will be displayed on the broadcaster's screen.

[0114] 2) Multi-dimensional Priority: Interactivity Priority: Among highly rated viewers, if the ratings are the same, those viewers who are highly interactive during the live stream will be given priority. For example, users who frequently comment and like.

[0115] New users prioritized: To encourage new user participation, new and existing users will be sorted separately based on ratings. New users are those entering the virtual space for the first time. Videos of top-rated new users will be displayed periodically.

[0116] Age and geographic balance: Users are grouped by age and region, and ranked separately based on their scores. This ensures that users of different ages and regions have the opportunity to be showcased, increasing the breadth of interaction. For example, the Nth place in Nanjing, the Nth place in Shanghai, the Nth place in the senior group or children's group, etc. These rankings will be displayed on the screen during the carousel of viewers practicing.

[0117] When implementing the above priority ranking, different weights are first assigned to interactivity and new users. The weighted values ​​for user interactivity and new user status are then summed to obtain a comprehensive score for each user: Overall Score

[0118] in, , These are the weights for interactivity and new user status, respectively. , It is a standardized score corresponding to interactivity and new user status.

[0119] After calculating the users' overall scores, all users are sorted according to their scores. Then, they are grouped by age and region, and each group is allocated one display slot. During actual display, the system first shows users with high interactivity and high scores for new users, and then allocates the remaining display slots to other users according to the rules of age and region balance.

[0120] Based on the above-set filtering criteria, the audience is sorted and filtered to select the audience practice footage that will be displayed on the broadcaster's screen.

[0121] Step 4: Dynamic Video Carousel Display: Set up a carousel system in the backend service of the live streaming platform. At regular intervals (e.g., 10 seconds), switch and display a group of viewer videos on the user's virtual interface. This allows for showcasing more viewer participation. The display duration and frequency of each video are determined by viewer ratings and interactivity. Viewers with high ratings and high interaction can have their videos displayed for longer periods or more frequently. The specific carousel display mechanism is as follows: Set up a dynamic carousel display system that adjusts the display duration and frequency of each video based on viewer ratings and interactivity. Specifically, set the base display time (the aforementioned initial display duration) to T = 10 seconds. If a viewer's rating is higher than the preset baseline score... This will increase the additional display time. If the audience's interactivity is higher than the preset benchmark interactivity This will increase the additional display time. The mathematical expressions for display duration D and display frequency F are as follows:

[0122]

[0123] Among them, the rating adjustment factor and the interaction adjustment factor are preset parameters of the system, which are used to adjust the display duration based on the audience's rating and interactivity.

[0124] Step 5: Virtual Space Interaction with the Audience: The virtual space provides real-time commentary and interaction on the displayed audience videos, encouraging audience participation and enhancing the interactivity and fun of the live stream. Rewards are offered to the displayed audience, such as virtual gifts or points. Virtual gifts can be used in games, and points can be redeemed for game time, incentivizing more viewers to actively participate and follow along.

[0125] In summary, the solution provided in this public disclosure is as follows: By acquiring action videos from multiple users, performing action recognition on the videos to obtain user action data for each video, and analyzing the matching degree between user action data and standard action data to obtain an action score for each user, the subjectivity of the scoring can be reduced, improving the user experience. Based on the action score and each user's attribute information, the target video is determined from multiple action videos. The target video is then pushed to the client associated with the virtual space for display, showcasing excellent action videos to users in the virtual space and enhancing the interactivity and fun of the live stream.

[0126] To implement the video display method provided in this disclosure, this disclosure also provides a video display device, such as... Figure 6 As shown. Figure 6 This is a schematic diagram of the structure of a video display device provided in an embodiment of the present disclosure. The video display device 600 includes: Acquisition unit 601 is used to acquire action videos of multiple users; The recognition unit 602 is used to perform action recognition on the action video to obtain user action data corresponding to each action video. Analysis unit 603 is used to analyze the degree of matching between user action data and standard action data to obtain an action score for each user; The determination unit 604 is used to determine the target video from multiple action videos based on action scores and attribute information of each user; Display unit 605 is used to push the target video to the client associated with the virtual space for display.

[0127] In one embodiment, the analysis unit 603 is specifically used for: Based on key information of the human skeleton, the similarity between user action data and standard action data in spatial form is analyzed to obtain a spatial similarity score. Align the time series of user action data with standard action data to obtain a time synchronization score for each action video; The user's action data is compared with the action stages in the standard action data to determine the action integrity score for each action video; Each user's action score is determined based on spatial similarity score, synchronization score, and integrity score.

[0128] In one embodiment, the determining unit 604 is specifically used for: Based on action scores, candidate videos that meet the target criteria are selected from multiple action videos. Based on each user's attribute information, users in the candidate videos are sorted by priority from high to low. The target video is determined from the candidate videos based on the priority ranking results.

[0129] In one embodiment, the display unit 605 is specifically used for: Obtain the initial display duration of the target video; Based on action scores, interaction data in attribute information, and initial display duration, determine the target display duration and display frequency; The target video is pushed to the client associated with the virtual space and displayed according to the target display duration and frequency.

[0130] In one embodiment, the acquisition unit 601 is specifically used for: Send follow-up participation requests to multiple users; In response to receiving confirmation requests from multiple users, a recording authorization request is sent to the user who sent the confirmation request. In response to receiving recording authorization confirmation messages from multiple users, the system acquires the action videos of the multiple users who sent the recording authorization confirmation messages.

[0131] In one embodiment, the video display device 600 further includes a push unit, which is used for: Obtain evaluation information for the target video and push the evaluation information to the client associated with the virtual space; Distribute virtual rewards to the user accounts corresponding to the target video.

[0132] It should be noted that the video display device provided in the above embodiments is only illustrated by the division of the above program modules. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the video display device can be divided into different program modules to complete all or part of the processing described above. In addition, the video display device provided in the above embodiments and the video display method embodiments provided in this disclosure belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0133] Figure 7 This is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiments of this disclosure, such as... Figure 7 As shown, the electronic device 700 includes at least one processor 702; and a memory 701 communicatively connected to the at least one processor 702; wherein the memory 701 stores instructions executable by the at least one processor 702, the instructions being executed by the at least one processor 702 to implement the steps of the video display method of the present disclosure embodiment.

[0134] Optionally, the electronic device may specifically be a video display device in the embodiments of this application, and the electronic device may implement the corresponding processes implemented by the video display device in the various methods of the embodiments of this application. For the sake of brevity, it will not be described in detail here.

[0135] It is understood that the electronic device also includes a communication interface 703. Various components in the electronic device are coupled together via a bus system 704. It is understood that the bus system 704 is used to implement communication between these components. In addition to a data bus, the bus system 704 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 7 The general designated all buses as Bus System 704.

[0136] It is understood that memory 701 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 701 described in this embodiment of the invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0137] The methods disclosed in the above embodiments can be applied to or implemented by processor 702. Processor 702 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the hardware of processor 702 or by instructions in software form. Processor 702 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 702 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, specifically memory 701. Processor 702 reads information from memory 701 and, in conjunction with its hardware, completes the steps of the aforementioned methods.

[0138] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components to perform the aforementioned method.

[0139] This disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the steps of the video display method of the present invention.

[0140] Optionally, the computer-readable storage medium can be applied to the video display device in the embodiments of this application, and the computer instructions cause the computer to execute the corresponding processes implemented by the video display device in the various methods of the embodiments of this application. For the sake of brevity, they will not be described in detail here.

[0141] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the video display method provided in this embodiment of the invention.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0143] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0144] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0145] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0146] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0147] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A video display method, characterized in that, include: Acquire action videos from multiple users; Perform motion recognition on the motion videos to obtain user motion data corresponding to each motion video; The user action data is matched with standard action data to obtain an action score for each user. Based on the action score and each user's attribute information, a target video is determined from multiple action videos; The target video is pushed to the client associated with the virtual space for display.

2. The method according to claim 1, characterized in that, The process of analyzing the matching degree between the user action data and standard action data to obtain an action score for each user includes: Based on the key point information of the human skeleton, the similarity between the user action data and the standard action data in terms of spatial shape is analyzed to obtain a spatial similarity score. The time series of the user action data and the standard action data are aligned to obtain the time synchronization score for each action video. The user action data is compared with the action stages in the standard action data to determine the action integrity score corresponding to each action video; Based on the spatial similarity score, the synchronization score, and the integrity score, an action score for each user is determined.

3. The method according to claim 1, characterized in that, The step of determining the target video from multiple action videos based on the action score and each user's attribute information includes: Based on the action score, candidate videos that meet the target conditions are selected from multiple action videos; Based on each user's attribute information, the users in the candidate videos are sorted from high to low priority; The target video is determined from the candidate videos based on the priority ranking results.

4. The method according to claim 1, characterized in that, The step of pushing the target video to a client associated with the virtual space for display includes: Obtain the initial display duration of the target video; Based on the action score, the interaction data in the attribute information, and the initial display duration, the target display duration and display frequency are determined. The target video is pushed to the client associated with the virtual space and displayed according to the target display duration and the display frequency.

5. The method according to claim 1, characterized in that, The acquisition of action videos from multiple users includes: Send follow-up participation requests to multiple users; In response to receiving confirmation requests from multiple users, a recording authorization request is sent to the user who sent the confirmation request. In response to receiving recording authorization confirmation information from the multiple users, the system acquires the action videos of the multiple users who sent the recording authorization confirmation information.

6. The method according to claim 1, characterized in that, After pushing the target video to the client associated with the virtual space for display, the method includes: Obtain evaluation information for the target video and push the evaluation information to the client associated with the virtual space; Virtual rewards are distributed to the user accounts corresponding to the target video.

7. A video display device, characterized in that, include: The acquisition unit is used to acquire action videos from multiple users; The recognition unit is used to perform action recognition on the action video to obtain user action data corresponding to each action video; The analysis unit is used to analyze the degree of matching between the user action data and the standard action data to obtain an action score for each user. A determining unit is configured to determine a target video from a plurality of action videos based on the action score and the attribute information of each user; The display unit is used to push the target video to the client associated with the virtual space for display.

8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.