Live Follow - up Video Detection Method, Device, Equipment and Storage Medium

By calculating the playback frequency and detecting key points, generating and practicing ratings, the problem that users and practicing videos cannot match live videos is solved, and the evaluation accuracy is improved.

CN115588231BActive Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211093077.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2025-07-18
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

In the sports live broadcast platform, the user's follow-up video cannot match the live sports video, resulting in the inability to accurately evaluate the user's sports condition.

Method used

By calculating the playback frequency of live sports videos and training videos, extracting standard video sequences and training action sequences, detecting key points and action angles, and generating training ratings.

Benefits of technology

Improve the accuracy of the rating of the rating, ensure the matching of the standard video sequence and the rating action sequence, and avoid the problem that users cannot keep up with the coach's movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115588231B_ABST
    Figure CN115588231B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence, and provides a live follow-along video detection method, device, equipment and storage medium. The method obtains a live sports video and a follow-along video, calculates a playback frequency according to the first video duration of the live sports video and the second video duration of the follow-along video, extracts a standard video sequence based on the playback frequency, and extracts a follow-along action sequence. It detects the standard key points of each standard video frame and the user key points of each follow-along frame, identifies the standard action angles of each standard video frame based on the standard key points, and identifies the follow-along action angles of each follow-along frame based on the user key points. According to the standard action angles and the follow-along action angles, a follow-along rating can be accurately generated. In addition, the present invention also relates to blockchain technology, and the follow-along rating can be stored in the blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a live follow-up video detection method, device, equipment and storage medium. Background Art

[0002] Currently, in a sports live broadcast platform, since the playback frequency of the live sports video cannot be controlled, when the training actions of the user cannot keep up with the actions of the coach in the live sports video, there is a problem that the follow-up video of the user cannot be successfully matched with the live sports video, resulting in the inability to accurately evaluate the sports situation of the user in the follow-up video. Summary of the Invention

[0003] In view of the above, it is necessary to provide a live follow-up video detection method, device, equipment and storage medium, which can solve the technical problem of being unable to accurately evaluate the sports situation of the user in the follow-up video.

[0004] On the one hand, the present invention proposes a live follow-up video detection method, and the live follow-up video detection method includes:

[0005] Obtain a live sports video, and obtain a follow-up video based on the live sports video;

[0006] Calculate the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-up video;

[0007] Extract a standard video sequence from the live sports video based on the playback frequency, and extract a follow-up action sequence from the follow-up video;

[0008] Detect the standard key points of each standard video frame in the standard video sequence, and detect the user key points of each follow-up frame in the follow-up action sequence;

[0009] Identify the standard action angle of each standard video frame based on the standard key points, and identify the follow-up action angle of each follow-up frame based on the user key points;

[0010] Generate a follow-up rating of the follow-up video according to the standard action angle and the follow-up action angle.

[0011] According to a preferred embodiment of the present invention, the calculating the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-up video includes:

[0012] Identify the start generation time point of the start video frame in the live sports video, and identify the end generation time point of the end video frame in the live sports video;

[0013] Calculate the difference between the end generation time point and the start generation time point to obtain the first video duration;

[0014] Calculate the ratio of the second video duration to the first video duration to obtain the playback frequency.

[0015] According to a preferred embodiment of the present invention, the extracting of the standard video sequence from the live sports video based on the playback frequency includes:

[0016] Compare the playback frequency with a preset frequency;

[0017] If the playback frequency is greater than the preset frequency, convert the live sports video based on the playback frequency to obtain a standard sports video, and generate the standard video sequence according to the standard video frames in the standard sports video and the frame positions of the standard video frames in the standard sports video; or

[0018] If the playback frequency is less than or equal to the preset frequency, extract the standard video frames from the live sports video based on the playback frequency as the standard video sequence.

[0019] According to a preferred embodiment of the present invention, the detecting of the standard key points of each standard video frame in the standard video sequence includes:

[0020] For each standard video frame, obtain the pixel values and pixel positions of each pixel point in the standard video frame;

[0021] Detect the pixel values and the pixel positions based on a pre-trained human detection model to obtain the detection category corresponding to each pixel point and the detection probability that the pixel point belongs to the detection category;

[0022] Screen target pixel points from multiple pixel points based on the detection probability;

[0023] Based on the pixel positions and the detection categories, determine the region formed by adjacent pixel points corresponding to the same detection category among the target pixel points as the standard key points.

[0024] According to a preferred embodiment of the present invention, the identifying of the standard action angles of each standard video frame based on the standard key points includes:

[0025] Taking any standard key point as the origin and taking the image side parallel to any standard video frame as the coordinate axis to construct a plane rectangular coordinate system;

[0026] According to the pixel positions corresponding to the standard key points and the plane rectangular coordinate system, identify the key point coordinate values corresponding to the standard key points;

[0027] Obtain the connection key point pairs of any one of the standard key points from the multiple standard key points, where the connection key point pairs include a first connection point and a second connection point;

[0028] Determine the connection edge formed by the any standard key point and the first connection point as the first connection edge, and calculate the angle between the first connection edge and the coordinate axis based on the key point coordinate value of the first connection point to obtain a first angle;

[0029] Determine the connection edge formed by the any standard key point and the second connection point as the second connection edge, and calculate the angle between the second connection edge and the coordinate axis based on the key point coordinate value of the second connection point to obtain a second angle;

[0030] Generate the standard action angle based on the first angle, the second angle, and a preset integer value.

[0031] According to a preferred embodiment of the present invention, the calculation formula of the standard action angle is:

[0032] θ = β - α ± (2kπ);

[0033] Where θ represents the standard action angle, β represents the second angle, α represents the first angle, k is the preset integer value, and θ > 0.

[0034] According to a preferred embodiment of the present invention, the follow - up rating for generating the follow - up video based on the standard action angle and the follow - up action angle includes:

[0035] Obtain the key point weight threshold corresponding to the any standard key point;

[0036] Generate the overlap degree of each follow - up frame according to the standard action angle, the follow - up action angle, and the corresponding key point weight threshold, and the calculation formula of the overlap degree is:

[0037] Where S represents the overlap degree, W = [w1, w2, …, w n represents the key point weight thresholds corresponding to the multiple standard key points, M = [m1, m2, …, m n represents the multiple standard action angles, and T = [t1, t2, …, t n represents the multiple follow - up action angles;

[0038] Generate the follow - up rating according to the overlap degree and a preset mapping relationship.

[0039] On the other hand, the present invention also proposes a live follow - up video detection device, and the live follow - up video detection device includes:

[0040] An acquisition unit, configured to acquire a live sports video and acquire a follow - along video based on the live sports video;

[0041] A calculation unit, configured to calculate the playing frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow - along video;

[0042] An extraction unit, configured to extract a standard video sequence from the live sports video based on the playing frequency, and extract a follow - along action sequence from the follow - along video;

[0043] A detection unit, configured to detect standard key points of each standard video frame in the standard video sequence, and detect user key points of each follow - along frame in the follow - along action sequence;

[0044] An identification unit, configured to identify the standard action angle of each standard video frame based on the standard key points, and identify the follow - along action angle of each follow - along frame based on the user key points;

[0045] A generation unit, configured to generate a follow - along rating of the follow - along video according to the standard action angle and the follow - along action angle.

[0046] On the other hand, the present invention also provides an electronic device, which includes:

[0047] A memory, storing computer - readable instructions; and

[0048] A processor, configured to execute the computer - readable instructions stored in the memory to implement the live follow - along video detection method.

[0049] On the other hand, the present invention also provides a computer - readable storage medium, in which computer - readable instructions are stored, and the computer - readable instructions are executed by a processor in an electronic device to implement the live follow - along video detection method.

[0050] It can be seen from the above technical solutions that the present application can accurately quantify the playing frequency through the first video duration and the second video duration, and then based on the playing frequency, it can ensure the matching relationship between the standard video sequence and the follow - along action sequence, thereby avoiding the matching problem caused when the user's follow - along actions cannot keep up with the coach's actions in the live sports video, and improving the accuracy of the follow - along rating. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a flowchart of a preferred embodiment of the live follow - along video detection method of the present invention.

[0052] Figure 2It is a visual diagram of the standard action angle in the present invention.

[0053] Figure 3 It is a functional module diagram of a preferred embodiment of the live follow - up video detection device of the present invention.

[0054] Figure 4 It is a schematic structural diagram of an electronic device of a preferred embodiment of the method for implementing the live follow - up video detection of the present invention. Detailed implementation manners

[0055] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0056] As Figure 1 shown, it is a flowchart of a preferred embodiment of the method for live follow - up video detection of the present invention. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.

[0057] The live follow - up video detection method can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theory, method, technology and application system.

[0058] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0059] The live follow - up video detection method is applied to one or more electronic devices. The electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre - set or stored computer - readable instructions. Its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field - programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0060] The electronic device can be any kind of electronic product that can perform human-computer interaction with users. For example, personal computers, tablet computers, smart phones, personal digital assistants (PDAs), game consoles, Internet Protocol Television (IPTV), smart wearable devices, etc.

[0061] The electronic device may include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network electronic device, a group of electronic devices composed of multiple network electronic devices, or a cloud composed of a large number of hosts or network electronic devices based on cloud computing.

[0062] The network where the electronic device is located includes, but is not limited to: the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.

[0063] 101. Obtain a live sports video and obtain a follow-up video based on the live sports video.

[0064] In at least one embodiment of the present invention, the live sports video is usually a guidance video filmed by a sports coach for a certain sport. The live sports video refers to a video that is played in real time.

[0065] The follow-up video refers to a training video generated after the training user imitates the actions based on the live sports video.

[0066] 102. Calculate the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-up video.

[0067] In at least one embodiment of the present invention, the first video duration refers to the total playback duration of the live sports video, and the second video duration refers to the total playback duration of the follow-up video. The playback frequency refers to the ratio of the second video duration to the first video duration.

[0068] In at least one embodiment of the present invention, the electronic device calculates the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-up video, including:

[0069] Identify the start generation time point of the start video frame in the live sports video, and identify the end generation time point of the end video frame in the live sports video;

[0070] Calculate the difference between the end generation time point and the start generation time point to obtain the first video duration;

[0071] Calculate the ratio of the second video duration to the first video duration to obtain the playback frequency.

[0072] Among them, the start video frame refers to the first video frame in the live sports video, and the last video frame in the end video frame. The start generation time point refers to the time point when the start video frame is captured, and the end generation time point refers to the time point when the end video frame is captured.

[0073] Through the end generation time point and the start generation time point, the first video duration can be accurately counted, so that by combining the second video duration and the first video duration, the accuracy of the playback frequency can be improved.

[0074] Specifically, the generation method of the second video duration is similar to that of the first video duration, and this application will not elaborate on this.

[0075] In at least one embodiment of the present invention, the method further includes:

[0076] Obtain the required frequency;

[0077] Based on the required frequency, perform conversion processing on the live sports video to obtain a target sports video;

[0078] Play the target sports video.

[0079] Among them, the required frequency is used to indicate the requirement of the training user for the playback speed of the live sports video.

[0080] Converting the live sports video through the required frequency can not only avoid the problem that users cannot learn the action details from the live sports video, but also avoid the problem that users cannot keep up with the coach's actions.

[0081] 103. Extract a standard video sequence from the live sports video based on the playback frequency, and extract a follow-up action sequence from the follow-up video.

[0082] In at least one embodiment of the present invention, the standard video sequence can be part or all of the video frames in the live sports video, the follow-up action sequence refers to the video frames corresponding to the standard video sequence, and the number of video frames in the standard video sequence is equal to the number of video frames in the follow-up action sequence.

[0083] In at least one embodiment of the present invention, the electronic device extracts a standard video sequence from the live sports video based on the playback frequency, which includes:

[0084] Compare the playback frequency with a preset frequency;

[0085] If the playback frequency is greater than the preset frequency, convert the live sports video based on the playback frequency to obtain a standard sports video, and generate the standard video sequence according to the standard video frames in the standard sports video and the frame positions of these standard video frames in the standard sports video; or

[0086] If the playback frequency is less than or equal to the preset frequency, extract the standard video frames from the live sports video as the standard video sequence.

[0087] Wherein, the preset frequency is usually set to 1.

[0088] Through the relationship between the playback frequency and the preset frequency, different methods can be adopted to generate the standard video sequence, improving the generation accuracy of the standard video sequence.

[0089] In at least one embodiment of the present invention, the electronic device generates the follow-up action sequence according to the follow-up video frames in the follow-up video and the frame positions of these follow-up video frames in the follow-up video.

[0090] 104. Detect the standard key points of each standard video frame in the standard video sequence, and detect the user key points of each follow-up frame in the follow-up action sequence.

[0091] In at least one embodiment of the present invention, the standard key points refer to the human body posture key points of the sports coach, and the user key points refer to the human body posture key points of the training user. The standard key points and the user key points respectively include human body posture key points such as nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip joint, right hip joint, left knee, right knee, left ankle, and right ankle.

[0092] In at least one embodiment of the present invention, the electronic device detecting the standard key points of each standard video frame in the standard video sequence includes:

[0093] For each standard video frame, obtain the pixel values and pixel positions of each pixel point in this standard video frame;

[0094] Detect the pixel values and pixel positions based on a pre-trained human body detection model to obtain the detection category corresponding to each pixel point and the detection probability that the pixel point belongs to the detection category;

[0095] Filter target pixel points from multiple pixel points based on the detection probability;

[0096] Based on the pixel positions and the detection categories, determine the area formed by adjacent pixel points corresponding to the same detection category among the target pixel points as the standard key points.

[0097] Among them, the human body detection model is a model generated by training and testing with crawled pictures related to people. The detection categories may include, but are not limited to: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip joint, right hip joint, left knee, right knee, left ankle, right ankle and other categories.

[0098] The target pixel points refer to the pixel points whose detection probability is greater than a preset probability, and the preset probability is set according to actual needs.

[0099] The adjacent pixel points refer to the target pixel points with adjacent pixel positions.

[0100] By combining the pixel values and the pixel positions, the detection category and the corresponding detection probability can be accurately identified. Based on the detection probability, interference pixel points can be extracted, thereby improving the screening accuracy of the target pixel points. Further, by analyzing the adjacent pixel points of the target pixel points, the recognition accuracy of the standard key points can be improved.

[0101] Specifically, the generation method of the user key points is similar to the generation method of the standard key points, and this application will not elaborate on this.

[0102] 105. Based on the standard key points, identify the standard action angles of each standard video frame, and based on the user key points, identify the follow-up action angles of each follow-up frame.

[0103] In at least one embodiment of the present invention, the standard action angle refers to the angle formed by a first connecting side and a second connecting side, where the first connecting side refers to the connecting side formed by any standard key point and a first connection point, and the second connecting side refers to the connecting side formed by the any standard key point and a second connection point. The first connection point and the second connection point respectively refer to the remaining standard key points connected to the any standard key point.

[0104] In at least one embodiment of the present invention, the standard action angle of each standard video frame recognized by the electronic device based on the standard key points includes:

[0105] Taking any standard key point as the origin and taking the image side parallel to any standard video frame as the coordinate axis to construct a plane rectangular coordinate system;

[0106] According to the pixel position corresponding to the standard key point and the plane rectangular coordinate system, the key point coordinate value corresponding to the standard key point is recognized;

[0107] Obtaining the connection key point pair of the any standard key point from multiple standard key points, the connection key point pair includes a first connection point and a second connection point;

[0108] Determining the connection edge formed by the any standard key point and the first connection point as the first connection edge, and calculating the included angle between the first connection edge and the coordinate axis based on the key point coordinate value of the first connection point to obtain a first angle;

[0109] Determining the connection edge formed by the any standard key point and the second connection point as the second connection edge, and calculating the included angle between the second connection edge and the coordinate axis based on the key point coordinate value of the second connection point to obtain a second angle;

[0110] Generating the standard action angle based on the first angle, the second angle and a preset integer value.

[0111] Wherein, the key point coordinate value refers to the coordinate value of the standard key point in the plane rectangular coordinate system. For example, if the horizontal pixel position of the standard key point A is the 8th pixel point and the vertical pixel position is the 5th pixel point, then the key point coordinate value is (8, 5).

[0112] The preset integer value is used to control the value of the standard action angle. When the first angle is greater than the second angle, the preset integer value is greater than 0. When the first angle is less than or equal to the second angle, the preset integer value is 0. For example, if α = 90°, β = 315°, then θ = β - α = 225°. If α = 90°, β = 45°, then θ = β - α + 2kπ = 315°.

[0113] Through the key point coordinate value, the first angle and the second angle can be accurately calculated. Furthermore, through the control of the standard action angle by the preset integer value, the accuracy of the standard action angle can be improved.

[0114] As Figure 2 shown, Figure 2 is the visual diagram of the standard action angle in the present invention. Among them,Figure 2 In this, α represents the first angle, β represents the second angle, and θ represents the standard action angle.

[0115] In this embodiment, by controlling the first angle and the second angle to have the same rotation direction, therefore, the accuracy of the standard action angle can be improved.

[0116] Specifically, the calculation formula for the first angle is:

[0117] α = arccos(c);

[0118] α = arcsin(s);

[0119] Wherein, α represents the first angle, c represents the horizontal coordinate value in the key point coordinate values of the first connection point, and s represents the vertical coordinate value in the key point coordinate values of the first connection point.

[0120] By combining the horizontal coordinate value and the vertical coordinate value, the determination accuracy of the first angle can be improved.

[0121] Specifically, the calculation method of the second angle is similar to that of the first angle, and this application will not elaborate on it here.

[0122] Specifically, the calculation formula for the standard action angle is:

[0123] θ = β - α ± (2kπ);

[0124] Wherein, θ represents the standard action angle, β represents the second angle, α represents the first angle, k is the preset integer value, and θ > 0.

[0125] By adjusting the standard action angle through the preset integer value, the accuracy of the standard action angle can be improved.

[0126] In at least one embodiment of the present invention, the recognition method of the follow-up action angle is similar to that of the standard action angle, and this application will not elaborate on it here.

[0127] 106. Generate a follow-up rating for the follow-up video according to the standard action angle and the follow-up action angle.

[0128] It should be emphasized that to further ensure the privacy and security of the above-mentioned follow-up rating, the above-mentioned follow-up rating can also be stored in a node of a blockchain.

[0129] In at least one embodiment of the present invention, the follow-up rating is used to quantify the exercise training situation of the training user. The follow-up rating includes, but is not limited to: Fail rating, Pass rating, Great rating, Perfect rating, etc.

[0130] In at least one embodiment of the present invention, the follow-up rating of the follow-up video generated by the electronic device according to the standard action angle and the follow-up action angle includes:

[0131] Obtain the key point weight threshold corresponding to any one of the standard key points;

[0132] Generate the coincidence degree of each follow-up frame according to the standard action angle, the follow-up action angle and the corresponding key point weight threshold, and the calculation formula of the coincidence degree is:

[0133] where S represents the coincidence degree, W = [w1, w2,..., w n represents the key point weight thresholds corresponding to multiple standard key points, M = [m1, m2,..., m n represents multiple standard action angles, and T = [t1, t2,..., t n represents multiple follow-up action angles;

[0134] Generate the follow-up rating according to the coincidence degree and the preset mapping relationship.

[0135] Among them, the key point weight threshold and the preset mapping relationship can be set according to actual needs. For example, if the coincidence degree is 0%-40%, it is a Fail rating; if the coincidence degree is 41%-60%, it is a Pass rating; if the coincidence degree is 61%-80%, it is a Great rating; if the coincidence degree is 81%-100%, it is a Perfect rating.

[0136] By setting the key point weight threshold for each standard key point, the quantization accuracy of the coincidence degree can be improved, thereby improving the accuracy of the follow-up rating.

[0137] Specifically, the electronic device generating the follow-up rating according to the coincidence degree and the preset mapping relationship includes:

[0138] Calculate the average value of multiple coincidence degrees;

[0139] Based on the average value and the preset mapping relationship, identify the follow-up rating.

[0140] As can be seen from the above technical solutions, the present application can accurately quantify the playback frequency through the first video duration and the second video duration, and then based on the playback frequency, ensure the matching relationship between the standard video sequence and the follow-up action sequence, thereby avoiding the matching problem caused when the user's follow-up actions cannot keep up with the coach's actions in the live sports video, and improving the accuracy of the follow-up rating.

[0141] As Figure 3 shown, it is a functional module diagram of a preferred embodiment of the live follow-up video detection device of the present invention. The live follow-up video detection device 11 includes an acquisition unit 110, a calculation unit 111, an extraction unit 112, a detection unit 113, an identification unit 114, a generation unit 115, a conversion unit 116, and a playback unit 117. The module / unit referred to in the present invention means a series of computer-readable instruction segments that can be acquired by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0142] The acquisition unit 110 acquires a live sports video and acquires a follow-up video based on the live sports video.

[0143] In at least one embodiment of the present invention, the live sports video is usually a guidance video shot by a sports coach for a certain sport. The live sports video refers to a video that is played in real time.

[0144] The follow-up video refers to a training video generated by a training user after imitating actions based on the live sports video.

[0145] The calculation unit 111 calculates the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-up video.

[0146] In at least one embodiment of the present invention, the first video duration refers to the total playback duration of the live sports video, and the second video duration refers to the total playback duration of the follow-up video. The playback frequency refers to the ratio of the second video duration to the first video duration.

[0147] In at least one embodiment of the present invention, the calculation unit 111 calculates the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-up video, including:

[0148] Identifying the start generation time point of the start video frame in the live sports video and identifying the end generation time point of the end video frame in the live sports video;

[0149] Calculate the difference between the end generation time point and the start generation time point to obtain the first video duration.

[0150] Calculate the ratio of the second video duration to the first video duration to obtain the playback frequency.

[0151] Among them, the start video frame refers to the first video frame in the live sports video, and the last video frame in the end video frame. The start generation time point refers to the time point when the start video frame is captured, and the end generation time point refers to the time point when the end video frame is captured.

[0152] Through the end generation time point and the start generation time point, the first video duration can be accurately counted, so that by combining the second video duration and the first video duration, the accuracy of the playback frequency can be improved.

[0153] Specifically, the generation method of the second video duration is similar to that of the first video duration, and details are not described herein again in this application.

[0154] In at least one embodiment of the present invention, the acquisition unit 110 acquires a required frequency.

[0155] The conversion unit 116 performs a conversion process on the live sports video based on the required frequency to obtain a target sports video.

[0156] The playback unit 117 plays the target sports video.

[0157] Among them, the required frequency is used to indicate the requirement of the training user for the playback speed of the live sports video.

[0158] Converting the live sports video through the required frequency can not only avoid the problem that users cannot learn action details from the live sports video, but also avoid the problem that users cannot keep up with the coach's actions.

[0159] The extraction unit 112 extracts a standard video sequence from the live sports video based on the playback frequency, and extracts a follow-up action sequence from the follow-up video.

[0160] In at least one embodiment of the present invention, the standard video sequence may be some video frames or all video frames in the live sports video. The follow-up action sequence refers to the video frames corresponding to the standard video sequence, and the number of video frames in the standard video sequence is equal to the number of video frames in the follow-up action sequence.

[0161] In at least one embodiment of the present invention, the extracting unit 112 extracts a standard video sequence from the live sports video based on the playing frequency, which includes:

[0162] Compare the playing frequency with a preset frequency;

[0163] If the playing frequency is greater than the preset frequency, convert the live sports video based on the playing frequency to obtain a standard sports video, and generate the standard video sequence according to the standard video frames in the standard sports video and the frame positions of these standard video frames in the standard sports video; or

[0164] If the playing frequency is less than or equal to the preset frequency, extract the standard video frames from the live sports video as the standard video sequence.

[0165] Wherein, the preset frequency is usually set to 1.

[0166] Based on the relationship between the playing frequency and the preset frequency, different methods can be adopted to generate the standard video sequence, improving the generation accuracy of the standard video sequence.

[0167] In at least one embodiment of the present invention, the extracting unit 112 generates the follow-up action sequence according to the follow-up video frames in the follow-up video and the frame positions of these follow-up video frames in the follow-up video.

[0168] The detecting unit 113 detects the standard key points of each standard video frame in the standard video sequence, and detects the user key points of each follow-up frame in the follow-up action sequence.

[0169] In at least one embodiment of the present invention, the standard key points refer to the human body posture key points of the sports coach, and the user key points refer to the human body posture key points of the training user. The standard key points and the user key points respectively include human body posture key points such as nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip joint, right hip joint, left knee, right knee, left ankle, and right ankle.

[0170] In at least one embodiment of the present invention, the detecting unit 113 detecting the standard key points of each standard video frame in the standard video sequence includes:

[0171] For each standard video frame, obtain the pixel values and pixel positions of each pixel point in this standard video frame;

[0172] Detect the pixel values and the pixel positions based on a pre-trained human body detection model to obtain the detection category corresponding to each pixel point and the detection probability that the pixel point belongs to the detection category;

[0173] Screen target pixel points from the multiple pixel points based on the detection probability;

[0174] Based on the pixel positions and the detection categories, determine the region formed by adjacent pixel points corresponding to the same detection category among the target pixel points as the standard key points.

[0175] Wherein, the human body detection model is a model generated by training and testing using the crawled pictures related to people. The detection categories may include, but are not limited to: nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip joint, right hip joint, left knee, right knee, left ankle, right ankle and other categories.

[0176] The target pixel points refer to the pixel points whose detection probability is greater than a preset probability, and the preset probability is set according to actual requirements.

[0177] The adjacent pixel points refer to the target pixel points with adjacent pixel positions.

[0178] By combining the pixel values and the pixel positions, the detection category and the corresponding detection probability can be accurately identified. Based on the detection probability, interference pixel points can be extracted, thereby improving the screening accuracy of the target pixel points. Further, by analyzing the adjacent pixel points of the target pixel points, the recognition accuracy of the standard key points can be improved.

[0179] Specifically, the generation method of the user key points is similar to the generation method of the standard key points, and this application will not elaborate on this.

[0180] The recognition unit 114 recognizes the standard action angle of each standard video frame based on the standard key points, and recognizes the follow-up action angle of each follow-up frame based on the user key points.

[0181] In at least one embodiment of the present invention, the standard action angle refers to the angle formed by a first connecting side and a second connecting side, wherein the first connecting side refers to the connecting side formed by any standard key point and a first connection point, and the second connecting side refers to the connecting side formed by the any standard key point and a second connection point. The first connection point and the second connection point respectively refer to the remaining standard key points connected to the any standard key point.

[0182] In at least one embodiment of the present invention, the standard action angles of each standard video frame recognized by the recognition unit 114 include:

[0183] Taking any standard key point as the origin and taking the image side parallel to any standard video frame as the coordinate axis to construct a plane rectangular coordinate system;

[0184] According to the pixel position corresponding to the standard key point and the plane rectangular coordinate system, the key point coordinate value corresponding to the standard key point is recognized;

[0185] Obtaining the connection key point pair of the any standard key point from multiple standard key points, the connection key point pair includes a first connection point and a second connection point;

[0186] Determining the connection edge formed by the any standard key point and the first connection point as the first connection edge, and calculating the included angle between the first connection edge and the coordinate axis based on the key point coordinate value of the first connection point to obtain a first angle;

[0187] Determining the connection edge formed by the any standard key point and the second connection point as the second connection edge, and calculating the included angle between the second connection edge and the coordinate axis based on the key point coordinate value of the second connection point to obtain a second angle;

[0188] Generating the standard action angle based on the first angle, the second angle and a preset integer value.

[0189] Wherein, the key point coordinate value refers to the coordinate value of the standard key point in the plane rectangular coordinate system. For example, if the horizontal pixel position of the standard key point A is the 8th pixel point and the vertical pixel position is the 5th pixel point, then the key point coordinate value is (8, 5).

[0190] The preset integer value is used to control the value of the standard action angle. When the first angle is greater than the second angle, the preset integer value is greater than 0. When the first angle is less than or equal to the second angle, the preset integer value is 0. For example, if α = 90°, β = 315°, then θ = β - α = 225°. If α = 90°, β = 45°, then θ = β - α + 2kπ = 315°.

[0191] Through the key point coordinate value, the first angle and the second angle can be accurately calculated. Furthermore, through the control of the standard action angle by the preset integer value, the accuracy of the standard action angle can be improved.

[0192] As Figure 2 shown, Figure 2 is the visual diagram of the standard action angle in the present invention. Among them,Figure 2 In this, α represents the first angle, β represents the second angle, and θ represents the standard action angle.

[0193] In this embodiment, by controlling the first angle and the second angle to have the same rotation direction, therefore, the accuracy of the standard action angle can be improved.

[0194] Specifically, the calculation formula for the first angle is:

[0195] α = arccos(c);

[0196] α = arcsin(s);

[0197] Wherein, α represents the first angle, c represents the horizontal coordinate value in the key point coordinate values of the first connection point, and s represents the vertical coordinate value in the key point coordinate values of the first connection point.

[0198] By combining the horizontal coordinate value and the vertical coordinate value, the determination accuracy of the first angle can be improved.

[0199] Specifically, the calculation method of the second angle is similar to that of the first angle, and this application will not elaborate on it.

[0200] Specifically, the calculation formula for the standard action angle is:

[0201] θ = β - α ± (2kπ);

[0202] Wherein, θ represents the standard action angle, β represents the second angle, α represents the first angle, k is the preset integer value, and θ > 0.

[0203] By adjusting the standard action angle through the preset integer value, the accuracy of the standard action angle can be improved.

[0204] In at least one embodiment of the present invention, the recognition method of the follow-up action angle is similar to the recognition method of the standard action angle, and this application will not elaborate on it.

[0205] The generating unit 115 generates the follow-up rating of the follow-up video according to the standard action angle and the follow-up action angle.

[0206] It should be emphasized that to further ensure the privacy and security of the above follow-up rating, the above follow-up rating can also be stored in a node of a blockchain.

[0207] In at least one embodiment of the present invention, the follow-up rating is used to quantify the exercise training situation of the training user. The follow-up rating includes, but is not limited to: Fail rating, Pass rating, Great rating, Perfect rating, etc.

[0208] In at least one embodiment of the present invention, the generating unit 115 generates the follow-up rating of the follow-up video according to the standard action angle and the follow-up action angle, including:

[0209] Obtain the key point weight threshold corresponding to any one of the standard key points;

[0210] Generate the coincidence degree of each follow-up frame according to the standard action angle, the follow-up action angle and the corresponding key point weight threshold, and the calculation formula of the coincidence degree is:

[0211] where S represents the coincidence degree, W = [w1, w2,..., w n represents the key point weight thresholds corresponding to multiple standard key points, M = [m1, m2,..., m n represents multiple standard action angles, and T = [t1, t2,..., t n represents multiple follow-up action angles;

[0212] Generate the follow-up rating according to the coincidence degree and the preset mapping relationship.

[0213] Among them, the key point weight threshold and the preset mapping relationship can be set according to actual needs. For example, if the coincidence degree is 0%-40%, it is a Fail rating; if the coincidence degree is 41%-60%, it is a Pass rating; if the coincidence degree is 61%-80%, it is a Great rating; if the coincidence degree is 81%-100%, it is a Perfect rating.

[0214] By setting the key point weight threshold for each standard key point, the quantization accuracy of the coincidence degree can be improved, thereby improving the accuracy of the follow-up rating.

[0215] Specifically, the generating unit 115 generates the follow-up rating according to the coincidence degree and the preset mapping relationship, including:

[0216] Calculate the average value of multiple coincidence degrees;

[0217] Based on the average value and the preset mapping relationship, identify the follow-up rating.

[0218] As can be seen from the above technical solutions, the present application can accurately quantify the playback frequency through the first video duration and the second video duration, and then, based on the playback frequency, ensure the matching relationship between the standard video sequence and the follow-up action sequence, thereby avoiding the matching problem caused when the user's follow-up actions cannot keep up with the coach's actions in the live sports video, and improving the accuracy of the follow-up rating.

[0219] As Figure 4 shown, it is a schematic structural diagram of an electronic device according to a preferred embodiment of the method for detecting a live follow-up video of the present invention.

[0220] In an embodiment of the present invention, the electronic device 1 includes, but is not limited to, a memory 12, a processor 13, and computer-readable instructions stored in the memory 12 and executable on the processor 13, such as a live follow-up video detection program.

[0221] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1, and does not constitute a limitation on the electronic device 1. It may include more or fewer components than shown, or combine some components, or different components. For example, the electronic device 1 may further include input / output devices, network access devices, buses, etc.

[0222] The processor 13 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 13 is the operation core and control center of the electronic device 1, connecting various parts of the entire electronic device 1 through various interfaces and lines, and executing the operating system of the electronic device 1 and various installed application programs, program codes, etc.

[0223] Exemplarily, the computer-readable instructions may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these computer-readable instruction segments are used to describe the execution process of the computer-readable instructions in the electronic device 1. For example, the computer-readable instructions may be divided into an acquisition unit 110, a calculation unit 111, an extraction unit 112, a detection unit 113, an identification unit 114, a generation unit 115, a conversion unit 116, and a playback unit 117.

[0224] The memory 12 may be used to store the computer-readable instructions and / or modules. By running or executing the computer-readable instructions and / or modules stored in the memory 12, and by invoking the data stored in the memory 12, the processor 13 realizes various functions of the electronic device 1. The memory 12 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the electronic device. The memory 12 may include non-volatile and volatile memories, such as: hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other storage devices.

[0225] The memory 12 may be an external memory and / or an internal memory of the electronic device 1. Further, the memory 12 may be a memory in a physical form, such as a memory stick, a TF card (Trans-flash Card), etc.

[0226] If the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it may also be completed by computer-readable instructions to instruct relevant hardware. The computer-readable instructions may be stored in a computer-readable storage medium, and when the computer-readable instructions are executed by the processor, the steps of the above method embodiments can be realized.

[0227] Among them, the computer-readable instructions include computer-readable instruction codes, which may be in the form of source codes, object codes, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer-readable instruction codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), and random access memories (RAMs).

[0228] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. A blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information on a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. A blockchain may include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0229] Combined Figure 1 , the memory 12 in the electronic device 1 stores computer-readable instructions to implement a live follow-along video detection method, and the processor 13 can execute the computer-readable instructions to implement:

[0230] Obtain a live sports video and obtain a follow-along video based on the live sports video;

[0231] Calculate the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-along video;

[0232] Extract a standard video sequence from the live sports video based on the playback frequency, and extract a follow-along action sequence from the follow-along video;

[0233] Detect the standard key points of each standard video frame in the standard video sequence, and detect the user key points of each follow-along frame in the follow-along action sequence;

[0234] Identify the standard action angles of each standard video frame based on the standard key points, and identify the follow-along action angles of each follow-along frame based on the user key points;

[0235] Generate a follow-along rating for the follow-along video according to the standard action angles and the follow-along action angles.

[0236] Specifically, for the specific implementation method of the above computer-readable instructions by the processor 13, reference may be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here.

[0237] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0238] Computer-readable instructions are stored on the computer-readable storage medium, wherein when the computer-readable instructions are executed by the processor 13, the following steps are implemented:

[0239] Obtain a live sports video and obtain a follow-along video based on the live sports video;

[0240] Calculate the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-along video;

[0241] Extract a standard video sequence from the live sports video based on the playback frequency, and extract a follow-along action sequence from the follow-along video;

[0242] Detect the standard key points of each standard video frame in the standard video sequence, and detect the user key points of each follow-along frame in the follow-along action sequence;

[0243] Identify the standard action angles of each standard video frame based on the standard key points, and identify the follow-along action angles of each follow-along frame based on the user key points;

[0244] Generate a follow-along rating for the follow-along video according to the standard action angles and the follow-along action angles.

[0245] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0246] In addition, the functional modules in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0247] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

[0248] In addition, it is obvious that the term "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices described may also be implemented by one element or device through software or hardware. The terms such as "first" and "second" are used to denote names and do not denote any particular order.

[0249] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A live follow-along video detection method, characterized in that, The live follow-along video detection method includes: Obtain a live sports video and obtain a follow-along video based on the live sports video; Calculate the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow-along video, including: identifying the start generation time point of the start video frame in the live sports video and identifying the end generation time point of the end video frame in the live sports video; calculating the difference between the end generation time point and the start generation time point to obtain the first video duration; calculating the ratio of the second video duration to the first video duration to obtain the playback frequency; Extract a standard video sequence from the live sports video based on the playback frequency, and extract a follow-along action sequence from the follow-along video; Detect the standard key points of each standard video frame in the standard video sequence, and detect the user key points of each follow-along frame in the follow-along action sequence; Identify the standard action angle of each standard video frame based on the standard key points, and identify the follow-along action angle of each follow-along frame based on the user key points; Generate a follow-along rating for the follow-along video according to the standard action angle and the follow-along action angle.

2. The live follow-along video detection method according to claim 1, wherein, The extracting a standard video sequence from the live sports video based on the playback frequency includes: Compare the playback frequency with a preset frequency; If the playback frequency is greater than the preset frequency, convert the live sports video based on the playback frequency to obtain a standard sports video, and generate the standard video sequence according to the standard video frames in the standard sports video and the frame positions of the standard video frames in the standard sports video; or If the playback frequency is less than or equal to the preset frequency, extract the standard video frames from the live sports video as the standard video sequence.

3. The live follow-along video detection method according to claim 1, wherein The detecting the standard key points of each standard video frame in the standard video sequence includes: For each standard video frame, obtain the pixel value and pixel position of each pixel point in the standard video frame; Detect the pixel value and the pixel position based on a pre-trained human body detection model to obtain the detection category corresponding to each pixel point and the detection probability of the pixel point belonging to the detection category; Screen target pixel points from multiple pixel points based on the detection probability; Determine the area formed by adjacent pixel points corresponding to the same detection category among the target pixel points as the standard key points based on the pixel position and the detection category.

4. The live follow-along video detection method according to claim 3, wherein The identifying the standard action angle of each standard video frame based on the standard key points includes: Construct a plane rectangular coordinate system with any standard key point as the origin and the image side parallel to any standard video frame as the coordinate axis; Identify the key point coordinate value corresponding to the standard key point according to the pixel position corresponding to the standard key point and the plane rectangular coordinate system; Obtain a connection key point pair of the any standard key point from multiple standard key points, and the connection key point pair includes a first connection point and a second connection point; Determine the connection edge formed by any one of the standard key points and the first connection point as the first connection edge, and calculate the angle between the first connection edge and the coordinate axis based on the key point coordinate value of the first connection point to obtain a first angle; Determine the connection edge formed by any one of the standard key points and the second connection point as the second connection edge, and calculate the angle between the second connection edge and the coordinate axis based on the key point coordinate value of the second connection point to obtain a second angle; Generate the standard action angle based on the first angle, the second angle, and a preset integer value.

5. The live follow-along video detection method according to claim 4, wherein, The calculation formula for the standard action angle is: ; Among them, represents the standard action angle, represents the second angle, represents the first angle, is the preset integer value, .

6. The live follow-along video detection method according to claim 4, wherein The generation of the follow - up rating of the follow - up video according to the standard action angle and the follow - up action angle includes: Obtain the key point weight threshold corresponding to any one of the standard key points; Generate the coincidence degree of each follow - up frame according to the standard action angle, the follow - up action angle, and the corresponding key point weight threshold, and the calculation formula for the coincidence degree is: , where represents the degree of coincidence, represents the key point weight thresholds corresponding to multiple said standard key points, represents multiple said standard action angles, represents multiple said following exercise action angles; Generate the follow - up rating according to the coincidence degree and a preset mapping relationship.

7. A live follow-along video detection device, characterized in that, The live follow - up video detection device includes: An acquisition unit, configured to acquire a live sports video and acquire a follow - up video based on the live sports video; A calculation unit, configured to calculate the playback frequency of the live sports video according to the first video duration of the live sports video and the second video duration of the follow - up video, including: identifying the start generation time point of the start video frame in the live sports video, and identifying the end generation time point of the end video frame in the live sports video; calculating the difference between the end generation time point and the start generation time point to obtain the first video duration; calculating the ratio of the second video duration to the first video duration to obtain the playback frequency; An extraction unit, configured to extract a standard video sequence from the live sports video based on the playback frequency, and extract a follow - up action sequence from the follow - up video; A detection unit, configured to detect the standard key points of each standard video frame in the standard video sequence, and detect the user key points of each follow - up frame in the follow - up action sequence; An identification unit, configured to identify the standard action angle of each standard video frame based on the standard key points, and identify the follow - up action angle of each follow - up frame based on the user key points; A generation unit, configured to generate the follow - up rating of the follow - up video according to the standard action angle and the follow - up action angle.

8. An electronic device, characterized in that, The electronic device includes: A memory, storing computer - readable instructions; and A processor, executing the computer - readable instructions stored in the memory to implement the live follow - up video detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer - readable storage medium stores computer - readable instructions, and the computer - readable instructions are executed by a processor in an electronic device to implement the live follow - up video detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Follow-up practice mode control method and display equipment

    CN112272324A

  • Video detection method, video detection device, storage medium and electronic equipment

    CN114170554A