Shake identification method and device, computer equipment and computer readable storage medium

By obtaining the position information of the key points of the human body in the video stream, calculating the shaking amplitude and frequency of the left and right shoulders and ankles, combined with Fourier transform and cross-correlation analysis, the problem of difficult to identify the shaking posture in the existing technology is solved, and accurate shaking recognition and quantification is achieved.

CN120472523APending Publication Date: 2025-08-12GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410157623.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-02
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing body language recognition technology is difficult to reliably identify and distinguish teachers' shaking postures. The unconscious shaking movements that are more common among teacher students or novice teachers are easily confused with other movements, resulting in increased recognition difficulty.

Method used

By obtaining video streams, we can identify the position information of the key points of the human body, including the position information of the key points of the left and right shoulders and the key points of the left and right ankles, calculate their shaking amplitude and frequency, and combine Fourier transform and cross-correlation analysis to generate shaking identification information.

Benefits of technology

Accurate identification and quantification of shaking postures is achieved, misjudgment is reduced, and it adapts to different environments and conditions, which improves the accuracy and generalization of the identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472523A_ABST
    Figure CN120472523A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a shake recognition method and device, computer equipment and a computer readable storage medium, and the method comprises the steps: obtaining a video stream; determining human body key point position information according to the video stream, wherein the human body key point position information comprises position information of left and right shoulder key points and position information of left and right ankle key points; determining a first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points; determining second shaking amplitudes of the left and right ankle key points according to the position information of the left and right ankle key points; and generating shake identification information according to the first shake amplitude and the second shake amplitude. According to the method, key body parts (shoulders and ankles) are analyzed in a centralized manner, so that whether the current body is in a shaking action or not can be identified and analyzed more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a shake recognition method, device, computer equipment, and computer-readable storage medium. Background Art

[0002] In the teaching environment, a teacher's body language, including head posture, facial expressions, gestures, and body distance, plays a crucial role. These non-verbal elements not only effectively convey information and emotions but also significantly enhance student engagement and improve the quality of classroom instruction. With technological advancements, artificial intelligence (AI) image recognition algorithms are increasingly being used to automatically identify teachers' body language, replacing traditional manual quantitative evaluation methods. The objectivity and efficiency brought by this automated recognition are extremely beneficial to the self-learning and growth of the teaching staff.

[0003] However, existing body language recognition technologies primarily focus on identifying gestures and postures, with insufficient attention paid to body "shaking." This "shaking" movement, particularly common among teacher trainees and novice teachers, is typically an unconscious and spontaneous physical expression. This behavior often leaves the observer with the impression of lack of confidence, instability, and even unprofessionalism, and requires practice to overcome and improve. While human observers can easily identify and judge this shaking, automatic recognition algorithms can easily confuse it with other similar movements, such as pacing or turning back and forth, making it more difficult to identify. Summary of the Invention

[0004] An object of the embodiments of the present invention is to provide a shake recognition method, apparatus, computer device, and computer-readable storage medium to solve the technical problem that related technologies cannot reliably recognize shake gestures.

[0005] In a first aspect, an embodiment of the present invention provides a shaking recognition method, comprising:

[0006] Get the video stream;

[0007] Determine the position information of key points of the human body according to the video stream, wherein the position information of key points of the human body includes the position information of the left and right shoulder key points and the position information of the left and right ankle key points;

[0008] Determining a first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points;

[0009] Determining a second shaking amplitude of the left and right ankle key points according to the position information of the left and right ankle key points;

[0010] The shake identification information is generated according to the first shake amplitude and the second shake amplitude.

[0011] In combination with the first aspect, in a possible implementation method, determining the first shaking amplitude of the left and right shoulder key points based on the position information of the left and right shoulder key points includes: determining a reference shaking frequency based on the position information of the left and right shoulder key points, the reference shaking frequency being the frequency at which the left and right shoulder key points swing at the same frequency to a maximum amplitude; determining the first shaking amplitude of the left and right shoulder key points based on the reference shaking frequency and the position information of the left and right shoulder key points.

[0012] It can be seen that this method can accurately capture and quantify the shaking behavior of the shoulder from the perspective of data analysis. By calculating the synchronous swing frequency and corresponding amplitude of the key points of the left and right shoulders, the shaking amplitude of the shoulder can be accurately determined, which helps to identify the shaking movement of the body.

[0013] In combination with the first aspect, in a possible implementation method, determining the baseline shaking frequency based on the position information of the left and right shoulder key points includes: performing cross-correlation calculation on the position information of the left shoulder key point and the position information of the right shoulder key point to obtain cross-correlation position information; performing Fourier transform on the cross-correlation position information to obtain a cross-correlation spectrum diagram; and the frequency corresponding to the maximum amplitude traversed in the cross-correlation spectrum diagram is the baseline shaking frequency.

[0014] It can be seen that this method calculates the relevant frequencies of the left and right shoulders through cross-correlation, and performs "shaking" identification and judgment based on the amplitude of the swing component of the left / right shoulder at this frequency. Compared with frequency analysis using a single point (such as the center of the left or right shoulder), this method is more discriminatory and can well distinguish similar movements such as pacing / turning back and forth.

[0015] In combination with the first aspect, in a possible implementation method, determining the first sway amplitude of the left and right shoulder key points based on the reference sway frequency and the position information of the left and right shoulder key points includes: performing Fourier transform on the position information of the left shoulder key point to obtain a left shoulder spectrum diagram; traversing the amplitude with a frequency equal to the reference sway frequency in the left shoulder spectrum diagram is the first left shoulder amplitude; performing Fourier transform on the position information of the right shoulder key point to obtain a right shoulder spectrum diagram; traversing the amplitude with a frequency equal to the reference sway frequency in the right shoulder spectrum diagram is the first right shoulder amplitude; and determining the first sway amplitude based on the first left shoulder amplitude and the first right shoulder amplitude.

[0016] It can be seen that this method uses the position information of the left and right shoulder key points combined with the baseline shaking frequency to determine the shaking amplitude of the left and right shoulders, thereby accurately quantifying and analyzing the individual shoulder shaking amplitude.

[0017] In combination with the first aspect, in a possible implementation method, determining the second swing amplitude of the left and right ankle key points based on the position information of the left and right ankle key points includes: performing Fourier transform on the position information of the left ankle key point to obtain a left ankle spectrum diagram; traversing the amplitude with a frequency equal to the reference swing frequency in the left ankle spectrum diagram is the first left ankle amplitude; performing Fourier transform on the position information of the right ankle key point to obtain a right ankle spectrum diagram; traversing the amplitude with a frequency equal to the reference swing frequency in the right ankle spectrum diagram is the first right ankle amplitude; and determining the second swing amplitude based on the first left ankle amplitude and the first right ankle amplitude.

[0018] It can be seen that in order to distinguish from similar movements such as pacing and turning back and forth, this method adds ankle key points to the frequency analysis, uses the position information of the left and right ankle key points, and combines it with the baseline shaking frequency to determine the shaking amplitude of the left and right ankles, which can accurately capture and analyze the shaking amplitude of individual ankles.

[0019] In combination with the first aspect, in a possible implementation method, determining the position information of key points of the human body based on the video stream includes: extracting multiple frames of character images from the video stream; extracting the position information of the key points of the parts from each frame of the character image according to a preset key point extraction algorithm; and normalizing the position information of each key point of the parts to obtain the position information of the key points of the human body.

[0020] It can be seen that this method can effectively extract and process the position information of key points of the human body from the video stream, providing a reliable data basis for subsequent motion analysis, posture evaluation or shake recognition.

[0021] In combination with the first aspect, in a possible implementation method, the normalizing the position information of each key point of the part to obtain the position information of the key points of the human body includes: obtaining multiple sets of left and right shoulder position information, the multiple sets of left and right shoulder position information are the position information of the left and right shoulder key points extracted from each frame of the character image; obtaining the average shoulder distance based on the multiple sets of left and right shoulder position information; and normalizing the position information of each key point of the part based on the average shoulder distance to obtain the position information of the key points of the human body.

[0022] It can be seen that this method uses the shoulder width in the swing direction as a reference for data normalization, effectively avoiding the influence of factors such as the distance and resolution of the camera installation on the judgment threshold, and improving the generalization of the shaking recognition method.

[0023] In combination with the first aspect, in a possible implementation method, the shake identification information includes shake presence information or shake non-existence information, and the generation of shake identification information based on the first shake amplitude and the second shake amplitude includes: judging whether the first shake amplitude is greater than a first preset amplitude, and judging whether the second shake amplitude is less than a second preset amplitude; when the first shake amplitude is greater than the first preset amplitude, and the second shake amplitude is less than the second preset amplitude, generating shake presence information; when the first shake amplitude is less than or equal to the first preset amplitude, or when the second shake amplitude is greater than or equal to the second preset amplitude, generating shake non-existence information.

[0024] It can be seen that the method adopts a composite judgment mechanism, which comprehensively considers the shaking of the shoulders and ankles to improve the accuracy and reliability of the shaking state judgment.

[0025] In a second aspect, an embodiment of the present invention provides a shaking recognition device, comprising:

[0026] An acquisition unit, used for acquiring a video stream;

[0027] a determining unit, configured to determine position information of key points of a human body according to the video stream, wherein the position information of key points of the human body includes position information of left and right shoulder key points and position information of left and right ankle key points;

[0028] The determining unit is further configured to determine a first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points;

[0029] The determining unit is further configured to determine a second shaking amplitude of the left and right ankle key points according to the position information of the left and right ankle key points;

[0030] A generating unit is configured to generate shake identification information according to the first shake amplitude and the second shake amplitude.

[0031] In a third aspect, an embodiment of the present invention provides a computer device, including:

[0032] at least one processor; and,

[0033] a memory communicatively connected to the at least one processor; wherein,

[0034] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the first aspect.

[0035] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to the first aspect.

[0036] In the scheme implemented by the above-mentioned motion identification method, device, computer device and computer-readable storage medium, a video stream is first obtained. Then, the position information of key points of the human body is determined based on the video stream. The position information of key points of the human body includes the position information of the left and right shoulder key points and the position information of the left and right ankle key points. Then, a first motion amplitude of the left and right shoulder key points is determined based on the position information of the left and right shoulder key points, and a second motion amplitude of the left and right ankle key points is determined based on the position information of the left and right ankle key points. Finally, motion identification information is generated based on the first motion amplitude and the second motion amplitude. By processing the video stream, this method accurately obtains and analyzes the position information of key points of the human body (such as shoulders and ankles), which can more reliably and accurately identify and quantify the motion amplitude of the body, and further accurately identify and analyze whether the person's body is shaking. By analyzing the data of multiple key points, it can adapt to different environments and conditions, such as different camera settings and lighting conditions, thereby improving the accuracy and generalization of motion identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0038] Figure 1 is a scene graph for shaking recognition in one embodiment of the present invention;

[0039] Figure 2A 1 is a flow chart of a shaking recognition method according to an embodiment of the present invention;

[0040] Figure 2B 1 is a schematic diagram of a processing result of a shaking recognition method according to an embodiment of the present invention;

[0041] Figure 3 1 is a schematic structural diagram of a shaking recognition device according to an embodiment of the present invention;

[0042] Figure 4 It is a structural diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0044] It should be noted that, unless there is a conflict, the various features of the embodiments of the present invention may be combined with each other and are all within the scope of protection of the present invention. In addition, although the functional modules are divided in the device schematics and the logical order is shown in the flow charts, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flow charts. Furthermore, the terms "first," "second," "third," etc. used in the present invention do not limit the data or execution order, but only distinguish between identical or similar items with substantially the same functions and effects.

[0045] The technical solution of this application can be applied to various shaking recognition scenarios, such as teaching scenarios, medical scenarios, etc.

[0046] For example, see Figure 1 , Figure 1 A schematic diagram of a teaching scenario is provided for an embodiment of the present application. In a classroom, there are multiple students and a teacher A. A camera is placed on the top of the classroom. The camera continuously collects videos of teacher A during class and sends the collected videos to a server. The server performs shaking recognition on teacher A according to the present method. The server obtains a video stream and determines the position information of key points of the human body according to the video stream. The position information of key points of the human body includes the position information of left and right shoulder key points and the position information of left and right ankle key points. The first shaking amplitude of the left and right shoulder key points is determined according to the position information of the left and right shoulder key points. The second shaking amplitude of the left and right ankle key points is determined according to the position information of the left and right ankle key points. Shake recognition information is generated according to the first shaking amplitude and the second shaking amplitude. Figure 1 In the video stream, it is detected that Teacher A is shaking. Furthermore, in real life, it can provide real-time recognition and reminders when teacher trainees experience body shaking during practical training, helping them learn and grow.

[0047] In view of this, the present application proposes a shaking recognition method to solve the above problem, which is described in detail below.

[0048] See also Figure 2A , Figure 2A A flowchart of a shake recognition method provided by an embodiment of the present invention is provided. The method includes the following steps:

[0049] S10. Obtain video stream.

[0050] The video stream may be a real-time captured video, such as a classroom scene recorded by a camera, or a pre-recorded video file.

[0051] Optionally, the captured video stream is transferred to a computer or server over a network or direct connection.

[0052] The acquired video stream may be pre-processed, such as adjusting the resolution, contrast, and brightness. This processing can better identify the subsequent process of identifying key points of the human body.

[0053] This method establishes the foundation for motion recognition by acquiring video streams, which provide the necessary visual information for extracting human motion and posture data. In real-time applications, such as classroom teaching, live video streams allow for immediate analysis of the teacher's movements, providing timely feedback and suggestions for improvement. In non-real-time scenarios, such as research and training, pre-recorded videos allow for more in-depth analysis of specific situations or behaviors.

[0054] S20. Determine position information of key points of the human body according to the video stream, where the position information of key points of the human body includes position information of left and right shoulder key points and position information of left and right ankle key points.

[0055] Among them, human body detection can be performed on the video stream first, and after the target person is determined, the target person is tracked, the human body frame is output, and the human body image is cropped from the original image according to the human body frame. 2D key point detection is performed on the human body image cropped from the human body frame, and the position information of the key points of the parts is extracted from each frame of the character image. The position information of the detected key points of the parts is stored in a queue (Q_kpts), and the position information of each key point of the parts is stored in a queue. The position information of each key point of the parts is further normalized to obtain the position information of the key points of the human body.

[0056] A queue is a data structure used to store and process keypoint data from consecutive frames. The length of the queue (Q_N) determines the time range that can be analyzed. At a fixed sampling rate (fps, or frames per second), the queue length determines the temporal and frequency resolutions of the data.

[0057] Specifically, the YOLOv7 network is used to detect the human body in the video stream, identify the position of the human body in the video frame, and generate a bounding box (bbox). Then, SiamRPN++ is used to track the detected human body to ensure that the human body can be continuously tracked even in consecutive video frames. The key point model of MediaPipe is used to detect 2D key points of the human body image (image_body) cropped from the human body frame (bbox), identify the key points of the body (such as left and right shoulders and left and right ankles), and generate the key point coordinates (body_kpts).

[0058] Among them, MediaPipe is an efficient computer vision tool for real-time identification of key points of the human body (such as head, shoulders, arms, ankles, etc.) in videos or images.

[0059] SiamRPN++ is an algorithm for tracking single targets in video streams. It is a combination of the Siamese network and the Region Proposal Network (RPN) and is used to track specific targets in videos in real time.

[0060] Among them, YOLOv7 is used for real-time target detection. YOLOv7 can be used to first quickly and accurately detect the human body from the video stream, and then use SiamRPN++ to perform continuous single-target tracking of the detected human body. This combined method ensures that the target can be tracked continuously and accurately even when the target moves or other interference factors appear in the video.

[0061] It can be seen that this method provides the necessary data for subsequent steps to analyze the shaking motion of the human body by identifying and locating the key points of the human body in the video.

[0062] S30. Determine a first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points.

[0063] The position information of the left and right shoulder key points may include their two-dimensional coordinates (x, y) in the video frame.

[0064] Here, "shaking" refers to the back and forth swinging of the left and right shoulders along the x-axis.

[0065] The first shake amplitude refers to the range of movement of the keypoint position information along a specific axis (usually the horizontal axis) within a certain period of time. The shoulder shake amplitude can be determined by calculating the positional changes of the left and right shoulder keypoints in a video frame sequence. This involves measuring the relative position changes of the shoulder keypoints between different frames.

[0066] The sway amplitude can be determined by measuring the extreme value of the position change of the shoulder key point over a certain period of time. For example, the difference between the maximum and minimum x-coordinate values of the shoulder key point over a period of time can be measured.

[0067] Among them, key point detection may be affected by noise. Filtering algorithms, such as Kalman filter, can be used to help smooth the data and improve the accuracy of shake amplitude calculation.

[0068] It can be seen that this method can more accurately identify the shaking movement of the body by accurately measuring the movement amplitude of the shoulder, thereby improving the accuracy and efficiency of shaking recognition.

[0069] S40: Determine a second shaking amplitude of the left and right ankle key points according to the position information of the left and right ankle key points.

[0070] The positions of the left and right ankle key points are usually represented by two-dimensional coordinates (x, y) in the video frame.

[0071] The second sway amplitude refers to the range of movement of the key point position information along a specific axis (usually the horizontal axis) within a certain period of time. The ankle sway amplitude can be quantified by analyzing the changes in the position of the ankle key points in consecutive video frames. This typically involves comparing the position differences of the ankle key points between consecutive frames to determine the range or amplitude of the ankle's movement.

[0072] This method helps distinguish ankle sway from other similar body movements (such as walking or turning) by analyzing the ankle sway amplitude. Different movements will result in different ankle sway amplitudes, which improves the accuracy of sway identification and provides important data support for understanding the overall movement and balance of the body.

[0073] S50: Generate shake identification information according to the first shake amplitude and the second shake amplitude.

[0074] The above-mentioned shaking identification information includes shaking presence information and shaking absence information.

[0075] The first shaking amplitude data of the shoulder and the second shaking amplitude data of the ankle are combined to determine whether the human body in the video stream is shaking.

[0076] Specifically, if the shoulder and ankle sway exceeds a preset threshold, sway is considered present; if it falls below the threshold, sway is considered absent. Setting sway thresholds is based on empirical data, user feedback, or expert advice, taking into account individual differences and the needs of different scenarios. The shoulder and ankle thresholds can be different.

[0077] It can be seen that this method can more accurately determine whether an individual has shaking behavior by comprehensively considering the shaking data of the upper part (shoulders) and lower part (ankles) of the body, reduce misjudgment, and improve the accuracy of shaking recognition.

[0078] By processing video streams, this method accurately obtains and analyzes the position information of key points of the human body (such as shoulders and ankles), which can more accurately identify and quantify the amplitude of body shaking, and further accurately identify and analyze whether the person's body is shaking. By analyzing data from multiple key points, it can adapt to different environments and conditions, such as different camera settings and lighting conditions, thereby improving the accuracy and generalization of shaking recognition.

[0079] In one embodiment, determining the first shaking amplitude of the left and right shoulder key points based on the position information of the left and right shoulder key points includes: determining a reference shaking frequency based on the position information of the left and right shoulder key points, the reference shaking frequency being the frequency at which the left and right shoulder key points swing at the same frequency to a maximum amplitude; determining the first shaking amplitude of the left and right shoulder key points based on the reference shaking frequency and the position information of the left and right shoulder key points.

[0080] The reference sway frequency is the maximum amplitude frequency generated when the left and right shoulder key points swing synchronously. This can be determined by performing a Fourier transform (FFT) on the position changes of the shoulder key points to obtain the main frequency components of the shoulder motion.

[0081] Among them, after determining the baseline shaking frequency, the shaking amplitude of the shoulder is calculated according to this frequency and the position information of the shoulder key points. Specifically, the shaking amplitude can be determined by analyzing the displacement of the shoulder key points at the baseline shaking frequency, and by calculating the average amplitude or extreme value of the key point position change at this frequency.

[0082] Among them, the above amplitude reflects the intensity of shoulder movement at the same frequency.

[0083] Optionally, in addition to the sway amplitude, the directionality and speed of the shoulder movement can be further analyzed to obtain a more comprehensive description of the sway characteristics.

[0084] It can be seen that this method can accurately capture and quantify the shaking behavior of the shoulder from the perspective of data analysis. By calculating the synchronous swing frequency and corresponding amplitude of the key points of the left and right shoulders, the shaking amplitude of the shoulder can be accurately determined, which helps to identify the shaking movement of the body.

[0085] In one embodiment, determining the reference晃动 frequency based on the position information of the left and right shoulder key points includes: performing a cross-correlation calculation on the position information of the left shoulder key point and the position information of the right shoulder key point to obtain cross-correlation position information; performing a Fourier transform on the cross-correlation position information to obtain a cross-correlation spectrogram; and traversing the cross-correlation spectrogram to find the frequency corresponding to the maximum amplitude as the reference晃动 frequency.

[0086] Among them, the above cross-correlation calculation is used to detect the similarity or correlation between two data sequences, especially to determine whether they change synchronously at the same frequency.

[0087] Among them, calculating the cross-correlation between the positions of the left shoulder key point and the right shoulder key point is to reveal whether the two key points move in a similar pattern or frequency, that is, whether there is synchronous晃动.

[0088] Among them, a fast Fourier transform (FFT) is performed on the data obtained from the cross-correlation calculation. This process converts the time series data into the frequency domain, so that the dominant motion frequency can be identified. The result of the FFT is a spectrogram, which is used to display the amplitudes at different frequencies.

[0089] Among them, in the cross-correlation spectrogram, each frequency point is traversed to find the point with the largest amplitude. This point is the frequency with the largest amplitude, and this frequency is considered the reference晃动 frequency of the synchronous swing of the left and right shoulders, that is, the main晃动 frequency of the left and right shoulders. That is, the reference晃动 frequency represents the main frequency of the synchronous swing of the left and right shoulder key points and can be regarded as the characteristic frequency of the shoulder晃动.

[0090] Optionally, before performing the FFT, it may be necessary to apply a window function (such as a Hanning window or a Hamming window) to process the data to reduce the boundary effect and improve the frequency resolution.

[0091] Among them, a frequency range can be set to exclude irrelevant frequencies. For example, very low or very high frequencies may not represent ordinary shoulder晃动.

[0092] Among them, when selecting the parameters of the FFT (such as the window size and the step size), it is necessary to balance the time resolution and the frequency resolution. A larger window can provide more accurate frequency information but reduces the temporal accuracy.

[0093] For example, perform a cross-correlation calculation on the position information of the left and right shoulder key points to obtain a cross-correlation sequence cor_5x6x, then perform an FFT transformation on it, and within the set frequency threshold range [th_f0, th_f1] (where th_f0 < th_f1, and in the example, th_f0 = 0.25Hz and th_f1 = 2Hz), find the frequency with the largest amplitude, which is recorded as the reference晃动 frequency of the same-frequency swing of the left and right shoulders.

[0094] It can be seen that this method calculates the relevant frequencies of the left and right shoulders through cross-correlation, and performs "shaking" identification and judgment based on the amplitude of the swing component of the left / right shoulder at this frequency. Compared with frequency analysis using a single point (such as the center of the left or right shoulder), this method is more discriminatory and can well distinguish similar movements such as pacing / turning back and forth.

[0095] In one embodiment, determining the first sway amplitude of the left and right shoulder key points based on the reference sway frequency and the position information of the left and right shoulder key points includes: performing Fourier transform on the position information of the left shoulder key point to obtain a left shoulder spectrum diagram; traversing the amplitude with a frequency equal to the reference sway frequency in the left shoulder spectrum diagram to be the first left shoulder amplitude; performing Fourier transform on the position information of the right shoulder key point to obtain a right shoulder spectrum diagram; traversing the amplitude with a frequency equal to the reference sway frequency in the right shoulder spectrum diagram to be the first right shoulder amplitude; and determining the first sway amplitude based on the first left shoulder amplitude and the first right shoulder amplitude.

[0096] Among them, the position information of the left shoulder key points is subjected to Fourier transform, and the time series data (the position change of the shoulder in the video frame) is converted into frequency domain data to generate a spectrum diagram of the left shoulder.

[0097] Similarly, the position information of the key points of the right shoulder is also subjected to Fourier transform to obtain the spectrum of the right shoulder.

[0098] Next, in the left shoulder spectrum, find a frequency point that matches the reference shake frequency and record the amplitude of that frequency point as the first left shoulder amplitude. Similarly, in the right shoulder spectrum, find an amplitude that matches the reference shake frequency and record that as the first right shoulder amplitude.

[0099] Before performing FFT, the key point position data is smoothed or noise removed to improve the accuracy and reliability of spectrum analysis.

[0100] The first sway amplitude can be achieved by simple averaging, weighted averaging, or other statistical methods, which are not limited here. The obtained left and right shoulder amplitude values are combined to obtain a comprehensive value representing the overall shoulder sway intensity, which is the first sway amplitude.

[0101] Optionally, in addition to spectral analysis, temporal analysis can be combined to more fully understand the characteristics of shoulder motion, such as the smoothness or irregularity of the motion.

[0102] It can be seen that this method uses the position information of the left and right shoulder key points combined with the baseline shaking frequency to determine the shaking amplitude of the left and right shoulders, thereby accurately quantifying and analyzing the individual shoulder shaking amplitude.

[0103] In one embodiment, determining the second swing amplitude of the left and right ankle key points based on the position information of the left and right ankle key points includes: performing Fourier transform on the position information of the left ankle key point to obtain a left ankle spectrum diagram; traversing the amplitude with a frequency equal to the reference swing frequency in the left ankle spectrum diagram is a first left ankle amplitude; performing Fourier transform on the position information of the right ankle key point to obtain a right ankle spectrum diagram; traversing the amplitude with a frequency equal to the reference swing frequency in the right ankle spectrum diagram is a first right ankle amplitude; and determining the second swing amplitude based on the first left ankle amplitude and the first right ankle amplitude.

[0104] Specifically, the position information of the left ankle is subjected to Fourier transform, the time series data is converted into frequency domain data, and a spectrum diagram of the left ankle is generated. In the spectrum diagram of the left ankle, a frequency point that matches the baseline shaking frequency is found, and the amplitude of the frequency point is recorded as the shaking amplitude of the left ankle at the baseline shaking frequency (the first left ankle amplitude).

[0105] Similarly, the position information of the right ankle is Fourier transformed to obtain the spectrum of the right ankle. In the spectrum of the right ankle, the frequency point that matches the reference shaking frequency is found, and the amplitude of the frequency point is recorded as the shaking amplitude of the right ankle at the reference shaking frequency (the first right ankle amplitude).

[0106] The second sway amplitude can be achieved by simple averaging, weighted averaging, or other statistical methods, which are not limited here. The obtained left and right ankle amplitude values are combined to obtain a comprehensive value representing the overall ankle sway intensity, which is the second sway amplitude.

[0107] It can be seen that in order to distinguish from similar movements such as pacing and turning back and forth, this method adds ankle key points to the frequency analysis, uses the position information of the left and right ankle key points, and combines it with the baseline shaking frequency to determine the shaking amplitude of the left and right ankles, which can accurately capture and analyze the shaking amplitude of individual ankles.

[0108] In one embodiment, determining the position information of key points of the human body based on the video stream includes: extracting multiple frames of character images from the video stream; extracting the position information of the key points of the parts from each frame of the character image according to a preset key point extraction algorithm; and normalizing the position information of each key point of the parts to obtain the position information of the key points of the human body.

[0109] The multiple frames of character images represent the postures and positions of the characters at different time points.

[0110] The number and frequency of frames extracted depend on the frame rate of the video and the required analysis accuracy. More frames can provide more continuous and detailed motion data.

[0111] Each frame of the image can be analyzed using a preset keypoint extraction algorithm (such as OpenPose, MediaPipe, etc.). These algorithms can identify various parts of the human body, such as the head, shoulders, arms, and legs. For each frame of the image, the algorithm identifies and extracts the location information of each keypoint on the human body, typically including the two-dimensional coordinates (x, y) of each keypoint in the image. For specific steps, refer to step S20.

[0112] Specifically, as mentioned in S20, the position information of each key point of the body part is stored in a queue, and the position information of each key point of the body part is further normalized to obtain the position information of the key points of the body part. The normalized key point position information is arranged in chronological order and can be used for dynamic analysis, such as calculation of motion trajectory, velocity, and acceleration.

[0113] Normalization is performed to eliminate the effects of factors such as image size, camera distance, and shooting angle on keypoint location information, making the analysis more accurate and universal. Normalization typically involves converting the coordinates of keypoints into proportions or ratios that are independent of body or image size. For example, a fixed length of the human body (such as shoulder width) can be used as a reference to adjust and scale the positions of other keypoints.

[0114] It can be seen that this method can effectively extract and process the position information of key points of the human body from the video stream, providing a reliable data basis for subsequent motion analysis, posture evaluation or shake recognition.

[0115] In one embodiment, the normalizing the position information of each key point of the part to obtain the position information of the key points of the human body includes: obtaining multiple sets of left and right shoulder position information, the multiple sets of left and right shoulder position information are the position information of the left and right shoulder key points extracted from each frame of the character image; obtaining the average shoulder distance based on the multiple sets of left and right shoulder position information; and normalizing the position information of each key point of the part based on the average shoulder distance to obtain the position information of the key points of the human body.

[0116] The process involves extracting the positional information of the key points of the left and right shoulders of a person from each frame of the video stream. This typically involves using computer vision algorithms (such as OpenPose and MediaPipe) to identify the human body structure and determine the coordinates (e.g., x and y coordinates) of the shoulder key points. This process is repeated for each frame in the video, thereby collecting a series of data on the positions of the left and right shoulder key points.

[0117] The shoulder distance can be calculated by calculating the distance between the left shoulder and the right shoulder based on the shoulder key points in each frame image, that is, by directly measuring the Euclidean distance between the two points.

[0118] The shoulder distance mean value represents the average level of shoulder width of the characters in the video. The shoulder distances calculated in all frames are averaged to obtain an average value of the shoulder distance.

[0119] Specifically, the calculated mean shoulder distance is used as a reference to scale and adjust the position information of each key point. For example, the coordinates of the key point can be divided by the mean shoulder distance, so that the position information is independent of the actual size of the individual.

[0120] For example, the specific steps of normalization processing can refer to the following formula:

[0121] w_shoulder=mean(abs(kpt_6x[front]-kpt_5x[front]));

[0122] kpt_*x=(kpt_*x-mean(kpt_*x)) / w_shoulder;

[0123] Among them, *∈{0,...,16}, 5 and 6 represent the key point numbers of the left and right shoulders respectively, and front represents the front view. That is, the shoulder width w_shoulder is the distance between the left and right shoulder key points when the body is front view, which is normalized to the key point coordinates minus their respective means and then divided by the shoulder width.

[0124] It can be seen that this method uses the shoulder width in the swing direction as a reference for data normalization, effectively avoiding the influence of factors such as the distance and resolution of the camera installation on the judgment threshold, and improving the generalization of the shaking recognition method.

[0125] In one embodiment, the shake identification information includes shake presence information or shake non-existence information, and the generation of shake identification information based on the first shake amplitude and the second shake amplitude includes: judging whether the first shake amplitude is greater than a first preset amplitude, and judging whether the second shake amplitude is less than a second preset amplitude; when the first shake amplitude is greater than the first preset amplitude, and the second shake amplitude is less than the second preset amplitude, generating shake presence information; when the first shake amplitude is less than or equal to the first preset amplitude, or when the second shake amplitude is greater than or equal to the second preset amplitude, generating shake non-existence information.

[0126] The shaking recognition information includes shaking presence information or shaking absence information. The shaking presence information indicates that the current character has shaking motion, and the shaking absence information indicates that the current character does not have shaking motion.

[0127] Here, whether the first shaking amplitude is greater than a preset first preset amplitude is determined, and the first preset amplitude is determined based on observation and analysis of normal movement and shaking state, and is used to distinguish normal movement from excessive shaking.

[0128] Among them, it is determined whether the second shaking amplitude is less than a preset second preset amplitude. The second preset amplitude confirms whether the shaking of the ankle is within a normal range, which helps to distinguish the overall shaking state of the body.

[0129] Specifically, if the first shaking amplitude exceeds its first preset amplitude (i.e., the shoulder shakes too much) and the second shaking amplitude is lower than its second preset amplitude (i.e., the ankle remains relatively stable), it is determined that the person in the video is in a shaking state.

[0130] The first preset amplitude and the second preset amplitude may be obtained based on empirical data, expert opinions, or by analyzing a large amount of sample data, and are not limited here.

[0131] For example, refer to Figure 2B , Figure 2B The diagram shows the processing results of body shaking, where the first column is the normalized key point coordinate sequence, the second column is the cross-correlation sequence of the left and right shoulders (all rows are consistent, where the numbers are the obtained correlation frequencies), and the third column is the amplitude diagram of the key point coordinates after FFT transformation, where the numbers are the amplitudes at the relevant frequencies. From top to bottom in the diagram, each row represents the key point processing results of the left and right shoulders and the left and right ankles. The minimum amplitude of the left and right shoulder swing is set to th_s, and the maximum amplitude of the left and right ankle swing is set to th_a (th_s = 0.1, th_a = 0.05 in the example). The following dual threshold judgment conditions are used to identify the body as "shaking". The judgment conditions are as follows:

[0132] min(am_5,am_6)>th_s

[0133] max(am_15,am_16) <th_a。

[0134] According to Figure 2B The current shaking judgment condition is met and it is identified as shaking.

[0135] It can be seen that this method adopts a composite judgment mechanism, which comprehensively considers the shaking of the shoulders and ankles to improve the accuracy and reliability of the shaking state judgment.

[0136] It should be noted that, in each of the above-mentioned embodiments, there is not necessarily a certain order between the above-mentioned steps. A person skilled in the art can understand, based on the description of the embodiments of this application, that in different embodiments, the above-mentioned steps may have different execution orders, that is, they may be executed in parallel, or may be executed interchangeably, etc.

[0137] As another aspect of the present invention, an embodiment of the present invention provides a shake recognition device. The shake recognition device may be a software module comprising a plurality of instructions stored in a memory, which a processor may access and execute to implement the shake recognition method described in each of the above embodiments.

[0138] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a shaking recognition device provided by an embodiment of the present application. Figure 3 As shown, the shaking recognition device includes:

[0139] An acquisition unit 301 is configured to acquire a video stream;

[0140] A determining unit 302 is configured to determine position information of key points of a human body according to the video stream, wherein the position information of key points of the human body includes position information of left and right shoulder key points and position information of left and right ankle key points;

[0141] The determining unit 302 is further configured to determine a first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points;

[0142] The determining unit 302 is further configured to determine a second shaking amplitude of the left and right ankle key points according to the position information of the left and right ankle key points;

[0143] The generating unit 303 is configured to generate shake identification information according to the first shake amplitude and the second shake amplitude.

[0144] By processing video streams, this method accurately obtains and analyzes the position information of key points of the human body (such as shoulders and ankles), which can more accurately identify and quantify the amplitude of body shaking, and further accurately identify and analyze whether the person's body is shaking. By analyzing data from multiple key points, it can adapt to different environments and conditions, such as different camera settings and lighting conditions, thereby improving the accuracy and generalization of shaking recognition.

[0145] In one embodiment, in determining the first shaking amplitude of the left and right shoulder key points based on the position information of the left and right shoulder key points, the determination unit 302 is further used to: determine a reference shaking frequency based on the position information of the left and right shoulder key points, the reference shaking frequency being the frequency at which the left and right shoulder key points swing at the same frequency to a maximum amplitude; and determine the first shaking amplitude of the left and right shoulder key points based on the reference shaking frequency and the position information of the left and right shoulder key points.

[0146] In one embodiment, in determining the reference shake frequency based on the position information of the left and right shoulder key points, the determination unit 302 is further used to: perform cross-correlation calculation on the position information of the left shoulder key point and the position information of the right shoulder key point to obtain cross-correlation position information; perform Fourier transform on the cross-correlation position information to obtain a cross-correlation spectrum diagram; and the frequency corresponding to the maximum amplitude traversed in the cross-correlation spectrum diagram is the reference shake frequency.

[0147] In one embodiment, in determining the first shake amplitude of the left and right shoulder key points based on the reference shake frequency and the position information of the left and right shoulder key points, the determination unit 302 is further used to: perform Fourier transform on the position information of the left shoulder key point to obtain a left shoulder spectrum diagram; traverse the amplitude with a frequency equal to the reference shake frequency in the left shoulder spectrum diagram to be the first left shoulder amplitude; perform Fourier transform on the position information of the right shoulder key point to obtain a right shoulder spectrum diagram; traverse the amplitude with a frequency equal to the reference shake frequency in the right shoulder spectrum diagram to be the first right shoulder amplitude; and determine the first shake amplitude based on the first left shoulder amplitude and the first right shoulder amplitude.

[0148] In one embodiment, in determining the second swing amplitude of the left and right ankle key points based on the position information of the left and right ankle key points, the determination unit 302 is further used to: perform Fourier transform on the position information of the left ankle key point to obtain a left ankle spectrum diagram; traverse the amplitude with a frequency equal to the reference swing frequency in the left ankle spectrum diagram to be the first left ankle amplitude; perform Fourier transform on the position information of the right ankle key point to obtain a right ankle spectrum diagram; traverse the amplitude with a frequency equal to the reference swing frequency in the right ankle spectrum diagram to be the first right ankle amplitude; and determine the second swing amplitude based on the first left ankle amplitude and the first right ankle amplitude.

[0149] In one embodiment, in determining the position information of key points of the human body based on the video stream, the determination unit 302 is further used to: extract multiple frames of character images from the video stream; extract the position information of the key points of the parts from each frame of the character image according to a preset key point extraction algorithm; and normalize the position information of each key point of the parts to obtain the position information of the key points of the human body.

[0150] In one embodiment, in the step of normalizing the position information of each key point of the part to obtain the position information of the key points of the human body, the determination unit 302 is further used to: obtain multiple sets of left and right shoulder position information, the multiple sets of left and right shoulder position information being the position information of the left and right shoulder key points extracted from each frame of the character image; obtain a mean shoulder distance based on the multiple sets of left and right shoulder position information; and normalize the position information of each key point of the part based on the mean shoulder distance to obtain the position information of the key points of the human body.

[0151] In one embodiment, the shake identification information includes shake presence information or shake non-existence information. In generating the shake identification information based on the first shake amplitude and the second shake amplitude, the generation unit 303 is further used to: determine whether the first shake amplitude is greater than a first preset amplitude, and determine whether the second shake amplitude is less than a second preset amplitude; when the first shake amplitude is greater than the first preset amplitude, and the second shake amplitude is less than the second preset amplitude, generate shake presence information; when the first shake amplitude is less than or equal to the first preset amplitude, or when the second shake amplitude is greater than or equal to the second preset amplitude, generate shake non-existence information.

[0152] It should be noted that the aforementioned shake recognition device can execute the shake recognition method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiments of the shake recognition device, please refer to the shake recognition method provided in the embodiments of this application.

[0153] See also Figure 4 , Figure 4 4 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device includes one or more processors 41 and a memory 42. The memory 42 is connected to the one or more processors 41, for example, via a bus.

[0154] The processor 41 is configured to support the computer device in executing the corresponding functions of the method in the above method embodiment. The processor 41 can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The above hardware chip can be an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0155] The memory 42 is used to store program code, etc. The memory 42 may include volatile memory (VM), such as random access memory (RAM); non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the aforementioned types of memory.

[0156] The memory 42 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the jitter recognition method in the embodiments of the present application. The processor 41 executes the non-volatile software programs, instructions, and modules stored in the memory 42 to execute the various functional applications and data processing of the jitter recognition method and jitter recognition device, thereby implementing the functions of the jitter recognition method and the various modules or units of the jitter recognition device provided in the above-mentioned method embodiments.

[0157] The memory 42 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data generated based on the use of the motion recognition device. In some embodiments, the memory 42 may optionally include a remote memory 42 located relative to the processor 41. These remote memories may be connected to the motion recognition device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0158] The one or more modules are stored in the memory 42. When executed by the one or more processors 41, the shaking recognition method in any of the above method embodiments is executed, for example, the method steps described in the above method embodiments are executed to realize the functions of the modules described in the above device embodiments.

[0159] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method as described in the above embodiment.

[0160] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0161] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A shaking recognition method, characterized in that: include: Get the video stream; Determine the position information of key points of the human body according to the video stream, wherein the position information of key points of the human body includes the position information of the left and right shoulder key points and the position information of the left and right ankle key points; Determining a first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points; Determining a second shaking amplitude of the left and right ankle key points according to the position information of the left and right ankle key points; The shake identification information is generated according to the first shake amplitude and the second shake amplitude.

2. The method according to claim 1, characterized in that Determining the first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points includes: Determine a reference shaking frequency according to the position information of the left and right shoulder key points, wherein the reference shaking frequency is the frequency at which the left and right shoulder key points swing at the same frequency to a maximum amplitude; The first shaking amplitude of the left and right shoulder key points is determined according to the reference shaking frequency and the position information of the left and right shoulder key points.

3. The method according to claim 2, characterized in that Determining the reference shaking frequency according to the position information of the left and right shoulder key points includes: Performing cross-correlation calculation on the position information of the left shoulder key point and the position information of the right shoulder key point to obtain cross-correlation position information; Performing Fourier transform on the cross-correlation position information to obtain a cross-correlation spectrum diagram; The frequency corresponding to the maximum amplitude in the cross-correlation spectrum is the reference shaking frequency.

4. The method according to claim 2, characterized in that Determining the first shaking amplitude of the left and right shoulder key points according to the reference shaking frequency and the position information of the left and right shoulder key points includes: Performing Fourier transform on the position information of the left shoulder key points to obtain a left shoulder spectrum graph; The amplitude whose frequency is equal to the reference shaking frequency in the left shoulder spectrum diagram is the first left shoulder amplitude; Performing Fourier transform on the position information of the right shoulder key points to obtain a right shoulder spectrum graph; The amplitude of the frequency equal to the reference shaking frequency traversed in the right shoulder spectrum diagram is the first right shoulder amplitude; A first shaking amplitude is determined according to the first left shoulder amplitude and the first right shoulder amplitude.

5. The method according to claim 2, characterized in that Determining the second shaking amplitude of the left and right ankle key points according to the position information of the left and right ankle key points includes: Performing Fourier transform on the position information of the left ankle key point to obtain a left ankle frequency spectrum; The amplitude whose frequency is equal to the reference shaking frequency in the left ankle spectrum diagram is a first left ankle amplitude; Performing Fourier transform on the position information of the right ankle key points to obtain a right ankle frequency spectrum; The amplitude whose frequency is equal to the reference shaking frequency in the right ankle frequency spectrum is a first right ankle amplitude; A second shaking amplitude is determined according to the first left ankle amplitude and the first right ankle amplitude.

6. The method according to any one of claims 1 to 5, characterized in that Determining the position information of key points of the human body according to the video stream includes: Extracting multiple frames of character images from the video stream; Extracting the position information of the key points of parts from each frame of the character image according to a preset key point extraction algorithm; The position information of each key point of the body part is normalized to obtain the position information of the key points of the body part.

7. The method according to claim 6, characterized in that The normalization process is performed on the position information of each key point of the body part to obtain the position information of the key points of the body part, including: Acquire multiple sets of left and right shoulder position information, wherein the multiple sets of left and right shoulder position information are position information of key points of the left and right shoulders extracted from each frame of the character image; Obtaining a mean shoulder distance based on the plurality of sets of left and right shoulder position information; Based on the mean value of the shoulder distance, the position information of each key point of the body part is normalized to obtain the position information of the key points of the body.

8. The method according to any one of claims 1 to 5, characterized in that The shake identification information includes shake presence information or shake absence information, and generating the shake identification information according to the first shake amplitude and the second shake amplitude includes: Determining whether the first shaking amplitude is greater than a first preset amplitude, and determining whether the second shaking amplitude is less than a second preset amplitude; When the first shaking amplitude is greater than the first preset amplitude and the second shaking amplitude is less than the second preset amplitude, generating shaking existence information; When the first shaking amplitude is less than or equal to the first preset amplitude, or when the second shaking amplitude is greater than or equal to the second preset amplitude, shaking non-existence information is generated.

9. A shaking recognition device, comprising: An acquisition unit, used for acquiring a video stream; a determining unit, configured to determine position information of key points of a human body according to the video stream, wherein the position information of key points of the human body includes position information of left and right shoulder key points and position information of left and right ankle key points; The determining unit is further configured to determine a first shaking amplitude of the left and right shoulder key points according to the position information of the left and right shoulder key points; The determining unit is further configured to determine a second shaking amplitude of the left and right ankle key points according to the position information of the left and right ankle key points; A generating unit is configured to generate shake identification information according to the first shake amplitude and the second shake amplitude.

10. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein the memory is connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the computer device implements the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 8.