Real-time wonderful picture snapshot method based on artificial intelligence
By collecting athletes' physiological and sports data, combining facial expressions and external attention data, building an intelligent evaluation model, and automatically adjusting the shooting angle, the problem of difficulty in capturing wonderful moments in existing technologies is solved, and the accurate capture and display of wonderful pictures is achieved.
Patent Information
- Application Number
- CN202511171339.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing event filming technology has difficulty tracking athletes' rapid movements in real time, resulting in blurred images or missed moments. It is impossible to accurately judge when a wonderful moment is, ignores the audience's feelings and feedback, and cannot fully evaluate the true excitement of the moment.
By collecting athletes' physiological and motion data, we build an explosive power evaluation model based on a recurrent neural network, collect facial expressions and external attention data in real time, set scoring thresholds, and combine intelligent camera technology to automatically adjust the shooting angle and capture wonderful scenes.
It achieves comprehensive perception and precise assessment of athletes' status, accurately identifies exciting moments, provides a better visual experience, adapts to the characteristics of different sports, and improves recognition effects.
Smart Images

Figure CN120730176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real-time capturing of wonderful pictures, and in particular to a real-time capturing method of wonderful pictures based on artificial intelligence. Background Art
[0002] As an integral part of human history, sports have consistently driven the progress of human civilization. Sports events are not only a competitive arena for various sports, but also a vital window into the physical, intellectual, and spiritual strength of humanity. From the ancient Greek Olympics to today's world-class competitions, sports have always captivated the attention and passion of countless spectators. This is because sports events showcase the boundless potential and unparalleled passion of humanity. On the field, athletes, with astonishing explosive power, intense focus, and a fierce competitive drive, constantly push the boundaries of human physical endurance. These moments of brilliance often elicit roars of cheers from the audience and resonate with countless audiences. It can be said that these precious "highlights" not only fully demonstrate the charm of sports but are also a key factor in the enduring appeal of sporting events.
[0003] However, in today's information age, it is difficult to fully capture these exciting moments by relying solely on existing on-site shooting technology. Traditional event shooting is mostly done by manually controlled cameras, which makes it difficult to track the athletes' rapid movements in real time, resulting in blurred images or missed exciting moments. Relying solely on observing and shooting the athletes' external movements, it is impossible to accurately judge when the exciting moments are. Only focusing on the athletes' own performance ignores the audience's feelings and feedback about the excitement, and it is impossible to fully evaluate the true excitement of a moment. Therefore, there is an urgent need for a real-time capture method of exciting scenes based on artificial intelligence. Through advanced perception technology, intelligent analysis algorithms and intelligent control technology, it can achieve comprehensive perception, accurate evaluation and intelligent capture of the athletes' status, in order to break through the limitations of existing technologies and bring a more exciting visual feast to the audience of sports events. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a method for capturing wonderful images in real time based on artificial intelligence.
[0005] To solve the above technical problems, the present invention provides the following technical solution: a method for capturing wonderful images in real time based on artificial intelligence, comprising: S1. Collect the athlete's physiological data and motion data. Physiological data includes the athlete's height, weight, heart rate, and blood pressure. Motion data includes the athlete's speed, range of motion, movement frequency, and joint angles.
[0006] S2. Extract features from the motion data to obtain motion features, extract features from the physiological data to obtain physiological features, build an explosive power evaluation model based on a recurrent neural network, input motion features and physiological features, and output the athlete's explosive power score.
[0007] S3. Collect the athlete's facial expression data in real time, extract expression features from the facial expression data, and establish a concentration evaluation index. According to the expression features, obtain the athlete's concentration score.
[0008] S4. Collect external attention data in real time. The external attention data includes the voice and body data of the audience. Extract features from the voice and body data, and calculate the athlete's attention score through a pattern recognition algorithm.
[0009] S5. Set an explosiveness score threshold, a concentration score threshold, and an attention score threshold. When the explosiveness score exceeds the explosiveness score threshold, the concentration score exceeds the concentration score threshold, and the attention score exceeds the attention score threshold, it is determined to be a wonderful moment.
[0010] S6. Focus on the athlete's face and automatically adjust the shooting angle to capture the wonderful scene.
[0011] According to a method for capturing wonderful images in real time based on artificial intelligence provided by the present invention, in step S2, the process of extracting features from motion data includes: S21. Use deep learning-based target detection algorithms to locate the contours and coordinates of key parts of athletes.
[0012] S22. Build a human dynamics model based on motion data to evaluate muscle strength.
[0013] S23. Calculate the motion speed and acceleration of key parts based on the streamer method.
[0014] S24. Evaluate the athlete's limb mass distribution based on motion data and calculate the moment of inertia based on the rotational inertia.
[0015] According to a method for real-time capturing of wonderful scenes based on artificial intelligence provided by the present invention, in step S2, the process of constructing an explosive power evaluation model based on a recurrent neural network includes: The motion data and physiological data are cleaned and standardized to obtain standardized data.
[0016] Build a basic recurrent neural network model and set the number of neural network layers and the number of neurons in each layer.
[0017] Collect historical data, extract the athletes' historical motion characteristics and historical physiological characteristics from the historical data, use the historical motion characteristics and historical physiological characteristics as input, and use the explosive power score as output to train the basic recurrent neural network model, retain the model parameters that meet the test accuracy, and obtain the explosive power evaluation model.
[0018] According to a method for capturing wonderful images in real time based on artificial intelligence provided by the present invention, in step S3, the process of extracting expression features from facial expression data includes: The Viola-Jones face detection algorithm is used to detect the face area, and the active shape model is used to set N key feature points in the face area.
[0019] According to the positions of key feature points, the face area is divided into facial feature areas, which include eyes, eyebrows, nose, lips and face area.
[0020] Calculate the geometric quantitative indicators between key feature points, including distance, angle and area.
[0021] Texture analysis is performed on each facial feature area to extract Gabor and LBP texture features.
[0022] Track the motion trajectory of key feature points in consecutive frames and calculate the time domain features of each key feature point as the dynamic features of the face area.
[0023] The expression vector is constructed through geometric quantization indicators, Gabor and LBP texture features and dynamic features, and the expression vector is used as the expression feature.
[0024] According to the present invention, the method for capturing highlights in real time based on artificial intelligence (AI) is described. In step S3, establishing a concentration evaluation index includes: establishing a concentration evaluation model; collecting facial expression data of athletes at different concentration levels and labeling them with corresponding concentration score labels; training the concentration evaluation model, and retaining model parameters that meet a preset accuracy rate; and using the facial expression features as input to output a concentration score.
[0025] According to a method for capturing wonderful images in real time based on artificial intelligence provided by the present invention, in step S4, the process of extracting features from sound data includes: The sound data is subjected to noise reduction processing and the continuous sound data is subjected to audio segmentation to obtain short-frequency sound signals.
[0026] The short-frequency sound signal is subjected to spectrum analysis to obtain short-frequency sound features, which include volume features, frequency features and rhythm features.
[0027] The recursive feature elimination method is used to screen the short-frequency sound features and obtain the relevant sound features.
[0028] According to a method for real-time capturing of wonderful images based on artificial intelligence provided by the present invention, in step S4, the process of extracting features from limb data includes: The limb data is calibrated and filtered, and continuous limb movements are segmented into independent action units, including clapping and waving.
[0029] The spatial features and temporal features of the action unit are extracted. The spatial features include the amplitude, direction and displacement distance of the action, and the temporal features include the duration and frequency of the action.
[0030] The spatial and temporal features are normalized to obtain relevant limb features.
[0031] According to the present invention, the method for capturing exciting scenes in real time based on artificial intelligence includes the following steps: calculating an athlete's attention score using a pattern recognition algorithm in step S4: quantifying relevant sound features and relevant body features, including the sound intervals of the relevant sound features and the movement ranges of the relevant body features; setting an index score for each sound interval and each movement range; and performing a weighted summation of the sound interval index scores and the movement range index scores to obtain an attention score.
[0032] According to a method for capturing wonderful images in real time based on artificial intelligence provided by the present invention, in step S6, the process of focusing on the athlete's face includes: using phase detection autofocus technology to detect high-frequency signals in the picture, determine the focal plane position, and combine face detection and eye detection technologies to achieve focus on the athlete's face.
[0033] According to a method for capturing wonderful images in real time based on artificial intelligence provided by the present invention, in step S6, the process of automatically adjusting the shooting angle includes: The three-axis accelerometer and gyroscope sensor are used to detect the motion state of the camera body.
[0034] The shooting angle is automatically adjusted through the gimbal control algorithm based on the position and movement of the athletes in the picture.
[0035] Combined with image recognition technology, it tracks and locks the position of athletes in the picture to capture wonderful scenes.
[0036] This invention provides an artificial intelligence-based method for capturing highlights in real time. By collecting physiological and motor data from athletes, it constructs a three-dimensional description of their condition. This data not only covers the athlete's basic physical characteristics but also key indicators such as their motor ability and physical condition. Based on this comprehensive sensory data, the athlete's explosive performance can be more accurately assessed, laying a solid foundation for identifying highlights. Secondly, the solution collects facial expression data in real time and uses computer vision technology to extract facial features to accurately assess the athlete's concentration score. Compared to relying solely on manual observation, this concentration analysis based on facial microexpressions is more objective and accurate, providing comprehensive insight into the athlete's psychological state and increasing the accuracy of highlight moment identification. Furthermore, it collects real-time audio and body language data from the audience, analyzing their attention through pattern recognition algorithms, adding a new dimension to highlight moment identification. Based on a comprehensive assessment of the athlete's explosive power, concentration, and audience attention, the solution sets corresponding scoring thresholds. A highlight moment is only identified when all three indicators meet the threshold. This strict logical relationship ensures more reliable and accurate recognition of highlight moments, avoiding false positives or missed negatives. At the same time, it also has the ability to flexibly adjust thresholds, allowing optimization based on the characteristics of different sports to continuously improve recognition results. Finally, the solution also integrates intelligent camera technologies such as phase detection autofocus and gimbal automatic adjustment, which can track and focus athletes in real time to ensure the clarity and continuity of exciting images. Combined with the exciting moments previously identified, it can automatically trigger continuous shooting to capture the best exciting moments and present an even better visual experience for the audience. In short, this AI-based real-time capture solution for exciting images combines advanced perception technology, intelligent analysis algorithms, and intelligent control technology to achieve comprehensive perception, accurate assessment, and intelligent capture of athlete status, bringing a brand new viewing experience to sporting events. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 This is a flow chart of a method for capturing wonderful images in real time based on artificial intelligence provided by an embodiment of the present invention; Figure 2 This is a flow chart of extracting features from motion data in a method for capturing wonderful images in real time based on artificial intelligence, provided by an embodiment of the present invention; Figure 3This is a flowchart of extracting expression features from a method for real-time capturing of wonderful images based on artificial intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0040] The following combination Figure 1-Figure 3 The present invention describes a method for capturing wonderful images in real time based on artificial intelligence.
[0041] Figure 1 It is a structural diagram of a method for real-time capturing of wonderful images based on artificial intelligence provided by an embodiment of the present invention.
[0042] like Figure 1 As shown, an embodiment of the present invention provides a method for capturing wonderful images in real time based on artificial intelligence, including: S1. Collect the athlete's physiological data and motion data. Physiological data includes the athlete's height, weight, heart rate, and blood pressure. Motion data includes the athlete's speed, range of motion, movement frequency, and joint angles.
[0043] In this embodiment, an athlete's physiological data, such as height and weight, helps to better understand their physical characteristics and potential. Kinematic data, such as speed, range of motion, and joint angles, can reflect their movement technique, explosive power, and coordination. Collecting this data provides basic data support for subsequent explosive power assessments and concentration analysis. These physiological indicators and motion data can more accurately portray an athlete's overall condition and provide a basis for identifying exciting moments.
[0044] S2. Extract features from the motion data to obtain motion features, extract features from the physiological data to obtain physiological features, build an explosive power evaluation model based on a recurrent neural network, input motion features and physiological features, and output the athlete's explosive power score.
[0045] Figure 2 This is a flow chart of extracting features from motion data in a method for capturing wonderful images in real time based on artificial intelligence provided by an embodiment of the present invention. Figure 2 As shown in FIG, the process of extracting features from motion data includes: S21. Use deep learning-based target detection algorithms to locate the contours and coordinates of key parts of athletes.
[0046] In this embodiment, a large amount of training data containing athlete movements is collected, and key joints, including the head, shoulders, elbows, wrists, hips, knees, ankles, etc., are annotated. A deep learning model based on a convolutional neural network is constructed, and a target detection algorithm is trained. In new video frames, the trained model is used to quickly locate the coordinates and contours of the athlete's key parts, and Kalman filtering is used to track the movement trajectory of the key parts in consecutive frames to improve detection stability.
[0047] S22. Build a human dynamics model based on motion data to evaluate muscle strength.
[0048] In this example, a multi-body human model is constructed using the athlete's anthropometric data, including height, weight, and limb length. A muscle-skeletal dynamics model is then developed based on anatomical features such as muscle attachment points and joint range of motion. The motion trajectory data of key parts is input into the model, and the joint torques are inversely solved using the Newton-Euler equation. Combined with muscle physiological characteristics, the force output of the major muscle groups is estimated.
[0049] S23. Calculate the motion speed and acceleration of key parts based on the streamer method.
[0050] In this embodiment, the position coordinates of key parts in the video are used to calculate the instantaneous velocity and acceleration using the central difference method. In order to suppress the influence of noise on the differential operation, the coordinate data can be smoothed and filtered first. By comparing the coordinate changes at different time points, the linear velocity and angular velocity of each joint can be obtained. Further differentiation can obtain the linear acceleration and angular acceleration, providing basic data for the next step of analysis.
[0051] S24. Evaluate the athlete's limb mass distribution based on motion data and calculate the moment of inertia based on the rotational inertia.
[0052] In this embodiment, the athlete's body measurement data is used in combination with a statistical model to predict the mass and inertia characteristics of each limb segment. The motion trajectory data of key parts is substituted into the data, and the torque of each joint is calculated according to Newton's second law. The torque is multiplied by the angular acceleration to obtain the inertia moment of each joint.
[0053] The process of building an explosive power evaluation model based on recurrent neural networks includes: The motion data and physiological data are cleaned and standardized to obtain standardized data.
[0054] Build a basic recurrent neural network model and set the number of neural network layers and the number of neurons in each layer.
[0055] Collect historical data, extract the athletes' historical motion characteristics and historical physiological characteristics from the historical data, use the historical motion characteristics and historical physiological characteristics as input, and use the explosive power score as output to train the basic recurrent neural network model, retain the model parameters that meet the test accuracy, and obtain the explosive power evaluation model.
[0056] In this embodiment, deep learning modeling of motion and physiological data can better capture the complex mechanisms of an athlete's explosive power. Recurrent neural networks can establish a nonlinear mapping relationship between motion and physiological characteristics and explosive power, resulting in higher prediction accuracy than traditional statistical models. Furthermore, this model can continuously learn and update, and as data accumulates, the accuracy of explosive power assessment will continue to improve.
[0057] The model can input an athlete's athletic and physiological data in real time and quickly output an explosive power score. Combined with preset explosive power thresholds, it can promptly detect when an athlete's explosive power reaches critical levels, providing coaches with real-time warnings. This helps coaches adjust training plans promptly, helping athletes compete at their best and improve their performance. It also provides coaches with objective feedback on their athletes' training status. Based on the explosive power scores output by the model, coaches can understand their athletes' training progress and potential and make targeted adjustments to their training plans. During competition, this assessment model can also provide coaches with real-time physiological status monitoring, facilitating more precise pre-match preparation and in-game adjustments.
[0058] S3. Collect the athlete's facial expression data in real time, extract expression features from the facial expression data, and establish a concentration evaluation index. According to the expression features, obtain the athlete's concentration score.
[0059] Figure 3 This is a flow chart of extracting expression features from a method for capturing wonderful images in real time based on artificial intelligence provided by an embodiment of the present invention. Figure 3 As shown in Figure 2, the process of extracting expression features from facial expression data includes: The Viola-Jones face detection algorithm is used to detect the face area, and the active shape model is used to set N key feature points in the face area.
[0060] Based on the locations of key feature points, the face is divided into facial feature regions, including the eyes, eyebrows, nose, lips, and face area. These key feature points accurately describe the geometric shape of the face. This region division facilitates the subsequent targeted feature extraction of each facial region.
[0061] Calculate geometric quantitative indicators between key feature points, including distance, angle, and area. These indicators can quantify the geometric shape changes of the face, including eyebrow height, eye opening, mouth opening, etc.
[0062] Texture analysis is performed on each facial feature region to extract Gabor and LBP texture features. Gabor features reflect the direction and frequency of facial muscle texture, while LBP features describe the microscopic pattern of texture. These texture features can capture subtle changes in facial expression.
[0063] Track the motion trajectory of key feature points in consecutive frames and calculate the time domain characteristics of each key feature point as the dynamic characteristics of the face area. Through the face tracking algorithm, the motion trajectory of key feature points in consecutive frames is tracked, and the time domain dynamic characteristics such as position, velocity, acceleration, etc. of each key feature point are calculated to reflect the rate and rhythm characteristics of facial expression changes.
[0064] Through geometric quantization indicators, Gabor and LBP texture features and dynamic features, expression vectors are constructed and used as expression features to fully depict the changing state of facial expressions.
[0065] The process of establishing a focus assessment metric includes: building a focus assessment model. Collecting facial expression data of athletes at different levels of focus and labeling them with corresponding focus scores. Training the focus assessment model and retaining model parameters that meet the preset accuracy. Using facial features as input, the model outputs a focus score.
[0066] In this embodiment, concentration is a key indicator for identifying exciting moments. Only when athletes are highly focused can they unleash their peak explosiveness and technical prowess. The concentration assessment model can monitor athletes' focus in real time, providing a reliable basis for identifying exciting moments. Combining the concentration score with the explosiveness and attention scores can more accurately capture truly exciting moments.
[0067] The focus assessment model also provides coaches with real-time feedback on their athletes' focus. Based on changes in focus scores, coaches can adjust training plans and help athletes develop positive focus habits. During competition, coaches can also use focus data to select appropriate player deployments and improve overall tactical execution.
[0068] When the system detects athletes entering a state of intense concentration, it automatically captures the highlights, providing viewers with an enhanced viewing experience. These captivating moments can not only be viewed live but also used in event broadcasts, significantly enhancing the appeal of the broadcast content. High-quality broadcast content attracts more sponsors, thereby increasing the commercial value of the event.
[0069] S4. Collect external attention data in real time. The external attention data includes the voice and body data of the audience. Extract features from the voice and body data, and calculate the athlete's attention score through a pattern recognition algorithm.
[0070] A microphone array collects real-time audience sound data, including cheers and applause. Infrared sensors also collect body language data, such as clapping and waving. This data reflects the audience's current level of attention and excitement.
[0071] The process of feature extraction of sound data includes: The sound data is subjected to noise reduction processing and the continuous sound data is subjected to audio segmentation to obtain short-frequency sound signals.
[0072] The short-frequency sound signal is subjected to spectrum analysis to obtain short-frequency sound features, which include volume features, frequency features and rhythm features.
[0073] The recursive feature elimination method is used to screen the short-frequency sound features and obtain the relevant sound features.
[0074] The process of feature extraction for limb data includes: The limb data is calibrated and filtered, and continuous limb movements are segmented into independent action units, including clapping and waving.
[0075] The spatial features and temporal features of the action unit are extracted. The spatial features include the amplitude, direction and displacement distance of the action, and the temporal features include the duration and frequency of the action.
[0076] The spatial and temporal features are normalized to obtain relevant limb features.
[0077] The process of calculating an athlete's attention score using a pattern recognition algorithm involves quantifying relevant sound features and body features, including the sound intervals and movement ranges of the relevant sound features. A score is assigned to each sound interval and movement range. The weighted sum of the sound interval and movement range scores is then used to calculate the attention score.
[0078] In this embodiment, by collecting on-site data in real time, the athlete's attention score can be quickly calculated, providing timely data support for the real-time identification of exciting moments. The combination of sound and body data can more comprehensively reflect the audience's attention status. The fusion of multi-source data can improve the accuracy and robustness of attention assessment. Compared with subjective judgment, attention assessment based on data analysis is more objective and fair, which helps to eliminate the interference of human factors and improve the accuracy of the entire system. At the same time, this attention assessment method based on audience behavior data is not only applicable to sports events, but can also be extended to other scenarios such as concerts and speeches.
[0079] S5. Set an explosiveness score threshold, a concentration score threshold, and an attention score threshold. When the explosiveness score exceeds the explosiveness score threshold, the concentration score exceeds the concentration score threshold, and the attention score exceeds the attention score threshold, it is determined to be a wonderful moment.
[0080] In this embodiment, the explosive power score threshold represents the minimum required level of explosive power required for an athlete's movements. Moves below this threshold are not considered exciting. The concentration score threshold represents the minimum required level of concentration required for an athlete. Moves below this threshold may lack the necessary concentration. The attention score threshold represents the minimum level of audience attention to the move. A move below this threshold indicates that the audience may not be paying attention. By collecting a large amount of actual event video data covering various sports, each move in the video is manually annotated and given an explosive power score, a concentration score, and an attention score. A labeled dataset is constructed to provide a basis for subsequent threshold determination. Statistical analysis is performed on the moves in the annotated dataset, and a probability distribution histogram of each score indicator is plotted. Based on the distribution of the histogram, appropriate quantiles are selected as thresholds for each score. For the real-time collected data, the explosive power score, concentration score, and attention score of each athlete are calculated. When an athlete's three scores simultaneously exceed their respective preset thresholds, the move is determined to be an exciting moment.
[0081] S6. Focus on the athlete's face and automatically adjust the shooting angle to capture the wonderful scene.
[0082] The process of focusing on the athlete's face includes: using phase detection autofocus technology to detect high-frequency signals in the picture, determine the focal plane position, and combine face detection and eye detection technology to achieve focus on the athlete's face.
[0083] In this embodiment, facial and eye detection technologies are combined. First, a deep learning model, such as MTCNN, is used to detect the athlete's face in the frame. The eye area is then located and used as the focus point, ensuring the athlete's face remains sharp. This face- and eye-based focusing method can maintain good focus even in scenes with fast-moving athletes.
[0084] The process of automatically adjusting the shooting angle includes: The camera uses a three-axis accelerometer and gyroscope to detect the camera's motion, including translation and rotation. Based on the athlete's position and movement in the frame, computer vision technology is used to track and predict the athlete's trajectory. A gimbal control algorithm adjusts the gimbal's translation and rotation in real time, ensuring the camera maintains focus on the athlete. This automatic angle adjustment, based on multi-sensor fusion, can handle a variety of complex sports scenarios and ensure accurate capture.
[0085] Integrating image recognition technology, it tracks and locks onto the athlete's position in the frame, capturing exciting moments. Using target detection and tracking algorithms, it continuously locks onto the athlete's position in the frame. When it detects a moment that exceeds the explosiveness, concentration, and attention score thresholds, it automatically triggers continuous shooting to capture the moment. A high-frame-rate continuous shooting mode ensures optimal capture. Captured images are intelligently sorted and filtered based on explosiveness, concentration, and attention scores.
[0086] Example 1: Athlete A, gender: male, age: 25, height: 180cm, weight: 75kg.
[0087] The heart rate during high-intensity exercise was 150 bpm, the blood pressure was 120 / 80 mmHg before exercise, and 130 / 85 mmHg after exercise.
[0088] The maximum movement speed is 6.2m / s. During the acceleration phase, the knee bending angle is 90°. When the speed reaches its peak, the movement frequency is 2.2Hz. The joint angle of the elbow when exerting force is 120°, and that of the shoulder is 150°.
[0089] Key parts positions: knee: coordinates (x=50,y=120), elbow: coordinates (x=70,y=150), shoulder: coordinates (x=60,y=160).
[0090] A human body dynamics model based on motion data was constructed to evaluate leg muscle strength: 300N, upper limb muscle strength: 200N, and average strength: (300N+200N) / 2=250N.
[0091] The acceleration estimated based on the streamer method is 3.5m / s².
[0092] Upper body mass: 30 kg (40% of body weight), lower body mass: 45 kg (60% of body weight). The distance from the center of gravity to the ground is 0.9 m. The calculated moment of inertia is 36.45.
[0093] Collect historical data, the table is as follows: Training Number Movement speed (m / s) Operating frequency (Hz) Knee joint angle (°) Heart rate (bpm) Blood pressure (mmHg) Explosive score 1 6.0 2.0 90 150 120 / 80 0.85 2 5.5 1.8 95 145 115 / 75 0.80 3 6.2 2.2 85 160 125 / 85 0.90 4 5.9 2.1 88 155 120 / 80 0.88 5 6.1 2.3 87 158 121 / 79 0.91 … … … … … … … 30 6.5 2.4 82 152 122 / 81 0.95 Extract motion features and physiological features from historical data and integrate them into input feature vectors: [0.75,0.60,0.50,0.40,0.48,0.55], explosiveness score = 0.95.
[0094] Build the structure of the basic recurrent neural network model, input layer: 6 feature nodes.
[0095] Hidden layer: 1 LSTM layer with 64 units and Tanh activation function.
[0096] Output layer: 1 node, activation function is Sigmoid (output concentration score).
[0097] Set the model hyperparameters: learning rate: 0.001, batch size: 32, training epochs: 100, loss function: mean squared error (MSE), and optimizer: Adam.
[0098] Use historical data to train the model and retain the model parameters that reach the preset loss function value.
[0099] The facial region of athlete A is detected using the Viola-Jones algorithm.
[0100] Face area coordinates: upper left corner (x=50, y=80), lower right corner (x=100, y=140).
[0101] Set N key feature points. Some key feature points and their coordinates are as follows: Left corner of the eye: (x=55,y=90), right corner of the eye: (x=65,y=90), left eyebrow: (x=52,y=85), right eyebrow: (x=68,y=85), tip of the nose: (x=60,y=110), left corner of the mouth: (x=55,y=130), right corner of the mouth: (x=65,y=130).
[0102] According to the definition of key feature points, the face area is divided into the following feature areas: Eye area: Left ((55,90), (60,100)), Right ((60,90), (70,100)).
[0103] Eyebrow area: left ((52,85), (55,90)), right ((65,85), (58,90)).
[0104] Nose area: (60,100)-(60,120).
[0105] Lip area: Left mouth ((55,130), (60,135)), right mouth ((65,130), (60,135)).
[0106] Face area: The main area contains all the features.
[0107] Interocular distance: .
[0108] Distance between eyebrows: .
[0109] Angle between left eye corner and nose tip: .
[0110] The area of the left eye region (rectangle): (70-55) × (100-90) = 15 × 10 = 150 cm 2 .
[0111] Gabor features = [0.3, 0.5, 0.2, 0.7].
[0112] LBP features = [1,0,1,1,0].
[0113] Motion trajectory data (in 5 consecutive frames) Frame 1: nose tip (60,110), left corner of mouth (55,130).
[0114] Frame 2: nose tip (60,112), left corner of mouth (56,132).
[0115] Frame 3: nose tip (60,113), left corner of mouth (57,134).
[0116] Frame 4: nose tip (60,114), left corner of mouth (58,135).
[0117] Frame 5: nose tip (60,115), left corner of mouth (59,136).
[0118] Dynamic features of the nose tip: Movement path: from (60,110) to (60,115).
[0119] Time interval: 1 frame = 0.04 seconds.
[0120] Average speed (calculated by frame-to-frame variation): .
[0121] Dynamic features of the left corner of the mouth: Movement path: from (55,130) to (59,136).
[0122] Average speed: .
[0123] Nose tip feature vector: dynamic feature = [25, 0.04] dynamic feature = [25, 0.04].
[0124] Left mouth corner feature vector: dynamic feature = [50, 0.04] dynamic feature = [50, 0.04].
[0125] All extracted features are integrated to form the final expression feature vector. The expression vector includes geometric quantization indicators, Gabor and LBP texture features, and dynamic features. Assume that the final constructed expression feature vector is as follows: Geometric quantitative indicators (standardization of distance, angle, and area): Geometric features = [normalized eye distance = 2010 = 0.5, normalized eyebrow distance = 2016 = 0.8, angle = 9075.96 ≈ 0.844, normalized left eye area = 300150 = 0.5].
[0126] Gabor features: [0.3, 0.5, 0.2, 0.7] [0.3, 0.5, 0.2, 0.7].
[0127] LBP features: [1,0,1,1,0][1,0,1,1,0].
[0128] Dynamic features (considering the tip of the nose and the left corner of the mouth): Dynamic feature = [nose tip speed = 25, left mouth corner speed = 50] Dynamic feature = [nose tip speed = 25, left mouth corner speed = 50].
[0129] The expression vector output combining all features: Expression feature vector = [0.5, 0.8, 0.844, 0.5, 0.3, 0.5, 0.2, 0.7, 1, 0, 1, 1, 0, 25, 50].
[0130] A single-layer feedforward neural network is used as the focus assessment model. The input layer contains the facial expression feature vector, and the output layer is the focus score (between 0 and 1).
[0131] Set the model parameters: Input nodes: 15 (from the expression feature vector), Hidden nodes: 5, Output node: 1 (focus score), Activation functions: ReLU (hidden layer), Sigmoid (output layer), Loss function: Mean Squared Error.
[0132] Collect facial expression data of athletes at different levels of concentration and their corresponding concentration score labels. The sample data is as follows: [0.5,0.8,0.844,0.5,0.3,0.5,0.2,0.7,1,0,1,0,0,25,50], 0.85.
[0133] [0.6,0.6,0.7,0.6,0.2,0.4,0.3,0.6,1,0,0,0,0,20,40], 0.70.
[0134] [0.5,0.7,0.8,0.4,0.4,0.5,0.5,0.5,0,1,0,1,1,30,55], 0.60.
[0135] [0.4,0.5,0.6,0.3,0.1,0.3,0.1,0.2,0,0,1,0,1,15,30], 0.45.
[0136] [0.7,0.9,0.9,0.8,0.5,0.6,0.4,0.8,1,1,1,1,0,35,60], 0.95.
[0137] [0.35,0.4,0.3,0.2,0.2,0.4,0.4,0.5,0,0,0,0,0,5,10], 0.20.
[0138] Use the collected data to train the focus assessment model. Use batch gradient descent and follow the following steps: Data preprocessing: Standardized expression features: Standardize the input data so that its mean is 0 and its variance is 1.
[0139] Normalize the label (concentration score) to between 0 and 1.
[0140] Training steps: The training set and test set are randomly divided into 70% training set and 30% test set.
[0141] Train the model, using cross-validation to save time and resources.
[0142] Convergence conditions: The training is stopped when the loss function of the model does not change by more than a threshold (such as 0.001) for 10 consecutive iterations.
[0143] Keep the model parameters that meet the preset accuracy (such as R²>0.85).
[0144] After completing model training and evaluating model performance using the test set, assume the following results are obtained: Mean absolute error (MAE): 0.05.
[0145] Root mean square error (RMSE): 0.07.
[0146] Coefficient of determination (R²): 0.88 (indicating that the model has a good ability to explain concentration).
[0147] Input feature vector: [0.5, 0.8, 0.844, 0.5, 0.3, 0.5, 0.2, 0.7, 1, 0, 1, 1, 0, 25, 50], The input feature vector is normalized and converted into a distribution with mean 0 and variance 1.
[0148] Use the trained neural network model to forward propagate and calculate the output: The standardized features are input to the hidden layer and the ReLU activation function is applied.
[0149] The output of the hidden layer is passed to the output layer, and the final concentration score is obtained through the Sigmoid activation function.
[0150] The output concentration score obtained by the model is: concentration score = 0.80.
[0151] Collect external attention data in real time and extract features from sound data: including noise reduction, audio segmentation, spectrum analysis and feature screening.
[0152] Perform noise reduction on the original sound data to eliminate background noise and irrelevant audio interference. The waveform of the original audio data is as follows: Length: 10 seconds.
[0153] Sampling frequency: 44.1kHz.
[0154] The noise-reduced audio signal is divided into multiple short frequency bands to obtain a short-frequency sound signal.
[0155] The denoised audio signal is divided into five short frequency bands. The data of each short frequency band is as follows: Short band number Audio sample segment (seconds) 1 [0.1,0.2,0.3,0.4] 2 [0.5,0.2,0.1,0.3] 3 [0.3,0.4,0.5,0.6] 4 [0.6,0.7,0.4,0.3] 5 [0.7,0.8,0.6,0.5]
[0156] Fast Fourier transform is used to perform spectrum analysis on the audio data of each short frequency band to extract volume features, frequency features and rhythm features.
[0157] The analysis results for short frequency bands are as follows: Short band number Volume characteristics (dB) Main frequency characteristics (Hz) Prosodic characteristics (Hz) 1 75 440 1.5 2 80 450 1.7 3 78 460 1.6 4 85 430 1.8 5 82 410 1.5
[0158] Recursive feature elimination is used to filter out the most relevant sound features for attention rating.
[0159] The relevant sound features of the final output are: Related sound features = [volume features, main frequency features] = [75,440].
[0160] The raw limb data is calibrated and filtered to remove noise and unnecessary interference.
[0161] Raw limb data: Sampling frequency: 100Hz.
[0162] Sampling time: 10 seconds.
[0163] After calibration and filtering, the smoothed limb data is obtained as follows: Time (seconds) X-axis acceleration (g) Y-axis acceleration (g) Z-axis acceleration (g) Angular velocity (rad / s) 0 0.02 0.05 -0.02 0.1 0.01 0.03 0.06 -0.01 0.12 0.02 0.01 0.07 0 0.10 … … … … … 9.99 0.02 0.03 -0.01 0.1 Continuous body movements are segmented into independent action units, and the start and end of actions are detected using motion-specific thresholds.
[0164] Two separate action units are identified: Applause action: Start time: 2 seconds.
[0165] End time: 3 seconds.
[0166] Waving action: Start time: 5 seconds.
[0167] End time: 7 seconds.
[0168] Calculate the maximum amplitude of a limb's movement. Calculate the primary direction of the movement. Calculate the displacement from the start to the end of the movement. Calculate the duration of the movement. Calculate the frequency of the movement.
[0169] For clapping and waving actions, the feature extraction results are as follows: Applause (2 to 3 seconds) Amplitude: 0.3g.
[0170] Direction: Y axis.
[0171] Displacement distance: 0.4m.
[0172] Duration: 1 second.
[0173] Frequency: 1 time / second.
[0174] Waving motion (5 to 7 seconds).
[0175] Amplitude: 0.5g.
[0176] Direction: X axis.
[0177] Displacement distance: 0.6m.
[0178] Duration: 2 seconds.
[0179] Frequency: 1 time / 2 seconds.
[0180] The Z-score standardization method was used to calculate the mean and standard deviation and standardize the features.
[0181] feature Applause (standardized feature) Waving (standardized feature) Amplitude (g) -1.0 1.0 Direction (X / Y) 0 1 Displacement distance (m) -0.5 0.5 Duration (seconds) -1.0 1.0 Frequency (times / second) 1.0 -0.5 After normalization, the final output of relevant limb features is summarized as follows: Action Type Amplitude (normalized) Direction (normalized) Displacement distance (normalized) Duration (normalized) Frequency (normalized) applaud -1.0 0 -0.5 -1.0 1.0 waving 1.0 1 0.5 1.0 -0.5 The relevant sound features are quantified and divided into specific sound intervals.
[0182] The sound characteristics are divided into the following intervals: Range 1: 0-50dB, index score 2.
[0183] Range 2: 51-75dB, index score 4.
[0184] Range 3: 76-90dB, index score 6.
[0185] Range 4: 91-110dB, index score 8.
[0186] Set the motion characteristics to the following range: Range 1 (0-0.2m), indicator score 2.
[0187] Range 2 (0.21-0.5m), indicator score 4.
[0188] Range 3 (0.51-0.8m), indicator score 6.
[0189] Range 4 (0.81-1.0m), indicator score 8.
[0190] The weights of voice and body features are set as follows: Sound feature weight: 0.6.
[0191] Limb feature weight: 0.4.
[0192] Attention score = (voice feature score × voice feature weight) + (body feature score × body feature weight).
[0193] Attention score = (6×0.6)+(4×0.4)=3.6+1.6=5.2.
[0194] Set the explosiveness score thresholds to 0.8, the concentration score threshold to 0.7, and the attention score threshold to 5.0. Athlete A's current explosiveness score is 0.95, the concentration score is 0.80, and the attention score is 5.2, so this is considered a highlight.
[0195] In summary, this embodiment provides an AI-based method for capturing exciting moments in real time. By collecting an athlete's physiological and motion data, including height, weight, heart rate, speed, and range of motion, a comprehensive perception of the athlete's physical and mental state is constructed. This data not only reflects the athlete's basic physical characteristics but also describes their movement ability and physical condition, laying the foundation for subsequent explosive power assessment. Comprehensive perception data facilitates a more accurate assessment of the athlete's overall condition and provides a reliable basis for identifying exciting moments.
[0196] Using a recurrent neural network model, combined with athletic and physiological data, we achieve intelligent assessment of an athlete's explosive power. This comprehensive analysis, based on kinesiology and physiological characteristics, accurately captures this key metric, providing an important basis for identifying spectacular moments. Compared to relying solely on subjective judgment, this data-based approach to explosive power assessment is more objective and accurate, eliminating the influence of human factors.
[0197] By collecting real-time facial expression data from athletes and using computer vision technology to extract facial features, we can assess their concentration scores. Concentration is a crucial psychological factor influencing an athlete's performance, and this concentration assessment based on facial micro-expressions is more accurate and objective than simple human observation, providing data support for optimal athlete deployment.
[0198] Real-time audio and body language data from live audiences is collected, and through feature extraction and pattern recognition, athlete attention scores are assessed. Audience feedback and behavior directly reflect their level of interest in the athletes' performance. This multimodal data-based attention analysis is more comprehensive and objective. This not only facilitates the identification of exciting moments but also provides higher-quality content for event broadcasts, enhancing the audience's viewing experience.
[0199] Thresholds for explosiveness, concentration, and attention are set, and only when all three scores meet the threshold is a highlight moment identified. This strict logical relationship ensures more reliable recognition of highlight moments, avoiding false positives or missed negatives. Furthermore, this threshold setting is flexible and can be optimized and adjusted based on the characteristics of different sports, continuously improving recognition accuracy.
[0200] Utilizing phase detection autofocus and gimbal automatic adjustment technologies, the system tracks and focuses on athletes in real time, ensuring clarity and continuity in the action. This intelligent camera technology, based on multi-sensor fusion, can handle complex scenarios such as rapid athlete movement, ensuring high-quality footage. Furthermore, based on previously identified highlights, the system automatically triggers continuous shooting to capture the best moments.
[0201] This AI-based solution for capturing highlights in real time intelligently and automatically automates the entire process, from sensing athlete status, assessing key metrics, identifying highlights, to ultimately capturing them. This significantly improves the accuracy and reliability of highlight recognition and provides viewers with a superior viewing experience. This innovative practice, integrating cutting-edge technology with the sports industry, is poised to become a key driver of smart sports and inject new momentum into the sporting industry.
[0202] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for capturing wonderful images in real time based on artificial intelligence, characterized in that: include: S1. Collecting the athlete's physiological data and motion data, wherein the physiological data includes the athlete's height, weight, heart rate, and blood pressure; and the motion data includes the athlete's speed, movement amplitude, movement frequency, and joint angle; S2. Extracting features from the motion data to obtain motion features, extracting features from the physiological data to obtain physiological features, constructing an explosive power evaluation model based on a recurrent neural network, inputting the motion features and the physiological features, and outputting an athlete's explosive power score; S3. Collecting the athlete's facial expression data in real time, extracting expression features from the facial expression data, and establishing a concentration evaluation index, and obtaining the athlete's concentration score based on the expression features; S4. Collecting external attention data in real time, the external attention data including voice data and body data of the audience, extracting features from the voice data and body data, and calculating the athlete's attention score using a pattern recognition algorithm; S5. Setting an explosiveness score threshold, a concentration score threshold, and an attention score threshold. When the explosiveness score exceeds the explosiveness score threshold, the concentration score exceeds the concentration score threshold, and the attention score exceeds the attention score threshold, determining it as a wonderful moment; S6. Focus on the athlete's face and automatically adjust the shooting angle to capture the wonderful scene.
2. The method for capturing wonderful images in real time based on artificial intelligence according to claim 1, characterized in that: In step S2, the process of extracting features from the motion data includes: S21. Use deep learning-based target detection algorithms to locate the contours and coordinates of key parts of athletes; S22. constructing a human body dynamics model based on the motion data to evaluate muscle strength; S23, calculating the movement speed and acceleration of the key part based on the streamer method; S24. Evaluate the athlete's limb mass distribution based on the motion data, and calculate the moment of inertia in combination with the rotational inertia.
3. The method for capturing wonderful images in real time based on artificial intelligence according to claim 1, characterized in that: In step S2, the process of constructing an explosive power evaluation model based on a recurrent neural network includes: Cleaning and standardizing the motion data and the physiological data to obtain standardized data; Build a basic recurrent neural network model, set the number of neural network layers and the number of neurons in each layer; Collect historical data, extract the athlete's historical movement characteristics and historical physiological characteristics from the historical data, use the historical movement characteristics and historical physiological characteristics as input, and use the explosive power score as output to train the basic recurrent neural network model, retain model parameters that meet the test accuracy, and obtain an explosive power evaluation model.
4. The method for capturing wonderful images in real time based on artificial intelligence according to claim 1, characterized in that: In step S3, the process of extracting expression features from the facial expression data includes: Using the Viola-Jones face detection algorithm, the face area is detected, and an active shape model is used to set N key feature points in the face area; Dividing the face region into facial feature regions according to the positions of the key feature points, the facial feature regions including eyes, eyebrows, nose, lips and face regions; Calculating geometric quantitative indicators between the key feature points, wherein the geometric quantitative indicators include distance, angle and area; Performing texture analysis on each of the facial feature regions to extract Gabor and LBP texture features; Tracking the motion trajectory of the key feature points in consecutive frames, and calculating the time domain features of each key feature point as the dynamic features of the face area; An expression vector is constructed by using the geometric quantization index, the Gabor and LBP texture features and the dynamic features, and the expression vector is used as the expression feature.
5. The method for capturing wonderful images in real time based on artificial intelligence according to claim 1, characterized in that: In step S3, the process of establishing a concentration evaluation index includes: establishing a concentration evaluation model; collecting facial expression data of athletes at different concentration levels and marking corresponding concentration score labels; training the concentration evaluation model and retaining model parameters that meet the preset accuracy rate; using the expression features as input and outputting a concentration score.
6. The method for capturing wonderful images in real time based on artificial intelligence according to claim 1, characterized in that: In step S4, the process of extracting features from the sound data includes: Performing noise reduction processing on the sound data and performing audio segmentation on the continuous sound data to obtain short-frequency sound signals; Performing spectrum analysis on the short-frequency sound signal to obtain short-frequency sound features, wherein the short-frequency sound features include volume features, frequency features, and rhythm features; The short-frequency sound features are screened using a recursive feature elimination method to obtain relevant sound features.
7. The method for capturing wonderful images in real time based on artificial intelligence according to claim 6, characterized in that: In step S4, the process of extracting features from the limb data includes: Calibrate and filter the limb data, and segment the continuous limb movements into independent movement units, wherein the movement units include clapping movements and waving movements; Extracting spatial features and temporal features from the action unit, wherein the spatial features include the amplitude, direction, and displacement distance of the action, and the temporal features include the duration and frequency of the action; The spatial features and the temporal features are normalized to obtain relevant limb features.
8. The method for capturing wonderful images in real time based on artificial intelligence according to claim 7, characterized in that: In step S4, the process of calculating the athlete's attention score through the pattern recognition algorithm includes: quantifying the relevant sound features and the relevant limb features, including the sound interval of the relevant sound features and the movement range of the relevant limb features; setting an index score for each of the sound intervals and each of the movement ranges; and performing weighted summation of the sound interval index score and the movement range index score to obtain an attention score.
9. The method for capturing wonderful images in real time based on artificial intelligence according to claim 1, characterized in that: In step S6, the process of focusing on the athlete's face includes: using phase detection autofocus technology to detect high-frequency signals in the picture, determine the focal plane position, and combine face detection and eye detection technology to achieve athlete face focus.
10. The method for capturing wonderful images in real time based on artificial intelligence according to claim 1, characterized in that: In step S6, the process of automatically adjusting the shooting angle includes: Use a three-axis accelerometer and gyroscope sensor to detect the motion state of the camera body; Automatically adjust the shooting angle through the pan / tilt control algorithm based on the position and movement of the athletes in the picture; Combined with image recognition technology, it tracks and locks the position of athletes in the picture to capture wonderful scenes.
Citation Information
Patent Citations
Snapshot method and system applied to video shooting, camera and storage medium
CN109922266A
Wonderful picture real-time automatic snapshot method based on motion sensor
CN111541843A
Student activity wonderful instant shooting and analysis method based on AI algorithm
CN117278801A
Live broadcast switching method and system based on AI face recognition
CN120186388A
Sports event collection generation method and system based on artificial intelligence
CN120264103A