Multi-target sports classroom exercise density evaluation method and equipment under single view angle and medium

Through the single-view video acquisition and deep learning model combined with electronic fence, grid indexing and two-layer goal tracking mechanism, the accuracy and cost problems of exercise density evaluation in multi-objective physical education classrooms are solved, and efficient and accurate exercise density evaluation and teaching quality analysis are achieved, meeting national curriculum standards.

CN120544097APending Publication Date: 2025-08-26NANJING RONGSHI MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510633911.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

It is difficult for the existing technology to accurately evaluate sports density in physical education classrooms of 50-100 people from a single perspective. The traditional methods have strong subjectivity, low accuracy and high cost. The existing equipment and systems are insufficient in real-time and accuracy in multi-target crowd-intensive scenarios, and cannot meet the requirements of the national compulsory course sports and health curriculum standards.

Method used

A single-view video acquisition and deep learning model is used for human detection, combined with electronic fences, grid indexes and two-layer target tracking mechanisms, real-time multi-objective motion density evaluation is achieved through parallel acceleration of GPUs, a five-dimensional evaluation index system is built, and teaching quality evaluation is carried out in combination with speech recognition technology.

Benefits of technology

It realizes efficient and accurate assessment of multi-objective motion density from a single perspective, significantly reduces hardware costs and site requirements, provides objective quantitative basis, supports real-time teaching adjustments, and improves teaching quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544097A_ABST
    Figure CN120544097A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target sports classroom exercise density evaluation method and device under a single view angle and a medium. The method comprises the following steps: S110, a video acquisition and preprocessing stage; s120, constructing a data management framework; s210, constructing an electronic fence; s220, performing human body detection and ROI extraction by using a deep learning model; s230, designing a gridding index mechanism; s240, adopting a video stabilization method based on feature points; s310, designing a multi-target state updating mechanism; s320, establishing an identity mapping table and a target identity feature set; s330, constructing a target tracking model; according to the single-view-angle analysis technology, multi-target motion density evaluation based on a single-view-angle video is achieved, multi-angle camera arrangement is not needed, hardware cost and site requirements are remarkably reduced, the system can conduct off-line or real-time analysis through videos collected by portable devices such as a mobile phone and a common camera, and the analysis efficiency is improved. And the practicability and the universality of the system are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a method, device and medium for evaluating motion density of multiple targets in a physical education classroom under a single viewing angle. Background Art

[0002] The national compulsory physical education and health curriculum standards require teachers to scientifically set student exercise loads during practical physical education classes and evaluate teaching effectiveness using three levels of indicators: individual exercise density, average individual exercise density, and group exercise density. However, current physical education classes face the following challenges: First, a physical education class typically involves 50-100 students, making it difficult for teachers to accurately assess each student's exercise status through visual observation. Second, traditional manual observation methods are highly subjective and inaccurate, failing to provide a quantitative basis for teaching evaluation. Finally, the lack of a real-time feedback mechanism makes it difficult for teachers to adjust teaching content based on students' actual exercise status, making it difficult to meet the requirements of the curriculum standards.

[0003] Although some auxiliary tools, such as fitness trackers and heart rate monitors, can record individual motion data, these devices have the following limitations: First, it is difficult to collect and analyze data from 50-100 people simultaneously; second, wearing the devices can interfere with students' normal movements; and third, the purchase and maintenance costs are high. Furthermore, existing video analysis systems primarily focus on accurately tracking small-scale targets, and struggle to ensure real-time performance and accuracy in scenarios with multiple targets and dense crowds. Furthermore, existing video analysis systems typically require multiple cameras simultaneously capturing images from different angles to obtain accurate motion assessment results. This not only increases equipment costs but also imposes strict requirements on site layout, making it difficult to promote and apply in general teaching environments.

[0004] Therefore, to meet the requirements of the national compulsory physical education and health curriculum standards, a single-viewpoint movement density assessment system capable of simultaneously monitoring classrooms of 50-100 students is urgently needed. This system should be able to provide objective and accurate quantitative evidence for physical education instruction using video data collected by a single camera, combined with a comprehensive five-indicator evaluation system. This would help teachers understand classroom movement status in real time and adjust teaching content scientifically. Compared to traditional multi-viewpoint systems, a single-view solution significantly reduces hardware costs and site requirements, making the system more easily applicable in mainstream schools. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, device and medium for evaluating the exercise density of multiple targets in a physical education classroom under a single perspective, so as to solve the problems raised in the above-mentioned background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for evaluating the exercise density of a multi-objective physical education class under a single viewing angle, comprising the following steps:

[0007] S110. Video acquisition and preprocessing stage: Video data acquisition: Initial video data for student training is collected from a single perspective. GPU parallel acceleration is used to accelerate real-time acquisition and convert continuous video streams into image sequences with time stamps.

[0008] S120. Data Management Architecture Construction: Build a multi-layer data management architecture to manage data uniformly through video frame processing, audio stream processing, and gesture data processing;

[0009] S210. Geo-Fence Construction: Build a polygon-based geo-fence system. Adjust the area by setting fences and dragging thresholds, and use grid indexing for spatial management.

[0010] S220. Use deep learning models for human detection and ROI extraction, and achieve efficient model inference through ROI region stitching;

[0011] S230. Design a grid indexing mechanism to achieve uniform division of scene space and target management through grid division;

[0012] S240. Using a feature point-based video stabilization method to estimate and compensate for camera shake by tracking the movement of key feature points;

[0013] S310. Design a multi-target status update mechanism to achieve continuous tracking of multiple targets through position information and similarity matching;

[0014] S320. Establish an identity mapping table and a target identity feature set, and achieve continuous maintenance of the target identity through multi-feature association measurement;

[0015] S330. Build a target tracking model to handle occlusion issues in dense scenes through target visibility evaluation and multi-dimensional feature similarity mechanism;

[0016] S410. Calculate the three indicators of individual movement density, average individual movement density and group movement density, and generate a density distribution heat map by gridding the area;

[0017] S510. Collect teacher classroom voice data and perform real-time speech-to-text processing through dual-source collection and streaming voice processing architecture;

[0018] S520. Conduct structured analysis of the transcribed language content to identify teaching links, extract key instructions, and assess language quality;

[0019] S530. Construct a classroom evaluation model based on five core indicators, and analyze the effectiveness of teaching strategies and their correlation with student engagement by combining phonetic transcription;

[0020] S540. Integrate all analysis results to generate a comprehensive classroom evaluation report and targeted improvement suggestions, providing teachers with specific measures to improve teaching quality.

[0021] Preferably, in the target detection and ROI extraction steps, a deep learning model is used to extract image features from a video frame sequence to detect the position and bounding box of the moving target.

[0022] Preferably, in the motion density calculation step, the motion state of each target is determined, and the individual motion density is calculated based on the change in joint angle.

[0023] Preferably, in the multi-target tracking step, the global layer tracker is responsible for maintaining the position and timestamp information of all targets;

[0024] The individual-level tracker is responsible for recording the first appearance time, movement period, and movement status of each target;

[0025] Design identity management mechanisms to ensure target ID continuity.

[0026] Preferably, in the evaluation index system step: constructing a five-dimensional evaluation index including individual movement density, average individual movement density, group movement density, thermal distribution index and regional load balance.

[0027] Preferably, in the teacher language intelligent analysis step, a dual-source acquisition mechanism is used to collect the teacher's classroom voice, including a video file audio channel and real-time microphone input;

[0028] Build a streaming speech processing architecture to analyze and transcribe teacher language content and generate classroom transcripts with timestamps;

[0029] Identify teaching links based on natural language processing technology and analyze the rationality of link duration;

[0030] Construct a teacher language classification model, which divides teacher language into five categories: instructional, feedback, motivational, organizational, and evaluative;

[0031] Analyze the rationality of teachers' language structure, including the distribution ratio and quality of various languages;

[0032] Evaluate the conformity of teaching content based on curriculum standards and analyze the degree to which teachers pay attention to individual differences.

[0033] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements the multi-target physical education classroom motion density assessment method under a single perspective as described in any of the above technical solutions.

[0034] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for evaluating multi-target physical education classroom exercise density under a single perspective as described in any of the above technical solutions is implemented.

[0035] Compared with the existing technology, the technical effects and advantages of the present invention are as follows: the method, device and medium for evaluating the exercise density of multiple targets in physical education classroom under a single viewing angle are

[0036] 1. Single-viewpoint analysis technology: This technology enables multi-target motion density assessment based on single-viewpoint video, eliminating the need for multi-angle camera deployment. This significantly reduces hardware costs and site requirements, enabling the system to perform offline or real-time analysis of videos captured by portable devices such as mobile phones and ordinary cameras, greatly improving the system's practicality and accessibility.

[0037] 2. Efficient multi-target processing architecture:

[0038] Designed a video processing architecture based on GPU parallel acceleration to achieve real-time acquisition and processing, supporting multi-target detection in large crowds

[0039] Constructed an electronic fence system based on polygonal areas, and achieved rapid space management and target positioning through grid indexing

[0040] Adopted ROI data processing technology to optimize system processing efficiency and computing resource utilization

[0041] 3. Stable target tracking mechanism:

[0042] A motion determination method based on joint angle changes was designed to avoid complex skeleton point detection and significantly improve the system's computational efficiency in dense scenarios.

[0043] A two-layer target tracking mechanism is proposed, which solves the real-time tracking problem of multiple targets by updating individual location information and global location records in parallel.

[0044] 4. Complete evaluation system:

[0045] A five-index evaluation system was constructed, including individual movement density, average individual movement density, group movement density, thermal distribution index, and regional load balance, to achieve a comprehensive evaluation of classroom movement status.

[0046] A density distribution heat map based on gridded area division is realized, providing an intuitive visualization basis for teaching evaluation and decision-making.

[0047] 5. Quantification of multi-dimensional evaluation indicators:

[0048] In response to the needs of physical education teaching evaluation, a five-indicator evaluation system was constructed, including individual exercise density, average individual exercise density, group exercise density, thermal distribution index and regional load balance. For the first time, accurate quantitative calculation was achieved through mathematical models (D_i, D_avg, D_g, U, B), which solved the problems of strong subjectivity and low accuracy of traditional manual observation methods, and provided an objective quantitative basis for teachers to scientifically set exercise loads.

[0049] 6. Two-layer target tracking mechanism: An innovative target tracking strategy, combining a lightweight motion tracker and a deep feature extractor, achieves real-time target tracking through a two-layer collaborative mechanism. The first layer uses positional features for rapid tracking, while the second layer uses deep learning features to handle complex situations such as target blur and occlusion. This two-layer mechanism works together to ensure a balance between tracking performance and computational efficiency in dense scenes, ensuring stable operation in gym classes with 50-100 students.

[0050] 7. Student Location Visualization: Each tracked student is assigned a unique ID, and their real-time location and movement status are displayed on the interface. Through visual connecting lines and a color-coding system, teachers can quickly distinguish between active and stationary students, enabling precise monitoring of individual engagement and providing data support for personalized instruction.

[0051] 8. Heatmap Visualization Technology: A multi-level color grading display mechanism is designed to intuitively display the distribution of movement density, including smooth color transitions and time decay effects. By calculating the combined weights of the number of people stationary and active, the system can generate real-time heatmaps reflecting the actual flow of people, allowing teachers to intuitively identify venue utilization and activity hotspots.

[0052] 9. Multimodal Teaching Quality Assessment: Combining video analysis and speech recognition technologies, a comprehensive multimodal teaching quality assessment system has been constructed. By collecting, transcribing, and analyzing the teacher's speech in real time, combined with student movement intensity assessment results, comprehensive monitoring and analysis of the teaching process is achieved. The system assesses multiple dimensions, including the alignment of teaching content with curriculum standards, the degree to which teachers address individual differences, and the scientific nature of teaching strategies. This provides data support and decision-making basis for improving the quality of physical education and teaching, and is a powerful tool for the transformation of traditional physical education assessment towards digital, precise, and intelligent systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1Schematic diagram of a method for evaluating exercise density in a multi-objective physical education classroom based on a single perspective in an embodiment of the present application;

[0054] Figure 2 A schematic diagram of the structure of a physical education classroom exercise density assessment system module provided in an embodiment of the present application;

[0055] Figure 3 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0057] See also Figure 1-3 The present invention provides a technical solution: a method for evaluating the exercise density of a multi-target physical education class under a single viewing angle, comprising the following steps:

[0058] S110. Video acquisition and preprocessing stage: Video data acquisition: Initial video data for student training is collected from a single perspective. GPU parallel acceleration is used to accelerate real-time acquisition and convert continuous video streams into image sequences with time stamps.

[0059] Use a single-view camera to capture videos of 50-100 people in a physical education classroom. The camera must meet the following requirements:

[0060] Use a camera with a wide field of view to ensure complete coverage of the entire activity area;

[0061] The camera should be installed at an appropriate height and adjusted to a suitable downward angle to ensure the target detection effect;

[0062] Ensure that the environment is well lit;

[0063] Set an appropriate frame rate and common resolution to ensure image quality;

[0064] Reduce video jitter by fixing the camera and optimizing the ambient light;

[0065] Convert continuous video streams into time-stamped image sequences and build a multi-layer data management architecture, including a video frame processing layer and a posture data processing layer.

[0066] These layout requirements ensure that the system can obtain a complete and clear picture of the site from a single perspective, providing a reliable data basis for subsequent multi-target detection and motion density assessment.

[0067] Design of the video processing module. The system implements two operating modes: real-time camera acquisition and local video file playback. In practice, the system uses a standard camera for video acquisition, setting an appropriate frame rate and standard resolution to ensure image quality. The camera must be fixed to reduce jitter, and the shooting environment must be well-lit to avoid blurry or dark images. The system processes the collected video data in real time, which can be expressed as:

[0068] D_k={d_1k,d_2k,...,d_nk},n≤100;

[0069] Where d_ik represents the state vector of the i-th detected target in the k-th frame, which contains the target's position and size information. n≤100 means that the system supports processing up to 100 targets simultaneously;

[0070] During video capture, the system ensures video quality through optimized buffer design and automatic exposure control. For real-time capture, the system implements an adaptive frame rate adjustment mechanism to maintain a stable video stream while guaranteeing processing performance. For local video file playback, the system supports decoding of multiple mainstream video formats and provides comprehensive playback control capabilities.

[0071] The system decomposes the acquired video data into a frame sequence, converting the continuous video stream into a series of still images with time stamps. The decomposed sequence can be expressed as: F = {F_1, F_2, ..., F_T}

[0072] Where: F_i represents the i-th frame image, T is the total number of video frames, and each frame image is assigned a unique timestamp t_i.

[0073] This decomposition method facilitates the system to perform subsequent frame-by-frame analysis. The introduction of timestamps ensures the timing accuracy of data processing, which is important for evaluating motion density and analyzing action continuity.

[0074] S120. Data Management Architecture Construction: Build a multi-layer data management architecture to manage data uniformly through video frame processing, audio stream processing, and gesture data processing;

[0075] The system adopts a multi-level data management architecture, including:

[0076] Video frame processing layer: responsible for image data acquisition, decoding and caching;

[0077] Posture data processing layer: analyzes and stores target posture information;

[0078] This layered architecture ensures synchronous data collection and processing, ensuring coordinated operation between the various components of the system. In practical applications, the system coordinates different data streams through a unified time base to ensure the timing consistency of the processing results.

[0079] S210. Geo-Fence Construction: Build a polygon-based geo-fence system. Adjust the area by setting fences and dragging thresholds, and use grid indexing for spatial management.

[0080] Realize area adjustment by setting fences and drag thresholds;

[0081] Use grid index to achieve spatial management and divide the activity area into M×N grid cells;

[0082] Video stabilization is achieved through GPU parallel acceleration, and frame rate control and resolution settings are combined to ensure video quality;

[0083] To define the physical education activity area, the system designed an electronic fence mechanism, which needed to solve the problems of area definition, effective target screening, and density statistics.

[0084] Define the fence system: F = {B, R, M}

[0085] Where B is the bounding polygon: {(x_1, y_1),...,(x_n, y_n)}, which defines the boundaries of the active area; R is the valid area: R = {r|r∈B}, representing the valid detection area within the fence; and M is the mapping matrix: M[x,y] = I((x,y)∈R), which is used to quickly determine whether a point is within the area. This mechanism achieves precise demarcation of the active area and provides spatial constraints for target detection.

[0086] To effectively screen targets within an area, the system designs a location determination mechanism, which needs to solve the problems of target screening, location mapping, and density statistics.

[0087] Define the target screening function:

[0088] Valid(p)=Inside(p,B)∧Active(p)

[0089] Where: Inside(p,B)=∑(CrossProduct(p,e_i))>0, e_i is the edge vector of the boundary polygon, and p is the target position coordinate;

[0090] Active(p)=Movement(p)>τ_m, where Movement(p) is the target's motion amount and τ_m is the motion threshold.

[0091] This mechanism achieves precise screening of targets within the area and ensures the accuracy of density calculation.

[0092] S220. Use deep learning models for human detection and ROI extraction, and achieve efficient model inference through ROI region stitching;

[0093] Use deep learning models for human detection and preliminary positioning, filter targets using confidence thresholds, and extract candidate ROI regions using boundary processing technology. Use ROI data processing technology to optimize system processing efficiency and computing resource utilization.

[0094] The system first uses a deep learning model to detect people, and then extracts and processes ROI for each detected target. The algorithm is as follows:

[0095] Define target detection and ROI extraction functions:

[0096] D(F)={(b_i,r_i)|conf(b_i)>τ_c,i∈[1,N]}

[0097] Where: b_i = [x1_i, y1_i, x2_i, y2_i] is the detection box coordinate, conf(b_i) is the detection confidence, τ_c is the confidence threshold, r_i is the extracted ROI area, and N is the number of detected targets.

[0098] In order to improve the efficiency of multi-target analysis, the system has designed ROI data processing technology. Define ROI processing function:

[0099] M=Process({r_i|i∈[1,N]})

[0100] Where: r_i is the normalized ROI area, and M is the processed data structure. This technology optimizes system processing efficiency and computing resource utilization, significantly improving system processing efficiency.

[0101] To ensure the accuracy of the ROI area, the system has designed a boundary processing mechanism:

[0102] R(b)=Process_Boundary(b,params)

[0103] Among them: b is the original target box, params is the processing parameter, and the mechanism ensures that the ROI area contains complete human target information through appropriate boundary adjustment.

[0104] The system normalizes the extracted ROI for subsequent analysis:

[0105] S(roi) = Normalize(roi, size_constraints, aspect_ratio)

[0106] Where: size_constraints is the size constraint parameter, and aspect_ratio is the aspect ratio characteristic to be maintained. This function ensures that all ROI regions have consistent specification standards, facilitating subsequent feature extraction and analysis.

[0107] S230. Design a grid indexing mechanism to achieve uniform partitioning of the scene space and target management through grid partitioning;

[0108] To achieve spatial management of multiple targets, the system designs a grid indexing mechanism. This mechanism needs to solve problems such as spatial partitioning, target mapping, and status maintenance.

[0109] Define the scene grid partitioning:

[0110] G = (M, N), representing the number of grids in the horizontal and vertical directions;

[0111] C(i, j) = (W / M, H / N), representing the grid cell size;

[0112] Where: M and N are the number of grids in the horizontal and vertical directions, W and H are the scene sizes, C(i, j) is the grid cell size, and A_cell = (W / M) * (H / N) is the grid cell area,

[0113] i ∈ [0, M - 1], j ∈ [0, N - 1] are the grid indices. This mechanism achieves uniform partitioning of the scene space and provides a basic framework for target management.

[0114] To achieve fast mapping of targets to grids, the system designs a grid index calculation mechanism. This mechanism needs to solve problems such as coordinate conversion, boundary handling, and conflict resolution.

[0115] Define the target mapping function: GridID(p) = (x / C_w, y / C_h)

[0116] Where: p = (x, y) is the target position coordinate, C_w and C_h are the width and height of the grid cell, · represents the floor operation, and the mapping constraint is: 0 ≤ x / C_w < M, 0 ≤ y / C_h < N. This mechanism achieves an O(1) time complexity mapping from the target position to the grid index.

[0117] To achieve dynamic management of grid cells, the system designs a cell status maintenance mechanism. This mechanism needs to solve problems such as target recording, density calculation, and status update.

[0118] Define the grid cell state: Cell_ij, which contains three basic elements: target set Objects, density value ρ_ij and state flag s_ij.

[0119] Where: Objects = {o_1, o_2, ..., o_k} represents the target set contained in the current grid;

[0120] ρ_ij=|Objects| / A_cell represents the grid motion density, A_cell is the grid cell area;

[0121] s_ij=Active, when ρ_ij>τ_d; s_ij=Inactive, when ρ_ij≤τ_d, where τ_d is the density threshold.

[0122] This mechanism realizes the real-time status maintenance of grid cells and supports efficient target management.

[0123] S240. Using a feature point-based video stabilization method to estimate and compensate for camera shake by tracking the movement of key feature points;

[0124] Target stabilization processing: The system uses a feature point-based target stabilization method to estimate and compensate for target position fluctuations by tracking key points of the human body and target contour features.

[0125] Define the target stabilization function:

[0126] Stabilize(O_t)=T_t·O_t

[0127] Where: O_t is the target in the current frame, T_t is the stabilization transformation matrix;

[0128] T_t=argmin_T∑||p_i-T·p'_i||2, where p_i, p'_i are the feature point pairs matched by the target in adjacent frames.

[0129] The system uses GPU parallel acceleration to implement this processing, combined with appropriate key point filtering and position prediction algorithms to ensure the stability of target position tracking. At the same time, measures such as fixed camera installation and optimized ambient light further improve the accuracy of target detection and tracking.

[0130] S310. Design a multi-target status update mechanism to achieve continuous tracking of multiple targets through position information and similarity matching;

[0131] It mainly solves the target tracking problem in dense scenes with 50-100 people. The parallel processing architecture designs a multi-stage pipeline architecture and task partitioning strategy, and realizes real-time processing of multiple targets through GPU acceleration.

[0132] In multi-target detection, an efficient parallel processing architecture is required. The system designs a multi-stage pipeline, implements task partitioning, and optimizes resource scheduling.

[0133] Define the parallel processing architecture:

[0134] P={P_stream,P_block,P_schedule}

[0135] in:

[0136] P_stream is a multi-stage pipeline architecture, including detection, tracking, and analysis;

[0137] P_block is the task block strategy, which divides N targets into N / M blocks;

[0138] P_schedule is a resource scheduling mechanism that dynamically allocates computing resources;

[0139] This mechanism realizes real-time processing of multiple targets and provides basic guarantee for system performance.

[0140] Status update mechanism, the system designs a multi-target status update mechanism:

[0141] X_k={x_1k,x_2k,...,x_nk};

[0142] Where: x_ik = [p_ik] is the position state vector of the i-th target in the k-th frame;

[0143] p_ik=[x_ik,y_ik] is the position vector;

[0144] The state update is based on location similarity calculation and target identity association, and continuous tracking is achieved by identifying targets at the same location;

[0145] This mechanism enables continuous tracking of multiple targets and provides reliable trajectory data for motion density analysis.

[0146] In order to achieve efficient tracking and management of multiple targets, the system has designed a two-layer target tracking mechanism:

[0147] The global layer tracker (T_g) is responsible for maintaining the basic state of all targets:

[0148] T_g={P,U}

[0149] in:

[0150] P = [p_1, p_2, ..., p_n] is a set of position records;

[0151] U = [t_1, t_2, ..., t_n] is the set of timestamp records;

[0152] The individual layer tracker (T_i) is responsible for recording the detailed status of each target:

[0153] T_i={t_f,t_m,t_t,s_m,s_f};

[0154] in:

[0155] t_f is the first appearance time;

[0156] t_m is the set of exercise period records;

[0157] t_t is the total tracking time;

[0158] s_m is the motion state flag;

[0159] s_f is the site location mark;

[0160] This two-layer architecture achieves basic tracking of targets through the global layer and maintains detailed motion state information through the individual layer, thereby achieving accurate tracking and state management of multiple targets.

[0161] S320. Establish an identity mapping table and a target identity feature set, and achieve continuous maintenance of the target identity through multi-feature association measurement;

[0162] Identity management mechanism. To maintain the identity information of multiple targets, an identity mapping table is established:

[0163] ID_k={id_1k,id_2k,...,id_nk};

[0164] Define the target identity feature set:

[0165] I_k={i_1k,i_2k,...,i_nk};

[0166] Where: i_jk is the identity feature vector of the jth target in the kth frame, which contains multi-dimensional feature information of appearance, motion and position;

[0167] To solve the identity association problem of multiple targets, the system designs a cross-frame matching mechanism. Define the cross-frame matching cost matrix:

[0168] C_k={c_ij|i∈[1,n],j∈[1,m]};

[0169] in:

[0170] c_ij=w_1d_a(i,j)+w_2d_m(i,j)+w_3d_p(i,j);

[0171] d_a(i,j) is the appearance feature distance;

[0172] d_m(i,j) is the motion feature distance;

[0173] d_p(i,j) is the position feature distance;

[0174] w_1, w_2, w_3 are weight coefficients, and ∑w_i=1;

[0175] This mechanism realizes the multi-feature association measurement between targets and provides a decision basis for identity matching.

[0176] S330. Build a target tracking model to handle occlusion issues in dense scenes through target visibility evaluation and multi-dimensional feature similarity mechanism;

[0177] The target tracking mechanism builds a two-layer tracking architecture to achieve accurate tracking through global state maintenance and individual state recording; an identity management mechanism is designed to ensure the continuity of the target ID; step S330, occlusion processing mechanism. In dense scenes, the system builds a target tracking model:

[0178] T={F_model,O_state,ID_map};

[0179] in:

[0180] F_model is the target feature model, which includes appearance and motion features;

[0181] O_state is the occlusion state prediction, which evaluates the target visibility;

[0182] ID_map is the identity mapping table, which maintains the continuity of the target ID;

[0183] The system establishes a target visibility assessment model:

[0184] V_jk=σ(w_v^T[o_jk,s_jk,t_jk]);

[0185] in:

[0186] o_jk is the occlusion rate;

[0187] s_jk is the target scale change rate;

[0188] t_jk is the target tracking quality;

[0189] σ(·) is the sigmoid activation function;

[0190] w_v is the weight vector;

[0191] To handle occlusion, the system designs a multi-dimensional feature similarity mechanism:

[0192] S(t1,t2)=w_1·S_p+w_2·S_s+w_3·S_r+w_4·S_m;

[0193] in:

[0194] S_p=exp(-d / 200) is the position similarity, d is the distance between targets;

[0195] S_s=exp(-δ_s·2) is the size similarity, δ_s is the size change rate;

[0196] S_r = exp(-δ_r·2) is the aspect ratio similarity, δ_r is the aspect ratio change rate;

[0197] S_m = exp(-d_p / 100) is the motion prediction similarity, d_p is the predicted position deviation; w_1 to w_4 are weight coefficients, which are assigned according to the importance of each feature in different scenarios. To improve the tracking efficiency of multiple targets, the system designs a local search strategy:

[0198] Define the target search area:

[0199] R_j={(x,y)|||p_j-(x,y)||≤r_j};

[0200] in:

[0201] p_j is the predicted position of target j;

[0202] r_j is the search radius, and the dynamic adjustment formula is:

[0203] r_j=max(r_min,min(r_max,η||v_j||));

[0204] r_min is the minimum search radius;

[0205] r_max is the maximum search radius;

[0206] η is the speed impact factor;

[0207] v_j is the moving speed of target j;

[0208] This strategy achieves adaptive adjustment of the search space and significantly improves tracking efficiency;

[0209] Occlusion handling mechanism: Based on multi-dimensional feature similarity and target visibility evaluation, combined with a local search strategy, it solves the occlusion problem in dense scenes.

[0210] S410. Calculate the three indicators of individual movement density, average individual movement density and group movement density, and generate a density distribution heat map by gridding the area;

[0211] Exercise density is a core indicator for assessing student exercise participation. This phase requires calculating and analyzing various metrics based on single-perspective tracking results, including individual exercise density, average individual exercise density, group exercise density, thermal distribution indicators, and regional load balance. The system utilizes a grid-based spatial management approach to accurately calculate regional exercise density from a single perspective and has designed an adaptive density calculation method to accommodate scenarios of varying exercise intensities. These metrics provide a comprehensive, quantitative basis for teachers to evaluate teaching effectiveness.

[0212] Step S410: Movement density calculation. Multiple evaluation indicators are calculated: individual movement density, average individual movement density, group movement density, thermal distribution index, and regional load balance. A density distribution heat map is generated through gridded regional division to achieve a comprehensive assessment of movement status.

[0213] In individual motion assessment, the motion density of each target needs to be calculated. The system needs to solve the problems of time statistics, state update, and density calculation. Define the individual motion density function: D_i(t) = M_i(t) / T_i(t)×100%

[0214] Where: Mi(t) is the duration of movement of individual i within the fence, and T(t) is the total tracking time of individual i within the fence. Time update: Mi(t+1)=M(t)+Δt, if and only if F(t)=1 and S(t)=1, T(t+1)=T(t)+Δt, if and only if F(t)=1, F(t) is the fence state, 1 means inside the fence, S(t) is the motion state, 1 means motion, and Δt is the sampling time interval.

[0215] The system designs a motion state determination mechanism based on joint angle thresholds: Define the motion state function:

[0216] S_i(t)=

[0217] 1. Meet one of the following conditions:

[0218] For each type of joint j∈J, the change of its angle θ_j from the reference value exceeds the set threshold;

[0219] 0, otherwise;

[0220] Among them: the system sets appropriate threshold parameters and judgment criteria for different types of joints based on the kinematic characteristics of the human body.

[0221] Define individual motion judgment function:

[0222] M_i(t)=|{j|j∈J_i,S_j(t)=1}| / |J_i|>ε;

[0223] Where: J_i is the valid joint set of individual i, |J_i| is the number of valid joints, and ε is the motion threshold, which can be dynamically adjusted according to different teaching scenarios. When a certain proportion of joints are in motion, the individual is judged to be in motion.

[0224] This mechanism realizes the real-time determination of the target's motion state and provides a basic basis for motion density calculation.

[0225] When assessing regional activity, the system primarily considers two key indicators: the number of people stationary and the number of people moving. Using video analysis technology, the system identifies each person's movement status and divides the area into M×N grid cells.

[0226] The system uses a smooth transition mechanism to handle changes in the number of people, avoiding sudden changes in the heat map display. The function for calculating the actual number of people in a grid cell is defined as: C_real(t) = C_prev(t-1) + (C_target(t) - C_prev(t-1)) × T_speed, where C_prev is the number of people at the previous moment, C_target is the target number of people, and T_speed is the transition speed coefficient (a faster transition speed is used as activity increases).

[0227] Regional activity scores are primarily based on the number of people stationary and active, with the system prioritizing the impact of active people on activity. A simplified regional activity scoring function is defined as: A(r,t) = αS(r,t) + βM(r,t), where S(r,t) is the number of people stationary in the area, M(r,t) is the number of people active, and α and β are weighting coefficients, with α < β, reflecting the greater impact of activity on activity. This mechanism displays regional activity using different color depths, providing intuitive visualization indicators for analyzing regional movement characteristics.

[0228] The system implements a continuous decay mechanism for thermal effects, allowing activity to decay naturally over time. The thermal decay function is defined as: D(r,t) = D(r,t-1) × max(0,1-(t-t_last) / t_window), where t_last is the last active time in the area and t_window is the decay window. When an area transitions from occupied to unoccupied, activity doesn't disappear immediately. Instead, it gradually decays to zero within a preset time window, enhancing the stability and continuity of the system display.

[0229] The system uses color smoothing technology to ensure a natural and smooth presentation of the heat map. The color transition function is defined as: C_result = C_prev × (1-f) + C_target × f, where C_prev is the current color, C_target is the target color, and f is the transition factor. The speed of color change is dynamically adjusted based on the direction of change.

[0230] In terms of student location visualization, the system implements an intuitive personnel tag display mechanism. A unique ID is assigned to each tracked student, and their real-time location and motion status are displayed on the interface. Define the position visualization function: V(p,s) = (v_pos,v_text,v_color), where p is the student's position coordinate, s is the motion state (moving / stationary), v_pos is the display position, v_text is the status text, and v_color is the status color (moving is green, stationary is white). The system connects student ID tags to their actual locations through visual connecting lines and uses color coding to intuitively distinguish different motion states. This design enables teachers to quickly identify the real-time location and participation status of each student in the classroom.

[0231] For interactive adjustments to electronic fences, the system has designed a flexible boundary definition mechanism. Teachers can adjust the electronic fence range in real time by dragging and dropping markers on the interface to adapt to different teaching scenarios. The fence interaction function is defined as: F_adjust(p_i,p_new) = {p_1,p_2,...,p_new,...,p_n}, where p_i is the marker being adjusted and p_new is the new position. The system supports the dynamic addition and removal of fence points and implements visual marking based on bucket icons, making fence boundaries clear at a glance. When students enter or leave the fenced area, the system automatically updates relevant statistics to ensure the accuracy of motion density calculations.

[0232] In terms of grid unit status visualization, the system implements a detailed status display function. Each grid unit can simultaneously display the number of people stationary and the number of people exercising, allowing teachers to accurately grasp the utilization of the venue. The grid status display function is defined as: G(col,row)=f(S(col,row),M(col,row)), where S(col,row) and M(col,row) represent the number of people stationary and exercising, respectively. The system achieves composite status display by overlaying semi-transparent color blocks. The influence of people exercising on the heat map color is more weighted than that of people stationary. This design intuitively reflects the density distribution of different types of activities and provides an accurate spatial reference for organizing teaching activities.

[0233] A scientific evaluation index system is a crucial tool for ensuring the quality of physical education. This phase constructed a comprehensive, multi-level evaluation system based on single-perspective data, encompassing five dimensions: individual movement density, average individual movement density, group movement density, thermal distribution index, and regional load balance. By integrating spatiotemporal movement characteristics acquired from a single perspective, the system achieves a comprehensive assessment of classroom movement status. This evaluation system not only objectively reflects classroom movement participation but also provides teachers with timely feedback on their teaching.

[0234] In step S410, a classroom evaluation model is constructed. The classroom evaluation model PE(t) is constructed, which includes five core indicators:

[0235] PE(t)={D_i(t),D_avg(t),D_g(t),U(t),B(t)};

[0236] in:

[0237] -D_i(t) is the individual movement density, reflecting the movement state of a single student;

[0238] -D_avg(t) is the average individual motion density, reflecting the overall average motion level;

[0239] -D_g(t) is the group movement density, reflecting the degree of movement participation at a certain moment;

[0240] -U(t) is the thermal distribution index, reflecting the distribution of hot spots of spatial activity in the site;

[0241] -B(t) is the regional load balance, reflecting the degree of balance in space usage;

[0242] To calculate individual movement density, we need to assess each student's movement status. The system needs to calculate and track movement duration and ensure data validity. Define the individual movement density function: D_i(t) = M_i(t) / T_i(t) × 100%.

[0243] Among them: M_i(t) is the exercise time of student i in the fence, T_i(t) is the total tracking time of student i in the fence. This indicator reflects the individual exercise intensity of students by calculating the proportion of exercise time of a single student.

[0244] To calculate average individual activity density, we need to assess the cumulative activity of all students. The system needs to calculate total activity duration and total tracking duration to ensure scale compliance. Define the average individual activity density function: D_avg(t) = M_total(t) / T_total(t) × 100%.

[0245] Where: M_total(t) = ∑{i=1}^NM_i(t), represents the total cumulative exercise time for all students; T_total(t) = ∑{i=1}^NT_i(t), represents the total cumulative tracking time for all students; and N represents the total number of students. This metric reflects the average exercise intensity of all students during the entire class period by calculating the ratio of the total cumulative exercise time to the total tracking time.

[0246] To calculate group movement density, we need to assess the overall movement state at the current moment. The system needs to count the total number of students on the field, the number of students exercising, and calculate the real-time percentage. We define the group movement density function as: D_g(t) = N_m(t) / N_a(t) × 100%.

[0247] Where: N_m(t) is the number of students in motion and within the fence at time t, N_a(t) is the total number of students within the fence at time t, and the statistical interval is Δt = 1.0s. This metric reflects the overall physical activity participation at a given moment by calculating the proportion of students in motion at the current moment, providing teachers with real-time feedback on classroom activity. Unlike average individual motion density, group motion density reflects the instantaneous state rather than the cumulative state.

[0248] In the calculation of heat distribution index, it is necessary to evaluate the utilization efficiency of the active area. The system needs to divide the grid cells, calculate the cell density, and calculate the weighted utilization rate. Define the heat distribution index function:

[0249] U(t)=∑{i=1}^M∑{j=1}^N(w_ij×ρ_ij(t)) / (M×N);

[0250] Where: w_ij is the weight coefficient of the grid unit (i, j): w_ij∈[0,1], ρ_ij(t) is the target density of the grid unit (i, j) at time t, M and N are the number of grid divisions in the horizontal and vertical directions. The system sets the appropriate division granularity based on the site size and analysis accuracy requirements. This indicator reflects the utilization efficiency of the site space resources by calculating the weighted density distribution of the gridded area.

[0251] The system also designs a regional load balancing evaluation mechanism. Define regional load balancing degree:

[0252] B(t)=1-σ(ρ(t)) / μ(ρ(t));

[0253] in:

[0254] σ(ρ(t)) is the standard deviation of the density of all grid cells at time t;

[0255] μ(ρ(t)) is the average density of all grid cells at time t;

[0256] This indicator reflects the degree of balance in space utilization. The closer B(t) is to 1, the more balanced the space utilization is.

[0257] These space utilization indicators form a comprehensive space assessment system that not only quantifies the efficiency of venue resource use but also provides data support for teachers to adjust their teaching methods. By calculating and visualizing these indicators in real time, the system helps teachers optimize the planning and utilization of teaching spaces.

[0258] S510. Collect teacher classroom voice data and perform real-time speech-to-text processing through dual-source collection and streaming voice processing architecture;

[0259] Scientific teaching evaluation should include not only analysis of students' movement intensity but also analysis of teachers' teaching behaviors. This phase utilizes speech recognition and natural language processing technologies to achieve real-time collection, transcription, and intelligent analysis of teachers' classroom speech, providing a comprehensive basis for quality assessment of physical education.

[0260] Step S510: Teacher's voice collection and transcription. Collect teacher's classroom voice data, and through dual-source collection and streaming voice processing architecture, realize real-time voice-to-text processing to build a complete classroom language archive.

[0261] In order to achieve accurate collection of the teacher's speech, the system has designed a dual-source sound collection mechanism:

[0262] Define the sound source acquisition system: A = {V, M, P};

[0263] Among them: V is the audio channel of the video file, which is used for offline analysis; M is the real-time microphone input, which is used for online analysis; P is the audio preprocessing module, which is responsible for noise removal, volume equalization and signal enhancement.

[0264] The system builds a streaming voice processing architecture to achieve real-time processing of continuous voice data:

[0265] Define the speech processing pipeline: S = {C, T, B};

[0266] C is the audio capture module, which collects audio clips at regular intervals; T is the transcription processing module, which converts the audio into text; and B is the text buffer, which temporarily stores the transcription results. This mechanism supports long-term continuous speech transcription and is suitable for complete classroom teaching scenarios.

[0267] The system implements a language transcription mechanism with timestamps:

[0268] Define the transcription function: TR(a) = {(t_1,s_1),(t_2,s_2),...,(t_n,s_n)};

[0269] Where a represents the audio data, t_i represents the timestamp, and s_i represents the corresponding transcript. This mechanism ensures that each segment of speech has an accurate time index, providing foundational data for subsequent teaching segment identification and duration analysis. The system can generate time-stamped classroom transcripts like "[00:31] Hello, students, class is here!", accurately recording the timing of each teacher's utterance.

[0270] The system has designed a basic classroom information integration mechanism, linking classroom records with the following information:

[0271] 1. Classroom metadata: class time, teacher, level, teaching unit, teaching content, class number, class number, number of students, venue, etc.

[0272] 2. Exercise density data: overall individual exercise density, group exercise density, and individual student exercise density;

[0273] 3. Teaching standard requirements: the duration of each link, teaching content requirements, etc. specified in the curriculum standards.

[0274] This information provides the necessary background and reference for subsequent in-depth analysis.

[0275] S520. Conduct structured analysis of the transcribed language content to identify teaching links, extract key instructions, and assess language quality;

[0276] Step S520, classroom segment identification and analysis, processing the classroom transcript with time stamps, automatically identifying the teaching segments, and analyzing the rationality of segment duration allocation.

[0277] The system has built an automatic recognition model for teaching links:

[0278] Define the link identification function:

[0279] SS(T)={(s_1,t_start1,t_end1),(s_2,t_start2,t_end2),...,(s_k,t_startk,t_endk)};

[0280] Where T is the timestamped teacher language set, s_i is the identified teaching segment label, including introduction, warm-up, basic part, physical exercise, and relaxation. t_starti and t_endi are the start and end times of each segment, respectively. The system uses a large language model to analyze language content and temporal distribution, automatically identifying the boundaries of each teaching segment within a complete lesson.

[0281] The system has designed a mechanism for analyzing the duration of classroom sessions:

[0282] Define the segment duration function: D(s_i) = t_endi - t_starti;

[0283] Among them: The system compares the calculated actual duration of each session with the ideal duration specified in the curriculum standards:

[0284] δ(s_i)=D(s_i) / D_standard(s_i)×100%;

[0285] Where D_standard(s_i) is the ideal duration of segment s_i according to the curriculum standard, and δ(s_i) is the duration compliance of the segment. When δ(s_i) deviates significantly from the standard (for example, the actual duration of a teaching segment is significantly lower than the required duration in the curriculum standard), the system will mark the segment duration allocation as unreasonable and generate corresponding improvement suggestions.

[0286] S530. Construct a classroom evaluation model based on five core indicators, and analyze the effectiveness of teaching strategies and their correlation with student engagement by combining phonetic transcription;

[0287] Step S530: Teacher language quality analysis and classroom evaluation model construction. A multi-dimensional analysis of the teacher's language in the classroom recording is performed to assess the quality of the teacher's language and the rationality of the teaching strategy. A classroom evaluation model is also constructed to provide a quantitative basis for teaching quality assessment.

[0288] Construct a classroom evaluation model PE(t), which includes five core indicators:

[0289] PE(t)={D_i(t),D_avg(t),D_g(t),U(t),B(t)};

[0290] in:

[0291] -D_i(t) is the individual movement density, reflecting the movement state of a single student;

[0292] -D_avg(t) is the average individual motion density, reflecting the overall average motion level;

[0293] -D_g(t) is the group movement density, reflecting the degree of movement participation at a certain moment;

[0294] -U(t) is the thermal distribution index, reflecting the distribution of hot spots of spatial activity in the site;

[0295] -B(t) is the regional load balance, reflecting the degree of balance in space usage;

[0296] The system builds a teacher language classification model:

[0297] Define a language classification system: C = {C_guide, C_feedback, C_motivate, C_organize, C_evaluate}

[0298] Among them: C_guide is guiding language, such as "Raise your knees, fold the rope in quarters and hold it tightly"; C_feedback is feedback language, such as "Very good, your knees are raised quite high"; C_motivate is motivational language, such as "Keep up the good work, okay"; C_organize is organizational language, such as "Return to your position"; C_evaluate is evaluative language, such as "The boys' formation is a little messy." The system uses a large language model to classify each teacher's language:

[0299] Class(t)=argmax_c∈CP(c|t);

[0300] Where: t is a single sentence of teacher language, P(c|t) is the probability that t belongs to category c.

[0301] The system calculates the number and distribution ratio of teachers' various languages:

[0302] _c=|{t|Class(t)=c}|;

[0303] _c=N_c / |T|×100%;

[0304] Where |T| is the total number of sentences in the teacher's language, N_c is the number of languages ​​in category c, and R_c is the proportion of languages ​​in category c. The system compares this distribution with the language distribution in an ideal teaching model to assess the rationality of the teacher's language structure. For example, a high proportion of organizational language may indicate inefficient classroom organization; a low proportion of feedback language may indicate insufficient teacher-student interaction.

[0305] The system has designed a teaching content evaluation mechanism based on curriculum standards:

[0306] Define the curriculum standard compliance evaluation function: CS(T,S)=∑{i=1}^n(w_i×Match(T,s_i));

[0307] Where: T is the classroom language set, S = {s_1,s_2,...,s_n} is the set of curriculum requirements, Match(T,s_i) is the degree of match between the language content and the curriculum requirement s_i, and w_i is the weight coefficient. This mechanism uses deep semantic understanding to assess whether the teaching content meets the requirements of the new curriculum standards, including coverage of core competencies and core subject experiences.

[0308] The system has built an individual difference attention assessment model:

[0309] Define individual attention function: A(T,I)=|{t∈T|Ref(t,I)=True}| / |T|;

[0310] Where T is the classroom language set, I is the student set, and Ref(t, I) indicates whether language t incorporates individual differences. The system analyzes whether teachers pay attention to individual student differences, specifically providing targeted guidance and attention to students who exhibit significant differences in movement density data (e.g., student 13 in the example data, whose movement density is only 31%, significantly below average).

[0311] At the same time, the system uses the constructed evaluation model PE(t) and combines it with currently collected movement density data and teacher language analysis to conduct in-depth teaching strategy analysis, assessing the rationality of teaching design and its alignment with curriculum standards, and analyzing the correlation between teacher strategies and student movement density. This analysis comprehensively considers movement density indicators, space utilization, and teacher language quality, providing a comprehensive quantitative basis for subsequent teaching quality evaluation.

[0312] S540. Integrate all analysis results to generate a comprehensive classroom evaluation report and targeted improvement suggestions, providing teachers with specific measures to improve teaching quality.

[0313] Step S540: Multimodal data fusion and report generation. The motion density data from the video analysis and the language analysis results are integrated to generate a comprehensive teaching evaluation report and personalized improvement suggestions.

[0314] The system builds a multimodal feature fusion model:

[0315] Define feature fusion function: F(V,L)={f_1,f_2,...,f_q};

[0316] Where V is the video analysis feature set (including motion density data), L is the language analysis feature set, and f_i is the fused feature dimension. This mechanism enables temporal alignment and correlation analysis of motion and language data, providing multi-dimensional evidence for comprehensive evaluation of teaching quality.

[0317] The system has designed a teaching consistency evaluation mechanism:

[0318] Define the teaching consistency function: C(A,T)=sim(A,T);

[0319] Where A represents the actual teaching behavior (derived from video analysis), T represents the teaching language content (derived from speech analysis), and sim represents the similarity function. For example, the system can analyze whether the teacher provides targeted guidance to students with low exercise intensity; whether high-quality guidance improves the efficiency of physical training despite insufficient time, etc.

[0320] The system implements a mechanism for generating personalized teaching improvement suggestions:

[0321] Define the suggestion generating function: R(A) = {r_1, r_2, ..., r_p};

[0322] Where: A is the result set of comprehensive classroom analysis, and r_i is the generated improvement suggestions. Based on the comprehensive analysis of classroom records and exercise density data, the system generates specific and actionable teaching improvement suggestions, such as:

[0323] 1. "It is recommended to increase the duration of the physical training session to more than 10 minutes. The current duration is only 3 minutes, which does not meet the curriculum requirements."

[0324] 2. "It is recommended to increase feedback language tailored to individual differences, especially focusing on students whose exercise density is less than 50% (such as students 13 and 14)"

[0325] 3. "It is recommended to add more reflective questions to the teaching design to promote students' in-depth understanding of the key points of the movements."

[0326] The system ultimately generates a structured teaching analysis report, which includes the following core components:

[0327] 1. Overview of basic classroom information and exercise density data: class time, teacher, class number, class, exercise density data and other basic information;

[0328] 2. Teachers’ language type distribution and quality assessment: distribution ratio and quality assessment of various languages;

[0329] 3. Analysis of teaching duration: comparison of actual duration of each stage with standard requirements and rationality assessment;

[0330] 4. Curriculum Standards Conformity Analysis: Analysis of the degree of conformity between teaching content and the new curriculum standards;

[0331] 5. Teaching highlights and concerns: Identified teaching highlights and issues that require special attention;

[0332] 6. Personalized teaching improvement suggestions: specific and practical teaching improvement suggestions.

[0333] This technical solution fully demonstrates the advantages of this system in combining video analysis and voice analysis technologies. Through in-depth analysis of classroom recordings, it provides teachers with objective, comprehensive and specific teaching evaluation and improvement basis. It is a powerful tool to promote the improvement of physical education teaching quality, and also provides data support and decision-making reference for teachers' professional growth.

[0334] In the target detection and ROI extraction steps, a deep learning model is used to extract image features from the video frame sequence and detect the position and bounding box of the moving target.

[0335] In the motion density calculation step, the motion state of each target is determined, and the individual motion density is calculated based on the changes in joint angles.

[0336] In the multi-target tracking step, the global layer tracker is responsible for maintaining the position, velocity and timestamp information of all targets;

[0337] The individual-level tracker is responsible for recording the first appearance time, movement period, and movement status of each target;

[0338] Design identity management mechanisms to ensure target ID continuity.

[0339] In the evaluation index system step: constructing a five-dimensional evaluation index including individual movement density, average individual movement density, group movement density, thermal distribution index and regional load balance.

[0340] In the teacher language intelligent analysis step, a dual-source acquisition mechanism is used to collect the teacher's classroom voice, including a video file audio channel and real-time microphone input;

[0341] Build a streaming speech processing architecture to analyze and transcribe teacher language content and generate classroom transcripts with timestamps;

[0342] Identify teaching links based on natural language processing technology and analyze the rationality of link duration;

[0343] Construct a teacher language classification model, which divides teacher language into five categories: instructional, feedback, motivational, organizational, and evaluative;

[0344] Analyze the rationality of teachers' language structure, including the distribution ratio and quality of various languages;

[0345] Evaluate the conformity of teaching content based on curriculum standards and analyze the degree to which teachers pay attention to individual differences.

[0346] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements the multi-target physical education classroom motion density assessment method under a single perspective as described in any of the above technical solutions.

[0347] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for evaluating multi-target physical education classroom exercise density under a single perspective as described in any of the above technical solutions is implemented.

[0348] The core innovations of this invention can be divided into the following three aspects: video acquisition and processing, motion analysis and target tracking, and innovation of the teaching evaluation system.

[0349] In terms of video acquisition and processing, the present invention has the following innovations:

[0350] Single perspective multi-target analysis technology:

[0351] The system realizes multi-target motion density assessment based on single-view video without the need for multi-angle camera deployment, allowing the system to analyze videos captured by portable devices such as mobile phones and ordinary cameras, significantly reducing hardware costs and site requirements, making the technology widely applicable in ordinary school environments.

[0352] Efficient multi-target parallel processing architecture:

[0353] In response to the requirements of the national compulsory physical education and health curriculum standards for exercise load assessment, a parallel processing system for 50-100 people was designed. Through a multi-level pipeline and task blocking strategy, combined with GPU acceleration, real-time detection and tracking of multiple targets was achieved, solving the problem of traditional manual observation that is difficult to count individual exercise status in real time, and greatly improving data collection and processing efficiency.

[0354] Efficient ROI data processing technology:

[0355] It implements an efficient processing mechanism for multi-target ROI data. Through grid indexing and parallel processing strategies, it significantly optimizes system processing efficiency and computing resource utilization, and performs well in large-scale crowd analysis scenarios.

[0356] In terms of motion analysis and target tracking, the present invention has the following innovations:

[0357] Precise joint angle movement determination mechanism:

[0358] This system uses thresholds for key human joint angles as the basis for motion determination, setting specific rules for different joint types (flexion joint angles less than the threshold are considered motion; extension joint angles greater than the threshold are considered motion). This system then determines individual motion states based on the proportion of joint activity. Compared to traditional methods, this mechanism offers advantages such as high computational efficiency, applicability to densely packed scenarios, more accurate results, and strong robustness. It can operate stably in physical education classrooms with 50-100 students.

[0359] Two-layer target tracking mechanism:

[0360] An innovative target tracking strategy, combining a lightweight motion tracker and a deep feature extractor, achieves real-time target tracking through a two-layer collaborative mechanism. The first layer uses positional features for rapid tracking, while the second layer uses deep learning features to handle complex situations such as target blur and occlusion. This two-layer mechanism works together to balance tracking performance and computational efficiency in dense scenes, ensuring stable operation in large-scale physical education scenarios.

[0361] In terms of the teaching evaluation system, the present invention has the following innovations:

[0362] A quantitative evaluation system based on five indicators:

[0363] In response to the needs of physical education teaching evaluation, a five-indicator evaluation system was constructed, including individual exercise density, average individual exercise density, group exercise density, thermal distribution index and regional load balance. For the first time, accurate quantitative calculation was achieved through mathematical models (D_i, D_avg, D_g, U, B), which solved the problems of strong subjectivity and low accuracy of traditional manual observation methods, and provided an objective quantitative basis for teachers to scientifically set exercise loads.

[0364] Space utilization assessment based on electronic fence:

[0365] In view of the relatively fixed characteristics of physical education classroom venues, electronic fence technology is combined with computer vision. Through polygonal area division and grid indexing, accurate statistics of exercise density and real-time evaluation of space usage in different teaching areas are achieved. Combined with heat map distribution indicators and regional load balance, it provides a quantitative basis for teachers to optimize venue usage and teaching organization.

[0366] In terms of multimodal teaching quality assessment, the present invention has the following innovations:

[0367] Multimodal teaching quality assessment system:

[0368] Combining video analysis technology with intelligent analysis of teacher language, and using speech recognition and natural language processing technology, we can achieve real-time collection, transcription and intelligent analysis of teachers' classroom language, and build a complete teaching link identification, teacher language quality analysis and multimodal data fusion mechanism, providing a comprehensive quality assessment basis for physical education teaching.

[0369] Generate personalized teaching improvement suggestions:

[0370] Based on the fusion results of video analysis and language analysis, the system can generate specific and actionable teaching improvement suggestions, provide precise guidance on aspects such as exercise density, duration of teaching sessions, and attention to individual differences, and help teachers continuously improve teaching quality.

[0371] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 3As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities, supporting multi-target parallel processing and real-time video analysis. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store video data, video frame data, posture estimation results, motion density data and teacher language analysis results. The network interface of the computer device is used to communicate with an external terminal through a network connection to realize remote transmission and shared analysis of data.

[0372] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described embodiments. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0373] The technical features of the above embodiments may be combined in any manner. To simplify the description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there are no contradictions in the combination of these technical features, they should be considered to be within the scope of this specification. For example, appropriate hardware configurations and algorithm parameters can be selected based on the hardware conditions and teaching scale of different schools to ensure that the system can achieve optimal performance in different application scenarios.

[0374] In general, the multi-target motion density assessment system based on a single perspective proposed in the present invention provides a new technical means for physical education teaching evaluation through the deep integration of video analysis and voice analysis technology. The system can not only objectively quantify students' exercise load and venue utilization efficiency, but also conduct a comprehensive analysis of teachers' teaching behaviors, forming a complete classroom quality evaluation system. This technical solution effectively solves the problems of strong subjectivity, low accuracy and difficulty in quantification of traditional physical education teaching evaluation methods, and provides strong technical support and decision-making reference for improving the quality of physical education and implementing the national compulsory education physical education and health curriculum standards. It is of great significance to promote the scientific and modern development of physical education in my country. Therefore, the scope of protection of the patent of this invention shall be based on the attached claims.

[0375] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for evaluating the exercise density of a multi-objective physical education class under a single viewing angle, characterized in that: The following steps are involved: S110. Video acquisition and preprocessing stage: Video data acquisition: Initial video data for student training is collected from a single perspective. GPU parallel acceleration is used to accelerate real-time acquisition and convert continuous video streams into image sequences with time stamps. S120. Data Management Architecture Construction: Build a multi-layer data management architecture to manage data uniformly through video frame processing, audio stream processing, and gesture data processing; S210. Geo-Fence Construction: Build a polygon-based geo-fence system. Adjust the area by setting fences and dragging thresholds, and use grid indexing for spatial management. S220. Use deep learning models for human detection and ROI extraction, and achieve efficient model inference through ROI region stitching; S230. Design a grid indexing mechanism to achieve uniform division of scene space and target management through grid division; S240. Using a feature point-based video stabilization method to estimate and compensate for camera shake by tracking the movement of key feature points; S310. Design a multi-target status update mechanism to achieve continuous tracking of multiple targets through position information and similarity matching; S320. Establish an identity mapping table and a target identity feature set, and achieve continuous maintenance of the target identity through multi-feature association measurement; S330. Build a target tracking model to handle occlusion issues in dense scenes through target visibility evaluation and multi-dimensional feature similarity mechanism; S410. Calculate the three indicators of individual movement density, average individual movement density and group movement density, and generate a density distribution heat map by gridding the area; S510. Collect teacher classroom voice data and perform real-time speech-to-text processing through dual-source collection and streaming voice processing architecture; S520. Conduct structured analysis of the transcribed language content to identify teaching links, extract key instructions, and assess language quality; S530. Construct a classroom evaluation model based on five core indicators, and analyze the effectiveness of teaching strategies and their correlation with student engagement by combining phonetic transcription; S540. Integrate all analysis results to generate a comprehensive classroom evaluation report and targeted improvement suggestions, providing teachers with specific measures to improve teaching quality.

2. The method for evaluating the multi-objective physical education classroom exercise density under a single viewing angle according to claim 1 is characterized in that: In the target detection and ROI extraction steps, a deep learning model is used to extract image features from a video frame sequence to detect the position and bounding box of the moving target.

3. The method for evaluating the multi-objective physical education classroom exercise density under a single viewing angle according to claim 1 is characterized in that: In the motion density calculation step, the motion state of each target is determined, and the individual motion density is calculated based on the change in joint angle.

4. The method for evaluating the exercise density of a multi-objective physical education class under a single viewing angle according to claim 1, characterized in that: In the multi-target tracking step, the global layer tracker is responsible for maintaining the position, velocity and timestamp information of all targets; The individual-level tracker is responsible for recording the first appearance time, movement period, and movement status of each target; Design identity management mechanisms to ensure target ID continuity.

5. The method for evaluating the exercise density of a multi-objective physical education class under a single viewing angle according to claim 1 is characterized in that: In the evaluation index system step: constructing a five-dimensional evaluation index including individual movement density, average individual movement density, group movement density, thermal distribution index and regional load balance.

6. The method for evaluating the exercise density of a multi-objective physical education class under a single viewing angle according to claim 1, characterized in that: In the teacher language intelligent analysis step, a dual-source acquisition mechanism is used to collect the teacher's classroom voice, including a video file audio channel and real-time microphone input; Build a streaming speech processing architecture to analyze and transcribe teacher language content and generate classroom transcripts with timestamps; Identify teaching links based on natural language processing technology and analyze the rationality of link duration; Construct a teacher language classification model, which divides teacher language into five categories: instructional, feedback, motivational, organizational, and evaluative; Analyze the rationality of teachers' language structure, including the distribution ratio and quality of various languages; Evaluate the conformity of teaching content based on curriculum standards and analyze the degree to which teachers pay attention to individual differences.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, it implements the method for evaluating the multi-target physical education classroom movement density under a single perspective as described in any one of claims 1-6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for evaluating the multi-target physical education classroom movement density under a single perspective described in any one of claims 1 to 6 is implemented.