Attention state detection method, system, device, medium and product
By combining panoramic image detection and a lightweight deep learning model with a single-object tracking algorithm, the problem of insufficient real-time performance and accuracy of attention state detection in traditional detection methods is solved, and real-time and comprehensive detection of object attention states in multi-object scenes is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HONGHE INNOVATION INFORMATION TECH CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional attention state detection methods cannot accurately capture the dynamic changes in an object's attention, nor can they establish a following relationship between objects, resulting in insufficient real-time performance, comprehensiveness, and accuracy of the detection results.
By acquiring panoramic images of the target area, head pose estimation is performed based on the head regions of several objects in the panoramic images. Combined with the spatial location information of a second object, the attention state of the object is determined. A lightweight deep learning model and a single-object tracking algorithm are used for real-time monitoring.
It enables real-time, comprehensive, and accurate detection of the attention state of objects in multi-object scenarios, overcoming the limitations of coverage and accuracy in manual observation, and is suitable for scenarios such as regular classrooms and meetings.
Smart Images

Figure CN121904831A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and more specifically, relates to an attention state detection method, system, device, medium and product. Background Technology
[0002] In scenarios involving multiple participants, such as teaching, meetings, and training, it is necessary to monitor the attention of the audience (e.g., students, attendees) to the core objective (e.g., lecturers, speakers) in real time to ensure the effectiveness of the event.
[0003] Traditional attention state detection methods rely heavily on manual observation, which makes it difficult to cover all audiences in large or multi-object scenarios and cannot accurately capture the dynamic changes in audience attention. They are not only lacking in real-time performance and comprehensiveness, but also unable to establish the following relationship between objects, resulting in inaccurate state judgment. Summary of the Invention
[0004] The purpose of this application is to provide an attention state detection method, system, device, medium, and product, aiming to solve the technical problems that related attention state detection methods cannot accurately capture the dynamic changes in the attention of objects and cannot establish the judgment of the following relationship between objects, resulting in insufficient real-time performance, comprehensiveness, and accuracy of the detection results.
[0005] To achieve the above objectives, according to the first aspect of this application, an attention state detection method is provided, the method comprising: Acquire a panoramic image of the target area, wherein the panoramic image contains at least a number of head regions of a first object and a second object; Based on several head regions of the first object in the panoramic image, the head pose of the first object is estimated to obtain the head orientation information of the first object. Based on the area where the second object is located in the panoramic image, determine the spatial position information of the second object relative to the first object; Based on the head orientation information of the first object and the spatial position information of the second object relative to the first object, the attention state of the first object is determined, wherein the attention state is used to characterize whether the head orientation of the first object follows the movement of the second object.
[0006] According to a second aspect of this application, an attention state detection system is provided for performing the attention state detection method as described in any one of the claims, the system comprising: A visual sensor is disposed within the target area to acquire panoramic images of the target area, wherein the panoramic images include a number of head regions of a first object and the full-body regions of a second object; The processor, connected to the vision sensor, is configured to acquire the panoramic image and, based on the head regions of several first objects in the panoramic image, estimate the head pose of the first objects to obtain head orientation information of the first objects; determine the spatial position information of the second objects relative to the first objects based on the region where the second objects are located in the panoramic image; and determine the attention state of the first objects based on the head orientation information of the first objects and the spatial position information of the second objects relative to the first objects, wherein the attention state is used to characterize whether the head orientation of the first objects follows the movement of the second objects.
[0007] According to a third aspect of this application, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device causes the electronic device to perform the method as described in any one of the claims.
[0008] According to a fourth aspect of this application, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the method as described in any one of the claims.
[0009] According to a fifth aspect of this application, a computer program product is provided that, when run on an electronic device, causes the electronic device to perform the method described in any one of the first aspects above.
[0010] It is understandable that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0011] The beneficial effects of the embodiments in this application compared with the prior art are: By acquiring a panoramic image of the target area, which includes the head regions of all first objects and second objects, simultaneous monitoring of several objects can be achieved, overcoming the limitations of insufficient coverage and accuracy in manual observation. Based on the head regions of several first objects in the panoramic image, head pose estimation is performed to obtain the head orientation information of the first objects. Furthermore, based on the location of the second objects in the panoramic image, the spatial position information of the second objects relative to the first objects is determined. Finally, based on the head orientation information of the first objects and the spatial position information of the second objects relative to the first objects, the attention state of the first objects relative to the second objects is determined, improving the accuracy of determining the attention state of the first objects. Therefore, the embodiments of this application can achieve real-time, comprehensive, and accurate detection of the attention state of first objects within the target area. The detection process does not rely on high-cost specialized equipment and is highly adaptable to various environments, making it particularly suitable for scenarios such as regular classrooms and meetings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic flowchart of an attention state detection method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating an optional attention state detection method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating an optional attention state detection method provided in an embodiment of this application; Figure 4 This is a flowchart illustrating an optional attention state detection method provided in an embodiment of this application; Figure 5 This is a flowchart illustrating an optional attention state detection method provided in an embodiment of this application; Figure 6 This is a flowchart illustrating an optional attention state detection method provided in an embodiment of this application; Figure 7 This is a flowchart illustrating an optional attention state detection method provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an attention state detection device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0014] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0015] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0016] It should also be understood that, in the description of this application, unless otherwise stated, the " / " used in the specification and appended claims indicates that the related objects are in an "or" relationship. For example, A / B can mean A or B. The "and / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0017] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, but are only used for distinguishing descriptions, and the terms "first" and "second" do not necessarily imply that they are different, nor should they be construed as indicating or implying relative importance.
[0018] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0019] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0020] Before describing the specific implementation of this application, it should first be noted that the image acquisition process (such as face image acquisition process, body image acquisition process, etc.) / feature extraction process involved in this application is performed with the knowledge and permission of the user (or the user's guardian). That is, the image acquisition process / feature extraction and recognition process complies with the requirements of laws and regulations and does not constitute an act that harms the public interest.
[0021] In scenarios involving multiple participants, such as teaching, meetings, and training sessions, it is necessary to monitor the attention levels of the audience (e.g., students, attendees) towards the core objective (e.g., lecturers, speakers) in real time to ensure the effectiveness of the event. Traditional monitoring methods rely heavily on manual observation, which is insufficient to cover all audiences in large or multi-participant settings. Furthermore, they cannot accurately capture the dynamic changes in audience attention, lacking both real-time and comprehensiveness, and failing to establish follow-up relationships between participants, leading to inaccurate status assessments.
[0022] Among related technologies, some detection solutions based on behavior analysis and eye tracking either rely on high-cost specialized equipment or have strict requirements on ambient lighting and object posture, limiting their practicality. Other solutions based on image recognition can only make preliminary judgments about the audience's state, lacking accuracy and resulting in inaccurate state determination. Consequently, timely intervention measures cannot be taken, affecting the quality of the activity.
[0023] To address the aforementioned issues, this application provides an example of an attention state detection method. Please refer to [example provided]. Figure 1 As shown, Figure 1 A schematic flowchart of an attention state detection method provided in this application is shown. This is an example and not a limitation; the method can be applied to or run in electronic devices. The method includes: S101, acquire panoramic images of the target area.
[0024] The panoramic image contains at least several head regions of the first object and a second object.
[0025] S102, Based on the head regions of several first objects in the panoramic image, perform head pose estimation on several first objects to obtain the head orientation information of the first objects.
[0026] S103, Based on the area where the second object is located in the panoramic image, determine the spatial position information of the second object relative to the first object.
[0027] S104, based on the head orientation information of each first object and the spatial position information of the second object relative to the first object, determine the attention state of each first object.
[0028] The attention state is used to characterize whether the head orientation of the first object follows the movement of the second object.
[0029] In some embodiments, the attention state detection method provided in this application can be applied, but is not limited to, teaching scenarios (target area is a classroom, first object is a student, second object is a lecturer), meeting scenarios (target area is a meeting room, first object is a participant, second object is a speaker), corporate internal training scenarios (target area is a corporate training room, first object is a trainee, second object is a training lecturer / demonstration equipment), vocational skills training scenarios (such as driving training, equipment operation training, target area is a driving simulator / practical training workshop, first object is a student, second object is an instructor / practical demonstrator / equipment operation guidance sign), customer service training / service scenarios (target area is a customer service training room / service workstation area, first object is a customer service person (training period / on-the-job period), second object is a trainer / customer, etc., and is not limited to this example.
[0030] The attention state detection method provided in this application will be described in detail below with reference to a specific application scenario. This embodiment takes a classroom teaching scenario as a typical application scenario (other application scenarios are the same or similar). The target area is the classroom, the first object is the students in the classroom (the number of the first object can be one or more, which can be adjusted and set according to actual needs), and the second object is the lecturer. This method realizes real-time monitoring of students' attention state in the classroom. The specific implementation process is as follows: First, a wide-angle depth camera can be pre-deployed as a visual sensor in the target area, namely the top center of the front of the classroom (in actual use, it is not limited to one camera; one or more cameras can also be set up at the back and sides of the classroom). The field of view of this camera should be no less than 120 degrees, the resolution should be set to 1920×1080 pixels, and the frame rate should be 30 frames per second to ensure complete coverage of the seating areas of all students and the activity area of the lecturer (including the podium and the area in front of the classroom).
[0031] After the user (academic staff) activates the wide-angle depth camera via command or button, the camera captures panoramic images of the classroom in real time. During the capture process, image quality is optimized through functions such as automatic exposure and white balance adjustment. The captured panoramic images must clearly show the head area of each student (including facial features and head outline) as well as the lecturer's complete figure and location. The panoramic images are transmitted in real time to the back-end processing terminal (such as the networked monitoring equipment of the academic affairs office, specifically an industrial computer or edge computing device) in the form of digital signals.
[0032] Next, after receiving the panoramic image, the backend processing terminal first performs preprocessing operations on the panoramic image, including but not limited to grayscale correction, Gaussian filtering for noise reduction, and image scale normalization, to eliminate the interference of ambient light changes and image noise on subsequent processing. Subsequently, the backend processing terminal calls a preset deep learning model (such as a lightweight deep learning model) to perform recognition processing on the preprocessed panoramic image. This deep learning model uses MobileNetV3 as the backbone network and combines it with a 68-point facial landmark detection algorithm to identify and locate the pixel coordinate range of each student's head region, thus clarifying the specific position of each student's head in the panoramic image.
[0033] Based on the head region obtained from localization, the coordinate information of facial feature points (including key points such as the corners of the eyes, the tip of the nose, and the corners of the mouth) within this head region can be further extracted. By solving the three-dimensional head posture equation, three posture parameters of the student's head—pitch angle, yaw angle, and / or roll angle—are calculated. The pitch angle represents the vertical rotation angle of the head, the yaw angle represents the horizontal rotation angle, and the roll angle represents the lateral roll angle. Based on these three posture parameters, a head orientation vector for each student is constructed using a spatial vector transformation algorithm. The direction of this head orientation vector is consistent with the student's actual head orientation, and this is used as the head orientation information for the first object. It should be noted that in this embodiment, the inference delay of the posture estimation process is controlled within 50-100 milliseconds to meet the requirements of real-time monitoring.
[0034] Then, the backend processing terminal synchronously calls the human detection model to detect the second object in the panoramic image, namely the lecturer / teacher in this example. This human detection model uses a lightweight version of YOLOv5s. By recognizing the lecturer's human contour and limb features in the panoramic image, it locates the lecturer's initial two-dimensional coordinates (i.e., pixel coordinates) in the panoramic image and marks the bounding box of the lecturer's entire body area. To achieve continuous tracking of the lecturer's movement trajectory, after detecting the initial two-dimensional coordinates, a single-target tracking algorithm (using the kernel correlation filter KCF tracking algorithm) is started. Based on feature matching of the lecturer's entire body area, the two-dimensional pixel coordinates of the lecturer in each frame image are updated in real time, ensuring that the tracking process is not interrupted even when the lecturer moves or turns around in the podium area.
[0035] Simultaneously, using depth data collected by a wide-angle depth camera, combined with a pre-defined classroom spatial coordinate system (e.g., establishing a three-dimensional coordinate system with the bottom left corner of the classroom floor as the origin, the horizontal direction to the right as the X-axis, the vertical direction forward as the Y-axis, and the vertical direction upward as the Z-axis), the lecturer's two-dimensional pixel coordinates are converted into three-dimensional spatial coordinates to obtain the lecturer's actual position in the classroom. For each student, using the center point of the student's head region as a reference point, the spatial distance and direction angle between the lecturer's three-dimensional spatial coordinates and this reference point are calculated, forming the lecturer's spatial position information relative to the student. This spatial position information can be stored in the form of a spatial position vector, with the starting point of the spatial position vector being the center point of the student's head region and the ending point of the spatial position vector being the lecturer's position.
[0036] Finally, based on the head orientation information of the first object and the spatial position information of the second object relative to the first object, the attention state of the first object is determined. For each student, the backend processing terminal calls the vector operation module to calculate the angle between the student's head orientation vector and the corresponding lecturer's spatial position vector. For example, the angle between the two vectors can be calculated using the dot product formula: θ = arccos[(a b) / (|a|×|b|)], where a is the head orientation vector, b is the spatial position vector, and θ is the angle between the two vectors.
[0037] In this embodiment, the preset angle threshold can be determined through multiple classroom experiments. For example, it can be 60 degrees, used to accurately distinguish whether the student's head is facing the lecturer. If the calculated angle θ is less than the preset angle threshold, i.e., 60 degrees, it indicates that the student's head orientation is consistent with the lecturer's spatial position relative to the student, meaning the student's head orientation follows the movement of the second object. Therefore, the attention state of the first object is determined to be a state of focused attention. If the angle θ is greater than or equal to the preset angle threshold, i.e., 60 degrees, it indicates that there is a significant deviation between the student's head orientation and the lecturer's spatial position, meaning the student's head orientation does not follow the movement of the second object. Therefore, the attention state of the first object is determined to be a state of unfocused attention.
[0038] It should be noted that the above-mentioned angle calculation and state judgment process is executed synchronously and in real time with the image acquisition, pose estimation, and position tracking processes. The processing cycle of each frame does not exceed 100 milliseconds, ensuring that the changes in each student's attention state can be dynamically captured.
[0039] The attention state detection method provided in this application acquires a panoramic image of the target area, which includes the head regions of all first objects and second objects. This allows for the simultaneous monitoring of multiple first objects, overcoming the limitations of insufficient coverage and accuracy in manual observation. By estimating the head pose of multiple first objects based on their head regions in the panoramic image, the head orientation information of the first objects is obtained. Furthermore, based on the location of the second objects in the panoramic image, the spatial position information of the second objects relative to the first objects is determined. Finally, based on the head orientation information of the first objects and the spatial position information of the second objects relative to the first objects, the attention state of the first objects is determined. Correlation analysis is performed between the head orientation information of the first objects and the spatial position information of the second objects relative to the first objects to quantify the spatial following relationship between the first and second objects, significantly improving the accuracy of determining the attention state of the first objects. Therefore, it can achieve real-time, comprehensive, and accurate detection of the attention state of first objects within the target area. The detection process does not rely on high-cost professional equipment and is highly adaptable to various environments, making it particularly suitable for scenarios such as regular classrooms and meetings.
[0040] In some embodiments, such as Figure 2 As shown, based on the head regions of several first objects in the panoramic image, head pose estimation is performed on several first objects to obtain the head orientation information of the first objects, including: S201, a deep learning model is used to identify the head regions of several first objects in the panoramic image to obtain the pixel coordinate range of the head regions of the first objects.
[0041] S202, extract features from the image data corresponding to the pixel coordinate range of each head region to determine the pitch angle, yaw angle and / or roll angle of the head of the first object.
[0042] S203, Based on the pitch angle, yaw angle and / or roll angle of the first object's head, construct the head orientation vector of the first object as head orientation information.
[0043] In some embodiments, firstly, a deep learning model can be used to identify the head regions of several first objects in the panoramic image to obtain the pixel coordinate range of the head regions of the first objects. After receiving the preprocessed panoramic image (e.g., grayscale correction, Gaussian filtering for noise reduction, and scale normalization), the background processing terminal calls a preset deep learning model to perform head region identification on the first objects in the preprocessed panoramic image.
[0044] It should be understood that the deep learning model may, but is not limited to, use MobileNetV3 as the feature extraction backbone network, combined with an improved 68-point facial landmark detection algorithm, and be trained on public face datasets and application scenario-specific face datasets (such as classroom scenarios) through transfer learning. The number of parameters of the trained model is controlled within 5MB to ensure fast inference on edge computing devices.
[0045] During the model inference process, a multi-scale feature pyramid is first constructed on the panoramic image. Feature fusion is used to enhance the recognition ability of small head regions. Then, each possible head candidate region is located based on the anchor point matching strategy. Then, the overlapping candidate regions are removed by the non-maximum suppression algorithm. Finally, the pixel coordinate range of the first object's head region is output. Taking the upper left corner of the image as the origin, the upper left corner pixel coordinates (x1, y1) and the lower right corner pixel coordinates (x2, y2) of the bounding box of the head region are output to form a closed pixel coordinate range, ensuring that the head region of the first object is accurately selected and does not include the background or other object interference regions.
[0046] Next, feature extraction is performed on the image data corresponding to the pixel coordinate range of each head region to determine the pitch angle, yaw angle, and / or roll angle of the first object's head. For the head region of the first object, the backend processing terminal crops the local image corresponding to the pixel coordinate range from the panoramic image, scales the local image to a uniform size (e.g., 224×224 pixels), and then inputs it into the feature extraction branch of the aforementioned deep learning model. This feature extraction branch extracts deep features of the head region of the first object through a combination of convolutional layers, batch normalization layers, and activation functions (Hard-Swish), including the coordinate information of facial key points (corners of the eyes, brow bones, tip of the nose, alar of the nose, corners of the mouth, jaw angle, etc.).
[0047] Based on the extracted coordinates of 68 facial key points, a pose calculation algorithm, such as the Perspective-n-Point (EPnP) algorithm, is employed. Combined with a pre-defined 3D face model template (containing the 3D coordinates of the key points), a projection relationship between pixel coordinates and 3D coordinates is constructed. The rotation matrix of the first object's head is then solved using the least squares method. By decomposing the rotation matrix of the first object's head into Euler angles, the pitch, yaw, and / or roll angles of the first object's head are obtained.
[0048] In some embodiments, the pitch angle ranges from -90° to 90°, representing the angle of head rotation up and down (positive for tilting up, negative for tilting down); the yaw angle ranges from -90° to 90°, representing the angle of head rotation left and right (positive for left turn, negative for right turn); and the roll angle ranges from -180° to 180°, representing the angle of head roll (positive for left roll, negative for right roll). In the embodiments of this application, the calculation accuracy of the three angles is controlled within ±2° to ensure the accuracy of the head posture description.
[0049] Finally, based on the pitch angle, yaw angle, and / or roll angle of the first object's head, a head orientation vector of the first object is constructed as head orientation information. A local three-dimensional coordinate system is established with the center of the first object's head region (the center point calculated from the pixel coordinate range (x1, y1) and (x2, y2) ((x1+x2) / 2, (y1+y2) / 2)) as the origin, i.e., the positive Z-axis is the direction of the front of the head, the positive X-axis is the right side of the head, and the positive Y-axis is the top of the head.
[0050] Based on the pitch, yaw, and / or roll angles obtained above, the Z-axis (initial frontal direction) of the local coordinate system is transformed into the direction vector corresponding to the actual head orientation through a rotation matrix transformation. Specifically, the rotation matrices corresponding to the three angles are multiplied sequentially to obtain the total rotation matrix. Then, the initial Z-axis unit vector (0,0,1) is multiplied by the total rotation matrix to obtain the transformed three-dimensional vector, which is the head orientation vector of the first object. The head orientation vector is stored in three-dimensional coordinate form (vx,vy,vz), where vx,vy, and vz are the components of the vector on the X, Y, and Z axes, respectively. The vector is normalized (magnitude is 1) to ensure the accuracy of subsequent angle calculations with the spatial position vector of the second object. This head orientation vector directly serves as the head orientation information representing the actual head orientation of the first object.
[0051] This embodiment uses a deep learning model and pose estimation algorithm to significantly improve processing speed while ensuring the accuracy of head region recognition and pose estimation, thus meeting the needs of real-time monitoring, and avoiding dependence on high-configuration hardware.
[0052] In some embodiments, such as Figure 3 As shown, based on the region where the second object is located in the panoramic image, the spatial position information of the second object relative to the first object is determined, including: S301, a human detection model is used to detect the region where the second object is located in the panoramic image in order to determine the contour features of the second object.
[0053] S302, a single-target tracking algorithm is used to continuously track the contour features of the second object in order to synchronously update the real-time two-dimensional coordinates of the second object.
[0054] S303, based on the depth data collected by the visual sensor within the target area, converts the real-time two-dimensional coordinates of the second object into three-dimensional spatial coordinates.
[0055] S304, calculate the spatial distance and direction angle of the two objects relative to the center point of the head region of the first object in three-dimensional space, and obtain the spatial position information of the second object relative to the first object.
[0056] In some embodiments, firstly, a human detection model is used to detect the region where the second object is located in the panoramic image to determine the contour features of the second object. After receiving the panoramic image, the backend processing terminal first performs preprocessing operations on the image, including image size scaling to adapt to the model input requirements, color gamut normalization, and edge enhancement to improve the recognition contrast of the second object region. Then, the backend processing terminal calls a preset human detection model (such as a lightweight human detection model) to perform the detection task. This human detection model adopts the YOLOv5s lightweight architecture. As an example and not a limitation, the parameter size can be controlled within 8MB through pruning optimization, and the inference speed is not less than 50 frames / second to meet the real-time detection requirements.
[0057] Taking the application of this method in a classroom setting as an example, during the training of the human detection model, a publicly available human dataset and a classroom-specific human dataset are integrated. The latter includes samples of various postures of the lecturer, such as standing, walking, and turning, ensuring a high detection rate for the second object in a classroom environment. During the detection process, the human detection model extracts image features at different scales through a feature pyramid network, locates candidate regions of the second object based on anchor point matching, and outputs the bounding box coordinates of the region where the second object is located after removing redundant boxes using a non-maximum suppression algorithm. These bounding box coordinates are based on the upper left corner of the panoramic image as the origin and use a combination of the minimum and maximum values of two-dimensional pixel coordinates to identify the bounding box range. Simultaneously, through the feature extraction branch of the human detection model, the contour features of the second object are extracted, including the contour curve coordinates of the shoulders, torso, and legs, as well as the pixel coordinates of key limb points such as the acromion, hip center, and ankle, forming a feature set representing the shape features of the second object, providing a stable feature basis for subsequent tracking.
[0058] Secondly, a single-target tracking algorithm is used to continuously track the contour features of the second object to synchronously update its real-time two-dimensional coordinates. After obtaining the initial contour features and bounding box coordinates of the second object, the single-target tracking algorithm is started (in this embodiment, an improved algorithm combining kernel correlation filter tracking algorithm and deep learning features is used as the single-target tracking algorithm). The contour features obtained from the previous detection are used as the tracking template. In each subsequent frame of the panoramic image, the search area is delineated with the initial bounding box as the center. The correlation response value between each candidate area within the search area and the template features is calculated. The candidate area with the largest response value is determined as the current position of the second object.
[0059] To address scenarios involving movement, turning, or partial occlusion of the second object, the single-object tracking algorithm updates the tracking template in real time. For example, when the detected contour feature matching degree is higher than a preset threshold (e.g., 85%), a moving average strategy is used to update the template weights to maintain tracking stability. When the detected contour feature matching degree is lower than the preset threshold (e.g., 85%), the human detection model is triggered to re-detect, quickly recovering the position of the second object in the panoramic image and resetting the tracking template to avoid tracking drift. During the tracking of the second object, each frame outputs the real-time two-dimensional pixel coordinates of the second object. These real-time two-dimensional pixel coordinates use the center point of the bounding box as the core representation and simultaneously record the aspect ratio of the bounding box to ensure accurate reflection of the real-time position of the second object.
[0060] Next, based on the depth data collected by the visual sensor within the target area, the real-time two-dimensional coordinates of the second object are converted into three-dimensional spatial coordinates. The visual sensor deployed in the target area is a wide-angle depth camera, which simultaneously outputs the depth data of the corresponding pixels while acquiring panoramic images. It should be understood, but not limited to, that in this embodiment, the sampling accuracy of the depth data is not less than 1 mm, and the effective measurement distance range is 0.3-10 meters, adapting to the spatial scale of a classroom scene. The backend processing terminal reads the depth data collected by the visual sensor in real time through the sensor interface, and extracts the depth value of the pixel corresponding to the real-time two-dimensional pixel coordinates of the second object. Combining the intrinsic parameters of the visual sensor (including the horizontal focal length, vertical focal length, and principal point coordinates), coordinate conversion is performed through the camera imaging model. A three-dimensional world coordinate system is established with the optical center of the visual sensor as the origin, with the horizontal direction to the right as the X-axis, the vertical direction downward as the Y-axis, and the direction perpendicular to the image plane forward as the Z-axis.
[0061] When performing coordinate conversion using the camera imaging model, the difference between the two-dimensional pixel coordinates of the second object and the principal point coordinates is first calculated. Then, the horizontal difference is multiplied by the depth value and divided by the horizontal focal length to obtain the X-axis component of the three-dimensional spatial coordinates. The vertical difference is multiplied by the depth value and divided by the vertical focal length to obtain the Y-axis component of the three-dimensional spatial coordinates. The depth value is directly used as the Z-axis component of the three-dimensional spatial coordinates. During the conversion process, a coordinate calibration algorithm is used to eliminate errors caused by sensor installation deviations, ensuring that the three-dimensional spatial coordinates can accurately map the actual spatial position of the second object within the target area.
[0062] Finally, the spatial distance and orientation angle of the second object relative to the center point of the head region of the first object are calculated to obtain the spatial position information of the second object relative to the first object. For the first object, the pixel coordinate range of the head region of the first object is first obtained, the pixel coordinates of the center point of the head region are calculated, and then, based on the depth data and intrinsic parameters of the vision sensor, the same method as the three-dimensional coordinate conversion of the second object is used to obtain the three-dimensional spatial coordinates of the center point of the head region.
[0063] In some embodiments, when calculating the spatial distance, the differences between the three-dimensional spatial coordinates of the second object and the three-dimensional spatial coordinates of the center point of the head region of the first object in the X-axis, Y-axis and Z-axis directions are first determined. The three differences are squared and summed respectively. Then the summation result is squared to obtain the straight-line distance between the two. This straight-line distance represents the actual spatial distance between the second object and the corresponding first object.
[0064] In some embodiments, when calculating directional angles, a local coordinate system is established with the center point of the first object's head as the origin. The horizontal direction to the right is the X' axis, the vertical direction upward is the Y' axis, and the direction pointing forward to the target area is the Z' axis. The three-dimensional spatial coordinates of the second object are converted into relative coordinates under this local coordinate system. Based on the relative coordinates, two key directional angles are calculated: one is the azimuth angle, which is the angle between the projection of the relative coordinates onto the plane formed by the X' and Z' axes and the positive direction of the Z' axis, used to characterize the left and right orientation of the second object relative to the first object; the other is the pitch angle, which is the angle between the relative coordinates and the plane formed by the X' and Z' axes, used to characterize the up and down orientation of the second object relative to the first object.
[0065] The spatial distance, azimuth angle, and pitch angle calculated above are integrated to form the spatial position information of the second object relative to the first object. This spatial position information can be stored in the form of parameter combination, or further converted into a spatial position vector pointing from the center point of the first object's head to the second object, providing spatial parameter support for the subsequent judgment of the attention state of the first object.
[0066] In this embodiment, by combining detection and tracking algorithms with precise conversion of depth data, the spatial location of the second object can be located in real time and accurately. Furthermore, personalized spatial relationship calculations are performed on the first object to ensure the relevance and accuracy of the spatial location information. This provides key technical support for the reliability of attention state judgment, while also taking into account the lightweight nature and real-time performance of the algorithm.
[0067] In some embodiments, such as Figure 4 As shown, based on the head orientation information of the first object and the spatial position information of the second object relative to the first object, the attention state of the first object is determined, including: S401, convert the head orientation information of the first object into a spatial direction vector, and convert the spatial position information of the second object relative to the first object into a spatial position vector.
[0068] S402 calculates the angle between the spatial direction vector and the spatial position vector.
[0069] S403, if the included angle value is less than the preset angle threshold, then the attention state of the first object is determined to be the attention concentration state.
[0070] S404, if the included angle value is greater than or equal to the preset angle threshold, then the attention state of the first object is determined to be an inattentive state.
[0071] In some embodiments, firstly, the head orientation information of the first object is converted into a spatial direction vector, and the spatial position information of the second object relative to the first object is converted into a spatial position vector. For the first object, the head orientation information has been specifically converted into a head orientation vector through the preceding steps. This head orientation vector takes the center point of the head region of the first object as its origin and represents the three-dimensional direction of the actual head orientation. To adapt to the subsequent angle calculation, this head orientation vector needs to be converted into a spatial direction vector under a unified coordinate system: taking the center point of the head region of the first object as the local origin, following the three-dimensional spatial coordinate system rules of the target region (horizontal to the right is the X-axis, vertical upward is the Y-axis, and pointing forward of the target region is the Z-axis), the head orientation vector is mapped to ensure that the vector direction is completely consistent with the actual head orientation of the first object. At the same time, the vector is normalized so that the vector magnitude is uniformly 1, eliminating the influence of length differences on the angle calculation.
[0072] The spatial position information of the second object relative to the first object includes parameters such as spatial distance, azimuth, and pitch angle, or it may already be stored as relative coordinates originating from the center point of the first object's head. When converting this spatial position information into a spatial position vector, the center point of the first object's head region is used as the local origin. The horizontal direction of the vector in the plane formed by the X and Z axes is determined based on the azimuth angle, and the vertical height of the vector in the Y-axis direction is determined based on the pitch angle. Combining the axial components in the relative coordinates, a three-dimensional vector is constructed pointing from the center point of the first object's head to the three-dimensional spatial coordinates of the second object. This three-dimensional vector is the spatial position vector, which is also normalized to ensure consistency with the calculation basis of the spatial direction vector.
[0073] Secondly, the backend processing terminal can calculate the angle between the spatial direction vector and the spatial position vector. During the calculation, algorithms representing the consistency of vector directions are used, combined with geometric principles, to deduce the degree of directional deviation between the two vectors. For example, the components of the spatial direction vector and the spatial position vector on the X, Y, and Z coordinate axes can be obtained first. The directional similarity between the two vectors is calculated through the correspondence between the components, and then the similarity is converted into a specific angle value.
[0074] It should be noted that the entire calculation process described above does not rely on complex formulas. It quickly outputs results through a pre-defined vector angle mapping model, ensuring calculation accuracy is controlled within ±1°, and the calculation time for a single vector does not exceed 5 milliseconds, meeting the processing speed requirements of real-time monitoring. This angle value intuitively reflects the degree of deviation between the orientation of the first object's head and the direction of the second object. The smaller the angle value, the more consistent the directions of the two objects; the larger the angle value, the more significant the deviation.
[0075] Finally, the attention state of the first object is determined based on the comparison between the included angle value and the preset angle threshold. The preset angle threshold was determined through multiple classroom scenario tests and is used to distinguish whether the first object's head is facing the second object. For example, a value of 60 degrees can be used. This is because when the first object's head is facing the direction of the second object, the included angle value is usually less than 60 degrees; when the first object's head turns to the side, back, or looks down or up in a direction other than the direction of the second object, the included angle value will be greater than or equal to 60 degrees.
[0076] If the calculated angle is less than 60 degrees, it indicates that the head orientation of the first object is consistent with the spatial position of the second object relative to the first object, meaning that the head orientation of the first object follows the movement of the second object. Therefore, the attention state of the first object is determined to be a state of focused attention. If the angle is greater than or equal to 60 degrees, it indicates that the head orientation of the first object deviates significantly from the spatial position of the second object, meaning that the head orientation of the first object does not follow the movement of the second object. Therefore, the attention state of the first object is determined to be a state of unfocused attention.
[0077] It should be noted that the judgment process is executed synchronously and in real time with the preceding vector transformation and angle calculation. The attention state judgment of all first objects in each frame is completed within 10 milliseconds, ensuring that the instantaneous changes in the attention state of the first object can be dynamically captured, providing accurate and real-time status data support for subsequent attention distribution statistics and early warning feedback.
[0078] In some embodiments, the attention state of the first object includes: a focused attention state when the head of the first object is facing the direction of moving with the second object; and a disfocused attention state when the head of the first object is not facing the direction of moving with the second object.
[0079] like Figure 5 As shown, there are multiple first objects. After determining the attention state of each first object based on its head orientation information and the spatial position information of the second object relative to the first object, the method further includes: S501, Based on the attention state of each first object, determine the first number of first objects in the target area that are in a state of focused attention.
[0080] S502, Based on the attention state of each first object, determine the distribution location of the first object in the target area when it is in an inattentive state.
[0081] S503, Based on the distribution location and the proportion of the first quantity to the total number of first objects, obtain the attention state distribution data of all first objects.
[0082] S504, based on the attention state distribution data of all first objects and the attention state of each first object, trigger the corresponding warning operation.
[0083] In some embodiments, the attention state of the first object includes a focused attention state and a disfocused attention state, wherein the focused attention state corresponds to the first object’s head moving in the direction of the second object, and the disfocused attention state corresponds to the first object’s head not moving in the direction of the second object.
[0084] In some embodiments, after determining the attention state of each first object, the background processing terminal can determine the first number of first objects in the target area that are in a state of focused attention based on the attention state of each first object. For example, the background processing terminal can iterate through the attention state identifiers of all first objects in real time (each first object corresponds to a unique state identifier, which is associated with a focused or unfocused state respectively), count the first objects identified as being in a focused state, and obtain the first number.
[0085] Meanwhile, the back-end processing terminal determines the total number of first objects based on the overall detection results of the first objects in the target area during the preceding image recognition process (the total number is the total number of all first objects actually monitored in the target area, which can be dynamically updated based on the recognition results of the head area in the panoramic image to adapt to temporary increases or decreases in personnel).
[0086] Secondly, based on the attention state of each first object, the distribution location of the first object in an inattentive state within the target area is determined. For each first object identified as inattentive, the terminal retrieves the three-dimensional spatial coordinates of the center point of the head region of the inattentive first object (these coordinates are obtained from the preceding head pose estimation and spatial coordinate conversion steps, which can accurately map the actual position of the first object within the target area). Combining the spatial division rules of the target area (e.g., dividing the classroom into several rows and columns according to the seating arrangement, or dividing it into front, middle, and back areas), the three-dimensional spatial coordinates of each inattentive first object are converted into actual position markers within the target area (e.g., "3rd row, 4th column", "back row, left side area", etc.).
[0087] Meanwhile, the backend processing terminal records the unique number of each first object in a state of inattention, associates and stores this unique number with the actual location identifier of the corresponding first object in a state of inattention, and forms a distribution location list of the first objects in a state of inattention, ensuring that the location of each first object in a state of inattention can be accurately traced.
[0088] Next, based on the aforementioned distribution locations and the proportion of the first quantity to the total number of first objects, attention state distribution data for all first objects is obtained. On one hand, the backend processing terminal calculates the ratio of the first quantity to the total number, converting it to a percentage using division (e.g., if 18 out of 30 first objects are in a concentrated state, the proportion is 60%). This percentage directly reflects the overall proportion of concentrated attention among first objects within the target area. On the other hand, the backend processing terminal converts the list of distribution locations of non-concentrated first objects into a visual data format (such as structured data containing location coordinates and area identifiers). Integrating the percentage data with the visualized distribution data yields attention state distribution data. This data includes both statistical information on the overall level of attention concentration and the location details of individual non-concentrated objects, providing dual data support for subsequent overall control and precise intervention.
[0089] Finally, the backend processing terminal triggers corresponding alert operations based on the attention state distribution data of all first objects and the attention state of each first object. For example, the backend processing terminal can preset two types of alert triggering conditions, corresponding to the overall attention situation and the attention state of a single first object, respectively.
[0090] For each individual first object, the backend processing terminal records the duration of each first object's inattentive state in real time (the timer starts when the first object is first identified as inattentive; if the state switches to attentiveness midway, the timer is reset), and presets a duration threshold (e.g., 30 seconds, which can be flexibly adjusted according to the actual application scenario). If the duration of any first object's inattentive state exceeds the preset duration threshold, the backend processing terminal triggers an alert operation, specifically including two forms of alerts: first, outputting an audible alert signal, such as issuing a soft prompt through a speaker in the target area, or providing a prompt through the second object's headset, to avoid disturbing the overall environment; second, generating visual alert information, such as highlighting the current location of the inattentive first object with a red border or flashing icon in the visual display interface of the backend processing terminal, while also displaying the duration of the first object's inattentive state.
[0091] In some embodiments, for the attention state distribution data of all first objects, the backend processing terminal pre-sets a preset proportion threshold (e.g., 50%, determined through multiple scenario verifications to adapt to the attention management needs of scenarios such as classrooms). If the proportion of the first number to the total number is lower than the preset proportion threshold, it indicates that the overall attention concentration of all first objects in the target area is low, and the backend processing terminal also triggers an early warning operation. At this time, the visual early warning information will be presented in the visual display interface of the backend processing terminal in the form of a global prompt (e.g., a pop-up window at the top displays "The attention concentration ratio of the first objects in the target area is too low"), and at the same time, an audio early warning signal will be output to remind relevant personnel to pay attention to the overall status of the first objects in the target area.
[0092] In some embodiments, after the warning operation is triggered, the background processing terminal can also synchronize the attention state distribution data and the visual warning information to the associated target terminal (such as the teacher's computer or tablet terminal) in real time, so as to ensure that relevant personnel can obtain the overall attention distribution and individual warning object information in a timely manner, and provide accurate guidance for the implementation of subsequent intervention measures.
[0093] In some embodiments, such as Figure 6 As shown, based on the attention state distribution data of all first objects and the attention state of each first object, corresponding early warning operations are triggered, including: S601, based on the attention state of each first object, determine the first duration for which any first object remains in an inattentive state.
[0094] S602, if the first duration exceeds the preset duration threshold, then trigger the output of an audible warning signal and / or generate a visual warning message.
[0095] S603 generates personalized intervention suggestions, including interactive guidance prompts and / or seating adjustment suggestions, for each first subject in an inattentive state, based on the first duration and distribution location.
[0096] S604 displays real-time attention state distribution data, visual warning information, and personalized intervention suggestions in the visualization interface, and highlights the first subject in an inattentive state by highlighting or special marking.
[0097] In some embodiments, a corresponding warning operation is triggered based on the attention state distribution data of all first objects and the attention state of each first object. The specific implementation process is as follows: First, based on the attention state of each first object, the first duration for which any first object remains in an inattentive state is determined. The backend processing terminal may, but is not limited to, assign an independent timing module to each first object, and this timing module is linked in real-time with the attention state determination result of the first object. When a first object is first determined to be in an inattentive state, the timing module immediately starts timing and records the start time of this state; if, during the timing process, the attention state of the first object switches to an attentive state, the timing module automatically resets, stops the current timing, and clears the timing data; if the attention state of the first object remains in an inattentive state, the timing module continuously accumulates the timing duration to form the first duration for that first object.
[0098] In some embodiments, the timing module can have a timing accuracy at the second level to ensure the accuracy of duration statistics. At the same time, it can associate and store the first duration of each first object with the corresponding object identifier in real time, which is convenient for subsequent threshold comparison and generation of personalized intervention suggestions.
[0099] Secondly, if the first duration exceeds a preset duration threshold, an audible warning signal and / or a visual warning message will be triggered. It should be understood that the preset duration threshold can be flexibly set according to the needs of the actual application scenario; for example, it can be set to 30 seconds in a classroom teaching scenario. This preset duration threshold can be manually adjusted by the user through the target terminal. The backend processing terminal compares the first duration of each first object with the preset duration threshold in real time. When the first duration of any first object exceeds the threshold, the backend processing terminal triggers the corresponding warning operation according to the preset warning mode. Optionally, the above warning operation includes, but is not limited to, the following three optional forms: outputting only an audible warning signal, generating only a visual warning message, and outputting both an audible warning signal and generating a visual warning message simultaneously. The audible warning signal can be output through audio equipment deployed in the target area, using a gentle and non-disturbing tone (such as a short tone) to avoid triggering resistance from the first object. The visual warning message can be generated in the form of graphic symbols, such as red borders or flashing icons, to accurately indicate the corresponding inattentive first object.
[0100] Next, for each individual experiencing inattention, personalized intervention suggestions are generated based on the duration of inattention and their location, including interactive guidance prompts and / or seating adjustment suggestions. The backend processing terminal uses preset intervention suggestion generation rules, combining different duration intervals and location characteristics to generate differentiated suggestions. For example, if the duration is within a preset short range (e.g., 30 seconds to 1 minute), and the individual is located in the front or middle of the target area, the generated interactive guidance prompts include: asking the individual a question in class, inviting them to participate in classroom discussions, etc., to guide them to quickly regain focus. Conversely, if the duration is within a preset long range (e.g., more than 1 minute), or the individual is located in the back or edge of the target area, in addition to interactive guidance prompts, seating adjustment suggestions are generated, such as suggesting moving the individual to a front-row middle seat, suggesting placing the individual next to an individual who is focused, etc., to help improve the individual's attention through environmental adjustments.
[0101] For multiple adjacent individuals simultaneously experiencing inattention, targeted suggestions are generated to disperse and adjust the seating arrangements of these individuals, avoiding mutual interference. It should be noted that all personalized intervention suggestions are generated in concise, actionable text format and stored in association with the corresponding individual identifier and location.
[0102] Finally, the visualization interface displays real-time attention distribution data, visual warning information, and personalized intervention suggestions, highlighting the first individual exhibiting inattention with a high-key or special identifier. The visualization interface is deployed on a backend processing terminal (such as a teacher's computer or tablet). For example, it can be divided into three functional areas: an attention distribution data display area, a warning information and individual identifier area, and a personalized intervention suggestion area.
[0103] In some embodiments, in the attention state distribution data display area, statistical charts (such as pie charts) are used to show the percentage of the number of objects in the attention-focused state, and regional heat maps or seating distribution maps are used to show the distribution of the objects in the inattentive state, intuitively presenting the overall attention distribution; in the warning information and object identification area, the generated visual warning information (such as red borders and flashing icons) is superimposed on the corresponding object position in the seating distribution map, and highlighted or specially marked (such as different colors to distinguish different first duration intervals), allowing relevant personnel to quickly locate the warning object; in the personalized intervention suggestion area, personalized intervention suggestions are displayed in order of the warning priority of the first object (the longer the first duration and the more peripheral the position, the higher the priority), and each suggestion item is associated with the corresponding object identification and distribution position. Clicking on an item will simultaneously highlight the corresponding object in the seating distribution map.
[0104] It should be noted that the data in the visualization interface is updated in real time, with the update frequency consistent with the attention state determination frequency (updated synchronously after each frame of image processing is completed), ensuring that relevant personnel can obtain the latest attention distribution, early warning information and intervention suggestions in real time, providing intuitive and comprehensive support for the rapid implementation of precise interventions.
[0105] In some embodiments, such as Figure 7 As shown, based on the distribution location and the proportion of the first quantity to the total number of first objects, the attention state distribution data of all first objects is obtained, including: S701, based on the proportion of the first quantity to the total number of the first objects, obtain the percentage statistics of the first objects that are in a state of focused attention.
[0106] S702, based on the distribution location, obtain the region coordinate annotation data of the first object in a state of inattention.
[0107] S703 generates attention state distribution data in the form of statistical charts or heatmaps based on percentage statistics and regional coordinate annotation data.
[0108] Among them, the attention state distribution data is used to characterize the overall distribution of the attention state of the first object within the target area.
[0109] In some embodiments, firstly, based on the ratio of the first quantity to the total number of first objects, a percentage statistical data of the first objects in a state of focused attention is obtained. The background processing terminal retrieves the first quantity (i.e., the total number of first objects in a state of focused attention within the target area) and the total number of first objects (i.e., the total number of all first objects actually monitored within the target area) statistically from the previous steps, and obtains the percentage values of the two through proportional conversion.
[0110] During the conversion process, the first quantity can be used as the numerator, and the total number of the first objects as the denominator for division. The result is then multiplied by 100 to convert it into a percentage. For example, if the total number of the first objects in the target area is 40 and the first quantity is 28, the corresponding percentage is 70%. This percentage statistic intuitively quantifies the overall proportion of attention concentration among the first objects in the target area, clearly reflecting the overall level of attention concentration among all the first objects in the target area, providing data support for relevant personnel to quickly grasp the overall level of attention in the target area.
[0111] Secondly, based on the distribution locations of the first objects in an inattentive state, the region coordinate annotation data of this type of first object is obtained. For each first object determined to be in an inattentive state, the terminal retrieves the three-dimensional spatial coordinates of the center point of the head region of the first object (these coordinates are obtained from the previous spatial coordinate conversion step, mapping the actual position of the first object within the target region), and combines them with the pre-set standardized coordinate system of the target region (for example, a two-dimensional plane coordinate system with the lower left corner of the target region as the origin, establishing a horizontal to the right as the X-axis and a vertical forward as the Y-axis, or a three-dimensional coordinate system containing height information), converting the three-dimensional spatial coordinates into unified region coordinates.
[0112] Simultaneously, the backend processing terminal assigns a unique identifier to the first object in each inattentive state, and associates this identifier with the corresponding area coordinates to form structured area coordinate annotation data. This area coordinate annotation data is stored in list or matrix form, with each data entry containing the first object identifier and area coordinate information, ensuring that the specific location of each inattentive first object within the target area can be accurately traced. Finally, based on the aforementioned percentage statistics and area coordinate annotation data, attention state distribution data in the form of statistical charts or heatmaps is generated. This attention state distribution data is used to comprehensively characterize the overall distribution of the attention states of the first objects within the target area.
[0113] In some embodiments, the back-end processing terminal can integrate percentage statistics with regional coordinate annotation data to generate combined statistical charts. For example, percentage statistics can be presented as a pie chart or a donut chart to clearly show the proportion of first objects in a focused state versus the proportion of first objects in an unfocused state; regional coordinate annotation data can be presented as a scatter plot or coordinate distribution plot, on a planar schematic diagram of the target area, the regional coordinates of each first object in an unfocused state are marked with a specific icon (such as a dot or a square), and this specific icon can also distinguish the duration of the first object's unfocused state by different colors or sizes.
[0114] In some embodiments, the backend processing terminal can also, based on the region coordinate annotation data, first divide the target region into several grid cells of equal area, count the number of first objects in an inattentive state within each grid cell, and obtain the inattentive density value of each grid cell; then, according to a preset density mapping rule, assign different density values to different shades of color (e.g., the higher the density, the darker the color), and overlay the percentage statistics as text annotations at a designated location (such as the top or corner) of the heatmap. The heatmap intuitively reflects the distribution density of first objects in an inattentive state within the target region through color shades. For example, dark areas represent areas where first objects in an inattentive state are more concentrated, and light areas represent areas where first objects in an attentive state are more concentrated. Combined with the overlaid percentage data, the overall attention level and local attention level of all first objects within the target region can be quickly determined.
[0115] In some embodiments, statistical charts or heatmaps can be stored in standardized image formats (such as PNG and SVG) as complete attention state distribution data. This attention state distribution data includes both quantitative information on the overall level of attention concentration and distribution characteristics of local areas, enabling a comprehensive and intuitive representation of the attention state of the first object within the target area. This provides accurate and visualized data support for triggering subsequent early warning operations and formulating intervention measures.
[0116] In some embodiments, after triggering the corresponding warning operation based on the attention state distribution data of all first objects and the attention state of each first object, the method further includes: Obtain feedback data within the target area.
[0117] Based on the feedback data, the feature extraction parameters of the deep learning model, the tracking parameters of the single-target tracking algorithm, and the calculation coefficients of the included angle value are adjusted.
[0118] The feedback data includes: the results of manual verification of the attention state of the first subject and the data on the effectiveness of the implementation of personalized intervention suggestions.
[0119] In some embodiments, after triggering the corresponding warning operation based on the attention state distribution data and the attention state of each first object, feedback data within the target area is first acquired. This feedback data includes the manual verification results of the attention state of the first object and the execution effect data of the personalized intervention suggestions. The back-end processing terminal can also provide an interactive entry point, or provide an interactive entry point through a target terminal (such as a teacher's mobile phone or tablet terminal) associated with the back-end processing terminal, allowing relevant personnel (such as teachers) to input the manual verification results.
[0120] In some embodiments, the manual verification results specifically include: for each first object's attention state (focused or unfocused), marking whether the attention state is consistent with the first object's actual attention state. If the two are inconsistent, the actual attention state and the reason for the judgment deviation can be manually supplemented (such as the system misjudging as unfocused when it is actually focused, the system not detecting attention switching, etc.). At the same time, relevant personnel can record the implementation status of personalized intervention suggestions to form implementation effect data, including whether the corresponding intervention suggestion was implemented, the implementation time point, the change in the first object's attention state after implementation (such as switching to a focused state within 10 seconds after implementation, or remaining unfocused after implementation), and the reason for not implementing the suggestion.
[0121] In some embodiments, the feedback data is stored in the form of a structured form, with each data entry associated with a corresponding first object identifier, timestamp, attention state determination result, and intervention suggestion number, ensuring that the feedback data accurately corresponds to the original monitoring data.
[0122] Subsequently, the backend processing terminal can adjust the feature extraction parameters of the deep learning model, the tracking parameters of the single-object tracking algorithm, and the calculation coefficients of the included angle value based on the feedback data to achieve iterative optimization of the model and algorithm. For example, for the deep learning model, the backend processing terminal can select the panoramic image frames corresponding to the first object with judgment deviation in the manual verification results and mark these panoramic image frames as the optimization sample set. For the panoramic image frames in the optimization sample set, the reasons for the judgment deviation are analyzed: if the judgment deviation is due to inaccurate head region feature extraction (such as occlusion or extreme posture leading to head region recognition error), the convolution kernel size, stride, and number of the feature extraction branch in the deep learning model are adjusted to enhance the extraction ability of small-scale features and edge features; if the judgment deviation is due to head posture angle calculation error (such as inaccurate pitch angle and yaw angle estimation), the weight parameters of the key point detection module in the deep learning model are adjusted to strengthen the feature response of key points such as the corners of the eyes and the tip of the nose, thereby improving the accuracy of posture angle calculation.
[0123] It should be understood that, in the specific adjustment process described above in this application embodiment, the mini-batch gradient descent method can also be used to perform secondary training on the optimized sample set using the real state verified by humans as the label, and to iteratively update the feature extraction parameters to ensure that the adjusted deep learning model gradually improves the estimation accuracy of the head pose of the first object in complex scenes.
[0124] In some embodiments, for the tracking parameters of the single-target tracking algorithm, the backend processing terminal can analyze deviation cases related to the tracking of the second object in the feedback data (such as tracking drift, failure to recover after occlusion, etc.) and adjust the tracking parameters of the single-target tracking algorithm accordingly. For example, if there is a tracking drift problem (such as the tracking box deviating from its actual position when the second object moves), the update threshold of the tracking template is reduced, and the update frequency of the tracking template is increased to ensure that the tracking template can adapt to the pose changes of the second object in real time; if there is a problem of failure to quickly recover after occlusion, the search area of the monocular tracking algorithm is expanded, and the similarity threshold of feature matching is adjusted to reduce the interference of the occluded area on feature matching; if there is a tracking box jitter problem, the smoothing coefficient in the tracking parameters is optimized to enhance the stability of the tracking trajectory. Furthermore, successful tracking cases and deviation cases in the feedback data can be compared and verified to ensure that the tracking accuracy and continuity of the single-target tracking algorithm for the second object are improved.
[0125] In some embodiments, for the calculation coefficient of the included angle value, the backend processing terminal can statistically analyze cases in the feedback data where the attention state is misjudged due to the deviation in the included angle value calculation (such as the calculated included angle value being less than a preset threshold but the person is not actually in a state of focused attention, or the included angle value being greater than the preset threshold but the person is actually in a state of focused attention). The matching deviation between the spatial direction vector and the spatial position vector in the misjudged cases can be analyzed. By adjusting the vector normalization coefficient and direction weight coefficient in the included angle value calculation process, the calculation accuracy of the vector included angle can be optimized. For example, if there is a misjudgment of "being judged as focused but not actually focused," it indicates that the included angle value is calculated too small, and the direction weight coefficient can be appropriately adjusted to increase the weight of the deviation between the horizontal and vertical directions; if there is a misjudgment of "being judged as not focused but actually focused," it indicates that the included angle value is calculated too large, and the normalization coefficient can be optimized to reduce the influence of irrelevant noise on the vector direction. Furthermore, after adjusting the calculation coefficient of the included angle value, the included angle value of the misjudged cases is recalculated to ensure that the included angle value matches the actual state verified by humans, making the application of the attention state judgment threshold more accurate.
[0126] After each parameter adjustment, the backend processing terminal can store the adjusted new parameters as an optimized version and enable this version of parameters in subsequent state monitoring processes. As feedback data continues to accumulate, the above parameter adjustment processes iterate cyclically, gradually reducing the false judgment rate of the model and algorithm, improving the overall accuracy and reliability of attention state monitoring, and continuously adapting to the attention state detection needs of different target areas and different scenarios.
[0127] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0128] According to an embodiment of this application, an attention state detection system is also provided for performing the attention state detection method as described in any one of the above. The system includes: A visual sensor, placed within the target area, is used to acquire panoramic images of the target area.
[0129] The panoramic image contains the head regions of several first objects and the full-body regions of second objects.
[0130] The processor, connected to a vision sensor, is used to acquire panoramic images and, based on the head regions of several first objects in the panoramic images, to estimate the head pose of several first objects and obtain the head orientation information of the first objects; based on the region where the second object is located in the panoramic images, to determine the spatial position information of the second object relative to the first object; and based on the head orientation information of each first object and the spatial position information of the second object relative to the first object, to determine the attention state of each first object.
[0131] The attention state is used to characterize whether the head orientation of the first object follows the movement of the second object.
[0132] In some embodiments, the attention state detection system provided in this application can be applied, but is not limited to, teaching scenarios (target area is a classroom, first object is a student, second object is a lecturer), meeting scenarios (target area is a meeting room, first object is a participant, second object is a speaker), corporate internal training scenarios (target area is a corporate training room, first object is a trainee, second object is a training lecturer / demonstration equipment), vocational skills training scenarios (such as driving training, equipment operation training, target area is a driving simulator / practical training workshop, first object is a student, second object is an instructor / practical demonstrator / equipment operation guidance sign), customer service training / service scenarios (target area is a customer service training room / service workstation area, first object is a customer service person (training period / on-the-job period), second object is a trainer / customer, etc., and is not limited to this example.
[0133] The following describes in detail the specific implementation of the attention state detection system provided in this application embodiment, using a teaching scenario (the target area is the classroom, the first object is the student, and the second object is the lecturer). This system is used to execute the aforementioned attention state detection method.
[0134] In some embodiments, the visual sensor can be a wide-angle depth camera, which combines color image acquisition and depth data acquisition capabilities, adapting to the lightweight and low-cost requirements of classroom scenarios. For example, the visual sensor can be pre-deployed in the target area, namely the top center of the front of the classroom (in actual use, it is not limited to one unit; one or more units can also be set at the back and sides of the classroom), and fixed with an adjustable bracket. The bracket can be finely adjusted up, down, left, and right to ensure that the field of view of the visual sensor completely covers the seating area of all students inside the classroom (including the entire range from the front row to the back row) and the instructor's activity area (e.g., the podium and the walking area within 1.5 meters in front of the classroom).
[0135] In some embodiments, the image acquisition resolution of the visual sensor is set to 1920×1080 pixels, and the acquisition frame rate is 30 frames / second, ensuring that the acquired panoramic image clearly presents the head area (including facial features and head outline) of each student and the full body area of the lecturer (including limb movements and location); the acquisition accuracy of the depth data of the visual sensor can be, but is not limited to, ±1 mm, and the effective measurement distance range can be, but is not limited to, 0.3-10 meters, to adapt to the spatial scale of the target area such as the classroom, and to accurately acquire the depth information of each object in the teaching or meeting scene; the field of view can be, but is not limited to, 120 degrees to avoid image acquisition blind spots.
[0136] In some embodiments, the processor in this application can be an edge computing device. This edge computing device possesses powerful parallel computing capabilities while being small in size and low in power consumption, making it suitable for deployment requirements in scenarios such as meetings and classrooms, and also meeting the real-time inference requirements of deep learning models. The processor establishes a stable connection with the visual sensor via a wired interface or wireless communication. In addition to receiving image and depth data transmitted from the visual sensor, the processor can also send parameter adjustment commands (such as adjusting the acquisition frame rate and exposure) to the visual sensor. In some embodiments, the processor is pre-installed with Linux-based operating software, integrating an image preprocessing module, a head pose estimation module, a second object localization and tracking module, an attention state judgment module, and a data storage module. These modules work collaboratively to complete the entire process of panoramic image processing.
[0137] In some embodiments, after receiving the panoramic image transmitted from the vision sensor, the processor first performs grayscale correction on the panoramic image to eliminate image brightness differences caused by uneven ambient light; then, it performs Gaussian filtering denoising to filter out random noise (such as specks generated by light reflection) in the panoramic image; finally, it performs image scale normalization to uniformly adjust the size of the panoramic image to the input size adapted to the model (e.g., 640×480 pixels), while maintaining the proportional relationship of each object in the panoramic image. The preprocessed panoramic image is synchronously transmitted to the head pose estimation module and the second object localization and tracking module.
[0138] Subsequently, the processor employs a head pose estimation module, calling a pre-defined deep learning model (using MobileNetV3 as the backbone network, combined with a 68-point facial landmark detection algorithm) to perform recognition processing on the pre-processed panoramic image. This deep learning model first locates the head region of the first object (student) and outputs the pixel coordinate range of the head region (boundary box coordinates with the top left corner of the image as the origin). Then, it extracts the coordinate information of 68 facial landmarks (corners of the eyes, tip of the nose, corners of the mouth, etc.) within the head region of the first object. Using the perspective n-point algorithm (EPnP) combined with a pre-defined 3D face model template, it calculates the pitch angle, yaw angle, and / or roll angle of the student's head. Based on these three pose parameters, a spatial vector transformation algorithm is used to construct the head orientation vector of the first object. After normalization, this vector is stored as head orientation information in the data storage module.
[0139] Next, the processor uses the second object localization and tracking module to synchronously call the human detection model to detect the second object (lecturer) in the preprocessed panoramic image. The human detection model extracts information such as human contours and limb features from the panoramic image to locate the entire body region of the second object, outputs the two-dimensional pixel coordinates of the bounding box, and extracts the contour features of the second object (such as the contour curve coordinates of the shoulders and torso). Then, a single-object tracking algorithm (an improved kernel correlation filter (KCF) algorithm) is launched, using the contour features obtained from the initial detection as the tracking template. The search area is defined in each subsequent frame image, and the two-dimensional pixel coordinates of the lecturer are updated in real time through feature matching to avoid tracking drift (such as when the lecturer moves or turns, the tracking box can still be accurately locked). At the same time, the second object positioning and tracking retrieves the depth data transmitted by the vision sensor. Based on the lecturer's real-time two-dimensional pixel coordinates, the depth value of the corresponding position is extracted. Combined with the calibration intrinsic parameters of the vision sensor, the two-dimensional pixel coordinates are converted into three-dimensional spatial coordinates (a three-dimensional coordinate system established with the lower left corner of the classroom floor as the origin) to obtain the real-time spatial position of the second object. Finally, based on the three-dimensional spatial coordinates of the center point of the student's head region (converted from the head region coordinates provided by the head pose estimation module), the spatial distance and orientation angle of the lecturer relative to the student are calculated to form the spatial position information of the second object relative to the first object, and stored in the data storage module.
[0140] Finally, the processor uses an attention state determination module to retrieve the head orientation information of the first object and the spatial position information (i.e., spatial position vector) of the second object relative to the first object from the data storage module. It calculates the angle between the head orientation vector of the first object and the corresponding spatial position vector, and converts it into a specific angle value through a vector direction similarity algorithm. The angle value is compared with a preset angle threshold (e.g., 60 degrees). If the angle value is less than 60 degrees, the attention state of the first object is determined to be a focused state (the head orientation follows the lecturer's movement). If the angle value is greater than or equal to 60 degrees, the attention state of the first object is determined to be a disfocused state (the student's head orientation does not follow the lecturer's movement). The attention state result of the first object is then associated with and stored with the corresponding object identifier and timestamp.
[0141] It should be noted that the visual sensor acquires image and depth data within the target area in real time and transmits it to the processor at a frequency of 30 frames per second. The processor sequentially completes image preprocessing, head pose estimation, second object localization and tracking, and attention state judgment through various functional modules. The entire processing time for each frame is controlled within 100 milliseconds, ensuring that the output attention state results can dynamically reflect the student's real-time status. The processor stores the head orientation information of the first object, the spatial location information of the second object, the attention state results, and the corresponding timestamps to the built-in solid-state drive (e.g., storage capacity of 512GB), supporting local data retention and subsequent export analysis.
[0142] The attention state detection system in this embodiment fully implements all the steps of the aforementioned attention state detection method through the collaborative work of a visual sensor and a processor. It features lightweight, real-time performance, and high precision. It does not rely on high-cost professional equipment, is easy to deploy, and does not interfere with the classroom environment. It can accurately monitor the attention state of each student, providing stable system support for subsequent attention distribution statistics and early warning feedback. It is efficient, simple, and low-cost.
[0143] Corresponding to the attention state detection method in the above embodiments, Figure 8 This is a schematic diagram of an attention state detection device provided in an embodiment of this application. The device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be... Figure 9 The electronic device shown.
[0144] Reference Figure 8 The attention state detection device includes: The acquisition unit 801 is used to acquire a panoramic image within the target area, wherein the panoramic image contains at least a number of head regions of a first object and a second object.
[0145] The pose estimation unit 802 is used to estimate the head pose of several first objects based on the head regions of several first objects in the panoramic image, so as to obtain the head orientation information of the first objects.
[0146] The first determining unit 803 is used to determine the spatial position information of the second object relative to the first object based on the area where the second object is located in the panoramic image.
[0147] The second determining unit 804 is used to determine the attention state of each first object based on the head orientation information of each first object and the spatial position information of the second object relative to the first object, wherein the attention state is used to characterize whether the head orientation of the first object follows the movement of the second object.
[0148] It is understood that the embodiments of the attention state detection device and any implementation thereof correspond to the embodiments of the attention state detection method and any implementation thereof. The technical effects corresponding to the embodiments of the attention state detection device and any implementation thereof can be found in the technical effects corresponding to the aforementioned embodiments of the attention state detection method and any implementation thereof, and will not be repeated here.
[0149] It should be noted that the attention state detection device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0150] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.
[0151] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0152] This application also provides an electronic device, which includes one or more processors and a memory; The memory is coupled to one or more processors. The memory is used to store computer program code, which includes computer instructions. One or more processors call the computer instructions to cause the electronic device to perform the attention state detection method described above.
[0153] Figure 9This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 900 can be a mobile phone, smart screen, tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), projector, or a communication device such as a server, storage device, or base station, or a smart car, etc. This application embodiment does not impose any limitations on the specific type of electronic device.
[0154] The memory 901 can be used to store computer software programs 902 and modules. The processor 903 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 901. The memory 901 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, telephone directory, etc.). In addition, the memory 901 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0155] The processor 903 may include one or more processors such as a central processing unit (CPU), an application processor (AP), and a baseband processor. The processor can serve as the nerve center and command center of the wireless router. The processor 903 can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The memory 901 can be used to store executable program code, including instructions. The processor 903 executes various functional applications and data processing of the network device by running the instructions stored in the memory. The memory 901 may include a program storage area and a data storage area, such as storing data for audio signals to be played. For example, the memory may be Double Data Rate Synchronous Dynamic Random Access Memory (DDR) or Flash memory.
[0156] This application also provides a computer-readable storage medium storing computer instructions; when the computer-readable storage medium is used on an electronic device, it causes the electronic device to execute the attention state detection method described above.
[0157] The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or can include one or more data storage devices such as servers or data centers that can be integrated with media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media, or semiconductor media (e.g., solid-state disks (SSDs)).
[0158] This application also provides a computer program product containing computer instructions, which, when run on an electronic device, enables the electronic device to execute the attention state detection method described above.
[0159] The computer storage medium and computer program product provided in the above embodiments of this application are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects corresponding to the methods provided above, and will not be repeated here.
[0160] In the above embodiments, implementation can also be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, Digital Subscriber Line, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc., and the storage medium can also include combinations of the above types of memory.
[0161] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0162] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments claimed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0163] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0164] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0165] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for detecting attention states, characterized in that, include: Acquire a panoramic image of the target area, wherein the panoramic image contains at least a number of head regions of a first object and a second object; Based on the head regions of several first objects in the panoramic image, head pose estimation is performed on several first objects to obtain the head orientation information of the first objects; Based on the area where the second object is located in the panoramic image, determine the spatial position information of the second object relative to the first object; Based on the head orientation information of the first object and the spatial position information of the second object relative to the first object, the attention state of the first object is determined, wherein the attention state is used to characterize whether the head orientation of the first object follows the movement of the second object.
2. The method according to claim 1, characterized in that, The step of estimating the head pose of several first objects based on their head regions in the panoramic image to obtain their head orientation information includes: A deep learning model is used to identify the head regions of several first objects in the panoramic image to obtain the pixel coordinate range of the head regions of the first objects. Feature extraction is performed on the image data corresponding to the pixel coordinate range of the head region to determine the pitch angle, yaw angle and / or roll angle of the head of the first object; Based on the pitch angle, yaw angle, and / or roll angle of the first object's head, a head orientation vector of the first object is constructed as the head orientation information.
3. The method according to claim 1, characterized in that, Determining the spatial position information of the second object relative to the first object based on the region where the second object is located in the panoramic image includes: A human detection model is used to detect the region where the second object is located in the panoramic image in order to determine the contour features of the second object; A single-target tracking algorithm is used to continuously track the contour features of the second object in order to synchronously update the real-time two-dimensional coordinates of the second object; Based on the depth data collected by the visual sensor within the target area, the real-time two-dimensional coordinates of the second object are converted into three-dimensional spatial coordinates; The spatial distance and orientation angle of the two objects relative to the center point of the head region of the first object are calculated respectively to obtain the spatial position information of the second object relative to the first object.
4. The method according to claim 1, characterized in that, Determining the attention state of the first object based on the head orientation information of the first object and the spatial position information of the second object relative to the first object includes: The head orientation information of the first object is converted into a spatial direction vector, and the spatial position information of the second object relative to the first object is converted into a spatial position vector; Calculate the angle between the spatial direction vector and the spatial position vector; If the included angle value is less than a preset angle threshold, then the attention state of the first object is determined to be a state of focused attention. If the included angle value is greater than or equal to the preset angle threshold, then the attention state of the first object is determined to be an inattentive state.
5. The method according to any one of claims 1 to 4, characterized in that, The attentional state of the first object includes: a state of focused attention when the head of the first object moves in the same direction as the second object; and a state of unfocused attention when the head of the first object does not move in the same direction as the second object. The number of the first objects is multiple. After determining the attention state of each first object based on the head orientation information of each first object and the spatial position information of the second object relative to each first object, the method further includes: Based on the attention state of each of the first objects, determine a first number of first objects within the target area that are in the state of focused attention; Based on the attention state of each of the first objects, determine the distribution location of the first objects in the inattentive state within the target area; Based on the distribution location and the proportion of the first quantity to the total number of the first objects, the attention state distribution data of all the first objects is obtained; Based on the attention state distribution data of all the first objects, and the attention state of each first object, a corresponding warning operation is triggered.
6. The method according to claim 5, characterized in that, The step of triggering a corresponding early warning operation based on the attention state distribution data of all the first objects and the attention state of each first object includes: Based on the attention state of each of the first objects, determine a first duration for which any one of the first objects remains in the state of inattention. If the first duration exceeds a preset duration threshold, an audible warning signal will be triggered and / or a visual warning message will be generated. For each first subject in a state of inattention, based on the first duration and the distribution location, a personalized intervention suggestion is generated, which includes interactive guidance prompts and / or seat adjustment suggestions. The attention state distribution data, the visual warning information, and the personalized intervention suggestions are displayed in real time in the visualization interface, and the first object in the inattention state is highlighted by highlighting or special marking.
7. The method according to claim 5, characterized in that, Based on the distribution location and the proportion of the first quantity to the total number of the first objects, attention state distribution data of all the first objects is obtained, including: Based on the proportion of the first quantity to the total number of the first objects, we obtain statistical data on the percentage of the first objects in the state of focused attention. Based on the distribution location, obtain the region coordinate annotation data of the first object in the state of inattention; Based on the percentage statistics and the regional coordinate annotation data, attention state distribution data in the form of statistical charts or heat maps is generated, wherein the attention state distribution data is used to characterize the overall distribution of the attention state of the first object within the target area.
8. The method according to claim 5, characterized in that, After triggering the corresponding warning operation based on the attention state distribution data of all the first objects and the attention state of each first object, the method further includes: Obtain feedback data within the target area, including: manual verification results of the attention state of the first object and execution effect data of personalized intervention suggestions; Based on the feedback data, the feature extraction parameters of the deep learning model, the tracking parameters of the single-target tracking algorithm, and the calculation coefficients of the included angle value are adjusted.
9. An attention state detection system, characterized in that, The system is used to perform the attention state detection method as described in any one of claims 1 to 8, the system comprising: A visual sensor is disposed within the target area to acquire panoramic images of the target area, wherein the panoramic images include a number of head regions of a first object and the full-body regions of a second object; The processor, connected to the vision sensor, is configured to acquire the panoramic image and, based on the head regions of several first objects in the panoramic image, estimate the head pose of the first objects to obtain head orientation information of the first objects; determine the spatial position information of the second objects relative to the first objects based on the region where the second objects are located in the panoramic image; and determine the attention state of the first objects based on the head orientation information of the first objects and the spatial position information of the second objects relative to the first objects, wherein the attention state is used to characterize whether the head orientation of the first objects follows the movement of the second objects.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the method as described in any one of claims 1 to 8.
11. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method as described in any one of claims 1 to 8 to be performed.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.