State analysis method and device, computer device and storage medium

By collecting video data and extracting features from wearable devices, and using a trained state distribution recognition model to determine the difference between the current state features and the standard state features, the problem of limited sensor data information in wearable devices is solved, enabling more accurate service provision.

CN122391950APending Publication Date: 2026-07-14
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610462622.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-04-09
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Wearable devices cannot provide accurate services through sensor data analysis because sensor data information is limited.

Method used

By acquiring video data collected during the wearing of wearable devices, feature extraction is performed to generate scene, object, limb, and perspective motion features. The trained state distribution recognition model is then used to determine the difference between the current state features and the standard state features.

Benefits of technology

It achieves multi-dimensional feature acquisition, which can accurately determine the current state of an object and provide more precise personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391950A_ABST
    Figure CN122391950A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a state analysis method and device, computer equipment and a storage medium. The method comprises: acquiring video data collected by a wearable device during wearing; performing feature extraction on the video data to obtain scene features, object features, limb features and visual angle motion features; generating target features according to the scene features, object features, limb features and visual angle motion features; inputting the target features into a trained state distribution identification model to output a current state feature corresponding to the video data and a corresponding standard state, wherein the trained state distribution identification model is trained based on label states of sample video data corresponding to an object, and sample scene features, sample object features, sample limb features and sample visual angle motion features in the sample video data; determining a state difference between the current state feature and a standard state feature corresponding to the standard state, and determining a current state of the object according to the state difference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a state analysis method, apparatus, computer device, and storage medium. Background Technology

[0002] With the development of technology, wearable devices have become common electronic devices in daily life, such as smartwatches, smart glasses, and smart bracelets.

[0003] In related technologies, to provide services on wearable devices, sensor data captured by non-visual sensors on the device is acquired, and then analyzed using appropriate algorithms to provide the services. However, the information provided by this sensor data is limited, which can lead to the wearable device being unable to accurately provide the corresponding services to the user. Summary of the Invention

[0004] This application provides a state analysis method, apparatus, computer device, and storage medium, which can output current state characteristics through an object-specific state distribution recognition model and accurately determine the current state of the object based on the difference between the current state characteristics and standard state characteristics.

[0005] To achieve the above objectives, one embodiment of this application provides a state analysis method, including: Acquire video data collected during the wearing of wearable devices; Feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features; Target features are generated based on the scene features, object features, limb features, and viewpoint motion features; The target features are input into the trained state distribution recognition model, which outputs the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data. Determine the state difference between the current state feature and the standard state feature corresponding to the standard state, and determine the current state of the object based on the state difference.

[0006] To achieve the above objectives, one embodiment of this application provides a state analysis apparatus, including: The acquisition module is used to acquire video data collected by the wearable device during the wearing process; The extraction module is used to extract features from the video data to obtain scene features, object features, limb features, and viewpoint motion features; The generation module is used to generate target features based on the scene features, the object features, the limb features, and the viewpoint motion features; The input module is used to input the target features into the trained state distribution recognition model and output the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data. The determination module is used to determine the state difference between the current state feature and the standard state feature corresponding to the standard state, and to determine the current state of the object based on the state difference.

[0007] In some implementations, the generation module is used for: The acquisition time sequence corresponding to the limb features and the viewpoint motion features is determined, and the temporal features corresponding to the video data are generated according to the acquisition time sequence. Target features are generated based on the temporal features, scene features, object features, limb features, and viewpoint motion features.

[0008] In some implementations, the generation module is used for: The temporal features, scene features, object features, limb features, and viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time. Each of the splicing features is sequentially spliced ​​according to the acquisition time sequence to generate the target feature.

[0009] In some implementations, the generation module is used for: The scene corresponding to each collection time is determined based on the scene characteristics; Based on the scene, the object features, the limb features, and the viewpoint motion features are weighted and processed to obtain updated object features, updated limb features, and updated viewpoint motion features; The temporal features, scene features, updated object features, updated limb features, and updated viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time.

[0010] In some implementations, the extraction module is used for: The video data is subjected to privacy filtering to obtain the filtered first video data; The first video data is subjected to blurry image filtering to obtain the filtered second video data. The second video data is normalized according to preset screen parameters to obtain normalized target video data; Feature extraction is performed on the target video data to obtain scene features, object features, limb features, and viewpoint motion features.

[0011] In some implementations, the extraction module is used for: Multiple target images are determined from the target video data according to a preset time interval; Each target image is input into a pre-trained image recognition model to output the scene features and object features corresponding to the video data; Each target image is input into a pre-trained action recognition model to output the limb features and viewpoint motion features corresponding to the video data.

[0012] In some implementations, the determining module is used for: The current behavior pattern of the object is determined based on the state differences. Obtain multiple historical behavior patterns of an object, and determine the trend of behavior pattern changes from the multiple historical behavior patterns to the current behavior pattern; The object's health status is determined based on the preset health status corresponding to the standard state and the trend of behavioral pattern changes.

[0013] In some implementations, the determining module is used for: Determine the standard psychological state score of the object under the standard state; The adjustment score value is determined based on the preset scoring criteria and the state differences; The target psychological state score is obtained by updating the standard psychological state score based on the adjusted score. The psychological state of the object is determined based on the psychological state score and the preset mapping relationship between psychological states, as well as the target psychological state score.

[0014] In some implementations, the state analysis apparatus further includes a training module for: Before inputting the target features into the trained state distribution recognition model and outputting the current state features and the corresponding standard state of the video data, multiple sets of sample video data collected by the wearable device over a historical period and the label state corresponding to each set of sample video data are obtained. The sample video data is the data corresponding to the object in the standard state. Feature extraction is performed on each group of sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each group of sample video data; Sample features are generated based on the sample scene features, the sample object features, the sample limb features, and the sample viewpoint motion features; The sample features are input into the state distribution recognition model, and the predicted state corresponding to the sample video data is output. The state distribution recognition model is iteratively trained based on the difference between the predicted state and the labeled state to obtain the trained state distribution recognition model.

[0015] In some implementations, the training module is used for: Each set of sample video data is subjected to privacy filtering to obtain the first filtered sample video data. For each group of first sample video data, perform unclear image filtering to obtain filtered second sample video data; The second sample video data of each group is normalized according to the preset screen parameters to obtain the normalized target sample video data corresponding to each group of the second sample video data. Feature extraction is performed on each set of target sample video data to obtain sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each set of sample video data.

[0016] To achieve the above objectives, one aspect of this application provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the steps in the state analysis method provided in this application.

[0017] To achieve the above objectives, one aspect of this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, it implements the steps in the state analysis method provided in this application.

[0018] To achieve the above objectives, one aspect of this application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the state analysis method provided in this application.

[0019] In this embodiment, video data collected by the wearable device during its wearing process is acquired; features are extracted from the video data to obtain scene features, object features, limb features, and viewpoint motion features; target features are generated based on the scene features, object features, limb features, and viewpoint motion features; the target features are input into a trained state distribution recognition model, which outputs the current state features corresponding to the video data and the corresponding standard state. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data; the state difference between the current state features and the standard state features corresponding to the standard state is determined, and the current state of the object is determined based on the state difference.

[0020] Therefore, by acquiring video data collected during the wearing of a wearable device, feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features. Target features are then generated based on these features. This enables multi-dimensional feature acquisition, and the combination of these multi-dimensional features generates richer target features. The target features are then input into a trained state distribution recognition model, which outputs the current state features and corresponding standard state of the video data. The trained state distribution recognition model is based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data. Therefore, the trained state distribution recognition model in this application is highly adaptable to the current object and can analyze target features to accurately output the current state features and corresponding standard state of the video data. Finally, the state difference between the current state features and the standard state features corresponding to the standard state is determined, and the current state of the object is determined based on this difference. In this way, by determining the state difference between the current state characteristics and the standard state characteristics corresponding to the normal standard state, the state difference can accurately reflect the deviation between the current state and the standard state, thereby accurately determining the current state of the object.

[0021] Compared to related technologies that analyze sensor data using algorithms to provide services, this application can more accurately determine the current state characteristics of an object by using a trained state distribution recognition model specific to the object, and accurately determine the current state of the object based on the difference between the current state characteristics and the standard state characteristics, thereby providing more precise services to the object based on the current state.

[0022] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the system framework corresponding to the state analysis method provided in the embodiments of this application; Figure 2 This is a schematic diagram of a scenario for the state analysis method provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the state analysis method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the training process of the state distribution recognition model provided in the embodiments of this application; Figure 5 This is another schematic flowchart of the state analysis method provided in the embodiments of this application; Figure 6 This is a schematic diagram of the state analysis device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0025] To enable those skilled in the art to better understand the solutions of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] It should be noted that in various specific embodiments of this application, when processing data related to the identity or characteristics of an object, such as object information, object behavior data, object historical data, and object location information, the object's permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require obtaining sensitive personal information of an object, separate permission or consent from the object is obtained through pop-ups or redirection to a confirmation page. Only after obtaining the object's separate permission or consent is the necessary object-related data required for the proper functioning of the embodiments of this application obtained.

[0027] In some processes described in the specification, claims and the foregoing drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may be performed in any order or in parallel. The step numbers are only used to distinguish the different steps and do not represent any execution order.

[0028] The state analysis method provided in this application relates to the field of computer technology. The state analysis method provided in this application can be used in numerous general-purpose or special-purpose computer system environments or configurations, such as in terminals, servers, or software running on terminals or servers. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the state analysis method, but is not limited to the above forms.

[0029] The state analysis method provided in this application can be executed by a computer program module, which is integrated within the state analysis device. Generally, the program module includes routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, the program module can reside in local and remote computer storage media, including storage devices.

[0030] Before providing a further detailed description of the embodiments of this application, the nouns and terms used in the embodiments of this application are explained, and the nouns and terms used in the embodiments of this application shall be interpreted as follows: Wearable devices are portable smart electronic devices that are worn directly on the human body or integrated into clothing and accessories. They integrate sensing, computing, communication, and interaction, such as smart glasses and smartwatches. The core is to achieve real-time collection, processing, and feedback of human body data to achieve seamless collaboration of "human-machine symbiosis".

[0031] Artificial intelligence models (AI models) are mathematical / logical systems trained through algorithms, data, and computation. They can learn patterns from data and automatically complete tasks such as perception, understanding, reasoning, generation, and decision-making, simulating or even extending human intelligence.

[0032] The above is a detailed description of some of the relevant terms in this application. If other terms are involved later, they will be described in detail thereafter.

[0033] The technical problems existing in the relevant technology are as follows: With the development of technology, wearable devices have become common electronic devices in daily life, such as smartwatches, smart glasses, and smart bracelets.

[0034] In related technologies, to provide services on wearable devices, sensor data captured by non-visual sensors on the device is acquired, and then analyzed using appropriate algorithms to provide the services. However, the information provided by this sensor data is limited, which can lead to the wearable device being unable to accurately provide the corresponding services to the user.

[0035] To address the aforementioned problems, this application's embodiments acquire video data collected during the wearing of a wearable device. Feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features. Target features are then generated based on these features. This enables multi-dimensional feature acquisition, and the combination of these multi-dimensional features generates richer target features. The target features are then input into a trained state distribution recognition model, which outputs the current state features corresponding to the video data and the corresponding standard state. The trained state distribution recognition model is based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data. Therefore, the trained state distribution recognition model in this application is highly adaptable to the current object and can analyze target features to accurately output the current state features and the corresponding standard state of the video data. Finally, the state difference between the current state features and the standard state features corresponding to the standard state is determined, and the current state of the object is determined based on this difference. In this way, by determining the state difference between the current state characteristics and the standard state characteristics corresponding to the normal standard state, the state difference can accurately reflect the deviation between the current state and the standard state, thereby accurately determining the current state of the object.

[0036] Compared to related technologies that analyze sensor data using algorithms to provide services, this application can more accurately determine the current state characteristics of an object by using a trained state distribution recognition model specific to the object, and accurately determine the current state of the object based on the difference between the current state characteristics and the standard state characteristics, thereby providing more precise services to the object based on the current state.

[0037] The system architecture used in this application embodiment is as follows: Please see Figure 1 , Figure 1 This is a schematic diagram of the system framework corresponding to the state analysis method provided in this application embodiment. The state analysis method provided in this application embodiment can be applied to this system framework.

[0038] It includes terminal 140, Internet 130, gateway 120, server 110, etc.

[0039] Terminal 140 or server 110 can be a device that performs state analysis methods.

[0040] Terminal 140 includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, and aircraft. Embodiments of this application can be applied to various scenarios, including but not limited to the fields of health management, sports and fitness, and medical rehabilitation. Furthermore, it can be a single device or a collection of multiple devices. For example, multiple desktop computers can be interconnected via a local area network, sharing a single monitor to work collaboratively, forming a single terminal 140. Terminal 140 can communicate with the Internet 130 via wired or wireless means to exchange data.

[0041] Server 110 refers to a computer system that can provide certain services to terminal 140. Compared to ordinary terminal 140, server 110 has higher requirements in terms of stability, security, and performance. Server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0042] Gateway 120, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from terminal 140 to server 110 are forwarded to the corresponding server 110 via gateway 120. Messages sent from server 110 to terminal 140 are also forwarded to the corresponding terminal 140 via gateway 120.

[0043] The state analysis method in this application can be applied to various scenarios, including but not limited to health management, sports and fitness, and medical rehabilitation. This application does not impose any limitations on the scenarios in which the state analysis method in this application can be used.

[0044] The application scenarios of this application are as follows: Please see Figure 2 , Figure 2 This is a schematic diagram of a scenario for the state analysis method provided in the embodiments of this application.

[0045] In a motion scene, an object wears a wearable device. The video data collected by the wearable device during the wearing process is obtained. The wearable device can collect video data through a camera, and the video data can be video data from a first-person perspective.

[0046] Feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features. For example, scene features include moving scenes, object features include a soccer ball and a goal, limb features include the amplitude of alternating leg movements, and viewpoint motion features include the angle and frequency of viewpoint changes.

[0047] Then, target features are generated based on scene features, object features, limb features, and viewpoint motion features. For example, the above features can be concatenated to form a target feature, such as: motion scene + soccer ball and goal + leg alternating movement amplitude + viewpoint change angle and frequency.

[0048] The target features are input into the trained state distribution recognition model, which outputs the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data.

[0049] It's important to note that the trained state distribution recognition model is based on sample video data generated by an object in its normal state. This includes sample video data corresponding to normal states such as eating, walking, playing soccer, and running. This video data is used to train the state distribution recognition model, allowing it to learn the behavioral distribution patterns and normal states of an object. Therefore, by inputting the target features into the trained state distribution recognition model, it can extract the current state features and determine the corresponding standard state—that is, the normal state that the current state feature should correspond to.

[0050] Finally, the state differences between the current state features and the standard state features corresponding to the standard state are determined, and the current state of the object is determined based on these differences. For example, the standard state features corresponding to the standard state include the standard body temperature range and standard heart rate range of the object in a soccer-playing scenario, while the current state features include the object's current body temperature and heart rate. By comparing the two, the state differences between the standard state features are determined, and the current state of the object can be determined based on these differences. For example, if the object is currently in a state of strenuous exercise, its body temperature and heart rate will both exceed the standard range.

[0051] Wearable devices can provide relevant service information based on the user's current status. For example, a wearable device can provide a voice prompt such as, "Your current body temperature and heart rate are both above the standard range. Please take a break and replenish fluids." This can provide a personalized service to the user.

[0052] The following will describe in detail the state analysis method, apparatus, computer equipment, and storage medium provided in the embodiments of this application.

[0053] Please see Figure 3 , Figure 3 This is a flowchart illustrating the state analysis method provided in an embodiment of this application. The state analysis method provided in an embodiment of this application may include the following steps: Step 210: Acquire video data collected during the wearing of the wearable device; Step 220: Extract features from the video data to obtain scene features, object features, limb features, and viewpoint motion features; Step 230: Generate target features based on scene features, object features, limb features, and viewpoint motion features; Step 240: Input the target features into the trained state distribution recognition model and output the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data. Step 250: Determine the state difference between the current state characteristics and the standard state characteristics corresponding to the standard state, and determine the current state of the object based on the state difference.

[0054] Before describing steps 210 to 250 in detail, let's first describe the training process of the trained state distribution recognition model.

[0055] Combination Figure 4 , Figure 4 This is a schematic diagram of the training process of the state distribution recognition model provided in this application embodiment. Specifically, it includes the following steps: Step 301: Obtain multiple sets of sample video data collected by the wearable device within a historical time period and the tag status corresponding to each set of sample video data. The sample video data is the data corresponding to the object in the standard state. Step 302: Extract features from each group of sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each group of sample video data. Step 303: Generate sample features based on sample scene features, sample object features, sample limb features, and sample viewpoint motion features; Step 304: Input the sample features into the state distribution recognition model and output the predicted state corresponding to the sample video data; Step 305: Iteratively train the state distribution recognition model based on the difference between the predicted state and the labeled state to obtain the trained state distribution recognition model.

[0056] The training process of the state distribution recognition model will be described below through steps 301 to 305.

[0057] In step 301, multiple sets of sample video data collected by the wearable device over a historical period and the tag status corresponding to each set of sample video data are obtained. The sample video data is the data corresponding to the object in the standard state.

[0058] First, we acquire multiple sets of sample video data collected by the wearable device over a historical period, along with the corresponding tag status for each set. The sample video data consists of first-person perspective video data captured when the wearable device is in a normal behavioral state. The wearable device can be fixed to the head, collar, or brim of a hat to ensure that the image captures the wearer's true field of vision, and the acquisition process does not interfere with daily life. The sample video data also represents the data corresponding to the object in a standard state; that is, the sample video data records the object's data under normal behavioral conditions (standard state).

[0059] Furthermore, for each set of sample video data, the video data also needs to be labeled. For example, the sample video data of an object walking is labeled as "walking," and the sample video data of an object eating is labeled as "eating."

[0060] This allows us to use multiple sets of sample video data as sample data and the corresponding label state for each set of sample video data as label data. The sample data and label data are then used to form a training dataset.

[0061] In step 302, feature extraction is performed on each group of sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each group of sample video data.

[0062] For each set of sample video data, there may be some privacy information, some unclear scenes, and different brightness of different scenes. Therefore, each set of sample video data needs to be preprocessed to obtain target sample video data. Finally, various features are extracted based on the target sample video data to ensure the reliability of the extracted features, so as to efficiently train the state distribution recognition model.

[0063] In some implementations, feature extraction is performed on each set of sample video data to obtain sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each set of sample video data, including: (1.1) Perform privacy filtering on each group of sample video data to obtain the first sample video data after filtering; (1.2) Perform unclear image filtering on each group of first sample video data to obtain filtered second sample video data; (1.3) Normalize each group of second sample video data according to the preset screen parameters to obtain the normalized target sample video data corresponding to each group of second sample video data. (1.4) Perform feature extraction on each group of target sample video data to obtain the sample scene features, sample object features, sample limb features and sample viewpoint motion features corresponding to each group of sample video data.

[0064] In this process, privacy filtering is performed on each set of sample video data to obtain filtered first sample video data. For example, facial blurring and filtering of information such as phone numbers are performed. These key areas are masked to protect privacy, thus obtaining multiple sets of first sample video data that do not contain privacy information.

[0065] For example, using a YOLOv8-based object detection model, privacy regions are identified in each frame of each original video to accurately locate faces and license plates. Gaussian blurring is applied to the identified privacy regions, while the non-privacy regions retain their original image. The processed frames are then stitched together according to the original video frame rate to generate the first sample video data for each group. Example result: In a shopping mall video, pedestrian faces are blurred, retaining only body movements and environmental information, with no clear faces exposed.

[0066] Then, the first sample video data for each group is processed to filter out unclear images, resulting in filtered second sample video data. For example, occluded images, blurry images, and images with distortion are filtered to obtain the second sample video data.

[0067] For example, for each frame of the first sample video in each group, a sharpness index is calculated. Specifically, the image edge gradient variance is calculated using the Laplacian operator, and frames with a variance greater than 80 are considered sharp. A brightness index is calculated: the frames are converted to the HSV color space, the mean value of the brightness channel is extracted, and frames with a mean value in the range of [50, 200] are considered to have normal brightness. Frames that do not meet the above two indices are filtered out, and the remaining valid frames are reassembled according to the original time sequence to generate the second sample video data for each group (the number of frames may be less than the original video, but the time sequence is continuous). Example result: In a subway station video, 5 frames of overly dark scenes and 3 frames of motion blur caused by flickering lights were filtered out, and 282 sharp frames were ultimately retained.

[0068] Then, the second sample video data for each group is normalized according to preset image parameters to obtain the normalized target sample video data corresponding to each group of second sample video data. For example, the image parameters of all sample video data are unified to eliminate differences in resolution, frame rate, and color. For example, the preset image parameters are a resolution of 640×480, a frame rate of 25fps, and a color space of RGB.

[0069] Finally, feature extraction was performed on each group of target sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each group of sample video data. For example, the sample scene features are: indoor enclosed scene + artificial lighting + escalator facilities. The sample object features are: 3 signs + 1 escalator. The sample limb features are: moderate arm swing amplitude + no running motion. The sample viewpoint motion features are: slight camera translation (horizontal displacement of 5 pixels), no rotation, and no zoom.

[0070] The advantage of doing this is that it allows for the unification of all sample video data, thereby obtaining more accurate sample scene features, sample object features, sample limb features, and sample perspective motion features, which helps to improve the learning efficiency of the state distribution recognition model.

[0071] In step 303, sample features are generated based on sample scene features, sample object features, sample limb features, and sample viewpoint motion features.

[0072] Specifically, the sample acquisition time sequence corresponding to the sample limb features and sample viewpoint motion features can be determined, and the sample temporal features corresponding to the video data can be generated based on the sample acquisition time sequence; sample features can be generated based on the sample temporal features, sample scene features, sample object features, sample limb features, and sample viewpoint motion features.

[0073] Specifically, the temporal features, scene features, object features, limb features, and perspective motion features of each sample at each collection time are concatenated to obtain the concatenated features of each sample at each collection time; and each concatenated feature is concatenated sequentially according to the order of sample collection time to generate sample features.

[0074] For example, the sample scene corresponding to each sample collection time is determined based on the sample scene characteristics; the sample object features, sample limb features, and sample viewpoint motion features are weighted according to the sample scene to obtain updated sample object features, updated sample limb features, and updated sample viewpoint motion features; the sample temporal features, sample scene features, updated sample object features, updated sample limb features, and updated sample viewpoint motion features for each sample collection time are concatenated to obtain the sample concatenated features for each sample collection time. Each sample concatenated feature is then concatenated sequentially according to the sample collection time order to generate the sample features.

[0075] The advantage of doing this is that it allows for the combination of features from multiple dimensions to obtain semantically rich sample features. This enables the state distribution recognition model to learn more knowledge based on sample features during training, thereby improving its ability to recognize state distributions and accurately extract state features and determine the corresponding standard states.

[0076] In step 304, the sample features are input into the state distribution recognition model, and the predicted state corresponding to the sample video data is output.

[0077] Specifically, sample features can be input into the state distribution recognition model, which outputs the predicted state corresponding to the sample video data. The state distribution recognition model can extract features from the sample features to obtain the state distribution features of the object and determine the predicted state of the object.

[0078] In step 305, the state distribution recognition model is iteratively trained based on the difference between the predicted state and the label state to obtain the trained state distribution recognition model.

[0079] This involves determining the difference between the predicted state and the labeled state, for example, the loss value corresponding to the predicted and labeled states. When this loss value does not meet a preset difference condition, backpropagation is performed based on this loss value to update the parameters of the state distribution recognition model, thereby achieving iteration of the state distribution recognition model. This process continues until the loss value meets the preset difference condition, resulting in the trained state distribution recognition model.

[0080] As shown above, the trained state distribution recognition model can accurately identify the standard state of an object under normal behavior, and can accurately obtain the standard state features based on sample features. This indicates that the trained state distribution recognition model has accurate feature extraction capabilities.

[0081] The advantage of doing this is that the trained state distribution recognition model learns the standard state and standard state features of an object under normal behavior. Subsequently, the standard state and standard state features can be used as a benchmark to achieve accurate evaluation of the object's state.

[0082] The above describes the training process of the state distribution recognition model. The following section will introduce the state analysis method described in steps 210 to 250. That is, the state analysis method will be described from the perspective of the application scenarios of the trained state distribution recognition model.

[0083] In step 210, video data collected by the wearable device during the wearing process is acquired.

[0084] For example, wearable devices could be smart glasses that can acquire video data from a first-person perspective collected by the wearable device. This video data is acquired within a preset time period most recent to ensure the real-time nature of the video data.

[0085] In step 220, feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features.

[0086] This involves extracting features from video data to obtain scene features, object features, body features, and viewpoint motion features. For example, scene features include descriptions of the scene in which the object appears; object features include descriptions of objects within the scene; body features include descriptions of the object's body movements; and viewpoint motion features include descriptions of the angle and frequency of viewpoint changes.

[0087] In some implementations, feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features, including: (1.1) Perform privacy filtering on the video data to obtain the filtered first video data; (1.2) Perform unclear image filtering on the first video data to obtain the filtered second video data; (1.3) Normalize the second video data according to the preset screen parameters to obtain normalized target video data; (1.4) Extract features from the target video data to obtain scene features, object features, limb features and viewpoint motion features.

[0088] For example, a lightweight privacy recognition algorithm (based on a pre-trained face detection model and a sensitive region recognition model) is used to detect privacy-sensitive regions in a video frame by frame. For example, regarding face privacy: it detects all faces in the video that are not the wearer's own (by comparing against a pre-set face template, excluding the wearer's own face), including faces of passersby during commutes and faces of colleagues at the office. Regarding identifier privacy: it detects sensitive identifier information in the video, including home address numbers (a combination of numbers and text), account passwords displayed on the office computer screen, photocopies of ID cards on the desktop, and content on the mobile phone screen.

[0089] Then, Gaussian blurring is applied to the privacy areas to obtain the first video data. After privacy filtering is performed on all frames, the processed frames are stitched together in the original acquisition order to generate a complete video file, which is the filtered first video data. This first video data has removed all privacy-sensitive content and only retains core information such as scene, objects, wearer's body movements, and viewpoint motion. The frame rate and resolution are consistent with the original video.

[0090] Then, unclear scenes in the first video data are removed. For example, the grayscale variance (a clarity evaluation index) of each frame is calculated. The larger the grayscale variance, the clearer the picture. The preset clarity threshold is 80 (grayscale variance ≥ 80 is considered clear, < 80 is considered unclear). The grayscale variance is calculated frame by frame to filter out unclear scenes, which mainly include three categories: dark scenes caused by backlighting (grayscale variance < 50), blur caused by camera shake when walking (grayscale variance 50-79), and black / partial black screen caused by the camera being blocked (grayscale variance < 30).

[0091] Then, the unclear images selected are removed according to the following rules: single unclear frames are removed directly; for continuous unclear images (such as 5 or more consecutive frames), all consecutive frames are removed to avoid interference from invalid data; the remaining clear images are re-stitched according to the original acquisition time order, and inter-frame transitions are added (to avoid image stuttering) to generate the filtered second video data.

[0092] Then, the image parameters of the second video data are unified to eliminate the feature extraction deviation caused by differences in brightness, contrast, and resolution, and to ensure the consistency of subsequent feature extraction. This is achieved by normalizing the data, such as resolution normalization, frame rate normalization, brightness and contrast normalization, etc. Finally, after all parameters are normalized, normalized target video data is generated.

[0093] Finally, feature extraction is performed on each group of target sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each group of sample video data. For example, scene features: represent the environmental scene corresponding to the frame, such as "home restaurant", "subway car", "office", "outdoor sidewalk", and the vector value corresponds to the feature similarity of the scene; Object features: Characterize the core objects appearing in this frame, such as "bowl and chopsticks", "mobile phone", "computer" and "briefcase". The vector values ​​correspond to the object's category, quantity and location information. Body characteristics: Characterize the wearer's body movements, such as "grabbing", "tapping", "chewing" and "walking", with vector values ​​corresponding to the movement trajectory and range of motion of the limb joints; Visual motion characteristics: characterize the wearer's head / eye movement, such as "turning head", "looking down", "looking straight ahead", and "eye movement". The vector values ​​correspond to the frequency of visual rotation, coverage area, and position of eye gaze.

[0094] The advantage of doing this is that the four types of core features extracted fully cover the key information of scene, object, body movement, and viewpoint motion, providing a high-quality data foundation for subsequent target feature generation and state distribution recognition model input.

[0095] In some implementations, feature extraction is performed on the target video data to obtain scene features, object features, limb features, and viewpoint motion features, including: (1.4.1) Determine multiple target images in the target video data according to a preset time interval; (1.4.2) Input each target image into a pre-trained image recognition model to output the scene features and object features corresponding to the video data; (1.4.3) Input each target image into the pre-trained action recognition model to output the limb features and viewpoint motion features corresponding to the video data.

[0096] For example, representative images are selected from normalized target video data at preset time intervals as the core carriers for subsequent feature extraction, balancing feature integrity and computational efficiency. For instance, one target image is selected every second.

[0097] Each target image is then input into a pre-trained image recognition model to output scene and object features corresponding to the video data. For example, the model identifies the scene category corresponding to the frame image based on environmental details (such as wall decorations, furniture layout, and outdoor scenery) and outputs scene features. Simultaneously, the model identifies the core objects (objects related to the wearer's behavior) appearing in each target image frame and outputs object features.

[0098] Each target image is input into a pre-trained action recognition model to output limb features and viewpoint motion features corresponding to the video data. For example, the model identifies the movement trajectory and range of motion of the wearer's limb joints (such as hands, arms, and shoulders) and outputs limb features. The model determines the viewpoint motion state (such as looking straight ahead, looking down, or turning the head) by recognizing the wearer's head posture and eye gaze direction, and outputs viewpoint motion features.

[0099] The advantage of doing this is that it allows for the rapid and accurate extraction of scene features, object features, limb features, and viewpoint motion features.

[0100] In some implementations, target features are generated based on scene features, object features, limb features, and viewpoint motion features, including: (2.1) Determine the acquisition time sequence corresponding to limb features and perspective motion features, and generate the temporal features corresponding to the video data based on the acquisition time sequence; (2.2) Generate target features based on temporal features, scene features, object features, limb features and perspective motion features.

[0101] For example, by extracting limb features, viewpoint motion features, and their associated acquisition timestamps, multiple sets of "limb feature-viewpoint motion feature-timestamp" data pairs are obtained. These timestamps are arranged in sequence to generate temporal features corresponding to the video data. For instance, each set of "limb feature-viewpoint motion features" after sorting is assigned a temporal dimension identifier: taking the video's starting frame as the temporal starting point, the units are sequentially designated as the 1st temporal unit, the 2nd temporal unit, ..., the nth temporal unit, with each temporal unit corresponding to the features of one frame.

[0102] Then, based on the timestamp matching principle, the temporal feature sub-vectors, scene feature sub-vectors, object feature sub-vectors, limb feature sub-vectors, and viewpoint motion feature sub-vectors corresponding to the same timestamp (same frame) are concatenated dimensionally. According to the acquisition time order, the concatenated feature vectors corresponding to each sampling time are sequentially arranged to form a continuous feature sequence in chronological order; this feature sequence is the target feature. In this embodiment, the target feature is a temporal feature matrix. This matrix contains multi-dimensional visual feature information of the scene, object, limb, and viewpoint motion, as well as temporal pattern information of feature changes over time, which can be directly input into the trained state distribution recognition model for state analysis.

[0103] The advantage of this approach is that, by assigning a time sequence to limb features and perspective motion features and generating temporal features, this embodiment achieves the representation of dynamic behavioral patterns in video data. Furthermore, by using a method of concatenating multiple features with the same timestamp and serializing them in a time sequence, target features are generated, allowing the target features to simultaneously contain static visual feature information and dynamic temporal pattern information. This solves the problem that single visual features lack a time dimension and cannot represent behavioral rhythms, providing a complete feature data foundation for subsequent state distribution recognition models to accurately analyze the behavioral patterns, health status, and psychological state of objects.

[0104] In some implementations, target features are generated based on temporal features, scene features, object features, limb features, and viewpoint motion features, including: (2.2.1) The temporal features, scene features, object features, limb features and viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time; (2.2.2) Each splicing feature is spliced ​​sequentially according to the collection time order to generate the target feature.

[0105] For example, the acquisition time sequence is (T1→T2→T3→……→Tn), and each acquisition time corresponds to time-series features, scene features, object features, limb features, and viewpoint motion features. Taking acquisition time T1 as an example, its corresponding time-series features, scene features, object features, limb features, and viewpoint motion features are concatenated to obtain a concatenated feature F1. By analogy, the concatenated features corresponding to each sampling time can be determined.

[0106] Finally, each spliced ​​feature is sequentially spliced ​​according to the acquisition time order to generate the target feature. This target feature retains the multi-dimensional visual feature information of each acquisition time in the original video. At the same time, through the sequential splicing of the acquisition time order, it fully preserves the temporal pattern and behavioral continuity information of the feature changes over time. It can be directly input into the trained state distribution recognition model for subsequent state feature output and state determination.

[0107] The advantage of doing this is that by using a single acquisition time as the processing unit to stitch together five types of features, the integrity of multiple feature information in the same time dimension is ensured, allowing the model to accurately capture the wearer's comprehensive state of scene, behavior, movement, etc. at a certain moment. Based on the original acquisition time sequence, the spliced ​​features are serialized and spliced ​​to achieve continuous representation of feature information in the time dimension. The temporal rhythm of the wearer's behavior (such as continuous changes in limb movements, temporal switching of scenes, etc.) is fully preserved, and the problem of lack of temporal correlation of single features is solved. All features are concatenated as vectors of uniform dimension, resulting in standardized numerical vectors of target features that can be directly adapted to the input requirements of state distribution recognition models without the need for additional feature transformation processing, thus improving the efficiency of feature processing and model analysis.

[0108] In some implementations, the temporal features, scene features, object features, limb features, and viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time, including: (2.2.1.1) Determine the scene corresponding to each collection time based on scene characteristics; (2.2.1.2) Based on the scene, the object features, limb features and view motion features are weighted and processed respectively to obtain updated object features, updated limb features and updated view motion features; (2.2.1.3) The temporal features, scene features, updated object features, updated limb features and updated viewpoint motion features at each acquisition time are spliced ​​together to obtain the spliced ​​features at each acquisition time.

[0109] The emphasis on object features, limb features, and viewpoint motion features varies depending on the scenario. For example, in a home setting, the focus is primarily on object features and viewpoint motion features, while limb features receive less weight. In a motion scenario, the focus is primarily on limb features and viewpoint motion features, while object features receive less weight.

[0110] Therefore, the scene corresponding to each collection time can be determined based on scene characteristics. Then, object features, limb features, and viewpoint motion features are weighted according to the scene to obtain updated object features, updated limb features, and updated viewpoint motion features. For example, in a home scene, the weights for object features, limb features, and viewpoint motion features can be set to 0.3, 0.2, and 0.3, respectively. Then, the object features, limb features, and viewpoint motion features are weighted separately according to these weights.

[0111] Finally, the temporal features, scene features, updated object features, updated limb features, and updated viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time.

[0112] The advantages of this approach are that it accurately determines the scene at each collection time based on scene characteristics, ensuring that the weighted processing fits the scene characteristics and solves the problem of inaccurate feature representation caused by "consistent feature importance in different scenes" (e.g., in outdoor scenes, the focus is on strengthening the perspective motion features to match the frequent movement of the wearer's gaze outdoors). The weighted processing retains the core information of the original features, while highlighting key scene features through weight allocation, so that the updated features can better reflect the wearer's behavioral characteristics in the scene and improve the accuracy of subsequent model analysis.

[0113] In step 240, the target features are input into the trained state distribution recognition model, and the current state features and the corresponding standard state corresponding to the video data are output. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data.

[0114] The target features can be input into the trained state distribution recognition model, which outputs the current state features and the corresponding standard state of the video data. The current state features can be understood as a vector output by the trained state distribution recognition model after analyzing the target features. The standard state is actually the theoretically closest normal state to the target features, such as the state of running normally.

[0115] In step 250, the state difference between the current state feature and the standard state feature corresponding to the standard state is determined, and the current state of the object is determined based on the state difference.

[0116] Then, determine the state difference between the current state features and the standard state features corresponding to the standard state. For example, the current state features (F_current): a 320-dimensional temporal feature vector, corresponding to the average value of the stitched features of the wearer within the current hour (collection time T11-T1320, 0.5-second intervals), simplified to the first 5 dimensions: [0.25, 0.36, 0.24, 0.35, 0.26,...]. The standard state features (F_standard): a 320-dimensional temporal feature vector, corresponding to the average value of the stitched features of the wearer in the normal state, simplified to the first 5 dimensions: [0.261, 0.364, 0.246, 0.364, 0.27,...] (consistent with the average value of the stitched features after the update at time T1, fitting the normal home / office scenario). Then, determine the Euclidean distance between the two types of feature vectors as the difference, for example, a value of 0.12.

[0117] Then, based on the state difference, the current state of the object is determined, that is, its deviation from the normal standard state is determined, thereby determining the state of the object. For example, in a motion scene, if the analysis shows that the difference is greater than a preset difference value, the object is considered to be in a state of excessive motion.

[0118] In some implementations, the current state of an object is determined based on state differences, including: (1.1) Determine the current behavior pattern of the object based on the state differences; (1.2) Obtain multiple historical behavior patterns of the object and determine the trend of behavior pattern change from multiple historical behavior patterns to the current behavior pattern; (1.3) Determine the health status of the object based on the preset health status and behavioral pattern change trend corresponding to the standard state.

[0119] For example, the state difference is 0.22 (slightly abnormal range), and the core differences are concentrated in limb features (0.28) and viewpoint motion features (0.25), indicating that the current limb activity and viewpoint rotation deviate significantly from the standard state, while the scene (home) and object interaction (water cup, chair) do not deviate significantly.

[0120] Matching Behavioral Patterns: The model's preset "behavioral pattern-feature difference mapping relationship" is invoked to compare the current state features and differences with the standard features of the three types of behavioral patterns: Compared with "normal home mode": the current value of limb features (normal range 0.6-0.8) is 0.32, and the current value of visual motion features (normal range 0.5-0.7) is 0.25. The difference exceeds the normal fluctuation, so there is no match; Compared with "normal morning exercise mode": the scene feature is home (morning exercise scene is outdoors), so there is no match; Compared with "abnormal static mode": the limb features (0.32) and visual motion features (0.25) are both within the feature range of this mode (limb 0.2-0.4, visual 0.2-0.3), and the scene and object features have no obvious deviation, so there is a perfect match.

[0121] Determine the current behavior pattern: Based on the above comparison results, the current behavior pattern of the object is determined to be "abnormal static mode", which is specifically manifested as: sitting quietly in a chair for a long time in the home environment, with no obvious hand interaction, fixed perspective (looking straight ahead at the wall), and a limb movement frequency significantly lower than the standard state, only occasionally raising the hand to drink water, without any other active behavior.

[0122] Then, the device's storage was used to retrieve the wearer's historical behavioral patterns for the past 7 days (integrated into 3 time periods per day, taking the dominant behavioral pattern of each day), which were summarized as follows: Day 1 (dominant pattern: normal morning exercise + normal home): meets the standard state, no abnormalities; Day 2 (dominant pattern: normal morning exercise + normal home): consistent with Day 1, stable behavioral pattern; Day 3 (dominant pattern: normal morning exercise + abnormal static): the abnormal static pattern appeared for the first time, accounting for 1 / 3 of the time period of the day; Day 4 (dominant pattern: abnormal static): the abnormal static pattern accounted for 2 / 3 of the time period of the day, and the morning exercise pattern disappeared; Day 5-Day 6 (dominant pattern: abnormal static): the abnormal static pattern accounted for all time periods of the day, with no normal pattern.

[0123] Analyze the changing trend: Using time as the horizontal axis and behavior pattern as the vertical axis, trace the change trajectory from the historical pattern to the current pattern: Change trajectory: Normal morning exercise + normal home (Day 1-Day 2) → A small amount of abnormal static state (Day 3) → A large amount of abnormal static state (Day 4) → Continuous abnormal static state (Day 5-Day 6) → Current abnormal static state (Day 7). Determining the trend of change: Based on trajectory analysis, the current behavioral pattern is showing a "continuous deterioration trend"—from a completely normal behavioral pattern, abnormalities gradually appear, and the duration and frequency of abnormal patterns are constantly increasing, while normal behavioral patterns (morning exercise, normal home life) are gradually disappearing, with no signs of recovery.

[0124] Retrieve preset health status: From the status distribution recognition model, retrieve the preset health status corresponding to the standard status (the wearer's normal daily status) - "healthy and stable status". This status corresponds to the normal morning exercise + normal home behavior pattern, and the limb activity and visual movement are within the normal range. The current trend is "continuous deterioration trend". The behavior pattern is gradually changing from normal to continuous abnormal static. The difference in state is within the range of slight abnormality (overall difference 0.22), which has not reached the level of obvious abnormality, but the deterioration trend is obvious. Considering the physiological characteristics of middle-aged and elderly wearers (normal morning exercise is crucial for health, and continuous static may be associated with reduced limb activity and poor mental state), further analysis shows that the continuous abnormal static pattern is accompanied by obvious deviations in limb characteristics and visual movement characteristics, which does not conform to the behavioral requirements of "healthy and stable state".

[0125] Therefore, the subject's current health status is determined to be "mildly abnormal health". The status description is as follows: The current health status deviates from the standard stable state. The core reason is the continuous deterioration of the behavior pattern (from normal to continuous abnormal static). It is necessary to pay attention to the physical activity. It is recommended to resume normal behaviors such as morning exercise. Continue to monitor changes in behavior pattern. If it deteriorates further, professional intervention should be requested.

[0126] Once a person's health status is determined, wearable devices can be used to provide them with corresponding health reminders.

[0127] In some implementations, the current state of an object is determined based on state differences, including: (2.1) Determine the standard psychological state rating of the subject under standard conditions; (2.2) Determine the adjustment score value based on the preset scoring criteria and state differences; (2.3) Update the standard psychological state score based on the adjustment score to obtain the target psychological state score; (2.4) Determine the psychological state of the subject based on the pre-defined mapping relationship between the psychological state score and the psychological state, as well as the target psychological state score.

[0128] For example, the pre-defined mapping relationship between psychological state scores and psychological states is: 85-100 points → good; 70-84 points → calm; 55-69 points → mild anxiety; ≤54 points → moderate anxiety.

[0129] The current scenario is a home setting. The subject's behavior is characterized by frequent small movements and wandering gaze, which differs from the standard home state (slow movements and stable gaze). The preset "standard state-score" mapping relationship is retrieved. This mapping is generated based on the subject's past normal home samples. Combined with the current home scenario, the standard psychological state score is determined to be 80 points, corresponding to the "peaceful mindset" benchmark.

[0130] The state difference is 0.23, and the core single-dimensional differences are (limb 0.28, visual-motor 0.25). The deduction values ​​are calculated as follows: 23 points for state difference (0.23 × 100), and an additional deduction of (0.28 + 0.25) × 80 = 42.4 points for the core single-dimensional differences. The resulting moderating score is: -23 -42.4 = -65.4 points.

[0131] Finally, the target score was calculated as "target score = standard score + adjustment score", which is 80 + (-65.4) = 14.6 points. Based on the preset mapping relationship, the target score of 14.6 points ≤ 54 points, therefore the subject's current psychological state is determined to be "moderate anxiety".

[0132] Once the psychological state is determined, wearable devices can be used to provide corresponding mental health reminder services to the individual.

[0133] In this embodiment, video data collected by the wearable device during its wearing process is acquired; features are extracted from the video data to obtain scene features, object features, limb features, and viewpoint motion features; target features are generated based on the scene features, object features, limb features, and viewpoint motion features; the target features are input into a trained state distribution recognition model, which outputs the current state features corresponding to the video data and the corresponding standard state. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data; the state difference between the current state features and the standard state features corresponding to the standard state is determined, and the current state of the object is determined based on the state difference.

[0134] Therefore, by acquiring video data collected during the wearing of a wearable device, feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features. Target features are then generated based on these features. This enables multi-dimensional feature acquisition, and the combination of these multi-dimensional features generates richer target features. The target features are then input into a trained state distribution recognition model, which outputs the current state features and corresponding standard state of the video data. The trained state distribution recognition model is based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data. Therefore, the trained state distribution recognition model in this application is highly adaptable to the current object and can analyze target features to accurately output the current state features and corresponding standard state of the video data. Finally, the state difference between the current state features and the standard state features corresponding to the standard state is determined, and the current state of the object is determined based on this difference. In this way, by determining the state difference between the current state characteristics and the standard state characteristics corresponding to the normal standard state, the state difference can accurately reflect the deviation between the current state and the standard state, thereby accurately determining the current state of the object.

[0135] Compared to related technologies that analyze sensor data using algorithms to provide services, this application can more accurately determine the current state characteristics of an object by using a trained state distribution recognition model specific to the object, and accurately determine the current state of the object based on the difference between the current state characteristics and the standard state characteristics, thereby providing more precise services to the object based on the current state.

[0136] Please see Figure 5 , Figure 5This is another schematic flowchart of the state analysis method provided in this application embodiment. The state analysis method provided in this application embodiment may include the following steps: Step 401: Acquire video data collected during the wearing of the wearable device; Step 402: Perform privacy filtering on the video data to obtain the filtered first video data; Step 403: Perform unclear image filtering on the first video data to obtain the filtered second video data; Step 404: Normalize the second video data according to the preset screen parameters to obtain normalized target video data; Step 405: Extract features from the target video data to obtain scene features, object features, limb features, and viewpoint motion features; Step 406: Determine the acquisition time sequence corresponding to limb features and perspective motion features, and generate temporal features corresponding to video data based on the acquisition time sequence; Step 407: Determine the scene corresponding to each collection time based on scene characteristics; Step 408: Perform weighted processing on object features, limb features and viewpoint motion features according to the scene to obtain updated object features, updated limb features and updated viewpoint motion features; Step 409: The temporal features, scene features, updated object features, updated limb features, and updated viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time. Step 410: Concatenate each feature sequentially according to the acquisition time order to generate the target feature; Step 411: Input the target features into the trained state distribution recognition model, and output the current state features and the corresponding standard state corresponding to the video data; Step 412: Determine the state difference between the current state characteristics and the standard state characteristics corresponding to the standard state, and determine the current state of the object based on the state difference.

[0137] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed description of the state analysis method above, which will not be repeated here.

[0138] Please see Figure 6 , Figure 6This is a schematic diagram of the state analysis device provided in an embodiment of this application, which can execute the state analysis method described above. In this embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functions of that module or unit.

[0139] Condition analysis device 500, including: The acquisition module 510 is used to acquire video data collected by the wearable device during the wearing process; The extraction module 520 is used to extract features from video data to obtain scene features, object features, limb features, and viewpoint motion features; The generation module 530 is used to generate target features based on scene features, object features, limb features, and viewpoint motion features; The input module 540 is used to input the target features into the trained state distribution recognition model and output the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data. The determination module 550 is used to determine the state difference between the current state characteristics and the standard state characteristics corresponding to the standard state, and to determine the current state of the object based on the state difference.

[0140] In some implementations, the generation module 530 is used for: The acquisition time sequence corresponding to limb features and perspective motion features is determined, and temporal features corresponding to video data are generated based on the acquisition time sequence. Target features are generated based on temporal features, scene features, object features, limb features, and perspective motion features.

[0141] In some implementations, the generation module 530 is used for: The temporal features, scene features, object features, limb features, and viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time. Each feature is sequentially spliced ​​together according to the acquisition time order to generate the target feature.

[0142] In some implementations, the generation module 530 is used for: The scene corresponding to each collection time is determined based on scene characteristics; Based on the scene, the object features, limb features and viewpoint motion features are weighted separately to obtain updated object features, updated limb features and updated viewpoint motion features; The temporal features, scene features, updated object features, updated limb features, and updated viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time.

[0143] In some implementations, the extraction module 520 is used for: The video data is subjected to privacy filtering to obtain the first filtered video data; The first video data is filtered to remove unclear images, resulting in the filtered second video data. The second video data is normalized according to the preset screen parameters to obtain the normalized target video data. Feature extraction is performed on the target video data to obtain scene features, object features, limb features, and viewpoint motion features.

[0144] In some implementations, the extraction module 520 is used for: Multiple target images are determined in the target video data according to a preset time interval; Each target image is input into a pre-trained image recognition model to output scene features and object features corresponding to the video data; Each target image is input into a pre-trained action recognition model to output the limb features and viewpoint motion features corresponding to the video data.

[0145] In some implementations, the determining module 550 is used for: Determine the current behavior pattern of the object based on the state differences; Obtain multiple historical behavior patterns of an object and determine the trend of behavior pattern changes from multiple historical behavior patterns to the current behavior pattern; The object's health status is determined based on the preset health status and behavioral pattern change trends corresponding to the standard state.

[0146] In some implementations, the determining module 550 is used for: Determine the standard psychological state rating value of the object under standard conditions; The adjustment score value is determined based on the preset scoring criteria and the differences in status; The target psychological state score is obtained by updating the standard psychological state score based on the adjustment score; The psychological state of the subject is determined based on the pre-defined mapping relationship between psychological state scores and psychological states, as well as the target psychological state score.

[0147] In some embodiments, the state analysis device 500 further includes a training module for: Before inputting the target features into the trained state distribution recognition model and outputting the current state features and corresponding standard states corresponding to the video data, multiple sets of sample video data collected by the wearable device over a historical period and the label state corresponding to each set of sample video data are obtained. The sample video data is the data corresponding to the object in the standard state. Feature extraction is performed on each set of sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each set of sample video data; Sample features are generated based on sample scene features, sample object features, sample limb features, and sample viewpoint motion features; The sample features are input into the state distribution recognition model, and the predicted state corresponding to the sample video data is output. The state distribution recognition model is iteratively trained based on the difference between the predicted state and the labeled state to obtain the trained state distribution recognition model.

[0148] In some implementations, the training module is used for: Each set of sample video data is subjected to privacy filtering to obtain the first sample video data after filtering. For each group of first sample video data, perform unclear image filtering to obtain filtered second sample video data; The second sample video data of each group is normalized according to the preset screen parameters to obtain the normalized target sample video data corresponding to each group of second sample video data. Feature extraction is performed on each set of target sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each set of sample video data.

[0149] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed description of the state analysis method above, which will not be repeated here.

[0150] In this embodiment, the acquisition module 510 acquires video data collected by the wearable device during wear; the extraction module 520 extracts features from the video data to obtain scene features, object features, limb features, and viewpoint motion features; the generation module 530 generates target features based on the scene features, object features, limb features, and viewpoint motion features; the input module 540 inputs the target features into the trained state distribution recognition model and outputs the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data; the determination module 550 determines the state difference between the current state features and the standard state features corresponding to the standard state, and determines the current state of the object based on the state difference.

[0151] Therefore, by acquiring video data collected during the wearing of a wearable device, feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features. Target features are then generated based on these features. This enables multi-dimensional feature acquisition, and the combination of these multi-dimensional features generates richer target features. The target features are then input into a trained state distribution recognition model, which outputs the current state features and corresponding standard state of the video data. The trained state distribution recognition model is based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data. Therefore, the trained state distribution recognition model in this application is highly adaptable to the current object and can analyze target features to accurately output the current state features and corresponding standard state of the video data. Finally, the state difference between the current state features and the standard state features corresponding to the standard state is determined, and the current state of the object is determined based on this difference. In this way, by determining the state difference between the current state characteristics and the standard state characteristics corresponding to the normal standard state, the state difference can accurately reflect the deviation between the current state and the standard state, thereby accurately determining the current state of the object.

[0152] Compared to related technologies that analyze sensor data using algorithms to provide services, this application can more accurately determine the current state characteristics of an object by using a trained state distribution recognition model specific to the object, and accurately determine the current state of the object based on the difference between the current state characteristics and the standard state characteristics, thereby providing more precise services to the object based on the current state.

[0153] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned state analysis method. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0154] Please see Figure 7 , Figure 7 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the state analysis method of the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0155] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described state analysis method.

[0156] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0157] The state analysis method, state analysis device, computer equipment, and storage medium provided in this application acquire video data collected by a wearable device during wear; extract features from the video data to obtain scene features, object features, limb features, and viewpoint motion features; and generate target features based on these features. This enables multi-dimensional feature acquisition, and the combination of these multi-dimensional features generates richer target features. The target features are then input into a trained state distribution recognition model, which outputs the current state features and corresponding standard states corresponding to the video data. The trained state distribution recognition model is trained based on the label states of sample video data corresponding to the object, as well as sample scene features, sample object features, sample limb features, and sample viewpoint motion features in the sample video data. Therefore, the trained state distribution recognition model in this application is highly adaptable to the current object and can analyze target features to accurately output the current state features and corresponding standard states corresponding to the video data. Finally, the state difference between the current state features and the standard state features corresponding to the standard state is determined, and the current state of the object is determined based on this difference. In this way, by determining the state difference between the current state characteristics and the standard state characteristics corresponding to the normal standard state, the state difference can accurately reflect the deviation between the current state and the standard state, thereby accurately determining the current state of the object.

[0158] Compared to related technologies that analyze sensor data using algorithms to provide services, this application can more accurately determine the current state characteristics of an object by using a trained state distribution recognition model specific to the object, and accurately determine the current state of the object based on the difference between the current state characteristics and the standard state characteristics, thereby providing more precise services to the object based on the current state.

[0159] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0160] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0162] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0163] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0164] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0165] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0166] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0167] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0168] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0169] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A state analysis method, characterized in that, include: Acquire video data collected during the wearing of wearable devices; Feature extraction is performed on the video data to obtain scene features, object features, limb features, and viewpoint motion features; Target features are generated based on the scene features, object features, limb features, and viewpoint motion features; The target features are input into the trained state distribution recognition model, which outputs the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data. Determine the state difference between the current state feature and the standard state feature corresponding to the standard state, and determine the current state of the object based on the state difference.

2. The state analysis method according to claim 1, characterized in that, The step of generating target features based on the scene features, object features, limb features, and viewpoint motion features includes: The acquisition time sequence corresponding to the limb features and the viewpoint motion features is determined, and the temporal features corresponding to the video data are generated according to the acquisition time sequence. Target features are generated based on the temporal features, scene features, object features, limb features, and viewpoint motion features.

3. The state analysis method according to claim 2, characterized in that, The step of generating target features based on the temporal features, scene features, object features, limb features, and viewpoint motion features includes: The temporal features, scene features, object features, limb features, and viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time. Each of the splicing features is sequentially spliced ​​according to the acquisition time sequence to generate the target feature.

4. The state analysis method according to claim 3, characterized in that, The process of stitching together the temporal features, scene features, object features, limb features, and viewpoint motion features at each acquisition time to obtain the stitched features at each acquisition time includes: The scene corresponding to each collection time is determined based on the scene characteristics; Based on the scene, the object features, the limb features, and the viewpoint motion features are weighted and processed to obtain updated object features, updated limb features, and updated viewpoint motion features; The temporal features, scene features, updated object features, updated limb features, and updated viewpoint motion features at each acquisition time are stitched together to obtain the stitched features at each acquisition time.

5. The state analysis method according to claim 1, characterized in that, The step of extracting features from the video data to obtain scene features, object features, limb features, and viewpoint motion features includes: The video data is subjected to privacy filtering to obtain the filtered first video data; The first video data is subjected to blurry image filtering to obtain the filtered second video data. The second video data is normalized according to preset screen parameters to obtain normalized target video data; Feature extraction is performed on the target video data to obtain scene features, object features, limb features, and viewpoint motion features.

6. The state analysis method according to claim 5, characterized in that, The step of extracting features from the target video data to obtain scene features, object features, limb features, and viewpoint motion features includes: Multiple target images are determined from the target video data according to a preset time interval; Each target image is input into a pre-trained image recognition model to output the scene features and object features corresponding to the video data; Each target image is input into a pre-trained action recognition model to output the limb features and viewpoint motion features corresponding to the video data.

7. The state analysis method according to claim 1, characterized in that, Determining the current state of the object based on the state differences includes: The current behavior pattern of the object is determined based on the state differences. Obtain multiple historical behavior patterns of an object, and determine the trend of behavior pattern changes from the multiple historical behavior patterns to the current behavior pattern; The object's health status is determined based on the preset health status corresponding to the standard state and the trend of behavioral pattern changes.

8. The state analysis method according to claim 1, characterized in that, Determining the current state of the object based on the state differences includes: Determine the standard psychological state score of the object under the standard state; The adjustment score value is determined based on the preset scoring criteria and the state differences; The target psychological state score is obtained by updating the standard psychological state score based on the adjusted score. The psychological state of the object is determined based on the psychological state score and the preset mapping relationship between psychological states, as well as the target psychological state score.

9. The state analysis method according to claim 1, characterized in that, Before inputting the target features into the trained state distribution recognition model and outputting the current state features and corresponding standard states corresponding to the video data, the method further includes: The wearable device acquires multiple sets of sample video data collected over a historical period, as well as the tag status corresponding to each set of sample video data. The sample video data is the data corresponding to the object in a standard state. Feature extraction is performed on each group of sample video data to obtain the sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each group of sample video data; Sample features are generated based on the sample scene features, the sample object features, the sample limb features, and the sample viewpoint motion features; The sample features are input into the state distribution recognition model, and the predicted state corresponding to the sample video data is output. The state distribution recognition model is iteratively trained based on the difference between the predicted state and the labeled state to obtain the trained state distribution recognition model.

10. The state analysis method according to claim 9, characterized in that, The step of extracting features from each group of sample video data to obtain sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each group of sample video data includes: Each set of sample video data is subjected to privacy filtering to obtain the first filtered sample video data. For each group of first sample video data, perform unclear image filtering to obtain filtered second sample video data; The second sample video data of each group is normalized according to the preset screen parameters to obtain the normalized target sample video data corresponding to each group of the second sample video data. Feature extraction is performed on each set of target sample video data to obtain sample scene features, sample object features, sample limb features, and sample viewpoint motion features corresponding to each set of sample video data.

11. A state analysis device, characterized in that, include: The acquisition module is used to acquire video data collected by the wearable device during the wearing process; The extraction module is used to extract features from the video data to obtain scene features, object features, limb features, and viewpoint motion features; The generation module is used to generate target features based on the scene features, the object features, the limb features, and the viewpoint motion features; The input module is used to input the target features into the trained state distribution recognition model and output the current state features and the corresponding standard state corresponding to the video data. The trained state distribution recognition model is trained based on the label state of the sample video data corresponding to the object, as well as the sample scene features, sample object features, sample limb features and sample perspective motion features in the sample video data. The determination module is used to determine the state difference between the current state feature and the standard state feature corresponding to the standard state, and to determine the current state of the object based on the state difference.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to execute the state analysis method according to any one of claims 1 to 10.

13. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the computer program, it implements the state analysis method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement the state analysis method according to any one of claims 1 to 10.