Cognitive-behavior characteristic recording method based on distributed visual equipment
By installing distributed visual devices on multiple parts of the subject's body, collecting and analyzing video images and motion trajectory information, and utilizing deep learning neural networks, the problem of early identification of cognitive impairment was solved, and quantitative recording and analysis of cognitive impairment were achieved.
Patent Information
- Application Number
- CN202510794775.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-14
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies struggle to identify cognitive impairments at an early stage, especially Alzheimer's disease, and there is a lack of effective observation tools and methods to record and analyze the relationship between subjects' cognitive states and behavioral characteristics.
A cognitive-behavioral feature recording method based on distributed vision devices was adopted. By installing image and motion information recording devices on multiple parts of the subject, video images and motion trajectory information were collected. Deep learning neural networks were used to compare features and reflect the differences in behavioral state compared with the normal group.
It enables quantitative recording and analysis of early stages of cognitive impairment, and provides automated tools to compare the behavioral characteristics of subjects with those of normal populations, thus offering new methods for research on cognitive impairment.
Smart Images

Figure CN120895155A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of human behavior measurement, in particular to a cognitive-behavioral feature recording method based on distributed visual equipment. BACKGROUND
[0002] With the deepening of the aging process of the population, the burden on families and society caused by Alzheimer's disease is becoming heavier and heavier. Cognitive impairment is an early process of Alzheimer's disease. Subjects will show cognitive impairment such as memory, spatial orientation ability, reasoning ability, etc. in the first ten years of suffering from Alzheimer's disease. Studies have shown that if cognitive impairment subjects can be intervened as early as possible, the occurrence of Alzheimer's disease can be better delayed.
[0003] However, the manifestations of cognitive impairment are usually very hidden, and there is no obvious difference between the cognitive impairment elderly and the normal elderly in the early stage of the disease course. The scale score used in clinical practice is highly subjective, time-consuming and laborious, and it is difficult to identify early symptoms. In fact, the cognitive impairment of the subjects will inevitably be reflected in the behavior characteristics level. Due to the lack of good observation tools and feature analysis methods, it is difficult to record and analyze the relationship between the cognitive state and the behavior characteristics of the subjects. SUMMARY
[0004] The purpose of the present application is to solve the problems existing in the prior art, provide a cognitive-behavioral feature recording method based on distributed visual equipment, which can quantitatively record the behavior characteristics of the human body in the task state. The method gives the feature difference distribution between the measurement object and the normal group in the form of difference through visual images, and reflects the behavior state of different subjects.
[0005] The purpose of the present application is solved by the following technical scheme:
[0006] A cognitive-behavioral feature recording method based on distributed visual equipment, characterized in that the steps of the recording method are:
[0007] S1, wearing distributed image and motion information recording equipment on multiple parts of the subject according to the requirements;
[0008] S2, in a fixed scene environment, giving the subject a language instruction, requiring the subject to complete the task according to the language instruction;
[0009] S3, the distributed image and motion information recording equipment collects video images and motion trajectory information of the subject completing the task according to the language instruction, and forms a plurality of sets of video images and motion trajectory information pairs corresponding to the parts of the subject based on the language instruction;
[0010] S4, input the video image and motion trajectory information pair corresponding to the selected part from the multiple video image and motion trajectory information pairs in step S3 into the prediction module to obtain the video image and motion trajectory prediction information pair corresponding to the remaining parts based on the language instruction;
[0011] S5, input the video image and motion trajectory information pair corresponding to the remaining parts in step S4 and the video image and motion trajectory prediction information pair corresponding to the remaining parts obtained in step S4 into the image feature and trajectory feature comparison module for one-to-one comparison to obtain the feature difference of the subject;
[0012] S6, the above feature difference is recorded as the cognitive-behavioral feature of the subject.
[0013] The distributed image and motion information recording device in steps S1 and S3 includes a camera, an inertial sensor, a microphone, a single-chip microcomputer processor, a USB interface, a memory card, an indicator light, a Bluetooth module, a clock, and a battery. The camera is a miniature camera capable of shooting video images. The inertial sensor is used to sense and record the acceleration and angular velocity of the distributed image and motion information recording device. The microphone is used to receive sound information of the subject and the surrounding environment. The single-chip microcomputer processor is used to read the camera, microphone, inertial sensor, and clock data of the current time and save them in the memory card. The data saved in the memory card is output through the USB interface or the Bluetooth module. The Bluetooth module is used to transmit data or receive control instructions from external devices. The clock is used to match the clock data of the current time for the camera, inertial sensor, and microphone. The battery is used to power the distributed image and motion information recording device.
[0014] The installation position and requirement of the distributed image and motion information recording device in step S1 are as follows: one distributed image and motion information recording device is installed on the head of the human body, the camera view is parallel to the human body view, and can move freely with the movement of the human head; two distributed image and motion information recording devices are installed on the left and right wrists of the human body, the camera view is perpendicular to the palm center, and ensures that the camera can see the target object when the hand completes the task; two distributed image and motion information recording devices are installed on the surface of the subject's shoes, the view direction points to the front of the human body, and ensures that the subject can see the ground obstacles in front when walking; one distributed image and motion information recording device is installed in front of the subject's abdomen to explore the image of the ground in front of the subject; one distributed image and motion information recording device is installed in front of the subject's chest to shoot the view image when the subject's hand operates.
[0015] The distributed image and motion information recording device in step S1 is installed on the specified part of the subject through magic tape, or a bandage, or embedded in clothing.
[0016] The video sampling rate of the distributed image and motion information recording device in the steps S1 and S3 is 30 Fps, the inertial sensor sampling rate is 100 Hz, and the audio data sampling rate is 48 KHz; the data collected by the distributed image and motion information recording device contains timestamp information of the current time and human body position information where the distributed image and motion information recording device is located.
[0017] The prediction module in the step S4 and the image feature and trajectory feature comparison module in the step S5 are arranged in a feature processing computer, the prediction module is used to receive the collected information of a certain part and give the prediction result of the remaining parts; the image feature and trajectory feature comparison module is used to receive the collected information of the remaining parts and the prediction result of the remaining parts, and output the feature difference of the subject after one-to-one comparison.
[0018] The training data of the prediction module in the step S4 is derived from the data obtained by the healthy population performing the steps S2 and S3.
[0019] The prediction module in the step S4 includes a language encoder module and an image encoder module as input modules, a Transformer neural network module as a data processing module, an image decoder module for decoding the hidden feature vector output by the Transformer neural network module, and a full connection layer for outputting other perspective prediction information at the current time; the language encoder module is used to receive language instructions, the image encoder module is used to receive the information collected by the distributed image and motion information recording device and encode it into tokens data that can be accepted by the Transformer neural network module, the Transformer neural network module processes the received data and outputs a hidden feature vector, the image decoder module decodes the hidden feature vector and outputs other perspective prediction information at the current time through the full connection layer.
[0020] The language encoder module uses a CLIP ViT-B / 32 text encoder to encode the language instructions; the image encoder module uses a pre-trained ViT-Base encoder for encoding; the TransForms neural network module selects a multi-layer GPT-2 transformer block, selects a 24-layer transformer block according to the reference data, and each block has 384 hidden units and 12 attention heads; the image decoder module selects a transformer based on ViT.
[0021] When the prediction module in the step S4 is trained, the steps of weight training are as follows:
[0022] Q1, recruit healthy people as subjects, wear distributed image and motion information recording devices on multiple parts of each subject according to requirements, and set the pre-training weight of the prediction module;
[0023] Q2, in a fixed scene environment, give all subjects the same language instructions, and ask the subjects to complete the task according to the language instructions;
[0024] Q3, the distributed image and motion information recording device collects video images and motion trajectory information of the subjects completing the task according to the language instructions, and forms a plurality of video image and motion trajectory information pairs corresponding to the language instructions and the subjects and the parts of the subjects;
[0025] Q4, input the video image and motion trajectory information pair corresponding to one part of a selected subject from the plurality of video image and motion trajectory information pairs in step Q3 into the prediction module, to obtain video image and motion trajectory prediction information pairs corresponding to other subjects and the remaining parts of the selected subject based on the language instructions;
[0026] Q5, input the video image and motion trajectory information pairs corresponding to other subjects and the remaining parts of the selected subject in step Q4 and the video image and motion trajectory prediction information pairs corresponding to other subjects and the remaining parts of the selected subject obtained in step Q4 into the image feature and trajectory feature comparison module for one-to-one comparison, to obtain the similarity loss of the video image and motion trajectory prediction information pairs corresponding to other subjects and the remaining parts of the selected subject obtained in step Q4 relative to the video image and motion trajectory information pairs corresponding to other subjects and the remaining parts of the selected subject in step Q4;
[0027] Q6, adjust and optimize the pre-training weight of the prediction module in step Q1 according to the minimum similarity loss, return to step Q4 and repeat steps Q4-Q6 until the similarity loss decrease rate tends to 0, and complete the weight training step.
[0028] The similarity loss decrease rate is the difference between the similarity loss of the last cycle and the similarity loss of the current cycle, and the ratio between the similarity loss of the last cycle and the similarity loss of the current cycle.
[0029] Compared with the prior art, the present application has the following advantages:
[0030] The cognitive-behavioral feature recording method based on the distributed visual equipment provided by the application reflects the behavioral characteristics by using the visual images of different body positions of the subject, so as to use the image generation method based on deep learning to construct the neural network learning and the behavioral characteristics of the normal subject, and then compare the video signals generated by the subject when performing the task with the video signals generated by the normal group in the form of video data, and reflect the behavioral differences between the subject and the normal group.
[0031] The application provides a distributed human task state motion process visual image recording device formed by a plurality of distributed image and motion information recording devices and a feature processing computer with a built-in prediction module and an image feature and trajectory feature comparison module, which can automatically compare the feature differences between the cognitive-behavioral characteristics of the subject and the normal people, and provide a new research tool for further understanding the behavioral differences caused by cognitive impairment. BRIEF DESCRIPTION OF DRAWINGS
[0032] ATTACHMENT Figure 1 The structural schematic diagram of the distributed human task state motion process visual image recording device provided by the application is shown in the figure.
[0033] ATTACHMENT Figure 2 The structural schematic diagram of the single distributed image and motion information recording device provided by the application is shown in the figure.
[0034] ATTACHMENT Figure 3 The front view of the wearing scheme of the distributed image and motion information recording device provided by the application is shown in the figure.
[0035] ATTACHMENT Figure 4 The left view of the wearing scheme of the distributed image and motion information recording device provided by the application is shown in the figure.
[0036] ATTACHMENT Figure 5 The functional schematic diagram of the prediction module provided by the application is shown in the figure.
[0037] ATTACHMENT Figure 6 The principle schematic diagram of the prediction module provided by the application is shown in the figure.
[0038] ATTACHMENT Figure 7 The weight training flowchart of the prediction module provided by the application is shown in the figure.
[0039] ATTACHMENT Figure 8 The principle diagram of the cognitive-behavioral feature recording method based on the distributed visual equipment provided by the application is shown in the figure.
[0040] ATTACHMENT Figure 9 The principle diagram of the cognitive-behavioral feature recording method based on the distributed visual equipment provided by the application when using the visual synchronization positioning mapping algorithm is shown in the figure. DETAILED DESCRIPTION
[0041] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that the invention will be thorough and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The same reference numerals in the drawings denote the same or similar structures, and therefore their detailed description will be omitted.
[0042] The terms “a,” “one,” “the,” and “the” are used to indicate the existence of one or more elements / components / etc.; the terms “including” and “having” are used to indicate an open-ended meaning of inclusion and that there may be other elements / components / etc. in addition to the listed elements / components / etc.
[0043] This invention provides, for example Figure 1 The distributed human task-mode motion visual image recording device shown comprises several distributed image and motion information recording devices and a feature processing computer. A single distributed image and motion information recording device is shown below. Figure 2 As shown, this distributed image and motion information recording device includes a camera, inertial sensor, microphone, microcontroller processor, USB interface, memory card, indicator lights, Bluetooth module, clock, and battery. The camera is a miniature camera; the inertial sensor records the device's acceleration and angular velocity; the microphone collects ambient sound; the microcontroller reads data from the camera, microphone, inertial sensor, and the current clock and stores it on the memory card; the USB interface is used for data input and output; and external devices such as mobile phones and computers can control the device via Bluetooth. This distributed image and motion information recording device can be installed on a designated part and orientation of the subject using Velcro, straps, or embedded in clothing.
[0044] like Figure 3 , Figure 4As shown, seven distributed image and motion information recording devices are installed on different positions of the human body. One of the distributed image and motion information recording devices is installed on the head of the human body, the camera view is parallel to the human body view, and can move freely with the movement of the human head; two distributed image and motion information recording devices are installed on the left and right wrists of the human body, respectively, and the camera view is perpendicular to the palm center outward, ensuring that the camera can see the target object when the hand completes the task; two distributed image and motion information recording devices are installed on the shoe of the subject, respectively, and the view direction points to the front of the human body, ensuring that the ground obstacles in front can be seen when the person walks; one distributed image and motion information recording device is installed in front of the subject's abdomen, which is used to explore the image of the ground in front of the subject at a distance; one distributed image and motion information recording device is installed in front of the subject's chest, which is used to shoot a larger field of view image when the subject's hand operates. After installation, the seven distributed image and motion information recording devices can shoot the video of the position, as well as the audio and inertial sensor data, all data are saved at a fixed sampling rate, preferably the video sampling rate is 30Fps, the inertial sensor sampling rate is 100Hz, and the audio data sampling rate is 48KHz. The collected data contains the timestamp information of the current time and the position information of the distributed image and motion information recording device on the human body.
[0045] As Figures 5-7 shown, the principle, composition, and weight training process of the prediction module built-in the feature processing computer are the features of the present application. The data of the distributed image and motion information recording device can be imported into the feature processing computer. The feature processing computer contains a prediction module, an image feature and trajectory feature comparison module, which can predict the video image information that the distributed image and motion information recording device at other body positions should collect according to the video image information obtained by one or any of the distributed image and motion information recording devices on the subject in the task state. The formation of the prediction module includes data collection, module construction, and weight training, which are as follows:
[0046] A, data collection
[0047] The subject wears 7 distributed image and motion information recording devices according to the above requirements, language instructions are input to the subject in a fixed scene environment, and the subject completes the language instructions; the language instructions, video images and motion trajectory information of the 7 recording devices worn by the subject are obtained through the distributed image and motion information recording devices. By collecting data of multiple different subjects under different language instructions, multiple sets of video image and motion trajectory information pairs at different positions of the subjects can be formed. Since the video image and motion trajectory information pairs come from the perspective of the coordinated movement of the limbs of the subjects under task driving, and the scene environment is also fixed, there is a correlation in information between the video image and motion trajectory information pairs. The correlation information between the video image and motion trajectory information pairs is mined through a deep neural network, so as to realize the prediction of the video image and motion trajectory information pairs at other positions by using the video image and motion trajectory information pairs at one or several positions.
[0048] B, construction of the prediction module
[0049] The input of the prediction module is the language instruction in the experiment process of the subject, the video image and motion trajectory of the distributed image and motion information recording device, and the position number corresponding to the video image and motion trajectory. All video image and motion trajectory information, language instruction information is processed by the encoder into tokens data (input representation vector) acceptable to the TransForms network. Among them, the language encoder module can use the CLIP (Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.) ViT-B / 32 text encoder to encode the language instruction, and the image encoder module uses the pre-trained ViT-Base (Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022.) encoder for encoding. All language instructions, video images and motion trajectory information pass through the encoder. The TransForms neural network module can be selected as a multi-layer GPT-2 transformer block. According to the reference data, a 24-layer transformer block can be selected, each block can have 384 hidden units and 12 attention heads. The TransForms neural network module integrates information from other modules and generates comprehensive latent features to output in the form of a hidden feature vector. The hidden feature vector is input into the VIT image decoder module, which is a transformer based on ViT, followed by a linear fully connected layer for outputting the prediction information pair of the video image and motion trajectory of other perspectives at the current time.
[0050] C. Weight training of the prediction module
[0051] The weight training process of the prediction module is as follows Figure 7As shown, the above-mentioned video image and motion trajectory information pairs are taken as input to construct the model training process. The video image and motion trajectory information pairs collected by the distributed image and motion information recording device of the subject's head and the language instructions are taken as input of the prediction module to perform image prediction to obtain the video image and motion trajectory prediction information pairs of other positions and compare the similarity loss between the video image and motion trajectory information pairs of the corresponding positions of the subject, and optimize the weight of the prediction module by minimizing the similarity loss. The initial weight in the prediction module is set to the public pre-training weight.
[0052] In use, a sufficient amount of data of normal subjects is collected to train the prediction module, so that the prediction module can achieve the function of generating video image and motion trajectory information of other perspectives from single or multiple perspective video image and motion trajectory information. The perspective video image and motion trajectory information output by the prediction module and the perspective video image and motion trajectory information input together reflect the cognitive state of the normal subjects during the execution of the task, i.e. the characteristic difference, between each other.
[0053] The process of the cognitive-behavioral characteristic recording method based on the distributed visual device provided by the present application is as shown in Figure 8 As shown, the present application distinguishes the difference between the cognitive-behavioral characteristics of the subject and the normal population by comparing the video image and motion trajectory information collected by the distributed image and motion information recording device worn by the subject when performing different tasks with the differences in different perspective visual images trained based on the healthy population. The difference value is recorded as the cognitive-behavioral characteristics of the subject.
[0054] The following provides a specific embodiment to compare the cognitive-behavioral characteristics of the subject, as shown in Figure 9As shown, the subject completes the target task under voice instructions, and the video images and motion trajectory information collected by the distributed image and motion information recording device worn on the head of the subject, the language instructions as input, through the prediction module to predict the corresponding images of other distributed image and motion information recording devices(including chest video images and motion trajectory prediction information pairs); At the same time, the distributed image and motion information recording device worn on the chest of the subject itself will record, so that two time sequence information pairs can be obtained. The two time sequence information pairs are sent into the visual simultaneous localization and mapping algorithm(SLAM) for analysis and recording of the motion trajectory characteristics of the current position; Further, the motion trajectory characteristics can be analyzed, such as calculating the spatial fluctuation amount of the trajectory, the pause practice of the trajectory, the variance of the attitude angle change and the like; The motion trajectory characteristics are recorded and compared with each other; Thus, the motion differences of different parts of the body of the subject and the healthy people when performing the same task can be obtained; Further, the cognitive condition of the target object can be reflected by analyzing and comparing the motion characteristics of the subject.
[0055] Experimental verification of cognitive behavior characteristics evaluation based on kitchen shopping task
[0056] a. Test environment configuration
[0057] In a 6m×4m simulated kitchen environment, we arranged three functional areas: an entrance area(1.5m×1m), a shelf area(2m×1m), and a workbench area(1.5m×0.7m). The shelf area adopts a three-layer structure design, with tomatoes neatly arranged on the top layer, egg boxes placed on the middle layer, and seasonings configured on the bottom layer as interference items. A stainless steel bowl with a diameter of 20 centimeters is fixed on the left side of the workbench, and a knife rack is set on the right side. The ground is marked with yellow tape to include a standard path containing two right-angle turns, extending from the entrance to the shelf area, then to the workbench, and finally returning to the starting point. The environmental lighting is constant through LED panel light sources.
[0058] b. Deployment of distributed image and motion information recording devices
[0059] 7 distributed visual-motor recording devices are deployed on the key positions of the subject's body through special fixing devices:
[0060] Head device: elastic headband fixed 120-degree wide-angle lens, matched with dual microphone array, recording first-view video(1280×720@30fps) and voice instruction(48kHz / 16bit sampling);
[0061] Left and right wrist devices: magic tape wristband fixed 160-degree fisheye lens, integrated with 6-axis IMU sensor(±16g acceleration / ±2000dps angular velocity);
[0062] Left and right instep devices: elastic binding fixed optical image stabilization lens, matched with 3-axis accelerometer;
[0063] Abdominal device: waist belt clip to fix 120-degree wide-angle lens for long-distance floor monitoring;
[0064] Chest device: magnetic chest pin to fix 120-degree wide-angle lens, covering the hand operation area;
[0065] All devices use STM32F407 microcontrollers, equipped with dual-channel UHS-I SD card storage (90MB / s write speed) and 3000mAh fast-charging battery; the time synchronization system combines clock chip timestamp synchronization to ensure that the cross-device time error is less than 100 milliseconds.
[0066] c. Task execution flow
[0067] Before the test, a three-level calibration procedure was performed: the axis IMU was calibrated for 30 seconds of static + 30 seconds of dynamic; the camera completed automatic white balance locking based on 24 color cards; the acoustic system ensured that the microphone delay difference was <0.1 milliseconds through a 1kHz sine wave test. The language instruction for the task was "please walk to the shelf - pick up the tomatoes and eggs - put them in the left bowl - return to the starting point", and the language instruction was played at 0.8 times the normal speed. The average task time of the healthy control group (50 people) was 28.5 seconds, the path length was 5.2 meters, and the hand movement coherence index was 0.92. The suspected cognitive impairment group showed typical abnormalities: >8 seconds of hesitation in front of the shelf (difficulty in locating objects), >30% error rate of grabbing (visual-motor coordination disorder), and >40 cm deviation in the path (spatial navigation defects).
[0068] d. Prediction module construction
[0069] The training data set came from 10 task records of 50 healthy subjects (total duration 7.5 hours, 8.1 million frames of images). Data augmentation included random ±20% brightness adjustment, motion blur simulation, and ±15-degree view rotation. The core model used a multi-modal Transformer architecture: the language instruction was converted to a 512-dimensional vector by a CLIP text encoder; the head image was encoded by a ViT-Base encoder to generate a 768-dimensional feature; after fusion with a 7-dimensional position encoding, it was input into a 24-layer GPT-2 Transformer (384 hidden units / 12 attention heads / GELU activation function); the output was upsampled by a four-layer deconvolution decoder based on ViT, generating a 224x224 resolution prediction image. The RAdam optimizer (learning rate 5e-5) was used for training, and the loss function was 0.7xSSIM+0.3xLPIPS, which was trained on 4xV100 GPUs with a batch size of 64 to SSIM>0.85 on the validation set for 5 cycles.
[0070] e. Quantification of cognitive behavioral differences
[0071] The analysis process of subject P001 (suspected mild cognitive impairment) consists of three key stages:
[0072] 1. Spatio-temporal alignment: mutual correlation calculation between chest IMU data and predicted sequence of inertial features to determine frame-level synchronization offset
[0073] 2. Perception difference calculation: actual capture and predicted chest perspective images are converted to Lab color space, and a weighted difference algorithm (L channel weight 0.7, a / b channels each 0.15) is used to generate a difference heat map in the range of 0-255
[0074] 3. Behavioral trajectory analysis: ORB-SLAM3 algorithm processes actual chest video (2000 feature points / frame, minimum matching 20 points), extracts three core indicators: path fluctuation (path first derivative standard deviation), decision pause times (speed <0.1 m / s for >1 s), hand movement angle variance (based on wrist IMU data)
[0075] The analysis results show that P001's visual difference index reaches 52.7 (healthy mean 15.2), path fluctuation 0.39 meters (healthy mean 0.11 meters), decision pause 5 times (healthy mean 1.2 times), hand movement entropy 5.1 (healthy mean 2.8), and object recognition delay 4.2 seconds (healthy mean 1.5 seconds). Multimodal verification shows that wrist IMU angular velocity abnormalities are highly correlated with image difference peaks (r=0.93, p<0.001), head gaze analysis shows that abnormal area dwell time increases by 300%, and speech response delay increases from 1.2 seconds to 3.5 seconds.
[0076] f. System implementation and security
[0077] The hardware uses a hierarchical architecture: edge nodes use NVIDIA Jetson AGX Orin (64GB memory) for real-time preprocessing, and central servers are configured with dual Xeon 6338 processors / 512GB memory / 4x RTX6000 Ada graphics cards for model inference. Software implements device communication based on the ROS2 framework, PyTorch 2.0+OpenCV 4.8 constructs the processing pipeline, Three.js implements 3D behavioral trajectory reproduction, and Django+Plotly Dash generates interactive clinical reports. The privacy protection system includes four mechanisms: real-time YOLOv8 face detection and Gaussian blur (σ=5.0), AES-256 end-to-end encryption, 24-hour automatic deletion strategy for raw video, and physical emergency power button.
[0078] The test can complete the evaluation of the traditional scale in 5 minutes, and can detect 0.5 degree level hand angle deviation and other subtle abnormalities, so as to provide an objective and sensitive quantitative tool for early cognitive impairment screening.
[0079] The application provides a distributed human task state motion process visual image recording device formed by a plurality of distributed image and motion information recording devices and a feature processing computer with a built-in prediction module and a comparison module for image features and trajectory features, which can automatically compare feature differences between cognitive-behavioral features of a subject and normal people, and provides a new research tool for further understanding of the ontological behavior differences caused by cognitive impairment.
[0080] In the embodiments of the present application, the term "a plurality of" refers to two or more, unless otherwise explicitly limited. The terms "mounting", "connecting", "fixing" and the like should be understood in a broad sense, for example, "connecting" can be fixed connection, detachable connection or integral connection. For those skilled in the art, the specific meaning of the above terms in the embodiments of the present application can be understood according to the specific circumstances.
[0081] In the description of the embodiments of the present application, it should be understood that the positions or position relationships indicated by the terms "upper", "lower" and the like are based on the positions or position relationships shown in the drawings, and are only for the convenience of describing the embodiments of the present application and simplifying the description, and do not indicate or imply that the devices or units referred to must have a particular direction, be constructed and operated in a particular position, therefore, it cannot be understood as a limitation on the embodiments of the present application.
[0082] In the description of the present application, the description of the terms "one embodiment", "one preferred embodiment" and the like means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0083] The above embodiments only illustrate the technical idea of the present application, and cannot limit the protection scope of the present application, any modification made according to the technical idea of the present application on the basis of the technical scheme falls within the protection scope of the present application; the technologies not involved in the present application can be realized by the prior art.
Claims
1. A method for recording cognitive-behavioral features based on distributed vision devices, characterized in that: The steps of this recording method are as follows: S1. As required, wear distributed image and motion information recording devices on multiple parts of the subject's body; S2. In a fixed environment, give the subject verbal instructions and ask the subject to complete the task according to the verbal instructions; S3. Distributed image and motion information recording devices collect video images and motion trajectory information of the subject as they complete tasks according to verbal instructions, forming multiple pairs of video images and motion trajectory information based on verbal instructions and corresponding to the subject's body parts; S4. Select one video image and motion trajectory information pair corresponding to a part from the multiple sets of video image and motion trajectory information pairs in step S3 and input it into the prediction module to obtain video image and motion trajectory prediction information pairs corresponding to the remaining parts based on language instructions. S5. Input the video images and motion trajectory information corresponding to the remaining parts in step S4 and the video images and motion trajectory prediction information corresponding to the remaining parts obtained in step S4 into the image feature and trajectory feature comparison module for comparison one by one to obtain the feature differences of the subjects. S6. The above-mentioned differences in characteristics were recorded as cognitive-behavioral characteristics of the subjects.
2. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 1, characterized in that: The distributed image and motion information recording device in steps S1 and S3 includes a camera, an inertial sensor, a microphone, a microcontroller processor, a USB interface, a memory card, an indicator light, a Bluetooth module, a clock, and a battery. The camera is a miniature camera capable of capturing video images. The inertial sensor is used to sense and record the acceleration and angular velocity of the distributed image and motion information recording device. The microphone is used to receive sound information from the subject and the surrounding environment. The microcontroller processor is used to read the data from the camera, microphone, inertial sensor, and the current clock and save it to the memory card. The data saved to the memory card is output through the USB interface or Bluetooth module. The Bluetooth module is used to transmit data or receive control commands from external devices. The clock is used to match the current clock data of the camera, inertial sensor, and microphone. The battery is used to power the distributed image and motion information recording device.
3. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 1 or 2, characterized in that: The installation locations and requirements for the distributed image and motion information recording devices in step S1 are as follows: one distributed image and motion information recording device is installed on the human head, with the camera angle parallel to the human's angle of view, and can move freely with the movement of the human head; two distributed image and motion information recording devices are respectively installed on the left and right wrists of the human body, with the camera angle perpendicular to the palm and facing outward, ensuring that the camera can see the target object when the human hand is performing the task; two distributed image and motion information recording devices are respectively installed on the subject's shoes, with the angle of view pointing directly in front of the human body, ensuring that the subject can see obstacles on the ground in front of him / her while walking; one distributed image and motion information recording device is installed in front of the subject's abdomen to capture an image of the ground in front of the subject; and one distributed image and motion information recording device is installed on the subject's chest to capture an image of the subject's field of vision when his / her hands are operating.
4. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 3, characterized in that: The distributed image and motion information recording device in step S1 is installed on a designated part of the subject via Velcro, straps, or clothing.
5. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 2, characterized in that: The distributed image and motion information recording device in steps S1 and S3 uses a video sampling rate of 30Fps, an inertial sensor sampling rate of 100Hz, and an audio data sampling rate of 48KHz. The data collected by the distributed image and motion information recording device includes the current time stamp information and the human body position information of the distributed image and motion information recording device.
6. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 1, characterized in that: The prediction module in step S4 and the image feature and trajectory feature comparison module in step S5 are both located in the feature processing computer. The prediction module is used to receive the collected information at a certain part and give the prediction results for the other parts. The image feature and trajectory feature comparison module is used to receive the collected information and prediction results for the other parts, and output the feature differences of the subject after comparing them one by one.
7. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 1, characterized in that: The prediction module in step S4 uses training data derived from data obtained by healthy individuals performing steps S2 and S3.
8. The cognitive-behavioral feature recording method based on distributed vision devices according to any one of claims 1, 6-7, characterized in that: The prediction module in step S4 includes a language encoder module and an image encoder module as input modules, a Transformer neural network module as a data processing module, an image decoder module that decodes the hidden feature vectors output by the Transformer neural network module, and a fully connected layer that outputs prediction information from other perspectives at the current moment. The language encoder module is used to receive language commands, the image encoder module is used to receive information collected by the distributed image and motion information recording device and encode it into tokens data that the Transformer neural network module can accept, the Transformer neural network module processes the received data and outputs hidden feature vectors, and the image decoder module decodes the hidden feature vectors and outputs prediction information from other perspectives at the current moment through the fully connected layer.
9. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 8, characterized in that: The language encoder module uses a CLIP ViT-B / 32 text encoder to encode language instructions; the image encoder module uses a pre-trained ViT-Base encoder for encoding; the TransForms neural network module selects multi-layer GPT-2 transformer blocks, and selects 24-layer transformer blocks based on reference data, with each block having 384 hidden units and 12 attention heads; the image decoder module selects a ViT-based transformer.
10. The cognitive-behavioral feature recording method based on distributed vision devices according to claim 1, characterized in that: When the prediction module is trained in step S4, the weight training steps are as follows: Q1. Recruit healthy individuals as subjects, and have each subject wear distributed image and motion information recording devices on multiple parts of their body as required, and set the pre-training weights for the prediction module. Q2. In a fixed environment, all subjects are given the same verbal instructions and are asked to complete the task according to the verbal instructions. Q3. Distributed image and motion information recording devices collect video images and motion trajectory information of the subject as they complete tasks according to verbal instructions, forming multiple pairs of video images and motion trajectory information based on verbal instructions and corresponding to the subject and the subject's body parts; Q4. Select one video image and motion trajectory information pair corresponding to a part of the subject from the multiple sets of video image and motion trajectory information pairs in step Q3 and input it into the prediction module to obtain video image and motion trajectory prediction information pairs corresponding to other subjects and the remaining parts of the selected subject based on language instructions; Q5. Input the video images and motion trajectory information corresponding to the remaining parts of other subjects and selected subjects in step Q4 and the video images and motion trajectory prediction information corresponding to the remaining parts of other subjects and selected subjects obtained in step Q4 into the image feature and trajectory feature comparison module for comparison one by one, and obtain the similarity loss of the video images and motion trajectory prediction information corresponding to the remaining parts of other subjects and selected subjects obtained in step Q4 relative to the video images and motion trajectory information corresponding to the remaining parts of other subjects and selected subjects in step Q4; Q6. Adjust and optimize the pre-trained weights of the prediction module according to the minimum similarity loss, return to step Q4 and repeat steps Q4-Q6 until the similarity loss decrease rate approaches 0, and complete the weight training step.