Middle and long distance non-contact physiological signal recognition sample library construction method and system based on video recognition

By constructing a medium- and long-distance contactless physiological signal recognition sample library based on video recognition, the problems of sample missing and insufficient data diversity in the prior art are solved, and the effectiveness and data diversity of medium- and long-distance heart rate monitoring are achieved.

CN120148067APending Publication Date: 2025-06-13CAS SMARTCITY (GUANGZHOU) INFO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510111384.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, the construction of medium- and long-distance contactless heart rate identification sample library has problems such as sample deletion and insufficient data diversity, and cannot support the generalization ability of medium- and long-distance heart rate monitoring and real-world application scenarios.

Method used

A medium and long-distance contactless physiological signal recognition sample library is constructed using a method based on video recognition. A sample library containing medium and long-distance heart rate data is constructed through steps such as physiological signal calibration, video stream data preprocessing, time stamp alignment, video stream post-processing and sample storage.

Benefits of technology

The effectiveness of medium and long-distance heart rate monitoring has been achieved, the sample collection scenario has been broadened, the data diversity and accuracy have been improved, and the model's adaptability in practical application scenarios has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148067A_ABST
    Figure CN120148067A_ABST
Patent Text Reader

Abstract

The invention provides a method for constructing a medium-long distance non-contact physiological signal recognition sample library based on video recognition, which comprises the following steps: calibrating a physiological signal, and enabling a target person to wear physiological signal detection equipment to move according to a preset process under outdoor video monitoring equipment, the physiological signal detection equipment detects physiological signals of a human body in real time and transmits the physiological signal data to the background; obtaining video stream data, separating the video stream data and separating a video clip with a target person; data timestamp alignment: performing timestamp alignment on the acquired physiological signal data and video stream data; performing video stream post-processing, and performing face key point privacy coding on the recognized video stream of the target person; and sample storage: performing task identification numbering on the target person, and performing grouped storage on the physiological signal data of the target person and the associated video stream according to the person identification number to form a sample library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of non-contact physiological signal detection, and particularly relates to a method and system for constructing a medium and long-distance non-contact physiological signal recognition sample library based on video recognition. Background Art

[0002] With the rapid development of information technology and biomedical engineering, non-contact physiological signal monitoring technology has gradually become one of the research hotspots in the field of health monitoring. Most traditional heart rate measurement methods rely on physical contact devices, such as electrocardiogram (ECG), photoplethysmography (PPG), etc. Although these methods can provide accurate heart rate data, they have problems such as the need for the target to wear and discomfort during wearing, which limit their wide application in daily life.

[0003] In recent years, significant progress has been made in the research on non-contact heart rate detection using computer vision technology, and remote photoplethysmography (rPPG) has been proposed. This technology can estimate an individual's heart rate from a video captured by an ordinary camera by analyzing the minute color changes in the face or other visible skin areas. This method has the advantages of not requiring wearable devices, being simple to operate, and being easy to integrate into existing systems, providing a new solution for realizing continuous and remote heart rate monitoring.

[0004] However, in the actual application process, in order to train and verify these advanced rPPG algorithms, a high-quality and diverse sample library is particularly important. However, in constructing a standard sample library suitable for rPPG research, the current technology faces the following challenges:

[0005] Lack of medium and long-distance samples: The currently publicly available non-contact heart rate recognition datasets mainly focus on data obtained in close-range scenarios, that is, the distance between the camera and the object is short. However, in actual applications, such as in public place monitoring or remote medical consultation scenarios, it is necessary to accurately capture heart rate information at medium and long distances (within the range of 2 meters to 10 meters). In this case, the existing datasets are obviously insufficient to support relevant research and technology development because they do not cover a sufficient number and quality of medium and long-distance samples.

[0006] Insufficient data diversity: Most of the existing publicly available datasets are usually recorded in a controlled environment, such as in a laboratory environment where the participants remain stationary or move slightly. This results in insufficient data diversity in the sample library and cannot fully reflect the variable environmental conditions in the real world (such as different light intensities, background interferences, target postures, etc.), thus limiting the generalization ability of the model in actual application scenarios.

[0007] In summary, although certain achievements have been made in non-contact heart rate recognition technology, due to the above reasons, a complete and general mid- and long-distance non-contact heart rate recognition sample database has not been established in the industry so far. Summary of the Invention

[0008] In view of the deficiencies of the prior art, the present invention provides a method and system for constructing a mid- and long-distance non-contact physiological signal recognition sample database based on video recognition.

[0009] The present invention provides a method for constructing a mid- and long-distance non-contact physiological signal recognition sample database based on video recognition, including the following steps:

[0010] Step S1: Physiological signal calibration. The target person wears a physiological signal detection device and conducts activities under an outdoor video surveillance device according to a preset process. The physiological signal detection device real-time detects the physiological signals of the human body and transmits the physiological signal data to the background.

[0011] Step S2: Obtain video stream data. Separate the video stream data and separate the video segments with the target person. According to the preset human key points, locate the human key points of the target person to analyze the posture and trajectory of the target person, and transmit the preprocessed video stream data to the background.

[0012] Step S3: Data timestamp alignment. Align the timestamps of the obtained physiological signal data and video stream data so that the physiological signal data corresponds to the video segment of the target person on the time axis.

[0013] Step S4: Video stream post-processing. Perform privacy encoding on the face key points of the video stream of the recognized target person, and perform image removal or blurring processing on the positions of the face key points.

[0014] Step S5: Sample storage. Assign a task identification number to the target person, and group and store the physiological signal data of the target person and the associated video stream according to the person identification number to form a sample database.

[0015] Preferably, in step S1, the physiological signal detection device is connected to and communicates data with a Bluetooth device of the HID class. When communicating data, the physiological signal data is encoded and encrypted. The physiological signal data starts with a preset specific three-byte sequence and ends with the start flag of the next segment of physiological signal data.

[0016] Preferably, in step S2, the YOLO 11 model is used to detect the target person in the video stream and separate the video segment of the target person. The YOLO 11 model includes a backbone network Backbone, a neck network Neck, and a head network Head. Detecting the target person in the video stream specifically includes the following steps:

[0017] Step S21: In the backbone network Backbone, the acquired video stream is intercepted into image frames with human images, and the image frames are processed through the convolutional layer and the recognition layer multiple times to extract the image features of the image frames. The recognition processing results are output at different recognition layers to form preprocessed feature information and feature maps;

[0018] Step S22: In the neck network Neck, the preprocessed feature information is sampled and the resolution of the feature map is enhanced, and then it is processed through the convolutional layer and the recognition layer again and the features are fused, and the detection maps obtained at different processing stages are output;

[0019] Step S23: In the head network Head, the detection maps output at different stages are respectively detected to obtain images with target persons, and human key point recognition and marking are performed.

[0020] Preferably, for the images with target persons obtained from step S23, the target persons are tracked based on the BoT-SORT model to form a target person trajectory sequence, which specifically includes the following steps:

[0021] Step S24: After the first frame image of the target person is detected, an initial trajectory is created for each target person, and the newly detected target persons in the subsequent frame images are associated with the existing trajectories using an association algorithm;

[0022] Step S25: State estimation initialization, for the initial trajectory of each target person, the initial state of the target person is estimated, including the position, acceleration and speed of the target person;

[0023] Step S26: Motion prediction and appearance feature update, based on the BoT-SORT model, the position of the target person in the next frame is predicted according to the current state of the target person, and the appearance features of the newly detected target person after moving are fused with the existing appearance features in the trajectory;

[0024] Step S27: Data association optimization, combining the predicted position of the target person and the updated appearance features, the detection results in the current frame are re-associated with the existing trajectories;

[0025] Step S28: Trajectory termination and output: It is judged whether the target person meets the trajectory termination according to the preset trajectory termination conditions, and the generated trajectories are screened and output.

[0026] Preferably, after obtaining the human key points in step S23, pose estimation is performed based on a rule-based method, and the specific steps include:

[0027] Step S231: Set the human key points as the head H, the hip P, the foot or knee F, and the shoulder S; calculate the height differences of the human key points:

[0028] ΔH = H.y - P.y

[0029] ΔP = P.y - F.y

[0030] where H.y is the value of the head H on the Y-axis coordinate, P.y is the value of the hip P on the Y-axis coordinate, and F.y is the value of the foot or knee F on the Y-axis coordinate;

[0031] Step S232: Set the key points of the shoulder S and the hip P to calculate the angle θ between the target person's body and the ground upper , θ upper satisfies the following formula:

[0032] cos(θ upper ) = (S.x - P.x) / sqrt((S.x - P.x)^2 + (S.y - P.y)^2)

[0033] where S.x is the value of the shoulder S on the X-axis coordinate, P.x is the value of the hip P on the X-axis coordinate, and S.y is the value of the shoulder S on the Y-axis coordinate;

[0034] Step S233: Set the key points of the foot / knee F and the hip P to calculate the angle θ between the target person's lower body and the ground lower , θ lower satisfies the following formula:

[0035] cos(θ lower ) = (F.x - P.x) / sqrt((F.x - P.x)^2 + (F.y - P.y)^2)

[0036] where F.x is the value of the foot or knee F on the X-axis coordinate.

[0037] Preferably, the method for pose estimation of the target person includes:

[0038] Standing pose: ΔH is greater than the preset threshold T_H, ΔP is greater than the preset threshold T_P, and the ratio of ΔH to ΔP is 0.8 - 1.2; θ upper is 70° - 110°, θ lower is 70° - 110°;

[0039] Sitting pose: ΔH is greater than the preset threshold T_H, ΔP is less than the preset threshold T_P, and the ratio of ΔH to ΔP is greater than 1.2; θ upper is 70° - 110°, θ lower is less than 70°;

[0040] Lying position: ΔH is less than the preset threshold T_H, and ΔP is less than the preset threshold T_P; θ upper is less than 70° or greater than 110°, and θ lower is less than 70° or greater than 110°.

[0041] Preferably, when post-processing the video stream in step S4, after identifying the target person, the individual area of the target person is segmented and the background is removed, the key face area of the target person is identified, and then the face key points are privately encoded.

[0042] Preferably, in step S3 for data timestamp alignment, based on the WebRTC protocol, the images recognized by the video surveillance device are synchronized to the background, a unique number is assigned to each identified target person, and a corresponding synchronous acquisition button is set for each target person in the background. After confirmation of acquisition, video capture of the area where the target person with the specified number is located is started, and a timestamp is embedded in each frame of the image. At the same time, physiological signal data is collected, and after data processing, a timestamp is attached and transmitted to the background to align the timestamps of the video stream and the physiological signals.

[0043] The present invention also provides a system for constructing a medium- and long-distance contactless physiological signal recognition sample library based on video recognition, which is used to construct a sample library according to the method for constructing a medium- and long-distance contactless physiological signal recognition sample library based on video recognition described in any one of the above, including:

[0044] A video acquisition module, which is used to acquire video stream data of medium- and long-distance people. The video acquisition range is 2-10 meters, and the acquired video resolution is greater than or equal to 720P;

[0045] A physiological signal calibration module, which includes a physiological signal detection device and a Web port, and is used for the acquisition and data upload of biological signals;

[0046] An algorithm analysis module, which is deployed on the computing power service module, and is used to intercept physiological signal segments containing multiple target objects from the video stream data for data analysis, align the timestamps with the data obtained from the physiological signal calibration module, and construct a physiological signal video sample database;

[0047] A computing power service module, which is used to deploy the algorithm analysis module and store the acquired data, and forms a B / S architecture with the Web port of the physiological signal calibration module.

[0048] Preferably, the physiological signal detection device is a pulse oximeter.

[0049] The method and system for constructing a medium- and long-distance contactless physiological signal recognition sample library based on video recognition provided by the present invention have at least the following beneficial effects:

[0050] 1. By synchronously performing video stream detection and physiological signal calibration and aligning timestamps, the accuracy of the sample library data is ensured.

[0051] 2. Expand the monitoring range: Break through the distance limitation of traditional contact or short - range non - contact heart rate detection, and achieve effective monitoring within the range of 2 meters to 10 meters, greatly broadening the sample collection scenarios.

[0052] 3. Improve data diversity and accuracy: Through medium - to - long - distance monitoring cameras in various environments and multi - target human key point detection technology, enhance the adaptability to different environmental conditions and improve the accuracy and stability of physiological signal measurement. Brief Description of the Drawings

[0053] The above - mentioned and other objects, features, and advantages of the present invention will become clearer through the preferred embodiments of the present invention shown in the drawings. The same reference numerals in all the drawings indicate the same parts, and the drawings are not deliberately drawn to scale in actual size. The focus is on showing the gist of the present invention.

[0054] Figure 1 It is a structural diagram of a system for constructing a medium - to - long - distance non - contact physiological signal recognition sample library based on video recognition provided by an embodiment of the present invention.

[0055] Figure 2 It is a data flow and processing framework diagram of a method for constructing a medium - to - long - distance non - contact physiological signal recognition sample library based on video recognition provided by an embodiment of the present invention. Detailed Embodiments

[0056] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings.

[0057] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element and integrated with it, or there may be an intermediate element at the same time. The terms "installed", "one end", "the other end" and similar expressions used in this article are only for the purpose of illustration.

[0058] Unless otherwise defined, all technical and scientific terms used in this article have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs. The terms used in the description of this specification in this article are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this article includes any and all combinations of one or more of the related listed items.

[0059] An embodiment of the present invention provides a method for constructing a medium - to - long - distance non - contact physiological signal recognition sample library based on video recognition, including the following steps:

[0060] Step S1: Physiological signal calibration, the target person wears the physiological signal detection device and performs activities according to the preset process under the outdoor video monitoring device. The physiological signal detection device detects the physiological signals of the human body in real time and transmits the physiological signal data to the background; the outdoor video monitoring device can use a high-definition network camera with a resolution of not less than 720P, which can not only ensure the clarity of the acquired video stream, but also ensure that the multi-target person video can be accurately separated and the key points of the human body can be detected when the video stream is visually recognized, so as to effectively perform posture analysis; if the resolution of the outdoor video monitoring device is reduced, it may lead to a decrease in the accuracy of target person feature recognition, thereby affecting the effectiveness of privacy protection measures related to video-based human posture estimation, and may shorten the effective distance for identifying physiological signals from the video stream. The target person should be selected from different feature groups, including people of different ages, genders and skin colors. The preset process should preset different scenarios, including the activities of a single target person, the activities of multiple target persons, the activities of multiple target tasks and multiple non-target persons, etc., and data should be collected under different lighting conditions, background environments, etc. to enhance data diversity.

[0061] Step S2: acquiring video stream data, separating the video stream data and separating the video segments with the target person, locating the target person's human body key points according to the preset human body key points to analyze the target person's posture and trajectory, and transmitting the pre-processed video stream data to the background;

[0062] Step S3: aligning data timestamps, aligning the acquired physiological signal data and video stream data with timestamps, so that the physiological signal data and the video clips of the target person correspond to each other on the time axis;

[0063] Step S4: post-processing the video stream, performing privacy encoding of facial key points on the video stream of the identified target person, and performing image removal or blurring processing on the positions of facial key points;

[0064] Step S5: Sample preservation: assign a task identification number to the target person, and group and store the target person's physiological signal data and the associated video stream according to the person identification number to form a sample library.

[0065] In this embodiment, in step S1, the physiological signal detection device is connected and data communicated with a Bluetooth device of the HID class. HID (Human Interface Device) refers to a human-computer interface device, a device interface standard for human-computer interaction, which allows a computer to communicate with various input and output devices through a universal protocol. When performing data communication, the physiological signal data is encoded and encrypted, and the physiological signal data uses a preset specific three-byte sequence as a start mark and uses the start mark of the next physiological signal data as an end mark.

[0066] In a specific embodiment, taking the CONTEC PF-10AW Bluetooth oximeter as an example, for the data packets received from the PF-10AW device, its structure follows the following format: Each data packet starts with a specific three-byte sequence as the start flag, for example, [170, 85, 15], and the three-byte start flag of the next data packet is used as the end flag, generally with a length of 11 bits, to identify that this segment of data stream is physiological index data. According to the content characteristics of these data packets, they can be further subdivided into PPG (photoplethysmography) data and physiological parameter data. Specifically, two bytes located after the data packet prefix are defined as the category flags, used to distinguish different types of data records. When the category flag is [7, 2], the following five bytes represent a series of continuous PPG detection values; when the category flag is [8, 1], the first byte following represents the blood oxygen saturation measurement value, and the second byte corresponds to the heart rate detection result. The specific meanings of the remaining bytes may vary depending on the device model, and they can be used to represent other types of detection results or additional information.

[0067] To better understand the above data packet format, two example data packets are given below:

[0068] [170, 85, 15, 7, 2, 73, 74, 70, 68, 67, 176]: Among them, [7, 2] is the category flag, indicating that the next five numbers (from 73 to 67) are continuous PPG detection values, and 176 is the data check bit.

[0069] [170, 85, 15, 8, 1, 97, 53, 0, 89, 0, 192, 193, 170, 85, 240, 3, 3, 3, 246]: Here, [8, 1] is used as the category flag, indicating that the two numbers following (97 and 53) represent the blood oxygen saturation and heart rate detection values respectively, 192 is the data check bit, and the subsequent [170, 85, 240, 3, 3, 3, 246] are invalid data.

[0070] By performing specific encoding and encryption processing on the physiological signal data, the security of the physiological signal data during the transmission process can be ensured, and the accuracy and integrity of the data are guaranteed during the construction and storage of the sample library.

[0071] In this embodiment, in step S2, based on the YOLO 11 model, the target person in the video stream is detected and the video segment of the target person is separated. The YOLO 11 model includes a backbone network Backbone, a neck network Neck, and a head network Head. Detecting the target person in the video stream specifically includes the following steps:

[0072] Step S21: In the backbone network Backbone, the acquired video stream is intercepted into image frames with human images. The image frames are processed through the convolutional layer and the recognition layer multiple times to extract the image features of the image frames. The recognition processing results are output at different recognition layers to form preprocessed feature information and feature maps. As Figure 2 shown, multiple Conv convolutional layers and specific modules such as C3K2, SPFF, and C2PSA are set in the Backbone, and the recognized feature information and feature maps are output in different C3K2 and C2PSA.

[0073] Step S22: In the neck network Neck, the preprocessed feature information is sampled and the resolution of the feature map is enhanced, and then it is processed through the convolutional layer and the recognition layer again and the features are fused, and the detection maps obtained at different processing stages are output;

[0074] Step S23: In the head network Head, the detection maps output at different stages are respectively detected to obtain images with target persons and the human key points are recognized and marked.

[0075] In a further embodiment, for the images with target persons obtained from step S23, the target persons are tracked based on the BoT-SORT model to form a target person trajectory sequence, which specifically includes the following steps:

[0076] Step S24: After the first frame image of the target person is detected, an initial trajectory is created for each target person, and the newly detected target persons are associated with the existing trajectories using an association algorithm for the detection results in subsequent frame images;

[0077] Step S25: State estimation initialization, for the initial trajectory of each target person, the initial state of the target person is estimated, including the position, acceleration, and speed of the target person;

[0078] Step S26: Motion prediction and appearance feature update, based on the BoT-SORT model, the position of the target person in the next frame is predicted according to the current state of the target person, and the appearance features of the newly detected target person after moving are fused with the existing appearance features in the trajectory;

[0079] Step S27: Data association optimization, combining the predicted position of the target person and the updated appearance features, the detection results in the current frame are re-associated with the existing trajectories;

[0080] Step S28: Trajectory termination and output: It is judged whether the target person meets the trajectory termination according to the preset trajectory termination conditions, and the generated trajectories are screened and output.

[0081] In this embodiment, the recognition model based on YOLO 11 and the trajectory tracking model based on BoT-SORT can be implemented or reproduced within the Ultralytics framework. By using the Ultralytics framework, the model parameters can be quickly trained, deployed, and adjusted to meet the requirements of different application scenarios and ensure optimal performance. In some cases, such as long distances, multiple angles, and multiple interferences (masks, hats), the recognition accuracy may not meet the requirements. Then, it is necessary to further collect face and human key point datasets for re-strengthening training of the model.

[0082] In this embodiment, after obtaining the human key points in step S23, pose estimation is performed based on a rule-based method. The specific steps include:

[0083] Step S231: Set the human key points as the head H, hip P, foot or knee F, and shoulder S; calculate the height difference between the human key points. The head position is counted based on the highest facial feature position. If none are detected, the shoulder S position is counted. If the foot is not detected, the knee position is counted, and at this time, the height difference is doubled:

[0084] ΔH = H.y - P.y

[0085] ΔP = P.y - F.y

[0086] where H.y is the value of the head H on the Y-axis coordinate, P.y is the value of the hip P on the Y-axis coordinate, and F.y is the value of the foot or knee F on the Y-axis coordinate;

[0087] Step S232: Set the key points of the shoulder S and the hip P to calculate the angle θ between the target person's body and the ground upper , θ upper satisfies the following formula:

[0088] cos(θ upper ) = (S.x - P.x) / sqrt((S.x - P.x)^2 + (S.y - P.y)^2)

[0089] Then θ can be calculated upper = arccos(cos(θ upper ));

[0090] where S.x is the value of the shoulder S on the X-axis coordinate, P.x is the value of the hip P on the X-axis coordinate, and S.y is the value of the shoulder S on the Y-axis coordinate;

[0091] Step S233: Set the key points of the foot / knee F and the hip P to calculate the angle θ between the lower body of the target person and the ground lower , θ lower satisfies the following formula:

[0092] cos(θ lower ) = (F.x - P.x) / sqrt((F.x - P.x)^2 + (F.y - P.y)^2)

[0093] Then, θ can be calculated lower = arccos(cos(θ lower ));

[0094] Among them, F.x is the value of the foot or knee F on the X-axis coordinate.

[0095] Furthermore, the method for estimating the pose of the target person includes:

[0096] Standing pose: ΔH is greater than the preset threshold T_H, ΔP is greater than the preset threshold T_P, and the ratio of ΔH to ΔP is 0.8 - 1.2; θ upper is 70° - 110°, θ lower is 70° - 110°;

[0097] Sitting pose: ΔH is greater than the preset threshold T_H, ΔP is less than the preset threshold T_P, and the ratio of ΔH to ΔP is greater than 1.2; θ upper is 70° - 110°, θ lower is less than 70°;

[0098] Lying pose: ΔH is less than the preset threshold T_H, ΔP is less than the preset threshold T_P; θ upper is less than 70° or greater than 110°, θ lower is less than 70° or greater than 110°.

[0099] Still further, the method for estimating the pose of the target person further includes:

[0100] The target person faces the video surveillance device directly: the eye key point E or the shoulder key point S is in a relatively horizontal position;

[0101] The target person looks sideways at the video surveillance device: one side of the eye key point E or the shoulder key point S is missing or highly overlapped;

[0102] The target person looks up at the video surveillance device: the position of the facial feature key point and the shoulder key point S is less than the preset threshold.

[0103] In this embodiment, during the post-processing of the video stream in step S4, after identifying the target person, the individual area of the target person is segmented and the background is removed, the key facial area of the target person is identified, and then the facial key points are privately encoded. In this embodiment, privacy protection is achieved by blurring or encrypting the eyes, nose, and ears of the face. For example, the positions of the eyes, nose, and ears of the face are detected among the obtained human key points, and then, according to the requirements of different scenarios, for one-time data, the image is removed or blurred, while for repetitive data, the asymmetric encryption algorithm RSA is used for encryption, and the decryption permission can be submitted to the relevant functional departments to effectively ensure that the collected data is not leaked.

[0104] In the preferred embodiment, in step S3 for data timestamp alignment, based on the WebRTC protocol, the images recognized by the video surveillance device are synchronized to the background, a unique number is assigned to each identified target person, and a corresponding synchronous acquisition button is set for each target person in the background. When the acquisition is confirmed, the video capture of the area where the target person with the specified number is located is started, and a timestamp is embedded in each frame of the image. At the same time, physiological signal data is collected, and after data processing, a timestamp is added and transmitted to the background to align the timestamps of the video stream and the physiological signals.

[0105] Taking the CONTEC PF-10AW Bluetooth blood oxygen meter as an example, first, the Bluetooth device pairing operation is performed; once the pairing is completed, the system immediately sends a video recording instruction to the background server. In response to this instruction, the background server immediately starts the video capture of the area where the target person with the specified number is located, and an accurate timestamp and the detected human key point information are embedded in each frame of the image. At the same time, the Web side starts to collect Bluetooth signals, and after sorting and classification processing, a timestamp is also added and transmitted to the background server for storage.

[0106] This embodiment can minimize the alignment error caused by the inherent transmission delay during the data acquisition process, ensuring a high degree of consistency and synchronization between the video stream and the Bluetooth sensor data, not only improving the accuracy of data synchronization but also providing a reliable basis for subsequent data analysis.

[0107] In step S5 for sample preservation, for different application requirements, the embodiments of the present invention provide multiple processing options:

[0108] When human segmentation needs to be implemented, the segmentation (SEG) module in the YOLO 11 architecture can be called to further refine the original data;

[0109] To protect the privacy of the person, the asymmetric RSA encryption algorithm is used to specifically encrypt and protect the eye area determined by the human key point information;

[0110] For the need of pose analysis, the system can perform additional data extraction and analysis based on the existing human key point information.

[0111] An embodiment of the present invention further provides a system for constructing a mid-long distance non-contact physiological signal recognition sample library based on video recognition, which is used to construct a sample library according to the method for constructing a mid-long distance non-contact physiological signal recognition sample library described in any one of the above, including:

[0112] A video acquisition module, which is used to acquire video stream data of mid-long distance people. The video acquisition range is 2-10 meters, and the acquired video resolution is greater than or equal to 720P;

[0113] A physiological signal calibration module, including a physiological signal detection device and a Web port, which is used for the acquisition and data upload of biological signals;

[0114] An algorithm analysis module, which is deployed on the computing power service module, is used to intercept physiological signal segments containing multiple target objects from the video stream data and perform data analysis, align the time stamps with the data obtained from the physiological signal calibration module, and construct a physiological signal video sample database;

[0115] A computing power service module, which is used to deploy the algorithm analysis module and store the acquired data, and form a B / S architecture with the Web port of the physiological signal calibration module.

[0116] In this embodiment, the physiological signal detection device is a pulse oximeter, which can be used to accurately measure the HR (Heart Rate) and SPO 2 (blood oxygen saturation) of the target person in real time. In rPPG, for the measurement of HR, by analyzing the periodic fluctuations of the rPPG signal and calculating the number of fluctuations per minute, it is the heart rate; for the measurement of SPO 2 The measurement is based on the difference in the absorption characteristics of oxyhemoglobin and deoxyhemoglobin in blood for light of different wavelengths (usually red light and infrared light), and the blood oxygen saturation is calculated through the rPPG signal. By using the technology of video visual recognition of video source data for mid-long distance non-contact physiological signal recognition, and at the same time using a physiological signal detection device for accurate physiological signal detection as calibration, a sample library can be constructed, which can provide a good basis for training and validating the rPPG algorithm.

[0117] In the description of this specification, the description with reference to terms such as "preferred embodiment", "further embodiment", "other embodiment" or "specific example" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0118] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition, characterized in that: The following steps are involved: Step S1: physiological signal calibration, the target person wears a physiological signal detection device and performs activities according to a preset process under an outdoor video surveillance device. The physiological signal detection device detects the physiological signals of the human body in real time and transmits the physiological signal data to the background; Step S2: acquiring video stream data, separating the video stream data and separating the video segments with the target person, locating the target person's human body key points according to the preset human body key points to analyze the target person's posture and trajectory, and transmitting the pre-processed video stream data to the background; Step S3: aligning data timestamps, aligning the acquired physiological signal data and video stream data with timestamps, so that the physiological signal data and the video clips of the target person correspond to each other on the time axis; Step S4: post-processing the video stream, performing privacy encoding of facial key points on the video stream of the identified target person, and performing image removal or blurring processing on the positions of facial key points; Step S5: Sample preservation: assign a task identification number to the target person, and group and store the target person's physiological signal data and the associated video stream according to the person identification number to form a sample library.

2. The method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition as claimed in claim 1, characterized in that: In step S1, the physiological signal detection device is connected to and communicates data with a HID-type Bluetooth device. During data communication, the physiological signal data is encoded and encrypted. The physiological signal data uses a preset specific three-byte sequence as a start mark and uses the start mark of the next physiological signal data as an end mark.

3. The method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition as claimed in claim 1, characterized in that: In step S2, the target person in the video stream is detected based on the YOLO 11 model and the video clip of the target person is separated. The YOLO 11 model includes a backbone network Backbone, a neck network Neck and a head network Head. Detecting the target person in the video stream specifically includes the following steps: Step S21: In the backbone network Backbone, the acquired video stream is intercepted into image frames with human body images, the image frames are processed by the convolution layer and the recognition layer for multiple times, the image features of the image frames are extracted, and the recognition processing results are output in different recognition layers to form pre-processed feature information and feature maps; Step S22: In the neck network Neck, the feature information obtained by preprocessing is sampled and the resolution of the feature map is improved, and the convolution layer and recognition layer are processed and the features are fused again, and the detection maps obtained at different processing stages are output; Step S23: In the head network Head, the detection images output at different stages are detected respectively to obtain images with the target person and identify and mark the key points of the human body.

4. The method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition as claimed in claim 3, characterized in that: For the image with the target person obtained in step S23, the trajectory of the target person is tracked based on the BoT-SORT model to form a trajectory sequence of the target person, which specifically includes the following steps: Step S24: after the first frame image of the target person is detected, an initial track is created for each target person, and an association algorithm is used to associate the newly detected target person with the existing track for the detection results in subsequent frame images; Step S25: initializing state estimation, estimating the initial state of each target person's initial trajectory, including the position, acceleration and velocity of the target person; Step S26: motion prediction and appearance feature update, predicting the position of the target person in the next frame according to the current state of the target person based on the BoT-SORT model, and fusing the appearance features of the newly detected target person after movement with the existing appearance features in the trajectory; Step S27: data association optimization, combining the predicted position of the target person and the updated appearance features, and re-associating the detection results in the current frame with the existing trajectory; Step S28: trajectory termination and output: determine whether the target person meets the trajectory termination conditions according to the preset trajectory termination conditions, and filter and output the generated trajectory.

5. The method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition as claimed in claim 3, characterized in that: After obtaining the key points of the human body in step S23, posture estimation is performed based on a rule-based method, and the specific steps include: Step S231: Set the key points of the human body as the head H, hips P, feet or knees F, and shoulders S; calculate the height difference of the key points of the human body: ΔH=Hy-Py ΔP=Py-Fy Wherein, Hy is the value of the head H on the Y-axis coordinate, Py is the value of the hip P on the Y-axis coordinate, and Fy is the value of the foot or knee F on the Y-axis coordinate; Step S232: Set the shoulder S and hip P key points to calculate the angle θ between the target person's body and the ground upper ,θ upper Satisfies the following formula: cos(θ upper )=(Sx-Px) / sqrt((Sx-Px)^2+(Sy-Py)^2) Among them, Sx is the value of the shoulder S on the X-axis coordinate, Px is the value of the hip P on the X-axis coordinate, and Sy is the value of the shoulder S on the Y-axis coordinate; Step S233: Set the foot / knee F and hip P key points to calculate the angle θ between the target person's lower body and the ground lower ,θ lower Satisfies the following formula: cos(θ lower )=(Fx-Px) / sqrt((Fx-Px)^2+(Fy-Py)^2) Wherein, Fx is the value of the foot or knee F on the X-axis coordinate.

6. The method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition as claimed in claim 5, characterized in that: Methods for estimating the target person's posture include: Standing posture: ΔH is greater than the preset threshold T_H, ΔP is greater than the preset threshold T_P, and the ratio of ΔH to ΔP is 0.8-1.2; θ upper is 70°-110°, θ lower 70°-110°; Sitting posture: ΔH is greater than the preset threshold T_H, ΔP is less than the preset threshold T_P, and the ratio of ΔH to ΔP is greater than 1.2; θ upper is 70°-110°, θ lower is less than 70°; Lying posture: ΔH is less than the preset threshold T_H, ΔP is less than the preset threshold T_P; θ upper is less than 70° or greater than 110°, θ lower is less than 70° or greater than 110°.

7. The method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition as claimed in claim 1, characterized in that: When the video stream is post-processed in step S4, after the target person is identified, the individual area of ​​the target person is segmented and the background is removed, the key facial area of ​​the target person is identified, and then the key facial points are privacy encoded.

8. The method for constructing a sample library for mid- and long-distance contactless physiological signal recognition based on video recognition as claimed in claim 1, characterized in that: In step S3 data timestamp alignment, the images recognized by the video surveillance device are synchronized to the background based on the WebRTC protocol, a unique number is assigned to each identified target person, and a corresponding synchronization acquisition button is set for each target person in the background. After the acquisition is confirmed, the video capture of the area where the target person with the specified number is located is started and a timestamp is embedded in each frame of the image. At the same time, physiological signal data is collected, and after data processing, the timestamp is attached and transmitted to the background to align the timestamps of the video stream and the physiological signal.

9. A system for constructing a sample library of mid- and long-distance contactless physiological signal recognition based on video recognition, used to construct a sample library according to the method for constructing a sample library of mid- and long-distance contactless physiological signal recognition based on video recognition according to any one of claims 1 to 8, characterized in that: include: Video acquisition module, used to obtain video stream data of people at medium and long distances, with a video acquisition range of 2-10 meters and a captured video resolution greater than or equal to 720P; Physiological signal calibration module, including physiological signal detection equipment and Web port, used for biological signal collection and data upload; The algorithm analysis module is deployed on the computing service module and is used to intercept physiological signal segments containing multiple target objects from the video stream data and perform data analysis, align the timestamps with the data obtained from the physiological signal calibration module, and build a physiological signal video sample database; The computing power service module is used to deploy the algorithm analysis module and store the collected data, and form a B / S architecture with the Web port of the physiological signal calibration module.

10. The medium- and long-distance contactless physiological signal recognition sample library construction system based on video recognition as claimed in claim 9, characterized in that: The physiological signal detection device is a pulse oximeter.