Heart rate detection method and device based on vehicle-mounted camera, medium and electronic equipment
By acquiring facial video of occupants through an in-vehicle camera, extracting key pixel information to generate frequency and time domain signals, eliminating noise and filtering, the accuracy problem of occupant heart rate detection in automobiles is solved, and real-time and accurate heart rate monitoring is achieved.
Patent Information
- Application Number
- CN202310661292.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-06
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-06-06
AI Technical Summary
Prolonged stays inside a car can cause changes in the physiological state of passengers, especially those with underlying medical conditions who are prone to dangerous situations. Current technology is not effective in detecting the heart rate of passengers.
The system acquires occupant facial video using an in-vehicle camera, extracts key pixel information from a preset facial segmentation region of the facial image, generates frequency domain and time domain signals, eliminates motion noise and performs filtering, and generates the occupant's heart rate signal.
It improves the accuracy of heart rate detection, reduces the impact of motion interference, and enables real-time monitoring of occupant heart rate changes.
Smart Images

Figure CN116687367B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a heart rate detection method, apparatus, medium, and electronic device based on an in-vehicle camera. Background Technology
[0002] As cars become a common mode of transportation, people are spending increasingly more time inside them. Prolonged exposure to a confined space can easily alter the physiological state of passengers. This is especially true for individuals with underlying health conditions, increasing the risk of dangerous situations.
[0003] Therefore, this application provides a heart rate detection method based on an in-vehicle camera to solve the above-mentioned technical problems. Summary of the Invention
[0004] The purpose of this application is to provide a heart rate detection method, device, medium, and electronic device based on an in-vehicle camera, which can solve at least one of the aforementioned technical problems. The specific solution is as follows:
[0005] According to a specific embodiment of this application, in a first aspect, this application provides a heart rate detection method based on an in-vehicle camera, comprising:
[0006] Obtain facial video of the occupants;
[0007] Based on the facial video, key pixel information in each preset facial segmentation region of the occupant's face in each frame of the facial image is obtained. The key pixel information includes the pixel RGB value of the key pixel and the key time point. The pixel RGB value of the key pixel is within a preset green color gamut.
[0008] Based on the pixel RGB values of all key pixels in each preset facial segmentation region and the key time points, the frequency domain signal and time domain signal of the corresponding preset facial segmentation region are generated respectively.
[0009] The occupant's heart rate signal is generated based on the frequency domain and time domain signals of all preset facial segmentation regions.
[0010] Optionally, generating the occupant's heart rate signal based on the frequency domain and time domain signals of all preset facial segmentation regions includes:
[0011] Motion noise in the frequency domain signal of each preset facial segmentation region is eliminated to obtain the corresponding frequency domain denoised signal for the preset facial segmentation region.
[0012] The temporal signal of each preset facial segmentation region is filtered to obtain the corresponding preset facial segmentation region's temporal filtered signal.
[0013] The occupant's heart rate signal is generated based on the frequency domain denoised signal of all preset facial segmentation regions and the time domain filtered signal of the corresponding preset facial segmentation regions.
[0014] Optionally, the step of eliminating motion noise in the frequency domain signal of each preset facial segmentation region to obtain the frequency domain denoised signal of the corresponding preset facial segmentation region includes the following formula:
[0015]
[0016] Among them, SNR i S represents the denoised signal of the i-th preset facial segmentation region, and f represents the frequency of the frequency domain signal of the corresponding preset facial segmentation region; i (f) represents the frequency domain signal of the i-th preset facial segmentation region; U i (f) represents the three frequency domain signals following the maximum frequency domain signal of the i-th preset facial segmentation region, and j represents the sequence number of the three frequency domain signals.
[0017] Optionally, the step of filtering the temporal signal of each preset facial segmentation region to obtain the corresponding preset facial segmentation region includes:
[0018] The first signal difference is obtained by subtracting the time domain signal of the second key time point from the time domain signal of any preset facial segmentation region at any first key time point, wherein the first key time point is a key time point that is after and immediately adjacent to the second key time point.
[0019] For any of the preset facial segmentation regions, when the absolute value of the first signal difference is less than or equal to a preset empirical difference threshold, the second signal difference is obtained by subtracting the first signal difference from the time domain signal of the third key time point, wherein the third key time point is a key time point that is after and immediately adjacent to the first key time point.
[0020] The time-domain signal at the third key time point is updated based on the second signal difference.
[0021] Optionally, generating the occupant's heart rate signal based on the frequency-domain denoised signals of all preset facial segmentation regions and the time-domain filtered signals of the corresponding preset facial segmentation regions includes:
[0022] When the frequency domain denoised signal of each preset facial segmentation region is less than or equal to the preset limit threshold, the first adjustment coefficient of the corresponding preset facial segmentation region is generated based on the frequency domain denoised signal of each preset facial segmentation region.
[0023] The occupant's heart rate signal is generated based on the first adjustment coefficient of all preset facial segmentation regions and the temporal filtering signal of the corresponding preset facial segmentation region.
[0024] Optionally, the method further includes:
[0025] When the frequency domain denoising signal of at least one preset facial segmentation region is less than or equal to a preset limit threshold, and the frequency domain denoising signal of the remaining preset facial segmentation regions is greater than the preset limit threshold, each preset facial segmentation region among the at least one preset facial segmentation region is determined to be a preset first effective segmentation region.
[0026] A second adjustment coefficient corresponding to the preset first effective segmentation region is generated based on the frequency domain denoising signal of each preset first effective segmentation region;
[0027] The occupant's heart rate signal is generated based on the second adjustment coefficients of all preset first effective segmentation regions and the time-domain filtered signal of the corresponding preset first effective segmentation region.
[0028] Optionally, after generating the occupant's heart rate signal based on the frequency domain and time domain signals of all preset facial segmentation regions, the method further includes:
[0029] The time-domain signal in the heart rate signal is subjected to bandpass filtering to obtain a shaped heart rate signal.
[0030] According to a specific embodiment of this application, in a second aspect, this application provides a heart rate detection device based on an in-vehicle camera, comprising:
[0031] The first acquisition unit is used to acquire facial videos of the occupants;
[0032] The second acquisition unit is used to acquire key pixel information in each preset facial segmentation region of the occupant's face in each frame of the facial image based on the facial video, wherein the key pixel information includes the pixel RGB value of the key pixel and the key time point, and the pixel RGB value of the key pixel is within a preset green color gamut.
[0033] The first generation unit is used to generate the frequency domain signal and time domain signal of the corresponding preset facial segmentation region based on the pixel RGB values of all key pixels in each preset facial segmentation region and the key time points.
[0034] The second generation unit is used to generate the occupant's heart rate signal based on the frequency domain signal and time domain signal of all preset facial segmentation regions.
[0035] Optionally, generating the occupant's heart rate signal based on the frequency domain and time domain signals of all preset facial segmentation regions includes:
[0036] Motion noise in the frequency domain signal of each preset facial segmentation region is eliminated to obtain the corresponding frequency domain denoised signal for the preset facial segmentation region.
[0037] The temporal signal of each preset facial segmentation region is filtered to obtain the corresponding preset facial segmentation region's temporal filtered signal.
[0038] The occupant's heart rate signal is generated based on the frequency domain denoised signal of all preset facial segmentation regions and the time domain filtered signal of the corresponding preset facial segmentation regions.
[0039] Optionally, the step of eliminating motion noise in the frequency domain signal of each preset facial segmentation region to obtain the frequency domain denoised signal of the corresponding preset facial segmentation region includes the following formula:
[0040]
[0041] Among them, SNR i S represents the denoised signal of the i-th preset facial segmentation region, and f represents the frequency of the frequency domain signal of the corresponding preset facial segmentation region; i (f) represents the frequency domain signal of the i-th preset facial segmentation region; U i (f) represents the three frequency domain signals following the maximum frequency domain signal of the i-th preset facial segmentation region, and j represents the sequence number of the three frequency domain signals.
[0042] Optionally, the step of filtering the temporal signal of each preset facial segmentation region to obtain the corresponding preset facial segmentation region includes:
[0043] The first signal difference is obtained by subtracting the time domain signal of the second key time point from the time domain signal of any preset facial segmentation region at any first key time point, wherein the first key time point is a key time point that is after and immediately adjacent to the second key time point.
[0044] For any of the preset facial segmentation regions, when the absolute value of the first signal difference is less than or equal to a preset empirical difference threshold, the second signal difference is obtained by subtracting the first signal difference from the time domain signal of the third key time point, wherein the third key time point is a key time point that is after and immediately adjacent to the first key time point.
[0045] The time-domain signal at the third key time point is updated based on the second signal difference.
[0046] Optionally, generating the occupant's heart rate signal based on the frequency-domain denoised signals of all preset facial segmentation regions and the time-domain filtered signals of the corresponding preset facial segmentation regions includes:
[0047] When the frequency domain denoised signal of each preset facial segmentation region is less than or equal to the preset limit threshold, the first adjustment coefficient of the corresponding preset facial segmentation region is generated based on the frequency domain denoised signal of each preset facial segmentation region.
[0048] The occupant's heart rate signal is generated based on the first adjustment coefficient of all preset facial segmentation regions and the temporal filtering signal of the corresponding preset facial segmentation region.
[0049] Optionally, the device further includes:
[0050] The determining unit is used to determine each of the at least one preset face segmentation regions as a preset first effective segmentation region when the frequency domain denoising signal of at least one preset face segmentation region is less than or equal to a preset limit threshold, and the frequency domain denoising signal of the remaining preset face segmentation regions is greater than the preset limit threshold.
[0051] The third generation unit is used to generate a second adjustment coefficient corresponding to the preset first effective segmentation region based on the frequency domain denoising signal of each preset first effective segmentation region.
[0052] The fourth generation unit is used to generate the occupant's heart rate signal based on the second adjustment coefficients of all preset first effective segmentation regions and the time-domain filtered signal of the corresponding preset first effective segmentation region.
[0053] Optionally, the device further includes:
[0054] The filtering unit is used to generate the occupant's heart rate signal based on the frequency domain signal and time domain signal of all preset facial segmentation regions, and then perform bandpass filtering on the time domain signal of the heart rate signal to obtain a shaped heart rate signal.
[0055] According to a specific embodiment of this application, in a third aspect, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the heart rate detection method based on an in-vehicle camera as described in any of the preceding claims.
[0056] According to a specific embodiment of this application, in a fourth aspect, this application provides an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the heart rate detection method based on a vehicle-mounted camera as described in any of the preceding claims.
[0057] Compared with the prior art, the above-described solutions of this application have at least the following beneficial effects:
[0058] This application provides a heart rate detection method, apparatus, medium, and electronic device based on an in-vehicle camera. This application converts the RGB values of key pixels within a preset green color gamut and key time points in each preset facial segmentation region of the occupant's face into frequency domain and time domain signals for each preset facial segmentation region. The occupant's heart rate signal is then generated using these frequency and time domain signals. Extracting key pixel information from preset facial segmentation regions reduces the impact of motion on the detection results. The RGB values of pixels within the preset green color gamut are most sensitive to changes in the occupant's heartbeat, thus improving the accuracy of the detection results. Attached Figure Description
[0059] Figure 1 A flowchart of a heart rate detection method based on an in-vehicle camera according to an embodiment of this application is shown;
[0060] Figure 2 A unit block diagram of a heart rate detection device based on an in-vehicle camera according to an embodiment of this application is shown. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0062] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “said,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.
[0063] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0064] It should be understood that although the terms first, second, third, etc., may be used in the embodiments of this application, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, first may also be referred to as second without departing from the scope of the embodiments of this application, and similarly, second may also be referred to as first.
[0065] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to detection.” Similarly, depending on the context, the phrases “if determination” or “if detection (of the stated condition or event)” can be interpreted as “when determination” or “in response to determination” or “when detection (of the stated condition or event)” or “in response to detection (of the stated condition or event).”
[0066] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.
[0067] It should be noted that any symbols and / or numbers present in the specification that are not marked in the accompanying drawings are not reference numerals.
[0068] The optional embodiments of this application are described in detail below with reference to the accompanying drawings.
[0069] The embodiments provided in this application are embodiments of a heart rate detection method based on an in-vehicle camera.
[0070] The following is combined Figure 1 The embodiments of this application will be described in detail.
[0071] Step S101: Obtain facial video of the occupant.
[0072] Video refers to various technologies that capture, record, process, store, transmit, and reproduce a series of still images as electrical signals. When the continuous image changes exceed 24 frames per second, according to the principle of visual persistence, the human eye cannot distinguish a single still image; it appears as a smooth and continuous visual effect. Such continuous images are called video. In other words, video consists of multiple frames.
[0073] This application embodiment uses an in-vehicle camera to capture facial video of the occupants, and then uses the facial video to detect the occupants' heart rate signal.
[0074] The facial video may include a facial video of one occupant for heart rate detection; or it may include facial videos of multiple occupants for simultaneous heart rate detection. This application does not impose any limitations on the embodiments described.
[0075] The facial videos described in this application embodiment are all colored videos.
[0076] In some specific embodiments, acquiring the occupant's facial video includes the following steps: periodically acquiring the occupant's facial video for a preset duration.
[0077] In this specific embodiment, a pre-set duration facial video is acquired from the occupant's facial video captured by the vehicle-mounted camera at regular intervals. For example, a facial video segment from the 10 seconds prior to the current time point is acquired every 10 seconds. This allows for uninterrupted heart rate monitoring of the occupant.
[0078] Step S102: Based on the facial video, obtain key pixel information in each preset facial segmentation region of the occupant's face in each frame of the facial image.
[0079] The key pixel information includes the pixel RGB value of the key pixel and the key time point, wherein the pixel RGB value of the key pixel is within a preset green color gamut.
[0080] For example, the occupant's face is divided into three preset facial segmentation regions: the forehead region, the nose and cheek region, and other frontal regions; image information of the mouth and eyes is not included in any of the preset facial segmentation regions. Since the movement of hair, mouth, and eyes has a significant impact on the detection results, this embodiment of the application excludes image information of these parts to improve the accuracy of detection.
[0081] Each frame of a video image is composed of multiple pixels. Each pixel is characterized by its RGB value, and the colors of each pixel in the video image constitute the picture of the video image.
[0082] A key time point refers to the point in time when a key pixel's RGB value appears. Since the video contains timestamps, the image time point of each frame of the facial video can be obtained. The image time point of each frame of the facial image corresponds to the key time point of each key pixel in the facial image.
[0083] Since the green color gamut is most sensitive to changes in occupant's heart rate, this application embodiment limits the pixel RGB values to the green color gamut.
[0084] Step S103: Generate the frequency domain signal and time domain signal of the corresponding preset facial segmentation region based on the pixel RGB values of all key pixels and key time points in each preset facial segmentation region.
[0085] The different angles used to analyze signals are called domains. The time domain and frequency domain can clearly reflect the interaction between signals and interconnects.
[0086] The time domain describes the relationship between a mathematical function or a physical signal and time. For example, the time-domain waveform of a signal can express how the signal changes over time. Typically, when evaluating the performance of digital products, the analysis is performed in the time domain because the product's performance is ultimately measured in the time domain.
[0087] The frequency domain is a coordinate system used to describe the frequency characteristics of a signal. A frequency domain plot shows the amount of signal within each given frequency band over a given frequency range.
[0088] The most important property of the frequency domain is that it is not real, but a mathematical construct. The time domain is the only objectively existing domain, while the frequency domain is a mathematical category that follows specific rules.
[0089] Step S104: Generate the occupant's heart rate signal based on the frequency domain signal and time domain signal of all preset facial segmentation regions.
[0090] This application embodiment converts the RGB values of key pixels within a preset green color gamut and key time points in each preset facial segmentation region of the occupant's face into frequency domain and time domain signals for each preset facial segmentation region. The occupant's heart rate signal is then generated using these frequency and time domain signals. Extracting key pixel information from preset facial segmentation regions reduces the impact of motion on the detection results. The RGB values of pixels within the preset green color gamut are most sensitive to changes in the occupant's heart rate, thus improving the accuracy of the detection results.
[0091] In some specific embodiments, generating the occupant's heart rate signal based on the frequency domain and time domain signals of all preset facial segmentation regions includes the following steps:
[0092] Step S104-1: Eliminate motion noise in the frequency domain signal of each preset facial segmentation region to obtain the frequency domain denoised signal of the corresponding preset facial segmentation region, and filter the time domain signal of each preset facial segmentation region to obtain the time domain filtered signal of the corresponding preset facial segmentation region.
[0093] The purpose of obtaining the frequency domain denoised signal is to improve the signal-to-noise ratio, eliminate motion noise, and thus be able to determine whether there is significant noise interference.
[0094] In some specific embodiments, the step of eliminating motion noise in the frequency domain signal of each preset facial segmentation region to obtain the frequency domain denoised signal of the corresponding preset facial segmentation region includes the following formula:
[0095]
[0096] Among them, SNR i S represents the denoised signal of the i-th preset facial segmentation region, and f represents the frequency of the frequency domain signal of the corresponding preset facial segmentation region; i (f) represents the frequency domain signal of the i-th preset facial segmentation region; U i (f) represents the three frequency domain signals following the maximum frequency domain signal of the i-th preset facial segmentation region, where j represents the sequence number of the three frequency domain signals. In this specific embodiment, f is (30, 200).
[0097] For example, the denoised signal for the forehead area is SNR1, the denoised signal for the nose and cheek area is SNR2, and the denoised signal for other areas on the front is SNR3.
[0098] In some specific embodiments, to remove large jumps and motion noise (such as head rotation and shaking) from the raw heart rate information, the step of filtering the temporal signal of each preset facial segmentation region to obtain the corresponding preset facial segmentation region includes the following steps:
[0099] Step S104-1-1: Subtract the time domain signal of the second key time point from the time domain signal of any preset facial segmentation region at any first key time point to obtain the first signal difference.
[0100] The first critical time point is the critical time point that is immediately after and adjacent to the second critical time point.
[0101] For example, a1 = S(t+1) - S(t); where S(t+1) represents the time-domain signal at the first critical time point, S(t) represents the time-domain signal at the second critical time point, and a1 represents the first signal difference.
[0102] Step S104-1-2: For any of the preset facial segmentation regions, when the absolute value of the first signal difference is less than or equal to a preset empirical difference threshold, the second signal difference is obtained by subtracting the first signal difference from the time domain signal at the third key time point.
[0103] The third key time point is a key time point that is after and immediately adjacent to the first key time point.
[0104] The second critical time point, the first critical time point, and the third critical time point are three adjacent and consecutive critical time points.
[0105] The preset empirical difference threshold is an empirical value obtained through experimentation or simulation.
[0106] For example, continuing the above example, when abs(a1) > Th (i.e., the preset empirical difference threshold), a2 = S(t+2) - a1; where S(t+2) represents the time domain signal at the third key time point, and a2 represents the second signal difference.
[0107] Step S104-1-3: Update the time domain signal of the third key time point based on the second signal difference.
[0108] In some other specific embodiments, when the first signal difference is less than or equal to a preset empirical difference threshold, the time domain signal at the third key time point remains unchanged.
[0109] For example, continuing with the example above, when abs(a1) ≤ Th, S(t+2) remains unchanged.
[0110] Step S104-2: Generate the occupant's heart rate signal based on the frequency domain denoised signals of all preset facial segmentation regions and the time domain filtered signals of the corresponding preset facial segmentation regions.
[0111] In this embodiment, the temporal signals of each preset facial segmentation region are combined with the denoised signals of each preset facial segmentation region to generate the heart rate signal of the occupant, thereby improving the accuracy of the heart rate detection results.
[0112] In some specific embodiments, generating the occupant's heart rate signal based on the frequency domain denoised signal of all preset facial segmentation regions and the time domain filtered signal of the corresponding preset facial segmentation regions includes the following steps:
[0113] Step S104-2-1: When the frequency domain denoising signals of each preset facial segmentation region are all less than or equal to the preset limit threshold, the first adjustment coefficient of the corresponding preset facial segmentation region is generated based on the frequency domain denoising signals of each preset facial segmentation region.
[0114] Specifically, it includes the following formula for calculating the first adjustment factor:
[0115]
[0116] Where, k i SNR represents the first adjustment coefficient for the i-th preset facial segmentation region. i denoises the denoised signal of the i-th preset face segmentation region, and n represents the total number of all preset face segmentation regions.
[0117] For example, the denoised signal for the forehead region is SNR1, the denoised signal for the nose and cheek region is SNR2, and the denoised signal for other areas on the front is SNR3; SNR1, SNR2, and SNR3 are all less than or equal to the preset limit threshold Lm.
[0118] The formula for calculating the first adjustment factor k1 in the forehead region is:
[0119]
[0120] The formula for calculating the first adjustment factor k2 for the nose and cheek area is:
[0121]
[0122] The formula for calculating the first adjustment factor k3 for other areas on the front side is:
[0123]
[0124] Step S104-2-2: Generate the occupant's heart rate signal based on the first adjustment coefficient of all preset facial segmentation regions and the temporal filtering signal of the corresponding preset facial segmentation region.
[0125] Specifically, the formula for calculating the heart rate signal is:
[0126]
[0127] in, Let S(t+1), S(t), and S(t+2) be the time series sequence. It only includes the calculation process at three key time points: S(t+1), S(t), and S(t+2).
[0128] In some specific embodiments, the method further includes the following steps:
[0129] Step S104-2-11: When the frequency domain denoising signal of at least one preset facial segmentation region is less than or equal to a preset limit threshold, and the frequency domain denoising signal of the remaining preset facial segmentation regions is greater than the preset limit threshold, each preset facial segmentation region among the at least one preset facial segmentation region is determined to be a preset first effective segmentation region.
[0130] For example, the denoised signal for the forehead region is SNR1, the denoised signal for the nose and cheek region is SNR2, and the denoised signal for other regions on the front is SNR3. If SNR2 is greater than the preset limit threshold Lm, and SNR1 and SNR3 are both less than or equal to the preset limit threshold Lm, then the forehead region and other regions on the front are both preset first effective segmentation regions.
[0131] Step S104-2-12: Generate the second adjustment coefficient corresponding to the preset first effective segmentation region based on the frequency domain denoising signal of each preset first effective segmentation region.
[0132] The calculation formula for the second adjustment coefficient of each preset first effective segmentation area is the same as the calculation formula for the first adjustment coefficient. For example, the second adjustment coefficient for the forehead area is q1=k1, and the second adjustment coefficient for other areas on the front is q3=k3.
[0133] Step S104-2-13: Generate the occupant's heart rate signal based on the second adjustment coefficients of all preset first effective segmentation regions and the time-domain filtered signal of the corresponding preset first effective segmentation region.
[0134] Specifically, the formula for calculating the heart rate signal is:
[0135]
[0136] Where m represents the total number of all preset first effective segmentation regions, and qr represents the second adjustment coefficient of the r-th preset facial segmentation region. This represents the temporal signal of the r-th preset facial segmentation region.
[0137] In some specific embodiments, after generating the occupant's heart rate signal based on the frequency domain signals and time domain signals of all preset facial segmentation regions, the method further includes the following steps:
[0138] Step S105: Bandpass filtering is performed on the time-domain signal in the heart rate signal to obtain a shaped heart rate signal.
[0139] The shaped heart rate signal outputs a sine wave-like heart rate signal, where the peak value of the sine wave represents the heartbeat position, and the difference between adjacent heartbeat positions is the heartbeat interval. Providing the waveform of the shaped heart rate signal to the occupant allows them to clearly understand their heart rate. The obtained heartbeat interval data also has medical and practical value.
[0140] This application also provides an apparatus embodiment that follows the above embodiments, used to implement the method steps described in the above embodiments. The interpretation of the same names is the same as that in the above embodiments, and the same technical effects are achieved. Therefore, it will not be repeated here.
[0141] like Figure 2 As shown, this application provides a heart rate detection device 200 based on an in-vehicle camera, comprising:
[0142] The first acquisition unit 201 is used to acquire facial videos of the occupants;
[0143] The second acquisition unit 202 is used to acquire key pixel information in each preset facial segmentation region of the occupant's face in each frame of the facial image based on the facial video, wherein the key pixel information includes the pixel RGB value of the key pixel and the key time point, and the pixel RGB value of the key pixel is within a preset green color gamut.
[0144] The first generation unit 203 is used to generate frequency domain signals and time domain signals of the corresponding preset facial segmentation regions based on the pixel RGB values of all key pixels in each preset facial segmentation region and key time points.
[0145] The second generation unit 204 is used to generate the occupant's heart rate signal based on the frequency domain signal and time domain signal of all preset facial segmentation regions.
[0146] Optionally, generating the occupant's heart rate signal based on the frequency domain and time domain signals of all preset facial segmentation regions includes:
[0147] Motion noise in the frequency domain signal of each preset facial segmentation region is eliminated to obtain the corresponding frequency domain denoised signal for the preset facial segmentation region.
[0148] The temporal signal of each preset facial segmentation region is filtered to obtain the corresponding preset facial segmentation region's temporal filtered signal.
[0149] The occupant's heart rate signal is generated based on the frequency domain denoised signal of all preset facial segmentation regions and the time domain filtered signal of the corresponding preset facial segmentation regions.
[0150] Optionally, the step of eliminating motion noise in the frequency domain signal of each preset facial segmentation region to obtain the frequency domain denoised signal of the corresponding preset facial segmentation region includes the following formula:
[0151]
[0152] Among them, SNR i S represents the denoised signal of the i-th preset facial segmentation region, and f represents the frequency of the frequency domain signal of the corresponding preset facial segmentation region; i (f) represents the frequency domain signal of the i-th preset facial segmentation region; U i (f) represents the three frequency domain signals following the maximum frequency domain signal of the i-th preset facial segmentation region, and j represents the sequence number of the three frequency domain signals.
[0153] Optionally, the step of filtering the temporal signal of each preset facial segmentation region to obtain the corresponding preset facial segmentation region includes:
[0154] The first signal difference is obtained by subtracting the time domain signal of the second key time point from the time domain signal of any preset facial segmentation region at any first key time point, wherein the first key time point is a key time point that is after and immediately adjacent to the second key time point.
[0155] For any of the preset facial segmentation regions, when the absolute value of the first signal difference is less than or equal to a preset empirical difference threshold, the second signal difference is obtained by subtracting the first signal difference from the time domain signal of the third key time point, wherein the third key time point is a key time point that is after and immediately adjacent to the first key time point.
[0156] The time-domain signal at the third key time point is updated based on the second signal difference.
[0157] Optionally, generating the occupant's heart rate signal based on the frequency-domain denoised signals of all preset facial segmentation regions and the time-domain filtered signals of the corresponding preset facial segmentation regions includes:
[0158] When the frequency domain denoised signal of each preset facial segmentation region is less than or equal to the preset limit threshold, the first adjustment coefficient of the corresponding preset facial segmentation region is generated based on the frequency domain denoised signal of each preset facial segmentation region.
[0159] The occupant's heart rate signal is generated based on the first adjustment coefficient of all preset facial segmentation regions and the temporal filtering signal of the corresponding preset facial segmentation region.
[0160] Optionally, the device further includes:
[0161] The determining unit is used to determine each of the at least one preset face segmentation regions as a preset first effective segmentation region when the frequency domain denoising signal of at least one preset face segmentation region is less than or equal to a preset limit threshold, and the frequency domain denoising signal of the remaining preset face segmentation regions is greater than the preset limit threshold.
[0162] The third generation unit is used to generate a second adjustment coefficient corresponding to the preset first effective segmentation region based on the frequency domain denoising signal of each preset first effective segmentation region.
[0163] The fourth generation unit is used to generate the occupant's heart rate signal based on the second adjustment coefficients of all preset first effective segmentation regions and the time-domain filtered signal of the corresponding preset first effective segmentation region.
[0164] Optionally, the device further includes:
[0165] The filtering unit is used to generate the occupant's heart rate signal based on the frequency domain signal and time domain signal of all preset facial segmentation regions, and then perform bandpass filtering on the time domain signal of the heart rate signal to obtain a shaped heart rate signal.
[0166] This application embodiment converts the RGB values of key pixels within a preset green color gamut and key time points in each preset facial segmentation region of the occupant's face into frequency domain and time domain signals for each preset facial segmentation region. The occupant's heart rate signal is then generated using these frequency and time domain signals. Extracting key pixel information from preset facial segmentation regions reduces the impact of motion on the detection results. The RGB values of pixels within the preset green color gamut are most sensitive to changes in the occupant's heart rate, thus improving the accuracy of the detection results.
[0167] This embodiment provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method steps described in the above embodiment.
[0168] This application provides a non-volatile computer storage medium storing computer-executable instructions that can perform the steps described in the above embodiments.
[0169] Finally, it should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
[0170] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A heart rate detection method based on an in-vehicle camera, characterized in that, include: Obtain facial video of the occupants; Based on the facial video, key pixel information in each preset facial segmentation region of the occupant's face in each frame of the facial image is obtained. The key pixel information includes the pixel RGB value of the key pixel and the key time point. The pixel RGB value of the key pixel is within a preset green color gamut. Based on the pixel RGB values of all key pixels in each preset facial segmentation region and the key time points, the frequency domain signal and time domain signal of the corresponding preset facial segmentation region are generated respectively. The occupant's heart rate signal is generated based on the frequency domain and time domain signals of all preset facial segmentation regions; Specifically, motion noise in the frequency domain signal of each preset facial segmentation region is eliminated to obtain the frequency domain denoised signal of the corresponding preset facial segmentation region, and the time domain signal of each preset facial segmentation region is filtered to obtain the time domain filtered signal of the corresponding preset facial segmentation region. The occupant's heart rate signal is generated based on the frequency domain denoised signal of all preset facial segmentation regions and the time domain filtered signal of the corresponding preset facial segmentation regions. When the frequency domain denoising signal of each preset facial segmentation region is less than or equal to the preset limit threshold, the first adjustment coefficient of the corresponding preset facial segmentation region is generated based on the frequency domain denoising signal of each preset facial segmentation region. The occupant's heart rate signal is generated based on the first adjustment coefficient of all preset facial segmentation regions and the temporal filtering signal of the corresponding preset facial segmentation region. Specifically, the formula for the first adjustment factor is: Where, k i SNR represents the first adjustment coefficient for the i-th preset facial segmentation region. i Let represent the denoised signal of the i-th preset facial segmentation region, and n represent the total number of all preset facial segmentation regions. The traversal index represents the summation operation, referring to the sequence number of any preset facial segmentation region, and its value covers all preset facial segmentation regions; Furthermore, the occupant's heart rate signal is generated based on the first adjustment coefficient of all preset facial segmentation regions and the temporal filtering signal of the corresponding preset facial segmentation region; Specifically, the formula for calculating the heart rate signal is: in, Let S(t+1), S(t), and S(t+2) be the time series sequence. It only includes the calculation process at three key time points: S(t+1), S(t), and S(t+2). This represents the first adjustment coefficient for the i-th preset facial segmentation region.
2. The method according to claim 1, characterized in that, The process of eliminating motion noise in the frequency domain signal of each preset facial segmentation region to obtain the corresponding frequency domain denoised signal for the preset facial segmentation region includes the following formula: Among them, SNR i S represents the denoised signal of the i-th preset facial segmentation region, and f represents the frequency of the frequency domain signal of the corresponding preset facial segmentation region; i (f) represents the frequency domain signal of the i-th preset facial segmentation region; U i,j (f) represents the three frequency domain signals following the maximum frequency domain signal of the i-th preset facial segmentation region, and j represents the sequence number of the three frequency domain signals.
3. The method according to claim 1, characterized in that, The step of filtering the temporal signal of each preset facial segmentation region to obtain the corresponding preset facial segmentation region includes: The first signal difference is obtained by subtracting the time domain signal of the second key time point from the time domain signal of any preset facial segmentation region at any first key time point, wherein the first key time point is a key time point that is after and immediately adjacent to the second key time point. For any of the preset facial segmentation regions, when the absolute value of the first signal difference is less than or equal to a preset empirical difference threshold, the second signal difference is obtained by subtracting the first signal difference from the time domain signal of the third key time point, wherein the third key time point is a key time point that is after and immediately adjacent to the first key time point. The time-domain signal at the third key time point is updated based on the second signal difference.
4. The method according to claim 1, characterized in that, The method further includes: When the frequency domain denoising signal of at least one preset facial segmentation region is less than or equal to a preset limit threshold, and the frequency domain denoising signal of the remaining preset facial segmentation regions is greater than the preset limit threshold, each preset facial segmentation region among the at least one preset facial segmentation region is determined to be a preset first effective segmentation region. A second adjustment coefficient corresponding to the preset first effective segmentation region is generated based on the frequency domain denoising signal of each preset first effective segmentation region; The occupant's heart rate signal is generated based on the second adjustment coefficients of all preset first effective segmentation regions and the time-domain filtered signal of the corresponding preset first effective segmentation region.
5. The method according to claim 1, characterized in that, After generating the occupant's heart rate signal based on the frequency domain and time domain signals of all preset facial segmentation regions, the method further includes: The time-domain signal in the heart rate signal is subjected to bandpass filtering to obtain a shaped heart rate signal.
6. A heart rate detection device based on a vehicle-mounted camera, characterized in that, The apparatus is used to perform the detection method of any one of claims 1-5, comprising: The first acquisition unit is used to acquire facial videos of the occupants; The second acquisition unit is used to acquire key pixel information in each preset facial segmentation region of the occupant's face in each frame of the facial image based on the facial video, wherein the key pixel information includes the pixel RGB value of the key pixel and the key time point, and the pixel RGB value of the key pixel is within a preset green color gamut. The first generation unit is used to generate the frequency domain signal and time domain signal of the corresponding preset facial segmentation region based on the pixel RGB values of all key pixels in each preset facial segmentation region and the key time points. The second generation unit is used to generate the occupant's heart rate signal based on the frequency domain signal and time domain signal of all preset facial segmentation regions.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.
8. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Heart rate information acquisition method, device, computer equipment and storage medium
CN111134650A
Method for intelligently detecting physiological indexes of a human body and nursing equipment
CN112244796A