Vehicle-mounted physiological data detection method and device, storage medium, system and vehicle

By detecting and selecting stable video frame sequences in the vehicle cockpit, the problem of on-board physiological data detection of video instability in the vehicle driving environment is solved, and more accurate and efficient physiological data detection is achieved.

CN120220201APending Publication Date: 2025-06-27SHANGHAI AUTOMOTIVE FLEXIBLE ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311818780.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the driving environment of the vehicle, the vehicle-mounted non-contact physiological data detection method is difficult to ensure the stability of video recording, affecting the accuracy of user physiological data detection.

Method used

By obtaining the target video of the target occupant in the vehicle cockpit, an image sequence is generated, the stability of the image sequence is detected, and the valid frame sequence that meets the predetermined conditions is selected, and finally the physiological data of the target occupant is extracted based on the effective frame sequence.

Benefits of technology

In unstable environments such as vehicle driving, users' physiological data can be extracted accurately and efficiently, improving the accuracy and efficiency of user physiological data detection in on-board scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220201A_ABST
    Figure CN120220201A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted physiological data detection method and device, a storage medium, a system and a vehicle. The vehicle-mounted physiological data detection method comprises the steps that a target video of a target passenger in a vehicle cabin is acquired, and the target video comprises continuous multi-frame face images of the target passenger; generating at least one image sequence based on the target video of the target passenger, wherein each image sequence comprises continuous multiple frames of facial region-of-interest images; detecting the stability degree of the at least one image sequence so as to select a part of facial interested images with the stability degree meeting the requirement from the at least one image sequence to form an effective frame sequence; and extracting physiological data of the target occupant based on the effective frame sequence. The detection accuracy and the detection efficiency of the vehicle-mounted physiological data can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a method, device, storage medium, system and vehicle for detecting in-vehicle physiological data. Background Art

[0002] With the rapid development of intelligent technologies, intelligent cockpit systems have gradually become a popular research field. By integrating various sensors, cameras and artificial intelligence algorithms inside the vehicle, intelligent cockpit systems provide a more intelligent, comfortable and safe driving and riding experience for drivers and passengers. Applying non-contact physiological data detection methods to intelligent cockpit systems not only makes it more convenient to detect physiological data inside the vehicle cockpit, but also increases the intelligence of intelligent cockpit systems.

[0003] Non-contact physiological data detection methods applied to in-vehicle scenarios, such as Image photoplethysmography (IPPG), collect the facial video of the user through a camera and analyze the physiological data of the user based on the facial video, which have relatively high requirements for the quality and stability of the facial video. However, it is difficult to ensure the stability of video recording in the environment of vehicle driving. If the stability of the extracted video segment is insufficient, it will affect the detection accuracy of the user's physiological data in the in-vehicle scenario.

[0004] Therefore, there is an urgent need to propose a new in-vehicle physiological data detection technology to improve the detection accuracy of the user's physiological data in the in-vehicle scenario. Summary of the Invention

[0005] To solve or at least partially solve the above technical problems, the present application provides a method, device, storage medium, system and vehicle for detecting in-vehicle physiological data, which can effectively improve the detection accuracy of the user's physiological data in the in-vehicle scenario.

[0006] In the first aspect of the present application, a method for detecting in-vehicle physiological data is provided, including:

[0007] Obtaining a target video of a target occupant in a vehicle cockpit, the target video including a continuous multi-frame facial image of the target occupant;

[0008] Generating at least one image sequence based on the target video of the target occupant, each image sequence including a continuous multi-frame facial region of interest image;

[0009] Detecting the stability of the at least one image sequence, and selecting partial facial region of interest images with stability meeting a predetermined condition as an effective frame sequence;

[0010] Extracting the physiological data of the target occupant based on the effective frame sequence.

[0011] In some embodiments, before obtaining the target video of the target occupant in the vehicle cockpit, it further includes: obtaining the first image collected in real time by a camera group installed in the vehicle cockpit, where the camera group is used to collect the target video of the target occupant; detecting whether the first image conforms to a predetermined rule; if the first image does not conform to the predetermined rule, sending a message to cause the target occupant and / or vehicle-related devices to make adjustments so that the first image conforms to the predetermined rule.

[0012] In some embodiments, the detecting the stability of the at least one image sequence and selecting partial face region of interest images with stability meeting a predetermined condition as the valid frame sequence includes:

[0013] For each image sequence, dividing the image sequence into several subsequences of a predetermined length;

[0014] Calculating the stability of the face region of interest images in each subsequence;

[0015] Selecting one or more subsequences with stability meeting the predetermined condition as the valid frame sequence.

[0016] In some embodiments, the calculating the stability of the face region of interest images in each subsequence includes:

[0017] Calculating the fluctuation value between any two frames of the face region of interest images in each subsequence;

[0018] Selecting the maximum value of the fluctuation value as the stability of the subsequence.

[0019] In some embodiments, the selecting one or more subsequences with stability meeting the predetermined condition as the valid frame sequence includes:

[0020] Detecting whether the stability of one or more subsequences is less than or equal to a preset threshold;

[0021] If so, selecting the subsequence as the valid sequence;

[0022] If not, discarding the subsequence.

[0023] In some embodiments, the fluctuation value includes the light brightness difference value and / or the face position movement value.

[0024] In some embodiments, in the same image sequence, each subsequence contains consecutive frame face region of interest images of a predetermined length, adjacent subsequences overlap with each other and the overlapping length is a fixed value or the number of frames between the first frames of adjacent subsequences is fixed, and the predetermined length of the subsequence is at least greater than one physiological data sampling period.

[0025] In some embodiments, physiological data of a target occupant is extracted based on the effective frame sequence, including: extracting an IPPG signal based on the effective frame sequence to obtain the physiological data of the target occupant.

[0026] In a second aspect of the present application, a vehicle-mounted physiological data detection device is provided, including:

[0027] An acquisition unit, configured to acquire a target video of a target occupant in a vehicle cockpit, where the target video includes consecutive multiple frame facial images of the target occupant;

[0028] A region extraction unit, configured to generate at least one image sequence based on the target video of the target occupant, and each image sequence includes consecutive multiple frame facial region of interest images;

[0029] An image detection unit, configured to detect the stability degree of the at least one image sequence, so as to select partial facial region of interest images with a stability degree meeting requirements from the at least one image sequence to form an effective frame sequence;

[0030] An image analysis unit, configured to extract physiological data of the target occupant based on the effective frame sequence.

[0031] In a third aspect of the present application, a computing device is provided, including at least one processor and at least one memory, where the memory stores program instructions, and when the program instructions are executed by the at least one processor, the at least one processor is caused to execute the above-mentioned vehicle-mounted physiological data detection method.

[0032] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which program instructions are stored, and characterized in that when the program instructions are executed by a computer, the computer is caused to execute the above-mentioned vehicle-mounted physiological data detection method.

[0033] In a fifth aspect of the present application, a vehicle-mounted physiological data detection system is provided, including a computing device and a camera group;

[0034] The computing device includes at least one processor and at least one memory, and the memory stores program instructions, and when the program instructions are executed by the at least one processor, the at least one processor is caused to execute the vehicle-mounted physiological data detection method as claimed above;

[0035] The camera group includes one or more cameras installed in the vehicle cockpit, and is configured to collect a target video of the target occupant;

[0036] In a sixth aspect of the present application, a vehicle is provided, including the above-mentioned vehicle-mounted physiological data detection device, the above-mentioned computing device or the above-mentioned vehicle-mounted physiological data detection system.

[0037] Compared with the prior art, the present application has the following beneficial effects:

[0038] In the embodiments of the present application, by extracting multiple image sequences containing images of the facial region of interest from a target video, detecting the light change situation and the facial shaking situation of each frame of the image of the region of interest in these image sequences to select multiple frames of images of the facial region of interest to form an effective frame sequence, and finally using the effective frame sequence to obtain the physiological data of the target occupant. Thus, multiple frames of images of the region of interest with a small shaking amplitude of the user's face and stable light change can be found in the target video, and the physiological data of the target occupant can be obtained through these images of the region of interest. Therefore, in scenarios such as vehicle driving where the light changes frequently in the vehicle cockpit and the user's face is prone to shaking, the physiological data of the user can be accurately and efficiently extracted from the user's video, and the detection accuracy and detection efficiency of the user's physiological data in the vehicle-mounted scenario can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the implementation manners of the present application, the relevant drawings will be briefly introduced below. It can be understood that the drawings described below are only used to illustrate some implementation manners of the present application, and those of ordinary skill in the art can also obtain many other technical features and connection relationships not mentioned in this text based on these drawings.

[0040] Figure 1 is a schematic flowchart of the vehicle-mounted physiological data detection method provided by the embodiments of the present application;

[0041] Figure 2a is an example diagram of subsequence division in a preferred embodiment of the present application;

[0042] Figure 2b is an example diagram of subsequence division in a preferred embodiment of the present application.

[0043] Figure 3 is a schematic structural diagram of the vehicle-mounted physiological data detection device provided by the embodiments of the present application;

[0044] Figure 4 is a schematic structural diagram of the computing device provided by the embodiments of the present application;

[0045] Figure 5 is an example structural diagram of the vehicle-mounted physiological data detection system provided by the embodiments of the present application;

[0046] Figure 6 is an example diagram of the vehicle provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present application will be described in detail below with reference to the drawings.

[0048] As described above, when detecting the physiological data of a user in a vehicle scenario using a non-contact method such as IPPG, there are relatively high requirements for the brightness and stability of the user's video. However, in a vehicle scenario, especially during vehicle driving, the light changes in the vehicle cockpit and the shaking of the user's face caused by vehicle bumps will both lead to unstable recording of the user's video, making it difficult for the recorded video quality to meet the requirements of technologies such as IPPG. As a result, the accuracy and detection efficiency of the user's physiological data detection in the vehicle scenario are poor.

[0049] In view of this, the embodiments of the present application provide the following vehicle-mounted physiological data detection methods, devices, equipment, storage media, systems, and vehicles. By extracting multiple image sequences containing images of the face region of interest from the target video, detecting the light changes and face shaking conditions of each frame of the image of the region of interest in these image sequences to select multiple frames of images of the face region of interest to form an effective frame sequence, and finally using the effective frame sequence to obtain the physiological data of the target occupant. Thus, it is possible to find multiple frames of images of the region of interest with a small shaking amplitude of the user's face and stable light changes in the target video, and obtain the physiological data of the target occupant through these images of the region of interest. Therefore, in scenarios where the light changes frequently in the vehicle cockpit, such as during vehicle driving, and the user's face is prone to shaking, the physiological data of the user can be accurately and efficiently extracted from the user's video, improving the detection accuracy and detection efficiency of the user's physiological data in the vehicle scenario.

[0050] The embodiments of the present application are applicable to various vehicle scenarios. Exemplarily, the embodiments of the present application can be applied to scenarios with a large number of vehicle occupants and complex occupants, such as buses, coaches, and passenger cars, and can also be applicable to scenarios with a small number of vehicle occupants, such as cars. Of course, the embodiments of the present application can also be applied to other means of transportation, such as ships and airplanes. The specific application scenarios of the embodiments of the present application are not limited herein.

[0051] The following elaborates in detail the specific implementation manners of the present application.

[0052] Vehicle-mounted Physiological Data Detection Method

[0053] Figure 1 The flowchart of the vehicle-mounted physiological data detection method provided by the embodiments of the present application is shown. Refer to Figure 1 As shown, the vehicle-mounted physiological data detection method of the embodiments of the present application may include:

[0054] Step S110, obtaining a target video of a target occupant in a vehicle cockpit, where the target video includes multiple consecutive face images of the target occupant;

[0055] The target video contains consecutive multiple-frame images of the face of the target occupant. These images may include partial facial features of the target occupant or may include the complete facial features of the target occupant. At least some of the frame images in the target video contain facial features of the target occupant, and the facial features may be partial facial features of the target occupant or the facial contour of the target occupant.

[0056] In step S110, the length of the target video can be preset. For example, the target video can be consecutive multiple-frame images with a length of 30 seconds to 60 seconds. The target video can be transmitted to the computing device in real time after being captured by a camera group installed in the vehicle cockpit so that the computing device can obtain the target video of the target occupant. The camera group can include one or more cameras and can be installed at any suitable position in the vehicle cockpit as long as it can capture the target video of the target occupant. When the camera group includes multiple cameras, different cameras can be installed at different positions (such as the sun visor, the backrest, one side of the roof, etc.) so as to capture the target videos of occupants at different positions through different cameras, thereby ensuring that the physiological data of each occupant in the vehicle cockpit can be detected.

[0057] Considering that the image quality of the target video directly affects the accuracy of detecting the physiological data of the target occupant, therefore, before step S110, it may further include: a step of detecting the image quality of the face of the target occupant. The step of detecting the image quality of the face of the target occupant may include but is not limited to: obtaining a first image captured in real time by a camera group installed in the vehicle cockpit, where the camera group is used to capture the target video of the target occupant; detecting whether the first image conforms to a predetermined rule; if the first image does not conform to the predetermined rule, sending a message to cause the target occupant and / or vehicle-related devices to make adjustments so that the first image conforms to the predetermined rule. Among them, the predetermined rule may include but is not limited to whether the first image contains the facial contour of the target occupant, whether the clarity and / or brightness of the facial image of the target occupant meet the preset requirements, etc.

[0058] In some embodiments, detecting whether the field of view of the camera group covers the facial contour of the target occupant may include: obtaining a first image captured in real time by the camera group installed in the vehicle cockpit, where the camera group is used to capture the target video of the target occupant; detecting whether the first image contains the facial contour of the target occupant; if the first image does not contain the facial contour of the target occupant, sending a first prompt, where the first prompt is used to prompt the target occupant to adjust their own position or posture so that the first image contains the facial contour of the target occupant, and if the first image contains the facial contour of the target occupant, controlling the camera group to capture the target video so as to obtain the target video of the target occupant. If the first image contains the facial contour of the target occupant, it indicates that the field of view of the camera group can cover the facial contour of the target occupant; when the first image does not contain the facial contour of the target occupant, it indicates that the field of view of the camera group may not cover the facial contour of the target occupant. Thus, the target video can be recorded when the camera group can capture the face of the target occupant, thereby ensuring that the quality of the target video meets the requirements.

[0059] There can be various ways to detect whether the first image contains the facial contour of the target occupant. In one example, the first image and a pre-set detection frame can be shown to the target occupant in real time, and the target occupant can visually judge whether the first image contains their own facial contour and whether the facial contour is within the detection frame. If a confirmation instruction input by the target occupant is received, it indicates that the first image contains the facial contour of the target occupant at this time. If no confirmation instruction input by the target occupant is received, the first icon can be continuously captured by the camera group and shown in real time until a confirmation instruction input by the target occupant is received. In one example, the first image can be subjected to image recognition by means of a pre-configured image recognition algorithm or a machine learning model. If facial contour features can be recognized in the first image, it indicates that the first image contains the facial contour of the target occupant at this time, otherwise it indicates that the first image does not contain the facial contour of the target occupant at this time. In addition, other methods can also be used, and the embodiments of the present application do not limit this.

[0060] In some embodiments, detecting the clarity and / or brightness of the facial image of the target occupant may include: obtaining a first image captured in real time by a camera group installed in the vehicle cockpit, where the camera group is used to capture the target video of the target occupant; detecting the clarity and / or brightness of the facial area of the first image; if the clarity and / or brightness of the facial area in the first image do not meet the preset requirements (for example, the clarity is lower than the preset clarity threshold and / or the brightness is lower than the preset brightness threshold), a light adjustment instruction is sent to the vehicle's lighting system or the camera group, and the lighting system or the camera group adjusts its own working state according to the light adjustment instruction (for example, controlling some intelligent lights in the vehicle cockpit to turn on or off, adjusting the brightness of the fill light of the camera group, etc.) until the clarity and / or brightness of the facial area in the first image meet the preset requirements. Thus, the external light can be adjusted by adjusting the light environment in the vehicle cockpit or the fill light of the camera group, so as to adjust the vehicle cockpit environment to meet the recording conditions of the target video, and then record the target video when the face of the target occupant is clearly visible, ensuring that the quality of the target video meets the requirements.

[0061] There can also be various ways to detect the clarity and / or brightness of the facial area of the first image. For example, the clarity and / or brightness of the facial area of the first image can be obtained by performing an image quality evaluation on the first image, and the first image can also be shown to the target occupant in real time so that the target occupant can visually judge whether the clarity / brightness of the first image meets the requirements. The specific implementation method of detecting the clarity and / or brightness of the facial area of the first image is similar to the method of detecting whether the first image contains a facial contour described above, and will not be elaborated here.

[0062] After detecting the quality of the facial image of the target occupant, step S110 may include: sending a recording instruction to the camera group, and the camera group collects the target video according to the recording instruction. After the collection is completed, the camera group provides the target video to the computing device for the computing device to perform the processing of subsequent steps S120 to step 140. In step S110, it may also include: sending a first prompt until the end of the target video collection when ensuring that the current environment meets the recording requirements, and the first prompt is used to remind the target occupant to maintain the current posture or position. Thus, it can further ensure that the quality of the target video can meet the requirements of subsequent processing.

[0063] Step S120, generating at least one image sequence based on the target video of the target occupant, where each image sequence includes a continuous multi-frame facial region of interest image;

[0064] In step S120, the facial region of interest images in different image sequences contain different facial features, and the facial features contained in the facial frame region of interest images in the same image sequence are the same. The facial features contained in the facial region of interest images in different image sequences can be partially the same. For example, the facial region of interest of one image sequence can be set as the nasal bridge region between the two eyes, the facial region of interest of another image sequence can be set as the forehead region, the facial region of interest of the third image sequence can be set as the region around the lips, the facial region of interest of the fourth image sequence can be set as the entire facial region, etc. For another example, the facial region of interest of each image sequence can be set as the forehead region or the nasal bridge region.

[0065] In step S120, the above-mentioned multiple image sequences can be obtained by performing multiple rounds of facial region of interest detection on the target. Each image sequence contains consecutive facial region of interest image frames, and the length of the image sequence can be less than or equal to the length of the target video. For example, the target video is a sequence of consecutive multiple frames with a length of 30 seconds. A single image sequence can be a sequence formed by 30 consecutive multiple frames of images or a sequence formed by 25 consecutive multiple frames of images or 20 consecutive multiple frames of images.

[0066] As can be seen from the above, through the processing of step S120, multiple image sequences can be generated from the target video, that is, multiple ROI videos can be generated. The facial features contained in each ROI video can be exactly the same, partially the same, or completely different.

[0067] Step S130: Detect the stability degree of at least one image sequence, and select partial facial region of interest images with a stability degree meeting the requirements from at least one image sequence to form a valid frame sequence;

[0068] The stability degree can include but is not limited to the degree of light change and the degree of facial shaking. The degree of light change can be indicated by the brightness change of the image sequence, and the degree of facial shaking can be indicated by the change in the facial position of the image sequence. In one example, the degree of light change can be represented by the light brightness amplitude below, and the degree of facial shaking can be represented by the facial shaking amplitude below.

[0069] Specifically, step S130 may include: for each image sequence, dividing the image sequence into a number of subsequences of a predetermined length; calculating the stability of the facial region of interest images in each subsequence; and selecting one or more subsequences whose stability meets the predetermined conditions as the effective frame sequences. In one example, calculating the stability of the facial region of interest images in each subsequence may include: calculating the fluctuation value between any two frames of the facial region of interest images in each subsequence; and selecting the maximum value of the fluctuation values as the stability of the subsequence, where the maximum value of the fluctuation values is the highest one among the calculated fluctuation values of multiple pairs of any two frames of the facial region of interest images in each subsequence. Among them, the fluctuation value may include the light brightness amplitude and / or the facial shaking amplitude. In one example, selecting one or more subsequences whose stability meets the predetermined conditions as the effective frame sequences includes: detecting whether the stability of one or more subsequences is less than or equal to a preset threshold; if so, selecting the subsequence as the effective sequence; if not, discarding the subsequence.

[0070] For example, for each image sequence, the image sequence may be divided into subsequences of a predetermined length, and the light brightness amplitude and / or the facial shaking amplitude of the facial region of interest images in each subsequence may be calculated; one or more subsequences whose light brightness amplitude and / or shaking amplitude meet the predetermined conditions are selected from the subsequences in all image sequences, and the selected one or more subsequences are used as the effective frame sequences. Thus, by detecting the light brightness change and the facial position change, a part of the consecutive region of interest images with a smaller light brightness change and a smaller facial shaking amplitude can be selected to form an effective frame sequence for extracting the physiological data of the target occupant.

[0071] In the same image sequence, each subsequence contains consecutive frame facial region of interest images of a predetermined length. Adjacent subsequences may overlap with each other and the overlapping length is a fixed value or the number of frames between the first frames of adjacent subsequences is fixed. The length of the subsequence can be flexibly set according to needs. For example, the length of the subsequence can be set to 6 seconds, 0.1 seconds, or any other length. The number of frames between the first frames of adjacent subsequences and / or the fixed value representing the overlapping length between adjacent subsequences can be flexibly configured according to needs. Assuming the length of the subsequence is 60 frames, the fixed value representing the overlapping length can be set to 58 frames, 57 frames, or any other value, and the number of frames between the first frames of adjacent subsequences can be set to 1, 2, or others.

[0072] Figure 2a and Figure 2b shows an example diagram of the subsequence division of an image sequence. Refer to Figure 2a as shown, Figure 2aIn it, the abscissa represents time, or the number of frames, or the image sequence. If the length of the subsequence is set to 60 frames and the number of frames between the first frames of adjacent subsequences is set to 1, then the image sequence can be divided into the following subsequences: frames 1 - 60, frames 2 - 61, frames 3 - 62, and so on. See Figure 2b as shown in Figure 2b In it, the abscissa represents time or the number of frames, and each square represents one frame of the image. If the length of the subsequence is set to 60 frames and the number of frames between the first frames of adjacent subsequences is set to 5, then the image sequence can be divided into the following subsequences: frames 1 - 60, frames 5 - 65, frames 10 - 70, and so on. It should be noted that the number of frames of the image sequence is a preset fixed value, and the frame images within all subsequences are sourced from the image sequence. The predetermined length of the subsequence is at least greater than one physiological data sampling period.

[0073] There can be various ways to calculate the light brightness amplitude and / or the face shaking amplitude of the face region of interest images in each subsequence. In one example, the process of calculating the light brightness amplitude and / or the face shaking amplitude of the face region of interest images in each subsequence can include: calculating the light brightness difference between any two frames of the face region of interest images in the subsequence, and taking the maximum value of the light brightness differences as the light brightness amplitude of the subsequence; and / or, calculating the face position movement value between any two frames of the face region of interest images in the subsequence, and taking the maximum value of the face position movement values as the face shaking amplitude of the subsequence.

[0074] Before selecting one or more subsequences whose light brightness amplitude and / or shaking amplitude meet the predetermined conditions from the subsequences in all image sequences, it further includes: for each subsequence in each image sequence, determining whether all the light brightness differences of the subsequence are less than or equal to the first preset threshold, or the light brightness amplitude is less than or equal to the first preset threshold. If so, retain the subsequence. If any one or more of the light brightness differences of the subsequence are greater than the first preset threshold, or the light brightness amplitude is greater than the first preset threshold, then discard the subsequence; and / or, for each subsequence in each image sequence, determining whether all the face position movement values of the subsequence are less than or equal to the second preset threshold, or the face shaking amplitude is less than or equal to the second preset threshold. If so, retain the subsequence. If any one or more of the face position movement values of the subsequence are greater than the second preset threshold, or the face shaking amplitude is greater than the second preset threshold, then discard the subsequence. Thus, some image frames with too high brightness or too large face shaking amplitude can be discarded in advance to ensure that the quality of each frame of the face region of interest images in the effective frame sequence meets the requirements of technologies such as IPPG.

[0075] If subsequences of all image sequences are discarded or the total length of the remaining subsequences is less than a preset lower limit of the effective frame length, a second prompt can be sent to the target occupant, and the second prompt is used to remind the target occupant to reshoot the target video. Thus, when the physiological data detection is forced to terminate, the user can be reminded to reshoot in a timely manner, thereby improving the user experience.

[0076] In one example, the second prompt may include the reason why the target video is unavailable, so that the user can adjust the light in the vehicle cockpit in a timely manner or continue the physiological data detection after the vehicle is running smoothly. For example, if most subsequences are discarded because the difference in light brightness is greater than the first preset threshold, the content of insufficient brightness in the vehicle can be included in the second prompt, so that the user can adjust the brightness of the vehicle cockpit and then reshoot the target video. For another example, if most subsequences are discarded because the facial position movement value is greater than the second preset threshold, the content such as excessive vehicle shaking can be included in the second prompt, prompting the user to reshoot the target video after the vehicle is running smoothly.

[0077] In step S130, the predetermined conditions for selecting subsequences can be configured flexibly. For example, the predetermined conditions may include one or more of the following: the light brightness amplitude of the subsequence is the smallest among all subsequences; the facial shaking amplitude of the subsequence is the smallest among all subsequences. Thus, the frame sequences of the facial regions of interest with the smallest change in light brightness and the frame sequences of the facial regions of interest with the smallest facial shaking in all image sequences can be selected to form an effective frame sequence for extracting the physiological data of the target occupant. Of course, the predetermined conditions can also be set to other contents, and the embodiments of the present application do not limit this.

[0078] In one example, m subsequences with the smallest light brightness amplitude in all image sequences can be selected first, and then k subsequences with the smallest facial shaking amplitude are selected from the selected m subsequences, and the k subsequences are used as the effective frame sequence, where m is an integer greater than 1, and k is an integer less than m and greater than or equal to 1.

[0079] In one example, p subsequences with the smallest light brightness amplitude in all image sequences can be selected, and q subsequences with the smallest facial shaking amplitude are selected from all image sequences, and the p subsequences and the q subsequences are spliced in time sequence to form an effective frame sequence or directly used as the effective frame sequence, where p and q are integers greater than or equal to 1.

[0080] It should be noted that the above-mentioned predetermined conditions and the selection strategy of subsequences can be flexibly set according to the requirements of the actual application scenario. The embodiments of the present application do not limit how to select subsequences.

[0081] Considering that the length of the effective frame sequence required for extracting physiological data needs to meet certain requirements, the lower limit of the length of the effective frame sequence or the lower limit of the length of the subsequence can be pre-configured to avoid errors. For example, assuming that the length of each image sequence is 30 seconds, if the frequency of the camera or camera used to collect the target video is 20 hz, then 20 frames of images are collected per second, and 30 * 20 frames of images are collected in 30 seconds. That is, each target video contains 30 * 20 frames of images. Taking heart rate as an example, the highest normal heart rate frequency is 4 Hz, the period is 0.25 seconds, and at least 7.5 consecutive frames of images are required to analyze the heart rate information of one period. At this time, the lower limit of the length of the effective frame sequence or the lower limit of the length of the subsequence can be set to 8 frames or other values greater than 7.5. Of course, in actual collection, usually multiple periods of heart rate information need to be collected for analysis, and the lower limit of the length of the effective frame sequence or the lower limit of the length of the subsequence can be set longer according to the collection requirements.

[0082] Step S140, extract the physiological data of the target occupant based on the effective frame sequence.

[0083] In step S140, the IPPG signal can be extracted based on the effective frame sequence to obtain the physiological data of the target occupant. Specifically, the IPPG signal is calculated by performing image analysis on the effective frame sequence, and then the physiological data such as the heart rate of the target occupant is calculated based on the IPPG signal. In one example, if the effective frame sequence contains multiple subsequences, the image analysis can be performed on each of the subsequences to obtain the IPPG signal of the subsequence, and then the physiological data such as the heart rate of the target occupant is calculated using the IPPG signals of the multiple subsequences. Thus, the accuracy of the IPPG signal can be improved through the mutual verification of multiple subsequences, and further the accuracy of the physiological data can be enhanced.

[0084] To facilitate the occupant to view the physiological data, after step S140, it may further include: presenting the physiological data to the target occupant so that the target occupant can intuitively and clearly understand their own physical condition. For example, the physiological data can be presented to the target occupant through the central control screen or other display screens. The embodiments of the present application are not limited thereto.

[0085] After step S140, it may further include: issuing a third prompt according to the physiological data of the target occupant, and the third prompt is used to remind the target occupant to pay attention to their physical condition. For example, when the heart rate of the target occupant is higher than the pre-set normal heart rate value, a voice prompt with the content "high heart rate" is issued.

[0086] To facilitate the occupant's timely understanding of the progress of physiological data detection, during the execution of steps S110 to S140, it may further include: issuing a fourth prompt, which is used to remind the target occupant of the current progress of physiological data detection. For example, a voice prompt of "Physiological data detection starts" can be issued after step S110 to remind the target occupant that the physiological data detection has started. A voice prompt of "Physiological data detection completed" can also be issued after step S140 to remind the target occupant that their physiological data detection has been completed.

[0087] In the embodiments of the present application, the specific forms of the first prompt, the second prompt, the third prompt, and the fourth prompt may be, but are not limited to, voice, visual prompt, alarm, or light flashing, etc. The embodiments of the present application do not limit this.

[0088] The vehicle-mounted physiological data detection method provided by the embodiments of the present application extracts multiple image sequences containing facial region-of-interest images from the target video, detects the light change situation and facial shaking situation of each frame of the region-of-interest images in these image sequences to select multiple frames of facial region-of-interest images to form an effective frame sequence, and finally performs analysis such as IPPG on the effective frame sequence to obtain the physiological data of the target occupant. Thus, it is possible to find multiple frames of region-of-interest images with a small facial shaking amplitude and stable light change of the user in the target video, and obtain the physiological data of the target occupant through these region-of-interest images. Therefore, in scenarios where the light in the vehicle cockpit changes frequently and the user's face is prone to shaking, such as when the vehicle is driving, the physiological data of the user can be accurately and efficiently extracted from the user's video, improving the detection accuracy and detection efficiency of the user's physiological data in the vehicle-mounted scenario.

[0089] Vehicle-mounted physiological data detection device

[0090] Figure 3 The structural schematic diagram of the vehicle-mounted physiological data detection device 300 provided by the embodiments of the present application is shown. Refer to Figure 3 , the vehicle-mounted physiological data detection device 300 may include:

[0091] An acquisition unit 301, configured to acquire a target video of a target occupant in a vehicle cockpit, where the target video includes multiple consecutive facial images of the target occupant;

[0092] A region extraction unit 302, configured to generate at least one image sequence based on the target video of the target occupant, and each image sequence includes multiple consecutive facial region-of-interest images;

[0093] An image detection unit 303, configured to detect the stability degree of the at least one image sequence, so as to select partial facial region-of-interest images with a stability degree meeting the requirements from the at least one image sequence to form an effective frame sequence;

[0094] An image analysis unit 304, configured to extract physiological data of a target occupant based on the valid frame sequence.

[0095] In some embodiments, the stability degree includes the degree of light change and / or the degree of face shaking. The degree of light change can be indicated by the brightness change of the image sequence, and the degree of face shaking can be indicated by the change of the face position in the image sequence. In one example, the degree of light change can be represented by the light brightness amplitude, and the degree of face shaking can be represented by the face shaking amplitude.

[0096] In some embodiments, an image detection unit 303 is configured to: for each image sequence, divide the image sequence into a plurality of subsequences with a predetermined length; calculate the stability degree of the face region of interest images in each subsequence; and select one or more subsequences whose stability degree meets a predetermined condition as the valid frame sequence.

[0097] In some embodiments, the image detection unit 303 is configured to: the calculating the stability degree of the face region of interest images in each subsequence includes: calculating the fluctuation value between any two frames of face region of interest images in each subsequence; and selecting the maximum value of the fluctuation value as the stability degree of the subsequence.

[0098] In some embodiments, the image detection unit 303 is configured to: detect whether the stability degree of one or more subsequences is less than or equal to a preset threshold; if so, select the subsequence as the valid sequence; if not, discard the subsequence.

[0099] In some embodiments, the fluctuation value includes the light brightness amplitude and / or the face shaking amplitude.

[0100] For each image sequence, divide the image sequence into subsequences with a predetermined length, calculate the light brightness amplitude and / or the face shaking amplitude of the face region of interest images in each subsequence; select one or more subsequences whose light brightness amplitude and / or shaking amplitude meet a predetermined condition from the subsequences in all image sequences, and use the selected one or more subsequences as the valid frame sequence.

[0101] In some embodiments, the image detection unit 303 is configured to: calculate the light brightness difference between any two frames of face region of interest images in the subsequence, and use the maximum value of the light brightness difference as the light brightness amplitude of the subsequence; and / or calculate the face position movement value between any two frames of face region of interest images in the subsequence, and use the maximum value of the face position movement value as the face shaking amplitude of the subsequence.

[0102] In some embodiments, the image detection unit 303 is further configured to: for each subsequence in each image sequence, determine whether all the light brightness differences of the subsequence are less than or equal to a first preset threshold, or whether the light brightness amplitude is less than or equal to the first preset threshold. If so, retain the subsequence; if any one or more of the light brightness differences of the subsequence are greater than the first preset threshold, or the light brightness amplitude is greater than the first preset threshold, discard the subsequence; and / or, for each subsequence in each image sequence, determine whether all the facial position movement values of the subsequence are less than or equal to a second preset threshold, or whether the facial shaking amplitude is less than or equal to the second preset threshold. If so, retain the subsequence; if any one or more of the facial position movement values of the subsequence are greater than the second preset threshold, or the facial shaking amplitude is greater than the second preset threshold, discard the subsequence.

[0103] For other technical details of the in-vehicle physiological data detection device 300 in the embodiments of the present application, reference may be made to the method part above, which will not be elaborated here. In practical applications, the in-vehicle physiological data detection device 300 can be implemented by software, hardware, or a combination of both.

[0104] Computing device

[0105] Figure 4 It is a structural schematic diagram of a computing device 400 provided in the embodiments of the present application. The computing device 400 includes: a processor 410 and a memory 420.

[0106] Wherein, the processor 410 can be connected to the memory 420. The memory 420 can be used to store the program code and data. Therefore, the memory 420 can be an internal storage unit of the processor 410, an external storage unit independent of the processor 410, or a component including an internal storage unit of the processor 410 and an external storage unit independent of the processor 410.

[0107] The memory 420 can include a read-only memory and a random access memory, and provide instructions and data to the processor 410. A part of the processor 410 can also include a non-volatile random access memory. For example, the processor 410 can also store information about the device type.

[0108] The processor 410 may be a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a CPLD, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. When the computing device 400 is running, the processor 410 executes the computer-executable instructions in the memory 420 to perform the operation steps of the above vehicle-mounted physiological data detection method.

[0109] Optionally, the computing device 400 may further include components such as a communication interface and a bus.

[0110] It should be understood that the computing device 400 according to the embodiments of the present application may correspond to the corresponding subject executing the methods according to the embodiments of the present application, and the above and other operations and / or functions of each module in the computing device 400 respectively implement the corresponding processes of the methods in this embodiment. For the sake of brevity, they will not be described in detail here.

[0111] In a specific application, the computing device 400 may be implemented as a microcontroller in a vehicle central control system, a vehicle cockpit controller, or other similar devices.

[0112] Computer-readable storage medium

[0113] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the program is run by a processor, the processor is caused to execute the above vehicle-mounted physiological data detection method. Here, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0114] Computer program product

[0115] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is run by a processor, the processor is caused to execute the above-mentioned vehicle-mounted physiological data detection method. Here, the programming language of the computer program product can be one or more, and the programming language can include, but is not limited to, object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as the "C" language.

[0116] Vehicle-mounted physiological data detection system

[0117] An embodiment of the present application also provides a vehicle-mounted physiological data detection system 500. The vehicle-mounted physiological data detection system 500 may include the aforementioned computing device 400 and a camera group 510. Among them, the camera group 510 may include one or more cameras installed in the vehicle cockpit for collecting target videos of target occupants.

[0118] In a specific application, each camera in the camera group 510 can be connected to the computing device 400 in a wired or wireless manner respectively, and transmit the collected target video to the computing device 400.

[0119] In one example, the camera group 510 includes a first camera and a second camera. The first camera is used to collect the target video of the driver, and the second camera is used to collect the target videos of all occupants except the driver. The cameras in the camera group 510 can be installed at any set position in the vehicle cockpit, such as on the sun visor, the main in-vehicle display screen, the ceiling-mounted TV, etc.

[0120] In one example, the vehicle-mounted physiological data detection system 500 may further include: a display device 520, including a display screen, for displaying the physiological data of the target occupant on the display screen under the control of the computing device. Exemplarily, the display device can be, but is not limited to, an in-vehicle central control screen or a specially set display screen, or.

[0121] In one example, the camera can be embedded in the sun visor, and a data display screen can be provided in the sun visor. The data display screen is the display device 520 and can be used to display physiological data.

[0122] In one example, the vehicle-mounted physiological data detection system 500 may further include: an output device 540, for issuing a first prompt, a second prompt, a third prompt, and / or a fourth prompt under the control of the computing device. The first prompt is used to prompt the target occupant to adjust their own position or posture, the second prompt is used to remind the target occupant to re-shoot the target video, the third prompt is used to remind the target occupant to pay attention to their physical condition, and the fourth prompt is used to remind the target occupant of the current progress of the physiological data detection. Exemplarily, the output device can be a voice announcer, a speaker, a display screen, or a reminder light, etc.

[0123] In one example, the vehicle-mounted physiological data detection system 500 may further include: a lighting system 550 configured to adjust its own operating state according to a lighting adjustment instruction provided by the computing device. Exemplarily, the lighting system may be an intelligent lighting system in the vehicle cockpit.

[0124] Vehicle

[0125] An embodiment of the present application further provides a vehicle 600, in which a computing device 400, a vehicle-mounted physiological data detection system 500 or a vehicle-mounted physiological data detection device 300 is installed. Figure 6 An example of the vehicle 600 is shown.

[0126] Finally, it should be noted that those of ordinary skill in the art can understand that, in order to enable readers to better understand the present application, many technical details are presented in the embodiments of the present application. However, even without these technical details and various changes and modifications based on the above embodiments, the technical solutions required to be protected by the various claims of the present application can be basically implemented. Therefore, in practical applications, various changes can be made in form and details to the above embodiments without departing from the spirit and scope of the present application.

Claims

1. A vehicle-mounted physiological data detection method, characterized in that, including: Obtain a target video of a target occupant in a vehicle cockpit, where the target video includes consecutive multiple frames of facial images of the target occupant; Generate at least one image sequence based on the target video of the target occupant, where each image sequence includes consecutive multiple frames of facial region of interest images; Detect the stability of the at least one image sequence, and select partial facial region of interest images with stability meeting a predetermined condition as an effective frame sequence; Extract physiological data of the target occupant based on the effective frame sequence.

2. The vehicle-mounted physiological data detection method according to claim 1, characterized in that, Before obtaining the target video of the target occupant in the vehicle cockpit, it further includes: Obtain a first image collected in real time by a camera group installed in the vehicle cockpit, where the camera group is used to collect the target video of the target occupant; Detect whether the first image conforms to a predetermined rule; If the first image does not conform to the predetermined rule, send a message so that the target occupant and / or vehicle-related devices make adjustments to make the first image conform to the predetermined rule.

3. The vehicle-mounted physiological data detection method according to claim 1, characterized in that, The detecting the stability of the at least one image sequence and selecting partial facial region of interest images with stability meeting a predetermined condition as an effective frame sequence includes: For each image sequence, divide the image sequence into several subsequences of a predetermined length; Calculate the stability of the facial region of interest images in each subsequence; Select one or more subsequences with stability meeting a predetermined condition as an effective frame sequence.

4. The vehicle-mounted physiological data detection method according to claim 3, characterized in that, The calculating the stability of the facial region of interest images in each subsequence includes: Calculate the fluctuation value between any two frames of facial region of interest images in each subsequence; Select the maximum value of the fluctuation value as the stability of the subsequence.

5. The vehicle-mounted physiological data detection method according to claim 4, wherein The selecting one or more subsequences with stability meeting a predetermined condition as an effective frame sequence includes: Detect whether the stability of one or more subsequences is less than or equal to a preset threshold; If so, select the subsequence as an effective sequence; If not, discard the subsequence.

6. The vehicle-mounted physiological data detection method according to claim 5, wherein The fluctuation value includes the light brightness difference value and / or the facial position movement value.

7. The vehicle-mounted physiological data detection method according to claim 3, wherein In the same image sequence, each subsequence includes consecutive frame facial region of interest images of a predetermined length, adjacent subsequences overlap with each other and the overlap length is a fixed value or the number of frames between the first frames of adjacent subsequences is fixed, and the predetermined length of the subsequence is at least greater than one physiological data sampling period.

8. The vehicle-mounted physiological data detection method according to claim 1, wherein, Extracting physiological data of the target occupant based on the effective frame sequence includes: extracting IPPG signals based on the effective frame sequence to obtain physiological data of the target occupant.

9. An in-vehicle physiological data detection device, characterized in that, including: An obtaining unit, configured to obtain a target video of a target occupant in a vehicle cockpit, where the target video includes consecutive multiple frames of facial images of the target occupant; A region extraction unit, configured to generate at least one image sequence based on the target video of the target occupant, where each image sequence includes consecutive multiple frames of facial region of interest images; An image detection unit, configured to detect the stability degree of the at least one image sequence, so as to select partial facial region of interest images with stability degree meeting requirements from the at least one image sequence to form an effective frame sequence; An image analysis unit, configured to extract physiological data of the target occupant based on the effective frame sequence.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, The program instructions, when executed by a computer, cause the computer to execute the in-vehicle physiological data detection method according to any one of claims 1-8.

11. A vehicle-mounted physiological data detection system, characterized in that, It includes a computing device and a camera group; The computing device includes at least one processor and at least one memory. The memory stores program instructions which, when executed by the at least one processor, cause the at least one processor to execute the vehicle-mounted physiological data detection method according to any one of claims 1-8; The camera group includes one or more cameras installed in the vehicle cockpit for collecting target videos of the target occupant.

12. A vehicle, characterized in that, It includes the vehicle-mounted physiological data detection device according to claim 9 and the vehicle-mounted physiological data detection system according to claim 11.