Personnel vital sign data detection method based on voiceprint recognition

By collecting and identifying the vital signs and behavioral audio information of the target person, comparing current and historical characteristics, the problem of insufficient accuracy in traditional detection methods is solved, and accurate detection of individual physiological status is achieved, health risks are reduced, and the intelligent level of smart health care is improved.

CN120472907APending Publication Date: 2025-08-12HUIZHOU HECHENG INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510601305.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In traditional technology, when collecting relevant indicators such as pulse, heart rate, respiratory rate and other related indicators to conduct physiological status detection of relevant personnel, it is impossible to accurately describe the specific physiological status of the individual, resulting in low detection accuracy, especially for groups such as the elderly, which may lead to health risks.

Method used

Collect the current vital sign audio information data and behavioral audio information data of the target personnel, extract features through voiceprint recognition technology, compare current and historical features, and judge the stability of physiological state, including determining the stable state under similar circumstances, otherwise judge the abnormal state.

Benefits of technology

It improves the accuracy of detecting individual physiological and healthy states, can detect abnormal changes earlier, reduce health risks, and improve the intelligence level of smart health care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472907A_ABST
    Figure CN120472907A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent health medical treatment, provides a voiceprint recognition-based personnel vital sign data detection method, and aims to solve the problem of low accuracy of personnel physiological state detection in the prior art. Performing corresponding voiceprint recognition to obtain current vital sign audio information features and current behavior audio information features, and determining corresponding historical vital sign audio information features and historical behavior audio information features; detecting whether the current vital sign audio information feature is similar to the historical vital sign audio information feature or not, and detecting whether the current behavior audio information feature is similar to the historical behavior audio information feature or not; under the condition that the corresponding information features are similar, it is determined that the current physiological state of the target person is stable, and the accuracy of detecting whether the physiological state of the target person changes or not can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart health and medical technology, and in particular to a method for detecting vital signs data of a person based on voiceprint recognition. Background Art

[0002] Vital signs are key indicators used to assess the basic physiological functions of the human body. Vital signs include but are not limited to body temperature (Body Temperature), pulse / heart rate (Pulse / Heart Rate, HR), respiratory rate (Respiratory Rate, RR), and blood pressure (Blood Pressure, BP). Vital signs can reflect an individual's physiological state.

[0003] There is a deep physiological correlation between the voiceprint corresponding to the sound and vital signs. For example, the vibration frequency of the vocal cords is regulated by the autonomic nervous system and has mechanical and acoustic coupling with the heartbeat (such as the micro-vibration of the larynx caused by the carotid artery pulsation). When speaking, acoustic characteristics such as the vibration frequency of the vocal cords and respiratory airflow rate are affected by physiological parameters such as heart rate, respiratory rate, and blood oxygen saturation. In addition, including but not limited to the groaning of a person in pain and the groaning of a person in depression, they also reflect the physiological state or physical condition of the relevant person at a deep physiological level. When the relevant person emits the sound result corresponding to an abnormal sound, there must be a corresponding abnormal reason, whether large or small, in physiological or health. Therefore, there is a deep physiological correlation between voiceprints and vital signs, and therefore there is a deep physiological correlation and technical coupling between voiceprint recognition and vital signs. That is, through voiceprint recognition of corresponding vital signs, the physiological state, physical state, or health status of the relevant person's vital signs can be judged.

[0004] With the development of artificial intelligence and smart terminals, dynamic detection of vital signs, including voiceprint recognition, has become the core of smart healthcare. For example, by detecting the breathing, heartbeat, or voice of relevant personnel through smart terminals corresponding to smart watches and smart phones, and detecting the vital signs of relevant personnel through sound, dynamic detection of the physiological or physical state of relevant personnel can be achieved. When abnormalities are detected in relevant personnel, appropriate health measures such as reminding relevant personnel to pay attention to their health or to undergo appropriate physical examinations can be taken. It can also be used as a monitoring method for the elderly or people in health care.

[0005] In traditional technology, when the physiological state of relevant personnel is detected through sound, the physiological state of relevant personnel is generally detected by collecting relevant indicators such as volume, pulse, heart rate, respiratory rate, blood pressure, etc., and corresponding processing is performed when abnormalities of relevant personnel are detected.

[0006] However, the inventors have realized that in conventional technologies, when detecting the physiological state of a person by collecting relevant indicators such as pulse, heart rate, respiratory rate, or blood pressure, the aforementioned indicators are limited in that they are only applicable to the common characteristics of a group. These indicators cannot accurately describe the person's own consistent physiological state, thus reducing the accuracy of detecting the specific physiological state corresponding to the individual person. For example, including but not limited to the physiological natural aging of the elderly, although the aforementioned indicators may be used to determine that the person is abnormal, the person may have been in this state for a long time, with this physical constitution. This state is physiologically stable and balanced for the person's body. If the aforementioned indicators are simply used to detect the vital signs of the person and then determine the person's physiological state, physical condition, or health status, the judgment of the specific physiological state of the person will be inaccurate, which may in turn bring health risks to the person.

[0007] Therefore, how to improve the accuracy of physiological status detection of personnel has become an urgent problem that needs to be solved in the field of smart health care. Summary of the Invention

[0008] The technical problem solved by the present invention is to solve the problem of low accuracy in detecting the physiological status of a person in traditional technologies.

[0009] In order to solve the above technical problems, the present invention provides the following technical solutions: collecting the current vital signs audio information data and the corresponding current behavior audio information data of the target person; performing voiceprint recognition on the current vital signs audio information data to obtain the current vital signs audio information features, and performing voiceprint recognition on the current behavior audio information data to obtain the current behavior audio information features; determining the corresponding historical vital signs audio information features and historical behavior audio information features; detecting whether the current vital signs audio information features are similar to the historical vital signs audio information features, and detecting whether the current behavior audio information features are similar to the historical behavior audio information features; when each of the above sets of information features are similar, determining that the current physiological state of the target person is in a stable state.

[0010] As a preferred embodiment of the method for detecting vital signs data of a person based on voiceprint recognition according to the present invention, the method further includes: when the current vital signs audio information feature is not similar to the historical vital signs audio information feature, determining the preset standard vital signs audio information feature; calculating the characteristic distance between the current vital signs audio information feature and the preset standard vital signs audio information feature to obtain a first characteristic distance; calculating the characteristic distance between the historical vital signs audio information feature and the preset standard vital signs audio information feature to obtain a second characteristic distance; judging whether the first characteristic distance is greater than the second characteristic distance; and when the first characteristic distance is greater than the second characteristic distance, determining that the current physiological state of the target person is in an abnormal state.

[0011] The beneficial effects of the present invention are as follows: by collecting the current vital sign audio information data and the current behavior audio information data of the target person, and performing corresponding voiceprint recognition, the current vital sign audio information characteristics and the current behavior audio information characteristics are obtained, and the corresponding historical vital sign audio information characteristics and the historical behavior audio information characteristics are determined; when the above corresponding information characteristics are similar, it is determined that the current physiological state of the target person is stable, so as to detect whether the vital sign data of the target person is consistent and stable from the perspective of the vital sign audio information and behavior audio information of the target person, the current audio and the corresponding historical audio, so as to detect the target person's own consistent physiological state, physical state or Whether the health status of the target person has changed, since the target person's vital signs audio information and behavioral audio information, current audio and corresponding historical audio are all personalized sounds adapted to the target person himself, and vital signs constrain behavior, behavior reflects vital signs, that is, behavior and vital signs are related to each other, and are sounds with certain behavioral and performance rules. Therefore, with the help of the target person's vital signs audio information and behavioral audio information, current audio and corresponding historical audio, the accuracy of detecting whether the target person's consistent physiological state and health status has changed can be improved, that is, the accuracy of the target person's vital signs data detection is improved, which helps to improve the intelligence level of smart health care. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A flow chart of a method for detecting vital signs of a person based on voiceprint recognition provided by an embodiment of the present invention;

[0013] Figure 2 A schematic diagram of the first sub-flow of the method for detecting vital signs data of a person based on voiceprint recognition provided by an embodiment of the present invention;

[0014] Figure 3 This is a schematic diagram of the second sub-flow of the method for detecting vital signs data of a person based on voiceprint recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0015] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.

[0016] An embodiment of the present invention provides a method for detecting vital signs data of a person based on voiceprint recognition. The method can be applied to devices including but not limited to wearable devices, smart phones, tablet computers, and the like, and can be used when detecting the physiological state, physical state or health state of a person based on but not limited to voiceprint recognition.

[0017] In the face of the technical problem of low accuracy in detecting the physiological state of a person in traditional technologies, the inventors proposed a method for detecting vital signs data of a person based on voiceprint recognition according to an embodiment of the present invention. The core idea of the embodiment of the present invention is: to detect the current vital signs audio information and behavioral audio information of the target person, and compare the vital signs audio information and behavioral audio information with the corresponding historical audio of the target person, so as to detect whether the vital signs data of the target person are consistent and stable from the perspective of the vital signs audio information and behavioral audio information of the target person, the current audio and the corresponding historical audio, so as to detect whether the target person's own consistent physiological state, physical state or health state is consistent. Whether there is any change, since the target person's vital signs audio information and behavioral audio information, current audio and corresponding historical audio are all personalized sounds adapted to the target person himself, and vital signs constrain behavior, behavior reflects vital signs, that is, behavior and vital signs are related to each other, and are sounds with certain behavioral and performance rules. Therefore, with the help of the target person's vital signs audio information and behavioral audio information, current audio and corresponding historical audio, the accuracy of detecting whether the target person's consistent physiological state and health state has changed can be improved, that is, the accuracy of the target person's vital signs data detection is improved, which helps to improve the intelligence level of smart health care.

[0018] The present invention is described in detail below through specific examples.

[0019] Example 1, please refer to Figure 1 and Figure 2 , Figure 1 A flow chart of a method for detecting vital signs of a person based on voiceprint recognition provided by an embodiment of the present invention is provided. Figure 2The overall flow chart of the method for detecting vital signs data of a person based on voiceprint recognition provided by an embodiment of the present invention is as follows. Figure 1 As shown, in this embodiment, the method includes but is not limited to the following steps S101-S106:

[0020] S101. Collect the current vital sign audio information data and the corresponding current behavior audio information data of the target person.

[0021] Explanatory speaking, the target person's vital signs audio information and behavioral audio information are closely related to and adapted to the target person's physiological state, physical state or health state. That is, no matter what physiological state, physical state or health state there is, there will be corresponding vital signs audio information and behavioral audio information. The vital signs audio information and behavioral audio information are the external manifestations of the target person, and the physiological state, physical state or health state is the internal expression of the target person's physiological functions. For example, for people with an excessively high heart rate, the stress response is an internal manifestation of the physiological state, physical state or health state, and the excessively high heart rate is the corresponding vital sign audio information; for people with unbalanced gait, crutches can be used to assist in movement, and gait imbalance is a manifestation of the physiological state, physical state or health state, and the sound of the crutches touching the ground is the corresponding walking behavior audio information. Therefore, the vital sign audio information and the behavioral audio information are closely related to and adapted to the physiological state, physical state or health state of the target person. Therefore, based on the vital sign audio information and the behavioral audio information, the corresponding physiological state can be detected, including but not limited to whether the corresponding physiological state is normal, abnormal or has changed, and the dynamic and static conditions of the physiological state corresponding to the change can be detected.

[0022] According to the above description, the current vital sign audio information data of the target person and the corresponding current behavior audio information data are collected to monitor whether the physiological state of the target person has changed. The change in physiological state is a physiological state dynamic or static condition that may pose a health risk compared to the stability of the physiological state. Among them, the current vital sign audio information data represents the data of the current vital sign audio information of the target person, and the vital sign audio information represents the vital sign information based on sound audio. The vital sign audio information includes but is not limited to the audio information of the physiological function in the natural state corresponding to the heart rate, respiratory rate, non-speech sounds (such as wheezing sounds), coughing sounds, groaning sounds, and swallowing sounds emitted by the throat; the current behavior audio information data represents the information data corresponding to the sound generated by the current behavior of the target person, and the current behavior audio information data corresponding to the current vital sign audio information data represents the audio information of the sound generated by the current behavior dominated by the behavioral habits and behavioral laws of the target person under the physiological state corresponding to the current vital sign audio information data. Therefore, the current behavior audio information data can be more adapted to and more closely matched with the physiological state of the target person. The current behavior audio information data includes but is not limited to Limited to the audio information data of sounds generated by behaviors corresponding to voice, the sound of crutches hitting the ground, footsteps, the sound of lights turning on and off, the sound of eating and drinking, and the sound of urinating during sleep interruption (i.e., the sound of getting up at night), the current behavior audio information data is generally the sound of behavior in an unnatural state. Therefore, the current vital signs audio information data of the target person and the corresponding current behavior audio information data are all information data that are closely related to and adapted to the current physiological state, physical state or health state of the target person, thereby realizing the consistency and stability of the vital signs data of the target person from the perspective of the vital signs audio information and behavior audio information of the target person, so as to detect whether the target person's own consistent physiological state, physical state or health state has changed, which can improve the accuracy of detecting whether the target person's own consistent physiological state and health state has changed, that is, improve the accuracy of detecting the vital signs data of the target person.

[0023] S102. Perform voiceprint recognition on the current vital sign audio information data to obtain current vital sign audio information features, and perform voiceprint recognition on the current behavior audio information data to obtain current behavior audio information features.

[0024] Explanatory, voiceprint recognition is generally performed based on a corresponding voiceprint recognition model. Thus, a vital sign voiceprint recognition model is pre-set, i.e., a preset vital sign voiceprint recognition model. The preset vital sign voiceprint recognition model represents a model for recognizing the corresponding voiceprint of vital signs. The preset vital sign voiceprint recognition model includes but is not limited to a convolutional neural network (CNN), a recurrent neural network (RNN / LSTM), and a Transformer model. Similarly, a behavioral audio voiceprint recognition model is pre-set, i.e., a preset behavioral audio voiceprint recognition model. The preset behavioral audio voiceprint recognition model represents a model for recognizing the corresponding voiceprint of behavioral audio. The preset behavioral audio voiceprint recognition model includes but is not limited to a convolutional neural network (CNN), a recurrent neural network (RNN / LSTM), and a Transformer model.

[0025] According to the above-mentioned conception and setting, based on the preset vital sign voiceprint recognition model, the current vital sign audio information data is subjected to voiceprint recognition to obtain the current vital sign audio information features, and the current vital sign audio information features represent the information features corresponding to the current vital sign audio information; and based on the preset behavioral audio voiceprint recognition model, the current behavioral audio information data is subjected to voiceprint recognition to obtain the current behavioral audio information features, and the current behavioral audio information features represent the information features corresponding to the current behavioral audio information. For example, the current behavioral audio information features include but are not limited to the sound features corresponding to the sound of a crutch touching the ground and the sound features corresponding to footsteps.

[0026] S103: Determine corresponding historical vital sign audio information features and historical behavior audio information features.

[0027] Explanatory, the historical vital sign audio information features corresponding to the current vital sign audio information features are determined. For example, when the current vital sign audio information features represent the respiratory features of the target person's current respiratory rate, the historical vital sign audio information features represent the respiratory features corresponding to the target person's past respiratory rate; and the historical behavior audio information features corresponding to the current behavior audio information features are determined. For example, when the current behavior audio information features represent the sound features corresponding to the current crutch hitting the ground, the historical behavior audio information features are the sound features corresponding to the same crutch hitting the ground in the past of the target person. Therefore, from the perspective of the target person's current audio and the corresponding historical audio, it is possible to detect whether the target person's vital sign data is consistent and stable, so as to detect whether the target person's own consistent physiological state, physical state or health state has changed. This can improve the accuracy of detecting whether the target person's own consistent physiological state and health state have changed, that is, improve the accuracy of detecting the target person's vital sign data.

[0028] S104: Detecting whether the current vital sign audio information feature is similar to the historical vital sign audio information feature, and detecting whether the current behavior audio information feature is similar to the historical behavior audio information feature;

[0029] S105: If the current vital sign audio information feature is similar to the historical vital sign audio information feature, and the current behavior audio information feature is similar to the historical behavior audio information feature, determining that the current physiological state of the target person is in a stable state;

[0030] S106. When the current vital sign audio information feature is not similar to the historical vital sign audio information feature, or the current behavior audio information feature is not similar to the historical behavior audio information feature, it is determined that the current physiological state of the target person is in an abnormal state.

[0031] Explanatoryly, it is detected whether the current vital sign audio information features are similar to the historical vital sign audio information features, and it is detected whether the current behavior audio information features are similar to the historical behavior audio information features, wherein the current vital sign audio information features and the historical vital sign audio information features are a corresponding set of information features, and the current behavior audio information features and the historical behavior audio information features are a corresponding set of information features. Euclidean distance, Manhattan distance, and cosine similarity can be used to determine whether each of the above corresponding sets of information features are similar.

[0032] When the current vital sign audio information features are similar to the historical vital sign audio information features, and the current behavioral audio information features are similar to the historical behavioral audio information features, it indicates that the corresponding vital signs of the target person have not changed significantly and are relatively stable, and the corresponding behavioral habits have not changed significantly and are relatively stable. It is determined that the current physiological state of the target person is in a stable state, that is, the stable physiological state determines the consistency and stability of the corresponding vital signs and behavioral habits. The stable physiological state means that the corresponding physiological functions are in a stable state. The stable physiological state is a broader and more universal human status measurement benchmark than the health standard. The stable physiological state includes the healthy physiological state, but the range is wider than the healthy physiological state. Especially when it is impossible to require everyone to meet the established health standards, the target person is pursued. The most desired state is a stable physiological state without large fluctuations. On the contrary, when the current vital sign audio information characteristics are not similar to the historical vital sign audio information characteristics, or the current behavioral audio information characteristics are not similar to the historical behavioral audio information characteristics, it indicates that the corresponding vital signs of the target person have changed significantly and are not stable, or the corresponding behavioral habits have changed significantly and are not stable, which indirectly indicates that the physiological state of the target person is unstable and has changed significantly, and it is determined that the current physiological state of the target person is in an abnormal state. At this time, it is generally necessary to further verify whether the physiological state of the target person is changing for the better or for the worse. In order to avoid the occurrence of health risks, especially when it changes in a worse direction, it is necessary to intervene in advance and take corresponding health care measures to ensure that the target person is in the best condition.

[0033] In summary, it is possible to detect whether the current physiological state of the target person is stable or abnormal based on the current vital signs audio information characteristics and historical vital signs audio information characteristics, as well as the current behavior audio information characteristics and historical behavior audio information characteristics. Detection refers to the process of identifying, measuring, analyzing or monitoring the corresponding physiological state of the target person through the above-mentioned technical means to discover its specific attributes, status or abnormal conditions.

[0034] It should be noted that physiological state refers to the functional status of the human body as a whole or a local organ system, while vital signs refer to the most basic, directly measurable physiological parameters. Vital signs are the "window" to physiological state. Vital signs (such as heart rate and blood pressure) are the direct external manifestations of physiological state. By monitoring these indicators, the overall physiological state can be indirectly inferred. That is, changes in vital signs trigger physiological state adjustments, and vice versa. For example, a fever (increased body temperature) will trigger an increase in metabolic rate (a change in physiological state), which further leads to an increase in heart rate (a change in vital signs). Moreover, corresponding vital signs and behavioral habits can only correspond to the functional status of the corresponding whole or local organ system of the human body, that is, to the corresponding physiological state. For example, increased blood pressure may reflect the physiological state corresponding to stress or cardiovascular abnormalities, while an increased respiratory rate may indicate the physiological state corresponding to lung discomfort. In other words, different vital signs and behavioral habits correspond to different functional states of the whole or local organ system of the human body, that is, to different overall or local physiological states. It is impossible to measure the overall or global physiological state based on a limited number of vital signs and behavioral habits.

[0035] In an embodiment of the present invention, the current vital signs audio information data and the corresponding current behavior audio information data of the target person are collected; the current vital signs audio information data are subjected to voiceprint recognition to obtain the current vital signs audio information features, and the current behavior audio information data are subjected to voiceprint recognition to obtain the current behavior audio information features; then the corresponding historical vital signs audio information features and historical behavior audio information features are determined; then the current vital signs audio information features are detected to see whether they are similar to the historical vital signs audio information features, and the current behavior audio information features are detected to see whether they are similar to the historical behavior audio information features; when each of the above groups of information features are similar, it is determined that the current physiological state of the target person is in a stable state; otherwise, it is determined that the current physiological state of the target person is in an abnormal state, thereby obtaining the target person's current physiological state from the target person's vital signs audio information and behavior audio information, current From the perspective of audio and corresponding historical audio, the consistency and stability of the target person's vital signs data are detected to detect whether the target person's own consistent physiological state, physical state or health state has changed. Since the target person's vital signs audio information and behavioral audio information, current audio and corresponding historical audio are all personalized sounds adapted to the target person himself, and vital signs constrain behavior, and behavior reflects vital signs, that is, behavior and vital signs are interrelated, and are sounds with certain behavioral and performance rules. Therefore, with the help of the target person's vital signs audio information and behavioral audio information, current audio and corresponding historical audio, the accuracy of detecting whether the target person's own consistent physiological state and health state has changed can be improved, that is, the accuracy of the target person's vital signs data detection is improved, which helps to improve the intelligence level of smart health care.

[0036] In one embodiment, see Figure 2 , Figure 2 This is a schematic diagram of the first sub-flow of the method for detecting vital signs data of a person based on voiceprint recognition provided by an embodiment of the present invention. Figure 2 As shown, in this embodiment, when the current vital sign audio information feature is not similar to the historical vital sign audio information feature, determining that the current physiological state of the target person is in an abnormal state includes:

[0037] S201: determining a preset standard vital sign audio information feature when the current vital sign audio information feature is not similar to the historical vital sign audio information feature;

[0038] S202: Calculate the characteristic distance between the current vital sign audio information characteristic and the preset standard vital sign audio information characteristic to obtain a first characteristic distance;

[0039] S203, calculating a characteristic distance between the historical vital sign audio information feature and the preset standard vital sign audio information feature to obtain a second characteristic distance;

[0040] S204, determining whether the first characteristic distance is greater than the second characteristic distance;

[0041] S205: If the first characteristic distance is greater than the second characteristic distance, determining that the current physiological state of the target person is abnormal;

[0042] S206: When the first characteristic distance is smaller than the second characteristic distance, it is not determined whether the current physiological state of the target person is abnormal.

[0043] Explanatoryly, standard vital signs audio information features are pre-set, that is, preset standard vital signs audio information features. The preset standard vital signs audio information features represent standard vital signs audio information features. The preset standard vital signs audio information features reflect the information features corresponding to the corresponding vital signs of the target person or universal person under normal circumstances. The preset standard vital signs audio information features can be the information features corresponding to the vital signs of the target person when the corresponding physiological state is stable and normal. The preset standard vital signs audio information features can also be the common information features corresponding to the vital signs of the corresponding universal person group when the corresponding physiological state is stable and normal. The preset standard vital signs audio information features are only used as a relative calculation benchmark for subsequent calculations, and are not taken as an absolute measurement standard. Therefore, the preset standard vital signs audio information features are not limited to the above description, and other corresponding contents can also be taken as relative calculation benchmarks.

[0044] According to the above description and settings, when the current vital sign audio information characteristics are not similar to the historical vital sign audio information characteristics, the preset standard vital sign audio information characteristics are determined; and the characteristic distance between the current vital sign audio information characteristics and the preset standard vital sign audio information characteristics is calculated to obtain a first characteristic distance, and the characteristic distance between the historical vital sign audio information characteristics and the preset standard vital sign audio information characteristics is calculated to obtain a second characteristic distance, wherein the characteristic distance represents the degree of similarity, the greater the similarity, the smaller the characteristic distance, and conversely, the smaller the similarity, the larger the characteristic distance, and the preset standard vital sign audio information characteristics are only used as a relative calculation basis for calculating the characteristic distance, and are only used as a relative reference quantity, and are not taken as an absolute measurement standard, and then it is judged whether the first characteristic distance is greater than the second characteristic distance, so as to compare the first characteristic distance with the second characteristic distance. The size of the distance between the feature distances, when the first feature distance is greater than the second feature distance, indicates that the current vital sign audio information feature is relatively far from the preset standard vital sign audio information feature, and the historical vital sign audio information feature is relatively close to the preset standard vital sign audio information feature, which also indicates that the corresponding physiological state and its corresponding vital sign are developing in a bad direction, and it is determined that the current physiological state of the target person is in an abnormal state. On the contrary, when the first feature distance is less than the second feature distance, it indicates that the current vital sign audio information feature is relatively close to the preset standard vital sign audio information feature, and the historical vital sign audio information feature is relatively far from the preset standard vital sign audio information feature, which also indicates that the corresponding physiological state and its corresponding vital sign are developing in a good direction, and it is not certain that the current physiological state of the target person is in an abnormal state.

[0045] It should be noted that, since the premise at this time is that the current vital sign audio information features are not similar to the historical vital sign audio information features, the first feature distance will not be equal to the second feature distance.

[0046] In an embodiment of the present invention, when the current vital sign audio information feature is not similar to the historical vital sign audio information feature, the preset standard vital sign audio information feature is determined; and the feature distance between the current vital sign audio information feature and the preset standard vital sign audio information feature is calculated to obtain a first feature distance; and the feature distance between the historical vital sign audio information feature and the preset standard vital sign audio information feature is calculated to obtain a second feature distance; then the distance between the first feature distance and the second feature distance is compared; and when the first feature distance is greater than the second feature distance, it is determined that the current physiological state of the target person is in an abnormal state, thereby achieving It is now possible to accurately judge the vital sign development trend of the target person's current vital signs, that is, the vital sign development trend of the target person represented by the current vital sign audio information corresponding to the current vital sign audio information characteristics of the target person, and when the vital sign development trend develops in a bad direction, accurately determine that the current physiological state of the target person is in an abnormal state, which can further improve the accuracy of detecting whether the target person's consistent physiological state and health state have changed, that is, improve the accuracy of the target person's vital sign data detection, and improve the accuracy of detecting that the target person's current physiological state is in an abnormal state, which helps to improve the intelligence level of smart health care.

[0047] In one embodiment, when the current behavior audio information feature is not similar to the historical behavior audio information feature, determining that the current physiological state of the target person is abnormal includes:

[0048] In a case where the current behavior audio information feature is not similar to the historical behavior audio information feature, determining a preset standard behavior audio information feature;

[0049] Calculating a characteristic distance between the current behavior audio information feature and the preset standard behavior audio information feature to obtain a third characteristic distance;

[0050] Calculating a characteristic distance between the historical behavior audio information feature and the preset standard behavior audio information feature to obtain a fourth characteristic distance;

[0051] determining whether the third characteristic distance is greater than the fourth characteristic distance;

[0052] When the third characteristic distance is greater than the fourth characteristic distance, determining that the current physiological state of the target person is in an abnormal state;

[0053] When the third characteristic distance is smaller than the fourth characteristic distance, it is not determined whether the current physiological state of the target person is abnormal.

[0054] Explanatoryly, the standard behavioral audio information features are pre-set, that is, the preset standard behavioral audio information features, the preset standard behavioral audio information features represent the standard behavioral audio information features, the preset standard behavioral audio information features reflect the information features corresponding to the corresponding behaviors of the target person under normal circumstances, the preset standard behavioral audio information features are behavioral features that reflect the general behavioral habits and behavioral regularities of the target person, the preset standard behavioral audio information features are information features corresponding to the behavioral habits and behavioral regularities of the target person when the corresponding physiological state is stable and under normal circumstances, the preset standard behavioral audio information features include but are not limited to the sound features corresponding to the sound of a cane touching the ground and the sound features corresponding to the footsteps, and the above-mentioned sound features include but are not limited to sound frequency and sound loudness.

[0055] According to the above description and setting, when the current behavior audio information feature is not similar to the historical behavior audio information feature, the preset standard behavior audio information feature is determined; and the feature distance between the current behavior audio information feature and the preset standard behavior audio information feature is calculated to obtain a third feature distance, and the feature distance between the historical behavior audio information feature and the preset standard behavior audio information feature is calculated to obtain a fourth feature distance, wherein the feature distance also represents the degree of similarity, the greater the similarity, the smaller the feature distance, and conversely, the smaller the similarity, the larger the feature distance; then determine whether the third feature distance is greater than the fourth feature distance, thereby comparing the distance between the third feature distance and the fourth feature distance, and when the third feature distance is greater than the fourth feature distance, When the third characteristic distance is less than the fourth characteristic distance, it indicates that the current behavior audio information feature is relatively far from the preset standard behavior audio information feature, and the historical behavior audio information feature is relatively close to the preset standard behavior audio information feature, which also indicates that the corresponding physiological state and its corresponding behavior are developing in a bad direction, and it is determined that the current physiological state of the target person is in an abnormal state. On the contrary, when the third characteristic distance is less than the fourth characteristic distance, it indicates that the current behavior audio information feature is relatively close to the preset standard behavior audio information feature, and the historical behavior audio information feature is relatively far from the preset standard behavior audio information feature, which also indicates that the corresponding physiological state and its corresponding behavior are developing in a good direction, and it is not certain that the current physiological state of the target person is in an abnormal state.

[0056] It should be noted that, since the premise at this time is that the current behavior audio information features are not similar to the historical behavior audio information features, the third feature distance will not be equal to the fourth feature distance.

[0057] In an embodiment of the present invention, when the current behavior audio information feature is not similar to the historical behavior audio information feature, a preset standard behavior audio information feature is determined; and a feature distance between the current behavior audio information feature and the preset standard behavior audio information feature is calculated to obtain a third feature distance; and a feature distance between the historical behavior audio information feature and the preset standard behavior audio information feature is calculated to obtain a fourth feature distance; then the distance between the third feature distance and the fourth feature distance is compared; and when the third feature distance is greater than the fourth feature distance, it is determined that the current physiological state of the target person is in an abnormal state, thereby accurately judging the behavior development trend to which the corresponding current behavior of the target person belongs, that is, the corresponding behavior development trend of the target person represented by the current behavior audio information corresponding to the current behavior audio information feature of the target person, and when the corresponding behavior development trend develops in a bad direction, accurately determining that the current physiological state of the target person is in an abnormal state can further improve the accuracy of detecting whether the target person's consistent physiological state and health state have changed, that is, further improve the accuracy of detecting the vital sign data of the target person indirectly reflected by the current behavior audio information feature, and improve the accuracy of detecting that the current physiological state of the target person is in an abnormal state, which helps to improve the intelligence level of smart health care.

[0058] In one embodiment, collecting the current vital sign audio information data of the target person and the corresponding current behavior audio information data includes:

[0059] Determine the internal and external time and space states corresponding to the target person;

[0060] According to the internal and external space-time states, the current vital sign audio information data of the target person and the corresponding current behavior audio information data are collected.

[0061] Explanatory, the internal and external space-time states represent the target person's own state and the state of his or her surrounding environment, wherein the target person's own state is his or her internal space-time state, and the target person's surrounding environment state is his or her external space-time state. The internal space-time state depends on, but is not limited to, the target person's gender, age, genetics, medical history, and dynamic and static states, which are determined by the current individualized internal physiological state of the body. The dynamic and static states include dynamic and static states, which include general activities, aerobic exercise, and strenuous exercise, and the static states include quiet states and resting states. The quiet state means not moving, and the resting state means entering a state of tranquility, which means not moving but also not seeing, thinking, or hearing. The external space-time state depends on, but is not limited to, the time and environmental states corresponding to day, night, and seasons, and includes, but is not limited to, the living room, square, etc. , quiet, noisy, and noisy spatial environment states corresponding to the internal and external space-time states, that is, the internal and external space-time states jointly determine the target person's current vital signs audio information and the corresponding current behavior audio information. That is, for the same target person, different internal and external space-time states will correspond to different current vital signs audio information and the corresponding current behavior audio information. For example, young men and young women, when exercising and at rest, including but not limited to heart rate and breathing are different, and in the above different situations, including but not limited to footsteps, eating sounds, and talking sounds, the corresponding behaviors are also different. Therefore, different internal and external space-time states will correspond to different current vital signs audio information and the corresponding current behavior audio information.

[0062] According to the above description, when collecting the current vital signs audio information data and the corresponding current behavior audio information data of the target person, the internal and external time and space states corresponding to the target person are first determined, and then the current vital signs audio information data and the corresponding current behavior audio information data of the target person are collected based on the internal and external time and space states, and then the corresponding historical vital signs audio information features and historical behavior audio information features are determined based on the internal and external time and space states, so that the current vital signs audio information data and the historical vital signs audio information features, the current behavior audio information data and the historical behavior audio information features of the target person are constrained to the same internal and external time and space states, so that each set of the above information features is more comparable, that is, each set of the above information features is more corresponding and more adapted to the target person himself, which can further improve the detection accuracy of whether the target person's consistent physiological state and health state have changed in the corresponding internal and external time and space states, that is, the accuracy of the target person's vital signs data detection is improved.

[0063] Furthermore, determining the internal and external space-time state corresponding to the target person includes at least one of the following:

[0064] Determining the internal space-time state corresponding to the target person;

[0065] Determine the external space-time state of the target person.

[0066] Specifically, as described above, determining the internal and external space-time states corresponding to the target person includes at least one of the following: determining the internal space-time state corresponding to the target person; determining the external space-time state of the target person, thereby combining the target person's own internal environment with the external environment to jointly determine the internal and external space-time states corresponding to the target person. Since the above-mentioned internal space-time state and external space-time state jointly affect the target person's current vital sign audio information data and the corresponding current behavior audio information data, for example, the heart rate is generally higher during the day than at night, and generally at night there will be sounds including but not limited to the sound of switching lights, the sound of sleep-interrupting urination behavior (i.e., the sound of getting up at night), and And then, according to the internal and external time and space states, the corresponding historical vital signs audio information features and historical behavior audio information features are determined, so that the current vital signs audio information data and historical vital signs audio information features, the current behavior audio information data and historical behavior audio information features of the target person are constrained to the same internal and external time and space states, so that each of the above-mentioned information features is more comparable, that is, each of the above-mentioned information features has a stronger correspondence and is more adapted to the target person himself, which can further improve the accuracy of detecting whether the target person's consistent physiological state and health state have changed under the corresponding internal and external time and space states, that is, improve the accuracy of the target person's vital signs data detection.

[0067] The embodiment of the present invention determines the internal and external time and space states corresponding to the target person, and collects the target person's current vital sign audio information data and the corresponding current behavior audio information data based on the internal and external time and space states. This can make the current vital sign audio information data and the corresponding current behavior audio information data more in line with the internal and external time and space scenes in which the target person is located, so that the current vital sign audio information data and the current behavior audio information data can more accurately reflect the target person's own consistent behavior patterns and performance patterns, and be more adapted to the target person himself. This can further improve the accuracy of detecting whether the target person's own consistent physiological state and health state have changed under the corresponding internal and external time and space states, that is, improve the accuracy of detecting the target person's vital sign data.

[0068] In one embodiment, see Figure 3 , Figure 3 This is a schematic diagram of the second sub-flow of the method for detecting vital signs data of a person based on voiceprint recognition provided by an embodiment of the present invention. Figure 3 As shown, in this embodiment, before collecting the current vital sign audio information data and the corresponding current behavior audio information data of the target person according to the internal and external time and space state, the following is also included:

[0069] S301, collecting initial audio information data of the target person;

[0070] S302, identifying the initial audio information data to obtain initial audio information features;

[0071] S203, identifying vital sign audio information features contained in the initial audio information features to obtain target vital sign audio information features;

[0072] S204, identifying the human behavior audio information features contained in the initial audio information features to obtain target behavior audio information features;

[0073] S305: Establish a correspondence between the target vital sign audio information features and the target behavior audio information features to obtain the target person's vital sign audio information features and their corresponding behavior audio information features.

[0074] Explanatoryly, the initial audio information data of the target person is collected. The initial audio information data generally represents the original audio information data of the target person. The initial audio information data includes vital signs audio information and personnel behavior audio information. Then, the initial audio information data is subjected to voiceprint recognition to obtain initial audio information features, and the vital signs audio information features contained in the initial audio information features are identified to obtain target vital signs audio information features, wherein the vital signs audio information features represent the features of the audio information corresponding to the vital signs of the target person, and the personnel behavior audio information features contained in the initial audio information features are identified to obtain target behavior audio information features, wherein the personnel behavior audio information features represent the features of the audio information corresponding to the corresponding behavior of the target person. Finally, the target vital signs audio information features and the target behavior audio information features are formed into a correspondence relationship to obtain the target person's vital signs audio information features and the corresponding behavior audio information features, thereby combining the target person's vital signs with the behavior information to achieve highly available vital signs and their corresponding physiological state detection, which can be applied to active health monitoring including but not limited to people with mobility difficulties and health care people. Among them, voiceprint recognition is generally performed based on a corresponding pre-set voiceprint recognition model. The voiceprint recognition model includes but is not limited to convolutional neural network (CNN), recurrent neural network (RNN / LSTM), and Transformer model. This is a commonly used technical means and will not be described in detail here.

[0075] Furthermore, identifying the personnel behavior audio information features contained in the initial audio information features to obtain the target behavior audio information features includes:

[0076] Determining a number of initial audio information features based on time series;

[0077] Identify common human behavior audio information features contained in the plurality of initial audio information features to obtain target behavior audio information features.

[0078] Specifically, a number of initial audio information features based on time series are determined, that is, the number of initial audio information features are a series of initial audio information features, and then the common personnel behavior audio information features contained in the number of initial audio information features are identified to obtain target behavior audio information features. The common personnel behavior audio information features are the behavioral information corresponding to behavioral habits and behavioral laws, that is, the target behavior audio information features are the common behavioral features under different time and space behaviors corresponding to the identified behavioral habits and behavioral laws, thereby filtering out noise behavioral interference in a high-noise environment, screening out behaviors that can reflect the behavioral habits and behavioral laws of the target personnel, and making the target behavior audio information features focus on the behavioral commonalities corresponding to the behavioral habits and behavioral laws of the target personnel, and making the target vital signs audio information of the target personnel correspond to and combine with the target behavior audio information, so that the vital signs audio information features of different target personnel and the corresponding behavioral audio information features are more individualized and personalized, so that the corresponding audio information features are more adapted to the target personnel themselves, and so that the corresponding audio information features can more accurately express the physiological state, behavioral habits and behavioral laws of the target personnel themselves.

[0079] The embodiment of the present invention collects the initial audio information data of the target person; and identifies the initial audio information data to obtain the initial audio information features; and identifies the vital signs audio information features contained in the initial audio information features to obtain the target vital signs audio information features; then identifies the person behavior audio information features contained in the initial audio information features to obtain the target behavior audio information features; and then forms a correspondence between the target vital signs audio information features and the target behavior audio information features to obtain the vital signs audio information features of the target person and the corresponding behavior audio information features. The target vital signs audio information of the target person and the target behavior audio information can be corresponded and combined, so that the vital signs audio information features and the corresponding behavior audio information features of different target persons are more individualized and personalized, and the corresponding audio information features are more adapted to the target person himself, so that the corresponding audio information features can more accurately express the physiological state, behavioral habits and behavioral laws of the target person himself, thereby further improving the detection accuracy of whether the target person's consistent physiological state and health state have changed, that is, improving the accuracy of the vital signs data detection of the target person.

[0080] In one embodiment, based on the internal and external spatiotemporal states, collecting the current vital sign audio information data and the corresponding current behavior audio information data of the target person includes at least one of the following:

[0081] During the preset sleep period, the target person's current behavior audio information data is collected, including at least one of the following: audio information data corresponding to the sound of urinating during sleep interruption, footsteps, and the sound of turning on and off lights;

[0082] During the preset non-sleep time period, the current behavior audio information data of the target person is collected, including at least one of the following: auxiliary action audio information data corresponding to the target action auxiliary device, footsteps, and door opening and closing sounds.

[0083] Explanatory, since the vital signs audio information and the vital signs corresponding to it are the physiological functions of the target person themselves and are unchangeable, therefore, after the target person's current vital signs audio information data is collected, the corresponding vital signs may change regardless of their internal and external time and space states, but the vital signs will not change. For example, the heart rate value may change in different internal and external time and space, but the heart rate as a vital sign indicator will not change in different internal and external time and space. However, since the behavioral audio information data is the behavior of the target person, the changes in behavior can be large. For example, behavior during the day and at night are generally different. Therefore, according to the internal and external time and space states, the target person's current vital signs audio information data and the corresponding current behavioral audio information data are collected, including at least one of the following:

[0084] During the preset sleep time period, mainly referring to the sleep time at night, the target person's current behavior audio information data collected includes at least one of the following: audio information data of getting up in the middle of the night corresponding to the sound of sleep-interrupted urination behavior, footsteps, and the sound of switching lights on and off. Among them, sleep-interrupted urination behavior is commonly known as "getting up in the middle of the night", and getting up in the middle of the night will produce noise and movement. Therefore, the corresponding audio information data of getting up in the middle of the night can be collected. The audio information data of getting up in the middle of the night represents the information data of the sound of getting up in the middle of the night behavior. Footsteps and the sound of switching lights on and off are generally inevitable sounds generated at night. The sound of switching lights on and off includes the sound of turning on the lights and the sound of turning off the lights. Therefore, during the preset sleep time period, the target person's current behavior audio information data collected includes at least one of the following: audio information data of getting up in the middle of the night corresponding to the sound of sleep-interrupted urination behavior, footsteps, and the sound of switching lights on and off.

[0085] During the preset non-sleep time period, which generally refers to the non-sleep activity time, such as the daytime activity time, the current behavioral audio information data collected of the target person includes at least one of the following: auxiliary action audio information data corresponding to the target action assistive device, footsteps, and door opening and closing sounds, among which the target action assistive devices include crutches, walkers, and wheelchairs; crutches can be divided into straight-handled crutches and cranked crutches; walkers include but are not limited to four-point support frames, including split-wheel type and wheelless type, which can provide higher stability; wheelchairs can be divided into manual wheelchairs, electric wheelchairs, and sports wheelchairs, and door opening and closing sounds include door opening sounds and door closing sounds. The behavioral audio information data collected during the above-mentioned preset sleep time period and the preset non-sleep time period reflect the behavioral information of the target person in different internal and external time and space states, and are behavioral information corresponding to behavioral habits and behavioral laws. They are generally closely related to and adapted to the vital signs and physiological state, physical state or health state of the target person, and are suitable for physiological state detection of target persons in scenarios such as but not limited to home-based elderly care and rehabilitation monitoring.

[0086] The embodiment of the present invention, by collecting the corresponding current behavior audio information data of the target person in different internal and external time and space state scenarios, can make the collected current behavior audio information data more in line with the internal and external time and space state scenarios of the target person from the perspective of the corresponding behavior sound of the target person, with the help of the target person's behavioral habits and behavioral patterns, and make the current behavior audio information data more adapted to the target person himself, and then indirectly detect whether the target person's consistent physiological state and health state have changed in the corresponding internal and external time and space state through the corresponding current behavior audio information data, which can improve the detection accuracy of whether the target person's consistent physiological state and health state have changed in the corresponding internal and external time and space state, that is, improve the accuracy of the detection of the target person's vital signs data.

[0087] In one embodiment, determining corresponding historical vital sign audio information features and historical behavior audio information features includes:

[0088] Determining historical vital sign audio information features corresponding to the current vital sign audio information features;

[0089] Determine the historical behavior audio information feature corresponding to the current behavior audio information feature.

[0090] Explanatoryally, as described above, determining the corresponding historical vital signs audio information features and historical behavior audio information features includes: determining the historical vital signs audio information features corresponding to the current vital signs audio information features, that is, the current vital signs audio information features and the historical vital signs audio information features have a corresponding and echoing relationship, and are a corresponding set of information features; determining the historical behavior audio information features corresponding to the current behavior audio information features, that is, the current behavior audio information features and the historical behavior audio information features have a corresponding and echoing relationship, and are a corresponding set of information features. For example, when the current vital sign audio information feature represents the respiratory feature of the target person's current respiratory frequency, the corresponding historical vital sign audio information feature is determined to be the respiratory feature corresponding to the target person's past respiratory frequency; when the current behavioral audio information feature represents the sound feature corresponding to the current crutch touching the ground, the corresponding historical behavioral audio information feature is determined to be the sound feature corresponding to the same crutch touching the ground in the past of the target person. In this way, from the perspective of the target person's current audio and the corresponding historical audio, it is possible to detect whether the target person's vital sign data is consistent and stable, so as to detect whether the target person's own consistent physiological state, physical state or health state has changed. This can improve the accuracy of detecting whether the target person's own consistent physiological state and health state have changed, that is, improve the accuracy of detecting the target person's vital sign data.

[0091] The embodiment of the present invention determines the historical vital sign audio information features corresponding to the current vital sign audio information features, and determines the historical behavior audio information features corresponding to the current behavior audio information features, and then detects whether the current physiological state of the target person has changed based on the current vital sign audio information features and the historical vital sign audio information features, as well as the current behavior audio information features and the historical behavior audio information features. Therefore, from the perspective of the target person's vital sign audio information and behavior audio information, the current audio and the corresponding historical audio, it detects whether the vital sign data of the target person is consistent and stable, so as to detect whether the target person's own consistent physiological state, physical state or health state has changed. Since the target person's vital sign audio information and behavior audio information, the current audio and the corresponding historical audio are all personalized sounds adapted to the target person himself, it can improve the accuracy of detecting whether the target person's own consistent physiological state and health state has changed, that is, it improves the accuracy of the target person's vital sign data detection, which helps to improve the intelligence level of smart health care.

[0092] It should be noted that the methods for detecting vital signs data of personnel based on voiceprint recognition described in the above embodiments can recombine the technical features contained in different embodiments as needed to obtain a combined implementation plan, but they are all within the scope of protection required by the present invention.

[0093] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium may be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0094] The relevant data collection appearing in the embodiments of the present invention complies with the requirements of relevant laws and regulations, such as China's "Personal Information Protection Law", GDPR (EU General Data Protection Regulation) or information security standards of other countries and regions.

[0095] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for detecting vital signs of personnel based on voiceprint recognition, characterized in that: include: Collect the target person's current vital signs audio information data and the corresponding current behavior audio information data; Performing voiceprint recognition on the current vital sign audio information data to obtain current vital sign audio information features, and performing voiceprint recognition on the current behavior audio information data to obtain current behavior audio information features; Determine corresponding historical vital sign audio information features and historical behavior audio information features; Detecting whether the current vital sign audio information feature is similar to the historical vital sign audio information feature, and detecting whether the current behavior audio information feature is similar to the historical behavior audio information feature; When each of the above-mentioned groups of information features are similar, it is determined that the current physiological state of the target person is in a stable state.

2. The method for detecting vital signs of a person based on voiceprint recognition according to claim 1, wherein: The method further comprises: When the current vital sign audio information feature is not similar to the historical vital sign audio information feature, or the current behavior audio information feature is not similar to the historical behavior audio information feature, it is determined that the current physiological state of the target person is in an abnormal state.

3. The method for detecting vital signs of a person based on voiceprint recognition according to claim 2, wherein: When the current vital sign audio information feature is not similar to the historical vital sign audio information feature, determining that the current physiological state of the target person is abnormal includes: In a case where the current vital sign audio information feature is not similar to the historical vital sign audio information feature, determining a preset standard vital sign audio information feature; Calculating a characteristic distance between the current vital sign audio information feature and the preset standard vital sign audio information feature to obtain a first characteristic distance; Calculating a characteristic distance between the historical vital sign audio information feature and the preset standard vital sign audio information feature to obtain a second characteristic distance; determining whether the first characteristic distance is greater than the second characteristic distance; When the first characteristic distance is greater than the second characteristic distance, it is determined that the current physiological state of the target person is in an abnormal state.

4. The method for detecting vital signs of a person based on voiceprint recognition according to claim 2, wherein: When the current behavior audio information feature is not similar to the historical behavior audio information feature, determining that the current physiological state of the target person is abnormal includes: In a case where the current behavior audio information feature is not similar to the historical behavior audio information feature, determining a preset standard behavior audio information feature; Calculating a characteristic distance between the current behavior audio information feature and the preset standard behavior audio information feature to obtain a third characteristic distance; Calculating a characteristic distance between the historical behavior audio information feature and the preset standard behavior audio information feature to obtain a fourth characteristic distance; determining whether the third characteristic distance is greater than the fourth characteristic distance; When the third characteristic distance is greater than the fourth characteristic distance, it is determined that the current physiological state of the target person is in an abnormal state.

5. The method for detecting vital signs of a person based on voiceprint recognition according to any one of claims 1 to 4, characterized in that: Collect the target person's current vital signs audio information data and the corresponding current behavior audio information data, including: Determine the internal and external time and space states corresponding to the target person; According to the internal and external space-time states, the current vital sign audio information data of the target person and the corresponding current behavior audio information data are collected.

6. The method for detecting vital signs of a person based on voiceprint recognition according to claim 5, characterized in that: Before collecting the target person's current vital sign audio information data and the corresponding current behavior audio information data based on the internal and external space-time state, the method further includes: Collect initial audio information data of the target person; Identifying the initial audio information data to obtain initial audio information features; Identifying vital sign audio information features contained in the initial audio information features to obtain target vital sign audio information features; Identifying the human behavior audio information features contained in the initial audio information features to obtain target behavior audio information features; The target vital sign audio information features and the target behavior audio information features are formed into a correspondence relationship to obtain the target person's vital sign audio information features and the corresponding behavior audio information features.

7. The method for detecting vital signs of a person based on voiceprint recognition according to claim 6, wherein: Identifying the human behavior audio information features contained in the initial audio information features to obtain target behavior audio information features includes: Determining a number of initial audio information features based on time series; Identify common human behavior audio information features contained in the plurality of initial audio information features to obtain target behavior audio information features.

8. The method for detecting vital signs of a person based on voiceprint recognition according to claim 5, wherein: Determine the target person's corresponding internal and external time and space status, including at least one of the following: Determining the internal space-time state corresponding to the target person; Determine the external space-time state of the target person.

9. The method for detecting vital signs of a person based on voiceprint recognition according to claim 5, wherein: According to the internal and external space-time states, the target person's current vital sign audio information data and the corresponding current behavior audio information data are collected, including at least one of the following: During the preset sleep period, the target person's current behavior audio information data is collected, including at least one of the following: audio information data corresponding to the sound of urinating during sleep interruption, footsteps, and the sound of turning on and off lights; During the preset non-sleep time period, the current behavior audio information data of the target person is collected, including at least one of the following: auxiliary action audio information data corresponding to the target action auxiliary device, footsteps, and door opening and closing sounds.

10. The method for detecting vital signs of a person based on voiceprint recognition according to claim 1, wherein: Determine the corresponding historical vital sign audio information features and historical behavior audio information features, including: Determining historical vital sign audio information features corresponding to the current vital sign audio information features; Determine the historical behavior audio information feature corresponding to the current behavior audio information feature.

Citation Information

Patent Citations

  • Identification method, first-aid decision-making method, medium and intelligent life health monitoring system

    CN115271002A

  • Breathing behavior habit monitoring system based on voice recognition and action analysis

    CN117672526A

  • Cough detection method and system based on artificial intelligence and computer equipment

    CN117711429A

  • Vital sign monitoring method and system, storage medium and equipment

    CN118177730A

  • System and Method for Realtime Examination of the Patient during Telehealth Video Conferencing

    US20240282468A1