An infant milk-sputum AI sub-scene first-aid flow process guidance method and system
Patent Information
- Application Number
- CN202611032737.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]本申请提供一种婴幼儿呛奶AI分场景急救流程指导方法及系统,其主要目的在于解决婴幼儿呛奶个人展开急救效率较低的问题
1.本申请提供的一种婴幼儿呛奶AI分场景急救流程指导方法及系统,通过构建包含意识状态推理层、气道阻塞程度推理层、咳嗽效力推理层和生理风险指标推理层的分层递阶式场景特征推理模型,将婴幼儿呛奶场景识别从传统的单一特征简单分类重新定义为多维度分层递阶推理问题,各推理层基于不同的多模态特征组合进行独立推理,并将上层推理结果作为下层推理的前置约束条件;能够精确区分婴幼儿呛奶的多种细分场景,避免了传统单一阈值分类导致的误判,显著提升了场景分类的精细化程度和准确率,为后续急救流程的精准匹配提供了可靠依据,大幅提高了急救指导的针对性和有效性。
Smart Images

Figure CN122822221A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of scenario-based emergency rescue process guidance technology, and in particular to an AI-based scenario-based emergency rescue process guidance method and system for infants choking on milk. Background Technology
[0002] Infant choking on milk is a common emergency in daily life, which can lead to airway obstruction, hypoxia, and even suffocation and death in severe cases. Currently, first aid guidance for infant choking on milk mainly relies on two methods: online searches or first aid manuals. Users need to independently search and determine the applicable first aid method in emergencies, which suffers from time-consuming searches, difficulty in judgment, and lack of intuitive operation instructions. Existing technology lacks intelligent means to finely categorize infant choking scenarios, failing to differentiate between combinations of different dimensions such as state of consciousness, degree of airway obstruction, and cough effectiveness. This results in first aid guidance that is broad and lacks specificity, easily leading to delays in optimal rescue due to misjudgment or untimely guidance. Therefore, how to achieve fine-grained classification of infant choking scenarios and dynamically adjust the first aid guidance process based on real-time changes in the situation has become an unresolved problem. Summary of the Invention
[0003] This application provides an AI-based, scenario-based emergency response method and system for infants choking on milk, the main purpose of which is to solve the problem of low efficiency in individual emergency response for infants choking on milk.
[0004] To achieve the above objectives, this application provides an AI-based scenario-based emergency response guidance method for infants choking on milk, comprising: S1, acquiring real-time multimodal state data of infants and preprocessing it to obtain preprocessed multimodal data.
[0005] S2. Based on the preprocessed multimodal data, construct each inference layer in the scene feature inference model and extract the inference indicators of each inference layer feature.
[0006] S3. Based on the inference indicators of each inference layer, generate scene feature judgment vectors, analyze and determine the type of choking scene and the severity level of choking, and generate corresponding personalized first aid guidance instruction sequences.
[0007] S4. Real-time acquisition of infant state change data, and performance evaluation based on state change data. In a preferred embodiment, the acquisition of real-time multimodal state data of infants and young children and preprocessing it to obtain preprocessed multimodal data includes: acquiring real-time multimodal state data of infants and young children after a choking incident, wherein the real-time multimodal state data includes audio data, video data and acceleration data, and preprocessing to extract effective audio segments, keyframe image sets and motion feature data segments to obtain preprocessed multimodal data.
[0008] In a preferred embodiment, each reasoning layer includes a consciousness state reasoning layer, an airway obstruction degree reasoning layer, a cough efficacy reasoning layer, and a physiological risk indicator reasoning layer; the reasoning indicators of each reasoning layer include consciousness state determination results, airway obstruction degree determination results, cough efficacy determination results, and physiological risk level indicators.
[0009] In a preferred embodiment, the step of constructing each inference layer in the scene feature inference model based on preprocessed multimodal data and extracting inference indicators of each inference layer features includes: S401, constructing a consciousness state inference layer, extracting eye state features from the keyframe image set, extracting sound response features from the effective audio segments, calculating consciousness state indicators through a consciousness state inference function, and obtaining a consciousness state determination result.
[0010] S402, construct an airway obstruction degree inference layer, select the corresponding feature extraction strategy based on the consciousness state determination result, extract airway state features from the effective audio segments or key frame image set, calculate the airway obstruction degree index through the airway obstruction degree inference function, and obtain the airway obstruction degree determination result.
[0011] S403, when the consciousness state determination result is clear consciousness and the airway obstruction degree determination result is partial obstruction, construct a cough efficacy inference layer, extract cough features from effective audio segments and motion feature data segments, calculate the cough efficacy index through the cough efficacy inference function, and obtain the cough efficacy determination result.
[0012] S404 constructs a physiological risk index inference layer, extracts physiological risk features from keyframe image sets, effective audio segments, and motion feature data segments, and calculates physiological risk level indicators through a comprehensive physiological risk assessment function.
[0013] In a preferred embodiment, the step of calculating the consciousness state index through the consciousness state inference function to obtain the consciousness state determination result includes: ; in, As an indicator of state of consciousness, This represents the normalized value for the degree of eyelid opening and closing. This is the normalized value of eye movement frequency. This is a normalized value for the response intensity to a call. and These are the weights for eye features and the weights for voice features, respectively.
[0014] In a preferred embodiment, the step of selecting the corresponding feature extraction strategy based on the consciousness state determination result includes: S601, when the consciousness state determination result is that the person is conscious, extracting respiratory sound features and wheezing sound features from the effective audio segment; the calculation process of the airway obstruction degree inference function is as follows: ; in, The characteristic value of wheezing. Let be the instantaneous power spectral density of the wheezing sound at time t. Let be the instantaneous power spectral density of the breath sound at time t. To analyze the length of the time window, when the wheezing characteristic value is greater than the preset wheezing threshold, it is determined to be a complete blockage state; otherwise, it is determined to be a partial blockage state.
[0015] S602, when the consciousness state determination result is loss of consciousness, then extract the lip color features and chest rise and fall amplitude features from the keyframe image set; and the calculation process of the airway obstruction degree inference function is as follows: ; in, This is an indicator of abnormal chest wall movement. This represents the maximum value of the current chest cavity rise and fall amplitude in the keyframe image sequence. The baseline value for chest wall movement amplitude during normal breathing is used. When the abnormal chest wall movement index exceeds the preset abnormal chest wall threshold, it is determined to be a complete obstruction state; otherwise, it is determined to be a partial obstruction state.
[0016] In a preferred embodiment, the inference index based on the features of each inference layer generates a scene feature judgment vector, analyzes and determines the type of choking scenario and the severity level of choking, and generates a corresponding personalized first aid guidance instruction sequence, including: S701, vectorizing and combining the consciousness state judgment result, airway obstruction degree judgment result, cough efficacy judgment result, and physiological risk level index to generate a scene feature judgment vector, matching it through a scene classification mapping table, analyzing and determining the type of choking scenario corresponding to the scene feature judgment vector; and determining the level threshold based on the physiological risk level index to obtain the severity level of choking.
[0017] S702, based on the type and severity of the choking scenario, searches and matches in a pre-built emergency procedure protocol library to obtain the corresponding basic emergency procedure protocol. It then makes adaptive adjustments based on real-time collected preprocessed multimodal data to generate a personalized emergency guidance instruction sequence for the current scenario.
[0018] In a preferred embodiment, the real-time acquisition of infant and toddler state change data and the performance evaluation based on the state change data include: S801, after generating a personalized first aid guidance instruction sequence, real-time monitoring of infant and toddler state change data; and calculating a state improvement index based on physiological state characteristics and initial physiological state characteristics before executing the personalized first aid guidance instruction sequence through a state improvement evaluation function.
[0019] S802, compare the state improvement index with the preset improvement threshold to obtain the performance evaluation result.
[0020] S803: If the evaluation result of the execution effect is that the expected improvement effect has not been achieved, S2 is executed again to obtain an updated personalized first aid guidance instruction sequence and switch to it.
[0021] In a preferred embodiment, calculating the state improvement index through the state improvement evaluation function includes: ; in, The index represents the improvement in condition. The initial physiological risk level index value before executing the personalized emergency care guidance sequence. The current physiological risk level index value after executing the personalized emergency medical guidance sequence, among which, It is a very small positive number, with a default value of 0.001.
[0022] To address the aforementioned issues, this application also provides a system for providing AI-based, scenario-specific emergency response guidance for infants choking on milk, the system comprising: The multimodal data processing module is used to acquire real-time multimodal state data of infants and young children and perform preprocessing to obtain preprocessed multimodal data.
[0023] The inference layer construction module constructs each inference layer in the scene feature inference model based on preprocessed multimodal data and extracts the inference indicators of each inference layer feature.
[0024] The personalized first aid guidance instruction sequence generation module generates scene feature judgment vectors based on the inference indicators of each inference layer, analyzes and determines the type and severity of the choking scenario, and generates the corresponding personalized first aid guidance instruction sequence.
[0025] The effectiveness evaluation module is used to collect data on changes in the infants' and toddlers' states in real time and to perform effectiveness evaluations based on this data.
[0026] Compared with the prior art, this application has the following beneficial effects: 1. This application provides an AI-based scenario-based emergency response guidance method and system for infant choking on milk. By constructing a hierarchical scenario feature reasoning model that includes a consciousness state reasoning layer, an airway obstruction degree reasoning layer, a cough efficacy reasoning layer, and a physiological risk indicator reasoning layer, the identification of infant choking on milk scenarios is redefined from the traditional simple classification of single features into a multi-dimensional hierarchical reasoning problem. Each reasoning layer performs independent reasoning based on different combinations of multimodal features, and the reasoning results of the upper layer are used as the preconditions for the reasoning of the lower layer. It can accurately distinguish multiple sub-scenarios of infant choking on milk, avoid misjudgments caused by traditional single threshold classification, significantly improve the refinement and accuracy of scenario classification, provide a reliable basis for the accurate matching of subsequent emergency response procedures, and greatly improve the pertinence and effectiveness of emergency guidance.
[0027] 2. This application constructs an adaptive dynamic emergency rescue process matching and feedback adjustment mechanism. During the execution of the emergency rescue guidance instruction sequence, it monitors the infant's status change data in real time, evaluates the execution effect based on the status improvement index, and automatically returns to the hierarchical reasoning process when the evaluation result shows that the expected improvement effect has not been achieved. Based on the new reasoning result, it switches to the updated emergency rescue process guidance instruction sequence. Through this three-layer feedback mechanism, which includes status monitoring, effect evaluation, scenario re-judgment, and process switching, the adaptive dynamic adjustment of the emergency rescue guidance process during operation is realized. This effectively solves the problem in the prior art where the guidance content is not updated in time due to changes in the infant's status, and ensures the adaptability and safety of the emergency rescue guidance.
[0028] 3. The wheezing feature extraction in this application is applied to acute emergency scenarios of infant choking on milk. The analysis object is the real-time respiratory sound signal after the choking event, and the purpose of the analysis is to determine the degree of airway obstruction in real time to guide emergency operations. Existing lung sound analysis methods are used for the diagnosis and severity grading of chronic respiratory diseases (such as asthma and COPD). Their analysis object is the respiratory sound signal in a resting state, and the purpose of the analysis is long-term disease assessment. Wheezing feature extraction is constrained by the reasoning results of the state of consciousness. Only when the state of consciousness reasoning layer determines that the patient is "conscious" will the power spectral density ratio calculation based on wheezing be performed. When the determination is "loss of consciousness", a completely different assessment path based on abnormal chest wall movement indicators is switched. This dynamic feature selection strategy is a unique design of this application, while existing lung sound analysis methods do not have this hierarchical constraint mechanism. Attached Figure Description
[0029] Figure 1 A flowchart illustrating an AI-based, scenario-specific emergency response method for infant choking on milk, provided as an embodiment of this application; Figure 2 This application provides a functional module diagram of a system for guiding an AI-based emergency response procedure for infant choking on milk, based on different scenarios, according to an embodiment of the present application. Figure 3 This is a schematic diagram illustrating the scenario-based emergency rescue process guidance for infants choking on milk in this application.
[0030] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0031] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0032] This application provides an AI-powered, scenario-based emergency response guidance method for infants choking on milk. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the AI-powered, scenario-based emergency response guidance method for infants choking on milk can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0033] Reference Figure 1 The diagram shown is a flowchart illustrating an AI-powered, scenario-based emergency response method for infant choking on milk, provided in an embodiment of this application. In this embodiment, the AI-powered, scenario-based emergency response method for infant choking on milk includes: S1, acquiring real-time multimodal state data of the infant and preprocessing it to obtain preprocessed multimodal data; In this embodiment of the application, the step of acquiring and preprocessing the real-time multimodal state data of infants and young children to obtain preprocessed multimodal data includes: acquiring the real-time multimodal state data of infants and young children after a choking incident, wherein the real-time multimodal state data includes audio data, video data and acceleration data, and performing preprocessing to extract effective audio segments, key frame image sets and motion feature data segments, thereby obtaining preprocessed multimodal data.
[0034] It should be noted that the audio data is the sound signal of the infant after choking on milk, collected by the microphone, including crying and coughing signals; the video data is the image sequence of the infant's face and torso, collected by the camera, including facial image sequence and limb movement image sequence; and the acceleration data is the acceleration signal of the infant's torso movement, collected by the accelerometer, including torso movement acceleration signal.
[0035] It should be noted that the audio data undergoes noise reduction and endpoint detection to extract valid audio segments; the video data undergoes frame sequence extraction and keyframe filtering to obtain a keyframe image set; and the acceleration data undergoes filtering and segmentation to obtain motion feature data segments.
[0036] It should be noted that the audio data, video data, and acceleration data have different sampling rates during acquisition. Before noise reduction, frame extraction, and filtering, the three data streams need to be timestamped. The specific process of timestamping is as follows: During the data acquisition phase, the main control clock module generates a unified global time reference signal. This global time reference signal is sent to the microphone, camera, and accelerometer at a preset synchronization frequency. When each data frame is acquired, each sensor appends the count value of the current global time reference signal as the hardware timestamp of that data frame to the header metadata of the data frame. Using the frame rate of the video data as the reference time axis, the timestamps of the audio and acceleration data frames are mapped onto the reference time axis using nearest neighbor interpolation, resulting in an initially aligned data frame sequence. Fine-tuning alignment is performed using cough events as cross-modal synchronization anchors. The method for identifying cross-modal synchronization anchors is as follows: the start time of the cough sound is detected from the effective audio segment to obtain the audio event timestamp; the peak time of trunk vibration acceleration within the corresponding time window is detected from the motion feature data segment to obtain the acceleration event timestamp; the time offset between the audio event timestamp and the acceleration event timestamp is calculated, and the time offset of the acceleration data timestamp sequence is corrected as a whole to complete the fine alignment.
[0037] Furthermore, the specific process of time gridding is as follows: using the frame interval of the video data as the step size of the reference time grid, for each reference time grid node, search for the audio data frame and acceleration data frame that are closest in time to that node, and use the data of that frame as the corresponding data of that time grid node; when the same time grid node corresponds to multiple data frames, take the data of the frame with the closest time distance; when the same time grid node has no corresponding data frame, take the two closest data frames before and after and perform linear interpolation to obtain the data of that node.
[0038] It should be noted that the noise reduction process uses an adaptive filtering algorithm to filter out environmental noise in the audio data using a preset noise reference signal; the endpoint detection process is based on a dual threshold decision method of short-time energy and short-time zero-crossing rate to detect and extract effective sound signal segments from the noise-reduced audio signal to obtain effective audio segments.
[0039] It should be noted that frame sequence extraction involves extracting image frames one by one from the video data at a preset frame rate to obtain the original frame sequence. Keyframe selection processing is based on the inter-frame difference method to calculate the pixel change between adjacent frames. When the pixel change exceeds a preset threshold, the frame is identified as a keyframe and retained to obtain a keyframe image set. Further, the specific calculation of the inter-frame change is as follows: First, convert two adjacent frames in the original frame sequence to grayscale images; calculate the absolute difference between the two adjacent grayscale images to obtain a difference image; calculate the sum of the grayscale values of all pixels in the difference image to obtain the inter-frame change.
[0040] It should be noted that the filtering process employs a low-pass filtering algorithm to suppress high-frequency noise components in the acceleration data, resulting in a smoothed acceleration signal. The segmented processing uses a sliding time window to divide the smoothed acceleration signal into multiple time-segment segments, yielding motion signal data segments for motion feature analysis. Furthermore, the low-pass filtering algorithm uses a Butterworth low-pass filter with a cutoff frequency set to 5 Hz to filter out high-frequency interference signals caused by infant limb tremors. The sliding time window length is set to 2 seconds, and the window sliding step size is set to 0.5 seconds to ensure continuous monitoring of trunk movements during choking incidents.
[0041] S2. Based on the preprocessed multimodal data, construct each inference layer in the scene feature inference model and extract the inference indicators of each inference layer feature.
[0042] In this embodiment of the application, each reasoning layer includes a consciousness state reasoning layer, an airway obstruction degree reasoning layer, a cough efficacy reasoning layer, and a physiological risk indicator reasoning layer; the reasoning indicators of each reasoning layer feature include consciousness state determination results, airway obstruction degree determination results, cough efficacy determination results, and physiological risk level indicators.
[0043] In this embodiment of the application, the step of constructing each inference layer in the scene feature inference model based on preprocessed multimodal data and extracting the inference index of each inference layer features includes: S401, constructing a consciousness state inference layer, extracting eye state features from the key frame image set, extracting sound response features from the effective audio segments, calculating consciousness state index through the consciousness state inference function, and obtaining the consciousness state determination result.
[0044] S402, construct an airway obstruction degree inference layer, select the corresponding feature extraction strategy based on the consciousness state determination result, extract airway state features from the effective audio segments or key frame image set, calculate the airway obstruction degree index through the airway obstruction degree inference function, and obtain the airway obstruction degree determination result.
[0045] S403, when the consciousness state determination result is clear consciousness and the airway obstruction degree determination result is partial obstruction, construct a cough efficacy inference layer, extract cough features from effective audio segments and motion feature data segments, calculate the cough efficacy index through the cough efficacy inference function, and obtain the cough efficacy determination result.
[0046] S404 constructs a physiological risk index inference layer, extracts physiological risk features from keyframe image sets, effective audio segments, and motion feature data segments, and calculates physiological risk level indicators through a comprehensive physiological risk assessment function.
[0047] It should be noted that cough characteristics include the temporal characteristics of the sound pressure level of the cough sound, the characteristics of the cough interval duration, and the characteristics of the trunk vibration amplitude during coughing; the calculation process of the cough efficacy inference function is as follows: ; in, It is a cough efficacy index used to determine whether a cough is an effective cough or an ineffective cough; The sound pressure level time-series function of the cough sound is obtained by converting the sound pressure level of the cough sound signal in the effective audio segment; The baseline cough sound pressure level squared over time mean is obtained by statistically averaging the sound pressure level squared over time mean of a large number of valid cough samples; The duration of a single cough is determined by detecting the start and end times of the cough sound signal in the effective audio segment; The peak amplitude of trunk vibration acceleration during coughing is obtained by peak detection of the acceleration signal in the motion feature data segment corresponding to the time period of the cough sound; The baseline amplitude of trunk vibration acceleration during normal crying is obtained from pre-stored standard crying vibration amplitude data; and These are the contribution weights for sound pressure level and trunk vibration, respectively. The default values for the contribution weights for sound pressure level and trunk vibration are 0.7 and 0.3, respectively. When the cough efficacy index is greater than the preset effective cough threshold, it is determined to be a valid cough; otherwise, it is determined to be an invalid cough.
[0048] Furthermore, the effective cough threshold and consciousness threshold are defined in the following table:
[0049] It should be noted that physiological risk characteristics include facial skin color characteristics, respiratory rate and rhythm characteristics, and trunk tension characteristics; the calculation process of the comprehensive physiological risk assessment function is as follows: ; in, It is a physiological risk level indicator used to quantitatively assess the current physiological risk level of infants and young children; The index for the degree of cyanosis of the lips and face is calculated by converting the images of the lips and face regions in the keyframe images from the RGB color space to the Lab color space, and then extracting the color difference feature values of the a and b channels. The value range is [0,1], where 0 indicates no cyanosis and 1 indicates severe cyanosis. This is an indicator of respiratory distress, with a value range of [0,1], where 0 indicates no respiratory distress and 1 indicates severe respiratory distress. This is an indicator of abnormal muscle tone, calculated by analyzing the proportion of low-frequency, high-amplitude components in the motion characteristic data segment. The value range is [0,1], where 0 indicates normal muscle tone and 1 indicates severe abnormal muscle tone. , , These are the weights for cyanosis severity, respiratory distress, and abnormal muscle tone, respectively. The default values for these weights are 0.4, 0.4, and 0.2, respectively.
[0050] Furthermore, the calculation process of the respiratory distress index is as follows: Respiratory audio segments are extracted from the effective audio clips; the number of breaths within a preset time window is counted using a peak detection algorithm to obtain the measured respiratory rate; the difference between the measured respiratory rate and the normal respiratory rate benchmark for the corresponding age group is calculated to obtain the respiratory rate deviation; simultaneously, the standard deviation between adjacent respiratory cycles is calculated to obtain the respiratory rhythm variability; finally, the respiratory rate deviation and respiratory rhythm variability are weighted and summed to obtain the respiratory distress index.
[0051] In this embodiment of the application, the step of calculating the consciousness state index through the consciousness state inference function to obtain the consciousness state determination result includes: ; in, As an indicator of state of consciousness, This represents the normalized value for the degree of eyelid opening and closing. This is the normalized value of eye movement frequency. This is a normalized value for the response intensity to a call. and These are the weights for eye features and the weights for voice features, respectively.
[0052] It should be noted that the consciousness status index is used to determine whether an infant or young child is in a state of consciousness or loss of consciousness; the normalized value of eyelid opening and closing is in the range of [0,1], where 0 indicates that the eyelids are completely closed and 1 indicates that the eyelids are completely open; the normalized value of eye movement frequency is in the range of [0,1], where 0 indicates that the eyes do not move and 1 indicates that the eye movement frequency reaches a normal level; the normalized value of the response intensity of the call is in the range of [0,1], where 0 indicates that there is no response and 1 indicates that the response intensity reaches a normal level; the default values for the weight of eye features and the weight of voice features are 0.6 and 0.4, respectively.
[0053] Furthermore, when the consciousness state index is greater than the preset consciousness awakening threshold, it is determined to be a state of consciousness awakening; otherwise, it is determined to be a state of loss of consciousness. The preset consciousness awakening threshold is an empirical value obtained based on clinical data statistics and experimental calibration. The specific process is as follows: collect a large number of multimodal sample data of infants of different ages in both awake and loss of consciousness states, extract the normalized values of eyelid opening and closing degree, eye movement frequency, and response intensity, and construct positive and negative sample datasets; substitute the three feature values from the positive and negative sample datasets into the consciousness state inference function for calculation to obtain the consciousness state index values of all samples. By analyzing the receiver operating characteristic (ROC) curves, the optimal cutoff point was determined with the goal of maximizing the Youden index (the sum of sensitivity and specificity minus 1). Finally, the optimal cutoff point was used as the default value of the preset consciousness threshold. In practice, this threshold can be fine-tuned according to the different age groups of infants. For example, the threshold value is 0.45 for infants aged 0 to 6 months, 0.50 for infants aged 6 to 12 months, and 0.55 for infants aged 12 to 24 months, to adapt to the differences in neurodevelopment of infants of different ages.
[0054] It should be noted that the eye state features include the degree of eyelid opening and closing and the frequency of eye movements, while the sound response features include the intensity of the response to a call. The normalized value of the eyelid opening and closing degree is obtained by performing edge detection on the eye region images in the keyframe image set, extracting the distance between the upper and lower eyelid edges, and taking the ratio of this distance to the maximum distance when the eyes are normally open. The normalized value of the eye movement frequency is obtained by tracking the eye position in multiple consecutive frames in the keyframe image sequence, counting the number of eye movements per unit time, and taking the ratio of the number of eye movements to the baseline number of movements in the normal state. The normalized value of the response intensity to a call is obtained by detecting the energy of the sound signal within a preset time window after the call is emitted in the effective audio segment, and taking the ratio of the detected response sound energy to the baseline energy of the normal response.
[0055] In this embodiment of the application, the step of selecting the corresponding feature extraction strategy based on the consciousness state determination result includes: S601, when the consciousness state determination result is that the person is conscious, extracting respiratory sound features and wheezing sound features from the effective audio segment; the calculation process of the airway obstruction degree inference function is as follows: ; in, The characteristic value of wheezing. Let be the instantaneous power spectral density of the wheezing sound at time t. Let be the instantaneous power spectral density of the breath sound at time t. To analyze the length of the time window, when the wheezing characteristic value is greater than the preset wheezing threshold, it is determined to be a complete blockage state; otherwise, it is determined to be a partial blockage state.
[0056] S602, when the consciousness state determination result is loss of consciousness, then extract the lip color features and chest rise and fall amplitude features from the keyframe image set; and the calculation process of the airway obstruction degree inference function is as follows: ; in, This is an indicator of abnormal chest wall movement. This represents the maximum value of the current chest cavity rise and fall amplitude in the keyframe image sequence. The baseline value for chest wall movement amplitude during normal breathing is used. When the abnormal chest wall movement index exceeds the preset abnormal chest wall threshold, it is determined to be a complete obstruction state; otherwise, it is determined to be a partial obstruction state.
[0057] It should be noted that the wheezing sound feature value is used to determine whether the airway is partially or completely obstructed; the instantaneous power spectral density of the wheezing sound is obtained by performing a short-time Fourier transform on the wheezing audio segment in the effective audio segment; the instantaneous power spectral density of the breathing sound is obtained by performing a short-time Fourier transform on the breathing audio segment in the effective audio segment; the default value of the analysis time window is 3 seconds. The abnormal chest wall movement index is used to determine whether the airway is partially or completely obstructed; the maximum value of the current chest wall fluctuation amplitude in the keyframe image sequence is obtained by measuring the change in the longitudinal pixel distance of the chest wall region in the keyframe image set; the baseline value of the chest wall fluctuation amplitude during normal breathing is obtained from pre-stored standard chest wall fluctuation amplitude data.
[0058] Furthermore, the values for the wheezing threshold and the chest wall abnormality threshold are shown in the table below:
[0059] S3. Based on the inference indicators of each inference layer, generate scene feature judgment vectors, analyze and determine the type of choking scene and the severity level of choking, and generate corresponding personalized first aid guidance instruction sequences.
[0060] In this embodiment of the application, the process of generating a scene feature judgment vector based on the inference indicators of each inference layer features, analyzing and determining the type and severity level of the choking scenario, and generating a corresponding personalized first aid guidance instruction sequence includes: S701, combining the consciousness state judgment result, airway obstruction degree judgment result, cough efficacy judgment result, and physiological risk level indicator into a vectorized combination to generate a scene feature judgment vector, matching it through a scene classification mapping table, analyzing and determining the type of choking scenario corresponding to the scene feature judgment vector; determining the level threshold based on the physiological risk level indicator, and determining the severity level of the choking.
[0061] S702, based on the type and severity of the choking scenario, searches and matches in a pre-built emergency procedure protocol library to obtain the corresponding basic emergency procedure protocol. It then makes adaptive adjustments based on real-time collected preprocessed multimodal data to generate a personalized emergency guidance instruction sequence for the current scenario.
[0062] It should be noted that the scene feature determination vector is a four-dimensional vector. Its first dimension is obtained by numerically encoding the consciousness state determination result; when the consciousness state determination result is "conscious," the encoded value is 1. The second dimension is obtained by numerically encoding the airway obstruction degree determination result; when the airway obstruction degree determination result is "partial obstruction," the encoded value is 1; when the airway obstruction degree determination result is "complete obstruction," the encoded value is 0. The third dimension is obtained by numerically encoding the cough efficacy determination result; when the cough efficacy determination result is "effective cough," the encoded value is 1; when the cough efficacy determination result is "ineffective cough," the encoded value is 0. The fourth dimension is directly assigned from the physiological risk level index. Furthermore, when the consciousness state determination result is "loss of consciousness" or the airway obstruction degree determination result is "complete obstruction," the encoded value of the cough efficacy determination result is set to null, indicating that this dimension does not participate in the scene type matching calculation.
[0063] It should be noted that the scene classification mapping table is a pre-built mapping relationship database, which stores a variety of preset choking milk scene types and the standard scene feature template vector corresponding to each scene type. The specific matching process of the scene classification mapping table is as follows: calculate the dimensional distance between the scene feature judgment vector and each standard scene feature template vector; and select the choking milk scene type corresponding to the standard scene feature template vector with the smallest dimensional distance to the scene feature judgment vector as the matched choking milk scene type.
[0064] Furthermore, the standard scene feature template vectors in the scene classification mapping table include, but are not limited to, the templates corresponding to the scene types in the following tables:
[0065] Among them, a value of 0 for the fourth dimension component indicates that the physiological risk level indicator does not participate in the matching calculation of the scenario type, but is only used for the determination of the severity level.
[0066] Furthermore, the dimensional distance is calculated as follows: when the scene feature determination vector and the standard scene feature template vector are both valid values in the corresponding dimension, the square of the difference in that dimension is calculated; when there is a null dimension in the standard scene feature template vector, that dimension is skipped and does not participate in the distance calculation; the square root of the sum of the squares of the differences in all valid dimensions is taken to obtain the dimensional distance.
[0067] It should be noted that the specific process for determining the severity level threshold is as follows: A first severity level threshold and a second severity level threshold are preset, with the first severity level threshold being less than the second severity level threshold. The physiological risk level index is compared with both the first and second severity level thresholds. When the physiological risk level index is less than the first severity level threshold, it is determined to be a mild case of choking on milk. When the physiological risk level index is greater than or equal to the first severity level threshold and less than the second severity level threshold, it is determined to be a moderate case of choking on milk. When the physiological risk level index is greater than or equal to the second severity level threshold, it is determined to be a severe case of choking on milk.
[0068] Furthermore, the default value for the first severity level threshold is 0.3, and the default value for the second severity level threshold is 0.7. The mild choking level indicates that the infant's physiological indicators are basically normal and there is no significant cyanosis or respiratory distress. The moderate choking level indicates that the infant has a certain degree of cyanosis or respiratory distress and needs to take immediate emergency measures. The severe choking level indicates that the infant has severe cyanosis, respiratory distress or abnormal muscle tone, which is a critical condition and requires emergency medical intervention.
[0069] It should be noted that the emergency procedure protocol library is a pre-built relational database that stores standard emergency procedure protocols corresponding to various milk choking scenarios and milk choking severity levels. Each standard emergency procedure protocol contains multiple emergency steps arranged in execution order. Each emergency step includes fields such as operation video, operation type, operation site, operation force range, operation frequency range, and operation duration range. The retrieval and matching adopts an exact matching strategy, using the milk choking scenario type as the primary key and the milk choking severity level as the secondary key to perform a joint retrieval and obtain the basic emergency procedure protocol that completely corresponds to the current scenario.
[0070] It should be noted that adaptive adjustment is based on the current state information in the preprocessed multimodal data collected in real time. It dynamically corrects the step parameters in the basic emergency rescue process protocol and generates a personalized guidance instruction sequence that is adapted to the actual state of the infant.
[0071] Furthermore, the specific process of adaptive adjustment is as follows: Extract the infant's current position information from the keyframe image set, including the infant's orientation angle and tilt angle; extract the current respiratory sound energy value and cough sound feature value from the effective audio segments; then, based on the infant's current position information, perform spatial coordinate mapping correction on the operation site description in the basic first aid procedure protocol; based on the respiratory sound energy value and cough sound feature value, dynamically scale and adjust the operation force range in the basic first aid procedure protocol; finally, fill the corrected operation parameters into the templates of each step of the basic first aid procedure protocol to generate a personalized first aid guidance instruction sequence.
[0072] Furthermore, the dynamic scaling adjustment of the operational force range is achieved as follows: preset a baseline respiratory sound energy value and a baseline cough sound characteristic value; calculate the ratio of the real-time extracted respiratory sound energy value to the baseline respiratory sound energy value to obtain the respiratory sound intensity ratio; calculate the ratio of the real-time extracted cough sound characteristic value to the baseline cough sound characteristic value to obtain the cough sound intensity ratio; take the weighted average of the respiratory sound intensity ratio and the cough sound intensity ratio as the force adjustment coefficient; multiply the upper and lower limits of the operational force range in the basic emergency procedure protocol by the force adjustment coefficient to obtain the adjusted operational force range.
[0073] It should be noted that the personalized first aid guidance instruction sequence contains multiple guidance steps arranged in chronological order. Each guidance step instruction includes a step number, operation type, operation site description, operation force value, operation frequency value, and expected effect description field. The personalized first aid guidance instruction sequence is output in a structured data format to drive the voice broadcast module and visual display module of the first aid guidance terminal for synchronous guidance.
[0074] S4. Collect data on changes in the state of infants and young children in real time, and evaluate the performance based on the data on changes in state.
[0075] In this embodiment of the application, the real-time acquisition of infant and toddler state change data and the performance evaluation based on the state change data include: S801, after generating a personalized first aid guidance instruction sequence, real-time monitoring of infant and toddler state change data; and calculating a state improvement index based on physiological state characteristics and initial physiological state characteristics before executing the personalized first aid guidance instruction sequence through a state improvement evaluation function.
[0076] S802, compare the state improvement index with the preset improvement threshold to obtain the performance evaluation result.
[0077] S803: If the evaluation result of the execution effect is that the expected improvement effect has not been achieved, S2 is executed again to obtain an updated personalized first aid guidance instruction sequence and switch to it.
[0078] It should be noted that the current physiological state characteristics are physiological state parameters extracted from multimodal data collected and preprocessed in real time at the current moment. The initial physiological state characteristics are the initial cyanosis level indicators, initial respiratory distress level indicators, and initial muscle tone abnormality indicators extracted from the multimodal data at the moment the personalized emergency care guidance instruction sequence is generated.
[0079] In this embodiment of the application, the step of calculating the state improvement index through the state improvement evaluation function includes: ; in, The index represents the improvement in condition. The initial physiological risk level index value before executing the personalized emergency care guidance sequence. The current physiological risk level index value after executing the personalized emergency medical guidance sequence, among which, It is a very small positive number, with a default value of 0.001.
[0080] It should be noted that when and hour, , indicating no change. When and When: the numerator is negative and the denominator is negative. , A negative value indicates a deterioration in the state, which is reasonable. When When very small: the denominator contains Item, not because Too small a value can lead to an abnormally large ratio. The range of values is limited to Within the range, the indicator has clearly defined upper and lower bounds. When When: the molecule is positive. A positive indicates an improvement in the situation; conversely, a negative indicates a deterioration.
[0081] It should be noted that the Condition Improvement Index is used to quantitatively assess the degree of improvement in the infant's condition after following the current first aid instructions. The Condition Improvement Index ranges from negative infinity to 1. When the Condition Improvement Index is positive, it indicates that the infant's physiological risk has decreased and their condition has improved. When the Condition Improvement Index is negative, it indicates that the infant's physiological risk has increased and their condition has worsened.
[0082] Furthermore, the preset improvement threshold is a pre-set constant with a default value of 0.1. When the status improvement index is greater than or equal to the improvement threshold, the execution effect evaluation result is determined to have achieved the expected improvement effect, and the subsequent steps in the current personalized emergency care instruction sequence are continued. When the status improvement index is less than the improvement threshold, the execution effect evaluation result is determined to have not achieved the expected improvement effect, and execution step S803 is triggered.
[0083] It should be noted that the switching process is implemented through an interrupt mechanism, and the specific process is as follows: A switching trigger signal is generated, which contains a pointer to the storage address of the updated personalized first aid guidance instruction sequence; the switching trigger signal is sent to the instruction execution engine of the first aid guidance terminal; after receiving the switching trigger signal, the instruction execution engine immediately pauses the currently executing personalized first aid guidance instruction sequence, reads the updated personalized first aid guidance instruction sequence pointed to by the storage address pointer, and starts execution from the first guidance step of the updated personalized first aid guidance instruction sequence.
[0084] like Figure 2 The diagram shown is a functional block diagram of a system for providing AI-based emergency rescue guidance for infants choking on milk according to an embodiment of this application.
[0085] The system 100 of the AI-based scenario-based emergency response guidance method for infant choking on milk described in this application can be installed in an electronic device. Depending on the functions implemented, the system 100 may include a multimodal data processing module 101, an inference layer construction module 102, a personalized emergency guidance instruction sequence generation module 103, and an effect evaluation module 104. The modules described in this application can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, and are stored in the memory of the electronic device.
[0086] In this embodiment, the data flow relationships between the modules are as follows: the audio, video, and acceleration acquisition modules transmit raw data to the preprocessing module; the preprocessing module processes the data to generate valid audio segments, keyframe image sets, and motion feature data segments, which are then transmitted to the four inference modules. The consciousness state inference module outputs the judgment result to the airway obstruction degree inference module and the scenario determination module; the airway obstruction degree inference module outputs the result to the cough efficacy inference module and the scenario determination module; the cough efficacy inference module outputs the result to the scenario determination module; and the physiological risk inference module outputs the level index to the scenario determination module. The scenario determination module transmits the determined scenario type and severity level to the instruction generation module, which combines the emergency procedure protocol library to generate a personalized emergency guidance instruction sequence and sends it to the monitoring and evaluation module. The effect evaluation module judges the execution effect; if the expected result is not achieved, a feedback re-inference instruction is triggered, returning to the four inference modules for re-inference, forming an adjustment mechanism.
[0087] In this embodiment, the functions of each module / unit are as follows: The multimodal data processing module is used to acquire real-time multimodal state data of infants and young children and perform preprocessing to obtain preprocessed multimodal data.
[0088] The inference layer construction module constructs each inference layer in the scene feature inference model based on preprocessed multimodal data and extracts the inference indicators of each inference layer feature.
[0089] The personalized first aid guidance instruction sequence generation module generates scene feature judgment vectors based on the inference indicators of each inference layer, analyzes and determines the type and severity of the choking scenario, and generates the corresponding personalized first aid guidance instruction sequence.
[0090] The effect evaluation module is used to collect data on changes in the state of infants and young children in real time, and to perform effect evaluation based on the data on changes in state.
[0091] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0092] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0093] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0094] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application.
[0095] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of this application without departing from the spirit and scope of the technical solutions of this application.
Claims
1. A method for providing AI-powered, scenario-based emergency response guidance for infants choking on milk, characterized in that... The method includes: S1. Acquire real-time multimodal state data of infants and young children and preprocess it to obtain preprocessed multimodal data; S2. Based on the preprocessed multimodal data, construct each inference layer in the scene feature inference model and extract the inference indicators of each inference layer feature. S3. Based on the inference indicators of each inference layer, generate scene feature judgment vectors, analyze and determine the type of choking scene and the severity level of choking, and generate corresponding personalized first aid guidance instruction sequences. S4. Collect data on changes in the state of infants and young children in real time, and evaluate the performance based on the data on changes in state.
2. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 1, characterized in that, The process of acquiring and preprocessing real-time multimodal state data of infants and young children to obtain preprocessed multimodal data includes: Real-time multimodal state data of infants and young children after a choking incident is obtained. The real-time multimodal state data includes audio data, video data and acceleration data. Preprocessing is performed to extract effective audio segments, key frame image sets and motion feature data segments to obtain preprocessed multimodal data.
3. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 1, characterized in that, Each reasoning layer includes a consciousness state reasoning layer, an airway obstruction degree reasoning layer, a cough efficacy reasoning layer, and a physiological risk indicator reasoning layer; the reasoning indicators of each reasoning layer include the consciousness state determination result, the airway obstruction degree determination result, the cough efficacy determination result, and the physiological risk level indicator.
4. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 1, characterized in that, Based on the preprocessed multimodal data, each inference layer in the scene feature inference model is constructed, and inference metrics of each inference layer are extracted, including: S401, Construct a consciousness state inference layer, extract eye state features from the keyframe image set, extract sound response features from the effective audio segments, calculate consciousness state indicators through the consciousness state inference function, and obtain the consciousness state determination result. S402, construct an airway obstruction degree inference layer, select the corresponding feature extraction strategy based on the consciousness state judgment result, extract airway state features from the effective audio segments or key frame image set, calculate the airway obstruction degree index through the airway obstruction degree inference function, and obtain the airway obstruction degree judgment result. S403, when the consciousness state determination result is that the consciousness is clear and the airway obstruction degree determination result is that the obstruction is partial, a cough efficacy inference layer is constructed, cough features are extracted from the effective audio segments and motion feature data segments, and the cough efficacy index is calculated through the cough efficacy inference function to obtain the cough efficacy determination result. S404 constructs a physiological risk index inference layer, extracts physiological risk features from keyframe image sets, effective audio segments, and motion feature data segments, and calculates physiological risk level indicators through a comprehensive physiological risk assessment function.
5. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 4, characterized in that, The process of calculating the consciousness state index through the consciousness state inference function to obtain the consciousness state determination result includes: ; in, As an indicator of state of consciousness, This represents the normalized value for the degree of eyelid opening and closing. This is the normalized value of eye movement frequency. This is a normalized value for the response intensity to a call. and These are the weights for eye features and the weights for voice features, respectively.
6. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 4, characterized in that, The feature extraction strategy selected based on the consciousness state determination result includes: S601, when the consciousness status determination result is that the person is conscious, extract the respiratory sound features and wheezing sound features from the effective audio segment; the calculation process of the airway obstruction degree inference function is as follows: ; in, The characteristic value of wheezing. Let be the instantaneous power spectral density of the wheezing sound at time t. Let be the instantaneous power spectral density of the breath sound at time t. To analyze the length of the time window; when the wheezing characteristic value is greater than the preset wheezing threshold, it is determined to be a complete obstruction state; otherwise, it is determined to be a partial obstruction state. S602, when the consciousness state determination result is loss of consciousness, then extract the lip color features and chest rise and fall amplitude features from the keyframe image set; and the calculation process of the airway obstruction degree inference function is as follows: ; in, This is an indicator of abnormal chest wall movement. This represents the maximum value of the current chest cavity rise and fall amplitude in the keyframe image sequence. The baseline value for chest wall movement amplitude during normal breathing is used. When the abnormal chest wall movement index exceeds the preset abnormal chest wall threshold, it is determined to be a complete obstruction state; otherwise, it is determined to be a partial obstruction state.
7. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 1, characterized in that, The inference index based on the features of each inference layer generates a scene feature judgment vector, analyzes and determines the type and severity of the choking scene, and generates a corresponding personalized first aid guidance instruction sequence, including: S701, the results of consciousness status assessment, airway obstruction degree assessment, cough efficacy assessment, and physiological risk level indicators are vectorized and combined to generate a scene feature assessment vector. The scene is matched with a scene classification mapping table to analyze and determine the type of choking scenario corresponding to the scene feature assessment vector. The level threshold is judged based on the physiological risk level indicators to determine the severity level of choking. S702, based on the type and severity of the choking scenario, searches and matches in a pre-built emergency procedure protocol library to obtain the corresponding basic emergency procedure protocol. It then makes adaptive adjustments based on real-time collected preprocessed multimodal data to generate a personalized emergency guidance instruction sequence for the current scenario.
8. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 1, characterized in that, The real-time collection of infant and toddler state change data, and the evaluation of performance based on the state change data, include: S801, after generating a personalized first aid guidance instruction sequence, monitors the infant's status change data in real time; based on physiological status characteristics and initial physiological status characteristics before executing the personalized first aid guidance instruction sequence, it calculates the status improvement index through a status improvement assessment function; S802, compare the state improvement index with the preset improvement threshold to obtain the performance evaluation result; S803: If the evaluation result of the execution effect is that the expected improvement effect has not been achieved, S2 is re-executed to obtain an updated personalized first aid guidance instruction sequence and switch to it.
9. The AI-based scenario-based emergency response guidance method for infant choking on milk as described in claim 8, characterized in that, The calculation of the state improvement index through the state improvement evaluation function includes: ; in, The index represents the improvement in condition. The initial physiological risk level index value before executing the personalized emergency care guidance sequence. The current physiological risk level index value after executing the personalized emergency medical guidance sequence, among which, It is a very small positive number, with a default value of 0.
001.
10. A system for implementing the AI-based scenario-based emergency rescue procedure guidance method for infant choking on milk as described in any one of claims 1-9, characterized in that, include: The multimodal data processing module is used to acquire real-time multimodal state data of infants and young children and perform preprocessing to obtain preprocessed multimodal data. The inference layer construction module constructs each inference layer in the scene feature inference model based on preprocessed multimodal data and extracts the inference indicators of each inference layer feature. The personalized first aid guidance instruction sequence generation module generates scene feature judgment vectors based on the inference indicators of each inference layer, analyzes and derives the choking scene type and choking severity level, and generates the corresponding personalized first aid guidance instruction sequence. The effectiveness evaluation module is used to collect data on changes in the infants' and toddlers' states in real time and to perform effectiveness evaluations based on this data.