Pre-hospital stroke face recognition method and system based on hybrid deep learning model

By constructing a personalized baseline response model using a hybrid deep learning model, the risk of stroke in a dynamic process is quantified, which solves the problem of low sensitivity in early stroke identification in existing technologies and achieves efficient and accurate identification of early stroke symptoms.

CN121661689APending Publication Date: 2026-03-13THE AFFILIATED SIR RUN RUN SHAW HOSPITAL OF SCHOOL OF MEDICINE ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing pre-hospital stroke identification methods rely on subjective observation and static snapshots, making it difficult to identify early or mild stroke symptoms. They also lack personalized baseline comparison and dynamic process analysis, resulting in low identification sensitivity.

Method used

A personalized baseline response model is constructed using a hybrid deep learning model. Through multimodal data analysis, the latent functional failure potential (LFFP) and failure potential accumulation rate (FPAR) are calculated to quantify the degree of neural functional instability and the rate of deterioration in the dynamic process.

Benefits of technology

It improves the sensitivity and objectivity of identifying early, subtle stroke symptoms, and can provide high-risk warnings before functional failure, reducing misjudgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661689A_ABST
    Figure CN121661689A_ABST
Patent Text Reader

Abstract

The invention discloses a pre-hospital stroke face recognition method and system based on a hybrid deep learning model, and belongs to the field of medical artificial intelligence. The method comprises the following steps: processing multi-modal baseline data when a user executes a standardized challenge in a healthy state by adopting a hybrid deep learning model so as to construct a personalized baseline response model; when the suspected stroke occurs, processing multi-modal instantaneous data when the user executes the same challenge by adopting a hybrid deep learning model to generate an instantaneous response trajectory; the hybrid deep learning model calculates a potential function failure potential representing a neural function instability degree based on a dynamic difference between the transient response trajectory and the baseline response model, and determines a time change trend of the potential function failure potential to obtain a failure potential accumulation rate; and finally, the model generates a face recognition result representing the stroke risk level based on the potential function failure potential and the failure potential accumulation rate. According to the invention, early and fine stroke-related face and speech abnormalities can be recognized with high sensitivity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical artificial intelligence technology, and in particular to a method and system for pre-hospital stroke facial recognition based on a hybrid deep learning model. Background Technology

[0002] Stroke, commonly known as "cerebrovascular accident," is an acute cerebrovascular disease characterized by high incidence, high disability rate, and high mortality. Based on different pathologies, stroke is mainly divided into ischemic stroke and hemorrhagic stroke, with ischemic stroke accounting for the vast majority. For patients with ischemic stroke, the first few hours after onset are the "golden window" for receiving reperfusion therapy such as intravenous thrombolysis. Timely and effective treatment can significantly improve patient prognosis and reduce mortality and disability. Therefore, the ability to quickly and accurately identify stroke symptoms in the pre-hospital environment and prompt patients to seek medical attention as soon as possible is a crucial aspect of improving stroke treatment outcomes.

[0003] Currently, the most widely used pre-hospital stroke rapid identification tool internationally is the "FAST" scale, which stands for Face, Arms, Speech, and Time. This scale assesses the likelihood of stroke by observing typical symptoms such as facial asymmetry (e.g., drooping corner of the mouth), unilateral arm weakness, and slurred speech. Emergency medical systems in many countries and regions use the FAST scale as a standard screening procedure.

[0004] However, existing technologies still have some inherent limitations in practical applications: First, it relies on subjective observation and obvious symptoms. The FAST scale primarily depends on the visual observation and subjective judgment of non-professionals (such as family members or bystanders), and its accuracy is affected by the observer's experience and attentiveness. More importantly, this method is mainly for those who have developed relatively obvious neurological deficits, such as significant facial drooping or severe slurred speech. For subtle symptoms in early stroke or minor stroke (TIA), such as mild muscle weakness, decreased motor coordination, or slowed reaction, traditional methods often fail to detect them, potentially leading to missed opportunities for optimal treatment.

[0005] Second, the assessment method is a static snapshot, lacking dynamic process analysis. Existing methods typically assess a static state; for example, asking the patient to smile and maintain it, then observing whether their face is symmetrical. This "static snapshot" approach fails to capture the dynamic information during the execution of the action. However, many early neurological functional impairments are precisely manifested in the dynamic process. For example, a patient may eventually be able to complete a smiling action, but their speed in completing the action is slower than usual, the trajectory is unstable, or the coordination of muscle exertion on both sides is poor. This dynamic, process-oriented information is crucial for early diagnosis but is overlooked by existing methods.

[0006] Third, there is a lack of personalized baseline comparison. Existing methods typically compare a patient's presentation to a universal "normal" standard. However, individual differences exist; some people may be born with mild facial asymmetry. In such cases, comparing a patient's current state to a general standard can lead to misdiagnosis. The ideal assessment method is to compare the patient's current state with their state when they were healthy—a "self-before-and-after comparison"—to most sensitively detect new changes caused by pathological events.

[0007] Although some mobile applications attempt to digitize FAST assessments, such as analyzing facial symmetry by taking photos with a phone camera, most of these applications still focus on analyzing static images and do not fundamentally overcome the second and third core limitations mentioned above. Summary of the Invention

[0008] The purpose of this application is to provide a pre-hospital stroke facial recognition method and system based on a hybrid deep learning model, aiming to improve the technical problem that existing pre-hospital stroke recognition methods are unable to achieve objective, quantitative, and dynamic process analysis of stroke symptoms based on personalized baselines, resulting in low sensitivity in the recognition of early and subtle symptoms.

[0009] To achieve the above objectives, in a first aspect, this application provides a pre-hospital stroke facial recognition method based on a hybrid deep learning model. The method includes: acquiring a user's multimodal baseline data and processing the multimodal baseline data using the hybrid deep learning model to construct a baseline response model representing the user's health status; acquiring the user's multimodal instantaneous data and processing the multimodal instantaneous data using the hybrid deep learning model to generate an instantaneous response trajectory; using the hybrid deep learning model to determine the dynamic difference between the instantaneous response trajectory and the baseline response model to obtain the potential functional failure potential (LFFP); determining the temporal variation trend of the LFFP to obtain the failure potential accumulation rate (FPAR); and generating a stroke facial recognition result for the user based on the LFFP and the FPAR, wherein the recognition result represents the stroke risk level. This application moves beyond simple static image classification by constructing a personalized dynamic baseline model and calculating the two novel dynamic risk indicators, LFFP and FPAR, elevating the facial recognition task from "morphological recognition" to the level of "functional stability recognition." This method can quantify and capture dynamic process anomalies that are undetectable by traditional methods, thereby greatly improving the sensitivity and objectivity of identifying early, subtle stroke symptoms.

[0010] In one possible implementation of the first aspect, the hybrid deep learning model includes a spatiotemporal feature extraction module for extracting spatiotemporal features and a sequence analysis module for performing sequence modeling on the spatiotemporal features; the step of constructing the baseline response model includes: extracting a sequence of spatiotemporal feature vectors from the multimodal baseline data through the spatiotemporal feature extraction module; and inputting the sequence of spatiotemporal feature vectors into the sequence analysis module to output the baseline response model. This application, through a decoupled modular design, enables the model to efficiently process spatiotemporal information and long-range dependencies in multimodal data such as video and audio, providing a solid foundation for constructing accurate response models.

[0011] In one possible implementation of the first aspect, the multimodal baseline data and the multimodal instantaneous data are collected while the user performs at least one facial movement challenge and at least one speech pronunciation challenge. This application, by combining facial and speech—two of the most stroke-affected functions—for comprehensive evaluation, can more fully cover potential neurological deficits, improving the coverage and accuracy of model recognition.

[0012] In one possible implementation of the first aspect, the step of determining the dynamic difference using the hybrid deep learning model includes: calculating the trajectory deviation of the instantaneous response trajectory relative to the baseline response model; calculating the dynamic asymmetry represented by the multimodal instantaneous data; calculating the response delay represented by the multimodal instantaneous data; and weightedly combining the trajectory deviation, the dynamic asymmetry, and the response delay to obtain the LFFP. The LFFP calculated by the model in this application is a multi-dimensional composite index that can comprehensively quantify neural function deficits from three orthogonal dimensions: action morphology, spatial symmetry, and temporal response speed, making the final recognition result more accurate and reliable.

[0013] In one possible implementation of the first aspect, the step of determining the FPAR includes: calculating the rate of change of the LFFP over time to obtain the FPAR. This application uses FPAR to directly reflect the rate of functional degradation, making it a dynamic indicator with significant early warning value. Compared to the static LFFP value, it allows the model to identify the collapse trend of system stability earlier.

[0014] In one possible implementation of the first aspect, the method further includes: if the LFFP output by the hybrid deep learning model exceeds a first preset threshold, adaptively adjusting the parameters of subsequent challenges, and re-acquiring the multimodal instantaneous data based on the adjusted challenges for processing by the hybrid deep learning model. This application, by endowing the model with a closed-loop feedback capability, enables deeper probing in critical cases by increasing the challenge difficulty, thereby improving the confidence and accuracy of identification.

[0015] In one possible implementation of the first aspect, the method further includes: the hybrid deep learning model predicting the future trend of the LFFP based on the current value of the FPAR; when the future trend exceeds a preset collapse threshold, the model directly generates a high-risk stroke facial recognition result. This application, through feedforward prediction, enables the model to identify high risk before complete functional collapse, further improving the timeliness of assessment and reducing the burden on users.

[0016] Secondly, this application provides a pre-hospital stroke facial recognition system based on a hybrid deep learning model, which includes a processor and a memory, wherein the memory stores a program configured to perform the method as described in the first aspect.

[0017] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the functional module structure of a pre-hospital stroke facial recognition system based on a hybrid deep learning model, provided in one embodiment of this application. Figure 2 This is a flowchart of a pre-hospital stroke facial recognition method based on a hybrid deep learning model, provided in one embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0020] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0021] In this application, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral part; it can be a direct connection or an indirect connection through an intermediate medium. Furthermore, the term "electrical connection" can refer to the manner in which an electrical connection is used to achieve signal transmission.

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0023] Unlike traditional methods, the "facial recognition" defined in this application does not simply determine whether a face is symmetrical, but rather performs a deep, quantitative identification of the functional state of the neuromuscular response system composed of facial features, speech, and other factors. This method achieves a leap from "static morphological recognition" to "dynamic functional stability recognition" through a specially designed hybrid deep learning model. This model constructs a personalized "baseline response model" for each user in their healthy state, and accurately identifies early, subtle signs of stroke by calculating the potential functional failure potential (LFFP) and the rate of failure potential accumulation (FPAR) at the time of suspected attack.

[0024] See attached document Figure 1This application provides an embodiment of a pre-hospital stroke facial recognition system 200 based on a hybrid deep learning model. This system can be implemented on smartphones, tablets, dedicated health monitoring devices, or any electronic device equipped with a processor, camera, microphone, and display screen. The core of the system is a program running on the processor. Logically, the system can be divided into several functional units.

[0025] The data acquisition and preprocessing unit 205 is responsible for interacting with the device's hardware (camera, microphone) to acquire raw multimodal data. This unit is also responsible for standardizing the raw data through preprocessing, providing high-quality input for subsequent model analysis. For video data, preprocessing includes, but is not limited to: face detection and alignment to ensure the face is centered and its pose is consistent within the video frame; illumination normalization to reduce the interference of ambient light changes on subsequent analysis; and facial keypoint detection, such as using existing mature technologies like MediaPipe Face Mesh to extract the coordinate sequence of hundreds (e.g., 468) 3D facial keypoints in each frame in real time. For audio data, preprocessing includes standard speech signal processing steps such as noise reduction, silence removal, framing, and windowing, followed by extraction of acoustic features, such as Mel-frequency cepstral coefficients (MFCCs) and their first and second-order differences.

[0026] In this embodiment, the core computing function of the system is undertaken by a hybrid deep learning model core 215. This model core 215 logically includes multiple internal units that work together, specifically including: a baseline model building unit 210, an instantaneous response analysis unit 220, a dynamic difference calculation unit 230, and a decision sub-model unit 250.

[0027] The baseline model building unit 210's core task is to use a hybrid deep learning model to process the user's multimodal baseline data in a healthy state to construct a personalized baseline response model. Upon first use, this unit guides the healthy user through a series of standardized challenges and invokes the data acquisition and preprocessing unit 205 to obtain high-quality baseline data. Then, the model's internal spatiotemporal feature extraction module and sequence analysis module work together to compress and represent these complex spatiotemporal data streams, representing the user's health state, into a low-dimensional, compact mathematical model—the user's personal baseline response model. Once established, this model is stored as the "gold standard" for all future identification tasks for that user.

[0028] The transient response analysis unit 220 is activated when a user or their family suspects a stroke and triggers an identification task. It guides the user through the same standardized challenge used to build the baseline and acquires transient multimodal data. Crucially, this unit uses the exact same model structure and parameter weights as the baseline model building unit 210 to process the transient data, mapping it to the same feature space as the baseline response model, thereby generating a directly comparable transient response trajectory. This strict consistency is fundamental to ensuring the reliability of the identification results.

[0029] The dynamic difference calculation unit 230 is the most innovative computational core of this hybrid deep learning model. It receives the baseline response model generated by the baseline model building unit 210 and the instantaneous response trajectory generated by the instantaneous response analysis unit 220 as input. Instead of performing a simple Euclidean distance comparison, this unit comprehensively quantifies the difference between the instantaneous response and the healthy baseline through three orthogonal dimensions: trajectory deviation, dynamic asymmetry, and response delay. Its outputs are two key dynamic indicator time series: LFFP(t) and FPAR(t). LFFP(t) quantifies the degree of neural instability at each time point, while FPAR(t) quantifies the rate of deterioration of this instability.

[0030] The risk identification unit 250 is the final decision-making module of the model. It receives two dynamic indicator sequences, LFFP and FPAR, as input. Internally, this unit integrates a decision sub-model, which can be a pre-trained classifier based on a large amount of clinical data (such as a support vector machine, gradient boosting tree, or a small fully connected neural network), or a system of logical rules derived from medical expert knowledge. This unit performs statistical analysis on the input indicator sequences (such as calculating peak values ​​and mean values) and inputs the analysis results into the decision sub-model, ultimately outputting a clear and easy-to-understand "stroke facial recognition result," which directly corresponds to a specific risk level, such as "low risk," "moderate risk, attention recommended," or "high risk, immediate medical attention recommended."

[0031] See attached document Figure 2 The specific steps of the pre-hospital stroke facial recognition method based on a hybrid deep learning model provided in the embodiments of this application will be described in detail below.

[0032] S100: Construct a baseline response model.

[0033] The goal of this step is to use a hybrid deep learning model to build a highly personalized baseline response model for users in a healthy state. This step is typically performed when a user first installs and sets up the application, or during a regular health check, ensuring that it is done while the user's neural functions are intact.

[0034] First, the system guides the user through a set of standardized neuromuscular challenges (SNCs) using clear voice and interface prompts. This comprehensively and repeatedly stimulates the user's facial and speech-related neuromuscular systems, thereby acquiring high-quality, information-rich multimodal baseline data. For example, this set of challenges is designed to cover different facial muscle groups and articulation functions, and may include: Maximum Smile Challenge: The system prompts, "Please do your best to make a bright smile and hold it for 3 seconds." This challenge primarily assesses the symmetry and strength of the core facial muscles that make you smile, such as the levator labii superioris and zygomaticus major.

[0035] Cheek puffing exercise: The system prompts, "Please puff out your cheeks as if blowing up a balloon, and hold for 3 seconds." This challenge primarily assesses the closure ability of the cheek muscles and the control of the orbicularis oris muscle, and is highly sensitive to detecting weakness in the perioral muscles.

[0036] Close your eyes tightly: The system prompts, "Please close your eyes tightly with all your might and hold for 3 seconds." This challenge primarily assesses the function of the orbicularis oculi muscle, an important indicator of the function of the upper branch of the facial nerve.

[0037] Clear Pronunciation: The system prompts, "Please read the sentence clearly at a normal speaking speed: The weather is very nice today." This challenge assesses the coordinated movement of the tongue, lips, palate, and other speech organs, as well as the presence of articulation disorders.

[0038] Throughout the user's execution of the aforementioned challenge, the system's data acquisition and preprocessing unit 205 records video at a frame rate of at least 60fps and a resolution of 1080p, while simultaneously recording audio at a sampling rate of at least 44.1kHz. These high-specification acquisition parameters are designed to ensure that sufficiently fine dynamic changes can be captured.

[0039] After acquiring the data, the core hybrid deep learning model of this application begins to process these multimodal baseline data to construct a baseline response model. In a preferred embodiment, the hybrid deep learning model is architecturally composed of a spatiotemporal feature extraction module and a sequence analysis module.

[0040] The goal of the spatiotemporal feature extraction module is to extract meaningful features from the raw data stream. For video data, this module first uses a facial keypoint detection algorithm (such as MediaPipe Face Mesh) to extract 468 3D coordinate points per frame. However, a simple sequence of coordinate points lacks description of local dynamic textures. Therefore, the model in this application adopts a 3D convolutional neural network (3D-CNN) structure. It takes consecutive video frame segments (e.g., 16 frames as an input unit) as input and learns to capture the motion and deformation patterns of local regions around facial keypoints within a short time window, i.e., spatiotemporal features, through multiple 3D convolutional and pooling layers. For example, a 3D convolutional kernel (such as...) This module can slide simultaneously in space (the x and y axes of the image) and time (the t axis of the frame sequence), effectively learning dynamic information such as the speed and acceleration of a smile. For audio data, the extracted MFCCs (e.g., 39 dimensions per frame) are themselves a type of spatiotemporal feature. Finally, the module aligns and catenates the spatiotemporal features extracted from the video with the audio features in the time dimension, forming a higher-dimensional, unified sequence of spatiotemporal feature vectors, for example, a sequence of length... ( (for time frames), dimension is (For example, The characteristic sequence of ).

[0041] In a preferred embodiment, the architecture of the 3D-CNN module is designed to handle a standardized video clip input. Specifically, the input video clip is preprocessed into a shape of... The tensor has dimensions corresponding to frame count (time), height (space), width (space), and RGB color channels, respectively. The detailed structure of this module is as follows: Convolutional Block 1: The input tensor first passes through a 3D convolutional layer, which has... There are filters, and the kernel size is . Step size is The padding method is 'same' to maintain spatial dimensions. A batch normalization layer follows the convolution operation to accelerate model convergence and improve stability. Next is a rectified linear unit (ReLU) activation function to introduce non-linearity. Finally, a 3D max-pooling layer (MaxPooling3D) is applied. The pooling size is applied to the feature map. This operation does not downsample in the time dimension, but halves the height and width in the spatial dimension, thus reducing computation while preserving temporal dynamics. After this block, the shape of the output feature map becomes... .

[0042] Convolutional Block 2: Similar to Convolutional Block 1, this block contains a... Each filter, kernel size is The 3D convolutional layer is followed by a batch normalization layer and a ReLU activation function. The subsequent 3D max-pooling layer uses... The pooling size is adjusted, and downsampling is performed in both time and space dimensions. The output feature map shape of this block becomes... .

[0043] Convolutional Block 3: This block further deepens the network and extracts more abstract features, containing a... Each filter, kernel size is A 3D convolutional layer, followed by batch normalization and ReLU. Also using a... The 3D max-pooling layer. The output feature map shape of this block becomes... .

[0044] Feature vectorization: The final output of the 3D-CNN module is a shape of The four-dimensional tensor. To input it into the subsequent sequence analysis module, it needs to be converted into a two-dimensional feature sequence. Specifically, at each time step (a total of 4 time steps), the... The spatial feature map is flattened into a one-dimensional vector. This yields a vector of length [missing information]. Each element has a dimension of The video feature sequence.

[0045] For audio data, the extracted ( The MFCCs features (number of audio frames) are mapped to the same temporal length (4 in this example) and similar feature dimensions as the video feature sequence through a simple fully connected layer or a one-dimensional convolutional network, facilitating subsequent fusion. Finally, this module concatenates the processed video feature sequence with the audio feature sequence along the feature dimensions, forming a unified, higher-dimensional spatiotemporal feature vector sequence, for example, a sequence of length... , dimension The characteristic sequence will be sent to the sequence analysis module.

[0046] After obtaining the spatiotemporal feature sequence, a module capable of understanding long-range dependencies within the sequence is needed. This application employs a Transformer encoder as the sequence analysis module. Its key advantage lies in its ability, through a self-attention mechanism, to process all elements in the sequence in parallel and directly model the relationship between any two time points. In a preferred embodiment, this Transformer encoder consists of... One (for example, It consists of stacked identical encoder layers. The input processing and internal structure of a single encoder layer will be described in detail below.

[0047] The Transformer model itself does not contain any information about the sequence order. To enable the model to utilize the sequence order, positional information must be injected into the input features. Therefore, when using a dimensional... Before the feature sequence is fed into the encoder stack, positional encoding needs to be added to it. Positional encoding is an encoding that has the same dimension as the input feature sequence. The matrix is ​​generated using a predefined set of sine and cosine functions of different frequencies. Specifically, for positions in the sequence... and feature dimensions Location coding The value can be determined by the following formula:

[0048]

[0049] This positional encoding matrix is ​​added element-wise to the original feature sequence matrix, and the result is used as the input to the first encoder layer. This allows the model to learn the absolute and relative positional information of the elements.

[0050] Each encoder layer consists of two main sub-layers: a multi-head self-attention sub-layer and a position-wise feed-forward network sub-layer.

[0051] Sub-layer 1: Multi-head self-attention mechanism The purpose of this sub-layer is to allow each element in the sequence to "pay attention" to all other elements in the sequence and compute a weighted, context-dependent representation.

[0052] a. Linear projection generates Q, K, V: The sequence input to this sub-layer (dimension 1) First, the query is generated through three independent linear (fully connected) layers. ), key ) and Value Three matrices.

[0053] b. Multi-head attention calculation: The "multi-head" mechanism is to... Segmented along the feature dimension One (for example, The smaller part, the "head," performs a scaled dot-product attention computation independently on each head. The core idea is that for each query vector... Calculate its relationship with all key vectors The dot product similarity is calculated, then scaled using a scaling factor (the square root of the key vector dimension) to stabilize the gradient. The similarity score is then converted into attention weights using a softmax function, and finally, these weights are applied to all value vectors. Perform a weighted summation. This process allows each head to focus on different aspects of the input sequence.

[0054] c. Result assembly and output: The outputs of each head are concatenated together and then projected through the last linear layer to obtain the final output of the multi-head self-attention sublayer.

[0055] For example, suppose the feature vector of an input sequence at one time step is The model has 8 heads. The sequence analysis module will employ a Transformer encoder. Its key advantage lies in its ability, through a self-attention mechanism, to process all elements in the sequence in parallel and directly model the relationship between any two time points. In a preferred embodiment, this Transformer encoder is composed of… One (for example, It consists of stacked identical encoder layers. The input processing and internal structure of a single encoder layer will be described in detail below.

[0056] The Transformer model itself does not contain any information about the sequence order. To enable the model to utilize the sequence order, positional information must be injected into the input features. Therefore, when using a dimensional... Before the feature sequence is fed into the encoder stack, positional encoding needs to be added to it. Positional encoding is an encoding that has the same dimension as the input feature sequence. The matrix is ​​generated using a predefined set of sine and cosine functions of different frequencies. Specifically, for positions in the sequence... and feature dimensions Location coding The value can be determined by the following formula:

[0057]

[0058] This positional encoding matrix is ​​added element-wise to the original feature sequence matrix, and the result is used as the input to the first encoder layer. This allows the model to learn the absolute and relative positional information of the elements.

[0059] Each encoder layer consists of two main sub-layers: a multi-head self-attention sub-layer and a position-wise feed-forward network sub-layer.

[0060] Sub-layer 1: Multi-head self-attention mechanism The purpose of this sub-layer is to allow each element in the sequence to "pay attention" to all other elements in the sequence and compute a weighted, context-dependent representation.

[0061] a. Linear projection generates Q, K, V: The sequence input to this sub-layer (dimension 1) First, the query is generated through three independent linear (fully connected) layers. ), key ) and Value Three matrices.

[0062] b. Multi-head attention calculation: The "multi-head" mechanism is to... Segmented along the feature dimension One (for example, The smaller part, the "head," performs a scaled dot-product attention computation independently on each head. The core idea is that for each query vector... Calculate its relationship with all key vectors The dot product similarity is calculated, then scaled using a scaling factor (the square root of the key vector dimension) to stabilize the gradient. The similarity score is then converted into attention weights using a softmax function, and finally, these weights are applied to all value vectors. Perform a weighted summation. This process allows each head to focus on different aspects of the input sequence.

[0063] c. Result assembly and output: The outputs of each head are concatenated together and then projected through the last linear layer to obtain the final output of the multi-head self-attention sublayer.

[0064] For example, suppose the feature vector of an input sequence at one time step is The model has 8 heads. It will be projected onto 8 different weight matrices to obtain 8 smaller sets. Each head independently computes attention, and then the eight results are concatenated and projected again to obtain the final context vector. .

[0065] A residual connection (i.e., the two are added together) is applied between the output and input of the multi-head self-attention sublayer, followed by a layer normalization step. This "Add & Normalize" step is crucial for training deep Transformer models, as it effectively prevents the vanishing gradient problem and stabilizes the training process.

[0066] Sublayer 2: Location-fed forward network The output after processing by the first sub-layer is fed into a position-feedforward network. This network applies the same transformation independently to each position in the sequence. It consists of two linear layers and a ReLU activation function located between them.

[0067] For example, the first linear layer can reduce the feature dimension from Expand to The second linear layer then projects it back. .

[0068] Final residual connections and layer normalization: Another "Add & Norm" step is applied between the output and input of the position feedforward network.

[0069] The complete processing flow of an encoder layer can be summarized as: Input -> Multi-head self-attention -> Add & Norm -> Feedforward network -> Add & Norm -> Output.

[0070] The input sequence is processed After stacking several such encoder layers, the final output is a dimension of (in It can be equal to For example, a low-dimensional latent representation sequence (e.g., 64). The set of these low-dimensional latent representation sequences generated when a user performs all challenges in a healthy state, in... In the latent space of 3D, a smooth, compact, bounded geometric structure (such as a manifold or a Gaussian mixture model distribution) is formed. This mathematical object, learned and parameterized by the hybrid deep learning model, which can accurately describe the user's personal health response pattern, is defined as the user's "baseline response model." This model (including its network weights and geometric description in the latent space) is stored for subsequent comparisons.

[0071] S200: Generate instantaneous response trajectory.

[0072] When a user feels unwell and initiates the recognition process, the system guides them through a standardized challenge identical to that in the S100. The system collects multimodal instantaneous data and feeds it into the same hybrid deep learning model that has already been loaded with the user's personal baseline model parameters. The model, through spatiotemporal feature extraction and sequence analysis modules, processes the user's real-time response into a sequence within the same... The trajectory in the latent space. This new trajectory is the transient response trajectory. Due to potential neural deficits, this transient response trajectory will theoretically deviate from the healthy region defined by the baseline response model constructed in S100.

[0073] S300: Calculate dynamic difference indices (LFFP and FPAR).

[0074] The model's dynamic difference calculation unit 230 quantifies the difference between the instantaneous response trajectory and the baseline response model through a series of internal calculations.

[0075] S310: Calculate the latent functional failure potential (LFFP).

[0076] The LFFP model is defined as a comprehensive, time-varying function. The higher the value, the greater the degree of neurological dysfunction. It is composed of a weighted combination of three components:

[0077] in, These are the weights obtained by training the model on a large amount of data, used to balance the importance of each component. For example, one could take... .

[0078] trajectory deviation Used to measure the degree of abnormality in the overall form of the movement. For each time point on the instantaneous response trajectory. Corresponding latent vector The model needs to find a reference point that best matches the healthy region defined by the baseline response model. This can be achieved by searching for [something] in the potential space. The Euclidean distance is to the nearest baseline data point, or if the baseline model is modeled as a generative model (such as a variational autoencoder). It can be The orthogonal projection onto the model manifold. For example, assume that in At that time, the instantaneous latent vector output by the model is The closest point found by the model in its stored baseline response model is... Then the trajectory deviation The normalized L2 norm of these two vectors is calculated as follows: .

[0079] Dynamic asymmetry This component is specifically designed to capture typical unilateral facial paralysis symptoms of stroke. The model's dynamic difference calculation unit directly selects symmetrical key point pairs (such as the left and right corners of the mouth) from the facial key point coordinate sequence obtained in the preprocessing stage. The analysis was performed. Dynamic asymmetry was defined as a normalized difference value, calculated using the following formula:

[0080] In this formula, Indicates a point in time. Indicates a point in time The kinematic parameter values ​​of key points on the left side (such as the left corner of the mouth) relative to a stationary state. This parameter can be displacement, velocity, or acceleration; in this embodiment, displacement in the vertical direction is taken as an example. Indicates a point in time The corresponding kinematic parameter values ​​of the right keypoint (such as the right corner of the mouth) that are symmetrical to the left keypoint. This indicates the absolute value operation. It is a numerical stability constant, whose technical function is to prevent the denominator from being zero. In some cases, such as at the beginning or end of an action, the displacement of key points on the left and right sides... and Both could be zero, which would result in a zero denominator and lead to a calculation error (dividing by zero). Adding a very small positive number to the denominator can solve this. This ensures that the denominator is always greater than zero, thus guaranteeing the stability and robustness of the calculation. The value of should be small enough to have a substantial impact on the calculation result of the asymmetry, but it needs to be greater than the floating-point precision of machine calculation. For example, in a preferred embodiment, It can be set to a very small floating-point number, for example (Right now ).

[0081] Assume that for the smile challenge, the model tracks the vertical displacement of the left and right corners of the mouth. At time point... The left corner of the mouth was detected to be upturned. ( (), while the right corner of his mouth only turned up. ( At this point, the dynamic asymmetry is calculated internally by the model as follows: The calculation result is a... Dimensionless number of the interval Indicates perfect symmetry. It indicates extreme asymmetry.

[0082] Response latency Used to quantify the speed of neural conduction and muscle response. The model automatically detects predefined "feature events" in motion or speech data streams. For example, for the smile challenge, a feature event might be defined as "the speed at which the corners of the mouth turn upwards reaches its peak." The model learns the user's healthy baseline time from stored baseline data. for In this identification, the model detected the time of occurrence of the instantaneous event. for Then the original delay To ensure dimensional consistency in the LFFP formula, the model must be adapted to... Dimensionless processing is performed. The model uses a reference time constant. (For example, Normalize: .

[0083] Substitute the values ​​from the above examples into the LFFP formula: (Assuming here) and (Calculated at the same time) S320: Calculate the failure potential accumulation rate (FPAR).

[0084] Achieved continuous Following the sequence, the model's dynamic difference calculation unit further calculates its rate of change over time, or FPAR, to reveal the dynamic trend of functional deterioration. For example, the model calculates in hour, ,exist hour, The sampling time interval is The instantaneous FPAR calculated by the model is: This positive value indicates that the user's neural function stability is improving. The potential unit deteriorates at a rate per second.

[0085] S400: Generate stroke facial recognition results.

[0086] Finally, the model's risk identification unit 250 performs a comprehensive analysis of the two complete dynamic indicator sequences, LFFP and FPAR, and outputs the final identification result. The S400 includes: S401: Construction of risk feature vectors.

[0087] The risk identification unit processes two complete time series as input. and Statistical analysis was conducted to extract a set of key statistical features that comprehensively summarize user performance during this challenge. These features collectively constitute a fixed-dimensional risk feature vector. The advantage of this approach is that it transforms variable-length time-series data into fixed-length input, facilitating processing by subsequent classification models. In a preferred embodiment, the risk feature vector... Including but not limited to the following components: The peak value reached by the LFFP sequence throughout the challenge reflects the most severe degree of neurological instability.

[0088] The arithmetic mean of the LFFP sequence reflects the overall level of instability.

[0089] The integral of the LFFP sequence over time (i.e., the area under the curve) represents the total amount of functional impairment accumulated throughout the entire challenge.

[0090] The positive peak reached by the FPAR sequence reflects the fastest instantaneous rate of functional deterioration.

[0091] The arithmetic mean of the FPAR series reflects the average trend of functional deterioration.

[0092] The time it takes for LFFP to reach its peak can reflect whether the damage appears rapidly or develops slowly.

[0093] For example, suppose that the model, after calculation, yields the following statistical characteristics: , , , These values ​​will be combined into a vector, for example... .

[0094] S402: The pre-trained decision sub-model classifies the risk feature vector.

[0095] Constructed risk feature vector The data is fed into the core of the risk identification unit: a pre-trained decision sub-model (classifier). The classifier's task is to map this high-dimensional risk feature vector to a predefined, discrete risk category.

[0096] In a preferred embodiment, the decision sub-model is a Support Vector Machine (SVM) classifier. The core idea of ​​SVM is to find one or more optimal separating hyperplanes in a high-dimensional feature space, which can separate data points of different classes with the largest possible margin. In this application, the data points are risk feature vectors. The categories are different stroke risk levels (e.g., low risk, moderate risk, high risk). For multi-class problems (three or more risk levels), a "one-vs-one" or "one-vs-rest" strategy is usually adopted, which is achieved by combining multiple binary SVM classifiers.

[0097] Since the relationship between the risk feature vector and the actual stroke risk is likely to be highly nonlinear, directly searching for a linear hyperplane in the original feature space may not yield good results. Therefore, the SVM in this embodiment employs a kernel trick. Through a kernel function, the SVM can implicitly transform the original feature vector... This is mapped to a higher-dimensional or even infinite-dimensional Hilbert space, and a linear separating hyperplane is found in this higher-dimensional space. This is equivalent to finding a complex nonlinear decision boundary in the original space.

[0098] For example, this embodiment preferably uses a radial basis function (RBF) kernel, the formula of which is: in, It is an adjustable hyperparameter that determines the complexity of the decision region.

[0099] In order for the SVM classifier to make accurate recognitions, it must be trained offline. The training process is as follows: A large-scale, diverse training dataset was collected. This dataset contains thousands of multimodal videos and audio recordings of different individuals (including confirmed stroke patients and confirmed healthy controls) performing standardized challenges (SNCs) in a controlled environment. Each data point was independently evaluated by at least two senior neurologists and labeled with a clear risk level, such as 0 for "low risk / healthy", 1 for "moderate risk / suspected TIA", and 2 for "high risk / clear stroke symptoms".

[0100] For each piece of multimodal data in the training set, the complete process described in S100 to S401 of this application is used for processing, that is, calculating the corresponding risk feature vector for each sample. .

[0101] Risk feature vectors of all samples The data, along with its corresponding risk level labels (0, 1, 2), are used as input to train the SVM classifier. The goal of training is to find the hyperplane parameters (including support vectors and Lagrange multipliers) that optimally separate these three classes of data points.

[0102] During training, cross-validation is typically used to find the optimal combination of hyperparameters (e.g., RBF kernel). Parameters and SVM penalty coefficients This is to prevent the model from overfitting on the training set and to improve its generalization ability on unknown data.

[0103] The penalty coefficient C controls the model's tolerance for misclassified samples, and its core role is to balance the objectives of "maximizing the classification margin" and "minimizing training error." A large C value means the model punishes misclassification heavily, trying its best to correctly classify all training samples. However, this can lead to an overly complex decision boundary that closely follows the training data, resulting in overfitting—the model performs well on the training set but poorly on unseen new data. Conversely, a smaller C value means the model is more tolerant of misclassification, allowing some training samples to be misclassified in exchange for a wider, smoother classification margin. This helps improve the model's generalization ability, but if the C value is too small, it can lead to underfitting—the model is too simple to capture complex patterns in the data. In practice, C is typically searched on a logarithmic scale, for example, choosing from a set {0.1, 1, 10, 100, 1000}.

[0104] The RBF kernel parameter γ defines the size of the influence range of a single training sample. A large γ value means the RBF kernel radius is small, and only data points very close to the support vectors will affect them. This can lead to a very "rough" decision boundary, highly dependent on individual support vectors. This is also highly prone to overfitting. Conversely, a small γ value means the RBF kernel radius is large, and the influence range of each support vector is wide, resulting in a very smooth decision boundary. If the γ value is too small, it can lead to underfitting, where the model cannot distinguish subtle differences in the data. The search for γ is also typically performed on a logarithmic scale, for example, by selecting from a set {0.0001, 0.001, 0.01, 0.1, 1}.

[0105] To find the optimal (C,γ) combination, this embodiment uses a standard grid search and k-fold cross-validation method.

[0106] First, define a two-dimensional grid consisting of candidate C and γ values. For example, C∈{0.1,1,10,100}, γ∈{0.001,0.01,0.1}.

[0107] For each pair (C) in the grid i ,γ j The training dataset is randomly divided into k similar, disjoint subsets. Then, k iterations are performed. In each iteration, one subset is selected as the validation set, and the remaining k-1 subsets are used as the training set. The current (C) is then used to perform k-fold cross-validation. i ,γ j Train the SVM model with parameters and evaluate its performance on the validation set (e.g., calculate accuracy or F1 score).

[0108] After k iterations, calculate (C) i ,γ j The average performance of the (C,γ) combination on the k validation sets is calculated. After iterating through all parameter combinations, the (C,γ) combination that maximizes the average performance is selected as the final optimal hyperparameter.

[0109] Use the found optimal hyperparameters (C) opt ,γ opt The SVM model is then retrained on the entire training dataset. This results in a model that utilizes all the information from the data and possesses optimal generalization ability. This final, fixed decision sub-model, containing all optimized parameters, is stored as part of the risk identification unit 250 for online inference.

[0110] When a new user undergoes stroke risk identification, the process is as follows: The system executes steps S100 to S401 to calculate the user's instantaneous risk feature vector. .

[0111] Will The input is fed into the pre-trained SVM decision sub-model.

[0112] The SVM model utilizes its defined decision boundary to... Perform the classification and output a category label. For example, suppose... The vector is input into the model, which calculates and determines that it is located within the decision region corresponding to the "high risk" category, and therefore outputs category label 2.

[0113] S403: Generate and present stroke facial recognition results.

[0114] After receiving the category labels from the decision sub-model, the risk identification unit converts them into a human-readable and explicit risk level description and presents it to the user.

[0115] For example, if the received tag is 2, the system will display the recognition result on the user interface in the most prominent way (such as a red highlighted background, bold font, and alarm icon): "Stroke facial recognition result: High risk. It is recommended to contact emergency services immediately or go to the nearest hospital." Through the detailed steps outlined above, the hybrid deep learning model of this application can not only extract advanced dynamic features, but also complete a reliable mapping from data to clinical decisions through a fully explained and rigorously trained decision sub-model, thereby solving the core problems in the prior art.

[0116] In an optional implementation, the method of the present invention further includes a closed-loop feedback adaptive challenge adjustment mechanism. This mechanism is designed to handle some critical states or cases with unclear outcomes by dynamically adjusting the difficulty of the challenge to conduct deeper exploration, thereby improving the confidence and accuracy of identification.

[0117] Specifically, the mechanism is implemented as follows: After executing step S402, if the statistical value of the latent functional failure potential (LFFP) output by the decision sub-model (e.g., LFFP) is... peakk If the LFFP value exceeds the first preset threshold (for example, the threshold is set to 0.4, representing the upper limit of the low-risk range), but does not reach the second preset threshold (for example, the threshold is 0.7, representing a clear high risk), that is, the LFFP value falls within a preset "fuzzy range", the system will not immediately output the final result.

[0118] At this point, the system will adaptively adjust the parameters of subsequent challenges. This adjustment can increase the difficulty of the challenge or change the form of the challenge. For example, the system can issue a more challenging new instruction to the user via voice: "Please quickly and alternately puff out your left and right cheeks, and repeat three times."

[0119] Then, based on this adjusted challenge, the system reacquires a second set of multimodal instantaneous data and executes the complete process from S200 to S402 again, which is processed by the hybrid deep learning model to calculate a new set of more discriminative LFFP and FPAR values. Because the "perturbation" intensity of the challenge is increased, potential, subtle neural function deficits are more easily amplified and exposed, allowing the newly calculated risk feature vectors to more clearly fall into a specific risk category, resulting in a final identification result with higher confidence.

[0120] In another alternative embodiment, the method of the present invention further includes a feedforward predictive control mechanism. This mechanism is designed to further improve the timeliness of the assessment and reduce the burden on users who may be in an uncomfortable state.

[0121] The core of this mechanism lies in the fact that the hybrid deep learning model not only analyzes historical data that has already occurred, but also predicts future states. Specifically, the model predicts the future trend of the potential functional failure potential (LFFP) based on the current value of the failure potential accumulation rate (FPAR).

[0122] This predictive function can be achieved by integrating a lightweight sequence prediction subnetwork (e.g., a recurrent neural network (RNN) or a long short-term memory network (LSTM)) into the risk identification unit 250. This subnetwork takes the time series of FPAR as input, learns the dynamics of LFFP changes, and predicts its numerical trend within a short time window in the future (e.g., the next 500 milliseconds).

[0123] During real-time identification, if the predictive model determines that the future trend of LFFP will exceed a preset threshold representing complete functional failure (e.g., the "failure threshold" can be set to 1.0), the system can proceed without waiting for the user to complete the entire standardized challenge. At this point, the model triggers an early termination mechanism, directly generating a high-risk stroke facial recognition result and immediately issuing the highest-level risk alert. This feedforward predictive capability allows the model to identify high-risk cases before complete functional failure, significantly reducing identification time.

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of this application.

[0125] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0126] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0127] The units described as separate components may or may not be physically separate. A component shown as a unit can be one physical unit or multiple physical units; that is, it can be located in one place or distributed in multiple different places. Depending on actual needs, some or all of the units can be selected to achieve the purpose of this embodiment.

[0128] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware.

[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A pre-hospital stroke facial recognition method based on a hybrid deep learning model, characterized in that, include: Acquire the user's multimodal baseline data and process the multimodal baseline data using the hybrid deep learning model to construct a baseline response model that characterizes the user's health status; The user's multimodal instantaneous data is acquired, and the multimodal instantaneous data is processed using the hybrid deep learning model to generate an instantaneous response trajectory; The hybrid deep learning model is used to determine the dynamic difference between the instantaneous response trajectory and the baseline response model in order to obtain the potential functional failure potential. Determine the temporal variation trend of the potential functional failure potential in order to obtain the failure potential accumulation rate; Based on the potential functional failure potential and the accumulation rate of the failure potential, the stroke facial recognition result of the user is generated, and the recognition result represents the stroke risk level.

2. The method according to claim 1, characterized in that, The hybrid deep learning model includes a spatiotemporal feature extraction module for extracting spatiotemporal features and a sequence analysis module for performing sequence modeling on the spatiotemporal features; the step of constructing the baseline response model includes: The spatiotemporal feature extraction module extracts a sequence of spatiotemporal feature vectors from the multimodal baseline data. The spatiotemporal feature vector sequence is input into the sequence analysis module to output the baseline response model.

3. The method according to claim 1, characterized in that, The multimodal baseline data and the multimodal instantaneous data are collected when the user performs at least one facial movement challenge and at least one speech pronunciation challenge.

4. The method according to claim 1, characterized in that, The step of determining the dynamic differences using the hybrid deep learning model includes: Calculate the deviation of the instantaneous response trajectory from the baseline response model; Calculate the dynamic asymmetry represented by the multimodal instantaneous data; Calculate the response delay characterized by the multimodal instantaneous data; The potential functional failure potential is obtained by weighting and combining the trajectory deviation, dynamic asymmetry, and response delay.

5. The method according to claim 1, characterized in that, The step of determining the failure potential accumulation rate includes: Calculate the rate of change of the potential functional failure potential over time to obtain the cumulative rate of the failure potential.

6. The method according to claim 1, characterized in that, The method further includes: If the potential functional failure output by the hybrid deep learning model exceeds a first preset threshold, the parameters of the subsequent challenge are adaptively adjusted, and the multimodal instantaneous data is reacquired based on the adjusted challenge for processing by the hybrid deep learning model.

7. The method according to claim 1 or 5, characterized in that, The method further includes: The hybrid deep learning model predicts the future trend of the potential functional failure potential based on the current value of the failure potential accumulation rate. When the future trend exceeds a preset collapse threshold, the model directly generates a high-risk stroke facial recognition result.

8. A pre-hospital stroke facial recognition system based on a hybrid deep learning model, characterized in that, include: At least one processor; The memory is communicatively connected to the processor; The memory stores a computer program that, when executed by the processor, implements the method as described in any one of claims 1 to 7.

9. The system according to claim 8, characterized in that, The hybrid deep learning model includes: A baseline model building unit is used to process multimodal baseline data to build a baseline response model; A transient response analysis unit is used to process multimodal transient data to generate transient response trajectories; A dynamic difference calculation unit is used to determine the dynamic difference between the instantaneous response trajectory and the baseline response model, so as to output the potential functional failure potential and the failure potential accumulation rate; A risk identification unit is used to generate stroke facial recognition results characterizing the stroke risk level based on the potential functional failure potential and the failure potential accumulation rate.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.