A method for identifying fall and collision injuries

Through video data processing and feature fusion technology, the location and severity of the fall and collision injury are accurately identified, and the problem of inaccurate identification in the prior art is solved, which improves rescue efficiency and reduces the risk of secondary injury.

CN116704413BActive Publication Date: 2025-05-16UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310695249.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-05-16
Estimated Expiration
2043-06-13

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately identify the location and severity of the fall and collision injury, which makes it impossible for rescue personnel to know the injury situation at the scene as soon as possible, which may lead to secondary injuries.

Method used

By collecting video data, preprocessing is performed to extract skeleton information and local area of ​​attention, combining human action flow and local image flow processing, a spatio-temporal graph convolution network and 3D convolution network are used to extract features, and feature fusion is performed through a multi-layer perceptron to identify the fall collision site and determine the severity of the injury.

Benefits of technology

It improves the accuracy of identifying damage parts after falling and collision, and effectively classifies the severity of the damage, providing more valuable reference information for timely alarms and subsequent rescue, reducing the probability of secondary damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704413B_ABST
    Figure CN116704413B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying injuries caused by falls and collisions. First, a video is collected to obtain images, and the obtained image data is preprocessed to extract skeleton information and local areas of interest. Then, human motion stream processing and local image stream processing are respectively performed to obtain motion features and local image features, which are fused through a multi-layer perceptron to obtain a collision site identification result. Finally, the collision site identification result sequence is further mined to obtain an injury severity classification result. The method of the present invention relies on ordinary video surveillance equipment and computers, and improves the accuracy of identifying the injury site of falls and collisions by fusing human posture and motion features with image features of local areas of interest, and grades the severity of injuries, providing more valuable reference information for timely warning and subsequent rescue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of electronic information technology, and in particular relates to a method for identifying fall and collision injuries. Background Art

[0002] Trauma is the main reason for hospital emergency visits, and falls and collisions account for a large proportion of them. Fall and collision injuries refer to different degrees of impact between the human body and the ground, resulting in minor to severe injuries. Fall injuries with serious consequences mainly occur in the elderly.

[0003] The part of the body where a person falls is an important factor affecting the severity of collision injuries. Among fall collision injuries, the most serious injuries are to the head and hip. If the collision injury occurs in the head, it is easy to cause symptoms such as concussion and cerebral hemorrhage, resulting in a high mortality rate; if the collision injury occurs in the hip, there is a high disability rate, and hip fractures often paralyze the person in bed and are difficult to operate on. In addition to the direct injuries, shock and prolonged collapse caused by collision injuries are also important factors of death.

[0004] The above hazards of fall injuries can be promptly warned through fall injury event detection, thereby enabling timely rescue. In addition, providing accurate injury information to rescuers is also of great significance. Especially when the patient is in a coma or other state where he cannot communicate, the rescuers cannot immediately know the location and severity of the injury on the scene, and operations such as checking the patient's injury and moving the patient may cause secondary injuries. If the injury location and severity can be automatically assessed at the same time as the injury is detected, rescuers can conduct targeted injury inspections and protective operations, thereby saving precious rescue time and reducing the probability of secondary injuries.

[0005] Regarding fall injury identification, the current related technologies are mainly divided into two categories: methods based on non-wearable sensors and methods based on wearable sensors.

[0006] Non-wearable methods mainly collect image data through fixed cameras (including depth cameras), and then identify the human body state in each frame of the image, such as calculating the coordinates of each body part, the center of gravity of the body and other data. Based on the calculated data or features, further detection is performed to determine whether a collision injury event has occurred. Among them, CN202110198589.0 predicts the type of fall injury based on skeleton data; CN201910222523.3 also detects fall events based on images.

[0007] Wearable methods mainly install accelerometers, gyroscopes and other instruments on the user's body in the form of belts, bracelets, etc., collect data such as acceleration of various axes in each direction through wearable instruments, and then calculate whether a collision injury has occurred. Among them, CN201910853809.1 identifies the type of fall, that is, the direction of fall, based on the data collected by wearable sensors.

[0008] Existing technologies usually only determine whether a fall injury event has occurred or not, and cannot provide more effective information to help rescue; some methods provide classification information on the type of injury, such as frontal falls, side falls, and backward falls in fall collisions, but this rough classification ignores the details of the collision process and does not consider the severity of the injury. It has low reference value for medical staff to implement subsequent clinical treatment; at the same time, there is a lack of human posture modeling during the fall process, resulting in a high false alarm rate. Summary of the invention

[0009] In order to solve the above technical problems, the present invention proposes a method for identifying fall and collision injuries, which identifies the site of fall and collision injury and determines the severity of the injury to provide a reference for subsequent treatment.

[0010] The technical solution adopted by the present invention is: a method for identifying fall and collision injuries, and the specific steps are as follows:

[0011] S1, collect video to obtain images;

[0012] S2, preprocessing the image data obtained in step S1 to extract skeleton information and local focus areas;

[0013] S3, human motion stream processing;

[0014] S4, local image stream processing;

[0015] S5, feature fusion and collision site identification;

[0016] S6. Determination of the degree of damage by serialized modeling.

[0017] Furthermore, the step S1 is specifically as follows:

[0018] The image is obtained by collecting video data from ordinary surveillance cameras. The image data is recorded as P and the expression is as follows:

[0019] P=[p1,p2,...,p T ]

[0020] Among them, p t represents the t-th frame image in P, and T represents the total number of video frames.

[0021] Furthermore, in step S2, the image data obtained in step S1 is input, the image is preprocessed frame by frame, and the skeleton information and the local focus area are obtained by posture estimation, as follows:

[0022] S21, extract skeleton;

[0023] The openpose framework is used to perform frame-level processing on the original video, extract the character positioning frame and obtain the skeleton node data stream corresponding to each frame of the video stream.

[0024] Then each frame image p gets the corresponding skeleton data s:

[0025] s=[J1,J2,...,J 13 ]

[0026] Among them, the skeleton nodes are set to 13, J i The numerical meaning of is the coordinate value in the rectangular coordinate system with the center of the character image as the origin (J i x , J i y ), and i=1, 2,...13.

[0027] The single-frame skeleton data s constitutes the skeleton information S of the entire video in time sequence.

[0028] S22, intercepting the local area of ​​interest;

[0029] Based on the recognized human posture, i.e., the single-frame skeleton data s, the image around the human joint is intercepted, i.e., a square image with the joint point as the center and a side length equal to 15% of the body height is obtained, and 13 RGB data streams are obtained.

[0030] Assume r i =R(J i ) represents the interception joint J i The square area with a side length equal to 15% of the height is the center. The set r of N = s = 13 local focus areas obtained in each frame image p is expressed as follows:

[0031] r=[r1,r2,...,r 13 ]

[0032] Among them, r i Represents the local area of ​​interest corresponding to the i-th joint.

[0033] Furthermore, the step S3 is specifically as follows:

[0034] S31, constructing a spatiotemporal action graph;

[0035] The skeleton nodes S of T frame images are [s1, s2, ..., s T] is taken as a point, and an action space-time graph G is constructed based on the human body structure and time sequence. The node set of the space-time graph is represented as V. The action space-time graph contains the action information of the video under the skeleton modeling dimension.

[0036] The human body structure refers to the connection between body parts or joints; and the time sequence refers to the connection between the same joints in adjacent frames.

[0037] The vertex set V of the action space-time graph G consists of vertices v ti Composition, t represents the tth frame in the time dimension, i represents the joint number, and 1-N represents different body parts or joints; the expression of V is as follows:

[0038] V={v ti |t=1,2,...,T;i=1,2,...,N}

[0039] S32, spatiotemporal graph convolution operation;

[0040] A time step in the action spatiotemporal graph is called a single layer, for which the graph convolution operation is defined.

[0041] The overall convolution formula is as follows:

[0042]

[0043] Among them, Output(·) represents the features extracted after spatiotemporal graph convolution; f in (·) represents the input feature location function, through which a sequence number can be mapped to the corresponding feature; sp(·) represents the sampling function, which is used for node v ti The neighbors of are enumerated to obtain the serial index; ω(·) represents the weight function, which provides the weight vector in the input feature real space and is used to calculate the inner product of the input feature vector after sampling; l ST (·) represents the label function; B(v ti ) represents node v ti The neighbor set of Z ti (v qj )=|{v rk |l ti (v rk )=l ti (v qj ),v rk ∈B(v ti )}|Used to eliminate the imbalance caused by the different number of elements in different types of subsets.

[0044] Where d(·) represents the distance between nodes; q represents the qth frame in the time dimension; the fixed value D=1 means that only a single node and the relationship between its directly adjacent nodes are considered in the same time step; the fixed value Γ=3 means that only the relationship between the current time step and the previous and next steps is considered in the same joint position.

[0045] At node v ti The neighbor set B(v ti ) defines the sampling function:

[0046] sp(v ti ,v qj )=v qj

[0047] The weight function ω(·) is obtained by the label function l ST (·) Subset the neighborhood of the node; for the node v with time step t and sequence number i ti For its neighboring node v qj The label function is defined as follows:

[0048]

[0049] Among them, the label function l in a single time step ti (·) Set node v ti The neighbor set B(v ti ) is divided into a fixed number of subsets K = 3, using a mapping l ti :B(v ti )→{0,1,2} maps neighbor nodes to their subsets in the following way:

[0050]

[0051] Among them, dist i Represents the distance from joint i to the center of the body.

[0052] After the single time step label function is defined, the cross-time step labels are added, and finally K×Γ labels are obtained.

[0053] S33, extracting action flow features;

[0054] Based on steps S31-S32, an action spatiotemporal graph is constructed, and the spatiotemporal graph convolution is used to define the Output stacking to obtain a spatiotemporal graph network F, and the action flow feature M is obtained by calculation:

[0055] M=F out (G)

[0056] Among them, F out(·) represents the calculation process of the spatiotemporal graph convolutional network, which is specifically formed by connecting multiple separate spatiotemporal graph convolutional layer Outputs in sequence. The first layer input is the action spatiotemporal graph G constructed in step S31, and the subsequent input is the feature graph with the same structure obtained after the previous layer is convolved with the spatiotemporal graph.

[0057] Furthermore, in step S4, the set of local focus areas in step S22 is used as input to output local image flow features, specifically as follows:

[0058] S41, for each of the N input local focus area image streams, perform convolution operation to extract features;

[0059] A 3D convolutional network is used to extract spatiotemporal features. The 3D convolutional network is stacked in a hierarchical structure. For the feature extraction calculation of the i-th layer, the height output by the i-1th layer is H i , width is W i , the time length is T i , characteristic thickness is M i The local characteristic space-time graph u (i-1) As input, use M i The 3D convolution kernels are used for convolution operation to obtain the feature map output of the i-th layer.

[0060] Among them, the operation formula of the jth convolution kernel of the i-th layer is expressed as follows:

[0061]

[0062] A convolution kernel performs a convolution on a pixel with spatial coordinates (x, y, z) with kernel weight w and offset weight b, and uses the ReLU activation function.

[0063] The local attention region set r obtained in step S22 is used as the input of the convolutional network, that is, it is regarded as the feature map of layer number i = 0. The N local attention regions are subjected to spatiotemporal convolution operations to obtain the image convolution feature stream C = [C1, C2, ..., C N ].

[0064] Among them, C i Represents the local focus area r i The feature vector obtained after dimensionality reduction of the feature map:

[0065] C i =Conv(Conv(...Conv(r i )))

[0066] S42, multi-head attention mechanism integrates body part features;

[0067] First, consider C as a feature sequence, where each part of the sequence corresponds to a corresponding body part. i Respectively with the learnable query weight matrix W q , key weight matrix W k , value weight matrix W v Multiply them together, so that the body part information is diffused into three different feature spaces. After dimensional scaling to stabilize the dot product value distribution, the similarity matrix is ​​calculated through softmax to obtain the final output matrix.

[0068] Each C i The calculation formula of attention scale is as follows:

[0069]

[0070] in,(·) T Represents the transpose of a matrix.

[0071] After obtaining the scaled attention processing features of all N parts, each part is regarded as a different channel dimension, and the multi-head attention formula is used to let different heads learn different information of the part, discover different levels of part features, and integrate this information to obtain the final representation. The multi-head attention formula is as follows:

[0072] MultiAttention(C)=Concat(head1,head2,...,head N )W e

[0073] Among them, W e Represents the fusion weight matrix of multi-head convolutional attention. Apply layer normalization to the output of multi-head attention to obtain the final output local image flow feature E.

[0074] Furthermore, the step S5 is specifically as follows:

[0075] An adaptive fusion method is used to fuse the motion flow feature output M in step S3 and the local image flow feature E in step S4.

[0076] First, M and E are concatenated to obtain a mixed feature vector H, and then the feature vector is input into a multi-layer perceptron to map the features to a new feature space, and is converted into a new feature representation through a ReLU activation function to enhance the feature expression capability.

[0077] After the motion features and local image features of a frame are fused in the multi-layer perceptron, a probability vector of the collision part is obtained, and the dimension of the vector is equal to the number of collision detection categories. The probability vector is normalized through the softmax layer, and the final collision part recognition result corresponding to the frame of video is output.

[0078] Furthermore, the step S6 is specifically as follows:

[0079] Based on steps S1-S5, the information of the falling collision position of each frame image is obtained. These collision position information totals T frames to form a collision position prediction sequence Y c , Y c Represents a sequence of length T.

[0080] Manually identify the sequence Y for the collision site c Mark the corresponding collision risk level Y d The degree of collision risk is divided into three levels: high, medium and low. The main basis for classification is:

[0081] 1. Whether there is any collision with the head, hips or other parts that are prone to serious injuries;

[0082] 2. The order in which the various parts collide;

[0083] 3. Specific collision posture, such as whether there is any supporting and cushioning action during the collision process.

[0084] Use the marked danger level Y d As labels, apply LSTM model to extract Y c The model can be used to identify the degree of danger of fall collision injuries by combining derivative information such as the fall collision site, collision sequence, and collision duration of specific parts.

[0085] Beneficial effects of the present invention: The method of the present invention first collects video to obtain images, pre-processes the obtained image data, extracts skeleton information and local focus areas, and then respectively processes the human motion stream and the local image stream to obtain motion features and local image features, which are fused through a multi-layer perceptron to obtain collision site recognition results, and finally further mine the collision site recognition result sequence to obtain injury severity classification results. The method of the present invention relies on ordinary video surveillance equipment and computers, and improves the accuracy of identifying the injury site of a fall collision by fusing the human posture and motion features with the image features of the local focus area, and grades the severity of the injury, providing more valuable reference information for timely warning and subsequent rescue. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 The present invention is a flowchart of a method for identifying fall and collision injuries.

[0087] Figure 2 Schematic diagram of the preprocessing process in an embodiment of the present invention.

[0088] Figure 3 It is a skeleton space-time diagram in an embodiment of the present invention. DETAILED DESCRIPTION

[0089] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0090] like Figure 1 As shown, a flow chart of a fall and collision injury identification method of the present invention, the specific steps are as follows:

[0091] S1, collect video to obtain images;

[0092] S2, preprocessing the image data obtained in step S1 to extract skeleton information and local focus areas;

[0093] S3, human motion stream processing;

[0094] S4, local image stream processing;

[0095] S5, feature fusion and collision site identification;

[0096] S6. Determination of the degree of damage by serialized modeling.

[0097] In this embodiment, the step S1 is specifically as follows:

[0098] The image is obtained by collecting video data from ordinary surveillance cameras. In this embodiment, the specification is 720*1280@30hz. The image data is recorded as P, and the expression is as follows:

[0099] P=[p1,p2,...,p T ]

[0100] Among them, p t represents the t-th frame image in P, and T represents the total number of video frames.

[0101] like Figure 2 As shown, in this embodiment, in step S2, the image data obtained in step S1 is input, the image is preprocessed frame by frame, and the skeleton information and the local focus area are obtained by posture estimation, as follows:

[0102] S21, extract skeleton;

[0103] The openpose framework is used to perform frame-level processing on the original video, extract the character positioning frame and obtain the skeleton node data stream corresponding to each frame of the video stream.

[0104] Then each frame image p gets the corresponding skeleton data s:

[0105] s=[J1,J2,...,J 13 ]

[0106] Among them, the skeleton nodes are set to 13, J iThe numerical meaning of is the coordinate value in the rectangular coordinate system with the center of the character image as the origin (J i x , J i y ), and i = 1, 2, ... 13. The skeleton node comparison table is shown in Table 1.

[0107] Table 1

[0108]

[0109] The single-frame skeleton data s constitutes the skeleton information S of the entire video in time sequence.

[0110] S22, intercepting the local area of ​​interest;

[0111] Based on the recognized human posture, i.e., the single frame skeleton data s, the image around the human joint is intercepted, i.e., a square image with the joint point as the center and the side length equal to 15% of the height is obtained to obtain 13 RGB data streams. In principle, the method of the present invention supports the appropriate increase or decrease of the parts, and the actual effect will change accordingly.

[0112] Assume r i =R(J i ) represents the interception joint J i The set r of N = |s| = 13 local focus areas is obtained from each frame of image p, which is a square area with a side length equal to 15% of the height. The expression is as follows:

[0113] r=[r1,r2,...,r 13 ]

[0114] Among them, r i Represents the local area of ​​interest corresponding to the i-th joint.

[0115] In this embodiment, step S3 is specifically as follows:

[0116] The input of human motion flow processing is the skeleton information S obtained in the preprocessing step, and the output is the human motion flow feature M.

[0117] S31, constructing a spatiotemporal action graph;

[0118] The skeleton nodes S of T frame images are [s1, s2, ..., s T ] is taken as a point, and an action space-time graph G is constructed based on the human body structure and time sequence. The node set of the space-time graph is represented as V. The action space-time graph contains the action information of the video under the skeleton modeling dimension.

[0119] Among them, the human body structure refers to the connection between body parts or joints, such as the edge connection between the corresponding nodes of the left wrist and the left elbow; the time sequence refers to the connection between the same joints between adjacent frames, such as the connection between the head node of the tth frame and the nodes of the t-1th and t+1th frames.

[0120] The vertex set V of the action space-time graph G consists of vertices v ti Composition, t represents the tth frame in the time dimension, i represents the joint number (1-N represents different body parts or joints). The expression of V is as follows:

[0121] V={v ti |t=1,2,...,T;i=1,2,...,N}

[0122] The skeleton space-time diagram is as follows Figure 3 shown.

[0123] S32, spatiotemporal graph convolution operation;

[0124] A time step in the action spatiotemporal graph is called a single layer, for which the graph convolution operation is defined.

[0125] The overall convolution formula is as follows:

[0126]

[0127] The graph convolution operation processes the input graph data and generates a higher-level feature map on the graph.

[0128] Among them, Output(·) represents the features extracted after spatiotemporal graph convolution; f in (·) represents the input feature location function. Through the feature location function, a serial number can be mapped to the corresponding feature; sp(·) represents the sampling function. Due to the complexity of the graph structure, the traversal of features cannot be indexed by simple numerical subscripts. It is necessary to use the sampling function sp(·) to index the node v ti The neighbors of are enumerated to obtain the serial index. ω(·) represents the weight function, which provides the weight vector in the input feature real space and is used to calculate the inner product of the input feature vector after sampling; l ST (·) represents the label function; B(v ti ) represents node v ti The neighbor set of Z ti (v qj )=|{v rk |l ti (v rk )=l ti (v qj ),v rk ∈B(v ti)}|Used to eliminate the imbalance caused by the different number of elements in different types of subsets.

[0129] Where d(·) represents the distance between nodes; q represents the qth frame in the time dimension; the fixed value D=1 means that only a single node and the relationship between its directly adjacent nodes are considered in the same time step. The reason is that the graph network formed by the human skeleton is small in size, and a sampling distance larger than 1 will not improve the effect and will increase the amount of calculation; the fixed value Γ=3 means that the same joint position only focuses on the relationship between the current time step and the previous and next steps.

[0130] At node v ti The neighbor set B(v ti ) defines the sampling function:

[0131] sp(v ti ,v qj )=v qj

[0132] The weight function ω(·) is obtained by the label function l ST (·) Subset the neighborhood of the node; for the node v with time step t and sequence number i ti For its neighboring node v qj The label function is defined as follows:

[0133]

[0134] Among them, the label function l in a single time step ti (·) Set node v ti The neighbor set B(v ti ) is divided into a fixed number of subsets K = 3, using a mapping l ti :B(v ti )→{0,1,2} maps neighbor nodes to their subsets in the following way:

[0135]

[0136] Among them, dist i Represents the distance from joint i to the center of the body.

[0137] After the single time step label function is defined, the cross-time step labels are added, and finally K×Γ labels are obtained.

[0138] S33, extracting action flow features;

[0139] Based on steps S31-S32, an action spatiotemporal graph is constructed, and the spatiotemporal graph convolution is used to define the Output stacking to obtain a spatiotemporal graph network F, and the action flow feature M is obtained by calculation:

[0140] M=F out (G)

[0141] Among them, F out (·) represents the calculation process of the spatiotemporal graph convolutional network, which is specifically formed by connecting multiple separate spatiotemporal graph convolutional layer Outputs in sequence. The first layer input is the action spatiotemporal graph G constructed in step S31, and the subsequent input is the feature graph with the same structure obtained after the previous layer is convolved with the spatiotemporal graph.

[0142] In this embodiment, in step S4, the set of local focus areas in step S22 is used as input to output local image flow features, specifically as follows:

[0143] Single skeleton motion information cannot reflect the complex human-environment interaction in real fall events, such as the positional relationship between the ground or other objects and the human body and the process of position change. For these image information that cannot be obtained from human posture movements, the corresponding RGB data is obtained and image features are extracted to enhance the accuracy of fall collision injury recognition.

[0144] Taking the local focus region r in the preprocessing step as input, it outputs the local image flow feature E.

[0145] S41, for each of the N input local focus area image streams, perform convolution operation to extract features;

[0146] A single image stream represents the information of the body part itself and its surrounding environment, and there is action information in the time dimension between the frames of the image stream. Therefore, a 3D convolutional network is used to extract spatiotemporal features.

[0147] The 3D convolutional network is stacked in a hierarchical structure. For the feature extraction calculation of the i-th layer, the height output by the i-1th layer is H i , width is W i , the time length is T i , characteristic thickness is M i The local characteristic space-time graph u (i-1) As input, use M i The 3D convolution kernels are used for convolution operation to obtain the feature map output of the i-th layer.

[0148] Among them, the operation formula of the jth convolution kernel of the i-th layer is expressed as follows:

[0149]

[0150] A convolution kernel performs a convolution on a pixel with spatial coordinates (x, y, z) with kernel weight w and offset weight b, and uses the ReLU activation function.

[0151] The local attention region set r obtained in step S22 is used as the input of the convolutional network, that is, it is regarded as the feature map of layer number i = 0. The N local attention regions are subjected to spatiotemporal convolution operations to obtain the image convolution feature stream C = [C1, C2, ..., C N ].

[0152] Among them, C i Represents the local focus area r i The feature vector obtained after dimensionality reduction of the feature map:

[0153] C i =Conv(Conv(...Conv(r i )))

[0154] S42, multi-head attention mechanism integrates body part features;

[0155] First, consider C as a feature sequence, where each part of the sequence corresponds to a corresponding body part. i Respectively with the learnable query weight matrix W q , key weight matrix W k , value weight matrix W v Multiply them together, so that the body part information is diffused into three different feature spaces. After dimensional scaling to stabilize the dot product value distribution, the similarity matrix is ​​calculated through softmax to obtain the final output matrix.

[0156] Each C i The calculation formula of attention scale is as follows:

[0157]

[0158] in,(·) T Represents the transpose of a matrix.

[0159] After obtaining the scaled attention processing features of all N parts, each part is regarded as a different channel dimension, and the multi-head attention formula is used to let different heads learn different information of the part, discover different levels of part features, and integrate this information to obtain the final representation. The multi-head attention formula is as follows:

[0160] MultiAttention(C)=Concat(head1,head2,...,head N )W e

[0161] Among them, W e Represents the fusion weight matrix of multi-head convolutional attention. Apply layer normalization to the output of multi-head attention to obtain the final output local image flow feature E.

[0162] In this embodiment, step S5 is specifically as follows:

[0163] An adaptive fusion method is used to fuse the motion flow feature output M in step S3 and the local image flow feature E in step S4.

[0164] First, M and E are concatenated to obtain a mixed feature vector H, and then the feature vector is input into a multi-layer perceptron to map the features to a new feature space, and is converted into a new feature representation through a ReLU activation function to enhance the feature expression capability.

[0165] After the motion features and local image features of a frame are fused in the multi-layer perceptron, a probability vector of the collision part will be obtained, and the dimension of the vector is equal to the number of collision detection categories. The probability vector is normalized through the softmax layer, and the final collision part recognition result corresponding to the frame of video is output.

[0166] In this embodiment, step S6 is specifically as follows:

[0167] Based on steps S1-S5, the information of the falling collision position of each frame image is obtained. These collision position information totals T frames to form a collision position prediction sequence Y c , Y c Represents a sequence of length T.

[0168] Manually identify the sequence Y for the collision site c Mark the corresponding collision risk level Y d The degree of collision risk is divided into three levels: high, medium and low. The main basis for classification is:

[0169] 1. Whether there is any collision with the head, hips or other parts that are prone to serious injuries;

[0170] 2. The order in which the various parts collide;

[0171] 3. Specific collision posture, such as whether there is any supporting and cushioning action during the collision process.

[0172] Use the marked danger level Y d As labels, apply LSTM model to extract Y c The model can be used to identify the degree of danger of fall collision injuries by combining derivative information such as the fall collision site, collision sequence, and collision duration of specific parts.

[0173] In summary, in the method of the present invention, the model input is video data, and the output, in addition to the conventional fall injury event detection results, can also obtain more specific detailed information that can guide on-site first aid. The specific detailed information includes the accurate identification results of the collision site and the dangerous degree of the fall collision injury. The model takes into account both human motion and local image features, and models them separately to improve the recognition accuracy. For human posture movements, a spatiotemporal graph convolutional network is used to extract features; for local images, a 3D convolutional network is used to extract features. The motion features and local image features are fused through a multi-layer perceptron to obtain the collision site identification results; the collision site identification result sequence is further mined to obtain the injury severity classification results.

[0174] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A method for identifying fall and collision injuries, the specific steps are as follows: S1, collect video to obtain images; S2, preprocessing the image data obtained in step S1 to extract skeleton information and local focus areas; The local focus region acquisition method is: based on a frame of human body posture identified, that is, a single frame of skeleton data s, an image around the human body joint is intercepted, that is, a square image with a joint point as the center and a side length equal to 15% of the body height is obtained, and 13 RGB data streams are obtained as the local focus region set r; S3, human motion stream processing; S4, local image stream processing; S41, for each of the N input local focus area image streams, perform convolution operation to extract features; The obtained local attention region set r is used as the input of the 3D convolutional network, that is, it is regarded as the feature map of layer number i = 0; the N local attention regions are subjected to spatiotemporal convolution operations to obtain the image convolution feature flow C = [C1, C2, ..., C N ]; S42, multi-head attention mechanism integrates body part features; First, consider C as a feature sequence, where each part of the sequence corresponds to a corresponding body part; i Respectively with the learnable query weight matrix W q , key weight matrix W k , value weight matrix W v Multiply them together, so that the body part information is diffused into three different feature spaces. After dimensional scaling to stabilize the dot product value distribution, the similarity matrix is ​​calculated through softmax to obtain the final output matrix. After obtaining the scaled attention processing features of all N parts, each part is regarded as a different channel dimension, and a multi-head attention formula is used to allow different heads to learn different information about the parts, discover different levels of part features, and integrate this information to obtain the final representation; the multi-head attention formula is as follows: MultiAttention(C)=Concat(head1,head2,......,head N )W e in, W e Represents the fusion weight matrix of multi-head convolutional attention; layer normalization is applied to the output results of multi-head attention to obtain the final output local image flow feature E; S5, feature fusion and collision site identification; S6. Determination of the degree of damage by serialized modeling.

2. A fall and collision injury identification method according to claim 1, characterized in that: The step S1 is specifically as follows: The image is obtained by collecting video data from ordinary surveillance cameras. The image data is recorded as P and the expression is as follows: P=[p1,p2,...,p T ] Among them, p t represents the t-th frame image in P, and T represents the total number of video frames.

3. A fall and collision injury identification method according to claim 1, characterized in that: In step S2, the image data obtained in step S1 is input, and the image is preprocessed frame by frame, and the skeleton information and the local focus area are obtained by posture estimation, as follows: S21, extract skeleton; Use the Openpose framework to process the original video at the frame level, extract the character positioning frame and obtain the skeleton node data stream corresponding to each frame of the video stream; Then each frame image p gets the corresponding skeleton data s: s=[J1,J2,...,J 13 ] Among them, the skeleton nodes are set to 13, J i The numerical meaning of is the coordinate value in the rectangular coordinate system with the center of the character image as the origin And i=1, 2,...13; The single-frame skeleton data s composes the skeleton information S of the entire video in time sequence; S22, intercepting the local area of ​​interest; Assume r i =R(J i ) represents the interception joint J i The set r of N = |s| = 13 local focus areas is obtained from each frame of image p, which is a square area with a side length equal to 15% of the height. The expression is as follows: r=[r1,r2,...,r 13 ] Among them, r i Represents the local area of ​​interest corresponding to the i-th joint.

4. A fall and collision injury identification method according to claim 1, characterized in that: The step S3 is specifically as follows: S31, constructing a spatiotemporal action graph; The skeleton nodes S of T frame images are [s1, s2, ..., s T ] as points, and an action spatiotemporal graph G is constructed based on the human body structure and time sequence. The node set of the spatiotemporal graph is represented as V. The action spatiotemporal graph contains the action information of the video in the skeleton modeling dimension; The human body structure refers to the connection between body parts or joints; the time sequence refers to the connection between the same joints in adjacent frames; The vertex set V of the action space-time graph G consists of vertices v ti Composition, t represents the tth frame in the time dimension, i represents the joint number, and 1-N represents different body parts or joints; the expression of V is as follows: V={v ti |t=1,2,...,T;i=1,2,...,N} S32, spatiotemporal graph convolution operation; A time step in the action spatiotemporal graph is called a single layer, for which a graph convolution operation is defined; The overall convolution formula is as follows: Among them, Output(·) represents the features extracted after spatiotemporal graph convolution; f in (·) represents the input feature location function, through which a sequence number can be mapped to the corresponding feature; sp(·) represents the sampling function, which is used for node v ti The neighbors of are enumerated to obtain the serial index; ω(·) represents the weight function, which provides the weight vector in the input feature real space and is used to calculate the inner product of the input feature vector after sampling; l ST (·) represents the label function; B(v ti ) represents node v ti The neighbor set of Z ti (v qj )=|{v rk |l ti (v rk )=l ti (v qj ),v rk ∈B(v ti )}|Used to eliminate the imbalance caused by the different number of elements in different types of subsets; Where d(·) represents the distance between nodes; q represents the qth frame in the time dimension; the fixed value D = 1 means that only a single node and the relationship between its directly adjacent nodes are considered in the same time step; the fixed value Γ = 3 means that only the relationship between the current time step and the previous and next time steps is considered in the same joint position; At node v ti The neighbor set B(v ti ) defines the sampling function: sp(v ti ,v qj )=v qj The weight function ω(·) is obtained by the label function l ST (·) Subset the neighborhood of the node; for the node v with time step t and sequence number i ti For its neighboring node v qj The label function is defined as follows: Among them, the label function l in a single time step ti (·) Set the node v ti The neighbor set B(v ti ) is divided into a fixed number of subsets K = 3, using a mapping l ti :B(v ti )→{0,1,2} maps neighbor nodes to their subsets in the following way: Among them, dist i represents the distance from joint i to the center of the body; After the single time step label function is defined, add the cross-time step labels, and finally get K×Γ labels; S33, extracting action flow features; Based on steps S31-S32, an action spatiotemporal graph is constructed, and the spatiotemporal graph convolution is used to define the Output stacking to obtain a spatiotemporal graph network F, and the action flow feature M is obtained by calculation: M=F out (G) Among them, F out (·) represents the calculation process of the spatiotemporal graph convolutional network, which is specifically formed by connecting multiple separate spatiotemporal graph convolutional layer Outputs in sequence. The first layer input is the action spatiotemporal graph G constructed in step S31, and the subsequent input is the feature graph with the same structure obtained after the previous layer is convolved with the spatiotemporal graph.

5. A fall and collision injury identification method according to claim 1, characterized in that: In step S4, the set of local focus areas in step S22 is used as input to output local image flow features, specifically as follows: In step S41, a 3D convolutional network is used to extract spatiotemporal features. The 3D convolutional network is stacked in a hierarchical structure. For the feature extraction calculation of the i-th layer, the height output by the i-1th layer is H i , width is W i , the time length is T i , characteristic thickness is M i The local characteristic space-time graph u (i-1) As input, M i The 3D convolution kernels are convolved to obtain the feature map output of the i-th layer; Among them, the operation formula of the jth convolution kernel of the i-th layer is expressed as follows: A convolution kernel performs convolution on the pixel point with spatial coordinate position (x, y, z), with kernel weight w and offset weight b, and the activation function used is ReLU; Among them, C i Represents the local focus area r i The feature vector obtained after dimensionality reduction of the feature map: C i =Conv(Conv(...Conv(r i ))) In step S42, each C i The calculation formula for scaling attention is as follows: in,(·) T Represents the transpose of a matrix.

6. A fall and collision injury identification method according to claim 1, characterized in that: The step S5 is specifically as follows: Adopting an adaptive fusion method to fuse the action stream feature output M in step S3 and the local image stream feature E in step S4; First, M and E are concatenated to obtain a mixed feature vector H, which is then input into a multi-layer perceptron to map the features to a new feature space, and converted into a new feature representation through a ReLU activation function to enhance the feature expression capability; After the motion features and local image features of a frame are fused in the multi-layer perceptron, a probability vector of the collision part is obtained, and the dimension of this vector is equal to the number of collision detection categories; the probability vector is normalized through the softmax layer, and the final collision part recognition result corresponding to the frame of video is output.

7. A method for identifying fall and collision injuries according to claim 1, characterized in that: The step S6 is specifically as follows: Based on steps S1-S5, the information of the falling collision position of each frame image is obtained. These collision position information totals T frames to form a collision position prediction sequence Y c , Y c represents a sequence of length T; Manually identify the sequence Y for the collision site c Mark the corresponding collision risk level Y d ; The degree of collision risk is divided into three levels: high, medium and low. The basis for the division is: A1. Whether there has been a collision with the head or hip, which are parts that are prone to serious injuries; A2. The order in which the various parts collide; A3. Specific collision posture, whether there is support and buffering action during the collision process; Use the marked danger level Y d As labels, apply LSTM model to extract Y c The model is used to identify the degree of danger of fall collision injuries by combining the fall collision location, collision sequence and duration of collision in specific locations.

Citation Information

Patent Citations

  • A pedestrian fall detection method based on skeleton detection

    CN109919132B

  • Fall type and injury part detection method based on feature classification

    CN110659595A

  • A method, system, and terminal for predicting the severity of fall injuries based on skeletal data.

    CN112998697B