Elevator interior fall behavior recognition method and system based on deep learning

By performing feature mining and multi-level feature enhancement on the monitoring images and audio data inside the elevator, the problem of low reliability in recognizing falls inside elevators has been solved, achieving more reliable fall recognition and alarm.

CN120599689BActive Publication Date: 2026-08-04SICHUAN SPECIAL EQUIP INSPECTION & RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN SPECIAL EQUIP INSPECTION & RES INST
Filing Date
2025-04-25
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, the reliability of recognizing falls in elevators is not high, and it is difficult to detect and identify falls in a timely manner in a closed and confined environment, which increases the risk of injury.

Method used

A deep learning-based approach is used to perform feature mining on the monitoring images and audio data inside the elevator, and to perform multi-level feature enhancement by combining the audio and image data. The results of fall behavior recognition are then analyzed and used as the triggering condition for the alarm device.

Benefits of technology

This improves the reliability of fall detection and enhances the reliability of alarm actions, ensuring that fall behavior is detected and alarms are triggered promptly inside elevators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599689B_ABST
    Figure CN120599689B_ABST
Patent Text Reader

Abstract

This application provides a deep learning-based method and system for recognizing falls inside elevators, relating to the field of deep learning technology. In this application, firstly, feature mining is performed on the target monitoring image data to output target image data features; then, feature mining is performed on the target monitoring audio data to output target audio data features; subsequently, based on the target audio data features, feature enhancement is performed on the target image data features to output enhanced image data features. During feature enhancement, multiple local data features in the target image data features are enhanced at multiple levels based on multiple local data features in the target audio data features; finally, based on the enhanced image data features, the target fall behavior recognition result is analyzed and output, and the target fall behavior recognition result serves as the alarm trigger condition for the target alarm device. Based on the above, the relatively low reliability of fall behavior recognition in existing technologies can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and more specifically, to a method and system for recognizing falling behavior inside elevators based on deep learning. Background Technology

[0002] With the increasing aging of the population, falls have become one of the most common accidental injuries among the elderly, especially in enclosed, confined spaces like elevators, where falls are often difficult to detect in time, greatly increasing the risk of injury. Therefore, real-time monitoring and recognition of falls in elevators has become an important research topic. However, current technologies typically use trained classifiers to identify videos and images to obtain corresponding fall recognition results. But falls are diverse and share similarities with other behaviors, making recognition based on existing classifiers prone to unreliable results. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method and system for recognizing fall behavior inside elevators based on deep learning, so as to improve the problem of relatively low reliability of fall behavior recognition in the prior art.

[0004] To achieve the above objectives, this application adopts the following technical solution:

[0005] A deep learning-based method for recognizing falls inside elevators, comprising:

[0006] Acquire target monitoring image data and target monitoring audio data inside the target elevator;

[0007] Feature mining is performed on the target surveillance image data to output the target image data features;

[0008] Perform feature mining on the target monitoring audio data and output the target audio data features;

[0009] Based on the target audio data features, feature enhancement is performed on the target image data features, and enhanced image data features are output. In the feature enhancement process, multiple local data features in the target image data features are enhanced in multiple levels based on multiple local data features in the target audio data features.

[0010] Based on the enhanced image data features, the target fall behavior recognition result is analyzed and output. The target fall behavior recognition result is used to characterize whether there is a fall behavior inside the target elevator, and the target fall behavior recognition result serves as the alarm action trigger condition for the target alarm device.

[0011] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator fall recognition method, the step of enhancing the target image data features based on the target audio data features and outputting the enhanced image data features includes:

[0012] Based on the continuity between corresponding audio frames and / or based on the similarity between local data features, the multiple local data features included in the target audio data features are combined to form multiple first data feature combinations, wherein each first data feature combination includes at least one local data feature, and when multiple local data features are included, the audio frames corresponding to the multiple local data features are continuous.

[0013] Based on the multiple first data feature combinations, the multiple local data features included in the target image data features are combined to form multiple second data feature combinations, wherein there is a one-to-one correspondence between the second data feature combinations and the first data feature combinations, and the included local data features have temporal consistency with respect to the corresponding audio frames and monitoring images.

[0014] Based on the multiple combinations of first data features, feature enhancement is performed on the multiple combinations of second data features to output enhanced image data features.

[0015] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator interior fall recognition method, the step of enhancing the multiple second data feature combinations based on the multiple first data feature combinations and outputting enhanced image data features includes:

[0016] Multiple enhancement levels are determined based on the multiple combinations of second data features, wherein there is a one-to-one correspondence between the multiple enhancement levels and the multiple combinations of second data features;

[0017] In each enhancement level, based on the output features corresponding to the previous enhancement level, the second data feature combination corresponding to the current enhancement level is enhanced by the first feature, and the enhancement feature corresponding to the current enhancement level is output. The output feature corresponding to the previous enhancement level of the first enhancement level is a feature formed by learning the corresponding sample data.

[0018] In each enhancement level, based on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level, the enhancement feature corresponding to the current enhancement level is enhanced by the second feature, and the output feature corresponding to the current enhancement level is output.

[0019] Based on the output features of this last enhancement level, the enhanced image data features are determined.

[0020] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator interior fall recognition method, the step of performing first feature enhancement on the second data feature combination corresponding to the current enhancement level based on the output features corresponding to the previous enhancement level in each enhancement level, and outputting the enhanced features corresponding to the current enhancement level, includes:

[0021] The output feature corresponding to the previous enhancement level and the second data feature corresponding to the current enhancement level are combined and spliced ​​to form a two-level spliced ​​feature. Based on the size of the output feature corresponding to the previous enhancement level, the two-level spliced ​​feature is subjected to sliding window processing.

[0022] Each sliding window feature formed by the sliding window process is subjected to self-attention processing to form the first attention feature corresponding to each sliding window feature;

[0023] The mean of each of the first attention features is calculated to form the reinforcement feature corresponding to the current reinforcement level.

[0024] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator interior fall behavior recognition method, the step of performing second feature enhancement on the enhancement features corresponding to the current enhancement level based on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level, and outputting the output features corresponding to the current enhancement level in each enhancement level, includes:

[0025] Perform feature space transformation on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level to form the corresponding first data transformation feature;

[0026] The first data transformation feature and the enhancement feature corresponding to the current enhancement level are concatenated to form a two-level concatenated feature. Based on the size of the first data transformation feature, the two-level concatenated feature is subjected to sliding window processing.

[0027] Each sliding window feature formed by the sliding window process is subjected to self-attention processing to form a second attention feature corresponding to each sliding window feature;

[0028] The mean of each second attention feature is calculated to form the output feature corresponding to the current reinforcement level.

[0029] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator fall recognition method, the step of performing feature mining on the target surveillance image data and outputting the target image data features includes:

[0030] Each frame of the target monitoring image data is convolved to obtain the first convolution feature corresponding to each frame of the monitoring image. For each frame of the target monitoring image data other than the first frame, the difference between the monitoring image and the previous frame is calculated to form the corresponding monitoring difference image. The monitoring difference image is convolved to obtain the corresponding second convolution feature.

[0031] The first convolutional feature corresponding to the first frame of the surveillance image is determined as the corresponding local data feature;

[0032] For each frame of the target monitoring image data other than the first frame, the first convolutional feature and the second convolutional feature corresponding to the monitoring image are fused to form the local data feature corresponding to the monitoring image.

[0033] The local data features corresponding to each frame of the monitoring image are combined to form the target image data features.

[0034] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator fall recognition method, the step of fusing the first and second convolutional features corresponding to each frame of the target monitoring image data (excluding the first frame) to form local data features corresponding to that monitoring image includes:

[0035] The first and second convolutional features corresponding to the surveillance image are concatenated to form the concatenated convolutional features corresponding to the surveillance image.

[0036] Self-attention mining is performed on the concatenated convolutional features to form self-attention features;

[0037] The self-attention features are compressed to form compressed features, and the compressed features and the first convolutional features are averaged to form corresponding local data features.

[0038] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator fall behavior recognition method, the step of performing feature mining on the target monitoring audio data and outputting the target audio data features includes:

[0039] For each audio frame in the target monitoring audio data, the corresponding spectrogram of the audio frame is determined, and the spectrogram is convolved to form the corresponding spectral convolution feature.

[0040] The spectral convolutional features corresponding to the first audio frame are determined as the corresponding local data features;

[0041] For each audio frame other than the first frame in the target monitoring audio data, the spectral convolutional features corresponding to the audio frame are fused with the spectral convolutional features corresponding to the previous audio frame to form corresponding local data features.

[0042] The local data features corresponding to each audio frame are combined to form the target audio data features.

[0043] In a preferred embodiment of this application, in the aforementioned deep learning-based elevator fall behavior recognition method, the step of fusing the spectral convolutional features corresponding to each audio frame other than the first frame in the target monitoring audio data with the spectral convolutional features corresponding to the previous audio frame to form corresponding local data features includes:

[0044] The spectral convolutional features corresponding to the audio frame and the spectral convolutional features corresponding to the previous audio frame are concatenated to form concatenated spectral convolutional features.

[0045] Self-attention mining is performed on the spliced ​​spectral convolution features to form corresponding self-attention spectral convolution features, and the self-attention spectral convolution features are compressed to form corresponding compressed spectral convolution features.

[0046] The compressed spectral convolution features and the spectral convolution features corresponding to the audio frames are averaged to form corresponding local data features.

[0047] Based on the above, this application also provides a deep learning-based method system for recognizing falls inside elevators, including:

[0048] Memory, used to store computer programs;

[0049] A processor connected to the memory is used to execute the computer program stored in the memory to implement the above-described deep learning-based method for recognizing falls inside elevators.

[0050] The elevator interior fall behavior recognition method and system provided in this application, based on deep learning, firstly, performs feature mining on the target monitoring image data to output target image data features; then, it performs feature mining on the target monitoring audio data to output target audio data features; next, based on the target audio data features, it performs feature enhancement on the target image data features to output enhanced image data features. During the feature enhancement process, it performs multi-level enhancement on multiple local data features in the target image data features based on multiple local data features in the target audio data features; finally, based on the enhanced image data features, it analyzes and outputs the target fall behavior recognition result, which serves as the alarm trigger condition for the target alarm device. Based on the above, on the one hand, since the target image data features are enhanced based on the target audio data features, the semantic information of the audio dimension is integrated into the target image data features, thereby compensating for other effective semantic information that is difficult to represent by the semantic information of the image dimension. On the other hand, since multi-level enhancement is performed, the enhancement accuracy can be relatively higher, that is, enhanced image data features with richer and more reliable semantic representation are obtained, which makes the reliability of the target fall behavior recognition result of the analysis output higher, thereby improving the problem of relatively low reliability of fall behavior recognition in the existing technology. In this way, the reliability of alarm action triggering can be further improved. Attached Figure Description

[0051] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.

[0052] Figure 1 The structural block diagram of the elevator interior fall behavior recognition system based on deep learning provided in the embodiments of this application is shown.

[0053] Figure 2 This is a flowchart illustrating the deep learning-based method for recognizing falls inside elevators, as provided in an embodiment of this application.

[0054] Figure 3 A schematic diagram illustrating feature enhancement provided in an embodiment of this application.

[0055] Figure 4 This is a schematic diagram illustrating the first feature enhancement provided in an embodiment of this application.

[0056] Figure 5 This is a schematic diagram of the sliding window processing provided in an embodiment of this application. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0058] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0059] like Figure 1 As shown in the figure, this application provides a deep learning-based elevator interior fall behavior recognition system. The deep learning-based elevator interior fall behavior recognition system may include a memory, a processor, and a deep learning-based elevator interior fall behavior recognition device.

[0060] Specifically, the memory and the processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, the memory and the processor can be electrically connected via one or more communication buses or signal lines. The deep learning-based elevator interior fall behavior recognition device includes at least one software functional module stored in the memory in the form of software or firmware. The processor is used to execute executable computer programs stored in the memory, such as the software functional modules and computer programs included in the deep learning-based elevator interior fall behavior recognition device, to implement the deep learning-based elevator interior fall behavior recognition method provided in this application embodiment.

[0061] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0062] Optionally, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0063] Optionally, the deep learning-based elevator fall recognition device may include:

[0064] The first module is used to acquire target monitoring image data and target monitoring audio data inside the target elevator;

[0065] The second module is used to perform feature mining on the target monitoring image data and output the target image data features.

[0066] The third module is used to perform feature mining on the target monitoring audio data and output the target audio data features.

[0067] The fourth module is used to perform feature enhancement on the target image data features based on the target audio data features and output enhanced image data features. In the feature enhancement process, multiple local data features in the target image data features are enhanced in multiple levels based on multiple local data features in the target audio data features.

[0068] The fifth module is used to analyze and output the target fall behavior recognition result based on the enhanced image data features. The target fall behavior recognition result is used to characterize whether there is a fall behavior inside the target elevator, and the target fall behavior recognition result serves as the alarm action trigger condition for the target alarm device.

[0069] Understandable. Figure 1 The structure shown is for illustrative purposes only. The deep learning-based elevator fall recognition system may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for exchanging information with other devices (such as monitoring equipment, alarm devices, etc.).

[0070] Combination Figure 2This application also provides a deep learning-based method for recognizing elevator falls inside the elevator, applicable to the aforementioned deep learning-based elevator fall behavior recognition system. The method steps defined in the relevant process of the deep learning-based elevator fall behavior recognition method can be implemented by the deep learning-based elevator fall behavior recognition system (hereinafter referred to as the recognition system). The following will discuss... Figure 2 The specific process shown will be explained in detail.

[0071] Step S110: Obtain target monitoring image data and target monitoring audio data inside the target elevator.

[0072] In this embodiment, the identification system can acquire target monitoring image data and target monitoring audio data inside the target elevator. For example, image acquisition devices and audio acquisition devices can be used to acquire images and audio data of the interior space of the target elevator, respectively, thereby forming corresponding target monitoring image data and target monitoring audio data. Alternatively, video acquisition devices can be used to acquire video data of the interior space of the target elevator, thereby forming corresponding video data. This video data can then be parsed to obtain the target monitoring image data and target monitoring audio data.

[0073] Step S120: Perform feature mining on the target monitoring image data and output the target image data features.

[0074] In this embodiment of the application, after obtaining the target surveillance image data, the recognition system can perform feature mining on the target surveillance image data and output target image data features. That is, it can mine semantic features in the target surveillance image data, such as potential semantic information related to behavior and actions, thereby obtaining target image data features.

[0075] Step S130: Perform feature mining on the target monitoring audio data and output the target audio data features.

[0076] In this embodiment of the application, after obtaining the target monitoring audio data, the recognition system can perform feature mining on the target monitoring audio data and output target audio data features. That is, it can mine semantic features in the target monitoring audio data, such as potential semantic information related to behavior and actions, thereby obtaining target audio data features.

[0077] Step S140: Based on the target audio data features, perform feature enhancement on the target image data features and output enhanced image data features.

[0078] In this embodiment, after mining the target image data features and the target audio data features, the recognition system can perform feature enhancement on the target image data features based on the target audio data features, and output enhanced image data features. That is, the semantic information in the target audio data features can be fused into the target image data features. Specifically, during feature enhancement, multiple local data features (which may correspond to audio frames) in the target audio data features are enhanced at multiple levels on multiple local data features (which may correspond to images) in the target image data features.

[0079] Step S150: Based on the enhanced image data features, analyze and output the target fall behavior recognition result.

[0080] In this embodiment, after obtaining the enhanced image data features, the recognition system can analyze and output a target fall behavior recognition result based on the enhanced image data features. The target fall behavior recognition result is used to characterize whether a fall occurs inside the target elevator, and it serves as the trigger condition for the alarm action of the target alarm device. For example, the enhanced image data features can be processed using a fully connected network, mapping them to a fully connected feature of size 1*1. Then, linear mapping or identity mapping can be performed on this fully connected feature to obtain a probability value, i.e., the probability that a fall occurs inside the target elevator. If this probability is greater than a preset value (e.g., 0.5, 0.6, etc.), a fall is confirmed, and the target alarm device can then take corresponding alarm actions, such as notifying relevant monitoring personnel.

[0081] Based on the above, on the one hand, since the target image data features are enhanced based on the target audio data features, the semantic information of the audio dimension is integrated into the target image data features, thereby compensating for other effective semantic information that is difficult to represent by the semantic information of the image dimension. On the other hand, since multi-level enhancement is performed, the enhancement accuracy can be relatively higher, that is, enhanced image data features with richer and more reliable semantic representation are obtained, which makes the reliability of the target fall behavior recognition result of the analysis output higher, thereby improving the problem of relatively low reliability of fall behavior recognition in the existing technology. In this way, the reliability of alarm action triggering can be further improved.

[0082] For example, in the target surveillance image data, if it contains the action of a person falling, it is obviously difficult to effectively determine whether the action is a fall or a normal squat (or other normal action based on certain needs). However, if the target surveillance audio data contains the sound of a fall and abnormal sounds caused by pain, it can help determine that it is a fall. In other words, fusing the semantic information of the image and audio dimensions allows for global analysis and recognition from both dimensions, thereby improving the reliability of the recognition.

[0083] Firstly, regarding step S120, it should be noted that the specific method for feature mining of the target monitoring image data is not limited and can be selected according to the actual situation.

[0084] For example, in an alternative implementation, the target surveillance image data can be convolved using a convolutional network to obtain target image data features.

[0085] For example, in another alternative implementation, considering that the falling behavior has a certain continuity in the image, in order to improve the semantic representation ability of the extracted target image data features, the above step S120 may further include steps S121, S122, S123 and S124, the specific contents of each step are as follows.

[0086] Step S121: Convolve each frame of the target monitoring image data to obtain the first convolution feature corresponding to each frame of the monitoring image; and for each frame of the target monitoring image data other than the first frame, perform difference calculation between the monitoring image and the previous frame of the monitoring image to form a corresponding monitoring difference image; and convolve the monitoring difference image to obtain the corresponding second convolution feature.

[0087] In this embodiment, each frame of the target surveillance image data can be convolved to obtain a first convolutional feature corresponding to each frame. Furthermore, for each frame of the target surveillance image data other than the first frame, a difference calculation is performed between the current frame and the previous frame (i.e., the difference in pixel values ​​between corresponding pixels in the two frames) to form a corresponding surveillance difference image. This surveillance difference image is then convolved to obtain a corresponding second convolutional feature. It should be noted that the specific parameters of the convolution kernels used to convolve the surveillance image and the surveillance difference image can be formed during training. Since the semantic information to be captured is different (e.g., the semantic information of the action itself represented by the surveillance image, and the semantic information of the action change represented by the surveillance difference image), the specific parameters of the convolution kernels can be different. However, the size of the convolution kernels can be the same, the corresponding convolution stride can be the same, and the sizes of the resulting first and second convolutional features can also be the same.

[0088] Step S122: Determine the first convolutional feature corresponding to the first frame of the monitoring image as the corresponding local data feature.

[0089] In this embodiment of the application, after obtaining the first convolutional feature, the first convolutional feature corresponding to the first frame monitoring image can be determined as the corresponding local data feature.

[0090] Step S123: For each frame of the target monitoring image data other than the first frame, fuse the first convolutional feature and the second convolutional feature corresponding to the monitoring image to form the local data feature corresponding to the monitoring image.

[0091] In this embodiment of the application, after obtaining the first convolutional feature and the second convolutional feature, for each frame of the target monitoring image data other than the first frame, the first convolutional feature and the second convolutional feature corresponding to the monitoring image are fused to form the local data feature corresponding to the monitoring image. That is, based on the semantic information of the action itself represented by the monitoring image, the semantic information of the action change between the monitoring image and the previous frame of the monitoring image is fused.

[0092] Step S124: Combine the local data features corresponding to each frame of the monitoring image to form the target image data features.

[0093] In this embodiment of the application, after obtaining the local data features corresponding to each frame of the monitoring image, the local data features corresponding to each frame of the monitoring image can be combined to form target image data features. For example, the target image data features can be a sequence formed by each local data feature, which can be arranged according to the time sequence between the corresponding monitoring images.

[0094] It is understood that in step S123 above, the specific method of fusing the first convolutional feature and the second convolutional feature corresponding to the monitoring image is not limited and can be selected according to actual needs. For example, in an alternative implementation, in order to achieve full feature fusion and avoid the loss of semantic information during the fusion process, step S123 above can further include the following specific implementation:

[0095] First, the first convolutional feature and the second convolutional feature corresponding to the monitoring image can be concatenated to form the concatenated convolutional feature corresponding to the monitoring image; for example, the concatenated convolutional feature can be (first convolutional feature, second convolutional feature);

[0096] Secondly, self-attention mining can be performed on the concatenated convolutional features to form self-attention features; specifically, this can be achieved through a corresponding self-attention network. In addition, through self-attention mining, the semantic information of the association between the first convolutional feature and the second convolutional feature in the concatenated convolutional features can be mined out.

[0097] Then, the self-attention features can be compressed (e.g., downsampled) to form compressed features, and the compressed features and the first convolutional features are averaged to form corresponding local data features; the compressed features and the first convolutional features have the same size, that is, the self-attention features can be compressed according to the size of the first convolutional features.

[0098] Secondly, it should be noted that the specific method for feature mining of the target monitoring audio data is not limited and can be selected according to actual needs.

[0099] For example, in an alternative implementation, the spectrograms corresponding to each audio frame in the target monitoring audio data can be convolved to obtain the corresponding target audio data features, thereby improving the efficiency of data mining.

[0100] For example, in another alternative implementation, considering that the sound of falling has a certain continuity in the actual dimension, in order to improve the semantic representation ability of the extracted target audio data features, the above step S130 can further include steps S131, S132, S133 and S134, the specific contents of each step are as follows.

[0101] Step S131: For each audio frame in the target monitoring audio data, determine the spectrogram corresponding to the audio frame, and perform convolution processing on the spectrogram to form the corresponding spectral convolution feature.

[0102] In this embodiment of the application, for each audio frame in the target monitoring audio data, the spectrum corresponding to the audio frame is determined (for example, the audio signal can be converted into a spectrum by Fourier transform), and the spectrum is convolved (which can be implemented by a convolutional network) to form the corresponding spectral convolution feature.

[0103] Step S132: Determine the spectral convolutional features corresponding to the first audio frame as the corresponding local data features.

[0104] In this embodiment of the application, after obtaining the spectral convolution features, the spectral convolution features corresponding to the first audio frame can be determined as the corresponding local data features.

[0105] Step S133: For each audio frame other than the first frame in the target monitoring audio data, the spectral convolutional feature corresponding to the audio frame is fused with the spectral convolutional feature corresponding to the previous audio frame to form the corresponding local data feature.

[0106] In this embodiment of the application, after obtaining the spectral convolution features, for each audio frame other than the first frame in the target monitoring audio data, the spectral convolution features corresponding to the audio frame and the spectral convolution features corresponding to the previous audio frame are fused to form corresponding local data features. In other words, the semantic information corresponding to the previous audio frame can be melted. In this way, the problem of semantic distortion caused by only focusing on the semantic information in the current audio frame can be avoided to a certain extent.

[0107] Step S134: Combine the local data features corresponding to each audio frame to form the target audio data features.

[0108] In this embodiment of the application, after obtaining the local data features corresponding to each audio frame, the local data features corresponding to each audio frame can be combined to form target audio data features. For example, the local data features can be sorted according to the temporal relationship between the corresponding audio frames to form a corresponding feature sequence, i.e., the target audio data features.

[0109] It is understood that in step S133 above, the specific method of fusing the spectral convolutional features corresponding to the audio frame and the spectral convolutional features corresponding to the previous audio frame is not limited and can be selected according to actual needs. For example, in an alternative implementation, in order to achieve full feature fusion and avoid the loss of semantic information during the fusion process, step S133 above can further include the following specific implementation:

[0110] First, the spectral convolutional features corresponding to the audio frame and the spectral convolutional features corresponding to the previous audio frame can be concatenated to form concatenated spectral convolutional features.

[0111] Secondly, the spliced ​​spectral convolutional features can be subjected to self-attention mining (for example, through a corresponding self-attention network) to form corresponding self-attention spectral convolutional features. Furthermore, the self-attention spectral convolutional features can be compressed (e.g., by downsampling) to form corresponding compressed spectral convolutional features. The size of the compressed spectral convolutional features can be the same as the size of the spectral convolutional features.

[0112] Then, the average of the compressed spectral convolution features and the spectral convolution features corresponding to the audio frame can be calculated to form the corresponding local data features.

[0113] Thirdly, regarding step S140, it should be noted that the specific method of feature enhancement for the target image data features is not limited and can be selected according to actual needs.

[0114] For example, in an alternative implementation, each local data feature in the target audio data features can be fused into the corresponding local data features in the target image data features, such as by performing a weighted summation calculation. Then, the results of the weighted summation calculation can be spliced ​​or averaged to obtain enhanced image data features.

[0115] For example, in another alternative implementation, considering that there is a certain semantic correlation between the local data features corresponding to different audio frames and monitoring images, in order to further improve the reliability of feature enhancement, the above step S140 may further include the following steps S141, S142 and S143, the specific contents of each step are as follows.

[0116] Step S141: Based on the continuity between corresponding audio frames and / or the similarity between local data features, combine multiple local data features included in the target audio data features to form multiple first data feature combinations.

[0117] In this embodiment, multiple local data features included in the target audio data features can be combined based on the continuity between corresponding audio frames and / or the similarity between local data features to form multiple first data feature combinations. Each first data feature combination includes at least one local data feature, and when multiple local data features are included, the audio frames corresponding to these multiple local data features are continuous. For example, considering that each audio frame generally contains 20ms to 40ms of audio data, the local data features corresponding to 10-20 consecutive audio frames can be determined as a first data feature combination. Alternatively, the similarity (e.g., cosine similarity) between the local data features corresponding to two adjacent audio frames can be calculated sequentially. If the similarity is less than a preset similarity (e.g., 0.4, 0.5, etc.), the local data features corresponding to the two audio frames are placed in two first data feature combinations; if the similarity is greater than or equal to the preset similarity, the local data features corresponding to the two audio frames are placed in one first data feature combination.

[0118] Step S142: Based on the multiple first data feature combinations, the multiple local data features included in the target image data features are combined to form multiple second data feature combinations.

[0119] In this embodiment, after forming the plurality of first data feature combinations, multiple local data features included in the target image data features can be combined based on the plurality of first data feature combinations to form multiple second data feature combinations. There is a one-to-one correspondence between the second data feature combinations and the first data feature combinations, and the included local data features have temporal consistency with respect to the corresponding audio frames and monitoring images. For example, if the first first data feature combination includes local data features corresponding to the first and second audio frames, then the first second data feature combination includes local data features corresponding to the first and second monitoring images; if the second first data feature combination includes local data features corresponding to the third, fourth, and fifth audio frames, then the second second data feature combination includes local data features corresponding to the third, fourth, and fifth monitoring images. Furthermore, the first audio frame and the first monitoring image have the same timestamp, and the second audio frame and the second monitoring image have the same timestamp.

[0120] Step S143: Based on the multiple first data feature combinations, perform feature enhancement on the multiple second data feature combinations and output enhanced image data features.

[0121] In this embodiment of the application, after forming the plurality of first data feature combinations and the plurality of second data feature combinations, feature enhancement can be performed on the plurality of second data feature combinations based on the plurality of first data feature combinations to output enhanced image data features. For example, each of the first data feature combinations can be sequentially fused into the corresponding second data feature combination.

[0122] It is understood that in step S143 above, the specific method of feature enhancement for the multiple combinations of second data features is not limited and can be selected according to actual needs. For example, in an alternative implementation, in order to fully realize the fusion of semantic information in the audio and image dimensions during the feature enhancement process at multiple levels, and to realize the fusion of semantic information at different times within the image dimension, so as to further improve the semantic representation capability of the determined enhanced image data features, step S143 above can further include steps S143a, S143b, S143c, and S143d, the specific contents of each step of which are as follows (in conjunction with...). Figure 3 ).

[0123] Step S143a: Determine the corresponding multiple enhancement levels based on the combination of the multiple second data features.

[0124] In this embodiment of the application, multiple enhancement levels can be determined based on the multiple combinations of second data features. There is a one-to-one correspondence between the multiple enhancement levels and the multiple combinations of second data features. For example, when there are two combinations of second data features, two enhancement levels are determined; when there are three combinations of second data features, three enhancement levels are determined.

[0125] Step S143b: In each enhancement level, based on the output features corresponding to the previous enhancement level, the second data feature combination corresponding to the current enhancement level is enhanced with the first feature, and the enhancement feature corresponding to the current enhancement level is output.

[0126] In this embodiment, after determining multiple enhancement levels, in each enhancement level, based on the output features corresponding to the previous enhancement level, the second data feature combination corresponding to the current enhancement level is enhanced with a first feature, and the enhanced feature corresponding to the current enhancement level is output. The output feature corresponding to the previous enhancement level of the first enhancement level is a feature formed by learning from the corresponding sample data (in the initial state of learning, this output feature can be a random vector, or it can be a vector of all zeros; thus, it can be continuously learned and updated during the training process of the corresponding neural network model).

[0127] Step S143c: In each enhancement level, based on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level, perform second feature enhancement on the enhancement feature corresponding to the current enhancement level, and output the output feature corresponding to the current enhancement level.

[0128] In this embodiment, after completing the first feature enhancement, in each enhancement level, a second feature enhancement can be performed on the enhancement feature corresponding to the current enhancement level based on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level, and the output feature corresponding to the current enhancement level can be output. That is, the fusion between the first data feature combinations and the fusion between the first data feature combination and the second data feature combination can be achieved through the first feature enhancement and the second feature enhancement, respectively.

[0129] Step S143d: Based on the output features of this last enhancement level, determine the enhanced image data features.

[0130] In this embodiment, after feature enhancement at the last enhancement level is completed, the enhanced image data features can be determined based on the output features of this last enhancement level. For example, the output features of the last enhancement level can be used as the enhanced image data features.

[0131] It is understood that in step S143b above, the specific method of strengthening the first feature of the second data feature combination corresponding to the current strengthening level is not limited and can be selected according to actual needs. For example, in an alternative implementation, in order to achieve full fusion of semantic information and capture more related semantic information, step S143b above can further include the following specific implementation content (in conjunction with...). Figure 4 As shown):

[0132] First, the output features corresponding to the previous enhancement level and the second data features corresponding to the current enhancement level can be combined and concatenated to form a two-level concatenated feature. Then, based on the size of the output features corresponding to the previous enhancement level, a sliding window process is applied to this two-level concatenated feature (combined with...). Figure 5As shown in the figure, the size of each sliding window feature formed by the sliding window processing is equal to the size of the output feature corresponding to the previous enhancement level. In this way, for each sliding window feature, the proportion of semantic features of the combination of the output feature corresponding to the previous enhancement level and the second data feature corresponding to the current enhancement level gradually changes, so that different related semantic information can be captured in the subsequent fusion. In addition, it should be noted that before splicing, the mean of each local data feature in the second data feature combination can be calculated to obtain the corresponding mean feature. Then, the mean feature is spliced ​​with the output feature corresponding to the previous enhancement level.

[0133] Secondly, each sliding window feature formed by the sliding window process can be processed with self-attention to form a first attention feature corresponding to each sliding window feature; for example, this can be achieved through a corresponding attention network, and the specific processing steps will not be described in detail here.

[0134] Then, the mean of each of the first attention features can be calculated to form the reinforcement feature corresponding to the current reinforcement level. For example, the mean of each of the first attention features can be directly used as the reinforcement feature corresponding to the current reinforcement level, or the mean of each local data feature in the combination of the mean feature and the second data feature corresponding to the current reinforcement level can be calculated to obtain the reinforcement feature corresponding to the current reinforcement level.

[0135] It is understood that in step S143c above, the specific method of performing second feature enhancement on the enhancement feature corresponding to the current enhancement level is not limited and can be selected according to actual needs. For example, in an alternative implementation, in order to achieve full fusion of semantic information and capture more related semantic information, step S143c above may include:

[0136] First, a feature space transformation can be performed on the first data feature combination corresponding to the second data feature combination of the current enhancement level to form the corresponding first data transformation feature. It should be noted that the second data feature combination corresponds to semantic information in the image dimension, while the first data feature combination corresponds to semantic information in the audio dimension. Since they are in different dimensions, to ensure the reliability of subsequent fusion, the first data feature combination can first undergo a feature space transformation so that it can be transformed into a feature space similar to that of the second data feature combination. Specifically, the local data features in the first data feature combination can be transformed first. The mean is calculated to obtain the corresponding mean feature. Then, a fully connected network can be used to implement the linear mapping. The fully connected network includes at least one linear unit (when multiple linear units are included, they can be cascaded, i.e., the output of the first linear unit is connected to the input of the second linear unit). Each linear unit can have a linear mapping function, such as y = Wx + b, where y represents the output (e.g., the first data transformation feature), x represents the input (e.g., the mean feature), W represents the weight matrix, and b represents the bias parameter. The weight matrix and the bias parameter are formed during the training of the corresponding neural network model.

[0137] Secondly, the first data transformation feature and the enhancement feature corresponding to the current enhancement level can be spliced ​​together to form a two-level spliced ​​feature. Based on the size of the first data transformation feature, a sliding window process is applied to the two-level spliced ​​feature. Thus, the size of each sliding window feature formed by the sliding window process is equal to the size of the first data transformation feature.

[0138] Then, self-attention processing can be performed on each sliding window feature formed by the sliding window processing to form a second attention feature corresponding to each sliding window feature; for example, this can be achieved through a corresponding attention network, and the specific processing steps will not be described in detail here.

[0139] Finally, the mean of each second attention feature can be calculated to form the output feature corresponding to the current enhancement level. For example, the mean of each second attention feature can be directly used as the enhancement feature corresponding to the current enhancement level, or the mean of the mean feature can be calculated with the enhancement feature corresponding to the current enhancement level to obtain the output feature corresponding to the current enhancement level.

[0140] In summary, the deep learning-based elevator interior fall behavior recognition method and system provided in this application firstly performs feature mining on the target monitoring image data to output target image data features; then, it performs feature mining on the target monitoring audio data to output target audio data features; next, based on the target audio data features, it performs feature enhancement on the target image data features to output enhanced image data features. During the feature enhancement process, it enhances multiple local data features in the target image data features at multiple levels based on multiple local data features in the target audio data features; finally, based on the enhanced image data features, it analyzes and outputs the target fall behavior recognition result, which serves as the alarm trigger condition for the target alarm device. Based on the above, on the one hand, since the target image data features are enhanced based on the target audio data features, the semantic information of the audio dimension is integrated into the target image data features, thereby compensating for other effective semantic information that is difficult to represent by the semantic information of the image dimension. On the other hand, since multi-level enhancement is performed, the enhancement accuracy can be relatively higher, that is, enhanced image data features with richer and more reliable semantic representation are obtained, which makes the reliability of the target fall behavior recognition result of the analysis output higher, thereby improving the problem of relatively low reliability of fall behavior recognition in the existing technology. In this way, the reliability of alarm action triggering can be further improved.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0142] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0143] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0144] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A deep learning-based method for recognizing falls inside elevators, characterized in that, include: Acquire target monitoring image data and target monitoring audio data inside the target elevator; Feature mining is performed on the target surveillance image data to output the target image data features; Perform feature mining on the target monitoring audio data and output the target audio data features; Based on the continuity between corresponding audio frames and / or the similarity between local data features, multiple local data features included in the target audio data features are combined to form multiple first data feature combinations. Each first data feature combination includes at least one local data feature, and when multiple local data features are included, the audio frames corresponding to these multiple local data features are continuous. Based on the multiple first data feature combinations, multiple local data features included in the target image data features are combined to form multiple second data feature combinations. The second data feature combinations and the first data feature combinations have a one-to-one correspondence, and the included local data features have temporal consistency with respect to the corresponding audio frames and monitoring images. Based on the multiple first data feature combinations, the multiple second data feature combinations are enhanced to output enhanced image data features. Based on the enhanced image data features, the target fall behavior recognition result is analyzed and output. The target fall behavior recognition result is used to characterize whether there is a fall behavior inside the target elevator, and the target fall behavior recognition result serves as the alarm action trigger condition for the target alarm device.

2. The method for recognizing falling behavior inside an elevator based on deep learning according to claim 1, characterized in that, The step of enhancing the multiple second data feature combinations based on the multiple first data feature combinations and outputting enhanced image data features includes: Multiple enhancement levels are determined based on the multiple combinations of second data features, wherein there is a one-to-one correspondence between the multiple enhancement levels and the multiple combinations of second data features; In each enhancement level, based on the output features corresponding to the previous enhancement level, the second data feature combination corresponding to the current enhancement level is enhanced by the first feature, and the enhancement feature corresponding to the current enhancement level is output. The output feature corresponding to the previous enhancement level of the first enhancement level is a feature formed by learning the corresponding sample data. In each enhancement level, based on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level, the enhancement feature corresponding to the current enhancement level is enhanced by the second feature, and the output feature corresponding to the current enhancement level is output. Based on the output features of this last enhancement level, the enhanced image data features are determined.

3. The method for recognizing falling behavior inside an elevator based on deep learning according to claim 2, characterized in that, The step of performing first feature enhancement on the second data feature combination corresponding to the current enhancement level based on the output features corresponding to the previous enhancement level in each enhancement level, and outputting the enhanced features corresponding to the current enhancement level, includes: The output feature corresponding to the previous enhancement level and the second data feature corresponding to the current enhancement level are combined and spliced ​​to form a two-level spliced ​​feature. Based on the size of the output feature corresponding to the previous enhancement level, the two-level spliced ​​feature is subjected to sliding window processing. Each sliding window feature formed by the sliding window process is subjected to self-attention processing to form the first attention feature corresponding to each sliding window feature; The mean of each of the first attention features is calculated to form the reinforcement feature corresponding to the current reinforcement level.

4. The method for recognizing falling behavior inside an elevator based on deep learning according to claim 2, characterized in that, The step of performing second feature enhancement on the enhancement features corresponding to the current enhancement level based on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level in each enhancement level, and outputting the output features corresponding to the current enhancement level, includes: Perform feature space transformation on the first data feature combination corresponding to the second data feature combination corresponding to the current enhancement level to form the corresponding first data transformation feature; The first data transformation feature and the enhancement feature corresponding to the current enhancement level are concatenated to form a two-level concatenated feature. Based on the size of the first data transformation feature, the two-level concatenated feature is subjected to sliding window processing. Each sliding window feature formed by the sliding window process is subjected to self-attention processing to form a second attention feature corresponding to each sliding window feature; The mean of each second attention feature is calculated to form the output feature corresponding to the current reinforcement level.

5. The method for recognizing falling behavior inside an elevator based on deep learning according to any one of claims 1-4, characterized in that, The step of performing feature mining on the target surveillance image data and outputting the target image data features includes: Each frame of the target monitoring image data is convolved to obtain the first convolution feature corresponding to each frame of the monitoring image. For each frame of the target monitoring image data other than the first frame, the difference between the monitoring image and the previous frame is calculated to form the corresponding monitoring difference image. The monitoring difference image is convolved to obtain the corresponding second convolution feature. The first convolutional feature corresponding to the first frame of the surveillance image is determined as the corresponding local data feature; For each frame of the target monitoring image data other than the first frame, the first convolutional feature and the second convolutional feature corresponding to the monitoring image are fused to form the local data feature corresponding to the monitoring image. The local data features corresponding to each frame of the monitoring image are combined to form the target image data features.

6. The method for recognizing falling behavior inside an elevator based on deep learning according to claim 5, characterized in that, The step of fusing the first and second convolutional features corresponding to each frame of the target surveillance image data (excluding the first frame) to form the local data features corresponding to that surveillance image includes: The first and second convolutional features corresponding to the surveillance image are concatenated to form the concatenated convolutional features corresponding to the surveillance image. Self-attention mining is performed on the concatenated convolutional features to form self-attention features; The self-attention features are compressed to form compressed features, and the compressed features and the first convolutional features are averaged to form corresponding local data features.

7. The method for recognizing falling behavior inside an elevator based on deep learning according to any one of claims 1-4, characterized in that, The step of performing feature mining on the target monitoring audio data and outputting the target audio data features includes: For each audio frame in the target monitoring audio data, the corresponding spectrogram of the audio frame is determined, and the spectrogram is convolved to form the corresponding spectral convolution feature. The spectral convolutional features corresponding to the first audio frame are determined as the corresponding local data features; For each audio frame other than the first frame in the target monitoring audio data, the spectral convolutional features corresponding to the audio frame are fused with the spectral convolutional features corresponding to the previous audio frame to form corresponding local data features. The local data features corresponding to each audio frame are combined to form the target audio data features.

8. The method for recognizing falling behavior inside an elevator based on deep learning according to claim 7, characterized in that, The step of fusing the spectral convolutional features corresponding to each audio frame other than the first frame in the target monitoring audio data with the spectral convolutional features corresponding to the previous audio frame to form corresponding local data features includes: The spectral convolutional features corresponding to the audio frame and the spectral convolutional features corresponding to the previous audio frame are concatenated to form concatenated spectral convolutional features. Self-attention mining is performed on the spliced ​​spectral convolution features to form corresponding self-attention spectral convolution features, and the self-attention spectral convolution features are compressed to form corresponding compressed spectral convolution features. The compressed spectral convolution features and the spectral convolution features corresponding to the audio frames are averaged to form corresponding local data features.

9. A deep learning-based elevator interior fall recognition system, characterized in that, include: Memory, used to store computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the deep learning-based elevator fall behavior recognition method according to any one of claims 1-8.