Compartment event detection method and device, computer equipment and storage medium

By obtaining images in the train compartment and applying emotional and behavior recognition strategies, and detecting emergencies in the car with audio data, the problem of difficult time handling emergencies in the train compartment is solved, and the effect of timely handling emergencies is achieved.

CN120388406APending Publication Date: 2025-07-29CRRC TANGSHAN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510440893.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

If an emergency occurs in the train car, such as a sudden illness or a fierce dispute, it will be difficult for technicians to detect and deal with it in time, resulting in the worsening of the incident.

Method used

By obtaining the target image in the train car, using emotion recognition and behavior recognition strategies, we can detect whether emergencies have occurred in the car, and verify them in combination with audio data.

Benefits of technology

Timely determine emergencies in the car to avoid further deterioration of the incident and ensure that technicians can handle it quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388406A_ABST
    Figure CN120388406A_ABST
Patent Text Reader

Abstract

The invention relates to a compartment event detection method and device, equipment and a storage medium, and the method comprises the steps: obtaining a target image in a train compartment, and carrying out the recognition of the target image according to an emotion recognition strategy, and obtaining an emotion recognition result; according to a behavior identification strategy, identifying the target image to obtain a behavior identification result; and the behavior recognition result and the emotion recognition result are sent to a train management platform, so that the train management platform detects whether an emergency occurs in the compartment or not according to the emotion recognition result and the behavior recognition result. Visibly, according to the embodiment of the invention, the target image in the train compartment can be identified to determine whether the emergency occurs in the compartment or not, so that the emergency in the compartment can be determined in time, technical personnel can deal with the emergency in time, and the problem that the emergency is further worsened is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vehicle control, and in particular, to a method, device, computer device, and storage medium for detecting carriage events. Background Art

[0002] In the modern transportation system, trains have become an important choice for people's long-distance travel due to their high efficiency, convenience, and large passenger capacity. As the main activity space for passengers, train carriages are designed to be more user-friendly. The spacious and bright interior space is equipped with comfortable seats, convenient luggage storage facilities, and modern ventilation and lighting systems, aiming to provide passengers with a good travel experience.

[0003] However, even though the facilities and equipment in train carriages are continuously iterated and upgraded, there are still potential problems. Limited by the number of train technicians, when emergencies occur in the carriage, such as sudden illness of passengers or intense disputes during the journey, it is often difficult for technicians to detect and intervene in a timely manner. This can easily lead to the continuous fermentation of emergencies that could have been resolved in a timely manner and develop into a worse situation. Therefore, there is an urgent need for a method for detecting carriage events. Summary of the Invention

[0004] In an embodiment of the present application, there is provided a method, device, computer device, and storage medium for detecting carriage events. The method can identify a target image in a train carriage to automatically detect whether an emergency has occurred in the carriage. In this way, emergencies in the carriage can be determined in a timely manner, enabling technicians to handle them in a timely manner and avoiding the further deterioration of emergencies.

[0005] In a first aspect of the embodiments of the present application, there is provided a method for detecting carriage events, the method comprising:

[0006] Obtain a target image in a train carriage;

[0007] Identify the target image according to an emotion recognition strategy to obtain an emotion recognition result;

[0008] Identify the target image according to a behavior recognition strategy to obtain a behavior recognition result;

[0009] Send the behavior recognition result and the emotion recognition result to a train management platform, so that the train management platform detects whether an emergency has occurred in the carriage according to the emotion recognition result and the behavior recognition result.

[0010] Optionally, the identifying the target image according to an emotion recognition strategy to obtain an emotion recognition result includes:

[0011] Identify a reference face image in the target image;

[0012] Preprocess the reference face image to obtain a target face image;

[0013] Call a pre-trained emotion recognition model to process the target face image and obtain an emotion recognition result.

[0014] Optionally, determining the reference face image in the target image includes:

[0015] Detect the face in the target image to obtain multiple rectangular frames;

[0016] Determine a target face detection frame within the multiple rectangular frames;

[0017] Determine the image included in the target face detection frame as the reference face image.

[0018] Optionally, according to the behavior recognition strategy, recognizing the target image to obtain a behavior recognition result includes:

[0019] Determine the reference human body image in the target image;

[0020] Preprocess the reference human body image to obtain a target human body image;

[0021] Call a pre-trained behavior recognition model to recognize the target human body image and obtain a behavior recognition result.

[0022] Optionally, determining the reference human body image in the target image includes:

[0023] Preprocess the target image to obtain a reference image;

[0024] Call a target detection model to determine the target area of the object in the reference image;

[0025] Determine the image corresponding to the target area in the reference image as the reference human body image in the target image.

[0026] Optionally, the method further includes:

[0027] Obtain the audio data in the train carriage;

[0028] Detect whether the audio data meets a first preset condition;

[0029] When the audio data meets the first preset condition, determine that an emergency has occurred in the carriage.

[0030] Optionally, detecting whether an emergency has occurred in the carriage according to the emotion recognition result and the behavior recognition result includes:

[0031] Determine a fusion recognition result according to the emotion recognition result and the behavior recognition result;

[0032] When the fusion recognition result meets a second preset condition, determine that an emergency has occurred inside the carriage;

[0033] When the fusion recognition result does not meet the second preset condition, determine that no emergency has occurred inside the carriage.

[0034] Optionally, the emotion recognition result includes a first category and a corresponding first confidence level, the behavior recognition result includes a second category and a corresponding second confidence level, and the determining the fusion recognition result according to the emotion recognition result and the behavior recognition result includes:

[0035] Determine a first weight corresponding to the first category and a second weight corresponding to the second category;

[0036] Multiply the first weight and the first confidence level to obtain a first value;

[0037] Multiply the second weight and the second confidence level to obtain a second value;

[0038] Add the first value and the second value to obtain the fusion recognition result.

[0039] In a second aspect of the embodiments of the present application, the embodiments of the present application provide a carriage event detection device, including:

[0040] An acquisition unit, configured to acquire a target image inside a train carriage;

[0041] A first recognition unit, configured to recognize the target image according to an emotion recognition strategy to obtain an emotion recognition result;

[0042] A second recognition unit, configured to recognize the target image according to a behavior recognition strategy to obtain a behavior recognition result;

[0043] A sending unit, configured to send the behavior recognition result and the emotion recognition result to a train management platform, so that the train management platform detects whether an emergency has occurred inside the carriage according to the emotion recognition result and the behavior recognition result.

[0044] Optionally, the first recognition unit is configured to:

[0045] Recognize a reference face image in the target image;

[0046] Preprocess the reference face image to obtain a target face image;

[0047] Call a pre-trained emotion recognition model to process the target face image and obtain an emotion recognition result.

[0048] Optionally, the first recognition unit is configured to:

[0049] Detect faces in the target image to obtain multiple rectangular frames;

[0050] Determine a target face detection frame within the multiple rectangular frames;

[0051] Determine the image included in the target face detection frame as the reference face image.

[0052] Optionally, the second recognition unit is configured to:

[0053] Determine a reference human body image in the target image;

[0054] Preprocess the reference human body image to obtain a target human body image;

[0055] Call a pre-trained behavior recognition model to recognize the target human body image and obtain a behavior recognition result.

[0056] Optionally, the second recognition unit is configured to:

[0057] Preprocess the target image to obtain a reference image;

[0058] Call a target detection model to determine the target area of the object in the reference image;

[0059] Determine the image corresponding to the target area in the reference image as the reference human body image in the target image.

[0060] Optionally, the device further includes a detection unit, and the detection unit is configured to:

[0061] Obtain audio data in the train carriage;

[0062] Detect whether the audio data meets a first preset condition;

[0063] When the audio data meets the first preset condition, determine that an emergency has occurred in the carriage.

[0064] Optionally, the sending unit is configured to:

[0065] Determine a fusion recognition result according to the emotion recognition result and the behavior recognition result;

[0066] When the fusion recognition result meets a second preset condition, determine that an emergency has occurred in the carriage;

[0067] When the fusion recognition result does not meet the second preset condition, it is determined that no emergency occurs in the carriage.

[0068] Optionally, the emotion recognition result includes a first category and a corresponding first confidence level. The sending unit is configured to:

[0069] Determine a first weight corresponding to the first category and a second weight corresponding to the second category;

[0070] Multiply the first weight and the first confidence level to obtain a first value;

[0071] Multiply the second weight and the second confidence level to obtain a second value;

[0072] Add the first value and the second value to obtain a fusion recognition result.

[0073] In a third aspect of the embodiments of the present application, a computer device is provided, including: a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above methods are implemented.

[0074] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program, when executed by a processor, implements the steps of any one of the above methods.

[0075] In the embodiments of the present application, a target image in a train carriage is acquired, and according to an emotion recognition strategy, the target image is recognized to obtain an emotion recognition result; according to a behavior recognition strategy, the target image is recognized to obtain a behavior recognition result; the behavior recognition result and the emotion recognition result are sent to a train management platform, so that the train management platform detects whether an emergency occurs in the carriage according to the emotion recognition result and the behavior recognition result. It can be seen that the embodiments of the present application can recognize the target image in the train carriage to automatically detect whether an emergency occurs in the carriage. In this way, emergencies in the carriage can be determined in a timely manner, and then technical personnel can handle them in a timely manner, avoiding the problem of further deterioration of emergencies. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0077] Figure 1 It is a schematic structural diagram of a carriage event detection method provided by an embodiment of the present application;

[0078] Figure 2 Flow chart of a carriage event detection method provided by an embodiment of the present application;

[0079] Figure 3 Flow chart of an emotion recognition result determination method provided by an embodiment of the present application;

[0080] Figure 4 Flow chart of a reference face image determination method provided by an embodiment of the present application;

[0081] Figure 5 Flow chart of a behavior recognition result determination method provided by an embodiment of the present application;

[0082] Figure 6 Flow chart of a reference human body image determination method provided by an embodiment of the present application;

[0083] Figure 7 Flow chart of an emergency verification method provided by an embodiment of the present application;

[0084] Figure 8 Flow chart of a carriage event detection method provided by an embodiment of the present application;

[0085] Figure 9 Schematic structural diagram of a carriage event detection device provided by an embodiment of the present application;

[0086] Figure 10 Schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0087] In the modern transportation system, trains have become an important choice for people's long-distance travel due to their high efficiency, convenience, and large passenger capacity. As the main activity space for passengers, the design of train carriages has become increasingly user-friendly. The spacious and bright interior space is equipped with comfortable seats, convenient luggage storage facilities, and modern ventilation and lighting systems, aiming to provide passengers with a good travel experience. However, even though the facilities and equipment in train carriages are continuously iterated and upgraded, there are still potential problems. Limited by the number of train technicians, when emergencies occur in the carriage, such as sudden illness of passengers or intense disputes during the journey, it is often difficult for technicians to detect and intervene in a timely manner. This can easily lead to the continuous fermentation of emergencies that could have been resolved in a timely manner, and the situation may develop into a worse state due to the lack of rapid response.

[0088] In view of the above problems, in the embodiments of the present application, a carriage event detection method is provided. First, a target image in the train carriage is obtained, and according to an emotion recognition strategy, the target image is recognized to obtain an emotion recognition result; according to a behavior recognition strategy, the target image is recognized to obtain a behavior recognition result; the behavior recognition result and the emotion recognition result are sent to a train management platform, so that the train management platform can detect whether an emergency occurs in the carriage according to the emotion recognition result and the behavior recognition result. It can be seen that the embodiments of the present application can recognize the target image in the train carriage to automatically detect whether an emergency occurs in the carriage. In this way, the emergency in the carriage can be determined in time, and then technicians can handle it in time, avoiding the problem of further deterioration of the emergency.

[0089] The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0090] In order to make the technical solutions and advantages in the embodiments of the present application clearer and more understandable, the following further details the exemplary embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0091] The following briefly describes the structural schematic diagram of the carriage event detection method provided by the embodiments of the present application:

[0092] As Figure 1As described above, multiple cameras are installed inside the train carriages. These cameras are connected to a power source through a power adapter to obtain energy from the power source. After the cameras capture images, they send these images to the on-vehicle server. The specific steps are as follows: The network adapter converts the images captured by the cameras into digital signals, encapsulates them according to the network communication protocol to form data frames, and converts the data frames into electrical signals suitable for transmission on the network transmission medium, and then sends them out through the network cable. When the electrical signals reach the port of the switch, the switch receives the data frames, checks the source MAC address of the data frames, associates it with the port that received the frame, and adds it to its own MAC address table. Then, it checks the destination MAC address of the data frame and looks up the corresponding port in the MAC address table. If a matching port is found, the switch forwards the data frame from the corresponding port and sends it to the corresponding device in the LAN; if no matching port is found, the switch sends the data frame as a broadcast frame to all other ports except the receiving port so that the target device in the LAN can receive it. Finally, the data frames are transmitted within the LAN through the network transmission medium, and then the data frames are transmitted to the on-vehicle server.

[0093] When the on-vehicle server receives these data frames, it analyzes the data frames to obtain the images included therein. Then, the on-vehicle server performs recognition on each image to obtain the emotion recognition result and the behavior recognition result, and sends the emotion recognition result and the behavior recognition result to the train management platform so that the user can receive the emotion recognition result and the behavior recognition result through the train management platform to determine whether a sudden abnormal event has occurred inside the carriage.

[0094] Among them, the on-vehicle server is a computer device installed on the vehicle, which is used to process, store and transmit various vehicle-related data and provide services for the vehicle electronic system and users. The train management platform is a platform with multiple functions. For example, the train management platform has an alarm function.

[0095] Please refer to Figure 2 , in the following embodiments, the above on-vehicle server is used as the execution entity, and the method provided in the embodiments of the present application is applied to the above on-vehicle server. The on-vehicle server can detect the target image and send the detection result to the train management platform, so that the train management platform can determine whether an emergency has occurred inside the train carriage according to the detection result. The carriage event detection method provided in the embodiments of the present application includes the following steps 201-step 204:

[0096] Step 201, obtain the target image captured inside the train carriage.

[0097] Among them, there is at least one object on the target image, and each object corresponds to a person.

[0098] In this step, in order to ensure that the facial information and limb information of all passengers in the carriage can be collected, cameras are installed at three positions in the front, middle, and rear on both sides of the train carriage, for a total of six cameras. After that, when the cameras in the same carriage capture images, these images can be sent to the on-vehicle server. When the on-vehicle server receives these images, it can detect the images to determine whether there are people in the images. When a person exists in an image, the image is determined as a target image and subsequent steps are performed on it.

[0099] In addition, after sending the images captured by the cameras to the on-vehicle server, for accurate recognition, the images corresponding to the same carriage can also be fused to obtain the image corresponding to the entire carriage, and subsequent processing can be performed based on the fused image.

[0100] Step 202: According to the emotion recognition strategy, recognize the target image to obtain the emotion recognition result.

[0101] Among them, the emotion recognition result can be the emotion recognition result of each object, the overall emotion recognition result reflected by all objects, or the emotion change trend. The emotion recognition result can be emotions such as happiness and sadness.

[0102] In this step, obtain the pre-stored emotion recognition strategy, and then use this emotion recognition strategy to recognize the target image to obtain the emotion recognition result.

[0103] Step 203: According to the behavior recognition strategy, recognize the target image to obtain the behavior recognition result.

[0104] Among them, the behavior recognition result can be the behavior recognition result of each object, the overall behavior recognition result reflected by all objects, or the behavior change trend. The behavior recognition result can be violent behavior or non-violent behavior.

[0105] In this step, obtain the pre-stored behavior recognition strategy, and then use this behavior recognition strategy to recognize the target image to obtain the behavior recognition result.

[0106] It should be noted that in this application, step 203 can also be executed first, followed by step 202, or step 202 and step 203 can be executed simultaneously. This application does not limit this.

[0107] Step 204: Send the behavior recognition result and the emotion recognition result to the train management platform so that the train management platform can detect whether an emergency occurs in the carriage according to the emotion recognition result and the behavior recognition result.

[0108] In this step, since the in-vehicle server may process multiple target images and thus obtain multiple sets of results, in order to enable the management platform to clearly understand the results corresponding to each target image, when sending the results to the train management platform, the corresponding target images can also be sent to the train management platform. Alternatively, the in-vehicle server first maps the results into the target images and then sends the target images including the results to the train management platform.

[0109] For example, when the emotion recognition result is the emotion recognition result of each object and the behavior recognition result is the behavior recognition result of each object, the emotion recognition result and the behavior recognition result can be mapped into the corresponding objects to obtain the mapped target images, and then these images are sent to the train management platform. For the mapped target images, each object may correspond to the emotion recognition result and the behavior recognition result, so that the train management platform can more clearly understand the situation in the corresponding carriage.

[0110] After that, after the train management platform receives the emotion recognition result and the behavior recognition result, the train management platform can also display the emotion recognition result and the behavior recognition result so that technicians can analyze these results to determine whether an emergency has occurred in the carriage. Alternatively, the train management platform detects these results to determine whether an emergency has occurred in the vehicle.

[0111] In the embodiment of the present application, a target image in the train carriage is obtained, and according to the emotion recognition strategy, the target image is recognized to obtain an emotion recognition result; according to the behavior recognition strategy, the target image is recognized to obtain a behavior recognition result; the behavior recognition result and the emotion recognition result are sent to the train management platform so that the train management platform can detect whether an emergency has occurred in the carriage according to the emotion recognition result and the behavior recognition result. It can be seen that the embodiment of the present application can recognize the target image in the train carriage to automatically determine whether an emergency has occurred in the carriage, so that the emergency in the carriage can be determined in time, and then technicians can handle it in time, avoiding the problem of further deterioration of the emergency.

[0112] In the embodiment of the present application, the in-vehicle server needs to first recognize a reference face image in the target image, and then preprocess the reference person image to obtain an image that meets the input requirements of the emotion recognition model. Then, the pre-trained emotion recognition model is called to process the target face image to obtain an emotion recognition result. Therefore, the embodiment of the present application provides a method for determining an emotion recognition result, and this method is as Figure 3 shown, which is a further limitation of step 202. The specific steps include:

[0113] Step 301, recognize the reference face image in the target image.

[0114] In this step, a face recognition algorithm can be used to identify the face region in the target image, and the image corresponding to the face region is determined as the face image.

[0115] Step 302: Preprocess the reference face image to obtain the target face image.

[0116] Among them, the preprocessing is to adjust the face image to obtain an image that meets the input requirements of the emotion recognition model, so as to make the output of the emotion recognition model more accurate. Specifically, it includes performing augmentation processing on the face in the face image to increase the diversity of face data and improve the generalization ability of the emotion recognition model. Performing grayscale processing on the face image to convert the color image into a grayscale image, thereby simplifying the data volume of the image and highlighting the structure and contour information of the image. Changing the size of the face image to conform to the image required by the input of the emotion recognition model.

[0117] In this step, preprocessing such as performing augmentation processing on the face in the reference face image, performing grayscale processing on the reference face image, and changing the size of the reference face image is carried out to obtain the target face image.

[0118] Step 303: Invoke the pre-trained emotion recognition model to process the target face image to obtain the emotion recognition result.

[0119] Among them, the emotion recognition model is pre-trained by technicians according to needs. Its input is the face image to be recognized, and the output is the emotion recognition result corresponding to the face image. When the emotion recognition result is the emotion change trend, the input of the emotion recognition model includes not only the image collected at the current time point, but also the images collected before, so that the emotion recognition model can recognize these images to obtain the emotion change trend. The specific training process of the emotion recognition model is similar to the existing training process and will not be elaborated here one by one.

[0120] In this step, when the emotion recognition result is the emotion recognition result corresponding to each object, each face image can be input into the pre-trained emotion recognition model to obtain the emotion recognition result corresponding to each face image, that is, the emotion recognition result corresponding to each object. When the emotion recognition result is the emotion recognition result corresponding to all objects, after obtaining the emotion recognition result corresponding to each object, these emotion recognition results can be summarized to obtain the emotion recognition result corresponding to all objects. For example, the most frequently occurring emotion recognition result is determined as the final emotion recognition result.

[0121] In an embodiment of the present application, the vehicle-mounted server can detect the human faces in the target image to obtain multiple rectangular frames. Then, within the multiple rectangular frames, a target face detection frame is determined. Finally, the image included in the target face detection frame is determined as the reference face image. Therefore, the embodiment of the present application provides a method for determining a reference face image, which is a further limitation of step 301. As Figure 4 shown, the specific steps include:

[0122] Step 401: Detect the human faces in the target image to obtain multiple rectangular frames.

[0123] In this step, the cascade algorithm is loaded into a variable to process the target image based on the loaded cascade classifier, search for and identify the positions of human faces in the target image, and mark the range of each human face with a rectangular frame to obtain multiple rectangular frames.

[0124] Furthermore, the length and width of the target image can be gradually reduced in proportion through the ScaleFactor parameter. Since the sizes of human faces in the image may vary, gradually reducing the image can increase the possibility of detecting human faces of various sizes, improve the comprehensiveness and accuracy of detection, and enable the cascade classifier to adapt to human faces of different sizes.

[0125] Among them, the ScaleFactor parameter is used to control the scaling ratio of the image in image detection.

[0126] Step 402: Determine a target face detection frame within the multiple rectangular frames.

[0127] In this step, there is a minNeighbors parameter set in image detection, which is used to form the minimum number of adjacent rectangular frames of the detection target. During the detection process, there may be a situation where multiple adjacent rectangular frames all identify human faces. Only when the number of adjacent rectangular frames reaches or exceeds this minimum number is it determined that there is indeed a human face in this area. Then, among these rectangular frames, the rectangular frame with the highest score is retained and determined as the target face detection frame.

[0128] Step 403: Determine the image included in the target face detection frame as the reference face image.

[0129] In this step, the image that the target face detection frame can include is the required reference face image. Therefore, the image included in the target face detection frame is determined as the reference face image.

[0130] It should be noted that if facial key points such as the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth are required in the subsequent detection process, after obtaining the reference facial image, the facial key points in the reference facial image need to be marked for subsequent detection based on these facial key points.

[0131] In the embodiment of the present application, the vehicle-mounted server can also determine the reference human image in the target image, and then preprocess the reference human image to obtain the target human image to meet the requirements of the behavior recognition model. Finally, the pre-trained behavior recognition model is called to recognize the target human image to obtain the behavior recognition result. Therefore, the embodiment of the present application provides a method for determining the behavior recognition result, and this method is as Figure 5 shown, which defines step 103, and the specific steps include:

[0132] Step 501, determine the reference human image in the target image.

[0133] In this step, the human recognition algorithm is used to recognize the reference human image in the target image.

[0134] Step 502, preprocess the reference human image to obtain the target human image.

[0135] In this step, based on the methods of computer vision and deep learning, the bone joint points in the reference human image are detected and located to obtain 18 bone joint points. Then, according to the Hungarian matching algorithm, the 18 human bone joint points are connected in sequence to form a skeleton feature map. Finally, according to this skeleton feature map, the reference human image is further cropped to obtain the corresponding slice heat map, that is, the target human image.

[0136] Step 503, call the pre-trained behavior recognition model to recognize the target human image to obtain the behavior recognition result.

[0137] Among them, the behavior recognition model is a violence recognition model pre-trained by technicians according to needs. Its input is the human image to be recognized, and the output is whether the behavior corresponding to the human image is a violent behavior. When the behavior recognition result is the behavior change trend, the behavior recognition model can be the action prediction part in the human pose estimation model, and then this prediction part is used to predict whether there is a violent behavior. The specific training process of this behavior recognition model is similar to the existing training process and will not be elaborated here one by one.

[0138] In this step, when the behavior recognition result is the behavior recognition result corresponding to each object, each human body image can be input into a pre-trained behavior recognition model to obtain the behavior recognition result corresponding to each human body image, that is, the behavior recognition result corresponding to each object. When the behavior recognition result is the behavior recognition result corresponding to all objects, after obtaining the behavior recognition result corresponding to each object, these behavior recognition results can be summarized to obtain the behavior recognition result corresponding to all objects. For example, the behavior recognition result that appears the most is determined as the final behavior recognition result.

[0139] In the embodiment of the present application, the target image is preprocessed to obtain a reference image, and then the target area where the object is located is obtained in the reference image. Finally, cropping is performed according to the target area to obtain the reference human body image in the target image. Therefore, the embodiment of the present application provides a method for determining a reference human body image, which is a further limitation of step 501, as Figure 6 shown, and the specific steps include:

[0140] Step 601, preprocess the target image to obtain a reference image.

[0141] In this step, due to factors such as shooting conditions and equipment, there are some problems that are not conducive to subsequent processing, such as poor lighting, noise interference, inappropriate size, etc. The preprocessing operation is to solve these problems. Common preprocessing operations include image grayscale conversion (converting a color image to a grayscale image to simplify calculations), noise reduction processing (removing noise in the image to make the image clearer), normalization (adjusting the brightness, contrast, etc. of the image to make the image have consistent features), adjusting the image size (scaling the image to a size suitable for subsequent algorithm processing), etc. By performing the above preprocessing operations on the target image, the quality of the target image is improved, and then a reference image is obtained.

[0142] Step 602, call the target detection model to determine the target area of the object in the reference image.

[0143] Among them, the target detection model is used to detect the area where the object is located in the image. Common target detection models include deep learning-based models, such as Faster R-CNN, YOLO (You Only Look Once) series, SSD (SingleShot MultiBox Detector), etc., as well as traditional target detection algorithms.

[0144] In this step, the reference image is input into the target detection model to obtain the target area of the object in the reference image.

[0145] Step 603, determine the image corresponding to the target area in the reference image as the reference human body image in the target image.

[0146] In this step, the minimum bounding matrix of the target area can be determined, and cropping is performed according to this matrix to obtain a human body image, which is the reference human body image in the target image.

[0147] In the embodiment of the present application, a microphone can also be arranged inside the train carriage. The microphone can collect audio data inside the train carriage. Then, the audio data can be detected to further verify whether an emergency occurs in the carriage. Therefore, the embodiment of the present application provides an emergency verification method, which is as Figure 7 shown, and the specific steps include:

[0148] Step 701, obtain the audio data inside the train carriage.

[0149] In this step, after the microphone collects the audio data, the microphone can send the audio data to an encoder, so that the encoder encodes the audio data and sends it to an in-vehicle server through a network adapter. In this way, after the in-vehicle server receives the data sent by the network adapter, a decoder in the in-vehicle server decodes it, so that the in-vehicle server obtains the audio data inside the train carriage.

[0150] Among them, the encoder is used to encode data to perform related processing such as compressing the data volume and format conversion. The decoder is used to decode data to obtain the audio data in the data. In this step, the encoder adaptively adjusts the code rate according to the scenario. Specifically, the code rate is adjusted according to the current network condition. The network adapter, also known as the network interface card (Network Interface Card, NIC) or network card, is mainly used to establish a communication connection between the in-vehicle device and the network.

[0151] In addition, to avoid data loss, the encoder can send the encoded data to a memory for storage, and then a dedicated decoder reads the data from the memory and decodes it, and then further processes or stores the decoded audio data as needed. For example, the stored data is sent to the in-vehicle server through a network adapter.

[0152] Step 702, detect whether the audio data meets the first preset condition.

[0153] Among them, the first preset condition is that keywords appear in the audio data and the number of keywords is greater than a preset value. For example, the keywords can be words such as fire, fainting, and injury. The preset value is set according to the experience of technicians. For example, the preset value is 4. Of course, the first preset condition can also be other conditions, which are not limited here.

[0154] In this step, an audio recognition model is used to convert the audio data into text data, and then it is detected whether there are keywords in the text data and whether the number of the keywords is greater than a preset value, so as to detect whether the audio data meets the first preset condition.

[0155] Step 703, when the audio data meets the first preset condition, it is determined that an emergency has occurred inside the carriage.

[0156] In this step, when the audio data meets the first preset condition, it indicates that an emergency has occurred inside the carriage, and the on-vehicle server can determine that an emergency has occurred inside the carriage.

[0157] After that, the on-vehicle server can send the above result to the train management platform. After the train management platform receives the result, it issues an alarm to notify the relevant personnel that an emergency has occurred inside the carriage.

[0158] In addition, in the embodiment of the present application, the data can also be directly sent to the train management platform through a network adapter. In this way, the train management platform decodes the data to obtain the audio data, and uses a speaker to convert the audio data into a sound signal for on-site playback.

[0159] In the embodiment of the present application, the train management platform can also detect the emotion recognition result and the behavior recognition result to detect whether an emergency has occurred inside the carriage, so that the relevant personnel can detect and handle the emergency in time. Therefore, the embodiment of the present application provides an emergency detection method, and the specific steps of the method include: determining a fusion recognition result according to the emotion recognition result and the behavior recognition result; when the fusion recognition result meets the second preset condition, it is determined that an emergency has occurred inside the carriage, and when the fusion recognition result does not meet the second preset condition, it is determined that no emergency has occurred inside the carriage.

[0160] In this step, in order to accurately identify whether an emergency has occurred inside the carriage, a fusion recognition result can be obtained based on the emotion recognition result and the behavior recognition result to comprehensively determine whether an emergency has occurred inside the carriage.

[0161] Further, the emotion recognition result not only includes the category, but also includes the corresponding confidence level. Similarly, the behavior recognition result includes the category and the corresponding confidence level. The specific steps for obtaining the fusion recognition result based on the emotion recognition result and the behavior recognition result can be: determining the first weight corresponding to the first category and the second weight corresponding to the second category; multiplying the first weight and the first confidence level to obtain a first value; multiplying the second weight and the second confidence level to obtain a second value; adding the first value and the second value to obtain the fusion recognition result.

[0162] The specific steps for determining the first weight corresponding to the first category and the second weight corresponding to the second category are as follows: First, the first weight can be determined based on the first category and the pre-established corresponding relationship between the category and the weight. The second weight can be determined based on the second category and the pre-established corresponding relationship between the category and the weight.

[0163] Among them, the above corresponding relationship is set based on the actual needs of technicians.

[0164] For example, when the categories in the emotion recognition result are extremely negative emotions such as anger, sadness, and fear, the weight is set to A1. When the categories in the emotion recognition result are disgust and surprise, the weight is set to A2. When the categories in the emotion recognition result are happiness and neutrality, the weight is set to A3. When the category in the behavior recognition result is violent behavior, the weight is set to B1. When the category in the behavior recognition result is normal behavior, the weight is set to B2. Then, multiply the confidence level in the emotion recognition result by the weight assigned to its category, multiply the confidence level in the behavior recognition result by the weight assigned to its category, and then add the two products to obtain a comprehensive result value, which becomes the fusion recognition result.

[0165] Of course, other methods can also be used in this application to obtain the fusion recognition result. For example, obtain the emotion recognition results and behavior recognition results of all objects in the image, add the confidence levels in these results to obtain a third value, and set this third value as the fusion recognition result. Or, when the emotion recognition result only includes categories, directly use the emotion recognition result and the behavior recognition result as the fusion recognition result.

[0166] In addition, this step can also set preset conditions. For example, when the fusion recognition result is a numerical value, set the preset condition as the fusion recognition result being within the first range. That is, when the fusion recognition result is within the first range, it is determined that an emergency has occurred in the carriage. When the fusion recognition result is not within the first range, it is determined that no emergency has occurred in the carriage. When the fusion recognition result is a combination of the emotion recognition result and the behavior recognition result, the preset condition is that the emotion recognition result is the first category or / and the behavior recognition result is the second category. The first category is negative emotions such as anger, sadness, fear, disgust, and surprise, and the second category is violent behavior.

[0167] Furthermore, in order to achieve more accurate alarm, technicians can also divide it into multiple abnormal levels, each abnormal level corresponding to a range. Then, according to the range where the fusion recognition result is located, determine the level corresponding to the fusion recognition result. Then, according to this level, determine whether an emergency has occurred in the carriage. In this case, the preset condition is that the level corresponding to the fusion recognition result is the target level

[0168] For example, set the exception levels C1, C2, and C3, corresponding to three levels respectively: no exception, ordinary exception, and severe exception, and set C2 and C3 as the target levels, and C1 as the non-target level. In this way, it is possible to determine whether an emergency has occurred in the carriage by determining the exception level corresponding to the fusion recognition result.

[0169] In addition, after determining the exception level, the exception level can also be notified to the user. For example, display the determined exception level so that the user can understand the severity of the emergency according to the exception level for subsequent processing.

[0170] In summary, the emergency detection method provided by the embodiments of the present application can more accurately and timely detect emergencies in the train carriage by comprehensively considering emotion and behavior information, using reasonable weight allocation and exception level setting, providing strong guarantee for the safe operation of the train and the safety of the lives and property of passengers.

[0171] Based on the above embodiments, the embodiments of the present application provide a flowchart of a carriage event detection method, as Figure 8 shown. The flowchart focuses on the monitoring of personnel inside the train carriage and the recognition of abnormal behaviors. The detailed process is as follows: First, collect the monitoring data of personnel inside the train carriage through a monitoring camera, and then preprocess the input image frames to prepare for subsequent analysis. Secondly, use the cascade classification algorithm to process and analyze the data of all personnel in the carriage, detect and label the face rectangular frames, and then perform non-maximum suppression and optimization to make the detection frames more in line with the actual situation. Then, label five key feature points on the face to locate the face. After changing the size of the face image, send it into the deep learning network model for real-time emotion prediction, and finally map the predicted emotion recognition result back to the original image. After that, use the object detection algorithm to detect all personnel in the image and crop them to an appropriate size. Perform 18 key body part detections on the cropped single-person images to generate slice heatmaps, and connect the key points into a skeleton map according to the Hungarian algorithm. Based on the skeleton map, call the deep learning network model to execute the violent behavior recognition prediction part, recognize each frame of the image, and finally map the predicted violent behavior recognition result and the category weighted result back to the original image. Finally, display the analysis results of emotion prediction and violent behavior recognition on the management control platform to facilitate technicians to timely grasp the status of personnel in the carriage and make a quick response to abnormal situations.

[0172] It should be understood that although the steps in the flowchart are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0173] Please refer to Figure 9 , an embodiment of the present application provides a carriage event detection device, including:

[0174] An acquisition unit 901, configured to acquire a target image in a train carriage,

[0175] A first recognition unit 902, configured to recognize the target image according to an emotion recognition strategy to obtain an emotion recognition result;

[0176] A second recognition unit 903, configured to recognize the target image according to a behavior recognition strategy to obtain a behavior recognition result;

[0177] A sending unit 904, configured to send the behavior recognition result and the emotion recognition result to a train management platform, so that the train management platform detects whether an emergency occurs in the carriage according to the emotion recognition result and the behavior recognition result.

[0178] Optionally, the first recognition unit 902 is configured to:

[0179] Recognize a reference face image in the target image;

[0180] Preprocess the reference face image to obtain a target face image;

[0181] Call a pre-trained emotion recognition model to process the target face image to obtain an emotion recognition result.

[0182] Optionally, the first recognition unit 902 is configured to:

[0183] Detect faces in the target image to obtain a plurality of rectangular frames;

[0184] Determine a target face detection frame within the plurality of rectangular frames;

[0185] Determine the image included in the target face detection frame as the reference face image.

[0186] Optionally, the second recognition unit 903 is configured to:

[0187] Determine a reference human body image in the target image;

[0188] Preprocess the reference human body image to obtain a target human body image;

[0189] Call a pre-trained behavior recognition model to recognize the target human body image and obtain a behavior recognition result.

[0190] Optionally, the second recognition unit 903 is configured to:

[0191] Preprocess the target image to obtain a reference image;

[0192] Call a target detection model to determine a target area of an object in the reference image;

[0193] Determine the image corresponding to the target area in the reference image as the reference human body image in the target image.

[0194] Optionally, the device further includes a detection unit 905, and the detection unit 905 is configured to:

[0195] Obtain audio data in the train carriage;

[0196] Detect whether the audio data meets a first preset condition;

[0197] When the audio data meets the first preset condition, determine that an emergency has occurred in the carriage.

[0198] Optionally, the sending unit 904 is configured to:

[0199] Determine a fusion recognition result according to the emotion recognition result and the behavior recognition result;

[0200] When the fusion recognition result meets a second preset condition, determine that an emergency has occurred in the carriage;

[0201] When the fusion recognition result does not meet the second preset condition, determine that no emergency has occurred in the carriage.

[0202] Optionally, the emotion recognition result includes a first category and a corresponding first confidence level, and the sending unit 904 is configured to:

[0203] Determine a first weight corresponding to the first category and a second weight corresponding to the second category;

[0204] Multiply the first weight and the first confidence level to obtain a first value;

[0205] Multiply the second weight and the second confidence level to obtain a second value;

[0206] Add the first value and the second value to obtain a fusion recognition result.

[0207] For the specific limitations of the above carriage event detection device, reference can be made to the limitations of the carriage event detection method in the foregoing text, which will not be elaborated here. Each unit in the above carriage event detection device can be implemented in whole or in part by software, hardware, and their combination. The above units can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0208] In one embodiment, a computer device is provided, and the internal structure diagram of the computer device can be as Figure 10 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it can implement a carriage event detection method as above. It includes: including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, any step in the above carriage event detection method is implemented.

[0209] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, any step in the above carriage event detection method can be implemented.

[0210] Those skilled in the art should understand that the embodiments of the present application can be provided as a carriage event detection method, system, or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0211] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0212] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0213] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0214] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0215] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A carriage event detection method, characterized in that, Including: Obtain a target image inside a train carriage; According to an emotion recognition strategy, recognize the target image to obtain an emotion recognition result; According to a behavior recognition strategy, recognize the target image to obtain a behavior recognition result; Send the behavior recognition result and the emotion recognition result to a train management platform, so that the train management platform detects whether an emergency occurs inside the carriage according to the emotion recognition result and the behavior recognition result.

2. The method according to claim 1, wherein The recognizing the target image according to the emotion recognition strategy to obtain an emotion recognition result includes: Recognize a reference face image in the target image; Preprocess the reference face image to obtain a target face image; Call a pre-trained emotion recognition model to process the target face image to obtain an emotion recognition result.

3. The method according to claim 2, wherein The determining the reference face image in the target image includes: Detect faces in the target image to obtain a plurality of rectangular frames; Determine a target face detection frame within the plurality of rectangular frames; Determine the image included in the target face detection frame as the reference face image.

4. The method according to claim 1, wherein The recognizing the target image according to the behavior recognition strategy to obtain a behavior recognition result includes: Determine a reference human body image in the target image; Preprocess the reference human body image to obtain a target human body image; Call a pre-trained behavior recognition model to recognize the target human body image to obtain a behavior recognition result.

5. The method according to claim 4, characterized in that, The determining the reference human body image in the target image includes: Preprocess the target image to obtain a reference image; Call a target detection model to determine a target area of an object in the reference image; Determine the image corresponding to the target area in the reference image as the reference human body image in the target image.

6. The method according to claim 1, characterized in that, The method further includes: Obtain audio data inside the train carriage; Detect whether the audio data meets a first preset condition; When the audio data meets the first preset condition, determine that an emergency occurs inside the carriage.

7. The method according to claim 1, wherein The detecting whether an emergency occurs inside the carriage according to the emotion recognition result and the behavior recognition result includes: Determine a fusion recognition result according to the emotion recognition result and the behavior recognition result; When the fusion recognition result meets a second preset condition, determine that an emergency occurs inside the carriage; When the fusion recognition result does not meet the second preset condition, determine that no emergency occurs inside the carriage.

8. The method according to claim 7, wherein The emotion recognition result includes a first category and a corresponding first confidence level, and the behavior recognition result includes a second category and a corresponding second confidence level. The determining a fusion recognition result according to the emotion recognition result and the behavior recognition result includes: Determine a first weight corresponding to the first category and a second weight corresponding to the second category; Multiply the first weight and the first confidence level to obtain a first value; Multiply the second weight and the second confidence level to obtain a second value; Add the first value and the second value to obtain a fusion recognition result.

9. A method and device for detecting carriage events, characterized in that, Including: An acquisition unit for acquiring a target image inside a train carriage; A first recognition unit for recognizing the target image according to an emotion recognition strategy to obtain an emotion recognition result; A second recognition unit for recognizing the target image according to a behavior recognition strategy to obtain a behavior recognition result; A sending unit for sending the behavior recognition result and the emotion recognition result to a train management platform, so that the train management platform detects whether an emergency occurs inside the carriage according to the emotion recognition result and the behavior recognition result.

10. A computer device, comprising: A memory and a processor, the memory stores a computer program, characterized in that when the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.