Method, device and equipment for detecting personnel conflict in train compartment and medium

By collecting and analyzing video and audio data in the train compartment in real time, identifying potential conflict behaviors and triggering early warnings, the problem of rapid detection and stopping of conflicts among personnel in the train compartment is solved, and safety and order are guaranteed.

CN120014611AActive Publication Date: 2025-05-16SHENZHEN ZHIHUI BOJIA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510458922.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

People conflicts in train cars may threaten the personal safety of passengers, disrupt the order of the carriage, lead to panic and uneasiness, affect the ride experience, and increase legal risks and pressure from railway companies.

Method used

Real-time video data is collected through the camera sensor, the video data quality is monitored in real time, frames are extracted and preprocessed, and keyframe images are obtained. Based on the keyframe images, potential conflict behavior is identified through pre-training models, targets need to be monitored, audio data and passenger behavior data are collected, conflict behavior is identified and judged based on multimodal information, and early warning mechanism is triggered to issue warning signals.

Benefits of technology

It has achieved rapid detection and stopped conflicting behaviors of personnel in the train, ensured passenger safety, maintained carriage order, reduced legal risks, and improved the accuracy and reliability of conflict detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014611A_ABST
    Figure CN120014611A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a device, equipment and a medium for detecting collision of personnel in a train compartment, and the detection method specifically comprises the steps: collecting real-time video data through a camera sensor, and monitoring the quality of the video data in real time; performing frame extraction on the video data according to a preset video frame sampling time interval, and performing preprocessing operation to obtain a key frame image; based on the key frame image, identifying a potential conflict behavior through a pre-training model; locking a to-be-monitored target according to an identification result, and collecting audio data of the to-be-monitored target, peripheral audio data and behavior data of other passengers; based on identification of multi-modal information, analyzing and judging whether multiple features of conflict behaviors exist or not; when the target needing to be monitored has a conflict behavior, an early warning mechanism is triggered, and an early warning signal is sent out, so that the purposes of guaranteeing the safety of passengers in train carriages, maintaining the order of the carriages and maintaining the harmonious running environment of the train are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of train safety monitoring, and in particular to a method, device, equipment and medium for detecting conflicts between people in train compartments. Background Art

[0002] The train compartment is a relatively closed and densely populated environment. If a conflict occurs in the train, it may directly threaten the personal safety of passengers; it may even disrupt the normal order in the train compartment. The noise and violence during the conflict will cause panic and anxiety among other passengers, leading to a tense atmosphere in the compartment, chaos and instability, affecting the passengers' riding experience. This will not only affect the people involved, but also bring additional legal risks and pressure to the management and operation of the train. Therefore, quickly detecting and stopping conflicts in the train is of great significance in terms of ensuring passenger safety, maintaining compartment order, avoiding legal disputes and responsibilities, and maintaining the reputation and image of the railway company. Summary of the invention

[0003] The main purpose of the present invention is to provide a method, device, equipment and medium for detecting personnel conflicts in train compartments, so as to achieve the purpose of ensuring the safety of passengers in train compartments, maintaining compartment order and maintaining a harmonious operating environment of the train.

[0004] To achieve the above object, the present invention provides a method for detecting a conflict between people in a train compartment, comprising the following steps: Collect real-time video data through camera sensors and monitor the quality of video data in real time; Sampling frames of the video data according to a preset video frame sampling time interval, and performing a preprocessing operation to obtain a key frame image; Based on the key frame images, identifying potential conflict behaviors through a pre-trained model; The target to be monitored is locked according to the recognition result, and the audio data of the target to be monitored, the surrounding audio data, and the behavior data of other passengers are collected; Based on the recognition of multimodal information, analyze and determine whether there are multiple features of conflicting behaviors; When there is conflict behavior in the monitored target, the early warning mechanism is triggered and an early warning signal is issued. The early warning signal includes the location of the carriage, the degree of congestion in the carriage, and the number of people involved in the conflict.

[0005] Furthermore, the step of collecting real-time video data through a camera sensor and monitoring the quality of the video data in real time includes: Continuously capture real-time video data inside the train compartment through a camera sensor, wherein the camera sensor is a wide-angle lens with night vision function, and the train compartment is equipped with multiple camera sensors; Real-time monitoring of the clarity and smoothness of the captured video; When video data loss or quality degradation to a set threshold is detected, an alarm mechanism is triggered and an alarm message is sent to the setting personnel. The alarm message includes the time, location and type of abnormality. When the alarm mechanism is triggered, the car audio data is collected in real time through the voice device, and other camera sensors collect the passenger behavior data in the car.

[0006] Furthermore, the step of sampling the video data at a preset video frame sampling time interval and performing a preprocessing operation to obtain a key frame image includes: Based on stable real-time video data, frames are extracted at set time intervals to obtain a series of initial key frame images; The initial key frame image is subjected to preprocessing operations of denoising, enhancement, and edge detection to obtain a final key frame image.

[0007] Furthermore, the step of identifying potential conflict behaviors through a pre-trained model based on the key frame image includes: Collect and preprocess video data of normal and abnormal human postures and behaviors in various compartment spaces to obtain a passenger behavior video dataset; Based on the passenger behavior video dataset, a three-dimensional convolutional neural network model is trained to obtain a potential conflict behavior model; The key frame image is input into the potential conflict behavior model to identify the potential conflict behavior of people in the car.

[0008] Furthermore, the step of training a three-dimensional convolutional neural network model based on the passenger behavior video dataset to obtain a potential conflict behavior model includes: Based on the preprocessed passenger behavior video dataset, the posture estimation algorithm is used to identify the key points of the human body and mark the corresponding position information; Convert key point location information into a data format suitable for processing by machine learning models; The converted key point position information is used as an additional feature channel to be spliced ​​with the spatiotemporal feature channel of the video data; the spatiotemporal feature is used to capture the motion information and spatial structure information between image frames; The spliced ​​behavior features are used as input and sent to the three-dimensional convolutional neural network model for deep learning on how to use the spliced ​​behavior features to perform conflict behavior identification tasks and obtain the final potential conflict behavior model.

[0009] Furthermore, the step of locking the target to be monitored according to the recognition result and collecting the audio data of the target to be monitored as well as the surrounding audio data and other passenger behavior data includes: According to the identification results of the model, the targets that need to be monitored with potential conflict behaviors are locked; According to the preset time length, the video data of the target to be monitored is continuously monitored, and the audio data and other passenger behavior data within the specified range of the target to be monitored are collected.

[0010] Furthermore, the step of analyzing and judging whether there are multiple features of conflicting behaviors based on the identification of multimodal information includes: Identify the collected multimodal information, wherein the multimodal information includes video data of the target to be monitored, audio data within a specified range of the target to be monitored, and other passenger behavior data; Respectively identifying the data in the multimodal information to obtain characteristic judgments of the conflicting behaviors; When the characteristics of the audio data or the passenger behavior data indicate potential conflicting behaviors, it is determined that the target to be monitored has conflicting behaviors.

[0011] The present invention also provides a device for detecting personnel conflicts in a train compartment, comprising: The data acquisition module is used to collect real-time video data through a camera sensor and monitor the quality of the video data in real time; A preprocessing module, used to extract frames of the video data according to a preset video frame sampling time interval, and perform preprocessing operations to obtain a key frame image; A model recognition module, used for recognizing potential conflict behaviors through a pre-trained model based on the key frame images; A monitoring module is used to lock the target to be monitored according to the recognition result, and collect the audio data of the target to be monitored, the surrounding audio data, and the behavior data of other passengers; An analysis module is used to analyze and determine whether there are multiple features of conflicting behaviors based on the recognition of multimodal information; The early warning module is used to trigger the early warning mechanism and send out an early warning signal when there is conflicting behavior in the target to be monitored. The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned method for detecting conflicts between people in train compartments when executing the computer program.

[0012] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for detecting conflicts between people in train compartments are implemented.

[0013] The method, device, equipment and medium for detecting conflicts between people in train carriages provided by the present invention have the following beneficial effects: real-time video data is collected through a camera sensor, and data quality is monitored to ensure timely alarm when video data is abnormal, thereby ensuring the continued effectiveness of the monitoring system; by combining video data, audio data and other passenger behavior data, the system can more comprehensively identify and analyze multiple features of conflict behavior, thereby improving the accuracy and reliability of conflict detection; when potential conflict behavior is detected, the system can quickly trigger an early warning mechanism and send out an early warning signal. The detailed information contained in the early warning signal, such as the location of the carriage, the degree of crowding and the number of people involved in the conflict, helps train staff to quickly locate the conflict scene, evaluate the conflict situation, and take corresponding handling measures. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 1 is a flow chart of a method for detecting a conflict between people in a train compartment according to an embodiment of the present invention; Figure 2 is a structural block diagram of a device for detecting conflicts between people in a train compartment according to an embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0015] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0017] Reference Figure 1 , which is a flow chart of a method for detecting a conflict between people in a train compartment proposed by the present invention, comprising the following steps: S1, collects real-time video data through camera sensors and monitors the quality of video data in real time; S2, sampling frames of the video data according to a preset video frame sampling time interval, and performing a preprocessing operation to obtain a key frame image; S3, based on the key frame image, identifying potential conflict behaviors through a pre-trained model; S4, locking the target to be monitored according to the recognition result, collecting audio data of the target to be monitored, surrounding audio data, and other passenger behavior data; S5, based on the recognition of multimodal information, analyzes and determines whether there are multiple features of conflicting behaviors; S6, when there is conflict behavior in the target to be monitored, the early warning mechanism is triggered and an early warning signal is issued. The early warning signal includes the position of the carriage, the degree of congestion in the carriage, and the number of people involved in the conflict.

[0018] As described in step S1 above, real-time video data is collected through camera sensors, and the quality of the video data is monitored in real time. Multiple camera sensors are configured in the train compartment to comprehensively monitor the situation in the compartment; through the camera sensors, real-time video data in the train compartment is continuously captured as the basis for subsequent conflict behavior identification and analysis; while collecting real-time video data, the quality of the video data needs to be monitored in real time, including two aspects: picture clarity and smoothness. Picture clarity refers to the resolution and detail expression ability of the video picture, and smoothness refers to the continuity and stability of video playback.

[0019] As described in step S2 above, the video data is extracted at a preset video frame sampling time interval, and a preprocessing operation is performed to obtain a key frame image. The video frame sampling time interval refers to the time interval for extracting image frames from continuous video data. The setting of this interval needs to be adjusted according to the actual situation to ensure that the amount of calculation and storage requirements are reduced as much as possible without affecting the accuracy of conflict behavior recognition; the frame extraction operation refers to extracting image frames from continuous video data at a preset time interval. These image frames will serve as the basis for subsequent analysis and processing. During the frame extraction process, the system will extract a series of image frames, namely the initial key frame images, from the video stream at the set time interval. The preprocessing operation is a series of processing on the initial key frame image to improve the quality and analyzability of the image. These preprocessing operations include denoising, enhancement and edge detection. The denoising operation aims to remove noise points in the image, which may be caused by sensor noise, distortion in the transmission process or other factors; the denoising operation helps to improve the clarity and recognizability of the image; the enhancement operation aims to improve the image's contrast, brightness and other attributes to make the details in the image clearer, which helps to better identify the characteristics of potential conflict behaviors in subsequent analysis; the edge detection operation aims to identify edge information in the image, which usually corresponds to the boundaries between objects or the contours of objects, which helps to more accurately identify objects and scenes in the image in subsequent analysis; the initial key frame image after the preprocessing operation will become clearer and easier to analyze. These preprocessed image frames are called key frame images, which are an important basis for subsequent conflict behavior identification and analysis.

[0020] As described in step S3 above, based on the key frame image, potential conflict behaviors are identified through a pre-trained model. The pre-trained model refers to a deep learning model that has been pre-trained on a large amount of data, and the model has learned prior knowledge about image recognition and classification; the pre-processed key frame image is input into the pre-trained model, and the pre-trained model will extract features of the input key frame image, and these features include information such as color, texture, shape, etc. in the image; based on the extracted features, the pre-trained model will classify the key frame image to determine whether it contains potential conflict behaviors. This classification process is based on the knowledge and rules learned by the model during the training phase; the model will output a recognition result indicating whether there is potential conflict behavior in the key frame image. The potential conflict behavior refers to behaviors that may cause conflicts or adverse consequences in train compartments, such as physical conflicts, verbal quarrels, overcrowding, etc. These behaviors usually have specific characteristics and patterns and can be detected and identified through image recognition and classification technology.

[0021] As described in step S4 above, the target to be monitored is locked according to the recognition result, and the audio data of the target to be monitored, as well as the surrounding audio data and other passenger behavior data are collected. According to step S3, the recognition result is obtained, and the recognition result may include but is not limited to physical conflict, verbal quarrel, overcrowding, abnormal behavior, etc.; based on these recognition results, those targets that are considered to need further monitoring are locked, and these targets may be specific individuals or groups, or specific areas or events; for the locked targets to be monitored, the system will give priority to collecting their audio data, which is achieved through the microphone array or single microphone arranged in the train car; in addition to the audio data directly targeting the monitoring target, the system will also collect audio data around the target, such as the conversations of other passengers, the sound of train operation, etc., so as to more comprehensively understand the environmental background and behavioral reasons of the target to be monitored; in addition to audio data, the system will also collect behavioral data of other passengers, such as tracking the movement trajectory of passengers, analyzing the posture and expression of passengers, etc.

[0022] As described in step S5 above, based on the recognition of multimodal information, multiple features of conflicting behaviors are analyzed and judged. Among them, multimodal information refers to the perception and understanding of the external world through a variety of different information representation forms. In the conflicting behavior detection of the present invention, the multimodal information includes the video data of the target to be monitored collected in step S4, the audio data within the specified range of the target to be monitored, and other passenger behavior data; image features are crucial in conflicting behavior detection, because the tension of the body, the anger or fear of the expression, and the aggressiveness of the action may be a precursor or manifestation of conflicting behavior, and the audio features are helpful to judge the emotional state of the person, such as anger, anxiety or tension, and whether there is quarrel or aggressive speech. After identifying the multimodal information, it is necessary to comprehensively analyze this information to judge whether there is conflicting behavior; for example, whether the person shows a tense body posture, such as leaning forward, clenching his hands, etc.; whether the person shows negative emotional expressions such as anger, fear or disgust; whether the intensity of the sound suddenly increases, and whether the tone becomes sharp or angry, which is used to judge the emotional state of the person. After comprehensively analyzing the multimodal information and the multi-features of conflicting behavior, a judgment can be made whether there is conflicting behavior.

[0023] As described in the above step S6, when the target to be monitored has conflicting behavior, the early warning mechanism is triggered and an early warning signal is issued, and the early warning signal includes the location of the car, the degree of crowding in the car, and the number of people involved in the conflict. Through comprehensive identification and judgment of multimodal information, it is determined that the target to be monitored has conflicting behavior and the behavior reaches the preset triggering conditions (such as the intensity of the conflict, the duration, etc.), the early warning mechanism will be triggered and the early warning signal will be sent. The early warning signal needs to contain sufficient information so that relevant personnel or systems can quickly understand the specific situation of the conflict and make a correct response, including the location of the car, clearly indicating the specific car where the conflict occurred, so that relevant personnel can quickly locate and go to the scene; the degree of crowding in the car, providing information on the degree of crowding in the car, which helps to determine whether additional rescue personnel or equipment are needed, and what method is most appropriate to enter the car; the number of people involved in the conflict, indicating the number of people involved in the conflict, which helps to assess the scale of the conflict and the possible degree of danger, thereby determining the level of response and the required resources. Warning signals can be issued in a variety of ways, including but not limited to: visual signals, such as display screens and indicator lights in the car, which are used to intuitively display warning information; auditory signals, such as alarms, broadcast notifications, etc., which are used to issue emergency notifications in the car and surrounding areas; wireless communications, such as through mobile phone text messages, e-mails or dedicated wireless communication systems, to send warning information to relevant personnel or systems.

[0024] In one embodiment, the step S1 of collecting real-time video data through a camera sensor and monitoring the quality of the video data in real time includes: S11, continuously capturing real-time video data in the train compartment through a camera sensor, wherein the camera sensor is a wide-angle lens with a night vision function, and the train compartment is equipped with multiple camera sensors; S12, real-time monitoring of the picture clarity and smoothness of the collected video; S13, when it is detected that the video data is lost or the quality drops to a set threshold, an alarm mechanism is triggered and an alarm message is sent to a setting person, wherein the alarm message includes the time, location and type of the abnormality; S14, while triggering the alarm mechanism, the car audio data is collected in real time through the voice device, and other camera sensors collect the passenger behavior data in the car.

[0025] As described in step S11 above, multiple wide-angle lens camera sensors with night vision functions are configured in the train compartment to continuously capture real-time video data in the compartment. The night vision function ensures that the camera sensor can capture clear video images even in a dark environment, and the wide-angle lens ensures a larger field of view, so that the entire compartment can be effectively monitored. As described in step S12 above, during the video data collection process, the system will monitor the clarity and smoothness of the video in real time to ensure that the quality of the collected video data meets certain standards, so that subsequent analysis and processing can be accurately performed.

[0026] As described in step S13 above, if video data is lost during real-time monitoring (for example, video signal interruption due to camera sensor failure or network problems) or video quality drops below a set threshold (for example, the image is blurred or the image is distorted, mosaic, delayed, etc. due to the camera being blocked or the light being too dark), the system will immediately trigger the alarm mechanism and send an alarm message to the set manager or monitoring center. The alarm message includes the time, location and type of abnormality (such as video loss, quality degradation, etc.). As described in step S14 above, when the alarm mechanism is triggered, in addition to sending the alarm message, the system will immediately start the voice device to collect audio data in the car in real time (such as conversations, quarrels, loud noises, etc. between passengers); at the same time, other camera sensors will continue to collect behavioral data of passengers in the car (such as passengers' movements, postures, expressions and interactions between passengers) so as to understand the situation in the car at any time.

[0027] In one embodiment, the step S2 of sampling the video data at a preset video frame sampling time interval and performing a preprocessing operation to obtain a key frame image includes: S21, based on the stable real-time video data, extracting frames at a set time interval to obtain a series of initial key frame images; S22, performing preprocessing operations of denoising, enhancement, and edge detection on the initial key frame image to obtain a final key frame image.

[0028] As described in step S21 above, in the embodiment of the present invention, a series of frames (i.e., initial key frame images) are extracted from the collected stable (with picture clarity and smoothness within a preset range) real-time video data at a specified time interval (e.g., 0.2 seconds) through opencv (an open source, cross-platform computer vision and machine learning software library). These images are evenly distributed in time and can roughly reflect the overall situation of the video content. As described in step S22 above, the initial key frame images are preprocessed, including denoising. Since the video data may be interfered by various noises (such as environmental noise, equipment noise, etc.) during the acquisition and transmission process, the initial key frame images need to be denoised to eliminate the impact of these noises on the image quality; enhancement, improving the contrast and brightness of the image to make it clearer, easier to observe and analyze; edge detection, identifying edge features in the image (such as the outline of the object, texture boundaries, etc.). These edge features have important reference value for subsequent image analysis, recognition and understanding tasks, and finally obtain key frame images.

[0029] In one embodiment, the step S3 of identifying potential conflict behaviors through a pre-trained model based on the key frame image includes: S31, collecting passenger behavior video data of various conflicting behaviors and normal behaviors in a train environment, and passenger audio data synchronously corresponding to the passenger behavior video data; S32, preprocessing the passenger behavior video data and the passenger audio data, and marking conflicting behaviors in the video and audio respectively; S33, based on the passenger behavior video data and the passenger audio data, selecting an existing model architecture for training to obtain a potential conflict behavior model; S34, inputting the key frame image into the potential conflict behavior model to identify the potential conflict behavior of the people in the car.

[0030] As described in step S31 above, the video data of passenger behavior in various train environments (such as crowded environments, various lighting conditions, various angles, etc.) are collected. These data include various conflict behaviors (such as quarrels, fighting, etc.) and normal behaviors (such as talking, reading, etc.). At the same time, it is also necessary to collect audio data that is synchronized with these video data to capture sound features, such as volume, speech speed, emotions, etc., where the synchronization of video and audio data is crucial because sound and image complement each other and together constitute a complete picture of passenger behavior. As described in step S32 above, the video and audio data collected in step S31 are cleaned, format converted, and feature extracted. The video data needs to be subjected to frame rate adjustment, size scaling, color correction, and other operations; the audio data needs to be subjected to noise reduction, volume standardization, and other processing; after preprocessing, the conflict behaviors in the video and audio data need to be marked, marking the start and end time of the conflict behavior, as well as the type of conflict behavior, etc. As described in step S33 above, a suitable machine learning or deep learning model architecture is selected based on the nature of the conflicting behavior to be identified and the characteristics of the collected data, and the selected model architecture is trained using the preprocessed video and audio data and the corresponding conflicting behavior markers. During the training process, the model will learn how to identify conflicting behaviors based on the input video and audio data. As described in step S34 above, the key frame images extracted in step S2 are input into the trained potential conflicting behavior model, and the model uses the knowledge and features learned during the training process to identify the potential conflicting behaviors of people in the car based on the input key frame images.

[0031] In the process of collecting, processing, storing and transmitting real-time video data, audio data and other passenger behavior data in train carriages, strict data privacy and security protection measures are implemented, including: encrypting all collected data, including video data, audio data and other passenger behavior data, ensuring the security of data during transmission and storage, and enhancing data protection strength by using encryption algorithms. In the data processing stage, sensitive information such as passengers' facial features and identity information are anonymized to avoid leaking passengers' personal privacy, and special data encryption equipment and software are used to achieve data encryption and anonymization. Establish a strict access control mechanism to ensure that only authorized personnel can access and process this data, and use multi-factor authentication (such as passwords, biometrics, etc.) to enhance the security of access control. Back up data regularly to prevent data loss or damage, and formulate a data recovery plan to ensure rapid recovery when data is lost or damaged. Regularly conduct security assessments and vulnerability scans on the system to promptly discover and repair potential security risks.

[0032] In one embodiment, the step S33 of selecting an existing model architecture for training based on the passenger behavior video data and the passenger audio data to obtain a potential conflict behavior model includes: S321, marking conflicting behaviors and normal behaviors in the passenger behavior video data and the passenger audio data, wherein the conflicting behaviors at least include the start and end time of the conflict, the locations of the conflicting participants, and the behaviors of the surrounding passengers when the conflict occurs; S322, on the video processing unit, extracting dynamic features of the labeled passenger behavior video data through a convolutional neural network, and selecting an existing video behavior recognition model for training; S323, extracting the frequency spectrum features of each frame of the audio signal in the labeled passenger audio data on the audio unit, building a model using a recurrent neural network, learning the time series features in the audio signal, and inputting the frequency spectrum features into the recurrent neural network model for training and optimization; S324, fusing the output features of each model in the video processing module and the audio processing module through a fusion module; S325, based on the annotated passenger behavior video data and passenger audio data, the fused model is trained and optimized to obtain a potential conflicting behavior model that simultaneously identifies whether there is a conflicting behavior in the video data and the audio data.

[0033] As described in the above step S321, the conflicting behaviors and normal behaviors in the passenger behavior video data and the passenger audio data are annotated, and the annotated information at least includes the start and end time of the conflict, the location of the conflict participants, and the behavior of the surrounding passengers when the conflict occurs. The annotated information is used as the basis for subsequent model training to help the model learn the characteristics of the conflicting behaviors. As described in the above step S322, on the video processing unit, a convolutional neural network (CNN) is used to extract the dynamic features of the annotated passenger behavior video data. CNN performs well in processing image and video data and can capture spatial and temporal features in the video. In an embodiment of the present invention, an existing YOLOv8 model (a target detection model that can identify and locate targets in an image) is selected, and the YOLOv8 model is trained based on the captured spatial and temporal features so that the model can recognize the behavior in the video. As described in step S323 above, on the audio unit, the spectral features of each frame of the audio signal in the annotated passenger audio data are extracted. The extracted spectral features reflect the characteristics of the audio signal in the frequency domain, which is helpful for identifying the information in the speech; a recurrent neural network (RNN) is used to build a model to learn the temporal features in the audio signal. RNN has advantages in processing sequence data (such as audio signals) and can capture the time dependency in the signal; the spectral features are input into the RNN model for training and optimization, so that the model can recognize the volume, speech speed, emotion, etc. in the audio. As described in step S324 above, the output features of each model in the video processing module and the audio processing module are fused through the fusion module, and the accuracy of the model in identifying conflicting behaviors is improved by combining video and audio information. As described in step S325 above, based on the annotated passenger behavior video data and passenger audio data, the fused model is trained and optimized to obtain a potential conflicting behavior model that can simultaneously identify whether there is a conflicting behavior in the video data and audio data. Through training and optimization, the model can learn the characteristics of the conflicting behavior for judgment on actual data. Optionally, the potential conflict behavior model may also adopt other conventional model training methods in the prior art.

[0034] In one embodiment, the step S4 of locking the target to be monitored according to the recognition result and collecting the audio data of the target to be monitored as well as the surrounding audio data and other passenger behavior data includes: S41, based on the recognition result of the model, lock the target to be monitored that has potential conflicting behaviors; S42, continuously monitoring the video data of the target to be monitored, collecting the audio data and other passenger behavior data within the specified range of the target to be monitored according to a preset time period.

[0035] As described in step S41 above, through the output of the potential conflict behavior recognition model in step S3, based on the analysis of video data, such as the characters' movements, expressions, positional relationships, etc., it is determined which targets (such as passengers, staff, etc.) may be showing conflict behaviors or have a tendency to conflict behaviors. Once a target with potential conflict behaviors is identified, the system will lock it as a target to be monitored; as described in step S42 above, after locking the target to be monitored, the system will enter a continuous monitoring stage, the purpose of which is to have a deeper understanding of the behavior patterns of the targets to be monitored and their interaction with the surrounding environment. The system continues to collect video data of the targets to be monitored to observe whether their movements, expressions, etc. have changed, and whether new conflict behaviors have appeared. In addition to video data, the system will also collect audio data within the specified range of the targets to be monitored to understand the voice content, tone, etc. of the targets to be monitored, so as to more comprehensively evaluate their emotional state and possible conflict behaviors. In addition, the system will also pay attention to the behavior data of other passengers around the targets to be monitored, including other passengers' movements, expressions, and interactions with the targets to be monitored, etc., to help the system more accurately judge the nature and possible consequences of the conflict behaviors. During this process, the preset duration of continuous monitoring is set according to specific scenarios and needs to ensure that the system can collect enough information to make accurate judgments.

[0036] In one embodiment, the step S5 of analyzing and determining whether there are multiple features of conflicting behaviors based on the recognition of multimodal information includes: S51, identifying the collected multimodal information, wherein the multimodal information includes video data of the target to be monitored, audio data within a specified range of the target to be monitored, and other passenger behavior data; S52, identifying the data in the multimodal information respectively through a potential conflict behavior model; S53, based on the recognition result, determining whether there is a potential conflicting behavior among the data in the multimodal information; S54, when there is a potential conflicting behavior in the audio data or the other passenger behavior data within the specified range of the target to be monitored, it is determined that there is indeed a conflicting behavior in the target to be monitored.

[0037] As described in step S51 above, the multimodal information collected in step S4 is identified, and the multimodal information includes video data (information such as actions, expressions, and positions) of the target to be monitored, audio data (information such as voice content, intonation, and speech speed) within the specified range of the target to be monitored, and behavioral data of other passengers (information such as actions, expressions, and interactions with the target to be monitored). As described in step S52 above, after the multimodal information is identified, the key information related to the conflicting behavior is identified for each modality of data using a pre-constructed potential conflicting behavior model. As described in step S53 above, it is determined whether each data in the multimodal information has potential conflicting behavior based on the recognition result of the model. The audio data or other passenger behavior data within the specified range of the target to be monitored is used as an auxiliary judgment of the conflicting behavior. If one or both of the two have potential conflicting behavior, it is considered that the target to be monitored does have conflicting behavior. For example, the model recognizes that the audio data contains obvious conflicting behavior signs such as quarrels and insults; and other passenger behavior data show that passengers suddenly gather in a certain place, etc.

[0038] Reference Figure 2 , is a structural block diagram of a train compartment personnel conflict detection device in one embodiment of the present invention, comprising: The data acquisition module is used to collect real-time video data through a camera sensor and monitor the quality of the video data in real time; A preprocessing module, used to extract frames of the video data according to a preset video frame sampling time interval, and perform preprocessing operations to obtain a key frame image; A model recognition module, used for recognizing potential conflict behaviors through a pre-trained model based on the key frame images; A monitoring module is used to lock the target to be monitored according to the recognition result, and collect the audio data of the target to be monitored, the surrounding audio data, and the behavior data of other passengers; An analysis module is used to analyze and determine whether there are multiple features of conflicting behaviors based on the recognition of multimodal information; The early warning module is used to trigger the early warning mechanism and send out an early warning signal when there is conflicting behavior in the target to be monitored.

[0039] For the specific implementation of each module in the above device example, please refer to the above method embodiment, which will not be repeated here.

[0040] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0041] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0042] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0043] In summary, real-time video data is collected through a camera sensor, and the quality of the video data is monitored in real time; the video data is framed according to a preset video frame sampling time interval, and a preprocessing operation is performed to obtain a key frame image; based on the key frame image, potential conflicting behaviors are identified through a pre-trained model; the target to be monitored is locked according to the identification result, and the audio data of the target to be monitored as well as the surrounding audio data and other passenger behavior data are collected; based on the recognition of multimodal information, multiple features of conflicting behaviors are analyzed and determined; when the target to be monitored has conflicting behaviors, the early warning mechanism is triggered and an early warning signal is issued to ensure the safety of passengers in the train compartment, maintain the order of the compartment, and maintain a harmonious operating environment of the train.

[0044] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.

[0045] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0046] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for detecting conflicts between people in a train compartment, characterized in that: The following steps are involved: Collect real-time video data through camera sensors and monitor the quality of video data in real time; Sampling frames of the video data according to a preset video frame sampling time interval, and performing a preprocessing operation to obtain a key frame image; Based on the key frame images, identifying potential conflict behaviors through a pre-trained model; The target to be monitored is locked according to the recognition result, and the audio data of the target to be monitored, the surrounding audio data, and the behavior data of other passengers are collected; Based on the recognition of multimodal information, analyze and determine whether there are multiple features of conflicting behaviors; When there is conflict behavior in the monitored target, the early warning mechanism is triggered and an early warning signal is issued. The early warning signal includes the location of the carriage, the degree of congestion in the carriage, and the number of people involved in the conflict.

2. The method for detecting conflicts between people in a train compartment according to claim 1, characterized in that: The step of collecting real-time video data through a camera sensor and monitoring the quality of the video data in real time includes: Continuously capture real-time video data inside the train compartment through a camera sensor, wherein the camera sensor is a wide-angle lens with night vision function, and the train compartment is equipped with multiple camera sensors; Real-time monitoring of the clarity and smoothness of the captured video; When video data loss or quality degradation to a set threshold is detected, an alarm mechanism is triggered and an alarm message is sent to the setting personnel. The alarm message includes the time, location and type of abnormality. When the alarm mechanism is triggered, the car audio data is collected in real time through the voice device, and other camera sensors collect the passenger behavior data in the car.

3. The method for detecting conflicts between people in a train compartment according to claim 1, characterized in that: The step of sampling the video data according to a preset video frame sampling time interval and performing a preprocessing operation to obtain a key frame image includes: Based on stable real-time video data, frames are extracted at set time intervals to obtain a series of initial key frame images; The initial key frame image is subjected to preprocessing operations of denoising, enhancement, and edge detection to obtain a final key frame image.

4. The method for detecting conflicts between people in a train compartment according to claim 1, characterized in that: The step of identifying potential conflict behaviors through a pre-trained model based on the key frame image includes: Collecting passenger behavior video data of various conflicting behaviors and normal behaviors in a train environment, and passenger audio data synchronously corresponding to the passenger behavior video data; Preprocessing the passenger behavior video data and the passenger audio data, and marking conflicting behaviors in the video and audio respectively; Based on the passenger behavior video data and the passenger audio data, an existing model architecture is selected for training to obtain a potential conflict behavior model; The key frame image is input into the potential conflict behavior model to identify the potential conflict behavior of people in the car.

5. The method for detecting conflicts between people in a train compartment according to claim 4, characterized in that: The step of selecting an existing model architecture for training based on the passenger behavior video data and the passenger audio data to obtain a potential conflict behavior model includes: Marking the conflicting behaviors and normal behaviors in the passenger behavior video data and the passenger audio data, wherein the conflicting behaviors at least include the start and end time of the conflict, the locations of the conflicting participants, and the behaviors of the surrounding passengers when the conflict occurs; On the video processing unit, the dynamic features of the annotated passenger behavior video data are extracted through a convolutional neural network, and an existing video behavior recognition model is selected for training; On the audio unit, extract the spectral features of each frame of the audio signal in the labeled passenger audio data, use a recurrent neural network to build a model, learn the temporal features in the audio signal, and input the spectral features into the recurrent neural network model for training and optimization; Through the fusion module, the output features of each model in the video processing module and the audio processing module are fused; Based on the labeled passenger behavior video data and passenger audio data, the fused model is trained and optimized to obtain a potential conflicting behavior model that can simultaneously identify whether there are conflicting behaviors in the video data and audio data.

6. The method for detecting conflicts between people in a train compartment according to claim 1, characterized in that: The step of locking the target to be monitored according to the recognition result and collecting the audio data of the target to be monitored as well as the surrounding audio data and other passenger behavior data includes: According to the identification results of the model, the targets that need to be monitored with potential conflict behaviors are locked; According to the preset time length, the video data of the target to be monitored is continuously monitored, and the audio data and other passenger behavior data within the specified range of the target to be monitored are collected.

7. The method for detecting conflicts between people in a train compartment according to claim 1, characterized in that: The step of analyzing and judging whether there are multiple features of conflicting behaviors based on the identification of multimodal information includes: Identify the collected multimodal information, wherein the multimodal information includes video data of the target to be monitored, audio data within a specified range of the target to be monitored, and other passenger behavior data; Respectively identifying data in the multimodal information through a potential conflict behavior model; Based on the recognition result, determining whether there is potential conflicting behavior among the data in the multimodal information; When there is a potential conflicting behavior in the audio data or the other passenger behavior data within the specified range of the target to be monitored, it is determined that the target to be monitored does have a conflicting behavior.

8. A device for detecting conflicts between people in a train compartment, characterized in that: include: The data acquisition module is used to collect real-time video data through a camera sensor and monitor the quality of the video data in real time; A preprocessing module, used to extract frames of the video data according to a preset video frame sampling time interval, and perform preprocessing operations to obtain a key frame image; A model recognition module, used for recognizing potential conflict behaviors through a pre-trained model based on the key frame images; A monitoring module is used to lock the target to be monitored according to the recognition result, and collect the audio data of the target to be monitored, the surrounding audio data, and the behavior data of other passengers; An analysis module is used to analyze and determine whether there are multiple features of conflicting behaviors based on the recognition of multimodal information; The early warning module is used to trigger the early warning mechanism and send out an early warning signal when there is conflicting behavior in the target to be monitored.

9. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method for detecting conflicts between people in a train compartment as claimed in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting conflicts between people in a train compartment as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Passenger behavior detection method and device, electronic equipment and computer readable medium

    CN116994198A

  • Network monitoring method and device, equipment and storage medium

    CN117834808A

  • Warning method, device, storage medium and server regarding physical conflict behavior

    WO2019127271A1