Face recognition and voice integrated monitoring system for train
By using a facial recognition and voice integrated monitoring system on the train, combined with visual and voice acquisition modules, the system analyzes personnel density and abnormal areas, solving the problem of limited information acquisition in existing train monitoring systems and enabling rapid and reliable anomaly handling and passenger guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU RAILWAY VOCATIONAL & TECH COLLEGE
- Filing Date
- 2023-05-22
- Publication Date
- 2026-04-24
AI Technical Summary
Existing train monitoring systems rely on limited information acquisition, have low reliability, and are not quick or efficient in handling abnormal events.
The integrated monitoring system combines a visual acquisition module and a voice module. It acquires visual information of the interior of the carriage and personnel characteristics, analyzes personnel density, and obtains voice information in case of anomalies. The system then processes and stores the information using a storage module.
It improves the reliability and speed of handling abnormal events, reduces the difficulty and time of identifying people, and enhances the passenger experience.
Smart Images

Figure CN121921817A_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of train monitoring technology, and specifically to a train facial recognition and voice integrated monitoring system. Background Technology
[0002] With the improvement of living standards, people's travel needs are increasing. With the development of rail transit, the comfort, safety and timeliness of rail trains have been significantly improved. Therefore, rail trains have become the preferred mode of transportation for people.
[0003] Currently, in order to further improve the safety of rail trains, all trains are equipped with monitoring systems. These systems can monitor and store the behavior of passengers and crew. On the one hand, the monitoring system serves as a deterrent and reminder to prevent illegal incidents. On the other hand, it can also monitor and store video footage of the train, providing evidence for handling such cases and thus greatly promoting the safe operation of the train.
[0004] However, most train monitoring systems currently use cameras, which is too simplistic, records limited information, and has low reliability. Furthermore, when handling abnormal events, it is necessary to retrieve stored video footage and then select the appropriate video based on the time period, which is not conducive to the rapid handling of abnormal events. Summary of the Invention
[0006] In view of the above problems, this application provides a train face recognition and voice integrated monitoring system. The integrated monitoring system, which combines visual acquisition and voice prompts, can collect facial and appearance information through the visual acquisition module and voice information of the target personnel through the voice acquisition module. By acquiring multiple information, the reliability of the reference is improved.
[0007] This invention provides a train facial recognition and voice integrated monitoring system, comprising: The visual acquisition module is installed inside the carriage and is used to identify image information inside the carriage and the facial and physical features of people inside the carriage. The analysis module is used to acquire and analyze the vehicle interior image information acquired by the visual acquisition module to obtain the personnel density distribution inside the vehicle compartment, and to obtain the location information of the abnormal personnel density area when the personnel density distribution inside the vehicle compartment is abnormal. A voice module, installed inside the carriage and connected to the analysis module, is used to obtain the location information when there is an area in the carriage where the distribution of personnel density is abnormal, so as to obtain the voice information of the area corresponding to the location information. A storage module, connected to the analysis module and the voice module, is used to store some information collected by the visual acquisition module and information collected by the voice module.
[0008] Furthermore, the analysis module is connected to the control center through the data transceiver module to obtain the identity information of the people entering the station obtained by the station control center. The identity information includes facial information, physical feature information, train number information, and seat position information.
[0009] Furthermore, when there is an abnormal distribution of personnel density inside the carriage, the visual acquisition module is also used to identify the facial and physical features of each person in the area where the abnormal distribution of personnel density occurs.
[0010] Furthermore, the analysis module is also used to process and analyze the voice information collected by the voice module to obtain sensitive words within the voice information. Based on the frequency of occurrence of the sensitive words, the corresponding voice information and video information are classified into levels, and the storage module stores them according to the classification levels.
[0011] Furthermore, the control center includes a second visual acquisition module located at the entrance. The second visual acquisition module is used to collect facial and body information of people entering the station, and compare the facial information with identity information to obtain the train number and seat information of the person. After determining the seat information, the seat information is bound to the body information, and then sent to the analysis module through the data transceiver module. The analysis module analyzes whether the seat position of the person inside the corresponding carriage is correct based on the body information corresponding to the seat information.
[0012] Furthermore, in cases where there is an anomaly where the body posture information does not match the seating arrangement information of passengers inside the carriage, the facial information of the passenger in the anomaly is collected by the visual acquisition module to reconfirm the passenger's identity. When the passenger and seating arrangement do not match, a voice reminder is issued by the voice module.
[0013] Furthermore, after the visual acquisition module collects the facial information of the passenger who is experiencing an abnormal situation, the identity information and seat information of the passenger are obtained through the facial information. The voice module issues a voice reminder message including the passenger's name, seat number, and seat number error information, thereby reminding the passenger to sit in the correct seat.
[0014] Furthermore, after acquiring the voice information, the analysis module binds and stores the voice information and the corresponding video information, and stores the information according to the level of the information classification by the analysis module. When the storage space is insufficient, the stored information is discarded according to the classification level.
[0015] Beneficial effects This invention provides a train facial recognition and voice integrated monitoring system, comprising: a visual acquisition module, installed inside the train carriage, for recognizing image information inside the carriage and facial and physical features of people inside the carriage; an analysis module, for acquiring and analyzing the image information inside the train carriage acquired by the visual acquisition module to obtain a passenger density threshold, and for acquiring location information of the area where the passenger density exceeds the threshold when the passenger density exceeds the threshold; and a voice module, installed inside the train carriage and connected to the analysis module, for acquiring the location information when the passenger density exceeds the threshold area, and acquiring voice information of the area corresponding to the location information. The system includes a storage module connected to the analysis module and the voice module. This module stores some information collected by the visual acquisition module and the information collected by the voice module. This configuration allows the visual acquisition module to recognize image and facial information inside the carriage. By analyzing and processing the information collected by the visual acquisition module, areas with abnormal conditions inside the carriage can be identified. Then, the voice module acquires the voice information of the abnormal areas and stores it in the storage module. The system can collect facial and appearance information through the visual acquisition module and voice information of the target personnel through the voice acquisition device. When handling abnormal conditions, multiple pieces of information can be referenced to improve the reliability of the reference.
[0016] On the other hand, by first obtaining the person's body posture information, and then analyzing the body posture information through the analysis module to determine the person's identity information, it is not necessary to continue obtaining the person's facial information, thereby reducing the difficulty and time required to determine the person's identity information. Attached Figure Description
[0017] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.
[0018] Figure 1 This is a schematic diagram of the structure of a train face recognition and voice integrated monitoring system provided by the present invention.
[0019] Figure 2 This invention provides a diagram of the voice module of a train face recognition and voice integrated monitoring system.
[0020] Figure 3 This is a flowchart illustrating a train face recognition and voice integrated monitoring system provided by the present invention. Detailed Implementation
[0021] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0023] Example 1 This invention provides a train facial recognition and voice integrated monitoring system, referenced Figure 1-3 As one specific implementation method, the system includes: The visual acquisition module is installed inside the carriage and is used to identify image information inside the carriage and the facial and physical features of people inside the carriage. The analysis module is used to acquire and analyze the vehicle interior image information acquired by the visual acquisition module to obtain the personnel density threshold inside the vehicle compartment, and to acquire the location information of the abnormal personnel density area when the personnel density distribution inside the vehicle compartment is abnormal. A voice module, installed inside the carriage and connected to the analysis module, is used to obtain the location information when a region in the carriage has a population density exceeding a threshold, so as to obtain the voice information of the region corresponding to the location information. A storage module, connected to the analysis module and the voice module, is used to store some information collected by the visual acquisition module and information collected by the voice module.
[0024] Specifically, the visual acquisition module consists of multiple cameras evenly distributed throughout each carriage. These cameras capture facial and body information of people entering and exiting the carriages, as well as images of the carriage interior. The facial and body information of people entering and exiting the carriages is used to obtain their identity information. Preferably, the order of acquiring this information is as follows: first, acquire body information, including clothing, accessories, body shape, and hairstyle. If the body information analysis module can determine the person's identity, then acquiring facial information is unnecessary, thus reducing the difficulty and time required for identification. Furthermore, when a person's identity cannot be determined through their body posture information, facial information is obtained through a camera. This facial information is then analyzed and processed to obtain the person's identity. After confirming the person's identity, their seating information is obtained. Video tracking is then used to verify if the person is seated correctly. If the person is seated incorrectly, a voice prompt is played via a voice module to remind them to change seats. The voice prompt includes the person's name, seat information, and information about the incorrect seating. This method guides passengers to their correct seats, improving the passenger experience. The voice prompt module includes several voice playback devices installed on each row of seats.
[0025] Furthermore, after passengers are correctly seated, during train operation, the visual acquisition unit continues to monitor the interior of the carriage via video, and the analysis module analyzes and processes the data to obtain the distribution of passengers inside the carriage. When an anomaly occurs in the distribution of passengers inside the carriage, the visual acquisition unit promptly acquires video information of the area with the anomaly, obtains facial and body information of each person in that area, and collects audio information from a voice module located in that area. This voice module includes multiple voice acquisition devices distributed throughout the carriage. The analysis module analyzes and processes the collected audio information, classifying the video and audio segments according to the results of the analysis and processing. The storage module then stores the video and audio segments according to the classification. The analysis module processes the video information inside the carriage to obtain the distribution of passengers inside the carriage. The method for processing the internal personnel distribution to obtain the personnel density threshold in the carriage includes the following steps: a) Acquire video information inside the carriage by the visual acquisition unit, and acquire video frame images within the video information at predetermined time intervals T. Perform grayscale processing on the video frame images to obtain a grayscale image of the personnel distribution inside the carriage; then use the planting filter method to perform image noise reduction processing; b) Perform background separation on the processed image to generate a binarized image, remove invalid information in the image, and obtain a grayscale image of the personnel distribution inside the carriage; c) Divide the processed grayscale image into n (n>2) regions according to the cell, then obtain the number of pixels in each region, then set the initial threshold γ, then obtain the number of regions with more than γ pixels n1 and the number of regions with less than γ pixels n2 in each grayscale image, and then calculate the reference coefficient μ, where μ is calculated as: μ=&[(n-n1) / (n-n2)) -3 ]*|n1- n2| / [(n1+n2) / n], where & is the adjustment coefficient, with a value range of 0.37-2.46. Specifically, when the value range of T is between 0 and 1 second, the value range of & is 0.37-0.65; when the value range of T is between 1 and 2 seconds, the value range of & is 0.66-1.78; and when the value range of T is between 2 and 3 seconds, the value range of & is 1.79-2.46.
[0026] d. Then, obtain the reference coefficient value of each frame image. Based on the distribution of people in the carriage in each frame image, establish a reference database of the distribution of people inside the carriage and the value of the reference coefficient. In practical applications, by obtaining the reference coefficient μ, it is possible to determine whether there is an abnormal distribution of people in the carriage within the corresponding time period of the current video frame. After an abnormality occurs, the area with the largest change in grayscale value in the video frame is obtained as the target area. Image information, facial information and voice information of people in this area are then obtained.
[0027] Furthermore, the analysis module is also used to process and analyze the voice information collected by the voice module to obtain sensitive words within the voice information. Based on the frequency of occurrence of the sensitive words, the corresponding voice information and video information are classified into levels, and the storage module stores them according to the classification levels.
[0028] Furthermore, the voice module includes several microphones evenly distributed throughout the carriage. When an area with abnormal personnel distribution is identified, the microphones in that area are controlled to acquire the voice information of the personnel within that area, and the voice information is analyzed and processed. The voice analysis and processing method includes the following steps: Step 1, establishing a sensitive word database; Step 2, analyzing the acquired voice information to obtain the number of times and frequency of sensitive words appearing in the voice information; Step 3, classifying the level according to the frequency and number of times sensitive words appear. Specifically, the level classification parameter is β, where the acquired voice information includes the number of speakers in the predetermined area (N), the number of speakers whose voices contain sensitive words (N1), the frequency of sensitive words appearing in the area within a unit time T1 (H), the relationship coefficient (θ) with a value range of 0.76-0.96, the value range of T1 (10-15 seconds), and the ratio of the area where the voice is acquired to the total area of the carriage (I). Then, β = I * [(N1 / N)H] 1 / 2 *θ, when the value of β is less than 0.4, the video and audio information acquired in this area is classified as the first level for storage; when the value of β is between 0.4 and 1, it is classified as the second level; when the value of β is greater than 1, it is classified as the third level. After classifying the information into levels, the information in the first level is stored using the highest level storage method. When the storage space of the subsequent storage system is insufficient, the data in the third level is discarded first, and then the data in the second level is discarded. The data in the first level needs to be confirmed by the administrator before it can be discarded. In this way, the storage device can automatically clean up the data, ensure the storage space for important data, and at the same time ensure the reliability of the data storage of important data.
[0029] Furthermore, as a specific implementation method, the analysis module described herein is connected to the control center through a data transceiver module to obtain the identity information of the people entering the station obtained by the station control center. The identity information includes facial information, physical feature information, train number information, and seat position information.
[0030] Specifically, when entering the station, passengers need to go through the station security check system. At this time, the second vision acquisition module set up at the entrance acquires the facial information, physical characteristics, and seat position information of the passengers entering the station. Then, based on each passenger's seating information, the identity information of each passenger entering the station is transmitted to the analysis module in the corresponding carriage. The passenger information is stored through the storage module, so that the information can be used to confirm the identity information of the passengers in the carriage during subsequent monitoring.
[0031] Furthermore, when the personnel density inside the carriage exceeds a predetermined threshold, the visual acquisition module is also used to identify the facial and physical features of individuals within the area where the personnel density threshold exceeds the predetermined threshold. Specifically, the visual acquisition module acquires images of the personnel distribution inside the carriage, processes the images to obtain the personnel density distribution, and when an abnormal personnel density distribution occurs, it acquires the facial and physical features of individuals within the target area of the abnormal density. First, it obtains the identity information of the corresponding individuals through physical features; if the identity information cannot be accurately determined through physical features, it obtains the identity information of the individuals based on facial features.
[0032] Furthermore, the control center includes a second visual acquisition module located at the entrance. The second visual acquisition module is used to collect facial and body information of people entering the station, and compare the facial information with identity information to obtain the train number and seat information of the person. After determining the seat information, the seat information is bound to the body information, and then sent to the analysis module through the data transceiver module. The analysis module analyzes whether the seat position of the person inside the corresponding carriage is correct based on the body information corresponding to the seat information. Specifically, the second vision acquisition module is a camera installed at the entrance. This module acquires facial and body information of passengers entering the station and connects to the gate to verify their identity. It obtains the identity and travel information of each passenger and binds the facial and body information to this data. This binding is then sent to the corresponding train's analysis module. To improve data transmission accuracy, the data transmission module includes the following steps: ① Binding the passenger's identity, travel, facial, and body information, packaging it into data, and generating an identification ID number based on the passenger's train information; ② Sending the processed data to a wireless transmitter via serial communication; ③ The wireless transmitter receives the packaged data, matches the identification ID number with the ID number of the corresponding carriage receiver, and sends it to the corresponding carriage; ④ The carriage receiver receives the data and uses the analysis module to decompress and restore it, thereby obtaining the internal information.
[0033] Example 2 Furthermore, as a specific implementation, the voice module also includes voice playback devices located throughout the carriage. In cases where there is an anomaly where the passenger's posture information does not match their seating position, the visual acquisition module captures the facial information of the passenger in the anomaly to reconfirm their identity. When the passenger and their seat do not match, the voice module issues a voice reminder. Specifically, when a passenger enters the carriage, the visual acquisition module captures their facial and posture information to determine their seating position. The system continuously monitors the passenger via video, analyzing whether their seat is correct when they are seated. If the passenger's actual seat does not match their ticket information, the voice playback device promptly plays a prompt to remind the passenger to change seats. This method ensures passengers can quickly and correctly board, avoiding unnecessary disputes caused by incorrect seating and improving the overall travel experience.
[0034] Furthermore, as a specific implementation, after the visual acquisition module collects the facial information of the passenger experiencing the abnormal situation, it obtains the passenger's identity and seat information through the facial information. The voice module then issues a voice reminder message including the passenger's name, seat number, and incorrect seat information, thereby reminding the passenger to take the correct seat. Playing the passenger's name through the voice module better reminds the passenger, and subsequently playing the information about the passenger's incorrect seat and correct seat allows the passenger to quickly change seats.
[0035] Furthermore, as a specific implementation, after acquiring the voice information, the analysis module binds and stores the voice information and the corresponding video information, and stores the information according to the level of the information classification by the analysis module, and discards the stored information according to the classification level when the storage space is insufficient.
[0036] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A train facial recognition and voice integrated monitoring system, characterized in that, include: The visual acquisition module is installed inside the carriage and is used to identify image information inside the carriage and the facial and physical features of people inside the carriage. The analysis module is used to acquire and analyze the vehicle interior image information acquired by the visual acquisition module to obtain the personnel density distribution inside the vehicle compartment, and to obtain the location information of the abnormal personnel density area when the personnel density distribution inside the vehicle compartment is abnormal. A voice module, installed inside the carriage and connected to the analysis module, is used to obtain the location information when a region with a personnel density exceeding a threshold appears in the carriage, so as to obtain the voice information of the region corresponding to the location information. A storage module, connected to the analysis module and the voice module, is used to store some information collected by the visual acquisition module and information collected by the voice module.
2. The train face recognition and voice integrated monitoring system according to claim 1, characterized in that, The analysis module is connected to the control center through the data transceiver module and is used to obtain the identity information of the people entering the station obtained by the station control center. The identity information includes facial information, physical feature information, train number information, and seat position information.
3. The train face recognition and voice integrated monitoring system according to claim 2, characterized in that, When an abnormal distribution of personnel density occurs inside the carriage, the visual acquisition module is also used to identify the facial and physical features of each person in the area where the abnormal personnel density occurs.
4. A train face recognition and voice integrated monitoring system according to claim 1 or 2, characterized in that, The analysis module is also used to process and analyze the voice information collected by the voice module to obtain sensitive words within the voice information. Based on the frequency of occurrence of the sensitive words, the corresponding voice information and video information are classified into levels, and the storage module stores them according to the classification levels.
5. The train face recognition and voice integrated monitoring system according to claim 4, characterized in that, The control center includes a second visual acquisition module located at the entrance. This module collects facial and body information of passengers entering the station and compares the facial information with their identity information to obtain their train number and seat number. After determining the seat number, the module binds it to the body information and then sends it to the analysis module via the data transceiver module. The analysis module analyzes whether the passengers' seats in the corresponding carriages are correct based on the body information corresponding to the seat number.
6. The train face recognition and voice integrated monitoring system according to claim 5, characterized in that, In the event of an anomaly where the body posture information does not match the seating arrangement information of passengers inside the carriage, the facial information of the passenger in the anomaly is collected by the visual acquisition module to reconfirm the passenger's identity. When the passenger and seating arrangement do not match, a voice reminder is issued by the voice module.
7. The train face recognition and voice integrated monitoring system according to claim 6, characterized in that, After the visual acquisition module collects the facial information of the passenger who is experiencing an abnormal situation, it obtains the passenger's identity information and seat information through the facial information. The voice module issues a voice reminder message including the passenger's name, seat number, and seat number error information, thereby reminding the passenger to sit in the correct seat.
8. A train face recognition and voice integrated monitoring system according to any one of claims 4-7, characterized in that, After acquiring the voice information, the analysis module binds and stores the voice information and the corresponding video information, and stores the information according to the level of the information classification by the analysis module. When the storage space is insufficient, the stored information is discarded according to the classification level.