A personnel and vehicle monitoring method and system based on capsule network
Through the capsule network fusion of sound and image features, the problem of difficult multimodal features in traditional security monitoring is solved, and high accuracy recognition is achieved in complex environments, and learning ability and adaptability are achieved.
Patent Information
- Application Number
- CN202111221927.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-10-20
AI Technical Summary
Among the existing security monitoring technologies, single-modal recognition methods are prone to misjudgment, especially in the case of light differences and unclear image acquisition, the recognition accuracy is difficult to improve, and traditional technologies are difficult to effectively integrate sound and image features, and lack learning ability and growth.
The capsule network is used to fusion process the sound and image features, and the low-level feature extraction and high-level abstract feature fusion are fusion, and the time-frequency change features are extracted using convolutional neural networks and recurrent neural networks, and the efficient fusion and recognition of multimodal features are achieved through iterative dynamic routing algorithms.
It improves the accuracy of personnel and vehicle identification, can accurately identify targets in complex environments, has learning ability and growth ability, and adapts to changes in different scenarios.
Smart Images

Figure CN114022726B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of security monitoring, and particularly to a method and system for monitoring personnel and vehicles based on capsule network. Background Art
[0002] With the continuous acceleration of the urbanization process in China, security monitoring and emergency management in urban streets, industrial parks, and residential communities are important tasks related to social security and stability. In this field, finely monitoring personnel, vehicles, and their interrelationships, and promptly discovering abnormal information and giving alarms are important methods for improving the level of security monitoring and the main direction of technological development.
[0003] Existing recognition technologies mostly use single-modal recognition methods based on visual feature analysis to recognize the facial features and clothing features of personnel; however, in practice, most monitored audio and video cannot capture clear face images; moreover, environmental differences such as indoor and outdoor lighting also affect the judgment of clothing colors. It can be seen that if only image recognition technology is used, it is easy to cause misjudgment, resulting in the inability to improve the accuracy of security monitoring recognition.
[0004] However, a large amount of practical data proves that natural human voices, such as accents, voiceprints, etc., are key information for verifying personnel identities and have more advantages in scenarios where personnel's appearance and clothing are disguised; on the other hand, vehicle engine sounds, tire-ground friction sounds, and horn sounds can directly or indirectly reflect key features such as the vehicle's condition and age. Using these features can improve the accuracy of vehicle type recognition.
[0005] However, the analysis of sound signals is ignored in existing recognition technologies, which limits the further improvement of recognition accuracy. Moreover, there is little research on recognition based on the fusion of audio and video features currently. The main difficulties encountered are that the features of different modalities are not easy to fuse; secondly, traditional technologies need to use relatively complete images for analysis, but the monitoring system usually only collects the front and side images of vehicles, with fewer features, reducing the recognition accuracy; furthermore, traditional technologies only use traditional image processing and matching technologies, which are difficult to perform fine classification and do not have learning ability and growth potential. Summary of the Invention
[0006] In order to overcome the deficiencies of the prior art, one of the purposes of the present invention is to provide a method for monitoring personnel and vehicles based on capsule network. Using image features and sound features for target recognition and classification can improve the recognition accuracy. At the same time, the sound and image features are fused, realizing the fusion of multi-modal features and solving the problem that it is difficult to fuse multi-modal features in traditional technologies.
[0007] Another purpose of the present invention is to provide a system for monitoring personnel and vehicles based on capsule network.
[0008] A third object of the present invention is to provide an electronic device.
[0009] A fourth object of the present invention is to provide a storage medium.
[0010] The first object of the present invention is achieved by the following technical solution:
[0011] A method for monitoring personnel and vehicles based on a capsule network, comprising:
[0012] Collecting monitoring audio and video, separating the monitoring audio and video to obtain sound data and image data;
[0013] Respectively extracting features from the sound data and the image data, the feature extraction including the extraction of low-level features and high-level abstract features; fusing the abstract features of the sound and the image obtained after feature extraction to generate a feature fusion vector, and using a capsule network to process the feature fusion vector to identify the feature information of personnel and vehicles in the monitoring audio and video;
[0014] Performing data comparison processing on the identified feature information of personnel and vehicles and pre-registered information to output the monitoring results of target personnel and target vehicles.
[0015] Further, the method for extracting features from the sound data is:
[0016] Analyzing the separated sound data to obtain the time, frequency, and amplitude parameters of the sound signal;
[0017] Generating a spectrogram of the sound signal in the frequency and time dimensions according to the time, frequency, and amplitude parameters of the sound signal, so that the two-dimensional spectrogram contains low-level features in the frequency domain and the time domain;
[0018] Using a convolutional neural network to extract the time-frequency change features in the spectrogram, and then using a recurrent neural network to extract the context-related features in the time domain of the time-frequency change feature map output by the convolutional neural network to output an abstract sound feature map.
[0019] Further, the method for extracting features from the image data is:
[0020] Performing image feature extraction on the image data to obtain low-level visual features and generating an image feature map carrying the low-level visual features;
[0021] Using a convolutional layer to process the image feature map to output an abstract image feature map.
[0022] Further, the method for fusing the abstract features of the sound and the image is:
[0023] Normalize the maximum and minimum values of each pixel point in the abstract sound feature map and the abstract image feature map respectively;
[0024] Project the normalized sound features and image features into a unified feature space to obtain the transformed sound and image features;
[0025] Fuse the transformed sound and image features to obtain a feature fusion vector.
[0026] Further, the method for processing the feature fusion vector using a capsule network is as follows:
[0027] Package the obtained feature fusion vector into low-level capsules and high-level capsules according to different dimensions;
[0028] Implement the transfer of feature vectors between low-level capsules and high-level capsules through an iterative dynamic routing algorithm to finally determine the output of the high-level capsules.
[0029] Further, after fusing the abstract features of sound and image, it further includes:
[0030] Adopt a cross-entropy function to analyze the semantic consistency classification deviation between sound and image, correct the obtained deviation, and then update the sound data and image data to re-identify and classify personnel and vehicles.
[0031] Further, after processing the feature fusion vector using a capsule network, it further includes:
[0032] Calculate the system loss using a margin loss function, correct the system loss, and then update the sound data and image data to re-identify and classify personnel and vehicles.
[0033] The second object of the present invention is achieved by the following technical solution:
[0034] A personnel and vehicle monitoring system based on a capsule network, which executes the personnel and vehicle monitoring method based on a capsule network as described above. The monitoring system includes:
[0035] A monitoring audio and video processing module, which is used to acquire monitoring audio and video, and separate the monitoring audio and video to obtain sound data and image data;
[0036] A feature extraction module, which is used to perform low-level feature extraction on the sound data and image data to obtain a spectrogram and an image feature map, then perform abstract feature processing on the spectrogram and the image feature map, and fuse the obtained abstract sound feature map and abstract image feature map to obtain a feature fusion vector;
[0037] A capsule network processing module, which is used to process the feature fusion vector by using the capsule network to identify the feature information of personnel and vehicles in the monitored audio and video;
[0038] A target object monitoring module, which is used to perform data comparison processing on the identified feature information of personnel and vehicles with pre-registered information to output the monitoring results of target personnel and target vehicles.
[0039] The third object of the present invention is achieved by the following technical solution:
[0040] An electronic device, which includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method for monitoring personnel and vehicles based on a capsule network is implemented.
[0041] The fourth object of the present invention is achieved by the following technical solution:
[0042] A storage medium, on which a computer program is stored. When the computer program is executed, the above-mentioned method for monitoring personnel and vehicles based on a capsule network is implemented.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] Low-level features and high-level abstract features of sound and image data are extracted, and then the high-level abstract features of sound and image are feature-fused, which solves the problem that it is difficult to perform multi-modal feature fusion in traditional technologies; at the same time, low-level features are extracted first, and then the capsule network is used to process the high-level abstract features of sound and image. Compared with the traditional recognition technology based on convolutional neural networks, the input of the capsule network used in the present invention is a fused high-level abstract feature vector / matrix, which carries the spatial relationship between local and local, local and global, so that the present invention can maximize the retention of local and global features in the original data at the same time; in addition, due to the increase in the dimension of the fused vector / matrix, it is easier to aggregate and classify features, so as to significantly improve the recognition accuracy and finally obtain accurate recognition results. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic flowchart of the method for monitoring personnel and vehicles based on a capsule network of the present invention;
[0046] Figure 2 It is an overall flowchart block diagram of the method for identifying personnel and vehicles of the present invention;
[0047] Figure 3 It is a block diagram of a two-layer network structure of CNN+RNN for sound feature extraction of the present invention;
[0048] Figure 4Block diagram of the capsule network for classifying high-dimensional fusion features in the present invention;
[0049] Figure 5 Module schematic block diagram of the personnel and vehicle monitoring system based on the capsule network in the present invention. Specific implementation manners
[0050] Next, in combination with the accompanying drawings and specific implementation manners, the present invention will be further described. It should be noted that, on the premise of non-conflict, the following-described embodiments or technical features can be arbitrarily combined to form new embodiments.
[0051] Embodiment 1
[0052] This embodiment provides a personnel and vehicle monitoring method based on a capsule network. This method is applied in the field of security monitoring and can finely identify the identities of personnel and vehicle models to achieve the purpose of timely finding and tracking target persons and vehicles. As Figure 1 、 Figure 2 shown, the personnel and vehicle monitoring method of this embodiment specifically includes the following steps:
[0053] Step S1: Collect monitoring audio and video, and separate the monitoring audio and video to obtain sound data and image data.
[0054] In this embodiment, video images can be collected through monitoring devices such as cameras, and the on-site audio needs to be included in the video images so that the monitoring audio and video contain image data and sound data; among them, the sound data contains natural voice information of personnel, such as the tone, timbre, and accent of speech; the sound data also contains unique audio information of vehicles, such as engine sounds, etc. The image data contains detailed visual information such as the appearance and clothing of people, as well as the appearance and license plate number of vehicles.
[0055] This embodiment uses a microphone array hardware system and combines a beamforming algorithm to perform sound source localization, thereby identifying the sources of different sounds.
[0056] Step S2: Respectively perform feature extraction on the sound data and image data. The feature extraction includes the extraction of low-level features and high-level abstract features; perform fusion on the abstract features of the sound and image obtained after feature extraction to generate a feature fusion vector, and use the capsule network to process the feature fusion vector to identify the feature information of personnel and vehicles in the monitoring audio and video.
[0057] The method for performing feature extraction on the sound data in this embodiment is:
[0058] Step S211: Analyze the separated sound data to calculate the time, frequency, and amplitude parameters of the sound signal;
[0059] Step S212: Generate a spectrogram of the sound signal in the frequency and time dimensions based on the time, frequency, and amplitude parameters of the sound signal, so that the two-dimensional spectrogram contains low-level features in the frequency domain and the time domain;
[0060] A spectrogram is a two-dimensional image of a sound signal in the frequency and time dimensions, containing low-level features in the frequency domain and the time domain, such as the tone of a person's speech and the voiceprint of an engine. The abscissa of the spectrogram represents time, the ordinate represents frequency, and the gray value represents the sound amplitude; in this embodiment, the time-frequency characteristics of the audio signal are visualized to generate a two-dimensional spectrogram, the purpose of which is to unify the format of the sound and the image feature map.
[0061] Step S213: Use a convolutional neural network to extract the time-frequency change features in the spectrogram, and then use a recurrent neural network to extract the context-related features in the time domain from the time-frequency change feature map output by the convolutional neural network to output an abstract sound feature map.
[0062] A spectrogram represents a composite image of the power distribution of sound information in the time domain and the frequency domain. Since a convolutional neural network (hereinafter referred to as CNN for short) can extract local features about time-frequency from the spectrogram through training data, but cannot express time context-related information. And a recurrent neural network (hereinafter referred to as RNN for short) solves this deficiency through a dynamically changing context window in the time domain. Therefore, this embodiment uses a two-layer network of CNN+RNN to extract the abstract features of sound. Specifically: first, use CNN to extract the time-frequency change features in the spectrogram, that is, extract the local texture features of the spectrogram. Then, use RNN to process and express the context-related features of the sound signal in the time domain.
[0063] In order to improve the recognition accuracy, this embodiment optimizes the fusion structure of CNN and RNN as follows:
[0064] As Figure 3 shown, a spectrogram with a size of v×h is input into CNN for processing, the ordinate represents the frequency range (0 to v), and the abscissa represents the time range (0 to h); the bottom layer of CNN consists of several time-domain convolutional layers (the number of layers is adjusted according to the actual situation), and each layer uses a convolutional kernel of size 1×a to perform convolutional operations along the time axis of the spectrogram, with a sliding step of 1. After the time-domain convolutional layer, a frequency-domain convolutional layer is connected, and three different sizes of convolutional kernels (b1×h, b2×h, b3×h) are used to perform convolutional operations along the frequency domain, with a sliding step of 1. Different sizes of convolutional kernels are beneficial to extracting complementary features in different regions of the feature spectrogram, improving the accuracy of the system for classifying and recognizing sound signals. The last layer of CNN is a pooling layer, which uses the maximum pooling method to reduce redundant information and reduce the computational amount of subsequent processing of the system.
[0065] The bottom layer of the RNN is a Bi-directional Long Short-Term Memory (BLSTM) layer, which extracts temporal context information from the feature map after CNN pooling. Then, the attention layer uses a weight allocation mechanism to make the network pay more attention to key information and improve the system performance. Finally, a Dropout operation is performed in the fully connected layer to avoid overfitting.
[0066] In this embodiment, in addition to processing the low-level features and high-level abstract features of the voice data, it is also necessary to process the low-level features and high-level abstract features of the image data. Specifically:
[0067] Step S221: Extract image features from the image data to obtain low-level visual features and generate an image feature map carrying the low-level visual features;
[0068] In this embodiment, several feature extraction methods are comprehensively used for the image data to generate a feature map expressing low-level features, which includes features such as the contour, texture, and local color of the target. Among them, the main methods for extracting image features include the HOG method (Histogram of Oriented Gradient), the LBP method (Local Binary Patterns), and the SIFT method (Scale-invariant feature transform). The Histogram of Oriented Gradient can effectively describe the gradient texture features of local images and is mainly used to extract the contour information of the target. The Local Binary Patterns calculates the gray comparison information of pixel points and their neighborhoods in the statistical image and extracts the texture information of the image. This method has strong stability, is less affected by illumination changes in most application scenarios, and also has strong rotation invariance. The Scale-invariant feature transform method processes the image with Gaussian functions having different standard deviations, finds extreme points in the spatial scale, and extracts their positions, scales, and rotation invariants, thereby extracting local features. Its advantage is that it can reduce the interference of problems such as scale and angle transformation, gray level, and noise of the image during the processing process.
[0069] The above image feature extraction methods can comprehensively use one or more of the above three methods according to the specific usage scenario to extract the low-level visual features in the image.
[0070] Step S222: Process the image feature map using a convolutional layer to output an abstract image feature map.
[0071] The image feature map carrying low-level features obtained in the above steps contains features such as the morphology and clothing of people, and the outline, local texture, and color of vehicles; then, the low-level feature map is processed through a convolutional layer to extract a high-level abstract feature map.
[0072] After the above feature extraction steps in this embodiment, an abstract sound feature map and an abstract image feature map can be obtained. The abstract sound feature map and the abstract image feature map are projected into a feature space with unified dimensions and formats for feature fusion. After fusion, the feature dimension is increased, which is more convenient for aggregation and classification.
[0073] Before feature fusion, the maximum and minimum values of each pixel point in the abstract sound feature map and the abstract image feature map are normalized respectively, that is, the original data values are mapped to the interval [0,1]. On the one hand, it is convenient for the fusion of sound and image data, and on the other hand, it reduces the negative impact of some large-value data on the overall model.
[0074] Since the prerequisite for feature fusion is the same dimension, therefore, the normalized sound features and image features need to be projected into a unified feature space to obtain the transformed sound and image features; specifically: the sound feature E A =[e a1 ,e a2 ,…,e am T is an m-dimensional vector, and the image feature is E V =[e v1 ,e v2 ,…,e vn T is an n-dimensional vector, and the dimension of the unified feature space is l, where l≥n>m. The following formula is used for transformation:
[0075] E′ A =W A ·E A
[0076] E′ V =W V ·E V ;
[0077] In the formula, E′ A and E′ V are the transformed sound and image features respectively, both of which are l-dimensional vectors. W A and W V are l×m and l×n dimensional transformation matrices respectively, which are determined through model training.
[0078] Subsequently, the transformed sound and image feature vectors are fused to obtain a feature fusion vector, and the fused vector is expressed as:
[0079] E′ = [E′ A , E′ V 。
[0080] In this embodiment, a capsule network is used to calculate the fused high-dimensional feature vectors. As Figure 4 shown, the obtained feature fusion vectors are encapsulated into low-level capsules and high-level capsules according to different dimensions; the transfer of feature vectors between low-level capsules and high-level capsules is realized through an iterative dynamic routing algorithm to finally determine the output of high-level capsules. The capsule network can make full use of the spatio-temporal relationships between various features, obtain the positional relationships between high-level features and low-level features, obtain more accurate personnel feature recognition and vehicle type classification results, and at the same time include clear spatial positional relationships.
[0081] The specific processing flow is as Figure 4 shown. E′ with dimension l = [E′ A , E′ V contains high-dimensional fusion features of personnel and vehicle sound information and is encapsulated into low-level capsules according to different dimensions. The high-level capsules finally represent the feature categories, the relationships between features, and are directly related to the final refined monitoring results.
[0082] The dynamic routing algorithm is between the low-level and high-level capsules, which helps the high-level capsules better learn the low-level features. u i is the local information represented by the i-th low-level capsule, is the predicted value of the overall information of the j-th high-level capsule under the i-th low-level capsule. The overall and local information is transformed through the affine transformation matrix W ij as follows:
[0083]
[0084] The input of the j-th high-level capsule comes from the weighted predicted values of all low-level capsules, and the weighting coefficient is defined as the coupling coefficient C ij , as follows:
[0085]
[0086] The coupling system number C ij is continuously updated with the dynamic routing algorithm. The algorithm uses the Softmax function, as follows:
[0087]
[0088] b ik The weight coefficient to be updated, and the iterative formula is as follows:
[0089]
[0090]
[0091] After the system is trained and iterated, the output V of the high-level capsule is finally determined. j .
[0092] In order to improve the accuracy and robustness of recognition in this embodiment, it is necessary to perform deviation analysis on the classification results, conduct training, and act on the data preprocessing stage, continuously update the coefficients of subsequent functional modules, and further improve the accuracy of subsequent processing.
[0093] The classification results of the system need to meet the semantic consistency of sound and image, otherwise obvious deviations will occur in the results. Semantic consistency requires that the sound and image data are first consistent in spatio-temporal coordinates. For example, the time difference between the two types of data will cause misjudgment of the results. Secondly, the image may interfere with the judgment of speech. For example, the expression and lip shape of the speaker may cause misreading of the sound content. After high-dimensional feature fusion, the present invention uses the cross-entropy function to analyze the classification deviation of semantic consistency and update the system parameters in real time. The function is expressed as:
[0094]
[0095] In the formula, y is the measured value of the semantic consistency classification task, 1 indicates consistency, and 0 indicates inconsistency. is the predicted value of the semantic consistency classification.
[0096] In addition, after the capsule network performs fine classification, the edge loss function is used to calculate the loss of the system. The function is as follows:
[0097]
[0098] In the formula, m + and m - are values set according to the actual situation. Generally, m + = 0.9, m - = 0.1. k is the number of classifications. If the feature belongs to the kth class, then T k = 1, otherwise T k = 0. ||v c || can represent the probability that the feature belongs to the kth class. λ is used to adjust the proportion, and the initial value is set to 0.5.
[0099] After correcting the system deviation through the above method in this embodiment, the sound data and image data are updated to re-recognize and classify personnel and vehicles, gradually improving the recognition accuracy.
[0100] Step S3: Compare the feature information of the recognized personnel and vehicles with the pre-registered information for data comparison processing to output the monitoring results of the target personnel and target vehicles.
[0101] The system in this embodiment uses a capsule network to finely identify people and vehicles. The main information includes personal identity information, accent characteristics and timbre, license plate numbers, vehicle brands and models, engine sound characteristics, tire-ground friction sound characteristics, and vehicle condition information indirectly reflected, etc. These information are comprehensively processed with the information registered by relevant departments, and the person-vehicle association information is compared to "portray" the target in all aspects, and abnormal information such as disguise, forgery, and suspicion can be identified in time and an alarm can be given. The method in this embodiment can provide detailed information about people and vehicles for units engaged in security monitoring, and can be used for feature analysis, trajectory query, timely discovery of abnormal vehicles and alarm.
[0102] Embodiment 2
[0103] This embodiment provides a personnel and vehicle monitoring system based on a capsule network, which executes the personnel and vehicle monitoring method based on a capsule network as described in Embodiment 1. As Figure 5 shown, its monitoring system includes:
[0104] A monitoring audio and video processing module, which is used to obtain monitoring audio and video, and separate the monitoring audio and video to obtain sound data and image data;
[0105] A feature extraction module, which is used to perform low-level feature extraction on the sound data and image data to obtain a spectrogram and an image feature map, then perform abstract feature processing on the spectrogram and the image feature map, and fuse the abstract sound feature map and the abstract image feature map obtained from the abstract feature processing to obtain a feature fusion vector;
[0106] A capsule network processing module, which is used to process the feature fusion vector by using a capsule network to identify the feature information of people and vehicles in the monitoring audio and video;
[0107] A target object monitoring module, which is used to perform data comparison processing on the feature information of the identified people and vehicles with the pre-registered information to output the monitoring results of the target people and target vehicles.
[0108] This system performs fusion processing on sound and image features, uses a capsule network for identification and fine classification, and has a simple system structure with the following advantages:
[0109] 1. In the data preprocessing stage of this system, the monitoring signal is separated into sound and image, and the sound feature is fully utilized in the subsequent process for target recognition and classification. The sound feature is added in the target monitoring process, which can improve the accuracy of target detection.
[0110] 2. In the low-level feature extraction stage of this system, the original data of sound and image are respectively subjected to low-level feature extraction. The sound signal is transformed into a two-dimensional spectrogram, and the frequency-domain and time-domain features are extracted; the image signal is processed by feature extraction methods such as HOG, LBP, and SIFT to form a low-level feature map of the image. Furthermore, CNN and RNN are used to process the spectrogram, and the convolutional layer is used to process the low-level feature map of the image to respectively obtain high-level abstract features and express them in the same form, completing the fusion of multi-modal features and solving the problem that it is difficult to perform multi-modal feature fusion in traditional technologies. In addition, compared with the traditional convolutional neural network-based recognition technology, the input of the capsule network used in the present invention is a fused high-level abstract feature vector / matrix, which can perform comprehensive classification on multi-modal features. Since the dimension of the fused vector / matrix increases, it is easier to perform aggregated classification of features, so the recognition accuracy is significantly improved.
[0111] 3. This system first extracts low-level features and then uses the capsule network to process the high-level abstract features of sound and image. Compared with the traditional neural network, the input of the capsule network is a vector / matrix, which carries the spatial relationship between local and local, and between local and global. Therefore, this system can maximize the simultaneous retention of local and global features in the original data and solve the problem of the deficiency in traditional technologies that it is difficult to take into account both local and global features. Further, this system uses a correction system to feedback and analyze the recognition results, improving the recognition accuracy.
[0112] 4. This system introduces a capsule network, which retains information such as the pose, position, size, and rotation of the input object. It can still correctly recognize the same object that has undergone operations such as translation, rotation, and scaling, and then generalize the learning results to new objects and scenarios. It has the advantage of a small amount of training data, strengthening the generalization ability to a certain extent and solving the problem of the weak generalization ability of traditional technologies.
[0113] Embodiment III
[0114] This embodiment provides an electronic device, which includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the personnel and vehicle monitoring method based on the capsule network in Embodiment I; in addition, this embodiment also provides a storage medium, on which a computer program is stored, and when the computer program is executed, it implements the above-mentioned personnel and vehicle monitoring method based on the capsule network.
[0115] The device and storage medium in this embodiment and the method in the foregoing embodiments are two aspects based on the same inventive concept. The process of the method implementation has been described in detail before, so those skilled in the art can clearly understand the structure and implementation process of the device and storage medium in this embodiment according to the foregoing description. For the sake of brevity of the specification, it will not be repeated here.
[0116] The above embodiments are only preferred embodiments of the present invention, and the scope of protection of the present invention cannot be limited thereby. Any non-substantive changes and substitutions made by those skilled in the art based on the present invention fall within the scope of protection required by the present invention.
Claims
1. A method for monitoring personnel and vehicles based on capsule network, characterized in that, Including: Collecting monitoring audio and video, separating the monitoring audio and video to obtain sound data and image data; Performing feature extraction on the sound data and image data respectively, and the feature extraction includes the extraction of low-level features and high-level abstract features; Fusing the abstract features of the sound and image obtained after feature extraction to generate a feature fusion vector, analyzing the semantic consistency classification deviation between the sound and image using the cross-entropy function, correcting the analyzed deviation, updating the sound data and image data to re-identify and classify personnel and vehicles, and processing the feature fusion vector using a capsule network to identify the feature information of personnel and vehicles in the monitoring audio and video, encapsulating the fused feature fusion vector into low-level capsules and high-level capsules according to different dimensions, and realizing the transfer of feature vectors between the low-level capsules and high-level capsules through an iterative dynamic routing algorithm to finally determine the output of the high-level capsules; comparing the identified feature information of personnel and vehicles with the pre-registered information for data comparison processing to output the monitoring results of target personnel and target vehicles.
2. The method for monitoring personnel and vehicles based on a capsule network according to claim 1, wherein The method for performing feature extraction on the sound data is: analyzing the separated sound data to obtain the time, frequency, and amplitude parameters of the sound signal; generating a spectrogram of the sound signal in the frequency and time dimensions according to the time, frequency, and amplitude parameters of the sound signal, so that the two-dimensional spectrogram contains low-level features in the frequency domain and time domain; using a convolutional neural network to extract the time-frequency change features in the spectrogram, and then using a recurrent neural network to extract the context-related features in the time domain of the time-frequency change feature map output by the convolutional neural network to output an abstract sound feature map.
3. The personnel and vehicle monitoring method based on capsule network according to claim 1, characterized in that The method for performing feature extraction on the image data is: performing image feature extraction on the image data to obtain low-level visual features and generating an image feature map carrying the low-level visual features; using a convolutional layer to process the image feature map to output an abstract image feature map.
4. The method for monitoring personnel and vehicles based on capsule network according to claim 2 or 3, characterized in that The method for fusing the abstract features of the sound and image is: respectively normalizing the maximum and minimum values of each pixel point in the abstract sound feature map and the abstract image feature map; projecting the normalized sound features and image features into a unified feature space to obtain the transformed sound and image features; Fusing the transformed sound and image features to obtain a feature fusion vector.
5. The method for monitoring personnel and vehicles based on a capsule network according to claim 1, characterized in that, After processing the feature fusion vector using a capsule network, it further includes: calculating the system loss using an edge loss function, correcting the system loss, and updating the sound data and image data to re-identify and classify personnel and vehicles.
6. A personnel and vehicle monitoring system based on a capsule network, characterized in that, Execute the personnel and vehicle monitoring method based on capsule network according to any one of claims 1 to 5. The monitoring system includes: a monitoring audio and video processing module for acquiring monitoring audio and video and separating the monitoring audio and video to obtain sound data and image data; a feature extraction module for performing low-level feature extraction on the sound data and image data to obtain a spectrogram and an image feature map, then performing abstract feature processing on the spectrogram and the image feature map, and fusing the obtained abstract sound feature map and abstract image feature map to obtain a feature fusion vector, analyzing the semantic consistency classification deviation of sound and image using the cross-entropy function, correcting the obtained deviation, and updating the sound data and image data to re-identify and classify personnel and vehicles; a capsule network processing module for processing the feature fusion vector using the capsule network to identify the feature information of personnel and vehicles in the monitoring audio and video, encapsulating the obtained feature fusion vector into low-level capsules and high-level capsules according to different dimensions, and realizing the transfer of feature vectors between the low-level capsules and the high-level capsules through an iterative dynamic routing algorithm to finally determine the output of the high-level capsules; a target object monitoring module for performing data comparison processing on the identified feature information of personnel and vehicles with pre-registered information to output the monitoring results of target personnel and target vehicles.
7. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it realizes the personnel and vehicle monitoring method based on capsule network according to any one of claims 1 to 5.
8. A storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed, it realizes the personnel and vehicle monitoring method based on capsule network according to any one of claims 1 to 5.
Citation Information
Patent Citations
Emotion recognition method and device based on voice
CN110288974A
Audio and video data processing method and system, electronic equipment and storage medium
CN111461235A
Campus safety management method, system and device and storage medium
CN111582042A