Passenger abnormal behavior recognition method and device in elevator monitoring night vision mode

Through the ‘teacher-student’ network framework and multi-scale category feature comparison learning method, the accuracy and cross-domain detection capabilities of passenger abnormal behavior detection in elevator night vision mode are solved, high-precision abnormal behavior recognition is achieved, and the practicality and safety of elevator monitoring are improved.

CN120452068AActive Publication Date: 2025-08-08ZHEJIANG NEW ZAILING TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510937758.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-08
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The existing elevator monitoring has low detection accuracy for passenger abnormal behavior in night vision mode and insufficient cross-domain detection capabilities, resulting in inaccurate identification and performance losses.

Method used

Using the ‘teacher-student’ network framework, high-quality pseudo-labels are generated by introducing state space auxiliary branches and shared encoder decoders, unsupervised comparison learning of multi-scale category features, and confrontation learning is carried out on the student network to build an object detection network model.

Benefits of technology

It improves the accuracy and robustness of abnormal behaviors of passengers in night vision mode, improves the practicality and safety of intelligent elevator monitoring, and achieves an average accuracy of 72.5%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452068A_ABST
    Figure CN120452068A_ABST
Patent Text Reader

Abstract

The invention discloses a passenger abnormal behavior recognition method and device in an elevator monitoring night vision mode. The method comprises the steps that a passenger abnormal behavior data set is constructed; constructing a target detection network model, which is obtained by the following steps: introducing a state space auxiliary branch on a teacher network, and obtaining a fused pseudo tag through an encoder and a decoder shared with the teacher network; obtaining category information according to the fused pseudo tag, respectively extracting category features of different scales of the backbone network based on the category information, and carrying out unsupervised comparative learning on the category features of different scales; on the student network, predicted category prototype vectors are extracted, and adversarial learning is carried out on the category prototype vectors of the source domain and the target domain through a domain discriminator; constructing a model total loss function; using the passenger abnormal behavior data set to train the passenger abnormal behavior; and identifying a to-be-detected image by using the trained target detection network model to obtain a detection result of the abnormal behavior of the passenger.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of anomaly detection technology, and in particular to a method and device for identifying abnormal passenger behavior in an elevator monitoring night vision mode. Background Art

[0002] With the acceleration of urbanization, vertical elevators have become an indispensable vertical transportation facility in high-rise buildings. However, when elevator monitoring is in night vision mode, the lighting conditions inside and outside the elevator are poor, resulting in low-quality images captured by the cameras. This poses a challenge to detecting unusual passenger behavior. Existing methods for detecting unusual passenger behavior in elevators are typically designed under normal lighting conditions. However, these methods are sensitive to background variations in lighting conditions, resulting in poor detection performance in low light and night vision modes, and their accuracy needs to be improved. In practical applications, the image quality of elevator monitoring in night vision mode is often lower than that in normal lighting. As a result, existing object detectors cannot accurately identify unusual passenger behavior in night vision mode, resulting in low recognition accuracy. Furthermore, the different distribution of image data in night vision mode and normal lighting, such as uneven lighting and lack of color information, can degrade the performance of typical object detectors, leading to suboptimal performance. Furthermore, the class imbalance problem in different scenarios can still affect detector accuracy. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide a method and device for identifying abnormal passenger behavior in the night vision mode of elevator monitoring, so as to solve the problems of inaccurate recognition of abnormal passenger behavior in the night vision mode by existing target detectors and the loss of detector performance caused by different data distribution, and to improve the robustness of abnormal passenger behavior detection in the night vision mode.

[0004] According to a first aspect of an embodiment of the present application, a method for identifying abnormal passenger behavior in an elevator monitoring night vision mode is provided, comprising: S1: Construct a passenger abnormal behavior dataset; S2: Build the target detection network model, which is obtained through the following steps: S21: Introducing a state-space auxiliary branch on the teacher network of the “teacher-student” network framework to obtain fused pseudo-labels through the encoder and decoder shared with the teacher network; S22: obtaining category information according to the fused pseudo-labels, extracting category features of different scales of the backbone network based on the category information, and performing unsupervised comparative learning on the category features of different scales; S23: Extract the predicted category prototype vector from the student network of the “teacher-student” network framework, and perform adversarial learning on the category prototype vectors of the source and target domains through the domain discriminator; S24: Construct the total loss function of the model to complete the construction of the target detection network model; S3: Using the passenger abnormal behavior dataset to train the target detection network model; S4: Use the trained target detection network model to identify the image to be detected and obtain the detection results of abnormal passenger behavior.

[0005] According to a second aspect of an embodiment of the present application, a device for identifying abnormal passenger behavior in an elevator monitoring night vision mode is provided, comprising: The first building module is used to build a passenger abnormal behavior dataset; The second construction module is used to construct a target detection network model, which is obtained by the following steps: introducing a state space auxiliary branch on the teacher network of the "teacher-student" network framework, and obtaining a fused pseudo-label through an encoder and decoder shared with the teacher network; obtaining category information according to the fused pseudo-label, extracting category features of different scales of the backbone network based on the category information, and performing unsupervised comparative learning on the corresponding category features under multiple views; extracting the predicted category prototype vector on the student network of the "teacher-student" network framework, and performing adversarial learning on the category prototype vectors of the source domain and the target domain through a domain discriminator; and constructing a model total loss function to complete the construction of the target detection network model. A training module, configured to train the target detection network model using the passenger abnormal behavior dataset; The recognition module is used to use the trained target detection network model to identify the image to be detected and obtain the detection results of abnormal passenger behavior.

[0006] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.

[0007] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0008] The technical solutions provided by the embodiments of the present application may have the following beneficial effects: This method utilizes a "teacher-student" network framework, introducing a state-space auxiliary branch into the teacher network. By sharing an encoder and decoder, it generates fused pseudo-labels, effectively improving the quality of the pseudo-labels. Unsupervised contrastive learning of multi-scale category features from the fused pseudo-labels is used to distinguish features across scales and categories. Furthermore, adversarial learning is performed on the student network by extracting predicted category prototype vectors to combine the category representations of the source and target domains, enabling the model to maintain consistent representational capabilities across different domains. Finally, through joint optimization of the entire network loss function, the model achieves an overall improvement in target detection capabilities in night vision scenarios. Extensive experiments demonstrate the effectiveness of the proposed method, achieving an average accuracy of 72.5% on a dataset of abnormal passenger behavior in night vision mode.

[0009] The present invention overcomes the problems of large image quality differences, inaccurate pseudo-labels, and weak cross-domain detection capabilities in elevator monitoring night vision environments, and achieves the technical effect of being able to stably and accurately identify abnormal passenger behavior in night vision mode, greatly improving the practicality and safety of intelligent monitoring in night vision scenarios.

[0010] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0012] Figure 1 The present invention is a flowchart of a method for identifying abnormal passenger behavior in an elevator monitoring night vision mode according to an exemplary embodiment.

[0013] Figure 2 The figure is a diagram of the entire network structure used in a method for identifying abnormal passenger behavior in an elevator monitoring night vision mode according to an exemplary embodiment.

[0014] Figure 3 A state space module used in an auxiliary branch of a state space model is shown according to an exemplary embodiment.

[0015] Figure 4 FIG. 1 is a structural diagram of a category prototype adversarial alignment method according to an exemplary embodiment.

[0016] Figure 5 The present invention is a block diagram of a device for identifying abnormal passenger behavior in an elevator monitoring night vision mode according to an exemplary embodiment.

[0017] Figure 6 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0018] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0019] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0020] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0021] Figure 1 This is a flow chart showing a method for identifying abnormal passenger behavior in an elevator monitoring night vision mode according to an exemplary embodiment. Figure 2 FIG. 1 is a diagram showing the entire network structure of a method for identifying abnormal passenger behavior in an elevator monitoring night vision mode according to an exemplary embodiment. Figure 1 and Figure 2 As shown, the method may include the following steps: S1: Construct a dataset of abnormal passenger behavior; this step may include the following sub-steps: S11: Acquire images of passenger abnormal behavior in normal lighting mode and night vision mode respectively; Specifically, a camera mounted on the elevator roof first collects image data under normal lighting conditions. During normal operation, images are automatically captured and selected for passengers in the elevator. Manual screening is then used to identify images of unusual behavior, such as those carrying dangerous items (batteries, electric bicycles, gas cylinders, etc.), generating image data of passengers exhibiting unusual behavior under normal lighting conditions. The camera is then switched to night vision mode, and the above steps are repeated to generate images of passengers exhibiting unusual behavior under night vision mode. This dual-mode image acquisition method is essential for constructing the source and target domains, enabling the model to effectively achieve cross-domain adaptation and accurate detection.

[0022] S12: Labeling dangerous objects in the passenger abnormal behavior image, wherein the labeled content includes a category label and the position coordinates of the four vertices of the dangerous object; Specifically, each image in both normal lighting and night vision modes was manually annotated with detailed labels, including 10 category labels: "motorcycle," "bicycle," "cart," "pet," "luggage," "stroller," "children's scooter," "gas tank," "toy," and "pedestrian," as well as the coordinates of the four vertices of each object's bounding box (i.e., top left, top right, bottom right, and bottom left). The category labels help the model learn to distinguish between different types of dangerous objects, while the coordinates of the four vertices are used to train the model for precise object localization and improve the accuracy of the detection box. Accurate manual labeling is the foundation for generating high-quality pseudo-labels and subsequent technical steps such as multi-scale feature alignment and unsupervised contrastive learning.

[0023] S13: Using the annotated passenger abnormal behavior images under normal lighting and the annotated passenger abnormal behavior images under night vision mode as source domain data and target domain data, respectively; Specifically, the collected and annotated normal-light images serve as the source domain data, serving as the foundation for model learning due to their high image quality and large data volume. Night-vision mode images, on the other hand, serve as the target domain data, representing the low-light challenges faced in actual model deployment environments. This division of the source and target domains helps guide the model through domain adaptation, leveraging features learned from normal-light images to transfer to night-vision imagery scenarios. This design provides a foundation for domain division in subsequent teacher-student training, enhancing the model's effectiveness in real-world nighttime environments.

[0024] S14: Using the source domain data and the target domain data as a passenger abnormal behavior dataset; Specifically, a unified index structure and image-annotation mapping specification are established for the above-mentioned annotated source domain and target domain image data, and the source domain and target domain image data are further divided into training set, validation set and test set respectively, so as to facilitate the call of the training data loading module in the deep learning framework.

[0025] S2: Build the target detection network model. This step may include the following sub-steps: S21: Introduce a state-space auxiliary branch on the teacher network of the “teacher-student” network framework, and obtain fused pseudo-labels through the encoder and decoder shared with the teacher network; this step may include the following sub-steps: S211: Introducing a state space auxiliary branch on the teacher network of the “teacher-student” network framework, wherein the state space auxiliary branch is composed of a state space module and an encoder and a decoder shared by the teacher network; Specifically, the teacher network model uses the ResNet-50 deep convolutional neural network as the backbone network and the Deformable DETR network model as the target detector; the student network model is entirely composed of the Deformable DETR network model. To improve the teacher network's ability to identify abnormal behaviors in complex night vision surveillance images, a state space auxiliary branch is introduced based on the teacher network. This auxiliary branch contains a state space module (refer to Figure 3 ), and shares its encoder and decoder network parameters with the teacher network. The state-space module models sequential features based on the spatiotemporal information state of the input image, enhancing the model's ability to model the relationship between local and global information in the input image, effectively improving the accuracy of pseudo-label generation. The introduction of the state-space auxiliary branch enables the teacher network to better capture contextual information.

[0026] S212: Utilizing the state-space module to receive features extracted by the backbone network of the teacher network, extracting and obtaining global features through the two-dimensional selective scanning mechanism of the state-space module, and feeding the global features into the encoder and decoder shared with the teacher network to obtain generated pseudo labels; Specifically, the state-space module receives intermediate features at different scales extracted by the teacher network backbone (ResNet-50) as state input vectors. Based on this, a two-dimensional selective scanning mechanism is introduced to perform spatial scanning of the input features, filter out key areas in the image, and aggregate spatial context information to extract enhanced global features. This global feature is then input into the encoder and decoder structure shared with the teacher network to complete the pseudo-label prediction process.

[0027] S213: Use non-maximum suppression to fuse the two sets of pseudo labels generated by the teacher network and the auxiliary branch, and finally obtain high-quality fused pseudo labels for training the student model. The calculation formula is: ; Among them, NMS stands for non-maximum suppression, represents the low-quality pseudo labels produced by the teacher network, represents the state space auxiliary branch, Represents the high-quality fused pseudo-label after fusion.

[0028] S22: Obtain category information according to the fused pseudo-labels, extract category features of different scales of the backbone network based on the category information, and perform unsupervised comparative learning on the category features of different scales. This step may include the following sub-steps: S221: On the target domain data, using the fused pseudo-labels and RoIAlign technology, extract the regional features of the target domain data and normalize them according to the category to obtain category information; Specifically, for the high-quality fused pseudo-labels from the teacher network, RoIAlign (Region of Interest Align) technology is used to precisely align and extract features from each bounding box region in the fused pseudo-labels. L2 normalization is then performed to construct a standardized category feature vector. This process ensures feature consistency and lays the foundation for subsequent contrastive learning.

[0029] S222: extracting feature map sets from multiple layers of the backbone network of the teacher network according to the category information, and scaling the bounding boxes accordingly to obtain category features of different scales; Specifically, the backbone of the teacher network typically contains feature maps at multiple scales. To ensure spatial alignment of extracted category features at different scales, the bounding boxes in the original fused pseudo-labels must be scaled accordingly to match the resolution of the corresponding feature maps. This global modeling of multi-scale, multi-category features enhances the model's robustness to changes in target scale, while also achieving more stable category representations through category aggregation, improving contrastive learning.

[0030] S223: Construct positive and negative sample pairs based on category features of different scales, where the positive sample pair is: the features of the same object generated in the teacher network and the student network for the same target image, and the negative sample pair is: any features of different categories or different images from the positive sample; Specifically, the target domain image is fed into the teacher model and the student model, respectively, to extract object features from two different perspectives. For each category, the same object features extracted from the teacher and student networks constitute a positive pair, while features from other categories or arbitrary features from other images serve as negative samples, forming a negative pair. Through the construction of positive and negative pairs, the model learns the differences between categories and the consistency between the same categories in the feature space, helping to improve the ability to express consistent representations across perspectives.

[0031] S224: performing similarity calculation on the positive and negative sample pairs to obtain a contrastive loss of local object features, then averaging the global representations of similar objects to calculate a global contrastive loss, weightedly combining the local loss with the global category loss to form an object-level contrastive learning loss, and finally constructing a final contrastive learning loss; Specifically, the similarity calculation of positive and negative sample pairs is performed to obtain the contrast loss of local object features. , the formula is as follows: ; Where, represents the positive sample, represents negative samples, It is a sample and The cosine similarity measure between is a hyperparameter, It is a , and the temperature The indicator function takes the value of 1, and N is the number of samples; Then the global representations of similar objects are averaged to calculate the global contrast loss , the formula is as follows: ; Where, Represents the contrast loss of local object features, and mean represents the average operation; The local loss is weightedly combined with the global category loss to form the object-level contrastive learning loss , the formula is as follows: ; in, is a hyperparameter; Finally, the final contrastive learning loss is constructed as follows: ; Where, represents the final contrastive learning loss, represents the object-level contrastive learning loss obtained by contrastive learning on the features after the backbone network, represents the object-level contrastive learning loss obtained by contrastive learning on the features after the encoder.

[0032] S23: On the student network of the “teacher-student” network framework, extract the predicted category prototype vector, and perform adversarial learning on the category prototype vectors of the source domain and the target domain through the domain discriminator; this step may include the following sub-steps: S231: The output features obtained by the backbone network of the student network are fed into the encoder-decoder structure of the student network to generate object queries, each query corresponding to a candidate detection object; Specifically, the target domain image is first input into the backbone network of the student model to extract multi-scale spatial semantic feature maps, and then sent to the Transformer encoder in the Deformable DETR architecture for feature modeling. The encoder output features are then interacted with a fixed number of learnable object queries. After processing by the Transformer decoder, each query corresponds to a high-dimensional representation vector of a candidate detection object, namely an object query. Each object query is dynamically matched with the multi-scale encoded features in the decoder to ultimately form a candidate detection object that expresses the potential target area. The benefit of this design is that it can fully utilize the global modeling capabilities of the Transformer, combine multi-scale information to achieve high-quality object modeling, and ensure the accuracy and robustness of subsequent category prototype extraction.

[0033] S232: Determine the predicted category of each candidate detection object query through the detection head, aggregate the query features of the same category, calculate the category feature center, and form a category prototype vector. The formula is as follows: ; in, represents the learned embedding of the object representation obtained by decoding the encoder output, whose dimension is d, the variable c represents the index of a certain category, and the function As an indicator, when The value is 1 when it is negative, otherwise it is 0; S233: Use domain discriminator to determine the source domain of the category prototype vector. This judgment process uses adversarial learning to make the prototype consistent in two domains, the source domain and the target domain; adversarial learning loss function Defined as: ; in, represents the alignment of multi-scale feature level and instance level, P represents the category prototype, D represents the domain discriminator, max represents the maximum value, min represents the minimum value, and the gradient reverse layer is used for minimum-maximum optimization.

[0034] S24: Construct the total loss function of the model to complete the construction of the target detection network model; Specifically, the teacher-student loss function is first constructed, and the teacher network T and the student network S have the same backbone network, encoder and decoder structure. The teacher network processes the weakly enhanced target domain image. , and generate pseudo labels The student network simultaneously processes the strongly enhanced source domain image and target domain images .exist Calculate the supervision loss on , the formula is as follows:

[0035] in, represents the bounding box loss, represents Giou loss, represents the classification loss.

[0036] Target domain Receive supervision from pseudo-labels:

[0037] The teacher network is updated only by the exponential moving average (EMA) of the student network without accumulating gradients:

[0038] in, and denote the parameters of the teacher network and the student network respectively, is a hyperparameter.

[0039] Finally, the overall loss function of the teacher-student framework is:

[0040] in, is a hyperparameter and the teacher network is only updated by the exponential moving average (EMA).

[0041] Secondly, the losses of multiple key modules are weighted and combined to form a complete loss function. The total loss function includes the teacher-student loss , final contrastive learning loss and adversarial learning loss , the specific form is as follows:

[0042] in, 、 is the weight hyperparameter of each loss term.

[0043] By integrating the three mechanisms of teacher-student supervised learning, contrastive learning, and adversarial learning, it can not only utilize the supervision information of the source domain, but also fully tap the structural potential of the target domain, greatly improving the cross-domain detection performance and robustness of the model.

[0044] S3: training the passenger abnormal behavior dataset using the passenger abnormal behavior dataset; Specifically, the constructed target detection network model is trained using the passenger abnormal behavior dataset. This dataset contains various types of abnormal behavior sample images, and collects labeled and unlabeled image data for the source domain and target domain respectively to construct a cross-domain dataset. The training process includes the following steps: First, the source domain image with complete annotations is input into the teacher network, and supervised training is performed using the standard target detection loss function; at the same time, the unlabeled target domain image is input into the student network, and the target domain samples are pseudo-supervised with the help of the high-quality pseudo labels generated by the teacher network; in parallel, the predicted category prototype vectors in the target domain and source domain are extracted and input into the domain discriminator for adversarial learning, so that the semantic feature distribution of the source domain and target domain are gradually aligned at the category granularity (reference Figure 4 ); Finally, an object-level contrastive learning mechanism is introduced into the student network to further enhance the model's discriminative ability for target region representation.

[0045] During the training process, a combination of methods such as teacher-student supervision, feature contrast learning, and category prototype adversarial learning are used to effectively alleviate the migration difficulties caused by the inconsistent data distribution between the source domain and the target domain, thereby improving the detection performance of the model on real target domain data.

[0046] S4: Use the trained target detection network model to identify the image to be detected and obtain the detection results of abnormal passenger behavior.

[0047] Specifically, the image to be detected is fed into a trained object detection model. The model then goes through a series of steps, including backbone network feature extraction, Transformer encoding and decoding, object query matching, and detection head classification and localization. The model ultimately outputs the behavior category and corresponding bounding box for each passenger in the image. The model identifies any abnormal behavior targets in the image and outputs the detection results as an "abnormal category label + location box," meeting the requirements of abnormal passenger behavior detection and enhancing the safety and control capabilities of elevator systems.

[0048] It can be seen from the above embodiments that this application constructs a target detection model for abnormal behavior of elevator passengers in night vision environments, innovatively combines the "teacher-student" framework, state space auxiliary branches, multi-scale category feature contrast learning and category prototype adversarial learning mechanisms, and systematically solves key technical problems such as low accuracy of abnormal behavior detection in night surveillance images, missing labels, and large inter-domain differences from multiple aspects. The trained model can accurately identify abnormal behavior of elevator passengers under night vision conditions and give high-precision position detection results, meeting the practical needs of elevator intelligent security systems for abnormality identification in complex scenarios. This solution not only breaks through the performance bottleneck of traditional single supervision methods in weak labeling environments, but also forms a more systematic technical solution in cross-domain image feature modeling, pseudo-label generation optimization, unsupervised comparative learning, etc., which significantly improves the practicality and reliability of abnormal behavior detection in elevators under night vision environments.

[0049] Corresponding to the aforementioned embodiment of the method for identifying abnormal passenger behavior in the elevator monitoring night vision mode, the present application also provides an embodiment of the device for identifying abnormal passenger behavior in the elevator monitoring night vision mode.

[0050] Figure 5 This is a block diagram of a device for identifying abnormal passenger behavior in an elevator monitoring night vision mode according to an exemplary embodiment. Figure 5 , the device comprises: The first construction module 1 is used to construct a passenger abnormal behavior dataset; The second construction module 2 is used to construct a target detection network model, which is obtained by the following steps: introducing a state space auxiliary branch on the teacher network of the "teacher-student" network framework, and obtaining a fused pseudo-label through an encoder and decoder shared with the teacher network; obtaining category information according to the fused pseudo-label, extracting category features of different scales of the backbone network based on the category information, and performing unsupervised comparative learning on the corresponding category features under multiple views; extracting the predicted category prototype vector on the student network of the "teacher-student" network framework, and performing adversarial learning on the category prototype vectors of the source domain and the target domain through a domain discriminator; constructing a model total loss function, thereby completing the construction of the target detection network model; Training module 3, used to train the target detection network model using the passenger abnormal behavior dataset; The recognition module 4 is used to use the trained target detection network model to recognize the image to be detected and obtain the detection results of abnormal passenger behavior.

[0051] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0052] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0053] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for identifying abnormal passenger behavior in the night vision mode of elevator monitoring. Figure 6 As shown in the figure, it is a hardware structure diagram of a device with data processing capability in which a device for identifying abnormal passenger behavior in an elevator monitoring night vision mode is provided in an embodiment of the present invention. Figure 6 In addition to the processor, memory, DMA controller, disk, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0054] Accordingly, the present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the above-mentioned method for identifying abnormal passenger behavior in the night vision mode of elevator monitoring. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.

[0055] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely as exemplary, and the true scope and spirit of the present application are indicated by the claims.

[0056] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for identifying abnormal passenger behavior in an elevator monitoring night vision mode, characterized in that: include: S1: Construct a passenger abnormal behavior dataset; S2: Build the target detection network model, which is obtained by the following steps: S21: Introducing a state-space auxiliary branch on the teacher network of the "teacher-student" network framework to obtain fused pseudo-labels through the encoder and decoder shared with the teacher network; S22: obtaining category information according to the fused pseudo-labels, extracting category features of different scales of the backbone network based on the category information, and performing unsupervised comparative learning on the category features of different scales; S23: Extract the predicted category prototype vector from the student network of the "teacher-student" network framework, and perform adversarial learning on the category prototype vectors of the source and target domains through the domain discriminator; S24: Construct the total loss function of the model to complete the construction of the target detection network model; S3: Using the passenger abnormal behavior dataset to train the target detection network model; S4: Use the trained target detection network model to identify the image to be detected and obtain the detection results of abnormal passenger behavior.

2. The method according to claim 1, characterized in that Construct a passenger abnormal behavior dataset, including: S11: Acquire images of passenger abnormal behavior in normal lighting mode and night vision mode respectively; S12: Labeling dangerous objects in the passenger abnormal behavior image, wherein the labeled content includes a category label and the position coordinates of the four vertices of the dangerous object; S13: Using the annotated passenger abnormal behavior images under normal lighting and the annotated passenger abnormal behavior images under night vision mode as source domain data and target domain data, respectively; S14: The source domain data and the target domain data are used as a passenger abnormal behavior dataset.

3. The method according to claim 1, characterized in that A state-space auxiliary branch is introduced on the teacher network of the "teacher-student" network framework to obtain fused pseudo labels through the encoder and decoder shared with the teacher network, including: S211: Introducing a state space auxiliary branch on the teacher network of the "teacher-student" network framework, wherein the state space auxiliary branch is composed of a state space module and an encoder and a decoder shared by the teacher network; S212: Utilizing the state-space module to receive features extracted by the backbone network of the teacher network, extracting and obtaining global features through the two-dimensional selective scanning mechanism of the state-space module, and feeding the global features into the encoder and decoder shared with the teacher network to obtain generated pseudo labels; S213: Use non-maximum suppression to fuse the two sets of pseudo labels generated by the teacher network and the auxiliary branch, and finally obtain high-quality fused pseudo labels for training the student model. The calculation formula is: ; Among them, NMS stands for non-maximum suppression, represents the low-quality pseudo labels produced by the teacher network, represents the state space auxiliary branch, Represents the high-quality fused pseudo-label after fusion.

4. The method according to claim 1, wherein Obtaining category information according to the fused pseudo-labels, extracting category features of different scales from the backbone network based on the category information, and performing unsupervised comparative learning on the category features of different scales, including: S221: On the target domain data, using the fused pseudo-labels and RoIAlign technology, extract the regional features of the target domain data and normalize them according to the category to obtain category information; S222: extracting feature map sets from multiple layers of the backbone network of the teacher network according to the category information, and scaling the bounding boxes accordingly to obtain category features of different scales; S223: Construct positive and negative sample pairs based on category features of different scales, where the positive sample pair is: the features of the same object generated in the teacher network and the student network for the same target image, and the negative sample pair is: any features of different categories or different images from the positive sample; S224: Perform similarity calculation on the positive and negative sample pairs to obtain the contrast loss of local object features, then average the global representations of similar objects, calculate the global contrast loss, weightedly combine the local loss with the global category loss to form an object-level contrastive learning loss, and finally construct the final contrastive learning loss, and optimize the overall network through backpropagation.

5. The method according to claim 4, characterized in that The similarity calculation of positive and negative sample pairs is performed to obtain the contrast loss of local object features. Then, the global representation of similar objects is averaged to calculate the global contrast loss. The local loss is weightedly combined with the global category loss to form the object-level contrastive learning loss. Finally, the final contrastive learning loss is constructed, including: Calculate the similarity of positive and negative sample pairs to obtain the contrast loss of local object features , the formula is as follows: ; Where, represents the positive sample, represents negative samples, It is a sample and The cosine similarity measure between is a hyperparameter, It is a , and the temperature The indicator function takes the value of 1, and N is the number of samples; Then the global representations of similar objects are averaged to calculate the global contrast loss , the formula is as follows: ; Where, Represents the contrast loss of local object features, and mean represents the average operation; The local loss is weightedly combined with the global category loss to form the object-level contrastive learning loss , the formula is as follows: ; in, is a hyperparameter; Finally, object-level contrastive learning loss is constructed at different scales to obtain the final contrastive learning loss, which is formulated as follows: ; Where, represents the final contrastive learning loss, represents the object-level contrastive learning loss obtained by contrastive learning on the features after the backbone network, represents the object-level contrastive learning loss obtained by contrastive learning on the features after the encoder.

6. The method according to claim 1, characterized in that On the student network of the "teacher-student" network framework, the predicted category prototype vector is extracted, and the category prototype vectors of the source domain and the target domain are subjected to adversarial learning through the domain discriminator, including: S231: The output features obtained by the backbone network of the student network are fed into the encoder-decoder structure of the student network to generate object queries, each query corresponding to a candidate detection object; S232: Determine the predicted category of each candidate detection object query through the detection head, aggregate the query features of the same category, calculate the category feature center, and form a category prototype vector. The formula is as follows: ; in, represents the learned embedding of the object representation obtained by decoding the encoder output, whose dimension is d, the variable c represents the index of a certain category, and the function As an indicator, when The value is 1 when it is negative, otherwise it is 0; S233: Use the domain discriminator to determine the source domain of the category prototype vector. This judgment process uses adversarial learning to make the prototype consistent in the source domain and the target domain; adversarial learning loss function Defined as: ; in, represents the alignment of multi-scale feature level and instance level, P represents the category prototype, D represents the domain discriminator, max represents the maximum value, min represents the minimum value, and the gradient reverse layer is used for minimum-maximum optimization.

7. In the method according to claim 1, the total model loss includes teacher-student loss, contrastive learning loss, and prototype adversarial loss, and each loss term uses a corresponding hyperparameter weight to control its contribution ratio.

8. A device for identifying abnormal passenger behavior in an elevator monitoring night vision mode, characterized in that: include: The first building module is used to build a passenger abnormal behavior dataset; The second construction module is used to construct a target detection network model, which is obtained by the following steps: introducing a state space auxiliary branch on the teacher network of the "teacher-student" network framework, obtaining a fused pseudo-label through an encoder and decoder shared with the teacher network; obtaining category information based on the fused pseudo-label, extracting category features of different scales of the backbone network based on the category information, and performing unsupervised comparative learning on the corresponding category features under multiple views; extracting the predicted category prototype vector on the student network of the "teacher-student" network framework, and performing adversarial learning on the category prototype vectors of the source domain and the target domain through a domain discriminator; and constructing a model total loss function to complete the construction of the target detection network model. A training module, configured to train the target detection network model using the passenger abnormal behavior dataset; The recognition module is used to use the trained target detection network model to identify the image to be detected and obtain the detection results of abnormal passenger behavior.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Remote sensing image unsupervised cross-domain target detection method based on progressive pseudo tag

    CN117830616A

  • Method for reconstructing abnormal behavior detection model of generative adversarial network

    CN117876959A

  • Lung adenocarcinoma wettability evolution classification method and system based on state space network

    CN118552758A

  • Arbitrary modal assistance-based camouflage target segmentation method

    CN120107975A

  • Passive field adaptive target detection model generation method based on class prototype alignment and target detection method

    CN120147606A