Method and system for detecting illegal residents in three small places based on lightweight multi-modal hybrid model
Through the lightweight multimodal hybrid model combined with timing and image data processing network, the accuracy and efficiency of shop residents' identification at night is solved, and efficient online identification and early warning is achieved, which is suitable for power data platforms and urban power supervision platforms.
Patent Information
- Application Number
- CN202510497214.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-15
AI Technical Summary
When identifying whether shops have night stays, the existing technology has problems such as limited coverage and inaccurate judgments. In particular, there is a lack of efficient identification methods that combine shop behavior characteristics and can be embedded in online systems, resulting in a high misjudgment rate.
The lightweight multimodal hybrid model is used to collect load data through non-invasive collection, combine research data to determine the night time period, and use the pre-trained timing-image fusion model to identify online identification and early warning of whether there is night-time human habitation in the shop, including the time series data processing network, image data processing network and multimodal fusion layer identification method.
It realizes accurate identification of shops at night, reduces the false alarm rate, and is convenient to deploy the model, with low floating-point calculation and parameters. It is suitable for large-scale promotion and application of power data platforms and urban power supervision platforms.
Smart Images

Figure CN120496121A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent electricity consumption monitoring, and specifically relates to a method and system for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model. Background Art
[0002] Traditional monitoring methods rely on manual inspections or fixed-time power thresholds, resulting in limited coverage and inaccurate judgments. Intelligent analysis of load data has become a new trend in recent years, but existing methods suffer from crude nighttime segmentation and a single identification model. This inability to accurately reflect the actual operation and occupancy of shops leads to a high rate of false positives. In particular, there is a lack of efficient identification methods that incorporate the behavioral characteristics of shops and can be embedded in online systems, making widespread application in real-world urban safety monitoring difficult.
[0003] By observing the nighttime power data of a large number of shops and combining it with the survey results, it was found that there are significant differences in the performance of nighttime power curves between occupied shops and unmanned shops. The nighttime power curves of unmanned shops are usually relatively stable or show regular fluctuations, while the nighttime power curves of occupied shops often show large-scale irregular fluctuations or abnormal power mutations. Based on this observation, it was found that by analyzing the nighttime power series of shops and their corresponding power curves, it is possible to establish a method to identify whether shops are occupied at night. At present, in order to identify whether shops or small places are occupied at night, the only way is to train an image binary classification model using labeled nighttime power curves of shops. However, this method is not accurate enough, and there is an urgent need for a recognition method that considers a wider range of dimensions. Summary of the Invention
[0004] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a method and system for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model. By non-invasively collecting load data and combining it with survey data to determine the night time period, the pre-trained time series-image fusion model is used to realize online identification and early warning of whether there is any nighttime occupancy in shops.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model, comprising the following steps:
[0007] Collecting target user data and preprocessing the target user data, wherein the target user data includes an original time series power sequence and a judgment label;
[0008] Divide the user's night time period and obtain the time series power sequence of the user during the night time period;
[0009] Identify illegal occupants based on the user's nighttime time series power sequence based on a pre-trained time series-image hybrid model, comprising a time series data processing network, an image data processing network, and a multimodal fusion layer.
[0010] If the recognition result is "resident", the detection result will be recorded and an alarm message will be issued.
[0011] As a preferred technical solution, the step of collecting target user data and preprocessing the target user data includes:
[0012] The missing values in the collected target user electricity consumption data are detected and filled with cubic spline interpolation, as shown in the following formula:
[0013] At the known data points (x0,y0),(x1,y1),……,(x n ,y n ), construct a piecewise cubic polynomial:
[0014] S i (x) = a i +b i (xx i )+c i (xx i ) 2 +d i (xx i ) 3
[0015] Among them, a i , b i , c i and d i is the segmentation coefficient;
[0016] Label each target user data sample and obtain the judgment label.
[0017] As a preferred technical solution, the method of dividing the user into nighttime periods and obtaining the time-series power sequence of the user in the nighttime period is specifically as follows:
[0018] The user's non-business hours are set as the user's nighttime electricity usage hours, and the electricity usage data within the user's nighttime electricity usage hours are selected to obtain the user's nighttime electricity usage data; the user's non-business hours include fixed hours on non-consecutive days and fixed hours on consecutive days, and the fixed hours are adjusted according to the season.
[0019] As a preferred technical solution, the pre-training of the time series-image hybrid model includes the following steps:
[0020] Perform data enhancement processing on the original time series power sequence to obtain the set fine-grained time series power data;
[0021] Performing format conversion on the set fine-grained time-series power data to obtain a power image set, wherein the format conversion includes: dynamically adjusting the size of the power image according to the sequence length, and filling the area between the power image and the coordinate axis with visual elements, wherein the dynamically adjusted size is proportional to the sequence length;
[0022] A training set is constructed using the power image set and its corresponding judgment labels, and the time series-image hybrid model is pre-trained using the training set.
[0023] As a preferred technical solution, the data enhancement processing is specifically as follows:
[0024] Random scaling: Randomly scale the original time series power sequence, with the scaling ratio range being [0.9, 1.1] or [0.95, 1.05];
[0025] Time jitter: Random noise is added to the original timing power sequence to simulate measurement errors or environmental interference.
[0026] Time warping: nonlinearly stretch the original time power sequence and set the maximum warping range;
[0027] Mirror flip: mirror flip the original timing power sequence;
[0028] Random truncation: The original time series power sequence is randomly truncated, and the truncation ratio does not exceed 5%.
[0029] As a preferred technical solution, the time series data processing network includes multiple groups of one-dimensional convolutions, activation functions, maximum pooling layers, and fully connected layers; the multiple groups of one-dimensional convolutions run in parallel, and the convolution kernel sizes of each group are different. The activation function is the ReLU activation function, and the maximum pooling layer is a global maximum pooling layer, which is used to reduce the dimension of the convolution layer output and obtain key time series features. The fully connected layer is used to map the time series features and align them with the image data processing network output;
[0030] Each group of one-dimensional convolutions is connected to an activation function, and the features output by the activation function are concatenated along the channel dimension.
[0031] As a preferred technical solution, the image data processing network includes a feature extraction module, a global average pooling layer and a fully connected layer; the feature extraction module includes a multi-layer feature extraction layer, an inverted residual structure and an SE attention module, which is used to extract high-level features of the image, the global average pooling layer is used to reduce the dimensionality of the high-level features of the image, and the fully connected layer is used to map the image features and align them with the output of the time series data processing network.
[0032] As a preferred technical solution, the multimodal fusion layer includes a splicing layer, a fusion feature optimization layer, a classification output layer and a loss function; the splicing layer is used to splice the outputs of the time series data processing network and the image data processing network to obtain fusion features, the fusion feature optimization layer includes a dropout layer and a fully connected layer for learning high-order interactions of fusion features, and the classification output layer includes a softmax activation function for outputting category probabilities.
[0033] As a preferred technical solution, the pre-trained time series-image hybrid model is used to identify illegal residents during the night time period based on the time series power sequence of the user, including:
[0034] Convert the user's time-series power sequence during the night time period into a format to obtain a power image set;
[0035] Using a time series data processing network to process the user's time series power sequence during the nighttime period to obtain a time series feature vector, and using an image data processing network to process the power image set to obtain an image feature vector; the time series feature vector is aligned with the image feature vector;
[0036] The multimodal fusion layer is used to concatenate the time series feature vector and the image feature vector to obtain the fused feature vector. After learning the high-order interactions of the fused feature vector, the fused feature vector is mapped to a low-dimensional classification space and the output category probability is calculated.
[0037] The cross entropy loss function is used as the optimization target to guide the parameter update during the training process. At the same time, the network parameters are continuously adjusted through the gradient descent optimization strategy until the parameters converge and the training is completed.
[0038] The trained time series-image hybrid model is used to identify the target user data and determine whether the user is living there illegally.
[0039] In a second aspect, the present invention further provides a system for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model, which is applied to the aforementioned method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model, and includes a preprocessing module, a data partitioning module, a detection and identification module, and an information communication module;
[0040] A preprocessing module, configured to collect target user data and preprocess the target user data, wherein the target user data includes an original time series power sequence and a judgment label;
[0041] The data partitioning module is used to divide the user's night time period and obtain the time series power sequence of the user during the night time period;
[0042] A detection and recognition module is used to identify illegal occupants based on the user's time series power sequence during the night time period based on a pre-trained time series-image hybrid model, which includes a time series data processing network, an image data processing network, and a multimodal fusion layer.
[0043] The information communication module is used to analyze the recognition results. If the recognition result is "people living", the detection result will be recorded and an alarm message will be issued.
[0044] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0045] The time-series-image hybrid model designed in this paper processes the user's time-series power sequence during the nighttime hours. It can integrate time-series and image information, fully exploit data features, and effectively distinguish between residents and non-residential businesses. Furthermore, this invention is robust to non-human fluctuations such as those caused by water heaters and air conditioners, reducing false alarm rates.
[0046] The time series-image hybrid model is easy to deploy: the model's floating-point computing capacity (FLOPs) is only 110MB, and the parameter capacity (Params) is only 1.01MB. It can be embedded in various power data platforms or urban electricity consumption supervision platforms, supporting large-scale promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0048] Figure 1 This is a flowchart of a method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to an embodiment of the present invention;
[0049] Figure 2 This is a schematic diagram of filling missing data values according to an embodiment of the present invention;
[0050] Figure 3 This is a diagram of the architecture of a time series-image hybrid model according to an embodiment of the present invention;
[0051] Figure 4 This is a diagram of the network structure of the feature fusion layer of an embodiment of the present invention;
[0052] Figure 5 A schematic diagram of a power sequence generation image according to an embodiment of the present invention;
[0053] Figure 6 The loss curve and accuracy curve of the model pre-training process in the embodiment of the present invention;
[0054] Figure 7 Schematic diagram of the structure of the three-small-space illegal occupancy detection system based on the lightweight multimodal hybrid model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0056] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0057] See also Figure 1 This embodiment provides a method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model, including the following steps:
[0058] S1. Collect target user data and pre-process the target user data, wherein the target user data includes an original time series power sequence and a judgment label.
[0059] First, it is necessary to collect the target user's electricity consumption data. In this embodiment, the system automatically reads the user's load data with a 1-minute sampling frequency from 8:01 the previous day to 8:00 the current day at 8:00 a.m. every day. The sampling frequency can also be adjusted according to actual needs.
[0060] Preprocess the target user data, such as Figure 2 As shown. The preprocessing method includes: detecting missing values in the collected target user electricity consumption data and filling them with cubic spline interpolation. Assume that the target user's power data sequence P for one day is {p1, p2, ..., p n}, where p i Is the power value at the i-th time point. For the short-term missing segment that meets the interpolation conditions, the power data at a certain time point is missing y t , select the known data points before and after this time period (x i ,y i) to fit and ensure smooth interpolation. n ,y n ), construct a piecewise cubic polynomial:
[0061] S i (x) = a i +b i (xx i )+c i (xx i ) 2 +d i (xx i ) 3
[0062] Among them, a i , b i , c i and d i It is a piecewise coefficient that ensures that the interpolation points at the missing values are continuous with the existing data points through interpolation constraints, which is better than the ability of linear interpolation to restore the fluctuation pattern.
[0063] Furthermore, regarding the acquisition of judgment tags:
[0064] In each user's sample data, samples in each night period are manually labeled as "occupied" or "unoccupied" (binary label, 1 means occupied, 0 means unoccupied), and the labels are derived from the survey results.
[0065] S2. Divide the user into nighttime periods and obtain the time-series power sequence of the user in the nighttime period.
[0066] The raw time-series power sequence acquired and preprocessed in step S1 above is extracted, and the data belonging to the nighttime period is separated. In the early stage, through on-site surveys of 210 target "three small" places, 10 days of non-business hours and occupancy / unoccupancy information data for these 210 stores were collected and recorded. This non-business period was set as the nighttime detection period for the stores, which served as the subsequent judgment window.
[0067] Each store sets a fixed night time in summer [T start夏 ,T end夏 ], set a fixed night time period in non-summer [T start非夏 ,T end非夏 ].
[0068] The power data corresponding to the nighttime period in the real-time user power data is extracted and a power curve image is generated. The input power curve image is normalized and resized (e.g., 224×224×3) and used as the model input. See step S3 for details.
[0069] S3. Identify illegal occupants based on the user's time series power sequence during the night time period based on a pre-trained time series-image hybrid model to obtain a recognition result. The time series-image hybrid model includes a time series data processing network, an image data processing network, and a multimodal fusion layer.
[0070] like Figure 3 As shown, in this embodiment, the lightweight InceptionTime network and MobileNetV3 network are selected to build a multimodal hybrid model. The model floating point calculation amount (FLOPs) is 110MB, and the parameter amount (Params) is only 1.01MB. Among them, the time series data processing network adopts the InceptionTime branch to extract the multi-scale features of the power time series data. The structure of the InceptionTime branch is as follows:
[0071] 1) Multi-scale convolutional feature extraction: Three sets of parallel one-dimensional convolutions (1D-CNN) are used, with kernel sizes of 3×1, 5×1, and 7×1, respectively, to capture power variation patterns at different time scales. The three sets of output features are concatenated along the channel dimension to obtain a multi-scale fused representation.
[0072] 2) Activation function: Improve nonlinear modeling capabilities through the ReLU activation function;
[0073] 3) Max Pooling Layer: Global Max Pooling is used to extract the most significant features globally, reducing the dimensionality of the convolutional layer output while retaining key features, improving model computational efficiency and reducing overfitting.
[0074] 4) Fully Connected Layer: Maps the temporal features output by InceptionTime to a 128-dimensional feature vector, aligning it with the output dimension of the image branch to facilitate subsequent multimodal fusion.
[0075] The image data processing network uses the MobileNetV3 branch to extract the spatial visual features of the power image. The main structure of MobileNetV3 includes:
[0076] 1) MobileNetV3 pre-trained feature extraction layer: Use the ImageNet pre-trained MobileNetV3-Small model, retain the 16-layer feature extraction layer of MobileNetV3-Small (including the inverted residual structure and SE attention module), remove the original fully connected classification layer, and extract high-level features of the image;
[0077] 2) Global Average Pooling: This layer reduces the dimension of the feature map to 576 dimensions while enhancing the stability of the model. The input is the load curve image (RGB image, size 224×224×3) generated by the nighttime load sequence.
[0078] 3) Fully connected layer (FC): Maps the image features extracted by MobileNetV3 to 128 dimensions and aligns them with the temporal feature vector output by the InceptionTime branch.
[0079] The multimodal fusion layer fuses the two branch features through feature splicing and inputs them into the fully connected layer to achieve classification. The specific process is as follows:
[0080] 1) Feature concatenation: concatenate the 128-dimensional temporal features and the 128-dimensional image features into a 256-dimensional fusion vector;
[0081] 2) Fusion feature optimization: A Dropout layer is added to suppress overfitting. The fusion vector is implemented through a mapping module consisting of two fully connected layers (ReLU activation) to achieve nonlinear high-order feature interactive learning. During the learning process, 256-dimensional features are mapped to 64-dimensional features.
[0082] 3) Classification output layer: The final fully connected layer maps the 64-dimensional features to a 2-dimensional classification space. The Softmax activation function is used to output the category probabilities, where the probability distribution of the two categories [P_No People, P_Occupied] is output. For example, a value of P_Occupied > 0.5 is considered to be occupied.
[0083] The cross-entropy loss function is used as the optimization objective to guide parameter updates during the training process, enabling the model to more accurately distinguish between illegal residents and non-illegal residents. The model continuously adjusts network parameters through the gradient descent optimization strategy to improve classification performance. After training is completed, the optimal model weights are saved and used as the online detection model, such as Figure 6 shown.
[0084] In order to accurately process the original time series power sequence and generate the corresponding image, this implementation pre-trains the time series-image hybrid model. The steps are as follows:
[0085] S31. Perform data enhancement processing on the original time series power sequence to obtain set fine-grained time series power data.
[0086] Among them, the specific data enhancement processing methods include:
[0087] (1) Random scaling: The time series is randomly scaled with a scaling ratio between [0.9, 1.1] or [0.95, 1.05] to simulate power fluctuations of different amplitudes;
[0088] (2) Time jitter: random noise is added to the time series with a noise level of 0.5 (occupied scene) and 0.1 (unoccupied scene) to simulate measurement errors or environmental interference;
[0089] (3) Time warping: nonlinear stretching of the time series with a maximum warping range of 5 (occupied scene) and 2 (unoccupied scene) to simulate changes in time scale;
[0090] (4) Mirror flip: Mirror flip the time series to increase data diversity;
[0091] (5) Random truncation: The time series is randomly truncated with a truncation ratio not exceeding 5% to simulate data loss or partial failure.
[0092] Through the above enhancement method, the dataset was expanded to 7760 items, including 3525 items of illegal residence data and 4235 items of non-illegal residence data.
[0093] S32. Perform format conversion on the set fine-grained time-series power data to obtain a power image set. The format conversion includes: dynamically adjusting the size of the power image according to the sequence length, and filling the area between the power image and the coordinate axis with visual elements. The dynamically adjusted size is proportional to the sequence length.
[0094] Furthermore, if Figure 5 As shown, the power image set is obtained by the following methods, including:
[0095] (1) Dynamically adjust image size: Dynamically adjust the image width according to the sequence length to ensure that the image can fully reflect the characteristics of the time series. The image height is fixed at 8, and the width is proportional to the sequence length (len(sequence) / 50);
[0096] (2) Draw the power curve: the curve color is blue, and the area between the curve and the X-axis is filled with light blue to enhance the visual effect of the image;
[0097] (3) Hiding coordinate axes: In order to reduce the interference of irrelevant information on model training, the coordinate axes of the image are hidden.
[0098] Through the above processing, the power series data is converted into an image data set.
[0099] S33. Construct a training set using the power image set and its corresponding judgment labels, and use the training set to pre-train the time series-image hybrid model.
[0100] Supervised learning is used during model training, using the above labels to guide the model to learn how to map input data to the correct category;
[0101] The training goal is to minimize the cross entropy loss function (Cross Entropy Loss), the formula is as follows:
[0102] Loss=-[y·log(P)+(1-y)·log(1-P)]
[0103] Where y is the true label (1 if someone lives there, 0 if no one lives there), and P is the probability that the model output is “someone lives there”.
[0104] After training is complete, when the model predicts new nighttime data for a store, it outputs a binary classification probability, for example:
[0105] [P_No one, P_Occupant] = [0.2, 0.8] → judged as occupied;
[0106] [P_No one, P_People] = [0.9, 0.1] → judged as no one;
[0107] The system can set a judgment threshold (such as 0.5) as an alarm condition, or adjust it to a stricter level (such as 0.6 or 0.7) based on actual business needs.
[0108] It's worth explaining that the pre-trained time-series-image hybrid model automatically learns distinct characteristic patterns. For example, it learns which are "natural fluctuations from appliances like refrigerators and lights" and which are "power usage fluctuations caused by human activity" (such as intermittent water heater startup, concentrated nighttime appliance startup, and high fluctuations). Therefore, in practice, the model won't mistakenly identify people as living simply because of regular fluctuations in power usage, such as from refrigerators or small lights. The "occupied" and "unoccupied" categories output by the model are equivalent to pattern recognition of "suspicious occupancy behavior characteristics" and "no obvious occupancy behavior characteristics."
[0109] S4. If the recognition result is "resident", the detection result is recorded and an alarm message is issued.
[0110] The extracted nighttime power data and the corresponding power curve Figure 1 The system inputs a pre-trained time-series-image hybrid model, which extracts and fuses features internally, outputting a judgment result (occupied / unoccupied). If the model determines "occupied," the system automatically marks the shop as suspected of illegal occupation, generates a record of illegal occupation, and sends an alert to the monitoring platform.
[0111] Example 2.
[0112] In order to better illustrate the beneficial effects of the present invention, this embodiment adopts the following embodiments:
[0113] 1. Dataset preparation.
[0114] During the 10-day survey, 1,552 valid data points were collected from occupants and unoccupied shops for model training. These data included 705 instances of illegal occupancy and 847 instances of non-illegal occupancy. Using augmentation methods, the dataset was expanded to 7,760 entries, including 3,525 instances of illegal occupancy and 4,235 instances of non-illegal occupancy. The augmented power series data were converted into images and merged with the series dataset to create a multimodal dataset.
[0115] The experiment uses a multimodal mixed dataset consisting of a time series dataset and an image dataset. The dataset is divided into a training set (6208 items) and a validation set (1552 items) in an 8:2 ratio to ensure that the distribution ratio of the two labels (occupied / unoccupied) is consistent.
[0116] 2. Model training results.
[0117] The original document of the loss curve and accuracy curve of the training process is already available.
[0118] The best accuracy of the model on the validation set is 98.65%. The classification performance such as precision, recall and F1 score are shown in the table.
[0119] Table 1 Classification performance of InceptionTime-MobileNetV3 model
[0120]
[0121] To further validate the effectiveness of the model design, we selected five models: InceptionTime, Transformer, MobileNetV3, ResNet-50, and Transformer-ResNet-50, and conducted a comprehensive performance comparison with the InceptionTime-MobileNetV3 model proposed in this example. Comparison metrics included accuracy, precision, recall, F1 score, floating-point operations (FLOPs), and parameter count.
[0122] Floating-point operations (FLOPs) represent the number of floating-point operations performed by the model during inference, and the number of parameters (Params) represents the total number of trainable parameters in the model. Both are key metrics for measuring model complexity and storage requirements. Lower FLOPs indicate higher computational efficiency and faster inference. Higher Params indicate greater expressiveness, but also consumes more storage space and is more prone to overfitting.
[0123] Comparison of classification effects of different models
[0124] The original document of the loss curve and accuracy curve of the training process is already available.
[0125] The best accuracy of the model on the validation set is 98.65%. The classification performance such as precision, recall and F1 score are shown in the table.
[0126] Table 2 Classification performance of InceptionTime-MobileNetV3 model
[0127]
[0128] To further validate the effectiveness of the model design, we selected five models: InceptionTime, Transformer, MobileNetV3, ResNet-50, and Transformer-ResNet-50, and conducted a comprehensive performance comparison with the InceptionTime-MobileNetV3 model proposed in this example. Comparison metrics included accuracy, precision, recall, F1 score, floating-point operations (FLOPs), and parameter count.
[0129] Floating-point operations (FLOPs) represent the number of floating-point operations performed by the model during inference, and the number of parameters (Params) represents the total number of trainable parameters in the model. Both are key metrics for measuring model complexity and storage requirements. Lower FLOPs indicate higher computational efficiency and faster inference. Higher Params indicate greater expressiveness, but also consumes more storage space and is more prone to overfitting.
[0130] Table 3 Comparison of classification effects of different models
[0131]
[0132] Experimental results show that InceptionTime-MobileNetV3 outperforms single-modality models (such as InceptionTime and MobileNetV3) in classification performance, and has obvious advantages over deeper models (such as Transformer-ResNet-50) in computational complexity.
[0133] In terms of accuracy, precision, recall, and F1 score, InceptionTime-MobileNetV3 significantly outperforms single time series or image processing models. Its accuracy is 21.59% higher than that of the InceptionTime model, 3.1% higher than that of the MobileNetV3 model, 16.69% higher than that of the Transformer model, and 4.06% higher than that of the ResNet-50 model, indicating that multimodal feature fusion can effectively improve the ability to identify illegal residents.
[0134] Compared with ResNet-50 and Transformer-ResNet-50, the model proposed in this embodiment significantly reduces computational overhead and improves computational efficiency while ensuring high classification performance. Its FLOPs are reduced by 44 times compared to Transformer-ResNet-50, and the number of parameters is reduced by 91 times, making it more suitable for application scenarios with limited computing resources.
[0135] In summary, InceptionTime-MobileNetV3 balances computational efficiency and model complexity while ensuring high classification performance, and shows strong application potential in the task of identifying illegal residents.
[0136] Experimental results show that InceptionTime-MobileNetV3 outperforms single-modality models (such as InceptionTime and MobileNetV3) in classification performance, and has obvious advantages over deeper models (such as Transformer-ResNet-50) in computational complexity.
[0137] In terms of accuracy, precision, recall, and F1 score, InceptionTime-MobileNetV3 significantly outperforms single time series or image processing models. Its accuracy is 21.59% higher than that of the InceptionTime model, 3.1% higher than that of the MobileNetV3 model, 16.69% higher than that of the Transformer model, and 4.06% higher than that of the ResNet-50 model, indicating that multimodal feature fusion can effectively improve the ability to identify illegal residents.
[0138] Compared with ResNet-50 and Transformer-ResNet-50, the model proposed in this embodiment significantly reduces computational overhead and improves computational efficiency while ensuring high classification performance. Its FLOPs are reduced by 44 times compared to Transformer-ResNet-50, and the number of parameters is reduced by 91 times, making it more suitable for application scenarios with limited computing resources.
[0139] In summary, InceptionTime-MobileNetV3 achieves a balance between computational efficiency and model complexity while ensuring high classification performance, demonstrating strong application potential in the task of identifying illegal occupants. It should be noted that for the sake of simplicity, the aforementioned method embodiments are presented as a series of actions. However, those skilled in the art should be aware that the present invention is not limited to the order of the actions described, as certain steps may be performed in a different order or simultaneously.
[0140] Based on the same concept as the method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model in the above-mentioned embodiment, the present invention also provides a system for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model. The system can be used to execute the above-mentioned method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model. For ease of explanation, the structural diagram of the embodiment of the system for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model only shows the parts related to the embodiment of the present invention. Those skilled in the art will understand that the illustrated structure does not constitute a limitation of the device, and it can include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0141] See also Figure 7 In another embodiment of the present application, a system 10 for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model is provided. The system includes a preprocessing module 11, a data partitioning module 12, a detection and identification module 13, and an information communication module 14.
[0142] A preprocessing module 11 is used to collect target user data and preprocess the target user data, wherein the target user data includes an original time series power sequence and a judgment label;
[0143] The data division module 12 is used to divide the user's night time period and obtain the time series power sequence of the user in the night time period;
[0144] A detection and recognition module 13 is configured to identify illegal occupants based on a pre-trained time-series-image hybrid model during the nighttime time period. The pre-trained time-series-image hybrid model includes a time-series data processing network, an image data processing network, and a multimodal fusion layer.
[0145] The information communication module 14 is used to analyze the recognition result. If the recognition result is "resident", the detection result is recorded and an alarm message is issued.
[0146] It should be noted that the three-small-place illegal occupancy detection system based on a lightweight multimodal hybrid model of the present invention corresponds one-to-one to the three-small-place illegal occupancy detection method based on a lightweight multimodal hybrid model of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the three-small-place illegal occupancy detection method based on a lightweight multimodal hybrid model are all applicable to the embodiment of the three-small-place illegal occupancy detection method based on a lightweight multimodal hybrid model. For specific details, please refer to the description in the embodiment of the method of the present invention, which will not be repeated here. This is hereby declared.
[0147] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0148] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model, characterized by: The steps include: Collecting target user data and preprocessing the target user data, wherein the target user data includes an original time series power sequence and a judgment label; Divide the user's night time period and obtain the time series power sequence of the user during the night time period; Identify illegal occupants based on the user's nighttime time series power sequence based on a pre-trained time series-image hybrid model, comprising a time series data processing network, an image data processing network, and a multimodal fusion layer. If the recognition result is "resident", the detection result will be recorded and an alarm message will be issued.
2. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 1 is characterized in that: The collecting target user data and preprocessing the target user data includes: Detect missing values in the collected target user data and fill them with cubic spline interpolation, as shown in the following formula: At the known data points (x0,y0),(x1,y1),……,(x n ,y n ), construct a piecewise cubic polynomial: S i (x)=a i +b i (x-x i )+c i (x-x i ) 2 +d i (x-x i ) 3 ; Among them, a i , b i , c i and d i is the segmentation coefficient; Label each target user data sample and obtain the judgment label.
3. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 1 is characterized in that: The method of dividing the user into nighttime periods and obtaining the time series power sequence of the user in the nighttime period is specifically as follows: The user's non-business hours are set as the user's nighttime electricity usage hours, and the electricity usage data within the user's nighttime electricity usage hours are selected to obtain the user's nighttime electricity usage data; the user's non-business hours include fixed hours on non-consecutive days and fixed hours on consecutive days, and the fixed hours are adjusted according to the season.
4. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 1 is characterized in that: The pre-training of the time series-image hybrid model includes the following steps: Perform data enhancement processing on the original time series power sequence to obtain the set fine-grained time series power data; Performing format conversion on the set fine-grained time-series power data to obtain a power image set, wherein the format conversion includes: dynamically adjusting the size of the power image according to the sequence length, and filling the area between the power image and the coordinate axis with visual elements, wherein the dynamically adjusted size is proportional to the sequence length; A training set is constructed using the power image set and its corresponding judgment labels, and the time series-image hybrid model is pre-trained using the training set.
5. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 4 is characterized in that: The data enhancement processing is specifically as follows: Random scaling: Randomly scale the original time series power sequence, with the scaling ratio range being [0.9, 1.1] or [0.95, 1.05]; Time jitter: Random noise is added to the original timing power sequence to simulate measurement errors or environmental interference. Time warping: nonlinearly stretch the original time power sequence and set the maximum warping range; Mirror flip: mirror flip the original timing power sequence; Random truncation: The original time series power sequence is randomly truncated, and the truncation ratio does not exceed 5%.
6. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 1 is characterized in that: The time series data processing network includes multiple groups of one-dimensional convolutions, activation functions, maximum pooling layers, and fully connected layers; the multiple groups of one-dimensional convolutions run in parallel, and the convolution kernel sizes of each group are different. The activation function is the ReLU activation function. The maximum pooling layer is a global maximum pooling layer, which is used to reduce the dimension of the convolution layer output and obtain key time series features. The fully connected layer is used to map the time series features and align them with the image data processing network output; Each group of one-dimensional convolutions is connected to an activation function, and the features output by the activation function are concatenated along the channel dimension.
7. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 1 is characterized in that: The image data processing network includes a feature extraction module, a global average pooling layer and a fully connected layer; the feature extraction module includes a multi-layer feature extraction layer, an inverted residual structure and an SE attention module, which is used to extract high-level features of the image, the global average pooling layer is used to reduce the dimensionality of the high-level features of the image, and the fully connected layer is used to map the image features and align them with the output of the time series data processing network.
8. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 1 is characterized in that: The multimodal fusion layer includes a splicing layer, a fusion feature optimization layer, a classification output layer and a loss function; the splicing layer is used to splice the outputs of the time series data processing network and the image data processing network to obtain fusion features, the fusion feature optimization layer includes a dropout layer and a fully connected layer for learning high-order interactions of fusion features, and the classification output layer includes a softmax activation function for outputting category probabilities.
9. The method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model according to claim 1 is characterized in that: The method of identifying illegal occupants based on the user's time series power sequence during the night time period based on the pre-trained time series-image hybrid model includes: Convert the user's time-series power sequence during the night time period into a format to obtain a power image set; Using a time series data processing network to process the user's time series power sequence during the nighttime period to obtain a time series feature vector, and using an image data processing network to process the power image set to obtain an image feature vector; the time series feature vector is aligned with the image feature vector; The multimodal fusion layer is used to concatenate the time series feature vector and the image feature vector to obtain the fused feature vector. After learning the high-order interactions of the fused feature vector, the fused feature vector is mapped to a low-dimensional classification space and the output category probability is calculated. The cross entropy loss function is used as the optimization target to guide the parameter update during the training process. At the same time, the network parameters are continuously adjusted through the gradient descent optimization strategy until the parameters converge and the training is completed. The trained time series-image hybrid model is used to identify the target user data and determine whether the user is living there illegally.
10. A system for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model, characterized by: A method for detecting illegal occupancy in three small places based on a lightweight multimodal hybrid model, applied to any one of claims 1-9, comprising a preprocessing module, a data partitioning module, a detection and recognition module, and an information communication module; A preprocessing module, configured to collect target user data and preprocess the target user data, wherein the target user data includes an original time series power sequence and a judgment label; The data partitioning module is used to divide the user's night time period and obtain the time series power sequence of the user during the night time period; A detection and recognition module is used to identify illegal occupants based on the user's time series power sequence during the night time period based on a pre-trained time series-image hybrid model, which includes a time series data processing network, an image data processing network, and a multimodal fusion layer. The information communication module is used to analyze the recognition results. If the recognition result is "resident", the detection result is recorded and an alarm message is issued.