Green Base Station Energy Saving Method Based on the Fusion of Self-Attention and Convolutional Deep Learning

By using a character recognition model that integrates self-attention and convolutional deep learning in the base station, dynamically counting and monitoring the number of people in the area and adjusting the base station resources, the problem of inaccurate number of people in the existing technology is solved, and more efficient base station resource management and energy consumption reduction are achieved.

CN119155779BActive Publication Date: 2025-05-30HUAXIN CONSULTATING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411604317.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-05-30
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

When the prior art allocates base station resources through the number of people statistics, there is a problem of inaccurate identification, which affects the accuracy of resource allocation, and the application of the model has limitations, limiting its application in actual scenarios.

Method used

A character recognition model based on the fusion of self-attention and convolutional deep learning is adopted, and video images are scanned periodically, character movement trends are analyzed, the number of people in the area is dynamically counted, and the base station resources are adjusted according to the number of people.

Benefits of technology

The accuracy of counting the number of people has been improved, more accurate allocation of base station resources has been achieved, and the energy consumption and operation costs of 5G base stations have been reduced, achieving the goal of green energy saving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119155779B_ABST
    Figure CN119155779B_ABST
Patent Text Reader

Abstract

The present invention discloses a green base station energy-saving method based on the fusion of self-attention and convolutional deep learning, which solves the problems in the prior art that the base station resources are allocated by counting the number of people, the counting of the number of people is inaccurate, affecting the allocation of benchmark resources, and there are limitations in the application of models for human recognition, restricting the application in actual scenarios. The method includes collecting and monitoring the first video image at the periodic scanning time node; using a human recognition model based on the fusion of self-attention and convolutional deep learning to perform human recognition and counting on the first video image; analyzing the movement trend of people to dynamically count the number of people in the monitoring area; calculating the number of users in the monitoring area, and activating the corresponding base station capacity resources according to the number of users. The present invention utilizes long-term monitoring equipment to identify the human situation in the image, combines the relevant settings of the base station, and reasonably allocates the resources of the base station, so that the 5G base station further reduces energy consumption and overall operating costs, thus achieving green energy saving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a green base station energy-saving method based on the fusion of self-attention and convolutional deep learning. Background Art

[0002] With the development of mobile communication technologies, the number of base stations has been increasing continuously, and their energy consumption problems have become increasingly prominent. At present, the energy consumption of 5G base station equipment accounts for a quite large proportion of the total energy consumption of the entire wireless network. In order to achieve green energy-saving of base stations, it is necessary to find an effective method to check the number of users and adjust the working mode of the base station accordingly.

[0003] In the prior art, there is a method of identifying people through a person recognition model and counting the number of users, and allocating 5G base station resources according to the number of people to achieve the purpose of energy saving. Existing person recognition models have been able to achieve the ability to identify and locate a variety of different targets. However, the application of the calculation methods adopted by these models has certain limitations, which restricts their application in actual scenarios. For example, feature pyramid structures are used for feature fusion. However, various feature pyramid network structures cannot effectively utilize the correlation between all pyramid feature maps. Another example is the inherent limitations of convolution, including local receptive fields and the independence of input content, which limit the ability of the model to capture long distances. Summary of the Invention

[0004] The present invention mainly solves the problems in the prior art that when allocating base station resources through the statistics of the number of people, the statistics of the number of people is inaccurate, which affects the allocation of base resources, and the application of the model for person recognition has limitations, restricting its application in actual scenarios. The present invention provides a green base station energy-saving method based on the fusion of self-attention and convolutional deep learning.

[0005] The above technical problems of the present invention are mainly solved by the following technical solutions: A green base station energy-saving method based on the fusion of self-attention and convolutional deep learning, comprising the following steps:

[0006] S1. Collect and monitor the first video image at the periodic scanning time node;

[0007] S2. Use a person recognition model that fuses self-attention and convolutional deep learning to perform person recognition and counting on the first video image;

[0008] S3. Analyze the movement trend of people and dynamically count the number of people in the monitoring area;

[0009] S4. Calculate the number of users in the monitoring area and activate the corresponding base station capacity resources according to the number of users.

[0010] The present invention uses long-term monitoring devices to identify the situation of people in images, combines relevant settings of base stations, and reasonably allocates the resources of base stations, so that 5G base stations further reduce energy consumption and lower the overall operating cost, thus achieving green energy conservation. By analyzing the movement trend of people, it predicts the appearance of people in the monitoring images, dynamically counts the number of users in the monitoring area according to the actual movement of people, improves the accuracy of counting the number of people, and can more accurately allocate the resources of the base station according to the number of people, improve the energy utilization efficiency of the base station, reduce the energy consumption of the base station, and lower the overall operating cost.

[0011] As a preferred solution, the person recognition model that fuses self-attention and convolutional deep learning includes:

[0012] The input image is converted into feature maps in multiple different-scale spaces through a backbone network, and the first features of each scale space are respectively extracted and output.

[0013] Through the feature fusion neck, shallow feature interaction, deep feature interaction are carried out on the first features of each scale space, and the first features of each scale space are spliced. According to the deep interaction and the spliced features, context fusion is carried out to respectively output the second features of each scale space.

[0014] Each second feature is input into the prediction head to generate multi-target outputs.

[0015] The backbone network converts the input image into first features with lower spatial resolution but rich semantic information in multiple different scales. Each first feature is respectively input into the feature fusion neck, and each first feature is input into a shallow feature interaction module for shallow feature interaction. The shallow feature interaction is responsible for cross-scale information exchange and fusion. The main goal is to integrate the upper-level and the same-level features from the backbone network and the high-resolution shallow features in the backbone network, aiming to retain rich localization details to enhance the spatial performance of the network. After passing through the shallow feature interaction module, it enters the cross-stage block module. The cross-stage block module is a multi-branch structure that allows interaction between features at different levels. After passing through the cross-stage block module, it enters the deep feature interaction module for deep feature interaction, using additional deep information to enhance the feature representation. It mainly aggregates the information of the shallow high-resolution layer, the shallow low-resolution layer, the same-level shallow layer, and the previous layer, and enriches the gradient information of the output layer through multi-directional connections. In addition, the first features of each scale space are spliced through a three-scale sequence fusion module. Specifically, the feature maps of different scales are horizontally stacked, and three-dimensional convolution is used to extract their scale sequence features.

[0016] As a preferred solution, the backbone network includes three sequentially connected efficient layer aggregation modules. The efficient layer aggregation modules extract features and output first features of different scales. Among them, the bottommost efficient layer aggregation module obtains a feature map through a local-global hybrid self-attention mechanism and a relative position enhanced self-attention mechanism, and stacks it with the feature maps of different scales in the upper layer of the feature map to output the first feature.

[0017] The backbone network converts the input image into multiple feature maps with lower spatial resolution but rich semantic information at different scales. The output feature maps are fed into three efficient layer aggregation modules K1, K2, and K3, which use attention mechanisms or other methods to selectively focus on important features.

[0018] As a preferred solution, the efficient layer aggregation module performs pointwise convolution and feature separation operations on the input information. After the operations, it is divided into a first branch and a second branch. The first branch enters the feature stack after passing through N inverted residual bottlenecks, and the second branch retains the original information and enters the feature stack. Finally, pointwise convolution is performed to output the information.

[0019] The local-global hybrid self-attention mechanism includes three input branches. The first input branch outputs global information after passing through a linear projection layer, layer normalization, one-dimensional convolution, layer normalization, and squeeze-and-excitation attention. The second input branch outputs local information after passing through a linear projection layer, layer normalization, multi-head self-attention, layer normalization, and squeeze-and-excitation attention. The global information, local information, and the original input are jointly input into the adaptive hybrid global and local information module to form a hybrid output. The hybrid output is transformed through a linear projection layer and added to the original input to form a residual connection, and the information is output after layer normalization processing.

[0020] The inputs of the relative position enhanced self-attention mechanism respectively enter the query matrix linear projection layer, key matrix linear projection layer, value matrix linear projection layer, parameter matrix linear projection layer, matrix linear projection layer, and key matrix linear projection layer. After the outputs are multiplied by matrices and added to the relative position of the bias amount, they are processed by the softmax function and then multiplied by the output of the value matrix linear projection layer to form the first output. The output of the parameter matrix linear projection layer acts on the depthwise separable convolution of the tokens and forms the second output through the tokens-level linear projection. The first output and the second output are added to obtain the final output.

[0021] In this solution, the efficient layer aggregation module can efficiently learn expressive multi-scale feature representations. The input information undergoes pointwise convolution and feature separation operations and is divided into two branches. One branch retains the original information and then directly enters the feature stacking operation, while the other branch is processed by N inverted residual bottleneck units. Due to the mechanism of efficient layer aggregation, the branches and the outputs passing through each inverted residual bottleneck unit are retained and finally concatenated together. For the specific structure of the inverted residual bottleneck unit, the input order expands the number of channels through pointwise convolution, followed by a k*k heterogeneous depthwise separable convolution operation, and finally reduces the number of channels through pointwise convolution to compensate for the information loss caused by the depthwise separable convolution. During training, the network runs n different-sized depth convolutions in parallel. By parallelizing large-kernel convolutions with several small-kernel convolutions, this module expands the receptive field without incurring additional inference costs and retains the information of small targets.

[0022] The local-global hybrid self-attention mechanism combines a local convolutional layer and a global self-attention layer to jointly model the long-term and short-term preferences of users. The local preferences (captured by CNN) and global preferences (captured by Transformer) are combined to more comprehensively model the dynamic preferences of users; adaptively mix global and local information, decouple the fusion process in different layers, enhance the expressive power, and adaptively aggregate long-term and short-term preferences (the mixing importance of the local and global dependence modules can be adjusted according to the personalized needs of users); squeeze-and-excitation attention is used to replace the softmax operation, allowing multiple related items to be considered simultaneously, enhancing the model's expressive power (allowing the model to simultaneously focus on multiple highly related objects instead of just a single object as in traditional softmax).

[0023] Squeeze-and-excitation attention: Receives an input tensor with the shape [B, N, D], where B is the batch size, N is the sequence length, and D is the feature dimension.

[0024] First, calculate the average value of the input tensor along the feature dimension D to obtain a tensor average with the shape [B, N, 1].

[0025] Transpose the dimensions of the input tensor average to obtain a tensor with the shape [B, 1, N].

[0026] Transform the transposed tensor through the first linear layer to obtain a hidden layer state tensor with the shape [B, 1, N / r], where N / r represents the compressed dimension.

[0027] Apply an activation function (ReLU) to further non-linearly activate the hidden layer state.

[0028] Transform the activated tensor back to the original sequence length [B, 1, N] through a second linear layer.

[0029] Apply the Sigmoid function to generate an attention weight tensor of attention scores, with a shape of [B, 1, N].

[0030] Transpose the attention scores again to multiply with the original input tensor, obtaining a tensor with a shape of [B, N, 1].

[0031] Adaptively blend global and local information, receiving three input tensors: the model input, global information, and local information.

[0032] Calculate the average value of the input tensor in the feature dimension for generating an adaptive weight.

[0033] Obtain the adaptive weight alpha through a linear projection layer and squeeze-and-excitation attention, and expand it to a shape of [B, 1, 1] (where B represents the batch size of input samples).

[0034] Set beta to 1 - alpha, representing the remaining proportion.

[0035] Based on these two weights, the global information and local information are blended together to form a mixed output.

[0036] After the mixed output is transformed through a linear projection layer, it is added to the original input to form a residual connection.

[0037] The final output is processed through a dropout layer and layer normalization.

[0038] Dynamically adjust the proportion between global information and local information. Utilize the residual connection to maintain the gradient flow and reduce the risk of gradient vanishing. Enhance the generalization ability of the model through layer normalization and dropout.

[0039] Relative Position Enhanced Self-Attention Mechanism, Importance of Position Information. For computer vision tasks, pixels or patches in an image are regarded as tokens, and their position information is crucial for constructing image semantics. A key feature of the self-attention mechanism is its insensitivity to the permutation order of tokens, which means it cannot directly utilize the position information of tokens. Introducing Relative Position Bias: Methods such as Swin Transformer introduce additional relative position biases in the attention matrix to guide the model to learn position information and content-based attention. The relative position enhanced self-attention mechanism further develops this idea by establishing an independent value representation space to capture position dependencies. Mathematical formula of the relative position enhanced self-attention mechanism: The relative position enhanced self-attention mechanism converts tokens into new representations by introducing a parameter matrix Vrel to mix information through relative positions. The relative position enhanced self-attention mechanism uses a weight tensor brel, which is indexed by relative positions. The relative position enhanced self-attention mechanism also retains the relative position bias B to give the model greater flexibility. Limiting the Maximum Absolute Distance: To enhance the generalization ability of the model and simplify the implementation, the relative position enhanced self-attention mechanism limits the maximum absolute relative position distance s on each axis. After doing so, the calculation of the relative position enhanced self-attention mechanism can be simplified to a token-based convolution with a kernel size of (2s + 1) * (2s + 1).

[0040] As a preferred solution, the feature fusion neck includes three layers of branches. Each layer of branch sequentially includes a shallow feature interaction module, a cross-stage block module, a deep feature interaction module, and a spatial context awareness module. The shallow feature interaction module receives the first features of the same layer and the previous layer, and outputs shallow features to the cross-stage block module. The cross-stage block module distributes the shallow features to each deep feature interaction module in the front and the shallow feature interaction module of the previous layer in the back. The deep feature interaction module outputs deep features to the spatial context awareness module. The three-scale sequence fusion module obtains the first features of the same layer of each layer of branch, and after feature splicing, inputs them into the first layer of spatial context awareness module. The spatial context awareness module outputs global context features to the deep feature interaction module of the next layer in the back and outputs them to the prediction head.

[0041] After each efficient layer aggregation module, there is a shallow feature interaction module responsible for cross-scale information exchange and fusion. The shallow feature interaction module is followed by a cross-stage block module, which is a multi-branch structure that allows interaction between features at different levels. The deep feature interaction module utilizes additional deep information to enhance the feature representation. The spatial context awareness module may capture spatial relationships and generate richer feature representations.

[0042] As a preferred solution, the shallow feature interaction module receives the first features of the upper layer, the same layer and the lower layer, and the shallow features of the next layer. The first features of the upper layer are processed by depthwise separable convolution to form deep information, which is stacked with the first features of the same layer and the shallow features of the next layer to output shallow features;

[0043] The deep feature interaction module receives the shallow features of the upper layer, the same layer, the lower layer, and the global context features of the upper layer in the front part. The shallow features of the upper layer, the shallow features of the lower layer, and the global context features of the upper layer in the front part are respectively processed by depthwise separable convolution and then stacked with the shallow features of the same layer to output deep features;

[0044] The triple-scale sequence fusion module includes three branches, which respectively receive the first features of three different scales. Each branch adjusts the channels through a convolution module. The first branch performs downsampling through max pooling and average pooling, and then is processed by a convolution module. The third branch performs upsampling using the nearest neighbor interpolation method and then is processed by a convolution module. Then, the features of the three branches are stacked, and then processed by a 3D convolution module to output the spliced features;

[0045] The context-aware module includes three branches, which respectively perform global average pooling, global max pooling, and pointwise convolution on the input features. After global average pooling and global max pooling, they jointly pass through the softmax function, matrix multiplication, and pointwise convolution processing. The pointwise convolution passes through the softmax function, matrix multiplication, and pointwise convolution processing. Finally, they jointly perform the Hadamard product to output the global context features.

[0046] The accurate positioning of the shallow feature interaction module depends on the detailed edge information obtained from the shallow network, while accurate classification requires a deeper network to capture coarse-grained information. An effective feature pyramid network should support the full and sufficient fusion of the information flow between the shallow and deep networks; retaining the shallow spatial information in the backbone network is crucial for enhancing the detection ability of smaller targets. However, the information provided by the backbone network is relatively basic and vulnerable to interference. Therefore, the shallow information is incorporated into the deep network as an auxiliary branch to ensure the stability of the subsequent layer learning; the main goal is to integrate the deep information with the high-resolution shallow features from the same layer and the next layer of the backbone network, aiming to retain rich positioning details to enhance the spatial performance of the network; 1*1 convolution is used to control the number of channels in the shallow information, ensuring that it accounts for a small proportion in the splicing operation and does not affect the subsequent learning. Maintaining the shallow backbone information through bidirectional connection enhances the network's ability to detect small targets.

[0047] Deep feature interaction module. To further improve the interactive utilization of feature layer information, multi-scale information fusion is performed at a deeper level; these connections span the high-resolution layers of the shallow layer, the low-resolution layers of the shallow layer, the shallow layers of the same level, and the front spatial context awareness for information aggregation. The gradient information of the output layer is enriched through multi-directional connections.

[0048] Triple-scale sequence fusion module. To identify densely overlapping small objects, the shape or appearance changes at different scales can be referred to and compared by magnifying the image. Since different feature layers of the backbone network have different sizes, the traditional feature pyramid fusion mechanism only upsamples the small-size feature maps and then segments or adds them to the features of the upper layer, ignoring the rich detail information of the large-size feature layers. The triple-scale sequence fusion module can segment the large, medium, and small-size features, add the large-size feature maps, and perform feature magnification to improve the detailed feature information. It can better combine the high-dimensional information of the deep feature maps and the detailed information of the shallow feature maps. The scale space is constructed along the scale axis of the image, which not only represents a scale but also represents the range of various scales that an object may have. The feature maps of different scales are horizontally stacked, and 3D convolution is used to extract their scale sequence features. The high-resolution feature maps contain most of the information crucial for the detection of small targets. Before feature encoding, the number of feature channels is first adjusted to be consistent with the main scale features. After the convolution module processes the large-size feature maps, its number of channels is adjusted to 1C, and then a hybrid structure of max pooling and average pooling is used for downsampling, which helps to retain the effectiveness and diversity of the high-resolution features and cell images. For the small-size feature maps, the convolution module is also used to adjust the number of channels, and then nearest neighbor interpolation is used for upsampling. This helps to maintain the richness of the local features of the low-resolution images and prevent the loss of small target feature information. Finally, the large, medium, and small feature maps of the same size are convolved once, then concatenated in the channel dimension, and finally 3D convolution, batch normalization, and activation function are performed.

[0049] Spatial context awareness module. It is designed to model the global context information in the cross-channel and spatial dimensions and suppress useless background features. The structure consists of three branches, and global average pooling (GAP) and global max pooling (GMP) are used to integrate the global information. Its functions are: extracting the context relationship between space and channels; enhancing the feature representation of small targets and reducing background confusion; implementing the linear transformation of simplifying the feature maps using 1x1 convolution; calculating the channel and spatial context information through matrix multiplication; and performing Hadamard product operation to output the feature maps containing global context.

[0050] As a preferred solution, step S3 specifically includes:

[0051] S31. Analyze the motion state of the person through deep learning, determine the moving direction of the person in the motion state, calculate the moving trajectory of the person, and judge whether the person is about to exceed the current video image range;

[0052] S32. If not, count the number of people in each monitoring area and enter step S4;

[0053] If so, determine the monitoring range where the person is about to appear according to the task moving trajectory, obtain the second video image of this monitoring after a set delay time, perform person recognition and counting on the second video image, count the number of people in each monitoring area, enter step S4, and at the same time judge whether the current time exceeds the current scanning period;

[0054] S33. If so, enter the next scanning period and return to step S1;

[0055] If not, return to step S31.

[0056] This method arranges cameras within a set monitoring area, using one or more cameras. The cameras should cover the monitoring area as much as possible. For example, one camera can be set with its coverage area defined as the monitoring area, or multiple cameras can be set, and the coverage areas of these multiple cameras form a monitoring area. The cameras are set with a scanning period to regularly scan or detect images for user population statistics. When it is analyzed that a person is in a moving state, that is, the person may walk out of the current monitoring area and enter other monitoring areas. Different monitoring areas may be within the coverage ranges of different base stations. Therefore, it is necessary to re - count the number of people in the monitoring area according to the movement dynamics of the person, so as to more accurately allocate the base station resources according to the number of people. In this solution, the movement trend of people is analyzed through deep learning to dynamically count the number of people in the monitoring area. By analyzing the movement trend of people in the current video image through deep learning, it is judged whether a person is about to exceed the current video image range according to the movement trend of the person. Additionally, it further includes analyzing the people in the video image who are not full - body, and the people who are not full - body are also considered to be about to exceed the video image range. The movement trend of people is determined by analyzing the movement state of people, determining the moving direction of the people in the moving state, calculating the movement trajectory of the people, so as to judge whether the person will exceed the video image range. At the same time, according to the moving direction and the movement trajectory, it is also possible to predict that the person will appear in the coverage range of the surrounding cameras, so as to be able to retrieve the video images of the surrounding cameras to determine whether the person has moved past. When it is judged that there are people in the video image who are about to exceed the video image range, the second video image of the monitoring range where the person will appear is obtained after a set delay time. The delay time is set according to requirements, and the delay time is less than the scanning period duration. For example, the delay time can be set according to the frequency of the camera capturing video frames. If the camera captures video frames every minute, the delay time is set to one minute. After obtaining the second video image, the people are identified and counted through a person recognition model, and the position of each person is marked. At the same time, the second video image is also analyzed through deep learning to identify the people who have moved into the current image range. At this time, the number of people calculated by the monitoring, that is, the cameras, is re - counted to statistically obtain the number of users in each monitoring area, and the base station resources are dynamically adjusted according to the number of users. At the same time, the movement trend of the people in the second video image is also analyzed to judge whether there are people who are about to exceed the second video image range. For example, if the people who have moved into the current image range continue to move, or other people have a movement trend, the above operations are repeated. The number of people in the monitoring area is dynamically counted according to the movement trend of the people, and the base station resources are adjusted in real - time until the people in the video images of the surrounding monitoring obtained after the delay time no longer move or the current time exceeds a scanning period, then the operation stops and enters the next scanning period for the operation of the next scanning period. When it is detected in the first video image that multiple people are about to exceed the image range, and in the second video image it is detected that other people are also about to exceed the image range, these people are tracked respectively, and step S3 is repeated.This solution uses a small monitoring area, which can be composed of one or a few monitoring cameras. The movement of people will not be too complex, and the base station resources can be adjusted in a timely manner according to the movement of people to achieve the purpose of reasonable resource allocation.

[0057] As a preferred solution, after step S2, it further includes:

[0058] Judge whether there are people in the person recognition result.

[0059] If so, enter step S3; if not, notify the 5G base station to enter the waiting, sleeping or deep sleeping state.

[0060] This solution first includes judging whether there are people after person recognition. When there are no tasks, the 5G base station is notified to enter the waiting, sleeping or deep sleeping state. When people are detected, the covered base station is activated, effectively improving the energy utilization efficiency of the base station, reducing the energy consumption of the base station, and reducing the overall operating cost.

[0061] As a preferred solution, calculating the number of users in the monitoring area specifically includes:

[0062] Obtain the operator user penetration rate in this service area and calculate the number of operator users in each monitoring area;

[0063] Obtain the proportion of 5G users of the operator and calculate the operator's 5G users in each monitoring area.

[0064] In order to count the number of 5G users, it is necessary to input the operator data in the service area as the basis for calculation and analysis. Calculate the number of operator users by inputting the operator user penetration rate in the service area, and calculate the number of 5G users by inputting the proportion of 5G users of the operator.

[0065] As a preferred solution, it further includes a real-time adjustment step, which specifically includes:

[0066] Calculate the expected capacity demand of users according to the number of users in the monitoring area, and at the same time obtain the actual capacity demand of users in the monitor area;

[0067] Compare the expected capacity demand of users with the actual capacity demand of users, and adjust the base station capacity resources in real time according to the actual capacity demand of users.

[0068] This solution is carried out after step S5. Due to the problem that users may have a large or small usage capacity in actual situations, it is necessary to comprehensively count and analyze the capacity requirements of users in order to better plan and adjust resources. Estimate the expected capacity requirements of users based on the number of users counted in the monitored area, and at the same time obtain the actual capacity requirements of users within the monitored range through the base station. Calculate the actual capacity requirements of users according to the positions of the people marked by the person recognition model, and then obtain the users in the monitored area through the positions. Count the actual capacity requirements of these users from the base station. Compare the actual capacity requirements of users with the expected capacity requirements of users. If the actual capacity requirements of users are greater than the expected capacity requirements of users, then increase the base station capacity resources according to the actual capacity requirements of users. If the actual user capacity requirements exceed the current base station resource capabilities, consider adding base station resources to meet the requirements. If the actual capacity requirements of users are less than the expected capacity requirements of users, no operation is performed. In addition, in order to further improve the energy utilization efficiency and reduce the energy consumption of the base station, if the actual capacity requirements of users are less than the expected capacity requirements of users, the base station capacity requirements can also be reduced according to the actual capacity requirements of users. After the real-time adjustment step ends, a round of adjustment of the base station resources according to the number of people ends.

[0069] Therefore, the advantages of the present invention are:

[0070] By using long-term monitoring devices to identify the people in the images and combining with the relevant settings of the base station, the resources of the base station are reasonably allocated, enabling the 5G base station to further reduce energy consumption and lower the overall operating cost, thus achieving green energy conservation. Through the analysis of the movement trends of people, predict the appearance of people in the monitoring images, dynamically count the number of users in the monitored area according to the actual movement of people, improve the accuracy of counting the number of people, and thus can more accurately allocate the base station resources according to the number of people, improve the energy utilization efficiency of the base station, reduce the energy consumption of the base station, and lower the overall operating cost.

[0071] The person recognition model can segment features of large, medium, and small sizes by adopting a ternary scale sequence fusion model, add large-size feature maps, and perform feature amplification to improve detailed feature information, and can better combine the high-dimensional information of the deep feature map and the detailed information of the shallow feature map; adopt a relative position enhanced self-attention mechanism to make up for the position blind spots of self-attention, and the model has a strong ability to capture dependence relationships from content and position; adopt a spatial context awareness module to extract the context relationship between space and channels, enhance the feature representation of small targets, and reduce background confusion; adopt a local-global hybrid self-attention mechanism to combine local preferences and global preferences to more comprehensively model the dynamic preferences of users. Description of the Drawings

[0072] Figure 1 It is a schematic flowchart of the method of the present invention.

[0073] Figure 2 It is an architecture diagram of the person recognition model in the present invention.

[0074] Figure 3 It is a structural diagram of the efficient layer aggregation module of the person recognition model in the present invention.

[0075] Figure 4 It is a structural diagram of the local-global hybrid self-attention mechanism of the person recognition model in the present invention.

[0076] Figure 5 It is a structural diagram of the relative position enhanced self-attention mechanism of the person recognition model in the present invention.

[0077] Figure 6 It is a structural diagram of the shallow feature interaction module of the person recognition model in the present invention.

[0078] Figure 7 It is a result diagram of the deep feature interaction module of the person recognition model in the present invention.

[0079] Figure 8 It is a structural diagram of the ternary scale sequence fusion module of the person recognition model in the present invention.

[0080] Figure 9 It is a structural diagram of the spatial context awareness module of the person recognition model in the present invention. Detailed implementation manners

[0081] The technical solution of the present invention will be further specifically described below through embodiments in conjunction with the accompanying drawings.

[0082] Embodiment:

[0083] An energy-saving method for a green base station based on the fusion of self-attention and convolutional deep learning in this embodiment is as Figure 1 shown, and includes the following steps:

[0084] S1. Collect and monitor the first video image at the periodic scanning time node;

[0085] S2. Use the person recognition model that fuses self-attention and convolutional deep learning to perform person recognition and counting on the first video image;

[0086] S3. Analyze the person movement trend and dynamically count the number of people in the monitoring area;

[0087] S4. Calculate the number of users in the monitoring area and activate the corresponding base station capacity resources according to the number of users.

[0088] As a preferred solution of this embodiment, step S1 specifically includes setting a scanning cycle arrangement, obtaining the first video image of the monitoring camera at the time node at the start of each scanning cycle, and performing subsequent monitoring and analysis. As the basis for the implementation of the method, a monitoring area is delimited, and monitoring cameras are installed at key positions in each monitoring area. Hereinafter referred to as monitoring, one or more monitors are installed in one monitoring area, and one or more monitors cover the monitoring area as much as possible. It is set that the monitor captures video frames at a fixed frequency, for example, once a minute.

[0089] As a preferred solution of this embodiment, in step S2, a person recognition model that fuses self-attention and convolutional deep learning is used to perform person recognition on the first video image, mark the pedestrians of each person, and calculate the number of people in the first video image. The person recognition model is a person recognition model for number detection based on the fusion of self-attention and convolutional deep learning.

[0090] As a preferred solution of this embodiment, first, it is judged whether there are people in the first video image according to the person recognition result;

[0091] If so, go to step S3 to perform further analysis of the person behavior pattern;

[0092] If not, notify the 5G base station to enter the waiting, sleeping or deep sleeping state.

[0093] In this solution, the operation of the base station is controlled by whether there are people. When there are no people, the waiting or sleeping state is selected, and the base station is activated again when people are detected, which effectively improves the energy utilization efficiency of the base station, reduces the energy consumption of the base station, and reduces the overall operation cost.

[0094] As a preferred solution of this embodiment, step S3 specifically includes:

[0095] S31. Analyze the person's movement state through deep learning, determine the moving direction of the person in the moving state, calculate the person's movement trajectory, and judge whether the person is about to exceed the current video image range;

[0096] S32. If not, count the number of people in each monitoring area and go to step S4;

[0097] If so, determine the monitoring range where the person is about to appear according to the task movement trajectory, obtain the second video image of this monitoring after a set delay time, perform person recognition and counting on the second video image, count the number of people in each monitoring area, go to step S4, and at the same time judge whether the current time exceeds the current scanning cycle;

[0098] S33. If so, enter the next scanning cycle and return to step S1;

[0099] If not, return to step S31.

[0100] Further analyze the human behavior patterns in the current video image. When it is analyzed that the person is in a moving state, that is, the person may walk out of the current monitoring area and enter other monitoring areas. Different monitoring areas may be within the coverage ranges of different base stations, which will cause changes in the user capacity requirements within the reference coverage range. Therefore, it is necessary to re-count the number of people in the monitoring area according to the dynamic movement of the person, so as to more accurately allocate the base station resources according to the number of people. In this solution, the deep learning is used to analyze the movement trend of the person and dynamically count the number of people in the monitoring area. The deep learning analysis adopts the following several technologies:

[0101] Object Tracking Algorithms:

[0102] DeepSORT: DeepSORT is an algorithm that combines deep learning feature extraction and the traditional SORT (Simple Online and Realtime Tracking) tracker. It can be used to track the positions and movement trajectories of multiple targets in a video sequence.

[0103] FairMOT: FairMOT is an algorithm for multi-object tracking. It can perform detection and tracking simultaneously in a single inference and is suitable for real-time application scenarios.

[0104] Action Recognition Algorithms:

[0105] Spatial-temporal Graph Convolutional Networks (ST-GCN): This type of algorithm analyzes human actions by constructing a spatio-temporal graph and is suitable for understanding complex behavior patterns.

[0106] Two-stream Convolutional Networks: This algorithm combines the RGB image stream and the optical flow to capture the spatial and temporal information in the video and is very effective for understanding the actions and directions of people.

[0107] Person Re-Identification, ReID:

[0108] Deep Person Re-ID: This type of algorithm can identify the same individual by learning the identity-invariant features across camera views, which is helpful for tracking the movement paths of people between different cameras.

[0109] Trajectory Prediction Algorithms:

[0110] Social LSTM: This algorithm takes into account the social interaction relationships among pedestrians and can predict the future trajectories of pedestrians.

[0111] ConvLSTM: Convolutional LSTM combines the advantages of convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) and can be used to predict the movement trajectories of people.

[0112] By analyzing the movement trends of people in the current video image through deep learning, it is determined whether a person is about to exceed the current video image range based on the movement trends of the people. Further determining whether a person is about to exceed the current video image range also includes analyzing people in the video image who are not full-body. People who are not full-body are also considered to be about to exceed the video image range. The movement trends of people are analyzed by analyzing the movement states of the people, determining the moving directions of the people in the movement states, calculating the movement trajectories of the people, and thus judging whether they will exceed the video image range. At the same time, based on the moving directions and movement trajectories, it is also possible to predict where the person will appear in the coverage range of the surrounding cameras, so as to be able to retrieve the video images of the surrounding cameras to determine whether the person has moved past. When it is determined that there is a person in the video image who is about to exceed the video image range, the second video image of the monitoring range where the task will appear is obtained after a set delay time. The delay time is set according to requirements and is less than the duration of the scanning cycle. For example, the delay time can be set according to the frequency at which the camera captures video frames. If the camera captures a video frame every minute, the delay time is set to one minute. After obtaining the second video image, the people are identified and counted through a person recognition model, and the positions of each person are marked. At the same time, the second video image is also analyzed through deep learning to identify the people who have moved into the current image range. At this time, the number of people in each monitoring area is re-counted based on the number of people calculated by the monitoring camera, that is, the camera, and the base station resources are dynamically adjusted according to the number of users. At the same time, the movement trends of the people in the second video image are also analyzed to determine whether there are people who are about to exceed the second video image range. For example, if the people who have moved into the current image range continue to move, or other people have movement trends, the above operations are repeated. The number of people in the monitoring area is dynamically counted based on the movement trends of the people, and the base station resources are adjusted in real time until the people in the video images of the surrounding monitoring do not move after the delay time or the current time exceeds a scanning cycle, then the operation stops and enters the next scanning cycle for the operations of the next scanning cycle. In this solution, small monitoring areas can be set up, which can be composed of one or a few monitoring cameras. The movement situations of people will not be too complex, and the base station resources can be adjusted in a timely manner according to the movement of people to achieve the purpose of reasonable resource allocation.

[0113] As a preferred solution of this embodiment, step S4 specifically includes:

[0114] Calculate the number of users based on the number of people recognized in each monitoring area. Specifically, obtain the user penetration rate of the operator in this business area and multiply it by the number of people in the monitoring area to calculate the number of operator users in each monitoring area; and obtain the proportion of 5G users of the operator and multiply it by the number of operator users to calculate the number of operator 5G users in each monitoring area.

[0115] Make prior preparations for the corresponding benchmarks. By inputting the number of base stations covering this venue, understand the base station coverage in each monitoring area to prepare for subsequent resource allocation. Further refine the base station coverage range for more accurate analysis. According to the obtained number of operator 5G users, activate the corresponding base station capacity resources to ensure smooth network operation.

[0116] Furthermore, it is also necessary to adjust the base station resources in real time according to the actual user situation. Since in actual situations, users may have problems of large or small usage capacity, it is necessary to comprehensively count and analyze the capacity requirements of users for better resource planning and adjustment. The real-time adjustment steps include:

[0117] Calculate the estimated capacity requirements of users based on the number of users in the monitoring area, and at the same time obtain the actual capacity requirements of users in the monitor area;

[0118] Compare the estimated capacity requirements of users with the actual capacity requirements of users, and adjust the base station capacity resources in real time according to the actual capacity requirements of users.

[0119] Specifically, estimate the estimated capacity requirements of users based on the number of users counted in the monitoring area, and at the same time obtain the actual capacity requirements of users within the monitoring range through the base station. Calculate the actual capacity requirements of users according to the positions of the people marked by the person recognition model, and then obtain the users within the monitoring area through the positions, and count the actual capacity requirements of these users from the base station. Compare the actual capacity requirements of users with the estimated capacity requirements of users. If the actual capacity requirements of users are greater than the estimated capacity requirements of users, then increase the base station capacity resources according to the actual capacity requirements of users. If the actual user capacity requirements exceed the current base station resource capabilities, consider increasing the base station resources to meet the requirements. If the actual capacity requirements of users are less than the estimated capacity requirements of users, no operation is performed. In addition, to further improve energy utilization efficiency and reduce the energy consumption of the base station, if the actual capacity requirements of users are less than the estimated capacity requirements of users, the base station capacity requirements can also be reduced according to the actual capacity requirements of users. After the real-time adjustment steps end, one round of adjustment of the base station resources according to the number of people ends.

[0120] The present invention utilizes long-term monitoring devices to identify the situation of people in images, combines relevant settings of base stations, and reasonably allocates the resources of base stations, enabling 5G base stations to further reduce energy consumption and lower the overall operating cost, thereby achieving green energy conservation. By analyzing the movement trends of people, it predicts the appearance of people in surveillance images, dynamically counts the number of users in the monitored area according to the actual movement of people, improves the accuracy of counting the number of people, and thus can more accurately allocate the resources of base stations according to the number of people, improve the energy utilization efficiency of base stations, reduce the energy consumption of base stations, and lower the overall operating cost.

[0121] The following uses a scenario example to illustrate the method steps of this embodiment. For example, a 5G base station energy-saving system near a shopping mall (concert venue, etc.). Assume that a green base station energy-saving system based on the fusion of self-attention and convolutional deep learning is deployed near a large shopping mall. The flow of people in this shopping mall is small on weekdays and surges on weekends and holidays.

[0122] Initial setting and monitoring

[0123] Install high-definition surveillance cameras at key positions around the shopping mall to cover as many pedestrian passages as possible. These surveillance cameras capture video frames at a fixed frequency, such as once a minute. Set a scanning period, such as automatically scanning once every 10 minutes as preset, to obtain the latest pedestrian flow data.

[0124] Use a pre-trained person recognition model to perform person recognition on the images captured by the surveillance. The model distinguishes pedestrians from background objects, marks the position of each pedestrian, and calculates the number of pedestrians in the image. Determine whether there are people. If there are people in the image, proceed to the next step of analysis. If there are no pedestrians, notify the 5G base station to enter an energy-saving mode, such as a sleep state.

[0125] Deep learning analysis is used to further analyze the behavior patterns of pedestrians, such as walking direction, speed, and degree of aggregation, etc., which helps to predict the pedestrian flow trend in the next period of time.

[0126] Based on the analysis results, correct and refine the number statistics, and count the number of people in each monitored area.

[0127] Output the number statistics, and send the final number statistics data to the control center of the system as the basis for adjusting the working state of the base station.

[0128] Service area and operator data input

[0129] Input the operator user penetration rate: collect the wireless signal strength data and other network performance indicators in the shopping mall area to calculate the number of operator users.

[0130] Input 5G user occupancy ratio: Understanding the 5G user ratios of different operators in the area and calculating the number of 5G users of the operators helps to allocate network resources more reasonably.

[0131] Input the number of covered base stations: Confirm all the base stations covering the area and understand their respective coverage ranges.

[0132] Resource allocation and notification

[0133] Based on the calculated total number of 5G users in the current area, activate the corresponding base station capacity resources, including: Dynamically adjusting the workload of the base station according to the change in the number of users, such as increasing the bandwidth or adjusting the signal transmission power.

[0134] Statistical user capacity requirements: Regularly update the user capacity requirements and optimize the base station configuration accordingly.

[0135] Increase base station resources: If it is found that the user demand exceeds the existing resources, the system will notify the administrator to add new base stations or enhance the functions of existing base stations.

[0136] Notify the 5G base station management platform: Ensure that all adjustments are synchronized to the 5G base station management platform so that it can respond to changes in user demand in a timely manner.

[0137] Through such a system, not only can unnecessary energy consumption be effectively reduced, but also the user experience can be improved, ensuring sufficient network capacity support during peak hours.

[0138] In step S2, a person recognition model for detecting the number of people based on the fusion of self-attention and convolutional deep learning is adopted, such as Figure 2 shown, the person recognition model includes:

[0139] Convert the input image into feature maps in multiple different scale spaces through the backbone network, and respectively extract and output the first features of each scale space;

[0140] Perform shallow feature interaction, deep feature interaction on the first features of each scale space through the feature fusion neck, and splice the first features of each scale space. According to the deep interaction and spliced features, perform context fusion and respectively output the second features of each scale space;

[0141] Input each second feature into the prediction head to generate multi-object output.

[0142] The backbone network converts the input image into multiple first features with lower spatial resolution but rich semantic information at different scales. Each first feature is respectively input into the feature fusion neck, and each first feature is input into a shallow feature interaction module for shallow feature interaction. The shallow feature interaction is responsible for cross-scale information exchange and fusion. Its main goal is to integrate the upper-level features with the features from the same level and the high-resolution shallow features in the backbone network, aiming to retain rich localization details to enhance the spatial performance of the network. After passing through the shallow feature interaction module, it enters the cross-stage block module. The cross-stage block module is a multi-branch structure that allows interaction between features at different levels. After passing through the cross-stage block module, it enters the deep feature interaction module for deep feature interaction, using additional deep information to enhance the feature representation. Specifically, it aggregates the information from the high-resolution layer of the shallow layer, the low-resolution layer of the shallow layer, the shallow layer at the same level, and the previous layer, enriching the gradient information of the output layer through multi-directional connections. In addition, the first features in the spatial dimension at each scale are stitched together through a three-scale sequence fusion module. Specifically, the feature maps at different scales are horizontally stacked, and 3D convolution is used to extract their scale sequence features.

[0143] Specifically, the backbone network includes three efficiently-layer-aggregated modules connected in sequence. The efficiently-layer-aggregated modules extract features and output first features at different scales. Among them, the bottommost efficiently-layer-aggregated module obtains the feature map through a local-global hybrid self-attention mechanism and a relative-position-enhanced self-attention mechanism, and stacks it with the feature maps at different scales of the upper layer of this feature map to output the first feature.

[0144] As Figure 3 shown, the efficiently-layer-aggregated module performs pointwise convolution and feature separation operations on the input information. After the operations, it is divided into a first branch and a second branch. The first branch enters the feature stacking after passing through N inverted residual bottleneck processes, and the second branch retains the original information and enters the feature stacking. Finally, pointwise convolution is performed to output the information.

[0145] The efficiently-layer-aggregated module can efficiently learn expressive multi-scale feature representations. The input information undergoes pointwise convolution and feature separation operations and is divided into two branches. One branch retains the original information and then directly enters the feature stacking operation, while the other branch passes through N inverted residual bottleneck units. Due to the mechanism of efficient layer aggregation, the branches and the outputs through each inverted residual bottleneck unit are retained and finally connected together. For the specific structure of the inverted residual bottleneck unit, the input sequence expands the number of channels through pointwise convolution, followed by a k*k heterogeneous depthwise separable convolution operation, and finally reduces the number of channels through pointwise convolution to compensate for the information loss caused by the depthwise separable convolution. During training, the network runs n different-sized depth convolutions in parallel. This module expands the receptive field without incurring additional inference costs while retaining the information of small targets by parallelizing large-kernel convolutions with several small-kernel convolutions.

[0146] As shown Figure 4 in the figure, the local-global hybrid self-attention mechanism includes three input branches. The first input branch passes through a linear projection layer, layer normalization, one-dimensional convolution, layer normalization, and squeeze-and-excitation attention, and then outputs global information. The second input branch passes through a linear projection layer, layer normalization, multi-head self-attention, layer normalization, and squeeze-and-excitation attention, and then outputs local information. The global information, local information, and the original input are jointly input into the adaptive hybrid global and local information module to form a hybrid output. The hybrid output is transformed through a linear projection layer and then added to the original input to form a residual connection, and the output information is obtained through layer normalization processing.

[0147] The local-global hybrid self-attention mechanism combines a local convolutional layer and a global self-attention layer to jointly model the long-term and short-term preferences of users. The local preferences (captured by CNN) and global preferences (captured by Transformer) are combined to more comprehensively model the dynamic preferences of users; it adaptively hybridizes global and local information, decouples the fusion process in different layers, improves the expressive ability, and adaptively aggregates long-term and short-term preferences (the hybrid importance of local and global dependence modules can be adjusted according to the personalized needs of users); the squeeze-and-excitation attention is used to replace the softmax operation, allowing multiple relevant items to be considered simultaneously, enhancing the expressive ability of the model (allowing the model to simultaneously focus on multiple highly relevant objects instead of only focusing on a single object as in traditional softmax).

[0148] Squeeze-and-excitation attention: Receives an input tensor with a shape of [B, N, D], where B is the batch size, N is the sequence length, and D is the feature dimension.

[0149] First, calculate the average value of the input tensor along the feature dimension D to obtain a tensor average value with a shape of [B, N, 1].

[0150] Transpose the dimensions of the input tensor average value to obtain a tensor with a shape of [B, 1, N].

[0151] Transform the transposed tensor through the first linear layer to obtain a hidden layer state tensor with a shape of [B, 1, N / r], where N / r represents the dimension after compression.

[0152] Apply an activation function (ReLU) to further non-linearly activate the hidden layer state.

[0153] Transform the activated tensor back to the original sequence length [B, 1, N] through the second linear layer.

[0154] The Sigmoid function is applied to generate the attention weight tensor attention scores, with a shape of [B, 1, N].

[0155] The attention scores are transposed again to multiply with the original input tensor, resulting in a tensor with a shape of [B, N, 1].

[0156] Adaptive hybrid of global and local information, receiving three input tensors: the model input, global information, and local information.

[0157] Calculate the average value of the input tensor in the feature dimension for generating the adaptive weight.

[0158] The adaptive weight alpha is obtained through the linear projection layer and squeeze-and-excitation attention and extended to a shape of [B, 1, 1] (where B represents the input sample batch size).

[0159] beta is set to 1 - alpha, representing the remaining proportion.

[0160] According to these two weights, the global information and local information are mixed together to form the mixed output.

[0161] After the mixed output is transformed by the linear projection layer, it is added to the original input to form the residual connection.

[0162] The final output is processed through the dropout layer and layer normalization.

[0163] Dynamically adjust the proportion between the global information and local information. Use the residual connection to maintain the gradient flow and reduce the risk of gradient vanishing. Enhance the generalization ability of the model through layer normalization and dropout.

[0164] As Figure 5 shown, the relative position enhanced self-attention mechanism inputs enter the query matrix linear projection layer, key matrix linear projection layer, value matrix linear projection layer, parameter matrix linear projection layer, matrix linear projection layer, and the output of the key matrix linear projection layer respectively. After matrix multiplication and addition with the bias relative position, it is processed by the softmax function and then multiplied with the output of the value matrix linear projection layer to form the first output. The output of the parameter matrix linear projection layer acts on the tokens depthwise separable convolution and forms the second output through the tokens-level linear projection. The first output and the second output are added together to obtain the final output.

[0165] Relative Position Enhanced Self-Attention Mechanism, Importance of Position Information. For computer vision tasks, pixels or patches in an image are regarded as tokens, and their position information is crucial for constructing image semantics. A key feature of the self-attention mechanism is that it is insensitive to the permutation order of tokens, which means it cannot directly utilize the position information of tokens. Introducing Relative Position Bias: Methods such as Swin Transformer introduce additional relative position biases in the attention matrix to guide the model to learn position information and content-based attention. The relative position enhanced self-attention mechanism further develops this idea by establishing an independent value representation space to capture position dependencies. Mathematical formula of the relative position enhanced self-attention mechanism: The relative position enhanced self-attention mechanism converts tokens into new representations by introducing a parameter matrix Vrel to mix information through relative positions. The relative position enhanced self-attention mechanism uses a weight tensor brel, which is indexed by relative positions. The relative position enhanced self-attention mechanism also retains the relative position bias B to give the model greater flexibility. Limiting the Maximum Absolute Distance: To enhance the generalization ability of the model and simplify the implementation, the relative position enhanced self-attention mechanism limits the maximum absolute relative position distance s on each axis. After doing so, the calculation of the relative position enhanced self-attention mechanism can be simplified to a token-based convolution with a kernel size of (2s + 1) * (2s + 1).

[0166] The feature fusion neck consists of three layers of branches. Each layer of branch sequentially includes a shallow feature interaction module, a cross-stage chunk module, a deep feature interaction module, and a spatial context awareness module. The shallow feature interaction module receives the first features of the same layer and the previous layer, and outputs shallow features to the cross-stage chunk module. The cross-stage chunk module distributes the shallow features to each deep feature interaction module in the front and the shallow feature interaction module of the previous layer in the back. The deep feature interaction module outputs deep features to the spatial context awareness module. The triple-scale sequence fusion module obtains the first features of the same layer of each layer of branch, and after feature concatenation, inputs them into the spatial context awareness module of the first layer. The spatial context awareness module outputs global context features to the deep feature interaction module of the next layer in the back and outputs them to the prediction head.

[0167] After each efficient layer aggregation module, there is a shallow feature interaction module responsible for cross-scale information exchange and fusion. The shallow feature interaction module is followed by the cross-stage chunk module, which is a multi-branch structure that allows interaction between features at different levels. The deep feature interaction module utilizes additional deep information to enhance the feature representation. The spatial context awareness module may capture spatial relationships and generate richer feature representations.

[0168] Such as Figure 6As shown, the shallow feature interaction module receives the first features of the previous layer, the same layer, and the shallow features of the next layer. The first feature of the previous layer forms deep information through depthwise separable convolution and is stacked with the first feature of the same layer and the shallow features of the next layer to output shallow features.

[0169] For the shallow feature interaction module, accurate localization depends on the detailed edge information obtained from the shallow network, while accurate classification requires a deeper network to capture coarse-grained information. An effective feature pyramid network should support the full and sufficient fusion of the information flow between the shallow and deep networks; retaining the shallow spatial information in the backbone network is crucial for enhancing the detection ability of smaller objects. However, the information provided by the backbone network is relatively basic and vulnerable to interference. Therefore, the shallow information is incorporated into the deep network as an auxiliary branch to ensure the stability of subsequent layer learning; the main goal is to integrate the deep information with the high-resolution shallow features from the same layer and the next layer of the backbone network, aiming to retain rich localization details to enhance the spatial performance of the network; 1×1 convolution is used to control the number of channels in the shallow information, ensuring that it accounts for a relatively small proportion in the concatenation operation and does not affect subsequent learning. The shallow backbone information is maintained through bidirectional connections, enhancing the network's ability to detect small objects.

[0170] As Figure 7 shown, the deep feature interaction module receives the shallow features of the previous layer, the same layer, the next layer, and the global context feature of the previous layer in the front. The shallow features of the previous layer, the shallow features of the next layer, and the global context feature of the previous layer in the front are each processed by depthwise separable convolution and then stacked with the shallow features of the same layer to output deep features.

[0171] For the deep feature interaction module, to further improve the interactive utilization of the feature layer information, multi-scale information fusion is performed at a deeper level; these connections span the high-resolution layer of the shallow layer, the low-resolution layer of the shallow layer, the shallow layer of the same level, and the front spatial context perception for information aggregation. The gradient information of the output layer is enriched through multi-directional connections.

[0172] As Figure 8 shown, the three-scale sequence fusion module includes three branches, which respectively receive the first features of three different scales. Each is adjusted in channels through a convolutional module. The first branch is downsampled through max pooling and average pooling and then processed by a convolutional module. The third branch is upsampled using the nearest neighbor interpolation method and then processed by a convolutional module. Then, the features of the three branches are stacked, and then processed by a three-dimensional convolutional module to output the concatenated features.

[0173] The three-scale sequence fusion module. To identify densely overlapping small objects, the shape or appearance changes at different scales can be referred to and compared by magnifying the image. Since different feature layers of the backbone network have different sizes, the traditional feature pyramid fusion mechanism only upsamples the small-sized feature maps and then segments or adds them to the features of the upper layer, ignoring the rich detailed information of the large-sized feature layers. The three-scale sequence fusion module can segment the large, medium, and small-sized features, add the large-sized feature maps, and perform feature magnification to improve the detailed feature information. It can better combine the high-dimensional information of the deep feature maps and the detailed information of the shallow feature maps. The scale space is constructed along the scale axis of the image, which not only represents one scale but also represents the range of various scales that an object may have. The feature maps of different scales are horizontally stacked, and three-dimensional convolution is used to extract their scale sequence features. The high-resolution feature maps contain most of the information crucial for the detection of small targets. Before feature encoding, the number of feature channels needs to be adjusted first to make it consistent with the main scale features. After the convolutional module processes the large-sized feature maps, its number of channels is adjusted to 1C, and then a hybrid structure of max pooling and average pooling is used for downsampling, which helps to retain the effectiveness and diversity of the high-resolution features and cell images. For the small-sized feature maps, the convolutional module is also used to adjust the number of channels, and then nearest neighbor interpolation is used for upsampling. This helps to maintain the richness of the local features of the low-resolution images and prevent the loss of small target feature information. Finally, the large, medium, and small feature maps of the same size are convolved once, then concatenated in the channel dimension, and finally three-dimensional convolution, batch normalization, and activation function are performed.

[0174] As Figure 9 shown, the context-aware module includes three branches, which respectively perform global average pooling, global max pooling, pointwise convolution on the input features. After global average pooling and global max pooling, they jointly go through the softmax function, matrix multiplication, and pointwise convolution processing. The pointwise convolution goes through the softmax function, matrix multiplication, and pointwise convolution processing, and finally, they jointly perform the Hadamard product to output the global context features.

[0175] The spatial context-aware module is designed to model global context information in the cross-channel and spatial dimensions and suppress useless background features. The structure consists of three branches, and global average pooling (GAP) and global max pooling (GMP) are used to integrate global information. Its functions are: to extract the context relationship between space and channels; to enhance the feature representation of small targets and reduce background confusion; to implement the linear transformation of simplifying the feature maps using 1x1 convolution; to calculate the channel and spatial context information through matrix multiplication; and to perform the Hadamard product operation to output the feature map containing global context.

[0176] Training of the person recognition model

[0177] Training data production: Find videos containing pedestrians at different locations, different times, and different weathers in the surveillance video. Vi represents the i-th video, and there are Ni video images in Vi. Select Mi video images from the Ni video images as training and test images.

[0178] Data augmentation of training data images: Brightness enhancement, contrast enhancement, image rotation, image flipping, affine transformation to expand images, shear transformation to expand images, HSV data augmentation, translation augmentation; eight methods are used for data augmentation.

[0179] Hyperparameters for network training

[0180] Default value of input resolution: The default input resolution of the network is 640x640.

[0181] Setting suggestions: Adjust according to actual needs. Increasing the input resolution can improve the accuracy of the model, but it will also increase the computational cost and inference time. If the application scenario has high requirements for speed, the input resolution can be appropriately reduced.

[0182] Batch size: 256

[0183] Influencing factors: The batch size affects the training speed and memory consumption of the model. Setting suggestions: Select an appropriate batch size according to the size of the graphics card's video memory and actual needs. A larger batch size can speed up the training, but it will also occupy more video memory.

[0184] Learning rate (Learning Rate)

[0185] Initial learning rate: Usually 0.01 when using the SGD optimizer and 0.001 when using the Adam optimizer.

[0186] Scheduling strategy: A learning rate decay strategy such as CosineAnnealing can be adopted to adjust the learning rate.

[0187] Setting suggestions: Adjust the initial learning rate according to the specific problem, dataset, and model size, and observe the change of the loss curve during training, and adjust the learning rate in a timely manner.

[0188] Momentum

[0189] Default value: In the SGD optimizer, YOLOv5 usually uses a relatively high momentum value, such as 0.937, to improve the training speed and stability. Setting suggestions: The choice of the momentum value should be adjusted according to the actual situation. A larger momentum can accelerate the parameter update, but it may also lead to overfitting.

[0190] Weight Decay

[0191] Function: A regularization method to prevent model overfitting. Default value: The commonly used weight decay value in the network is 0.0005.

[0192] Setting suggestion: Adjust the weight decay value according to the complexity of the model and the characteristics of the dataset. A larger value will increase the regularization strength, but may also lead to model underfitting.

[0193] Number of training epochs

[0194] Definition: The number of training epochs refers to the number of times the entire training set passes through the forward and backward propagation of the model to update the parameters.

[0195] Setting suggestion: Select the appropriate number of epochs by observing the changes in loss values and metrics on the training and validation sets. Too few epochs will lead to model underfitting, while too many may lead to model overfitting.

[0196] Data Augmentation

[0197] Function: Expand the dataset by randomly transforming the training data to improve the generalization ability and robustness of the model.

[0198] Setting suggestion: There are various data augmentation methods built into the network, such as flipping, scaling, random cropping, etc. Select the appropriate data augmentation method according to the specific problem and the characteristics of the dataset.

[0199] Other hyperparameters

[0200] Model depth and width: The network introduces depth_multiple and width_multiple coefficients to obtain models of different sizes. Select the appropriate model size according to the requirements of the application scenario.

[0201] Loss function weights: The loss functions in the network include box loss, cls loss, and obj loss. The importance of different tasks can be balanced by adjusting the weights of these losses.

[0202] IOU threshold: The IOU threshold used to determine the positive and negative sample bounding boxes. A higher IOU threshold can better filter out negative samples and inaccurate positive samples, but may also lead to a reduction in the number of positive samples.

[0203] Model Function: Continuously extract visual information from video data sources in real time, and then use a carefully trained and optimized neural network to perform in-depth analysis on each frame of the image. In particular, this model is designed to accurately identify and locate human head targets in the picture. It can not only accurately mark the positions of each human head (precisely defined by bounding boxes), but also evaluate the confidence of these detection results, that is, the degree of certainty of the model about the correctness of its judgment. Through this process, intuitive and reliable human head detection and tracking results can be provided for users, enabling various application scenarios, such as people counting, etc.

[0204] Model Iteration: First, a batch of new datasets are carefully collected. These datasets cover a wide range of scenarios and conditions, providing rich materials for the further training of the model. Subsequently, the current model is used to conduct a detailed detection of these data, and according to the detection results, they are finely divided into four categories: images that clearly contain target images, images misreported as target images, images that actually have no targets but may not be detected due to conditional limitations, and images that do not contain any targets at all.

[0205] Based on this classification, targeted strategies are adopted to process various types of images. In particular, images misreported as target images are regarded as valuable negative sample resources, which help the model learn how to distinguish real targets from background noise. For those images that contain targets but are not detected by the current model, they are regarded as key training samples, which directly point out the room for improvement in the model's performance.

[0206] Next, focus on the fine data annotation of these images that have not detected targets, ensure that each target is accurately marked, and at the same time adopt a variety of data augmentation techniques, such as rotation, scaling, color adjustment, etc., to enrich the diversity of training samples and improve the generalization ability of the model.

[0207] Based on the original model, use these newly annotated and augmented data to retrain the model. By continuously optimizing learning parameters and adjusting the network structure, strive to significantly improve its detection accuracy while maintaining the model's efficiency. After the training is completed, test the model, and evaluate the improvement of the model's performance by comparing the detection results of the new and old models.

[0208] Green base station energy-saving method based on the fusion of self-attention and convolutional deep learning, algorithm code business logic implementation method for predicting the number of people. First, the interface accepts parameters such as the camera address, algorithm type, and callback address. Second, the interface will start a new process to start capturing image frames from the camera video stream and store them in Redis, and at the same time notify the monitoring program. Third, the monitoring program receives the notification, retrieves the image frames from Redis, and at the same time starts creating an algorithm instance and calls the algorithm instance to start analyzing the image frames. Fourth, the algorithm instance analyzes the image frames according to the business logic, stores the analysis results in Redis, and at the same time notifies the monitoring program. Fifth, the monitoring program receives the notification, retrieves the results, and submits the analysis results to the business interface (callback).

[0209] The person recognition model of the present invention can segment features of large, medium, and small sizes by adopting a ternary scale sequence fusion model, add large-size feature maps, and perform feature amplification to improve detailed feature information, and can better combine the high-dimensional information of the deep feature map and the detailed information of the shallow feature map; adopt a relative position enhanced self-attention mechanism to make up for the position blind spot of self-attention, and the model has a strong ability to capture dependency relationships from content and position; adopt a spatial context awareness module to extract the context relationship between space and channels, enhance the feature representation of small targets, and reduce background confusion; adopt a local-global hybrid self-attention mechanism to combine local preference and global preference to more comprehensively model the dynamic preferences of users.

[0210] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A green base station energy-saving method based on the fusion of self-attention and convolutional deep learning, characterized in that: The following steps are involved: S1. Periodic scanning time node acquisition monitoring first video image; S2. Use a person recognition model that combines self-attention and convolutional deep learning to recognize and count people in the first video image; the model includes: The backbone network extracts the first features of each scale space from the input image; Through the feature fusion neck, shallow feature interaction and deep feature interaction are performed on the first features of each scale space, and the first features of each scale space are spliced. According to the deep interaction and splicing features, context fusion is performed to output the second features of each scale space respectively; Each second feature input is used to predict the head to generate a multi-target output; The backbone network includes three high-efficiency layer aggregation modules connected in sequence. The high-efficiency layer aggregation modules extract features and output first features of different scales. The lowest-level high-efficiency layer aggregation module obtains feature maps through local-global hybrid self-attention mechanism and relative position enhanced self-attention mechanism, and stacks them with feature maps of different scales on the feature map to output the first features. The ternary scale sequence fusion module includes three layers of branches, which receive the first features of three layers of different scales respectively, and each passes through a convolution module to adjust the channel. The first layer of branches is downsampled by maximum pooling and average pooling, and then processed by the convolution module. The third layer of branches is upsampled by the nearest neighbor interpolation method, and then processed by the convolution module. The three layers of branch features are stacked and then processed by the three-dimensional convolution module to output the spliced ​​features; The context-aware module consists of three branches, which respectively perform global average pooling, global maximum pooling, and point-by-point convolution on the input features. After global average pooling and global maximum pooling, they are processed by softmax function, matrix multiplication, and point-by-point convolution. After point-by-point convolution, they are processed by softmax function, matrix multiplication, and point-by-point convolution. Finally, Hadamard product is performed together to output global context features. S3. Analyze the movement trend of people and dynamically count the number of people in the monitoring area; S31 analyzes the motion state of the character to determine whether the character will exceed the current video image range; S32. If not, count the number of people in each monitoring area and proceed to step S4; If so, it is determined that the person will appear in the monitoring area, and the second monitoring video image is obtained after the delay time, and the person is identified and counted, and the number of people in each monitoring area is counted, and step S4 is entered, and at the same time, it is determined whether the scanning cycle is exceeded; S33. If yes, enter the next scanning cycle and return to step S1; If not, return to step S31; S4. Calculate the number of users in the monitoring area and activate the corresponding base station capacity resources according to the number of users.

2. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 1 is characterized in that The person recognition model integrating self-attention and convolutional deep learning includes: The input image is converted into feature maps of multiple scale spaces through the backbone network, and the first features of each scale space are extracted and output respectively.

3. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 1 is characterized in that: The efficient layer aggregation module performs point-by-point convolution and feature separation operations on the input information, and after the operation, it is divided into a first branch and a second branch. The first branch enters the feature stack after N inverted residual bottlenecks, and the second branch retains the original information and enters the feature stack. Finally, point-by-point convolution is performed and the information is output; The local-global hybrid self-attention mechanism includes three input branches. The first input branch outputs global information after passing through a linear projection layer, layer normalization, one-dimensional convolution, layer normalization, and compressed excitation attention. The second input branch outputs local information after passing through a linear projection layer, layer normalization, multi-head self-attention, layer normalization, and compressed excitation attention. The global information, local information, and original input are input into an adaptive hybrid global and local information module to form a hybrid output. The hybrid output is transformed by a linear projection layer and then added to the original input to form a residual connection. The output information is processed by layer normalization. The relative position enhanced self-attention mechanism inputs enter the query matrix linear projection layer, the key matrix linear projection layer, the value matrix linear projection layer, and the parameter matrix linear projection layer respectively. The outputs of the matrix linear projection layer and the key matrix linear projection layer are matrix multiplied and added to the relative position of the bias. After being processed by the softmax function, they are matrix multiplied with the output of the value matrix linear projection layer to form a first output. The output of the parameter matrix linear projection layer acts on the tokens depth-separable convolution, and is linearly projected at the tokens level to form a second output. The first output is added to the second output to obtain the final output.

4. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 2 is characterized in that: The feature fusion neck includes three layers of branches, each layer of branches includes a shallow feature interaction module, a cross-stage blocking module, a deep feature interaction module, and a spatial context perception module in sequence. The shallow feature interaction module receives the first features of the same layer and the previous layer, and outputs the shallow features to the cross-stage blocking module. The cross-stage blocking module distributes the shallow features to the front deep feature interaction modules and the rear upper shallow feature interaction module. The deep feature interaction module outputs the deep features to the spatial context perception module. The ternary scale sequence fusion module obtains the first features of the same layer of each branch, and after feature splicing, it is input into the first layer of spatial context perception module. The spatial context perception module outputs the global context features to the rear next layer of deep feature interaction module and to the prediction head.

5. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 4 is characterized in that: The shallow feature interaction module receives the first features of the previous layer and the same layer, and the shallow features of the next layer. The first features of the previous layer are subjected to depth-separable convolution to form deep information, which is stacked with the first features of the same layer and the shallow features of the next layer to output shallow features; The deep feature interaction module receives the shallow features of the previous layer, the same layer, and the next layer, and the global context features of the previous layer, the shallow features of the previous layer, the shallow features of the next layer, and the global context features of the previous layer, which are respectively subjected to deep separable convolution and then stacked with the shallow features of the same layer to output deep features.

6. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 2 is characterized by: The step S3 specifically includes: S31. Analyze the motion state of the person through deep learning, determine the moving direction of the person in motion, calculate the moving trajectory of the person, and determine whether the person is about to exceed the current video image range; S32. If not, count the number of people in each monitoring area and proceed to step S4; If yes, determine the range where the person is about to appear based on the task movement trajectory, obtain the second video image of the monitoring after the set delay time, perform person recognition and counting on the second video image, count the number of people in each monitoring area, and enter step S4, while judging whether the current time exceeds the current scanning cycle; S33. If yes, enter the next scanning cycle and return to step S1; If not, return to step S31.

7. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 1 or 6, characterized in that: After step S2, the method further includes: Determine whether there is a person in the character recognition result. If yes, go to step S3; if no, notify the 5G base station to enter the waiting, sleep or deep sleep state.

8. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 1, 2 or 6, characterized in that: The calculation of the number of users in the monitoring area specifically includes: Obtain the operator user penetration rate in this business area and calculate the number of operator users in each monitoring area; Get the operator's 5G user ratio and calculate the number of operators' 5G users in each monitoring area.

9. The green base station energy saving method based on the fusion of self-attention and convolutional deep learning according to claim 1 or 2, characterized in that It also includes real-time adjustment steps, including: Calculate the expected capacity requirements of users based on the number of users in the monitoring area, and obtain the actual capacity requirements of users in the monitoring area; Compare the user's expected capacity demand with the user's actual capacity demand, and adjust the base station capacity resources in real time according to the user's actual capacity demand.

Citation Information

Patent Citations

  • Ultra-dense network resource distribution method based on access user mobility prediction

    CN108123828A

  • Feature fusion behavior recognition method and system based on deformable deep convolution and attention mechanism

    CN118196904A