Image classification methods, devices, electronic devices, storage media, and software products

By constructing a spatiotemporal feature vector and multi-scale feature fusion network to process base station survey images, the problems of low efficiency and insufficient accuracy in base station survey photo classification are solved, and efficient and accurate classification is achieved in complex environments.

CN122090134APending Publication Date: 2026-05-26CHINA MOBILE GROUP DESIGN INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE GROUP DESIGN INST
Filing Date
2026-01-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing base station survey image classification methods have low classification efficiency, subjective and difficult-to-unify classification standards, and insufficient accuracy in classifying base station survey images in complex environments, failing to effectively adapt to diverse and dynamically changing scenario requirements.

Method used

By acquiring visual image data, shooting time information, and geographical location information of base station survey images, a spatiotemporal feature vector is constructed. A multi-scale feature extraction network and a cross-modal fusion network are used to generate a fused feature representation. Finally, the image classification result is determined based on the fused feature representation.

Benefits of technology

The model's generalization ability and adaptability to survey images across multiple time periods and scenarios have been enhanced, improving classification accuracy and robustness in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090134A_ABST
    Figure CN122090134A_ABST
Patent Text Reader

Abstract

This application discloses an image classification method, apparatus, electronic device, storage medium, and program product to address the problems of low classification efficiency, subjective and difficult-to-unify classification standards, insufficient accuracy in classifying base station survey images under complex environments, and inability to effectively adapt to diverse and dynamically changing scenario requirements in existing base station survey image classification methods. The method includes: acquiring visual image data, shooting time information, and geographic location information of a target base station survey image; constructing a spatiotemporal feature vector based on the shooting time information and geographic location information; processing the visual image data using a multi-scale feature extraction network to obtain a multi-scale image feature representation of the visual image data; inputting the spatiotemporal feature vector and the multi-scale image feature representation into a cross-modal fusion network to generate a fused feature representation of the target base station survey image; and determining the classification result of the target base station survey image based on the fused feature representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly to an image classification method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] Base station surveying is crucial for the planning and construction of mobile communication networks. With the increasing number of base stations being built, how to efficiently and accurately classify and manage the large volume of survey photos has become a key concern.

[0003] Traditional methods for classifying base station reconnaissance photos mainly rely on manual classification or deep learning-based automatic classification. However, manual classification methods are inefficient, have subjective and inconsistent classification standards, and are difficult to meet the needs of large-scale engineering projects. Secondly, while existing deep learning-based automatic classification methods can automatically extract high-level semantic features of images and achieve a certain degree of automated classification, their generalization ability in classifying reconnaissance images across multiple time periods and scenarios is insufficient, resulting in low accuracy and robustness in complex environments.

[0004] Therefore, how to overcome the limitations of existing technologies and improve the classification accuracy and robustness of base station survey photos in complex environments has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] This application provides an image classification method to address the problems of existing base station survey photo classification methods, such as low classification efficiency, subjective and difficult-to-unify classification standards, insufficient accuracy in classifying base station survey images in complex environments, and inability to effectively adapt to diverse and dynamically changing scenario requirements.

[0006] This application also provides an image classification device, an electronic device, a computer-readable storage medium, and a computer program product.

[0007] The embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide an image classification method, including: Acquire visual image data, shooting time information, and geographical location information of the target base station survey images; Spatiotemporal feature vectors are constructed based on shooting time and geographic location information; A multi-scale feature extraction network is used to process visual image data to obtain multi-scale image feature representations of the visual image data; The spatiotemporal feature vectors and multi-scale image feature representations are input into a cross-modal fusion network to generate a fused feature representation of the target base station survey image; The classification results of the target base station survey images are determined based on the fusion feature representation.

[0008] Secondly, embodiments of this application provide an image classification device, including an acquisition module, a construction module, a processing module, a fusion module, and a classification module, wherein: The acquisition module is used to acquire visual image data, shooting time information, and geographical location information of the target base station survey images; The module is used to construct spatiotemporal feature vectors based on shooting time and geographic location information; The processing module is used to process visual image data using a multi-scale feature extraction network to obtain multi-scale image feature representations of the visual image data. The fusion module is used to input spatiotemporal feature vectors and multi-scale image feature representations into the cross-modal fusion network to generate fused feature representations of the target base station survey images; The classification module is used to determine the classification result of the target base station survey image based on the fused feature representation.

[0009] Thirdly, embodiments of this application provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the image classification method as described above.

[0010] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the image classification method described above.

[0011] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the image classification method described above.

[0012] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: The method provided in this application extracts multi-scale visual features from the target base station survey image and performs cross-modal fusion by combining the spatiotemporal features constructed from its shooting time information and geographical location information. Finally, the classification result of the target base station survey image is determined based on the fused feature representation. In this way, the changes of the target base station survey image in time and space can be better understood, thereby enhancing the model's generalization ability and adaptability to survey images in multiple time periods and multiple scenes, and ensuring its classification accuracy and robustness for base station survey images in complex environments. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic diagram illustrating the implementation process of an image classification method provided in this application embodiment; Figure 2 This application provides a schematic diagram of the specific structure of an image classification device according to an embodiment of the present application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] As will be known to those skilled in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0016] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0017] It should be understood that the training and prediction processes of the AI ​​models involved in the various embodiments of this specification all adhere to multiple legal and compliant principles, including legal data sources, compliant data content, compliant data governance, compliant training objectives and schemes, compliant training processes, compliant training environments and tools, and compliant ethical verification of training results, and comply with the requirements of Article 5 of the Patent Law. Among them: Data source legitimacy: All datasets used for AI model training were obtained through legal means, covering three categories: publicly authorized data, data authorized by partners, and self-collected compliant data. Publicly authorized data comes from compliant data sources following open-source licenses such as Apache 2.0, with complete copyright attribution and authorization scope clearly marked, and no unauthorized open-source code or data reuse. Data authorized by partners has been subject to formal data usage agreements, clearly defining the scope, duration, and confidentiality obligations, and possessing a complete authorization chain. For self-collected data involving personal information, strict informed consent procedures have been followed, and anonymization processes (including but not limited to field masking, feature anonymization, and differential privacy technology applications) have been implemented to remove personally identifiable information, fully complying with the requirements of relevant laws and regulations such as the "Interim Measures for the Administration of Generative Artificial Intelligence Services" and the "Personal Information Protection Law."

[0018] Data content compliance: The AI ​​model's dataset undergoes multiple screenings and cleaning processes to remove all content that may violate social morality or harm public interests. It contains no obscene, pornographic, violent, discriminatory, or information that endangers national or public safety, nor does it involve the illegal acquisition or use of genetic resources. For data in sensitive fields (such as healthcare and finance), an additional privacy-preserving computation module (including federated learning and secure multi-party computation technologies) ensures that the data is "usable but not visible," avoiding compliance risks during the original data transmission process and ensuring that the data application scenarios and uses comply with public order and good morals and industry regulatory requirements.

[0019] Data governance norms: A complete data traceability system is established during the AI ​​model training process to automatically record the source, collection time, annotation process, cleaning rules, and permission allocation of training data, generating traceable compliance reports to ensure that the data is verifiable throughout its entire lifecycle. The dataset annotation process for AI models is completed by a professional human R&D team, clearly defining the proportion of human creative contributions and avoiding reliance on AI-generated data that has not undergone substantial human modification, thus meeting the examination requirements for "human main contributions" in AI patent applications.

[0020] Training objectives and plans are compliant: The AI ​​model training focuses on base station reconnaissance image classification scenarios, such as base station construction, mobile communication network optimization, and automated processing of reconnaissance data. It aims to improve the classification accuracy and robustness of base station reconnaissance images through deep learning technology to meet the needs of large-scale image data management in modern communication network construction. The training scheme and final output results do not violate any mandatory provisions of laws or administrative regulations, do not harm public interests or the legitimate rights and interests of others, and pose no potential risk of being used for illegal activities, privacy violations, or disruption of public safety. It strictly adheres to the ethical principle of "intelligent for good." In particular, during data processing and model training, all image data used is legally authorized reconnaissance data, and the collection, use, and storage of data strictly comply with relevant laws and regulations on privacy protection and data security to avoid any negative impact on personal privacy, public safety, or social order.

[0021] Training process compliance: A closed-loop training framework is adopted to ensure compliance and controllability of the training process. The specific process is as follows: First, training samples are obtained through compliant data sources. After the aforementioned data cleaning and desensitization, they are input into the neural network model to generate preliminary training results. Second, an expert system is introduced to verify the preliminary results. Based on preset rules and human expert experience, the feasibility of the results is evaluated, and outputs that may pose ethical risks or compliance hazards are corrected (such as removing decision-making logic that violates public order and good morals, and adjusting model parameters that do not comply with safety regulations). Finally, the loss function weights are dynamically optimized based on expert system feedback to strengthen the model's learning of compliant results, avoid overfitting errors or non-compliant labels, and form a closed-loop control of "data input - model training - expert verification - parameter optimization - result feedback" to ensure that the entire training process complies with A5 ethical review requirements.

[0022] Training environment and tool compliance: AI model training is implemented using nationally licensed chips and a compliant training platform. All open-source frameworks and components used in the training process have obtained their corresponding licenses, and copyright statements and patent citation information are fully retained, with no instances of infringement or reuse. The training environment is built using virtual devices (containers / virtual machines) with fixed random seeds and initial parameter configurations to ensure the reproducibility of the training process. Furthermore, through access control and operation log recording, risks such as data leakage and parameter tampering during training are prevented, ensuring the security and compliance of the training process.

[0023] Training results ethical verification compliance: After the model is trained, it undergoes additional third-party ethical compliance assessment and algorithm filing review to verify that the model output does not violate social morality or harm public interests. For potentially sensitive scenarios (such as public services and intelligent decision-making), a special result verification mechanism is established to ensure that the model always complies with Article 5 of the Patent Law and relevant laws and regulations in practical applications.

[0024] In summary, the data and training process used in the AI ​​model of this specification strictly comply with the relevant provisions of Article 5 of the Patent Law and the Patent Examination Guidelines (2023 Edition), and there are no violations of laws, social ethics, public interests, or illegal use of genetic resources. It fully meets the compliance requirements for patent authorization.

[0025] To address the problems of low classification efficiency, subjective and difficult-to-unify classification standards, insufficient accuracy in classifying base station survey images in complex environments, and inability to effectively adapt to diverse and dynamically changing scenario requirements, this application provides an image classification method.

[0026] The execution subject of this method can be various types of computing devices, or it can be an application or app installed on the computing device. The computing device can be a user terminal such as a mobile phone, tablet computer, or smart wearable device, or it can be a server.

[0027] For ease of description, this application uses a server as the execution subject of the method in its embodiments to illustrate the method. Those skilled in the art will understand that this embodiment uses a server as an example to describe the method, which is merely an illustrative example and does not limit the scope of protection of the corresponding claims.

[0028] Specifically, the implementation flow of the method provided in this application embodiment is as follows: Figure 1 As shown, it includes the following steps: Step 102: Obtain visual image data, shooting time information, and geographical location information of the target base station survey image.

[0029] Base station survey images refer to images taken during base station survey work. For example, these images can be taken by survey personnel and / or survey equipment (such as inspection cameras, drones, etc.) during base station construction or inspection.

[0030] Correspondingly, target base station survey images refer to base station survey images to be classified, obtained from storage systems or real-time acquisition devices.

[0031] Visual image data are RGB or grayscale images used to characterize the physical state of the base station, equipment layout, environmental conditions, etc.

[0032] Shooting time information is used to record the specific moment the image was acquired.

[0033] Geographic location information, including latitude and longitude coordinates, can be used to construct spatial distribution relationships.

[0034] In some embodiments, considering that visual image data, shooting time information, and geographic location information are usually stored in the form of image file metadata (EXIF), the visual image data, shooting time information, and geographic location information of the target base station survey image can be obtained from the metadata (EXIF) through image processing libraries such as OpenCV, PIL, or through image-specific parsing modules.

[0035] Step 104: Construct a spatiotemporal feature vector based on shooting time information and geographic location information.

[0036] In this embodiment of the application, encoding can be performed based on shooting time information and geographical location information to obtain time encoding and geographical encoding; then, the time encoding and geographical encoding are concatenated to obtain a spatiotemporal feature vector.

[0037] In some embodiments, the shooting time information can be encoded using methods such as periodic encoding, sine and cosine position encoding, learnable time encoding, or relative time difference encoding to obtain time encoding.

[0038] For example, taking the sine and cosine position encoding method as an example, the time code can be obtained by encoding based on the shooting time information as follows:

[0039] in, Image representing the target base station survey i The time code corresponding to the shooting time information; Image representing the target base station survey i The shooting time information; The dimension index representing the time encoding; This represents the total dimension of the time-encoded vector.

[0040] In some embodiments, geographic location information can be encoded using methods such as discretization coding, spherical projection coding, distance-based coding, learnable location embedding, or hybrid coding to obtain geocoding.

[0041] It should be noted that the time encoding and geocoding methods listed above are merely illustrative examples in the implementation of this application and do not impose any limitations on the embodiments of this application.

[0042] After obtaining the time code and geo code, the time code and geo code can be concatenated along the feature dimension to generate a spatiotemporal feature vector.

[0043] Step 106: Use a multi-scale feature extraction network to process the visual image data to obtain multi-scale image feature representations of the visual image data.

[0044] To fully acquire image details and semantic information at different scales in the target base station survey images, this embodiment employs a multi-scale feature extraction network, such as a Feature Pyramid Network (FPN), to process visual image data. This network extracts feature maps of visual image data at different resolutions through multiple convolutional layers, and then fuses these feature maps—including low-resolution high-semantic features and high-resolution detail features—through upsampling and feature fusion mechanisms, forming a multi-scale image feature representation with rich semantic and spatial information.

[0045] Specifically, in one implementation, a target base station survey image can be input into a multi-scale feature extraction network. After receiving the target base station survey image, the multi-scale feature extraction network can extract feature maps of the visual image data of the target base station survey image at different resolutions through multiple convolutional layers, such as... Then, by using transposed convolution or upsampling, the feature maps at different resolutions are unified to the original resolution to obtain multi-scale image feature representations.

[0046] For example, using the above sampling method, feature maps at different resolutions can be unified to the original resolution using the following calculation method to obtain a multi-scale image feature representation:

[0047] in, This represents multi-scale image feature representation; The fusion weight / weighting coefficient represents the feature at the i-th scale. It can be a fixed weight (such as a preset value) or an adaptively learned weight (such as generated through an attention mechanism), which is used to adjust the contribution of features at different scales in the final multi-scale image feature representation. This indicates that feature maps at different resolutions are upsampled.

[0048] Step 108: Input the spatiotemporal feature vector and multi-scale image feature representation into the cross-modal fusion network to generate the fused feature representation of the target base station survey image.

[0049] In this embodiment, the cross-modal fusion network can be a Transformer-based fusion network. When generating the fused feature representation of the target base station reconnaissance image, a feature vector mapping transformation can be performed on the multi-scale image feature representation and the spatiotemporal feature vector. Specifically, the multi-scale image feature representation is input into three independent fully connected layers (or linear transformation layers) to generate a first query vector, a first key vector, and a first value vector, respectively. Simultaneously, the spatiotemporal feature vector is input into three other independent fully connected layers (or linear transformation layers) to generate a second query vector, a second key vector, and a second value vector, respectively.

[0050] Next, using a scaled dot product attention mechanism, the first attention weights of the first query vector and the second key vector are calculated, and the second value vector is weighted and summed according to the first attention weights to obtain the first fused feature. These first attention weights represent the degree of attention the image features pay to different parts of the spatiotemporal features.

[0051] The mathematical expression for the scaling dot product attention mechanism is as follows:

[0052] in, This represents the output of the scaled dot product attention mechanism, also known as the attention aggregation result or fusion feature; These represent the query vector, key vector, and value vector, respectively. This represents the dimension of the key vector. Represents the normalization function; This represents the dot product similarity score matrix, which is the degree of matching or relevance between the i-th query vector and the j-th key vector.

[0053] Additionally, the second attention weights of the second query vector and the first key vector are calculated, and the first value vector is weighted and summed according to the second attention weights to obtain the second fusion feature; Finally, the obtained first and second fusion features are integrated. This integration includes either linear projection after feature concatenation or direct element-wise addition. Through this operation, the spatiotemporal feature vectors and multi-scale image feature representations can fully interact and complement each other under the guidance of a bidirectional attention mechanism, ultimately outputting a unified fusion feature representation rich in multimodal information for subsequent classification decisions.

[0054] Step 110: Determine the classification result of the target base station survey image based on the fused feature representation.

[0055] The fused feature representation can be a fixed-length feature vector or a feature sequence consisting of multiple tokens, such as a sequence formed by flattening multi-scale image features in the spatial dimension and fusing the spatiotemporal context to obtain the feature sequence.

[0056] For example, taking the fused feature representation as a feature sequence, when determining the classification result of the target base station survey image, the feature sequence can be converted into a fixed-dimensional global representation vector through feature aggregation operation; wherein, the feature aggregation operation can include global average pooling, max pooling or attention pooling, etc.

[0057] Subsequently, a classification mapping is performed on the global representation vector to output the predicted scores for each preset category. Optionally, the image classification model can employ a linear classification head or a multilayer perceptron (MLP) classification head containing at least one layer of nonlinear mapping. Specifically, the global representation vector can first undergo dimensionality transformation through a fully connected layer of the image classification model, then be given nonlinear expressive power through an activation function (such as ReLU or GELU), and finally output a vector or classification score consistent with the number of categories through an output layer. During the inference phase, the classification score is input into a Softmax normalization function to obtain the classification probability distribution of the target base station survey image belonging to each category, for example:

[0058] in, This represents the probability of classifying the target base station reconnaissance image into its category. This represents the fusion feature representation; The weight matrix representing the classification heads; This represents the bias vector of the classification head; Softmax is used to normalize the scores of each category to non-negative probability values ​​that sum to 1. Finally, the category with the highest probability is determined as the classification result of the target base station reconnaissance image.

[0059] For example, suppose three survey image categories are pre-defined, corresponding to Category 1, Category 2, and Category 3, respectively. Their specific meanings can be defined by the actual engineering labeling system. (This refers to the survey images targeting the base station.) i The spatiotemporal feature vectors and multi-scale image features are used to generate a cross-modal fusion network that outputs a fused feature representation. This fused feature representation is then converged to obtain a 256-dimensional global feature vector. Then, the image classification model uses a linear classification head (fully connected layer) to classify the data. This is mapped to a 3D classification score vector. Assume that in one inference iteration, we obtain: .

[0060] Will By inputting Softmax, we can obtain the class probability distribution: exp(1.2) = 3.32; exp(2.8) = 16.445; exp(0.4) = 1.492; The sum of these three is 3.32 + 16.445 + 1.492 = 21.257.

[0061] Therefore, the classification probabilities for each category are: Classification probability of category 1 ; Classification probability of category 2 ; Classification probability of category 3 ; The maximum probability is 0.774, corresponding to category 2. Therefore, the classification result of the target base station survey image can be determined as category 2.

[0062] It should be noted that the above-described methods for determining classification results based on fusion feature representation are merely illustrative examples in the embodiments of this application and do not impose any limitations on the embodiments of this application.

[0063] In one alternative implementation, considering that when determining the classification results of target base station survey images based on fused feature representations, problems may arise such as poor image quality, environmental interference, or limited viewpoint, leading to low model confidence or prediction results that contradict the common sense of spatiotemporal continuity. To avoid this problem, after obtaining the classification results of the target base station survey images, a spatiotemporal graph neural network can be further introduced to collaboratively optimize the classification results of batch images. Specifically, the shooting time and geographical location information of the target base station survey images can be used to construct a spatiotemporal graph structure (or spatiotemporal graph network) from the independent target base station survey images. Then, a graph convolutional network (GCN) is used to aggregate the information of neighboring nodes in this spatiotemporal graph structure to smooth and correct the classification results.

[0064] When constructing a spatiotemporal map structure based on shooting time and geographic location information, each target base station survey image can be used as a node in the spatiotemporal map structure. The features of each node are initialized as its fused feature representation or features derived from its preliminary classification probability.

[0065] Next, if two target base station survey images are relatively close in both time and space, an edge is considered to exist between them, indicating a strong spatiotemporal correlation. To quantify the strength of this correlation, edge weights can be introduced. These weights are calculated based on the temporal and spatial proximity of the target base station survey images, as follows:

[0066] in, Representing nodes in a spacetime graph structure i and nodes jThe edge weight between them; Represents the target node i and nodes j The spherical geographical distance between them; The weighting coefficient / scale balance coefficient of the time difference term is used to balance the relative contributions of spatial distance and time difference to edge weights; Represents a node i The shooting time information; Represents a node j The shooting time information; This represents the decay scale parameter, used to control how quickly the edge weights decay as the overall spatiotemporal difference increases.

[0067] in, The calculation method is as follows:

[0068] This represents the Earth's average radius; Represents a node i and nodes j The difference in latitude between them; Represents a node i Latitude; Represents a node j Latitude; Represents a node i and nodes j The difference in longitude between them.

[0069] The constructed spatiotemporal graph structure is input into a graph convolutional network (GCN). The GCN is used to propagate and update the spatiotemporal graph structure so that each node in the spatiotemporal graph structure aggregates the feature information of its neighboring nodes. The aggregation weight is determined by the aforementioned edge weights.

[0070] The node feature update of a single-layer graph convolutional network can be represented as:

[0071] in, Indicates the ( ) l +1) Nodes in layer graph convolution i eigenvectors; Indicates the first l Nodes in layer graph convolution j eigenvectors; Represents a non-linear activation function; Represents a node i The set of neighboring nodes; Indicates the first l Learnable weight matrix for layer graph convolution; and Representing nodes respectively i and nodes j The degree of a node refers to the number of edges it is connected to, reflecting its connectivity or influence in the spatiotemporal graph structure.

[0072] Through multi-layered GCN propagation, the features of each node incorporate information from its spatiotemporal neighbors. Ultimately, the updated node features are... An updated classification probability distribution is obtained by using a classification layer, such as a linear layer plus softmax. Then, the classification results of the target base station reconnaissance images are adjusted based on this updated probability distribution.

[0073] For example, the category with the highest probability in the updated classification probability distribution can be used as the final optimized classification result.

[0074] Alternatively, the initial distribution probability of the target base station survey image determined based on the fusion feature representation can be weighted and fused with the updated probability, and finally the adjusted classification result can be determined based on the weighted fusion.

[0075] In this way, by explicitly modeling the spatiotemporal map structure and utilizing the spatiotemporal correlation constraints between survey images, it is possible to effectively correct misclassifications caused by insufficient information or noise in a single image, significantly improving the overall consistency, robustness, and reliability of the classification results. The optimization effect is even more obvious for samples with low confidence or in complex environments.

[0076] In an optional implementation, before performing steps 106 and 108, the multi-scale feature extraction network and the cross-modal fusion network need to be trained first. Specifically, considering that base station survey images are usually collected by different personnel and equipment in complex outdoor scenes, they are easily affected by factors such as changes in lighting intensity, weather differences, shooting angle shifts, occlusion, blurring due to shaking, and cluttered backgrounds. At the same time, in actual engineering, the number of high-quality labeled samples is often limited, which can easily lead to problems such as overfitting, insufficient generalization ability to new scenes, and weak robustness to low-quality images during the model training stage. Therefore, when training the multi-scale feature extraction network and the cross-modal fusion network, they can be trained based on the base station survey image training dataset after image enhancement.

[0077] Optionally, image enhancement processing includes at least one of basic geometric transformation enhancement, color perturbation enhancement, blur simulation enhancement, and composite sample enhancement.

[0078] Basic geometric transformation enhancement refers to increasing the diversity of training data by transforming the geometry of the original image. Common geometric transformations include rotation, translation, scaling, and flipping.

[0079] In some embodiments, the original image can be randomly rotated, with the rotation angle controlled within ±15°, to enhance the adaptability of the multi-scale feature extraction network and the cross-modal fusion network to different camera angles; at the same time, a horizontal flipping operation of the image is introduced to improve the learning ability of the multi-scale feature extraction network and the cross-modal fusion network to structural symmetry features.

[0080] Color perturbation enhancement refers to simulating changes in images captured under different lighting conditions and environments by altering the color characteristics of the image (such as brightness, contrast, saturation, etc.).

[0081] In some embodiments, the brightness, saturation, and contrast of the image can be randomly changed, with an adjustment range set to ±20%, to simulate color changes caused by differences in weather, lighting, or shooting equipment. The processing function is shown below:

[0082] in, These are the pixel values ​​of the original image. These are the pixel values ​​after color perturbation enhancement. This is the contrast adjustment factor. Adjust the offset for brightness.

[0083] Blur simulation enhancement refers to simulating image blur caused by factors such as unstable shooting or focal length shift, in order to help the model improve its ability to process low-quality images.

[0084] In some embodiments, Gaussian blurring can be introduced to achieve blurred simulation enhancement. Specifically, a two-dimensional Gaussian kernel function can be used to convolve the image. The expression of the two-dimensional Gaussian kernel function is as follows:

[0085] Represents the planar coordinates with the center of the Gaussian kernel as the origin, that is, the horizontal and vertical distances relative to the center point. It is the standard deviation of the Gaussian distribution, which determines the degree of diffusion of the Gaussian kernel, i.e. the intensity of fuzziness; Indicates coordinates The weight values ​​at each point are used to perform a weighted average on the image to achieve a blurring effect. It is a normalization factor that ensures the total weight of the Gaussian kernel is 1, thus maintaining the overall brightness of the image.

[0086] Composite sample enhancement involves combining at least two images to generate new sample data, such as CutMix enhancement.

[0087] In some embodiments, a region can be randomly cropped from one image and replaced with a corresponding region from another image. In this way, the generated image not only contains visual information from both images but also combines features from both labels, as specifically represented below:

[0088]

[0089] in, This represents the generated augmented sample; The class label of the generated enhanced sample is represented by M; M is a binary mask matrix used to define the location of the cut-off region. , This refers to the two original images to be enhanced, specifically, This indicates the image that has been cut and replaced; This indicates an image that provides replacement content; , Two original images and Corresponding category tags; A random variable used to control the mixing ratio of labels.

[0090] Based on the above description, the base station survey image training dataset after image enhancement can be obtained in the following way: obtain the base station survey image training dataset; perform image enhancement processing on the base station survey image training dataset to obtain the base station survey image training dataset after image enhancement; wherein, the image enhancement processing includes at least one of basic geometric transformation enhancement, color perturbation enhancement, fuzzy simulation enhancement and composite sample enhancement.

[0091] The training dataset constructed in the above manner can significantly improve the diversity and coverage of training samples, reduce the risk of overfitting the model to specific lighting / angle / background, thereby improving the stable extraction capability of multi-scale feature extraction networks for details at different scales, and enhancing the generalization ability of cross-modal fusion networks to jointly model spatiotemporal context and visual semantics in complex environments, ultimately improving the quality of fused feature representation and the classification accuracy of target base station survey images.

[0092] In one alternative implementation, considering the often unbalanced class distribution in base station survey images during actual acquisition and annotation—for example, the number of common class samples far exceeds the number of minority / rare fault class samples—the model is more likely to favor the class with the larger sample size during training, resulting in insufficient learning of the minority class. This leads to phenomena such as low minority class recall, severe false positives and false negatives, and decision boundaries being interfered with by noisy samples, which may affect the convergence stability of the multi-scale feature extraction network and the cross-modal fusion network, as well as the final classification accuracy. Therefore, when training the multi-scale feature extraction network and the cross-modal fusion network, a balanced training dataset can also be used.

[0093] Specifically, a balanced training dataset can be obtained in the following way: First, the training dataset of base station survey images is denoised, for example, by using Tomek Link denoising to remove noisy samples located at class boundaries that are prone to confusion, thereby reducing the interference of overlapping regions on model training.

[0094] Then, minority class samples are identified, that is, samples of a class whose number of samples in the denoised training dataset is less than a preset threshold, and the local sparsity of these minority class samples is calculated. In this embodiment, for example, the sparsity of the region where the sample is located can be characterized by the Center Offset Weight (COW) value or the neighborhood density index.

[0095] Assigning a synthetic weight to each minority class sample based on local sparsity gives higher synthetic weights to minority class samples located in sparse or poorly covered regions, thereby prioritizing the enhancement of the coverage of that class of samples.

[0096] Subsequently, minority class samples are oversampled according to the synthetic weights to generate new synthetic samples. For example, nearest neighbors are selected from the sample neighborhood according to the weight ratio and interpolated to generate the samples. If necessary, semantic perturbations are introduced to improve diversity and realism, thereby expanding the number of minority class samples.

[0097] During oversampling, for each minority class sample, its local density can be used as a reference. Adaptive adjustment k Number of neighbors This allows for the use of more neighbors to generate diverse samples in sparse areas and fewer neighbors in dense areas to avoid generating redundant samples.

[0098] For the selected minority class samples and his neighbors It can be based on weight The proportion of randomly generated interpolated samples:

[0099] in, , This indicates a new synthetic sample.

[0100] The generated new samples are subjected to semantic preservation perturbation, and the sample feature space is decomposed using principal component analysis (PCA). Low-variance principal component directions are selected. And add Gaussian noise with an amplitude of λ in that direction. The enhanced new sample is obtained:

[0101] Finally, the new synthetic samples are merged with the denoised training dataset to obtain a balanced training dataset.

[0102] The balanced training dataset constructed in the above manner can reduce class bias during the training phase, enabling the model to obtain more balanced gradient updates and feature learning opportunities for each class, especially the minority class, during the optimization process. This improves the model's ability to identify minority / abnormal scenarios and its recall rate, enhances the reliability and robustness of the decision boundary, and ultimately improves the overall training effect and classification accuracy of the multi-scale feature extraction network and the cross-modal fusion network.

[0103] In one alternative implementation, to enhance the model's generalization ability and its understanding of spatiotemporal semantics, a multi-task loss function can be used for joint optimization when training the multi-scale feature extraction network and the cross-modal fusion network. The multi-task loss function includes at least two of the following: main classification loss, time period prediction loss, geographic location clustering triplet loss, and contrastive learning loss.

[0104] For example, multi-task loss function It can be represented as:

[0105] in, This represents the primary classification loss; Indicates the predicted loss over a given time period; This represents the loss from the geographic location clustering triplet; This represents the learning loss compared to the comparison.

[0106] in, Standard cross-entropy loss can be used:

[0107] Time-based prediction loss is used to improve the model's understanding of the time dimension; Geographic location clustering triplet loss is used to shorten the feature distance between samples from the same geographic location and widen the feature distance between samples from different locations. Contrastive learning loss is used to enhance the model's discriminative ability under different acquisition conditions.

[0108] By employing a multi-task loss function for joint optimization, the model can better understand the spatiotemporal distribution characteristics of survey photos while maintaining classification accuracy, thereby maintaining stable performance in complex environments.

[0109] In one alternative implementation, to enhance the interpretability and classification efficiency of the model, thereby further improving the overall classification accuracy and usability, an interpretable heatmap of the target base station reconnaissance image can also be generated; wherein, the interpretable heatmap is used to indicate the image regions that the multi-scale feature extraction network and the cross-modal fusion network focus on during classification.

[0110] The method provided in this application extracts multi-scale visual features from the target base station survey image and performs cross-modal fusion by combining the spatiotemporal features constructed from its shooting time information and geographical location information. Finally, the classification result of the target base station survey image is determined based on the fused feature representation. In this way, the changes of the target base station survey image in time and space can be better understood, thereby enhancing the model's generalization ability and adaptability to survey images in multiple time periods and multiple scenes, and ensuring its classification accuracy and robustness for base station survey images in complex environments.

[0111] To address the problems of low classification efficiency, subjective and inconsistent classification standards, insufficient accuracy in classifying base station survey images under complex environments, and inability to effectively adapt to diverse and dynamically changing scenario requirements in existing base station survey image classification methods, this application provides an image classification device. A schematic diagram of the specific structure of this device is shown below. Figure 2 As shown, it includes an acquisition module 21, a construction module 22, a processing module 23, a fusion module 24, and a classification module 25. The functions of each module are as follows: The acquisition module 21 is used to acquire visual image data, shooting time information and geographical location information of the target base station survey image; Module 22 is used to construct spatiotemporal feature vectors based on shooting time information and geographic location information; Processing module 23 is used to process visual image data using a multi-scale feature extraction network to obtain multi-scale image feature representations of the visual image data; The fusion module 24 is used to input spatiotemporal feature vectors and multi-scale image feature representations into the cross-modal fusion network to generate a fused feature representation of the target base station survey image; Classification module 25 is used to determine the classification result of the target base station survey image based on the fused feature representation.

[0112] Optionally, the multi-scale feature extraction network and the cross-modal fusion network are trained on a balanced training dataset; The balanced training dataset is obtained through the following steps: The base station survey image training dataset is denoised to obtain a denoised training dataset. Calculate the local sparsity of minority class samples in the denoised training dataset, and assign a synthetic weight to each minority class sample based on the local sparsity. Minority class samples are class samples in the denoised training dataset whose sample number is less than a preset threshold. Each minority class sample is oversampled according to the synthesis weights to generate new synthetic samples; The new synthetic samples are merged with the denoised training dataset to obtain a balanced training dataset.

[0113] Optionally, the multi-scale feature extraction network and the cross-modal fusion network are trained on a training dataset of base station reconnaissance images after image enhancement; The training dataset of enhanced base station reconnaissance images was obtained through the following steps: Obtain the training dataset of base station survey images; Image enhancement processing is performed on the base station survey image training dataset to obtain the image-enhanced base station survey image training dataset. Image enhancement processing includes at least one of the following: basic geometric transformation enhancement, color perturbation enhancement, fuzzy simulation enhancement, and composite sample enhancement.

[0114] Optionally, the image classification device is also used for: Construct a spatiotemporal diagram structure based on shooting time and geographic location information; Graph convolutional networks are used to propagate features of spatiotemporal graph structures in order to update the classification probabilities of nodes in the spatiotemporal graph structure. Adjust the classification results of the target base station survey images based on the updated classification probabilities; In this spatiotemporal graph structure, the nodes are the target base station survey images, the edges represent the spatiotemporal relationships between different target base station survey images, and the edge weights are calculated based on the temporal and spatial proximity of the target base station survey images.

[0115] Optional, building module 22, for: The shooting time information is encoded using sine and cosine position coding to obtain the time code; Geographic location information is encoded using a spherical geographic distance-based coding method to obtain a geographic code; The spatiotemporal feature vector is obtained by concatenating the time code and the geographic code.

[0116] Optionally, the cross-modal fusion network is a Transformer-based fusion network, with a fusion module used for: The multi-scale image feature representation is mapped to a first query vector, a first key vector, and a first value vector; Map the spatiotemporal feature vectors to a second query vector, a second key vector, and a second value vector; Using the scaling dot product attention mechanism, the first attention weights of the first query vector and the second key vector are calculated, and the second value vector is weighted and summed according to the first attention weights to obtain the first fused feature; Calculate the second attention weights between the second query vector and the first key vector, and then perform a weighted summation of the first value vector based on the second attention weights to obtain the second fusion feature; The first fusion feature and the second fusion feature are fused together to obtain the fusion feature representation.

[0117] Optionally, the multi-scale feature extraction network is a feature pyramid network, and the processing module is used for: Feature maps of visual image data at different resolutions are extracted using a feature pyramid network. By fusing feature maps at different resolutions, a multi-scale image feature representation is obtained.

[0118] Optionally, the image classification device further includes a training module, which is used for: When training the multi-scale feature extraction network and the cross-modal fusion network, a multi-task loss function is used for joint optimization. The multi-task loss function includes at least two of the following: main classification loss, time period prediction loss, geographic location clustering triplet loss, and contrastive learning loss.

[0119] Optionally, the image classification device is also used for: Generate interpretable heatmaps of target base station reconnaissance images; Interpretable heatmaps are used to indicate the image regions that multi-scale feature extraction networks and cross-modal fusion networks focus on during classification.

[0120] The apparatus provided in this application extracts multi-scale visual features from the target base station survey image and performs cross-modal fusion by combining the spatiotemporal features constructed from its shooting time information and geographical location information. Finally, the classification result of the target base station survey image is determined based on the fused feature representation. In this way, the changes of the target base station survey image in time and space can be better understood, thereby enhancing the model's generalization ability and adaptability to survey images in multiple time periods and multiple scenes, and ensuring its classification accuracy and robustness for base station survey images in complex environments.

[0121] Figure 3To illustrate the hardware structure of an electronic device according to various embodiments of this application, the electronic device 300 includes, but is not limited to, components such as: a radio frequency unit 301, a network module 302, an audio output unit 303, an input unit 304, a sensor 305, a display unit 306, a user input unit 307, an interface unit 308, a memory 309, a processor 310, and a power supply 311. Those skilled in the art will understand that... Figure 3 The electronic device structures shown are not intended to limit the electronic device. An electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. In the embodiments of this application, the electronic device includes, but is not limited to, mobile phones, tablets, laptops, PDAs, in-vehicle terminals, wearable devices, and pedometers.

[0122] The processor 310 is used to acquire visual image data, shooting time information, and geographical location information of the target base station survey image; construct a spatiotemporal feature vector based on the shooting time information and geographical location information; process the visual image data using a multi-scale feature extraction network to obtain a multi-scale image feature representation of the visual image data; input the spatiotemporal feature vector and the multi-scale image feature representation into a cross-modal fusion network to generate a fused feature representation of the target base station survey image; and determine the classification result of the target base station survey image based on the fused feature representation.

[0123] The memory 309 is used to store a computer program that can run on the processor 310, which, when executed by the processor 310, performs the functions described above by the processor 310.

[0124] It should be understood that, in this embodiment, the radio frequency unit 301 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink data from the base station and processes it with the processor 310; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 301 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 301 can also communicate with networks and other devices via a wireless communication system.

[0125] The electronic device provides users with wireless broadband internet access through network module 302, such as helping users send and receive emails, browse web pages, and access streaming media.

[0126] The audio output unit 303 can convert audio data received by the radio frequency unit 301 or the network module 302 or stored in the memory 309 into audio signals and output them as sound. Furthermore, the audio output unit 303 can also provide audio output related to specific functions performed by the electronic device 300 (e.g., call signal reception sound, message reception sound, etc.). The audio output unit 303 includes a speaker, a buzzer, and a receiver, etc.

[0127] Input unit 304 is used to receive audio or video signals. Input unit 304 may include a graphics processing unit (GPU) 3041 and a microphone 3042. The GPU 3041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on display unit 306. The image frames processed by GPU 3041 can be stored in memory 309 (or other storage media) or transmitted via radio frequency unit 301 or network module 302. Microphone 3042 can receive sound and process such sound into audio data. The processed audio data can be converted into a format that can be transmitted to a mobile communication base station via radio frequency unit 301 in telephone call mode.

[0128] The electronic device 300 also includes at least one sensor 305, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 3061 according to the ambient light level, and the proximity sensor can turn off the display panel 3061 and / or backlight when the electronic device 300 is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used to identify the posture of the electronic device (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. The sensor 305 may also include a fingerprint sensor, pressure sensor, iris sensor, molecular sensor, gyroscope, barometer, hygrometer, thermometer, infrared sensor, etc., which will not be described in detail here.

[0129] The display unit 306 is used to display information input by the user or information provided to the user. The display unit 306 may include a display panel 3061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0130] User input unit 307 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of electronic devices. Specifically, user input unit 307 includes touch panel 3071 and other input devices 3072. Touch panel 3071, also known as a touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 3071). Touch panel 3071 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to processor 310, which receives and executes commands from processor 310. In addition, touch panel 3071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to touch panel 3071, user input unit 307 may also include other input devices 3072. Specifically, other input devices 3072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here.

[0131] Furthermore, the touch panel 3071 can cover the display panel 3061. When the touch panel 3071 detects a touch operation on or near it, it transmits the information to the processor 310 to determine the type of touch event. Subsequently, the processor 310 provides corresponding visual output on the display panel 3061 based on the type of touch event. Although in Figure 3 In this embodiment, the touch panel 3071 and the display panel 3061 are two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 3071 and the display panel 3061 can be integrated to realize the input and output functions of the electronic device. The specific implementation is not limited here.

[0132] Interface unit 308 serves as an interface for connecting external devices to electronic device 300. For example, external devices may include a wired or wireless headphone port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 308 can be used to receive input from external devices (e.g., data, power, etc.) and transmit the received input to one or more components within electronic device 300, or it can be used to transmit data between electronic device 300 and external devices.

[0133] The memory 309 can be used to store software programs and various data. The memory 309 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 309 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0134] The processor 310 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 309, and by calling data stored in the memory 309, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. The processor 310 may include one or more processing units; preferably, the processor 310 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 310.

[0135] The electronic device 300 may also include a power supply 311 (such as a battery) that supplies power to various components. Preferably, the power supply 311 can be logically connected to the processor 310 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system.

[0136] In addition, the electronic device 300 includes some functional modules not shown, which will not be described in detail here.

[0137] Preferably, this application embodiment also provides an electronic device, including a processor 310, a memory 309, and a computer program stored in the memory 309 and executable on the processor 310. When the computer program is executed by the processor 310, it implements the various processes of the above-described image classification method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0138] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described image classification method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0139] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1The steps of the function specified in one or more boxes.

[0143] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0144] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0145] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0146] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0147] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An image classification method, characterized in that, include: Acquire visual image data, shooting time information, and geographical location information of the target base station survey images; A spatiotemporal feature vector is constructed based on the shooting time information and the geographical location information; The visual image data is processed using a multi-scale feature extraction network to obtain a multi-scale image feature representation of the visual image data; The spatiotemporal feature vector and the multi-scale image feature representation are input into a cross-modal fusion network to generate a fused feature representation of the target base station survey image; The classification result of the target base station survey image is determined based on the fused feature representation.

2. The method as described in claim 1, characterized in that, The multi-scale feature extraction network and the cross-modal fusion network are trained based on a balanced training dataset; The balanced training dataset is obtained through the following steps: The base station survey image training dataset is denoised to obtain a denoised training dataset. Calculate the local sparsity of minority class samples in the denoised training dataset, and assign a synthetic weight to each minority class sample based on the local sparsity. The minority class samples are the class samples in the denoised training dataset whose sample number is less than a preset threshold. Each minority class sample is oversampled according to the synthetic weights to generate a new synthetic sample; The new synthetic samples are merged with the denoised training dataset to obtain the balanced training dataset.

3. The method as described in claim 1 or 2, characterized in that, The multi-scale feature extraction network and the cross-modal fusion network are trained based on the base station reconnaissance image training dataset after image enhancement; The enhanced base station reconnaissance image training dataset is obtained through the following steps: Obtain the training dataset of base station survey images; Image enhancement processing is performed on the base station survey image training dataset to obtain the image-enhanced base station survey image training dataset. The image enhancement process includes at least one of basic geometric transformation enhancement, color perturbation enhancement, fuzzy simulation enhancement, and composite sample enhancement.

4. The method as described in claim 1, characterized in that, The method further includes: A spatiotemporal graph structure is constructed based on the shooting time information and the geographical location information; A graph convolutional network is used to perform feature propagation on the spatiotemporal graph structure in order to update the classification probability of the nodes in the spatiotemporal graph structure; The classification results of the target base station survey image are adjusted based on the updated classification probabilities. In this context, the nodes of the spatiotemporal graph structure are the target base station survey images, the edges represent the spatiotemporal relationships between different target base station survey images, and the edge weights are calculated based on the temporal and spatial proximity of the target base station survey images.

5. The method as described in claim 1, characterized in that, The construction of the spatiotemporal feature vector based on the shooting time information and the geographical location information includes: The shooting time information is encoded using sine and cosine position coding to obtain a time code; The geographic location information is encoded using a spherical geographic distance-based encoding method to obtain a geographic code; The spatiotemporal feature vector is obtained by concatenating the time code and the geographic code.

6. The method as described in claim 1, characterized in that, The cross-modal fusion network is a Transformer-based fusion network. The step of inputting the spatiotemporal feature vector and the multi-scale image feature representation into the cross-modal fusion network to generate the fused feature representation of the target base station reconnaissance image includes: The multi-scale image feature representation is mapped into a first query vector, a first key vector, and a first value vector; The spatiotemporal feature vector is mapped to a second query vector, a second key vector, and a second value vector; Using the scaling dot product attention mechanism, the first attention weights of the first query vector and the second key vector are calculated, and the second value vector is weighted and summed according to the first attention weights to obtain the first fusion feature; Calculate the second attention weights of the second query vector and the first key vector, and perform a weighted summation of the first value vector based on the second attention weights to obtain the second fusion feature; The first fusion feature and the second fusion feature are fused together to obtain the fusion feature representation.

7. The method as described in claim 1, characterized in that, The multi-scale feature extraction network is a feature pyramid network. The process of using the multi-scale feature extraction network to process the visual image data to obtain multi-scale image feature representations of the visual image data includes: Feature maps of the visual image data at different resolutions are extracted using a feature pyramid network. The feature maps at different resolutions are fused to obtain the multi-scale image feature representation.

8. The method as described in claim 1, characterized in that, When training the multi-scale feature extraction network and the cross-modal fusion network, a multi-task loss function is used for joint optimization. The multi-task loss function includes at least two of the following: main classification loss, time period prediction loss, geographic location clustering triplet loss, and contrastive learning loss.

9. The method as described in claim 1, characterized in that, The method further includes: Generate an interpretable heatmap of the target base station reconnaissance image; The interpretable heatmap is used to indicate the image regions that the multi-scale feature extraction network and the cross-modal fusion network focus on during classification.

10. An image classification device, characterized in that, It includes an acquisition module, a construction module, a processing module, a fusion module, and a classification module, among which: The acquisition module is used to acquire visual image data, shooting time information, and geographical location information of the target base station survey images; The construction module is used to construct a spatiotemporal feature vector based on the shooting time information and the geographical location information; The processing module is used to process the visual image data using a multi-scale feature extraction network to obtain a multi-scale image feature representation of the visual image data. The fusion module is used to input the spatiotemporal feature vector and the multi-scale image feature representation into the cross-modal fusion network to generate the fused feature representation of the target base station survey image; A classification module is used to determine the classification result of the target base station survey image based on the fused feature representation.

11. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the image classification method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image classification method as described in any one of claims 1 to 9.

13. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the image classification method according to any one of claims 1 to 9.