Wireless private network coverage map generation method, device, equipment, and storage medium

The AI-based method for generating wireless private network signal coverage maps automates network planning, reducing costs and technical barriers for small and medium-sized enterprises.

WO2026101447A1PCT designated stage Publication Date: 2026-05-15CLOUDRAN AI PTE LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CLOUDRAN AI PTE LTD
Filing Date
2024-11-08
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Current wireless network planning for small and medium-sized enterprises is costly and has high technical barriers due to reliance on professional teams and specialized simulation tools.

Method used

A method and device for generating a wireless private network signal coverage map using artificial intelligence, involving data preprocessing, feature extraction, and image generation to automate the process, reducing technical barriers and costs.

Benefits of technology

The method improves efficiency and versatility of wireless private network planning by automating the process, allowing users to input data for automated network planning and design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2024050724_15052026_PF_FP_ABST
    Figure SG2024050724_15052026_PF_FP_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for generating a wireless private network signal coverage map, belonging to the field of computer technology. The embodiment of the present application designs a standardized data preprocessing process for data heterogeneity. By performing different preprocessing operations on the input data according to different modalities and transforming it into standardized data of a preset format, features are extracted from the standardized data of each modality. This allows the extracted features to accurately represent the data of each modality, facilitating subsequent processing and fusion of different modality features within the same framework. The extracted features can be merged according to the target feature classification to obtain a feature vector that meets the model generation requirements. All the above processes are automated, and users only need to input building data or requirement data to automatically implement wireless private network planning. This reduces the threshold for network planning and design, improves the efficiency and versatility of wireless private network planning, and reduces its costs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Wireless Private Network Coverage Map Generation Method, Device, Equipment, and Storage Medium

[0002] Technical Field

[0003] This application relates to the field of computer technology, particularly to a method, device, equipment, and storage medium for generating a wireless private network coverage map.

[0004] Background Technology

[0005] With the widespread deployment of 4G and 5G cellular wireless technologies, their applications are gradually expanding into the core areas of various industries. An increasing number of industries are adopting wireless private networks to support business operations, thereby placing higher demands on network design and planning.

[0006] However, current wireless network planning typically relies on design schemes from professional teams or outputs from specialized simulation tools. For small and medium-sized enterprises, this approach is not only costly but also has relatively high technical barriers.

[0007] Invention Content

[0008] The embodiments of the present application provide a wireless private network signal coverage map generation method, device, equipment, and storage medium, achieving automated execution, reducing technical barriers and costs, and improving efficiency and versatility. The technical solution is as follows:

[0009] On the one hand, a method for generating a wireless private network signal coverage map is provided, which includes:

[0010] Obtaining input data, which includes target building data and / or user requirement data for wireless private network planning, and includes one or more modalities;

[0011] Preprocessing the input data according to the modality' of the input data and the preprocessing method corresponding to the modality, converting the input data into standardized data in the format corresponding to the modality;

[0012] Feature extracting based on the standardized data for each modality to obtain features of the standardized data;

[0013] Merging the extracted features according to the target feature classification to obtain a target feature vector;

[0014] Taking the target feature vector as input into an image generation model, which generates a wireless private network signal coverage map inside the target building based on the target feature vector.

[0015] In some embodiments, the input data includes multiple modalities, which include various types of modalities, one type of modality includes one or more modalities, and one type of modality corresponds to a unified format; According to the modality of the input data and the corresponding preprocessing method, preprocessing the input data according to the corresponding modality, converting the input data into standardized data in the format corresponding to the modality, including:

[0016] For input data of the same type of modality, preprocessing the input data of the multiple modalities through the corresponding preprocessing methods of the multiple modalities, adjusting the input data of the multiple modalities to standardized data in the unified format of the type of modality.

[0017] In some embodiments, the input data includes multiple modalities, which include at least two of the following: photographic image modality; planar image modality, graphic file modality, video modality, LiDAR modality, infrared induction modality, text modality, and voice modality: the input data includes at least two of tire following: building images photographed by users in the photographic image modality, building videos recorded by users in the video modality, floor plans of building interiors in the planar image modality, building sketches drawn by users in the planar image modality, graphic files output by professional drawing software in the graphic file modality, LiDAR data in the LiDAR modality, infrared induction data in the infrared induction modality, text information provided by users in the text modality, and voice information provided by users in the voice modality;

[0018] According to the modality of the input data and the corresponding preprocessing method, preprocessing the input data according to the corresponding modality, converting the input data into standardized data in the format corresponding to the modality, including at least two of the following:

[0019] In response to the input data including the building images photographed by the user, adjusting the resolution of the building images to the target resolution, converting the bitmap of the building images to grayscale to obtain the first image, performing edge detection on the first image to obtain the contour features of tire building structure, and using the first image and the contour features as the standardized data of the building image;

[0020] In response to the input data including the building video recorded by the user in the video modality, decoding the building video to convert it into a frame sequence; extracting key frames from the frame sequence to obtain the taiget key frames; performing resolution adjustment, grayscale processing, and edge detection processing on the target key frames to obtain the standardized data of the target key frames;

[0021] In response to the input data including the floor plan of the building interior and / or the building sketch drawn by the user, performing resolution adjustment and edge detection processing on the floor plan and / or building sketch to obtain the standardized data of the floor plan;

[0022] Tn response to the input data including the graphic file output by the professional drawing software, calling a professional conversion tool to perform format conversion on the graphic file to obtain the second image in the target image format, adjusting the resolution of the second image to the target resolution to obtain the standardized data of the graphic file;

[0023] In response to the input data including the text information input by the user, removing stop words and irrelevant symbols from tire text information, converting abbreviations, synonyms, and variant words in the text information, transforming the processed text information into the target text format to obtain the standardized data of the text information;

[0024] Tn response to the input data including the voice information input by the user, performing noise reduction processing on the voice information, performing speech recognition on the noise- reduced voice information to obtain the first text information corresponding to the voice information, performing preprocessing such as removing stop words and irrelevant symbols, converting abbreviations, synonyms, and variant words, and transforming the text format on the first text information to obtain the standardized data of the voice information;

[0025] In response to the input data including LiDAR data, filtering the LiDAR data, converting the filtered LiDAR data from the device coordinate system to tire global coordinate system to obtain the standardized data of the LiDAR data;

[0026] In response to the input data including infrared induction data, filtering the infrared induction data, performing temperature calibration on the filtered infrared induction data, and performing image enhancement on the calibrated infrared induction data to obtain the standardized data of the infrared induction data.

[0027] In some embodiments, the multiple modalities include at least two types of modalities from the categories of image modality, planar diagram modality, and language modality, where the image modality includes the photographic image modality, graphic file modality, video modality, LiDAR modality, and infrared induction modality; the planar diagram modality includes the planar image modality; and the language modality includes the text modality and voice modality.

[0028] The feature extraction performed on the standardized data of each modality to obtain the features of the standardized data includes: Aligning the standardized data of the same type of modality to obtain the target standardized data for each modality; Executing the feature extraction step on the aligned target standardized data to obtain the features of the target standardized data.

[0029] In some embodiments, the feature extraction performed on the standardized data of each modality to obtain the features of the standardized data includes at least one of the following:

[0030] Tn response to the standardized data or the target standardized data being image modality data, performing at least one of the following on the standardized data or target standardized data:

[0031] Using an image recognition model to perform object recognition on the standardized data or target standardized data, outputting building elements in the image, analyzing the texture of local areas in the image using a classifier, identifying building material information in the image, and obtaining corresponding reflection coefficients and / or absorption coefficients based on the building material information;

[0032] Using a semantic segmentation model to perfonn image segmentation on the standardized data or target standardized data, obtaining visual cues in the image, and obtaining spatial layout and connections between spaces in the image based on the visual cues;

[0033] Using monocular depth estimation technology to extract three-dimensional spatially related information from two-dimensional images; If the standardized data is LiDAR data, using ground segmentation technology to separate ground points from non-ground points in the LiDAR data, clustering point data to obtain objects and scene elements in the LiDAR data, and performing feature extraction on the objects and scene elements to obtain spatial layout features and / or distance features;

[0034] In response to the standardized data or the target standardized data being planar diagram modality data, performing at least one of the following on the standardized data or target standardized data:

[0035] Using optical character recognition technology to extract building material information and / or building space information from the planar diagram, and obtaining corresponding reflection coefficients and absorption coefficients based on the building material information;

[0036] Using a semantic segmentation model to segment the planar diagram to obtain different building areas and building elements, classifying the building areas and target building elements, and extracting building space information from the classification results;

[0037] Identifying taiget areas from the planar diagram, detecting the planar diagram using image recognition technology to obtain base station installation points marked in the planar diagram, with the target area being the area in the planar diagram that requires taigctcd network coverage;

[0038] Recording positional point information and / or reflection point locations in the planar diagram, and calculating at least one of direct path, reflection path, direct path phase angle, and reflection path phase angle based on the positional point information and / or reflection point locations;

[0039] In response to the standardized data or the target standardized data being language modality data, performing at least one of the following on the standardized data or target standardized data:

[0040] Using a named entity recognition model to extract keywords from the text to obtain at least one of building structure information, building material information, and building space information;

[0041] Determining entity' relationships in the text using a relationship extraction model.

[0042] In some embodiments, the method further includes:

[0043] Tn response to the input data including data of the target type, performing embedded encoding on the data of the target type to obtain an embedding vector, where the target type is a requirement feature related to the output target;

[0044] Writing the embedding vector into the input data of the image modality', and performing preprocessing operations on the input data after writing.

[0045] In some embodiments, the taiget feature vector includes multiple dimensions, where the target feature classification indicates the dimensions in which features of different classifications are located in the target feature vector;

[0046] Merging the extracted features according to the target feature classification to obtain the target feature vector includes:

[0047] Determining the dimension in the target feature vector corresponding to each extracted feature according to the target feature classification; Writing each extracted feature into tire corresponding dimension in tire taiget feature vector.

[0048] In some embodiments, the method further includes:

[0049] Using the input data and the target feature vector as training data for the image generation model, and using the wireless private network signal coverage map inside the target building as the annotated data for the training data.

[0050] On the one hand, a wireless private network signal coverage map generation device is provided, which includes:

[0051] An acquisition module, used for obtaining input data, which includes taiget building data and / or user requirement data for wireless private network planning, and the input data includes one or more modalities;

[0052] A preprocessing module, used for preprocessing the input data according to the modality of the input data and the corresponding preprocessing method, converting the input data into standardized data in the format corresponding to the modality;

[0053] An extraction module, used for feature extraction on the standardized data of each modality, to obtain the features of the standardized data;

[0054] A merging module, used for merging the extracted features according to the target feature classification, to obtain a target feature vector;

[0055] A generation module, used for inputting the taiget feature vector into the image generation model, and the image generation model generates a wireless private network signal coverage map inside the target building based on the target feature vector.

[0056] In some embodiments, the input data includes multiple modalities, which include various types of modalities, one type of modality includes one or more modalities, and one type of modality corresponds to a unified format;

[0057] The preprocessing module is used for input data of the same type of modality' with multiple modalities, to preprocess the input data of tire multiple modalities through the corresponding preprocessing methods of the multiple modalities, adjusting the input data of the multiple modalities to standardized data in the unified format of the type of modality.

[0058] Tn some embodiments, the input data includes multiple modalities, which include at least two of the following: photographic image modality, planar image modality, graphic file modality, video modality, LiDAR modality, infrared induction modality, text modality, and voice modality; the input data includes at least two of the following: building images photographed by users in the photographic image modality, building videos recorded by users in the video modality, floor plans of building interiors in the planar image modality, building sketches drawn by users in the planar image modality, graphic files output by professional drawing software in the graphic file modality, LiDAR data in the LiDAR modality, infrared induction data in the infrared induction modality, text information provided by users in the text modality, and voice information provided by users in the voice modality.

[0059] The preprocessing module is used to perform at least two of the following:

[0060] In response to the input data including building images photographed by the user, adjust the resolution of the building images to the target resolution, process tire bitmap of the building images to grayscale to obtain a first image, perform edge detection on the first image to obtain the contour features of the building structure, and use the first image and the contour features as the standardized data of the building image;

[0061] In response to the input data including building videos recorded by the user in the video modality, decode the building videos to convert them into a frame sequence; extract key frames from the frame sequence to obtain the target key frames; perform resolution adjustment, grayscale processing, and edge detection processing on the target key frames to obtain the standardized data of the taiget key frames;

[0062] In response to the input data including floor plans of the building interior and / or building sketches drawn by the user, perform resolution adjustment and edge detection processing on the floor plans and / or building sketches to obtain the standardized data of the floor plans;

[0063] In response to the input data including graphic files output by professional drawing software, call a professional conversion tool to perform fonnat conversion on the graphic files to obtain a second image in the target image format, adjust the resolution of the second image to the target resolution, and obtain the standardized data of the graphic files;

[0064] In response to the input data including text information input by the user, remove stop words and irrelevant symbols from the text information, convert abbreviations, synonyms, and variant words in the text information, transform the processed text information into the target text fonnat, and obtain the standardized data of the text information;

[0065] In response to the input data including voice information input by the user, perform noise reduction processing on the voice information, perform speech recognition on the noise-reduced voice information to obtain the first text information corresponding to the voice information, perform preprocessing such as removing stop words and irrelevant symbols, converting abbreviations, synonyms, and variant words, and transforming the text fonnat on the first text information to obtain the standardized data of the voice information;

[0066] In response to the input data including LiDAR data, filter the LiDAR data, convert the filtered LiDAR data from the device coordinate system to the global coordinate system, and obtain the standardized data of the LiDAR data;

[0067] In response to the input data including infrared induction data, filter the infrared induction data, perform temperature calibration on the filtered infrared induction data, and enhance the image of the calibrated infrared induction data to obtain the standardized data of the infrared induction data.

[0068] In some embodiments, the multiple modalities include at least two types of modalities from the categories of image modality, planar diagram modality, and language modality, where the image modality includes the photographic image modality, graphic file modality, video modality, LiDAR modality, and infrared induction modality; the planar diagram modality includes the planar image modality; and the language modality includes the text modality and voice modality.

[0069] The extraction module is used for: Aligning the standardized data of the same type of modality to obtain the target standardized data for each modality;

[0070] Executing the feature extraction step on the aligned target standardized data to obtain the features of the target standardized data.

[0071] In some embodiments, the extraction module is used to perform at least one of the following:

[0072] In response to the standardized data or the taigct standardized data being image modality data, performing at least one of the following on the standardized data or target standardized data:

[0073] Using an image recognition model to perform object recognition on the standardized data or target standardized data, outputting building elements in the image, analyzing the texture of local areas in the image using a classifier, identifying building material information in the image, and obtaining corresponding reflection coefficients and / or absorption coefficients based on the building material information;

[0074] Using a semantic segmentation model to perform image segmentation on the standardized data or target standardized data, obtaining visual cues in the image, and obtaining spatial layout and connections between spaces in the image based on the visual cues;

[0075] Using monocular depth estimation technology to extract three-dimensional spatially related information from two-dimensional images;

[0076] If the standardized data is LiDAR data, using ground segmentation technology to separate ground points from non-ground points in the LiDAR data, clustering point data to obtain objects and scene elements in the LiDAR data, and performing feature extraction on the objects and scene elements to obtain spatial layout features and / or distance features;

[0077] In response to the standardized data or the target standardized data being planar diagram modality data, performing at least one of the following on the standardized data or taiget standardized data:

[0078] Using optical character recognition technology to extract building material information and / or building space information from the planar diagram, and obtaining corresponding reflection coefficients and absorption coefficients based on the building material information;

[0079] Using a semantic segmentation model to segment the planar diagram to obtain different building areas and building elements, classifying the building areas and target building elements, and extracting building space information from the classification results;

[0080] Identifying taiget areas from the planar diagram, detecting the planar diagram using image recognition technology to obtain base station installation points marked in the planar diagram, with the target area being the area in the planar diagram that requires tai eted nelwork coverage;

[0081] Recording positional point information and / or reflection point locations in the planar diagram, and calculating at least one of direct path, reflection path, direct path phase angle, and reflection path phase angle based on the positional point information and / or reflection point locations;

[0082] In response to the standardized data or the target standardized data being language modality data, performing at least one of the following on tire standardized data or target standardized data:

[0083] Using a named entity recognition model to extract keywords from the text to obtain at least one of building structure information, building material information, and building space information:

[0084] Determining entity' relationships in the text using a relationship extraction model.

[0085] In some embodiments, the method further includes:

[0086] Tn response to the input data including data of the target type, perfonning embedded encoding on the data of the target type to obtain an embedding vector, where the target type is a requirement feature related to the output target;

[0087] Writing the embedding vector into the input data of the image modality’, and performing preprocessing operations on the input data after writing.

[0088] In some embodiments, the target feature vector includes multiple dimensions, where the target feature classification indicates the dimensions in which features of different classifications are located in the target feature vector;

[0089] Merging the extracted features according to the target feature classification to obtain the target feature vector includes:

[0090] Determining the dimension in the target feature vector corresponding to each extracted feature according to the target feature classification;

[0091] Writing each extracted feature into the corresponding dimension in the target feature vector.

[0092] In some embodiments, the generation module is also used to use the input data and the target feature vector as training data for the image generation model, and using the wireless private network signal coverage map inside the target building as the annotated data for the training data.

[0093] On the one hand, an electronic device is provided, which includes one or more processors and one or more memories, wherein the one or more memories store at least one computer program, and the at least one computer program is loaded and executed by the one or more processors to implement various optional embodiments of the aforementioned wireless private network signal coverage map generation method.

[0094] On the one hand, a computer-readable storage medium is provided, which stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement various optional embodiments of the aforementioned wireless private network signal coverage map generation method.

[0095] On the one hand, a computer program product or computer program is provided, which includes one or more program codes, and the one or more program codes are stored on a computer- readable storage medium. One or more processors of an electronic device read the one or more program codes from the computer-readable storage medium, and the one or more processors execute the one or more program codes, enabling the electronic device to perform any of the aforementioned possible implementations of the wireless private network signal coverage map generation method.

[0096] In the embodiments of the present application, considering the excellent advantages of artificial intelligence technology' in performing generation tasks based on a large amount of data, and that different modalities of data have different forms of expression, data types, and feature distributions, a data standardization preprocessing process is designed for data heterogeneity before applying the image generation model. By performing different preprocessing operations on the input data according to different modalities and transforming it into standardized data in a preset format, and then performing feature extraction on the standardized data of each modality', the extracted features can accurately represent the data of each modality, which is conducive to subsequent processing and fusion of features of different modalities within the same framework. It is precisely because of the aforementioned standardization process that the extracted features can be merged according to the target feature classification to obtain a feature vector that meets the model generation requirements, thereby generating a wireless private network signal coverage map. The above data processing, feature extraction, and image generation processes are all automated procedures. U sers only need to input building data or requirement data to automatically' implement wireless private network planning, which greatly reduces the threshold when planning and designing the network, improves the efficiency and versatility of wireless private network planning, and reduces its cost.

[0097] Figure Description

[0098] To clearly illustrate the technical solution of the present application embodiment, a brief introduction will be given to the drawings that need to be used in the description of the embodiment. It is evident that the drawings described below are merely some embodiments of the present application, and for those skilled in the art, it is possible to obtain other drawings without the need for creative effort based on these drawings.

[0099] Figure 1 is a schematic diagram of the implementation environment of a wireless private network signal coverage map generation method provided in an embodiment of the present application;

[0100] Figure 2 is a flowchart of a wireless private network signal coverage map generation method provided in an embodiment of the present application;

[0101] Figure 3 is a structural schematic diagram of a wireless private network signal coverage map generation device provided in an embodiment of the present application;

[0102] Figure 5 is a structural block diagram of an electronic device provided in an embodiment of the present application;

[0103] Figure 5 is a structural block diagram of a terminal provided in an embodiment of the present application (Note: There seems to be a repetition in numbering, it should be Figure 6);

[0104] Figure 6 is a structural schematic diagram of a server provided in an embodiment of the present application.

[0105] Detailed method of implementation To more clearly convey the objectives, technical solutions, and advantages of this application, further detailed descriptions of the implementation of this application will be provided in conjunction with the accompanying drawings.

[0106] Tn this application, terms such as "first," "second," etc., are used to distinguish between items or similar items that have essentially the same functions and roles. It should be understood that there is no logical or temporal dependency between "first," "second," "nth," nor do they limit the quantity or execution order. It should also be understood that even though the following description uses terms like "first," "second," etc., to describe various elements, these elements should not be restricted by these terms. These terms are merely used to differentiate one element from another. For example, without departing from the scope of the various examples given, the "first image" could be referred to as tire "second image," and similarly, tire "second image" could be referred to as the "first image." Both the first and second images are images and, in certain contexts, are separate and distinct images.

[0107] The tenn "at least one" in this application means one or more, and the term "multiple" means two or more, for example, "multiple data packets" refers to two or more data packets.

[0108] It should be understood that the terminology used in the description of various examples in this document is intended merely to describe specific examples and is not intended to be limiting. As used in the description of various examples and the appended claims, the singular forms "a" or "the" are intended to include the plural forms unless the context clearly dictates otherwise.

[0109] It should also be understood that the term "and / or" as used in this document means and covers any and all possible combinations of the listed items. The term "and / or" is a tenn of art used to describe a relationship between associated objects, indicating the existence of three relationships, for example, A and / or B, which means: A alone, both A and B together, or B alone. Additionally, the character " / " in this application generally indicates an "or" relationship between the associated objects before and after it.

[0110] It should also be understood that in the various embodiments of this application, the numbering of the processes does not imply the order of execution. The order of execution of the processes should be determined by their function and inherent logic, and should not limit the implementation process of the embodiments of this application.

[0111] It should also be understood that determining B based on A docs not mean that B is determined solely based on A, but also based on A and / or other information.

[0112] It should also be understood that the term "comprising" (also referred to as "includes," "including," "comprises," and / or "comprising") when used in this specification indicates the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groupings.

[0113] It should also be understood that the term "if can be interpreted to mean "when" ("when" or "upon") or "in response to determining" or "in response to detecting." Similarly, depending on the context, the phrases "if it is determined that..." or "if [the stated condition or event] is detected" can be interpreted to mean "when... is determined" or "in response to determining..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]."

[0114] The following reference drawings are descnbed to assist in the comprehensive understanding of the various embodiments of this application as defined by the claims and their equivalents. This description includes various specific details to aid in understanding but should be considered exemplary only. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of this application. Additionally, for clarity and brevity, the description of well-known functions and structures may be omitted.

[0115] The terminology and wording used in the following description and claims are not limited to their dictionary meanings but are used by the inventor solely to enable a clear and consistent understanding of this application. Therefore, it should be apparent to those skilled in the art that the descriptions provided below of the various embodiments of this application are for illustrative purposes only and are not intended to limit the scope of this application as defined by the appended claims and their equivalents.

[0116] It should be understood that the singular forms "one," "a," and "the" may also include plural references unless the context clearly indicates otherwise. For example, the term "component surface" includes one or more such surfaces. When we refer to an element being "connected" or "coupled" to another element, it can mean that one element is directly connected or coupled to another element or that one element is connected or coupled to another element through an intermediary element. Additionally, the tenns "connect" or "couple" used here can include wireless connections or wireless coupling.

[0117] The term "comprising" or "can comprise" indicates the presence of the corresponding disclosed functions, operations, or components in the various embodiments of this application, without limiting tire presence of one or more additional functions, operations, or features. Furthermore, the term "comprising" or "having" can be interpreted to represent certain characteristics, numbers, steps, operations, components, or their combinations, but should not be interpreted as excluding the possibility of the existence of one or more other characteristics, numbers, steps, operations, components, or their combinations.

[0118] The term "or" as used in the various embodiments of this application includes any listed term and all combinations thereof. For example, "A or B" can include A, can include B, or can include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not explicitly defined, it can refer to one, multiple, or all of the multiple items, for example, the description "parameter A includes Al, A2, A3" can be implemented as parameter A including Al or A2 or A3, and can also be implemented as parameter A including at least two of the three items Al, A2, A3.

[0119] Unless differently defined, all terms (including technical or scientific terms) used in this application have the same meaning understood by those skilled in the art in the context of this application. Common terms defined in dictionaries are interpreted to have meanings consistent with the context in the relevant technical field and should not be idealized or overly formalized unless explicitly defined as such in this application.

[0120] At least some of the functions of the devices or electronic devices provided in the embodiments of this application can be implemented through AT models, such as at least one module of the device or electronic device being implemented through Al models. Functions associated with Al can be executed through non-volatile memory, volatile memory, and processors.

[0121] The processor may include one or more processors. At this time, the one or more processors can be general processors, such as a Central Processing Unit (CPU), an Application Processor (AP), etc., or purely graphic processing units, such as a Graphics Processing Unit (GPU), a Vision Processing Unit (VPU), and / or Al-specific processors, such as a Neural Processing Unit (NPU).

[0122] Tire one or more processors control the processing of input data based on predefined operational rules or artificial intelligence (Al) models stored in non-volatile memory and volatile memory. Predefined operational rules or Al models are provided through training or learning.

[0123] Here, providing through learning means obtaining predefined operational rules or Al models with desired characteristics by applying learning algorithms to multiple learning data. The learning can be executed in the Al device or electronic device itself according to the embodiment, and / or can be implemented through a separate server / system.

[0124] Al models can include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations through calculations between the input data (such as the results of the previous layer's calculations and / or the input data of the Al model) and the multiple weight values of the current layer. Examples of neural networks include but are not limited to Convolutional Neural Networks (CNN), Deep Neural Networks (DNN), Recurrent Neural Networks (RNN), Restricted Boltzmann Machines (RBM), Deep Belief Networks (DBN), Bidirectional Recurrent Deep Neural Networks (BRDNN), Generative Adversarial Networks (GAN), and Deep Q Networks.

[0125] Learning algorithms are methods that use multiple learning data to train a predetermined target device (e.g., a robot) to enable, allow, or control the target device to make determinations or predictions. Examples of learning algorithms include but are not limited to supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0126] According to this application, at least one step in the methods performed by electronic devices, such as recognizing architectural elements, classifying planar map areas, and other steps, can be implemented using artificial intelligence models. The processor of the electronic device can perform pre-processing operations on data to convert it into a form suitable for input to an artificial intelligence model . Artificial intelligence models can be obtained through training . Here, "obtained through training" means obtaining predefined operational rules or artificial intelligence models configured to perform desired features (or purposes) by training an underlying artificial intelligence model with multiple training data.

[0127] Below is an explanation of the terminology involved in this application. Artificial Intelligence (Al) involves the use of digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, AT is an interdisciplinary technology in computer science that attempts to understand the essence of intelligence and to produce a new type of intelligent machine that can respond in a manner similar to human intelligence . Al is also the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the capabilities of perception, reasoning, and decision-making.

[0128] Al technology is an interdisciplinary field, involving a wide range of areas, including both hardware and software technologies. Basic Al technologies generally include sensors, dedicated Al chips, cloud computing, distributed storage, big data processing technologies, operating / interaction systems, mechatronics, and more. Al software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine leaming / deep learning.

[0129] Computer Vision (CV) is the science of making machines "see." More specifically, it involves using cameras and computers to replace the human eye for machine vision tasks such as identification, tracking, and measurement, and further for image processing to make the computer processing more suitable for the human eye to observe or transmit to instruments for detection. As a scientific discipline, computer vision studies the relevant theories and technologies, attempting to establish an Al system capable of extracting information from images or multidimensional data. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality; simultaneous localization and mapping, and more. It also includes common biometric recognition technologies such as facial and fingerprint recognition.

[0130] Key technologies in Speech Technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel represents the future direction of human-computer interaction, with speech being one of the most promising modes of human-computer interaction in the future.

[0131] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, the language people use daily, and thus is closely related to linguistic research. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robot Q&A, knowledge graphs, and more.

[0132] Machine Learning (ML) is an interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory; and more. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and is the fundamental way to make computers intelligent, with its applications throughout the field of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, Bayesian networks, reinforcement learning, transfer learning, inductive learning, and fonnula teaching.

[0133] With the research and advancement of Al technology, Al has been studied and applied in various fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, unmanned driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, Al will be applied in more fields and play an increasingly important value.

[0134] The solution provided by the embodiments of this application involves technologies such as Al's computer vision, speech technology, natural language processing, and machine learning, which are specifically illustrated through the following examples.

[0135] With the widespread deployment of 4G and 5G cellular wireless technologies, their applications are gradually expanding into the core areas of various industries. More and more industries are introducing wireless private networks to support business operations, thereby imposing higher demands on network design and planning. However, current wireless network planning typically relies on design schemes from professional teams or outputs from professional simulation tools. For small and medium -sized enterprises, this approach is not only costly but also has a relatively high technical barrier.

[0136] The rapid development of artificial intelligence technology offers new possibilities for solving these problems, hi particular, large Al models can better perform multimodal generation tasks and demonstrate strong data processing and analytical capabilities. By training these models on a vast amount of target scenario data, it is possible to extract key scenario information during the inference phase, thereby predicting the coverage and capacity of wireless networks. This method is expected to significantly reduce the barriers to obtaining high-performance wireless private network planning and design.

[0137] In the embodiment of the present application, the data processing phase is carefully designed, for example, different preprocessing operations are performed for multimodal data during the data processing phase to ensure data quality. Through the method provided by the embodiment of the present application, after the standardization treatment of the data, it is possible to establish a unified data processing and analysis model for the input data, despite the multimodal input data having different forms of expression, data types, and feature distributions. This results in high- quality feature vectors, thereby improving the performance of the model and the input efficiency, reliability, and accuracy of the output results. Below is an explanation of the implementation environment of the present application.

[0138] Figure 1 is a schematic diagram of the implementation environment of a wireless private network signal coverage map generation method provided in an embodiment of the present application. This implementation environment includes terminal 101 , or the implementation environment includes terminal 101 and wireless private network signal coverage map generation platform 102. Terminal 101 is connected to the wireless private network signal coverage map generation platform 102 via a wireless or wired network.

[0139] Terminal 101 is at least one of a smartphone, gaming console, desktop computer, tablet computer, e-book reader, MP3 (Moving Picture Experts Group Audio Layer III, dynamic image expert compression standard audio level 3) player, or MP4 (Moving Picture Experts Group Audio Layer IV, dynamic image expert compression standard audio level 4) player, laptop computer. Terminal 101 installs and runs an application that supports the generation of wireless private network signal coverage maps.

[0140] For example, the terminal 101 has data collection and processing functions, collects userinput data, preprocesses and extracts features from the collected data, and then generates a wireless private network signal coverage map inside the building based on an image generation model, achieving wireless private network planning and design to provide feedback to the user. Terminal 101 can independently complete this work, or terminal 101 can collect input data and transmit it to the wireless private network signal coverage map generation platform 102 to provide data processing services. The embodiment of the present application does not limit this.

[0141] The wireless private network signal coverage map generation platform 102 includes at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. The wireless private network signal coverage map generation platform 102 is used to provide backend services for applications that support the generation of wireless private network signal coverage maps.

[0142] Optionally, the wireless private network signal coverage map generation platform 102 undertakes the main processing work, and terminal 101 undertakes the secondary processing work; or, the wireless private network signal coverage map generation platform 102 undertakes the secondary processing work, and terminal 101 undertakes the main processing work; or, the wireless private network signal coverage map generation platform 102 or terminal 101 undertakes the processing work separately. Alternatively, the wireless private network signal coverage map generation platform 102 and terminal 101 adopt a distributed computing architecture for collaborative computing.

[0143] Optionally, the wireless private network signal coverage map generation platform 102 includes at least one server 1021 and a database 1022, which is used to store data. In this embodiment of the application, the database 1022 stores sample input data and sample annotated data, and can also store user input data to provide data sendees for at least one server 1021.

[0144] A server is a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services. CDN, and big data and artificial intelligence platforms. Terminals are smartphones, tablet computers, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these.

[0145] Those skilled in the art know that the number of terminals 101 and servers 1021 mentioned above can be more or less. For example, the terminals 101 and servers 1021 mentioned above can be just one, or they can be dozens or hundreds, or more, and the embodiment of the present application does not limit the number and type of terminals or servers.

[0146] Figure 2 is a flowchart of a wireless private network signal coverage map generation method provided in an embodiment of tire present application, which is applied to an electronic device, the electronic device being a terminal or a server. See Figure 2, the method includes the following steps.

[0147] 201. The electronic device obtains input data, which includes target building data and / or user requirement data for wireless private network planning, and the input data includes one or more modalities.

[0148] In the embodiment of the present application, a user can input data on the electronic device to enable the electronic device to perform wireless private network planning for the user based on the input data, and output a wireless private network signal coverage map inside a building. The input data can include data containing building information, user requirement-related data, or a combination of both types of data.

[0149] The user can be an individual or a corporate user. For example, some small and medium-sized enterprises that need to deploy a wireless private network can input their target building data and / or user requirement data that they need to deploy through the electronic device.

[0150] If the electronic device is a terminal, the user can input data on this electronic device, and the electronic device can collect the user's input data. If the electronic device is a server, the user can input data on the termrnal, the terminal can send the collected input data to the server, and the server receives the input data, that is, it obtains the input data. The embodiment of the present application does not limit the specific type of this electronic device.

[0151] In some embedments, the input data can be unimodal or multimodal input containing building information and / or user requirements obtained through a specific method (that is, the method of obtaining the input data). The input data includes multiple modalities, which can include one or more of photographic image modality, planar image modality, graphic file modality, video modality, LiDAR modality, infrared induction modality, text modality, and voice modality.

[0152] Tn a specific possible embodiment, the input data can be any one or a combination of user- photographed building images in the photographic image modality, user-recorded building videos in the video modality, floor plans of building interiors in the planar image modality, building sketches drawn by users in the planar image modality, graphic files output by professional drawing software in the graphic file modality, LiDAR data in the LiDAR modality, infrared induction data in the infrared induction modality, text information provided by users in the text modality, and voice information provided by users in the voice modality'. That is, the input data can include but is not limited to the aforementioned types of modal data or any combination thereof.

[0153] It should be noted that the aforementioned input data is just some types of data set by' relevant technical personnel according to requirements. In actual applications, relevant technical personnel can also set or add other types of input data according to needs or experience. In other embodiments, the input data can also be other types of information input containing building information and / or user requirements obtained from users. The embodiment of the present application does not limit this, that is, the input data can include but is not limited to the aforementioned types of data or combinations of several ty pes of data.

[0154] In some embodiments, the input data includes multiple modalities, and accordingly, the input data includes at least two of the following: user-photographed building images in the photographic image modality, user-recorded building videos in the video modality', floor plans of building interiors in the planar image modality, building sketches drawn by users in the planar image modality, graphic files output by professional drawing software in the graphic file modality', LiDAR data in the LiDAR modality', infrared induction data in the infrared induction modality', text information provided by' users in the text modality, and voice information provided by' users in the voice modality

[0155] Regarding the specific method (that is, the method of obtaining the input data), it includes but is not limited to the following methods or any combination thereof:

[0156] Method one: Obtained through user uploads, user online drawing, etc., on a website page.

[0157] Method two: Obtained through user uploads, user creation, etc., in mobile phone, computer, tablet (Pad) applications.

[0158] Method three: Obtained through outputs of other building, drawing professional software, etc.

[0159] Of course, the method of obtaining the input data can also include other methods, such as users entering a website address, and the electronic device accessing the website address to download, etc. The embodiment of the present application does not limit this.

[0160] 202. The electronic device performs preprocessing corresponding to the modality of the input data and the preprocessing method corresponding to that modality, converting the input data into standardized data in the format corresponding to that modality.

[0161] Considering that data of different modalities have different forms of expression, data types, and feature distributions, after the electronic device obtains the input data, it can first preprocess the input data to convert it into standardized data, adjusting it to a preset unified format to ensure the standardization of the processed data, facilitating subsequent processing and integration of features of different modalities within a unified framework.

[0162] Also, due to the different forms of expression, data types, and feature distributions of data from different modalities, the electronic device employs different preprocessing methods for input data of different modalities. In some embodiments, a correspondence between modalities and preprocessing methods can be pre-established, and the electronic device can identify the modality of the input data, determine the preprocessing method for that input data according to this correspondence, and then perform the preprocessing.

[0163] In some embodiments, a configuration file storing the correspondence between modalities and preprocessing methods can be pre-stored in the electronic device. When input data is obtained and its modality is recognized, this configuration file can be called or read, the preprocessing method corresponding to the modality can be matched from this configuration file, and then the execution instructions of this preprocessing method can be executed to perform the corresponding preprocessing on the input data, converting it into standardized data.

[0164] In a specific possible embodiment, different preprocessing instructions corresponding to different modalities can be pre-stored in the electronic device. The electronic device can recognize the modality of the input data, determine the corresponding preprocessing instructions based on the recognized modality, and execute these preprocessing instructions on tire input data. For different modalities in the input data, the electronic device performs different preprocessing operations on them, that is, executes different preprocessing instructions.

[0165] For example, modalities can include photographic image modality, planar image modality, graphic file modality, video modality, LiD AR modality, infrared induction modality, text modality, voice modality etc., and relevant technical personnel can freely add more according to needs; the embodiment of the present application does not limit this.

[0166] In some embodiments, the input data includes multiple modalities, including at least two of photographic image modality, planar image modality, graphic file modality, video modality, LiDAR modality; infrared induction modality, text modality, and voice modality'. The input data includes at least two of user-photographed building images in the photographic image modality; user-recorded building videos in the video modality, floor plans of building interiors in the planar image modality, building sketches drawn by users in the planar image modality; graphic files output by professional drawing software in the graphic file modality; LiDAR data in the LiDAR modality, infrared induction data in the infrared induction modality; text information provided by users in the text modality; and voice information provided by users in the voice modality.

[0167] Accordingly, step 202 may include any combination of at least two of the following situations. Of course, these multiple modalities are not limited to the aforementioned few, and the input data is not limited to the aforementioned few. Correspondingly; this step may' also include other situations or any combination of other situations with the following situations.

[0168] Case 1 : In response to the input data including building images photographed by' the user, the electronic device adjusts the resolution of the building images to the taiget resolution, processes the bitmap of the building images to grayscale to obtain a first image, performs edge detection on the first image to obtain the contour features of the building structure, and uses the first image and the contour features as the standardized data for the building image.

[0169] In this context, the target resolution is a preset unified resolution, which can be set by' relevant technical personnel based on needs or experience. The specific value of the target resolution is not limited in the embodiment of the present application. Adjusting the image resolution ensures the consistency of subsequent processing steps and the quality of the output data.

[0170] The grayscale processing is performed using image processing techniques, which can reduce memory usage, increase processing speed, and enhance visual contrast by reducing color information to highlight the target area. In some embodiments, the process of grayscale processing of the building image bitmap can be implemented in various ways, such as maximum value method, average value method, weighted average method, etc., without limitation in the embodiment of the present application.

[0171] The edge detection process uses image edge detection technology to extract the contour features of the building structure in the image, facilitating the visibility of key building information in the image.

[0172] Case 2: In response to the input data including building videos recorded by the user in the video modality, the electronic device decodes the building video, converts the building video into a frame sequence, and extracts keyframes from the frame sequence to obtain the target keyframes. The electronic device performs resolution adjustment, grayscale processing, and edge detection on the target keyframes to obtain the standardized data of the target keyframes.

[0173] In this case, the resolution adjustment, grayscale processing, and edge detection are the same as in Case 1 and will not be further elaborated.

[0174] In Case 2, extracting the target keyframes from the frame sequence is a keyframe extraction process, which can be implemented in various ways. For example, a method based on shot segmentation that selects the first or last frame of each shot as a keyframe by analyzing the video's shot boundaries. Or a method based on color features that calculates the color differences between different frames in the frame sequence to identify keyframes. Or a method based on deep learning that applies a neural network model to extract key points and local features from video frames and obtains keyframes by comparing feature changes between consecutive frames. Or a method based on changes in the viewing angle that extracts keyframes by detecting changes in the viewing angle of the scene in the video. Of course, the keyframe extraction process can also use other methods.

[0175] A specific possible implementation method is provided below', without limitation to its specific implementation in the embodiment of the present application. In a specific possible embodiment, for the first and second frames in the frame sequence, the electronic device generates a color histogram for the first and second frames, respectively, calculates the histogram similarity between the color histograms of the first and second frames, and in response to the histogram similarity being lower than a first threshold, extracts the second frame as a candidate keyframe. The first and second frames are any adjacent frames in the frame sequence, with the first frame preceding the second frame. For all candidate keyframes, similarity matching is performed, and one of the two candidate keyframes with similarity higher than a second threshold is removed to obtain the target keyframe. Then, resolution adjustment, grayscale processing, and edge detection can be performed on each target keyframe. That is, a corresponding color histogram is generated for each frame image, and the histogram similanty between the color histograms of two consecutive frames is calculated; when the histogram similarity is below a certain threshold, a scene change is judged to have occurred, and the frame is extracted as a candidate keyframe; the first and second thresholds can both be set by relevant technical personnel based on needs or experience, without limitation in the embodiment of the present application.

[0176] Case 3: Tn response to the input data including floor plans of the building interior and / or building sketches drawn by the user, the electronic device performs resolution adjustment and edge detection processing on the floor plan and / or building sketch to obtain the standardized data of the floor plan.

[0177] In Case 3, tire floor plans of the building interior and / or building sketches drawn by tire user are generally black and white floor plans, so there is no need for grayscale processing, only resolution adjustment and edge detection processing are needed to process them into the standardized data format unified in the previous two cases. The resolution adjustment and edge detection processing are the same as in Case 1 and will not be further elaborated.

[0178] Case 4: In response to the input data including graphic files output by professional drawing software, the electronic device calls a professional conversion tool to convert the graphic file format, obtaining a second image in the target image format. The electronic device adjusts the resolution of the second image to the target resolution, obtaining the standardized data of the graphic file.

[0179] Since graphic files are output by professional drawing software and their fonnats are not easily processable, they can be converted to more easily manageable image formats, i.e., target image formats. These target image formats can be set by relevant technical personnel based on experience, without limitation in the embodiment of the present application. The taiget image fonnat can be a single fonnat or multiple fonnats, set according to requirements.

[0180] Case 5: In response to the input data including text information input by the user, the electronic device removes stop words and irrelevant symbols from the text information. The electronic device also converts abbreviations, synonyms, and variant words in the text information, transforming the processed text information into the target text format, obtaining the standardized data of the text information.

[0181] In this case, stop words and irrelevant symbols in the text information are not helpful for semantic analysis and are considered redundant information. Therefore, they can be removed first. Abbreviations, synonyms, and variant words in the text information can be converted to a standard expression, i.e., the target text format, to facilitate the uniformity of subsequent data processing.

[0182] The target text fonnat can be set by relevant technical personnel based on needs or experience, without limitation in the embodiment of the present application.

[0183] Case 6: In response to the input data including voice information input by the user, the electronic device performs noise reduction processing on the voice information. The electronic device performs speech recognition on the noise-reduced voice information to obtain the first text information corresponding to the voice information. Hie electronic device performs preprocessing on the first text information, including removing stop words, irrelevant symbols, converting abbreviations, synonyms, and variant words, and transforming the text format, obtaining the standardized data of the voice information.

[0184] Noise reduction processing can apply audio processing techniques to remove background noise from voice information, improving voice clarity. This noise reduction process can be implemented in various ways, for example, by using amplifiers, filters, etc., to process voice information for optimized signal transmission and processing. It can also involve automatic gain control (AGC), high-pass filters, etc., to remove low-frequency noise and enhance signal quality. Advanced noise reduction algorithms can also be employed, such as adaptive filters (e.g., LMS (Least Mean Squares), NLMS (Normalized Least Mean Square) algorithms), spectral subtraction in the frequency domain, or deep learning algorithms (e.g., noise reduction based on neural networks). Of course, the noise reduction process can also be implemented in other ways, without limitation in the embodiment of the present application.

[0185] Speech recognition technology, also known as automatic speech recognition (ASR), aims to convert the vocabulary content in human speech into computer-readable input, such as buttons, binaiy codes, or character sequences. Unlike speaker recognition and speaker verification, which attempt to identify or confirm the speaker rather than the vocabulary content, speech recognition focuses on the words contained within the speech. After converting voice information to text information, preprocessing operations similar to those in Case 6 can be performed.

[0186] Case 7: In response to the input data including LiDAR data, the electronic device perfonns filtering on the LiDAR data. The electronic device converts the filtered LiDAR data from the device coordinate system to the global coordinate system, obtaining the standardized data of the LiDAR data.

[0187] In this case, filtering algorithms are used to remove invalid or noisy points contained in the LiDAR data to improve the quality and clarity of the LiDAR data. Converting the LiDAR data to the global coordinate system places all data under the same coordinate sy stem, making the data more closely related and accurate. The filtering algorithm can be any algorithm, without limitation in the embodiment of the present application.

[0188] Case 8: In response to the input data including infrared induction data, the electronic device perfonns filtering on the infrared induction data, perfonns temperature calibration on the filtered infrared induction data, and enhances the image of the calibrated infrared induction data to obtain the standardized data of the infrared induction data.

[0189] In this case, the filtering process, like Case 7, removes noisy points from the infrared induction data, which will not be further elaborated

[0190] Temperature calibration of the infrared induction data ensures the accuracy of the data. This temperature calibration process can be implemented in various ways. Of course, this temperature calibration process can also be implemented in other ways, without limitation in the embodiment of the present application. The image enhancement process involves adding some information to the original image or transforming data to selectively highlight features of interest in the image or suppress (mask) certain undesired features to match the image with visual response characteristics.

[0191] Tills image enhancement process can be implemented using frequency domain methods or spatial domain methods. The frequency domain method treats the image as a two-dimensional signal and enhances it based on two-dimensional Fourier transformation. Low-pass filtering (allowing only low-frequency signals to pass) can remove noise from the image, while high-pass filtering can enhance high-frequency signals such as edges, making blurry images clear. Representative algorithms in the spatial domain include local averaging and median filtering (taking the middle pixel value in a local neighborhood), which can be used to remove or reduce noise. This image enhancement process can employ any image enhancement method, without limitation in the embodiment of the present application.

[0192] In some embodiments, the multiple modalities include various types of modalities, with one type of modality including one or more modalities, and one ty pe of modality corresponding to a unified format. Accordingly, in step 202, for input data of multiple modalities of the same type, the electronic device can prcproccss the input data of the multiple modalities through the corresponding preprocessing methods, adjusting the input data of the multiple modalities to standardized data in the unified format of that modality type.

[0193] For example, in a specific example combining Case 1 and Case 2, the input data includes building images photographed by the user and building videos recorded by the user. The building images photographed by the user are in the photographic image modality, and the building videos recorded by the user are in the video modality. The electronic device can preprocess the building images and videos separately, following the preprocessing methods shown in Case 1 for images and Case 2 for videos. In both cases, the building images undergo resolution adjustment, grayscale processing, and edge detection, and after processing the building videos to obtain frame sequences and extracting target keyframes, the same preprocessing operations as for the building images are performed on the target keyframes. Thus, the input data of both modalities will ultimately be adjusted to standardized data in a unified format for image modalities.

[0194] 203. The electronic device performs feature extraction on the standardized data for each modality, obtaining the features of the standardized data.

[0195] In the aforementioned step 202, the input data for each modality is processed into a unified format through standardization, which allows for the effective meiging of multimodal features after extraction, enabling data processing within the same framework.

[0196] Tn some embodiments, the multiple modalities include at least two types of modalities from the categories of image modality, planar diagram modality, and language modality. The image modality includes photographic image modality, graphic file modality, video modality, LiDAR modality, and infrared induction modality. The planar diagram modality includes planar image modality, and the language modality includes text modality and voice modality. That is, the input data includes data from various types of modalities.

[0197] Accordingly, in the aforementioned step 202, the electronic device can align the standardized data of the same modality’ type to obtain the target standardized data for each modality. Then, in the aforementioned step 203, feature extraction can be performed on the aligned target standardized data to obtain the features of the target standardized data. Subsequent steps, such as step 204, can then merge the features of the target standardized data and proceed with the subsequent image generation process.

[0198] Data that has undergone standardization, if multimodal, can be aligned to ensure consistency in the features extracted from different modalities. Modality alignment refers to the process of merging features from different modalities during the feature extraction phase. This approach is suitable for modalities with relatively small differences in feature scale and property, facilitating the model's ability to learn the relationships between modalities. Through modality alignment, different modalities can be categorized into the following three typical types of modalities, which are divided into the three modality types mentioned below.

[0199] Image Modality: User-captured building images, keyframes extracted from user-recorded building videos, processed infrared induction images, LiDAR data, graphic files output by professional drawing software, etc. For image modalities, after the initial preprocessing step, the input data is converted into images or image sets with a unified format and resolution, thus allowing for direct modality alignment.

[0200] That is, in response to the standardized data being image modality data, the electronic device can perform modality alignment on the standardized data of the image modality to obtain the target standardized data of the image modality.

[0201] Planar Diagram Modality: Floor plans of building interiors, building sketches drawn by users, etc. For planar diagram modalities, since there are significant differences between user-drawn sketches and standard architectural floor plans in tenns of precision, standardization, and level of detail, it is necessary to input the sketches into a pre -trained artificial intelligence model (such as a generative adversarial network) to generate corresponding architectural floor plans with a standard fonnat before modality alignment.

[0202] That is, in response to the standardized data being of the planar diagram modality, the electronic device inputs the standardized data of the planar diagram modality into a drawing generation model, which generates corresponding architectural floor plans with a standard fonnat based on the standardized data. Then, modality’ alignment is performed on the architectural floor plans with the standard format to obtain the target standardized data of the planar diagram modality.

[0203] Language Modality: User-input voice, user-input text, etc. For language modalities, after the initial preprocessing step, text data of the same format can be obtained, which can be directly aligned modally.

[0204] That is, in response to the standardized data being language modality data, the electronic device can perform modality alignment on the standardized data of the language modality’ to obtain the target standardized data of tire language modality.

[0205] After modality alignment of the three types of modalities, there are significant differences in feature scale, data type, and feature distribution. Therefore, different artificial intelligence models can be used for feature extraction of different modalities, and then the feature results output by different models can be aligned and fused in terms of modality features to form the final output.

[0206] In some embodiments, the electronic device can have pre-trained artificial intelligence models, and the electronic device can input the standardized data or target standardized data of different modalities into the Al model of that modality. The Al model then performs feature extraction on the standardized data or target standardized data to obtain the features of that modality.

[0207] It should be noted that the input data obtained in step 201 includes target building data and / or user requirement data for wireless private network planning. Therefore, in the embodiment of the present application, the extracted features may include building-related features and / or user requirement features.

[0208] In some embodiments, building-related features may include at least one of building structure information, building material information, and building space information.

[0209] In some embodiments, building structure information and corresponding building material information include, but are not limited to, at least one of beam, column, stair, wall, partition, window, and other building structure information, as well as the material information of these structures, such as concrete, metal, glass, etc. Building structure information and building material information can affect the distribution of wireless signals to varying degrees. For example, glass and wooden lightweight partitions have less obstruction to wireless signal transmission, while metal and concrete walls can severely hinder it. Energy-saving glass with a metal coating may block more wireless signal transmission, etc. In addition, building material information can also include the reflection coefficient and absorption coefficient corresponding to each material.

[0210] In some embodiments, building space information includes, but is not limited to, building length and width information, floor height information, floor thickness, number of floors, etc. Among them, building length and width information and floor height information are crucial for determining the efficiency of use of interior space and indoor network planning. Floor thickness affects the signal penetration ability between floors, and the number of floors affects the final base station configuration. For example, high-rise buildings may require the installation of micro base stations or repeaters to ensure coverage.

[0211] In some embodiments, in the network planning and design of wireless private networks (such as 5th Generation Mobile Communication Technology (5G) private networks), user requirements are usually related to their expectations for network perfonnance, usage scenarios, and specific needs of individuals or organizations. These features are very important for customizing solutions to meet specific network requirements.

[0212] User requirement features can generally be divided into two types. One type is related to the network implementation method and can be referred to as the first requirement feature. This includes but is not limited to the desired location for base station installation, the model of network equipment used, and the frequency information of the network equipment expected to be used. These network devices include but are not limited to core network devices, wireless access network devices, transmission network devices, gateway devices, etc. These network devices all have features related to network performance indicators, which can be divided into numerical features and categorical features based on feature types. Numerical features, such as the maximum throughput of the device, coverage range, power consumption, etc. Categorical features, such as the frequency range supported by the device, specific functions supported, power consumption level, etc. For numerical features, they can be directly processed into standardized features using normalization. For categorical features, they are first converted into integer-type features using feature encoding, and then normalized.

[0213] The other type of requirement feature is related to the output target, that is, the capacity and coverage of the network (for example, 5G), and can be referred to as the second requirement feature. For example, users may specify specific areas that they hope the network will cover according to different actual needs, such as video conference rooms that require high-speed data transmission, industrial automation production lines that require instant response, key infrastructure areas and emergency service areas that require reliability, etc. For these types of features, there can be different processing methods. In some embodiments, they can also be preprocessed and feature extracted, thereby included in the aforementioned steps 201 to step 203. In other embodiments, they can be detected and processed separately without undergoing the aforementioned treatment, thereby assisting the image generation process in step 205. For more details, please refer to the corresponding content on the processing of input data of target types in step 205.

[0214] In some embodiments, different feature extraction steps can be performed according to the modal type of tire standardized data or target standardized data to extract different features. The following describes the feature extraction for different types of modal data, and the feature extraction is just an example. Relevant technical personnel can also add other feature extraction methods and types of extracted features according to needs, and the embodiment of the present application does not limit this.

[0215] F or image modality data, in response to the standardized data or the target standardized data being of the image modality, the electronic device can perform at least one of the following steps A to C on the standardized data or target standardized data:

[0216] Step A: The electronic device can use an image recognition model to perform object recognition on the standardized data or target standardized data, output the building elements in the image, analyze the texture of local areas in the image using a classifier, identify the building material information in the image, and obtain the corresponding reflection coefficient and / or absorption coefficient based on the building material information.

[0217] In Step A, a pre -trained image recognition model is used to identify building elements in the image; a pre-trained classifier is used to analyze the texture of local areas in the image to identify different building materials, and the corresponding reflection and absorption coefficients are queried or calculated based on the identified building materials.

[0218] Step B: The electronic device uses a semantic segmentation model to perform image segmentation on the standardized data or target standardized data, obtaining visual cues in the image, and based on these visual cues, obtaining the spatial layout and the connectivity between spaces in the image.

[0219] Tn Step B, a pre-trained semantic segmentation model is used to analyze visual cues in the image, thereby inferring the spatial layout and the connectivity between spaces, such as floor height.

[0220] Step C: The electronic device uses monocular depth estimation technology to extract three- dimensional spatial information from two-dimensional images.

[0221] The three-dimensional spatial information can include various types, such as floor thickness, without limitation in the embodiment of the present application.

[0222] Step D: If the standardized data is LiDAR data, the electronic device uses ground segmentation technology to separate ground points from non-ground points in the LiDAR data, clusters the point data to obtain objects and scene elements in the LiDAR data, and performs feature extraction on these objects and scene elements to obtain spatial layout features and / or distance features.

[0223] Ground segmentation techniques, such as height thresholds, are used to separate ground points from non-ground points in the LiDAR data, and clustering algorithms are applied to cluster the point data, thereby distinguishing different objects and scene elements. Key features are then extracted from the distinguished data for obtaining information on spatial layout and distance.

[0224] The aforementioned steps B to D are illustrative of the process of extracting building spatial information. They are only one example, and relevant technical personnel can set which features to extract based on what technology according to needs or experience. Tire embodiment of the present application does not limit this.

[0225] For planar diagram modality data, in response to the standardized data or the target standardized data being of the planar diagram modality the electronic device can perform at least one of the following steps E to H on the standardized data or target standardized data:

[0226] Step E: The electronic device uses Optical Character Recognition (OCR) technology to extract building material information and / or building spatial information from tire planar diagram, and obtains the corresponding reflection coefficients and absorption coefficients based on the building material information.

[0227] In Step E, OCR technology is utilized to extract key building information from the planar diagram, such as floor height, material markings, etc., and the corresponding reflection and absorption coefficients are queried or calculated based on the identified materials.

[0228] Step F: The electronic device segments the planar diagram based on a semantic segmentation model to obtain different building areas and building elements, classifies these building areas and target building elements, and extracts building spatial information from the classification results. In Step F, a pre-trained semantic segmentation model is used to segment the planar diagram areas and classify them based on building areas and specific building elements (beams, columns, windows, stairs, etc.), and building spatial information such as floor boundaries and wall thicknesses is extracted from the features of different classifications.

[0229] Step G: The electronic device identifies taiget areas from the planar diagram, detects the planar diagram using image recognition technology, and obtains the base station installation points marked in the planar diagram, with the taiget area being the area in the planar diagram that requires targeted network coverage.

[0230] In Step G, specific areas that require targeted network coverage are identified from the drawings, and image recognition technology is used to detect and locate the base station installation points marked on the drawings.

[0231] Step H: The electronic device records positional point information and / or reflection point locations in the planar diagram, and calculates at least one of direct path, reflection path, direct path phase angle, and reflection path phase angle based on this positional point information and / or reflection point locations.

[0232] In Step H, positional point information and reflection point locations in the image arc recorded for subsequent calculation of direct and reflected paths, as well as direct and reflected path phase angles.

[0233] For language modality data, in response to the standardized data or the target standardized data being of the language modality, the electronic device can perform at least one of the following steps I to J on the standardized data or target standardized data:

[0234] Step I: The electronic device uses a Named Entity Recognition (NER) model to extract keywords from the text to obtain at least one of building structure information, building material information, or building spatial information.

[0235] In Step I, a trained NER model is used to extract keywords from the text, such as information on building structures, materials, spaces, etc.

[0236] Step J: The electronic device uses a relationship extraction model to determine the entity relationships in the text.

[0237] In Step J, a relationship extraction model is used to determine the entity relationships in the text, such as which materials arc used for which specific building structures.

[0238] It should be noted that the aforementioned steps A to J are only exemplary and are illustrated using building-related features as an example. In actual applications, the electronic device can also extract user requirement features from standardized data of different modalities. User requirements may include the first requirement features, the second requirement features, or both, as detailed above. Of course, relevant technical personnel can set which features to extract based on what technology according to needs or experience, and the embodiment of the present application does not limit this.

[0239] 204. The electronic device meiges the extracted features according to the target feature classification to obtain a target feature vector. After preprocessing the input data from multiple modalities and then performing feature extraction to obtain their respective features, the extracted features can be combined into a single feature vector. This means aligning and fusing the multimodal features to obtain an overall feature vector, allowing the artificial intelligence model to process the feature vector and generate a wireless private network signal coverage map.

[0240] In some embodiments, the features of the standardized data include building structural information, building material information, building spatial information, and user requirement information.

[0241] In some embodiments, the target feature vector includes multiple dimensions, and the target feature classification is used to indicate the dimensions in which features of different classifications are located within the target feature vector. Accordingly, step 204 can involve: the electronic device determines the dimension in the target feature vector corresponding to each extracted feature according to the target feature classification, and writes each extracted feature into the corresponding dimension in the taiget feature vector.

[0242] This process means that relevant technical personnel can preset the dimensions of the target feature vector based on needs or experience. Subsequently, after extracting the features from the standardized data, they can fill them into the corresponding dimensions of the taiget feature vector. For example, if the target feature vector is set to have floor height in the third dimension, the extracted feature of floor height can be filled into the third dimension of the target feature vector. In some embodiments, for features not extracted from the input data, such as user requirement- related information, building material information, building structural information, or building spatial information, preset default values are used instead.

[0243] 205. The electronic device inputs the target feature vector into the image generation model, which generates a wireless private network signal coverage map for the interior of the target building based on the target feature vector.

[0244] After integrating the features of multimodal data to obtain the target feature vector, a pretrained artificial intelligence model can be applied to generate a wireless private network signal coverage map based on the integrated features.

[0245] In some embodiments, the wireless private network signal coverage map may take the form of a heatmap. By processing the features of each dimension in tire target feature vector, the image generation task can be accomplished to obtain a heatmap. Of course, the wireless private network signal coverage map can also take other forms, and the embodiment of the present application does not limit this.

[0246] Tn some embodiments, in response to the input data including data of the target type, the electronic device can perform embedded encoding on the data of the target type to obtain an embedding vector. The target type is a requirement feature related to the output target, and the embedding vector is written into the input data of the image modality. Preprocessing operations are then performed on the input data after writing. The input data of the target type is the second requirement feature described in step 203, which is related to the output target, namely, the demand features related to 5G capacity and coverage. For example, users may specify specific areas they wish the network to cover based on different actual needs, such as video conference rooms that require high-speed data transmission, industrial automation production lines that require instant response, key infrastructure areas and emergency service areas that require reliability, etc.

[0247] For these types of features, embedded encoding can be used to transform these di screte feature information into continuous numerical vectors, which are then superimposed onto the building images for which the signal coverage map needs to be generated. This approach can make demands that are similar to each other close in vector space, allowing areas with different demands to remain in different vector spaces even when other features are highly similar. This achieves the customization of output network design and planning schemes according to user requirements.

[0248] Regarding user requirement features, they can actually be extracted from language modalities, sketches drawn by users, or both simultaneously. Of course, they might also be extracted from the standardized data obtained after preprocessing other modalities of input data. The following serves as a supplementary explanation for the situation where the feature extraction process in the aforementioned step 203 involves user requirement features:

[0249] For example, in the case of language modality: A trained Named Entity Recognition (NER) model can be used to extract keywords from text, such as base station type, base station location, base station transmission power, and other information. A relationship extraction model can also be used to determine the entity relationships in the text, such as which type of base station is used at which location.

[0250] Another example is sketches drawn by users: Optical Character Recognition (OCR) technology can be employed to recognize textual information from sketches and convert it into corresponding text, after which the same process as for language modality can be followed.

[0251] However, it is not limited to extracting user requirement features from this type of modal input data. This is just an exemplary' illustration. The specifics can be set by relevant technical personnel according to requirements, and the present application does not limit this.

[0252] Regarding the aforementioned steps 201 to 205 and their various implementations, two specific examples are provided below to illustrate the process in detail.

[0253] In a specific possible example, the wireless private network signal coverage map generation method is applied to a computer system, executed by a node, and the method may include the following steps one to four:

[0254] Step one: Receive a first signal and a second signal. In this example, the first signal may refer to building images captured by the user, building videos recorded by the user, floor plans of building interiors, sketches drawn by the user, voice information input by' the user, text information input by the user, but not limited to these. Tire second signal may be voice information input by the user, text information input by the user, but not limited to these.

[0255] Optionally , the first signal is a floor plan of the building interior.

[0256] Optionally, adjust the resolution of the received floor plan. Typically, the resolution to be adjusted is predetermined. According to the predetermined resolution (i.e., the aforementioned target resolution), the image can be adjusted to the corresponding resolution, while using bicubic interpolation to ensure that the image quality is not reduced during the adjustment process.

[0257] Optionally, obtain the three primary color component information of the image, use a weighted method combined with the different importance of the three primary color components of the image, such as [0.299, 0.587, 0.114], to process the image into grayscale, and obtain the corresponding grayscale image. By convoluting the Sobel operator with the grayscale image, the horizontal and vertical grayscale differential values Gx and Gy of the image's pixel points can be obtained. By performing a square sum operation on the honzontal and vertical grayscale differential values, the grayscale magnitude of each pixel point is obtained. When the grayscale magnitude is greater than a first preset value, the pixel point is determined to be an edge point, and the edge image is ultimately obtained.

[0258] Optionally, the second signal is voice information input by the user.

[0259] Optionally, for the input voice information, Wiener filtering is used for noise reduction treatment to obtain higher quality voice output.

[0260] Through an Al network, the noise-reduced voice signal is transformed into the corresponding text sequence.

[0261] Optionally, through natural language processing techniques, irrelevant stop words and symbols are removed from the converted text. And through semantic analysis, key words and phrases are extracted, thereby segmenting the text into a collection of words or terms containing actual information, for subsequent feature extraction.

[0262] Step two: Extract building structure-related features from the first signal.

[0263] Based on the extracted edge image, use an Al network to perform pixel-level segmentation on the floor plan area and recognize and annotate the segmented areas.

[0264] For the segmented areas, use morphological operations (dilation, erosion, etc.) and object detection algorithms to further identify the building framework, architectural elements (beams, windows, columns, stairs, etc.).

[0265] Extract building-related features from the segmented areas and recognized architectural elements. Optionally, extract the length, width, and height information of the recognized building areas as the first feature vector.

[0266] Optionally, extract key structural parameters from the recognized building framework and architectural elements, such as node positions, wall thickness, and length, as the second feature vector.

[0267] From the annotated floor plan, extract the RGB channel information of different elements' locations, combine it with a preset color material lookup table, and extract the material information used by each element (concrete, metal, glass, etc.). Query the preset matenal database to obtain the reflection coefficient and absorption coefficient corresponding to the materials in the second feature vector as the third feature vector.

[0268] Using Optical Character Recognition (OCR) technology, extract key signal coverage area information and base station installation location information contained in the floor plan as the fourth feature vector.

[0269] From the annotated floor plan, extract the distance and angle information between the base station location and the wall elements of the floor plan as part of the fourth feature vector.

[0270] From the annotated floor plan, extract the distance and angle information between the base station location and the node positions extracted in the second feature vector as tire fifth feature vector.

[0271] Merge the extracted first, second, third, fourth, and fifth features into the sixth feature vector, then store them in the same data structure for subsequent analysis and processing.

[0272] Step three: Extract user requirement-related features from the second signal.

[0273] Utilizing word embedding technology; transform the preprocessed second signal into a numerical vector. These numerical vectors represent the semantic and syntactic features of words and can indicate relationships between different words.

[0274] Input the transformed numerical vectors into an Al network to extract two types of user requirement features: 1. Requirements related to the implementation method: Extract numerical features such as base station deployment locations, network device models, frequency information, maximum throughput, power consumption, and categorical features like the frequency range supported by the device, specific functions, etc., and arrange these features as the seventh feature vector; 2. Requirements related to the output taiget: Extract user-specified network coverage demand areas, such as video conference rooms, industrial automation production lines, etc.

[0275] Merge the seventh feature vector with the sixth feature vector to output the eighth feature vector (also known as the target feature vector). If the seventh feature vector and the sixth feature vector have similar features, the value of that feature from the seventh feature vector will be taken as the value for the sixth feature vector.

[0276] Combine user-input requirements related to the output target with the building's floor plan to generate a mask matrix of the same size as the pixels of the floor plan, where the mask values corresponding to each area are determined by the user-specified taiget areas. Merge the generated mask matrix as an additional channel of the floor plan for subsequent combination with the feature vector to output the target building's 5G signal coverage heatmap.

[0277] The data processing method of this invention provides a complete preprocessing and feature extraction method for multimodal data input by7users, transforming multimodal data containing complex information into standardized data suitable for Al model learning and inference. It also outputs feature vectors containing key information needed to generate 5G signal coverage maps through pre-trained Al models. Moreover, the multimodal data processing and feature extraction procedures provided are automated, requiring no additional operations from the user. Furthermore, users only need to input their desired requirements, and the user requirement feature extraction module can extract the desired requirement information, making the feature vectors output by the Al model more aligned with the user's actual needs. Some user requirements can also be used to guide the model to focus on covering specific areas when generating signal coverage hcatmaps. These operations reduce the barrier to obtaining network planning and design solutions, making it easier for small and medium-sized enterprises to deploy 5G private networks, and of course, not limited to 5G private networks.

[0278] In some embodiments, tire implementation of tire wireless private network signal coverage map generation method provided by the present application example helps to generate high-quality training data and annotated data, enhancing the performance and reliability of large models in the planning and design of wireless private networks (such as 5G networks). Specifically, the electronic device can use the input data and the target feature vector as training data for the image generation model, and use the wireless private network signal coverage map inside the target building as the annotated data for that training data.

[0279] In the present application example, considering the excellent advantage of artificial intelligence technology in performing generation tasks based on a large amount of data, and that different modalities of data have different forms of expression, data types, and feature distributions, a data standardization preprocessing process is designed for data heterogeneity before applying the image generation model. By performing different preprocessing operations on the input data according to different modalities and transforming it into standardized data in a preset format, and then performing feature extraction on the standardized data of each modality the extracted features can accurately represent the data of each modality, facilitating subsequent processing and fusion of features of different modalities within the same framework. It is precisely7because of the aforementioned standardization processing method that the extracted features can be merged according to the target feature classification to obtain a feature vector that meets the model generation requirements, thereby generating a wireless private network signal coverage map. The above data processing, feature extraction, and image generation processes arc all automated, and users only7need to input building data or requirement data to automatically7implement wireless private network planning, greatly reducing the threshold for network planning and design, improving the efficiency and versatility of wireless private network planning, and reducing its cost.

[0280] All the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described one by one again.

[0281] Figure 3 is a schematic diagram of the structure of a wireless private network signal coverage map generation device provided by an embodiment of the present application. See Figure 3, the device includes:

[0282] • Acquisition module 301, which is used to obtain input data, including target building data and / or user requirement data for wireless private network planning, and the input data includes one or more modalities;

[0283] • Preprocessing module 302, which is used to preprocess the input data according to the modality of the input data and the preprocessing method corresponding to the modality, converting the input data into standardized data in the format corresponding to the modality;

[0284] • Extraction module 303, which is used to extract features from the standardized data of each modality, obtaining the features of the standardized data;

[0285] • Merging module 304, which is used to merge the extracted features according to the target feature classification, obtaining tire target feature vector;

[0286] • Generation module 305, which is used to input the target feature vector into the image generation model, and the image generation model generates a wireless private network signal coverage map inside the target building based on the target feature vector.

[0287] In some embodiments, the input data includes multiple modalities, including various types of modalities, where one type of modality includes one or more modalities, and one type of modality corresponds to a unified format;

[0288] • The preprocessing module 302 is used to preprocess the input data of multiple modalities of the same type of modality, respectively, through the preprocessing methods corresponding to the multiple modalities, adjusting the input data of the multiple modalities to standardized data in the unified format of the type of modality.

[0289] In some embodiments, the input data includes multiple modalities, including at least two of the following: photographic image modality, planar image modality, graphic file modality, video modality, LiDAR modality, infrared induction modality, text modality, and voice modality; the input data includes at least two of tire following: building images photographed by users in the photographic image modality, building videos recorded by users in the video modality, floor plans of building interiors in the planar image modality, architectural sketches drawn by users in the planar image modality, graphic files output by professional drawing software in the graphic file modality, LiDAR data in the LiDAR modality, infrared induction data in the infrared induction modality, text information of users in the text modality, and voice information of users in the voice modality.

[0290] The preprocessing module 302 is used to perform at least two of the following:

[0291] • In response to the input data including building images photographed by the user, adjust the resolution of the building image to the target resolution, convert the bitmap of the building image to grayscale to obtain a first image, perform edge detection on the first image to obtain the contour features of the building structure, and use the first image and the contour features as the standardized data of the building image;

[0292] • In response to the input data including building videos recorded by users in the video modality, decode the building video, convert the building video into a sequence of frames; extract keyframes from the frame sequence to obtain the target keyframes; perform resolution adjustment, grayscale processing, and edge detection on the target keyframes to obtain the standardized data of the target keyframes;

[0293] • In response to the input data including floor plans of building interiors and / or architectural sketches drawn by users, perform resolution adjustment and edge detection processing on the floor plans and / or architectural sketches to obtain the standardized data of the floor plans;

[0294] • In response to the input data including graphic files output by professional drawing software, call a professional conversion tool to convert the graphic file, obtain a second image in the target image format, and adjust the resolution of the second image to the target resolution to obtain the standardized data of the graphic file;

[0295] • In response to the input data including text information input by the user, remove stop words and irrelevant symbols from the text information, convert abbreviations, synonyms, and variant words in the text information, and transform the processed text information into the target text format to obtain the standardized data of the text information;

[0296] • In response to the input data including voice information input by the user, perform noise reduction processing on the voice information, perform speech recognition on the noise-reduced voice information to obtain the first text information corresponding to the voice information, perform preprocessing such as removing stop words and irrelevant symbols, converting abbreviations, synonyms, and variant words, and text format conversion on the first text information to obtain the standardized data of the voice information;

[0297] • In response to the input data including LiDAR data, filter the LiDAR data, and convert the filtered LiDAR data from the device coordinate system to the global coordinate system to obtain the standardized data of the LiDAR data;

[0298] • In response to tire input data including infrared induction data, filter the infrared induction data, calibrate the filtered infrared induction data for temperature, and enhance the calibrated infrared induction data to obtain the standardized data of the infrared induction data.

[0299] Tn some embodiments, the multiple modalities include at least two types of modalities from the categories of image modalities, planar diagram modalities, and language modalities. The image modalities include photographic image modality’, graphic file modality, video modality, LiDAR modality, and infrared induction modality; the planar diagram modalities include planar image modality; and the language modalities include text modality and voice modality.

[0300] The extraction module 303 is used for:

[0301] • Aligning the standardized data of the same modality type to obtain the target standardized data for each modality;

[0302] • Executing the feature extraction step on the aligned target standardized data to obtain the features of the target standardized data. In some embodiments, the extraction module 303 is used to perform at least one of the following: • In response to the standardized data or the target standardized data being image modality data, perform at least one of the following on the standardized data or target standardized data: o U se an image recognition model to identify targets in the standardized data or target standardized data, output building elements in the image, analyze the texture of local areas in the image using a classifier, identify building material information in the image, and based on the building material information, obtain the corresponding reflection coefficient and / or absorption coefficient; o Use a semantic segmentation model to perform image segmentation on the standardized data or target standardized data, obtain visual cues in the image, and based on the visual cues, obtain the spatial layout and connections between spaces in the image; o Use monocular depth estimation technology to extract three-dimensional spatial information from two-dimensional images; o If the standardized data is LiDAR data, use ground segmentation technology to separate ground points from non-ground points in the LiDAR data, cluster the point data to obtain objects and scene elements in the LiDAR data, perform feature extraction on these objects and scene elements to obtain spatial layout features and / or distance features;

[0303] • In response to the standardized data or the target standardized data being planar diagram modality' data, perform at least one of the following on the standardized data or target standardized data: o Use optical character recognition technology to extract building material information and / or building spatial information from the planar diagram, and based on the building material information, obtain the corresponding reflection coefficient and absorption coefficient; o Use a semantic segmentation model to segment the planar diagram to obtain different building areas and building elements, classify these building areas and target building elements, and extract building spatial information from the classification results; o Identify target areas in the planar diagram, detect the planar diagram using image recognition technology to obtain the base station installation points marked in the planar diagram, with the target area being the area in the planar diagram that requires targeted network coverage; o Record positional point information and / or reflection point locations in the planar diagram, and based on this information, calculate at least one of direct path, reflection path, direct path phase angle, and reflection path phase angle;

[0304] • In response to the standardized data or the target standardized data being language modality data, perform at least one of the following on the standardized data or target standardized data: o Use a named entity recognition model to extract keywords from the text to obtain at least one of building structure information, building material information, and building spatial information; o Use a relationship extraction model to determine entity relationships in the text. In some embodiments, the device also includes:

[0305] • In response to the input data including data of the target type, perform embedded encoding on the data of the taiget type to obtain an embedding vector, with the target type being a demand feature rel ted to the output target:

[0306] • Write the embedding vector into the image modality input data and perform preprocessing operations on the written input data. In some embodiments, the target feature vector includes multiple dimensions, with the target feature classification indicating the dimensions in the target feature vector where features of different classifications are located.

[0307] The merging module 304 is used for:

[0308] • Determining the corresponding dimension in the target feature vector for each extracted feature according to the target feature classification;

[0309] • Writing each extracted feature into the corresponding dimension in the taiget feature vector.

[0310] • In some embodiments, the generation module 305 is also used to use the input data and the target feature vector as training data for the image generation model, using the wireless private network signal coverage map inside the taiget building as the annotated data for the training data.

[0311] • The device provided by the embodiment of the present application considers that artificial intelligence technology has excellent advantages in performing generation tasks based on a large amount of data, and different modalities of data have different forms of expression, data types, and feature distributions. Before applying the image generation model, a data standardization preprocessing process is designed for data heterogeneity. By performing different preprocessing operations on the input data according to different modalities and transforming it into standardized data in a preset format, and then performing feature extraction on the standardized data of each modality', the extracted features can accurately represent the data of each modality, which is conducive to processing and fusing features of different modalities within the same framework. It is precisely because of the above standardization process that the extracted features can be meiged according to the taiget feature classification to obtain a feature vector that meets the model generation requirements, and then generate a wireless private network signal coverage map. The above data processing, feature extraction, and image generation processes arc all automated processes. Users only need to input building data or requirement data to automatically implement wireless private network planning, greatly' reducing the threshold when planning and designing the network, improving the efficiency and versatility of wireless private network planning, and reducing its cost.

[0312] • It should be noted: The wireless private network signal coverage map generation device provided by the embodiment in the generation of the wireless private network signal coverage map is only used as an example to illustrate the division of the above functional modules. In actual applications, according to the needs, the above functions are allocated to different functional modules to be completed, that is, the internal structure of the wireless private network signal coverage map generation device is divided into different functional modules to complete all or part of the functions described above. In addition, the wireless private network signal coverage map generation device provided by the embodiment and the wireless private network signal coverage map generation method embodiment belong to the same conception, and the specific implementation process is detailed in the method embodiment, which is not repeated here.

[0313] Figure 4 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device 400 can vary significantly due to different configurations or performance, including one or more processors (Central Processing Units, CPUs) 401 and one or more memories 402. The memory 402 stores at least one computer program, which is loaded and executed by tire processor 401 to implement the wireless private network signal coverage map generation method provided by the various method embodiments described above. The electronic device also includes other components for implementing the device's functions. For example, the electronic device also has wired or wireless network interfaces and input / output interfaces, etc., for input and output. This embodiment of the application does not provide further elaboration.

[0314] The electronic device in the above method embodiments is implemented as a terminal. For example, Figure 5 is a structural block diagram of a terminal provided by an embodiment of the present application. The terminal 500 can be a portable mobile terminal, such as a smartphone, tablet computer, MP3 (Moving Picture Experts Group Audio Layer III, dynamic image expert compression standard audio level 3) player, MP4 (Moving Picture Experts Group Audio Layer IV, dynamic image expert compression standard audio level 4) player, laptop computer, or desktop computer. The terminal 500 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0315] Typically, the terminal 500 includes: a processor 501 and a memory 502.

[0316] The processor 501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 501 may be implemented using at least one hardware form of DSP (Digital Signal Processing, digital signal processing), FPGA (Field-Programmable Gate Array, field-programmable gate array), PLA (Programmable Logic Array, programmable logic array). The processor 501 may also include a main processor and a coprocessor, where the main processor is a processor used for processing data in the wake state, also known as the CPU (Central Processing Unit, central processing unit); the coprocessor is a low-power processor used for processing data in the standby state. In some embodiments, the processor 501 may integrate a GPU (Graphics Processing Unit, graphics processor), which is responsible for rendering and drawing the content required to be displayed on the display screen. Tn some embodiments, the processor 501 may also include an Al (Artificial Intelligence, artificial intelligence) processor, which is used to process computational operations related to machine loaming.

[0317] The memory 502 may include one or more computer-readable storage media, which may be non-volatile. The memory 502 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-volatile computer-readable storage medium in the memory7502 is used to store at least one instruction, which is executed by the processor 501 to implement the wireless private network signal coverage map generation method provided by the method embodiments in this application.

[0318] In some embodiments, the terminal 500 may also optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501 , memory 502, and peripheral device interface 503 may7be connected by a bus or signal line. Each peripheral device may7be connected to the peripheral device interface 503 via a bus, signal line, or circuit board. Specifically, the peripheral devices include at least one of: RF circuit 504, display7505, camera module 506, audio circuit 507, positioning component 508, and power supply 509.

[0319] The peripheral device interface 503 may be used to connect at least one peripheral device related to I / O (Input / Output, input / output) to the processor 501 and memory7502. In some embodiments, the processor 501, memory' 502, and peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 501, memory 502, and peripheral device interface 503 may be implemented on separate chips or circuit boards, and this embodiment does not limit this.

[0320] The RF circuit 504 is used to receive and transmit RF (Radio Frequency, radio frequency) signals, also known as electromagnetic signals. The RF circuit 504 communicates with communication networks and other communication devices through electromagnetic signals. The RF circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, tuners, oscillators, digital signal processors, codec chipsets, Universal Integrated Circuit Cards (UICCs), etc. The RF circuit 504 can communicate with other tenninals through at least one wireless communication protocol. This wireless communication protocol includes but is not limited to: the World Wide Web, metropolitan area networks, intranets, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity7, wireless fidelity) networks. Tn some embodiments, the RF circuit 504 may also include circuits related to NFC (Near Field Communication, ncar-ficld wireless communication), and this application docs not limit this.

[0321] The display 505 is used to display the UI (User Interface, user interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display 505 is a touch screen, the display 505 also has the ability to capture touch signals on the surface or above the surface of the display7505. The touch signal can be input as a control signal to the processor 501 for processing. At this time, the display 505 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display 505 may be one, set on the front panel of the terminal 500; in other embodiments, the display 505 may be at least tw o, set on different surfaces of the terminal 500 or in a folding design; in other embodiments, the display 505 may be a flexible display, set on the curved surface or folding surface of the terminal 500 Even the display 505 can also be set to an irregular shape that is not a rectangle, that is, an irregular screen. The display 505 can be made of materials such as LCD (Liquid Crystal Display, liquid crystal display) and OLED (Organic Light-Emitting Diode, organic light-emitting diode).

[0322] The camera module 506 is used to capture images or videos. Optionally, the camera module 506 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal, hi some embodiments, the rear camera is at least two, which are any one of the main camera, depth of field camera, wide- angle camera, and telephoto camera, to achieve background blurring functions by fusing the main camera and the depth of field camera, panoramic shooting and VR (Virtual Reality, virtual reality) shooting functions or other fusion shooting functions by fusing the main camera and the wide- angle camera. In some embodiments, the camera module 506 may also include a flash. The flash can be a monochromatic flash or a dual-color temperature flash. A dual-color temperature flash refers to the combination of wann light flash and cold light flash, which can be used for light compensation under different color temperatures.

[0323] The audio circuit 507 may include a microphone and a speaker. The microphone is used to capture user and environmental sound waves and convert the sound waves into electrical signals input to the processor 501 for processing, or input to the RF circuit 504 to achieve voice communication. For the purpose of stereo sound capture or noise reduction, there may be multiple microphones, respectively set in different parts of the terminal 500. The microphone can also be an array microphone or an omnidirectional capture-type microphone. The speaker is used to convert the electrical signals from the processor 501 or the RF circuit 504 into sound waves. The speaker can be a traditional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves that humans can hear but also convert electrical signals into sound waves that humans cannot hear for ranging and other purposes. In some embodiments, the audio circuit 507 may also include a headphone jack.

[0324] The positioning component 508 is used to locate the current geographical location of the terminal 500 to achieve navigation or LBS (Location Based Service, location-based service). The positioning component 508 can be a positioning component based on the American GPS (Global Positioning System, global positioning system), the Chinese Beidou system, or the Russian GLONASS system.

[0325] The power supply 509 is used to power various components in the terminal 500. The power supply 509 can be AC, DC, disposable batteries, or rechargeable batteries. When the power supply 509 includes a rechargeable battery, the rechargeable battery can be a wired battery or a wireless battery. The wired battery' is a battery that is charged through wired lines, and the wireless battery' is a battery' that is charged through wireless coils. The rechargeable battery’ can also be used to support fast charging technology.

[0326] In some embodiments, the terminal 500 may also include one or more sensors 510. The one or more sensors 510 include but are not limited to: an acceleration sensor 511, a gyroscope sensor

[0327] 512, a pressure sensor 513, a fingerprint sensor 514, an optical sensor 515, and a proximity sensor 516.

[0328] The acceleration sensor 511 can detect the acceleration magnitude on the three coordinate axes of the coordinate system established by the terminal 500. For example, the acceleration sensor 511 can be used to detect the components of gravitational acceleration on the three coordinate axes. Tire processor 501 can control the display 505 to display the user interface in landscape or portrait view based on the gravitational acceleration signal collected by the acceleration sensor 511. The acceleration sensor 511 can also be used for games or user motion data collection.

[0329] The gyroscope sensor 512 can detect the orientation and rotation angle of the terminal 500, and the gyroscope sensor 512 can work in conjunction with tire acceleration sensor 511 to capture the user's 3D movements with respect to the terminal 500. The processor 501, based on the data collected by the gyroscope sensor 512, can implement functions such as: motion sensing (e.g., changing the UI based on the user's tilting operation), image stabilization during photography, game control, and inertial navigation.

[0330] The pressure sensor 513 can be set on the side frame and / or under the display 505 of the terminal 500. When the pressure sensor 513 is set on the side frame of the terminal 500, it can detect the user's grip signal on the terminal 500, and the processor 501 can recognize left and right hand use or perform quick operations based on the grip signal collected by the pressure sensor

[0331] 513. When the pressure sensor 513 is set under the display 505, the processor 501 can control operable controls on the UI interface based on the pressure operations of the user on the display 505. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0332] The fingerprint sensor 514 is used to collect the user's fingerprints, and the processor 501 identifies the user's identity based on the fingerprints collected by tire fingerprint sensor 514, or the fingerprint sensor 514 identifies the user's identity based on the collected fingerprints. When the user's identity is recognized as a trusted identity, the processor 501 authorizes the user to perform related sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 514 can be set on the front, back, or side of the terminal 500. When the terminal 500 has physical buttons or a manufacturer's logo, the fingerprint sensor 514 can be integrated with the physical buttons or the manufacturer's logo.

[0333] The optical sensor 515 is used to collect ambient light intensity. In one embodiment, the processor 501 can control the display brightness of the display 505 based on the ambient light intensity collected by the optical sensor 515. Specifically, when the ambient light intensity is high, the display brightness of the display 505 is increased; when the ambient light intensity is low, the display brightness of the display 505 is decreased. In another embodiment, the processor 501 can also dynamically adjust the shooting parameters of the camera module 506 based on the ambient light intensity collected by the optical sensor 515. The proximity sensor 516, also known as the distance sensor, is typically set on the front panel of the terminal 500. The proximity sensor 516 is used to collect the distance between the user and the front of the terminal 500. In one embodiment, when the proximity sensor 516 detects that the distance between the user and the front of the tenninal 500 is gradually decreasing, the processor 501 controls the display 505 to switch from the bright screen state to the screen-off state; when the proximity sensor 516 detects that the distance between the user and the front of the tenninal 500 is gradually increasing, the processor 501 controls the display 505 to switch from the screen-off state to the bright screen state.

[0334] Persons skilled in the art can understand that the structure shown in Figure 5 does not limit the terminal 500 and can include more or fewer components than shown, combine certain components, or adopt different component layouts.

[0335] The electronic device in the above method embodiments is implemented as a server. For example, Figure 6 is a structural schematic diagram of a server provided by an embodiment of the present application. The server 600 can vary significantly due to different configurations or performance, including one or more processors (Central Processing Units, CPUs) 601 and one or more memories 602, wherein the memory 602 stores at least one computer program, which is loaded and executed by the processor 601 to implement the wireless private network signal coverage map generation method provided by the various method embodiments described above. Of course, the server also has wired or wireless network interfaces as well as input / output interfaces for input and output, and the server includes other components for implementing the device's functions, which are not further described here.

[0336] In an exemplary implementation, a computer-readable storage medium is also provided, such as a memory including at least one computer program, which can be executed by a processor to complete the wireless private network signal coverage map generation method described in the above embodiments. For example, the computer-readable storage medium is Read-Only Memory (ROM), Random Access Memory (RAM), Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage devices.

[0337] In an exemplary implementation, a computer program product or computer program is also provided, which includes one or more program codes stored on a computer-readable storage medium. One or more processors of an electronic device read the one or more program codes from the computer-readable storage medium, and the one or more processors execute the one or more program codes, causing the electronic device to perform the wireless private network signal coverage map generation method described above.

[0338] In some embodiments, the computer program involved in the present application embodiment may be deployed to be executed on a computer device, or executed on multiple computer devices located at a single site, or executed on multiple computer devices distributed across multiple sites and interconnected by a communication network, with the multiple computer devices distributed across multiple sites and interconnected by a communication network forming a blockchain system. The ordinary technician in the field understands that all or part of the steps of tire above embodiments are completed by hardware, and also by a program instructing the relevant hardware to complete, with the program stored on a computer-readable storage medium, the storage medium mentioned above being Read-Only Memory, disk, or disc, etc.

[0339] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present application, should be included within the protection scope of the present application.

Claims

CLAIMSWhat is claimed is:

1. A method for generating a wireless private network signal coverage map, characterized in that the method includes:Obtaining input data, which includes target building data and / or user requirement data for wireless private network planning, and includes one or more modalities:Preprocessing the input data according to the modality of the input data and the preprocessing method corresponding to the modality, converting the input data into standardized data in the format corresponding to the modality;Feature extracting based on the standardized data for each modality to obtain features of the standardized data;Meiging the extracted features according to the taiget feature classification to obtain a target feature vector;Inputting the target feature vector into an image generation model, which generates a wireless private network signal coverage map inside the target building based on the target feature vector.

2. The method according to claim 1, characterized in that the input data includes multiple modalities, which include multiple types of modalities, one type of modality includes one or more modalities, and one type of modality corresponds to a unified format;The preprocessing of the input data according to the modality of the input data and the preprocessing method corresponding to the modality, converting the input data into standardized data in the format corresponding to the modality, includes:For input data of multiple modalities of the same type, preprocessing is performed separately through the preprocessing methods corresponding to the multiple modalities, adjusting the input data of the multiple modalities to standardized data in the unified format of the modality type.

3. The method according to claim 1 or 2, characterized in that the input data includes multiple modalities, which include at least two of the following: photographic image modality, planar image modality, graphic file modality, video modality, LiDAR modality', infrared induction modality', text modality, and voice modality; the input data includes at least two of the following: architectural images captured by users in the photographic image modality, architectural videos recorded by users in the video modality, floor plans of building interiors in the planar image modality’, architectural sketches drawn by users in the planar image modality, graphic files output by professional drawing software in the graphic file modality, LiDAR data in the LiDAR modality, infrared induction data in theinfrared induction modality, text information of users in the text modality, and voice information of users in the voice modality7:The preprocessing of the input data according to the modality of the input data and the preprocessing method corresponding to the modality, converting the input data into standardized data in the format corresponding to the modality, includes at least two of the following:In response to the input data including architectural images captured by users, adjust the resolution of the architectural images to the target resolution, convert the raster of the architectural images to grayscale to obtain a first image, perform edge detection on the first image to obtain the contour features of the building structure, and use the first image and the contour features as the standardized data of tire architectural image;In response to the input data including architectural videos recorded by users in the video modality7, decode the architectural video to convert it into a frame sequence: extract keyframes from the frame sequence to obtain target keyframes; perform resolution adjustment, grayscale processing, and edge detection on the target keyframes to obtain the standardized data of the target keyframes;In response to the input data including floor plans of building interiors and / or architectural sketches drawn by users, perform resolution adjustment and edge detection processing on the floor plans and / or architectural sketches to obtain the standardized data of the floor plans;In response to the input data including graphic files output by professional drawing software, call a professional conversion tool to convert the graphic files into a target image format, adjust the resolution of the second image to the target resolution, and obtain the standardized data of the graphic files;In response to the input data including text information input by users, remove stop words and irrelevant symbols from the text information, convert abbreviations, synonyms, and variant words in the text information, transform the processed text information into the target text format, and obtain the standardized data of the text information;In response to the input data including voice information input by users, perform noise reduction processing on the voice information, perform speech recognition on the noise- reduced voice information to obtain the first text information corresponding to the voice information, remove stop words and irrelevant symbols from the first text information, convert abbreviations, synonyms, and variant words, and perform preprocessing such as text format conversion to obtain the standardized data of the voice information;In response to the input data including LiDAR data, filter the LiDAR data and convert the filtered LiDAR data from the device coordinate system to the global coordinate system to obtain the standardized data of the LiDAR data;In response to the input data including infrared induction data, filter the infrared induction data, calibrate the filtered infrared induction data for temperature, and enhancethe calibrated infrared induction data to obtain tire standardized data of tire infrared induction data.

4. The method according to claim 3, characterized in that the multiple modalities include multiple types of modalities, which include at least two of the following: image modality, planar diagram modality, and language modality, where the image modality includes at least two of the photographic image modality, graphic file modality, video modality, LiDAR modality; and infrared induction modality; the planar diagram modality' includes the planar image modality; and the language modality includes the text modality' and voice modality;The feature extraction of the standardized data for each modality to obtain the features of the standardized data includes:Aligning the standardized data of the same type of modality to obtain the target standardized data for each modality;Performing the feature extraction steps on the aligned target standardized data to obtain the features of the target standardized data.

5. The method according to claim 4, characterized in that the feature extraction of the standardized data for each modality to obtain the features of the standardized data includes at least one of the following:In response to the standardized data or the taiget standardized data being image modality data, performing at least one of the following on the standardized data or target standardized data:Using an image recognition model to perform object recognition on the standardized data or target standardized data, outputting building elements in tire image, using a classifier to analyze the texture of local areas in the image, identifying building material information in the image, and obtaining corresponding reflection coefficients and / or absorption coefficients based on the building material information;Using a semantic segmentation model to perform image segmentation on the standardized data or target standardized data to obtain visual cues in the image, and obtaining spatial layout and connections between spaces in the image based on the visual cues;Using monocular depth estimation technology to extract three-dimensional spatial information from two-dimensional images;If the standardized data is UiDAR data, using ground segmentation technology to separate ground points from non-ground points in the LiDAR data, clustering point data to obtain objects and scene elements in the LiDAR data, and performing feature extraction on the objects and scene elements to obtain spatial layout features and / or distance features;In response to the standardized data or the taiget standardized data being planardiagram modality data, performing at least one of the following on tire standardized data or target standardized data:Using Optical Character Recognition (OCR) technology to extract building material information and / or building spatial information from the planar diagram, and obtaining corresponding reflection coefficients and absorption coefficients based on the building material information:Using a semantic segmentation model to segment the planar diagram to obtain different building areas and building elements, classifying the building areas and target building elements, and extracting building spatial information from the classification results:Identifying target areas from the planar diagram, using image recognition technology to detect the planar diagram to obtain base station installation points marked in the planar diagram, with the taiget area being the area in the planar diagram that requires targeted network coverage:Recording position point information and / or reflection point locations in the planar diagram, and calculating at least one of direct path, reflection path, direct path phase angle, and reflection path phase angle based on the position point information and / or reflection point locations;In response to the standardized data or the target standardized data being language modality data, performing at least one of the following on the standardized data or target standardized data:Using a Named Entity Recognition (NER) model to extract key ords from the text to obtain at least one of building structure information, building material information, and building spatial information;Using a relationship extraction model to determine entity relationships in the text.

6. The method according to claim 1, characterized in that the method further includes:Tn response to the input data including data of the target type, perfonning embedded encoding on the data of the target type to obtain an embedding vector, where the target type is a requirement feature related to the output target;Writing the embedding vector into the input data of the image modality, and performing preprocessing operations on the input data after the writing.

7. The method according to claim 1, characterized in that the target feature vector includes multiple dimensions, and the target feature classification is used to indicate the dimensions in which different classified features are located within the taiget feature vector;The merging of the extracted features according to the target feature classification to obtain the target feature vector includes:According to the target feature classification, determining the dimension in the taigetfeature vector corresponding to each extracted feature;Writing each extracted feature into the corresponding dimension in the target feature vector.

8. A wireless private network signal coverage map generation device, characterized in that the device includes:An acquisition module, which is used to obtain input data, the input data includes target building data and / or user requirement data for wireless private network planning, and the input data includes one or more modalities;A preprocessing module, which is used to perform preprocessing corresponding to the modality of the input data and the preprocessing method corresponding to the modality, converting the input data into standardized data in tire format corresponding to the modality;A feature extraction module, which is used to perform feature extraction on the standardized data for each modality to obtain the features of the standardized data;A merging module, which is used to merge the extracted features according to the target feature classification to obtain a target feature vector;An image generation module, which is used to input the target feature vector into the image generation model, and the image generation model generates a wireless private network signal coverage map inside the target building based on the target feature vector.

9. An electronic device, characterized in that the electronic device includes one or more processors and one or more memories, and the one or more memories store at least one computer program, which is loaded and executed by the one or more processors to implement any of the wireless private network signal coverage map generation methods described in claims 1 to 7.

10. A computer-readable storage medium, characterized in that the storage medium stores at least one computer program, which is loaded and executed by a processor to implement any of the wireless private network signal coverage map generation methods described in claims 1 to 7.