A machine learning-based data processing method, system and computer device

By classifying and processing multiple data sources in parallel, the spectrum occupancy rate and the number of holes are obtained, which solves the problem of insufficient spectrum data quality control and realizes refined management and efficient utilization of spectrum resources.

CN120337040BActive Publication Date: 2025-10-17HANGZHOU IDEACOME INTERNET FINANCIAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510813440.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-17
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Existing machine learning algorithms lack effective data preprocessing and quality assessment mechanisms in spectrum data processing, leading to noise, bias, and missing values ​​due to differences in data sources. This makes it difficult to dynamically adapt to changes in spectrum resource usage patterns and affects the accuracy of analysis results.

Method used

By acquiring initial data from multiple data sources, classifying, dividing, and processing it in parallel, the spectrum occupancy rate and number of holes are obtained, idle frequency bands are filtered, mapped to a high-dimensional feature space, and inversely mapped to obtain frequency band resource features, thus realizing frequency band identification and processing.

Benefits of technology

It improves the accuracy and utilization of spectrum resource analysis, ensures the reliability and completeness of analysis results, and supports the refined management of spectrum resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337040B_ABST
    Figure CN120337040B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, in particular to a data processing method and system based on machine learning and computer equipment. The method comprises the following steps: obtaining multiple initial data based on multiple data sources; obtaining the data tightness of each initial data, classifying the initial data based on the data tightness to obtain multiple data classification libraries; dividing each data classification library based on a frequency band distributed node to obtain multiple data calculation nodes, and processing each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set. The initial data is obtained from multiple sources, the information of different channels is integrated, the limitation of a single data source is avoided, a comprehensive and rich data basis is provided for spectrum resource analysis, the reliability and integrity of the analysis result are ensured, the data processing speed and spectrum analysis accuracy are improved through data tightness classification, frequency band distributed node division and parallel processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data processing method, system and computer equipment based on machine learning. Background Art

[0002] Machine learning algorithms can automatically learn and discover underlying patterns and patterns from massive amounts of spectrum data, enabling dynamic monitoring, intelligent analysis, and optimized management of spectrum resources. By learning from historical spectrum data, machine learning models can predict spectrum usage trends and proactively allocate spectrum resources for new services or unexpected demands.

[0003] Existing machine learning algorithms lack the quality control and dynamic adaptability of data sources. Data formats, accuracy, and reliability may vary from source to source. Lack of effective data preprocessing and quality assessment mechanisms can lead to noise, bias, or missing values ​​in the initial data. Frequent parameter adjustments may be required for different types of spectrum data or application scenarios, and the configuration of distributed nodes across frequency bands may be inflexible, making it difficult to dynamically adapt to changes in spectrum resource usage patterns. This can lead to irrational data segmentation and affect the accuracy of subsequent analysis results. Summary of the Invention

[0004] The main purpose of the present invention is to provide a data processing method based on machine learning, aiming to solve the technical problems in the prior art.

[0005] The present invention proposes a data processing method based on machine learning, comprising:

[0006] Acquire multiple initial data based on multiple data sources;

[0007] Obtaining the data compactness of each initial data, and classifying the initial data based on the data compactness to obtain multiple data classification libraries;

[0008] Performing data division on each of the data classification libraries based on frequency band distributed nodes to obtain a plurality of data computing nodes, and performing data processing on each of the data computing nodes based on machine learning parallel data processing to obtain a first frequency band data set;

[0009] Acquire multiple spectrum occupancies according to the first frequency band dataset, and filter the first frequency band dataset based on the spectrum occupancies to obtain a second frequency band dataset;

[0010] Obtaining the number and distribution of spectrum holes in each frequency band data in the second frequency band data set, and obtaining spectrum data idleness based on the number and distribution of spectrum holes;

[0011] Filtering the second frequency band data set to obtain an idle frequency band set, and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data;

[0012] Obtaining frequency band resource features according to the fitting data through inverse mapping technology, and performing frequency band identification processing on the idle frequency band set based on the resource features.

[0013] As a preferred, the step of obtaining the data tightness of each initial data, classifying the initial data based on the data tightness, and obtaining a plurality of data classification libraries, comprises:

[0014] Obtaining frequency band feature information of each initial data, wherein the frequency band feature information includes signal strength and spectrum occupation amount;

[0015] Obtaining a correlation value corresponding to each initial data according to the signal strength and the spectrum occupation amount;

[0016] Taking the correlation value as the data tightness corresponding to each initial data;

[0017] Obtaining a plurality of preset threshold intervals according to the correlation value;

[0018] Establishing a data classification library according to the preset threshold interval;

[0019] Based on each preset threshold interval, all the correlation values are distinguished, and the initial data corresponding to each correlation value is classified into the corresponding data classification library.

[0020] As a preferred, the step of performing data division on each data classification library based on the frequency band distributed node to obtain a plurality of data computing nodes, and performing data processing on each data computing node based on machine learning parallel data processing to obtain a first frequency band data set, comprises:

[0021] Obtaining a plurality of frequency ranges according to the frequency band distributed node;

[0022] Obtaining the start node and the end node of each frequency range;

[0023] According to the start node and the end node corresponding to a plurality of frequency ranges, node allocation is performed on each data classification library to obtain a plurality of data computing nodes;

[0024] Performing data cleaning on each data computing node to obtain a cleaned data computing node;

[0025] Performing feature extraction on the data computing node to obtain key node feature information, wherein the key node feature information includes time slot distribution feature information, signal intermittent feature information, and spectrum occupation feature information;

[0026] normalizing the time slot distribution feature information to obtain a first feature;

[0027] normalizing the signal intermittent feature information to obtain a second feature;

[0028] normalizing the spectrum occupation feature information to obtain a third feature;

[0029] integrating the first feature, the second feature and the third feature to obtain integrated data of a corresponding data calculation node;

[0030] obtaining a first frequency band data set according to the integrated data of all data calculation nodes.

[0031] As a preferred, the step of obtaining a plurality of spectrum occupation rates according to the first frequency band data set, and screening the first frequency band data set based on the spectrum occupation rate to obtain a second frequency band data set, comprises:

[0032] obtaining a plurality of integrated data according to the first frequency band data set;

[0033] obtaining a plurality of first frequency bands based on the integrated data according to a frequency band distributed node;

[0034] obtaining a node duration of signal occupation in a specific time window for each first frequency band;

[0035] obtaining a total duration of signal occupation in a specific time window for all the first frequency bands;

[0036] obtaining a spectrum occupation rate according to the node duration and the total duration;

[0037] obtaining a preset spectrum occupation rate of the first frequency band data set;

[0038] judging whether the spectrum occupation rate conforms to the preset spectrum occupation rate;

[0039] if not, the first frequency band is excluded;

[0040] if yes, the first frequency band corresponding to the spectrum occupation rate is integrated to obtain a second frequency band data set.

[0041] As a preferred, the step of obtaining a spectrum data idle degree based on the number and distribution of spectrum holes comprises:

[0042] obtaining a plurality of frequency band data according to the second frequency band data set, and taking it as a second frequency band;

[0043] obtaining a number of spectrum holes according to the second frequency band;

[0044] Obtaining a distribution of each of the spectrum holes, and obtaining a start frequency and an end frequency of each spectrum hole based on the distribution of the spectrum holes;

[0045] Acquire a spectrum bandwidth corresponding to the spectrum hole according to the start frequency and the end frequency;

[0046] Obtaining the total spectrum bandwidth of each frequency band data;

[0047] The spectrum data idleness is calculated according to the spectrum bandwidth and the total spectrum bandwidth, wherein the calculation formula is:

[0048] ;

[0049] in, Indicates the idleness of spectrum data, Indicates the The spectrum bandwidth of the spectrum hole, represents the total spectrum bandwidth, and n represents the number of spectrum holes.

[0050] Preferably, the step of filtering the second frequency band data set to obtain an idle frequency band set, and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data includes:

[0051] Obtaining the preset spectrum data idleness;

[0052] determining whether a spectrum data idleness corresponding to each second frequency band in the second frequency band data set is greater than a preset spectrum data idleness;

[0053] If it is less than, the corresponding second frequency band will be eliminated;

[0054] If it is greater than or equal to, the corresponding second frequency band is used as an idle frequency band, and an idle frequency band set is generated;

[0055] Obtaining the frequency band mapping value of the idle frequency band in the high-dimensional feature space based on a preset kernel function;

[0056] Obtaining a feature vector of each idle frequency band in a high-dimensional space according to the frequency band mapping value;

[0057] Calculating the relationship value between the idle frequency band and other frequency band data by using a kernel function according to the feature vector;

[0058] Similar idle frequency bands are fitted based on the relationship values ​​to obtain fitting data.

[0059] Preferably, the step of obtaining frequency band resource characteristics according to the fitting data by using an inverse mapping technology includes:

[0060] Obtaining a plurality of fitting frequency bands according to the fitting data;

[0061] The spectrum feature of each fitting frequency band in the original space is obtained by inverse operation on the fitting frequency band through inverse mapping technology;

[0062] The number of dimension-reduced spectrum holes and the dimension-reduced spectrum occupancy rate are obtained according to the spectrum feature;

[0063] The correlation strength between each fitting frequency band is obtained according to the number of dimension-reduced spectrum holes and the dimension-reduced spectrum occupancy rate;

[0064] Each fitting frequency band is classified based on the correlation strength, and the classified fitting frequency band is marked to obtain a fitting frequency band set;

[0065] The frequency band resource feature is obtained according to the fitting frequency band set, wherein the frequency band resource feature includes frequency band hole density and frequency band occupancy rate.

[0066] The application also provides a data processing system based on machine learning, comprising:

[0067] The first acquisition module is configured to acquire a plurality of initial data based on a plurality of data sources;

[0068] The second acquisition module is configured to acquire the data tightness of each initial data, classify the initial data based on the data tightness, and obtain a plurality of data classification libraries;

[0069] The third acquisition module is configured to divide each data classification library into a plurality of data computing nodes based on frequency band distributed nodes, and process each data computing node based on machine learning parallel data processing to obtain a first frequency band data set;

[0070] The fourth acquisition module is configured to acquire a plurality of spectrum occupancy rates from the first frequency band data set, and filter the first frequency band data set based on the spectrum occupancy rates to obtain a second frequency band data set;

[0071] The fifth acquisition module is configured to acquire the number and distribution of spectrum holes of each frequency band data in the second frequency band data set, and acquire the spectrum data idle degree based on the number and distribution of spectrum holes;

[0072] The sixth acquisition module is configured to filter the second frequency band data set to obtain an idle frequency band set, and map the idle frequency band set to a high-dimensional feature space to obtain fitting data;

[0073] The processing module is configured to acquire the frequency band resource feature from the fitting data through inverse mapping technology, and perform frequency band identification processing on the idle frequency band set based on the resource feature.

[0074] The present invention also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned data processing method based on machine learning when executing the computer program.

[0075] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned data processing method based on machine learning.

[0076] The beneficial effects of this invention include: acquiring initial data from multiple sources, integrating information from different channels, avoiding the limitations of a single data source, providing a comprehensive and rich data foundation for spectrum resource analysis, and ensuring the reliability and integrity of analysis results. Through data density classification, frequency band distributed node division, and parallel processing, data processing speed and spectrum analysis accuracy are improved. For example, accurate calculation of spectrum occupancy and the number of spectrum holes enables refined management of spectrum resources and improves spectrum resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 Schematic diagram of a method flow according to an embodiment of the present invention.

[0078] Figure 2 FIG. 1 is a schematic diagram of the device structure according to an embodiment of the present invention.

[0079] Figure 3 This is a schematic diagram of the internal structure of a computer device according to an embodiment of the present application.

[0080] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0081] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0082] like Figures 1-3 As shown, the present application provides a data processing method based on machine learning, comprising:

[0083] S1. Acquire multiple initial data based on multiple data sources;

[0084] S2. Obtaining the data compactness of each initial data, and classifying the initial data based on the data compactness to obtain multiple data classification libraries;

[0085] S3. Dividing each of the data classification libraries based on frequency band distributed nodes to obtain multiple data computing nodes, and processing each of the data computing nodes based on machine learning parallel data processing to obtain a first frequency band data set;

[0086] S4, obtaining a plurality of spectrum occupancy rates according to the first frequency band data set, and screening the first frequency band data set based on the spectrum occupancy rates to obtain a second frequency band data set;

[0087] S5, obtaining a spectrum hole quantity and distribution of each frequency band data in the second frequency band data set, and obtaining a spectrum data idle degree based on the spectrum hole quantity and distribution;

[0088] S6, filtering the second frequency band data set to obtain an idle frequency band set, and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data;

[0089] S7, obtaining frequency band resource features according to the fitting data through inverse mapping technology, and performing frequency band identification processing on the idle frequency band set based on the resource features.

[0090] As described in steps S1-S7 above, a plurality of initial data is obtained based on a plurality of data sources (S1). Its beneficial effect lies in that it can comprehensively integrate spectrum data from different channels, ensuring the richness and diversity of data, such as collecting data from numerous sensors, base stations and various wireless devices, laying a solid foundation for subsequent accurate analysis. Then, the data tightness of each initial data is obtained, and based on this, a plurality of data classification libraries is obtained by classification (S2) so that spectrum data with similar features are aggregated, facilitating targeted analysis and improving processing efficiency, just like grouping data with similar signal strength and spectrum occupancy mode into a category, which can more efficiently analyze the spectrum characteristics of this category. It can solve the problem of disordered and disorganized spectrum data, which makes it difficult to conduct in-depth analysis, for example, disorganized data makes it extremely difficult to find spectrum rules, and classification can effectively improve it. The specific solution step is to first extract the frequency band feature information in the initial data, such as signal strength and spectrum occupancy, calculate the correlation value as the tightness measure, divide the threshold interval to establish the classification library, and classify the data into the corresponding library, for example, different classification libraries are divided according to the signal strength range;

[0091] Then, the data is divided based on the frequency band distributed nodes and processed in parallel to obtain a first frequency band data set (S3). Its beneficial effect is to accelerate spectrum data processing and respond to dynamic spectrum changes in a timely manner, such as quickly processing data in real-time spectrum monitoring scenarios to support dynamic spectrum allocation decisions; it can also deeply analyze the characteristics of each frequency band to provide a basis for fine management, such as optimizing communication parameters in different frequency bands. This can solve the problem of slow processing of massive spectrum data, which cannot meet the real-time needs, and overcome the performance bottleneck of traditional centralized processing in the face of large-scale data. The solution step is to determine the frequency range and node allocation according to the frequency band distributed nodes, clean the data to ensure quality, extract features and normalize the data, and then integrate the data, such as allocating data in different frequency band ranges to corresponding computing nodes for processing;

[0092] Afterwards, the spectrum occupancy is obtained according to the first frequency band data set and the second frequency band data set is screened (S4). This helps to accurately grasp the spectrum occupancy condition and provides a basis for reasonable allocation, such as screening the idle frequency band for new business development, while optimizing the spectrum configuration and improving the utilization rate. It solves the problem of unreasonable allocation caused by the inability to accurately know the spectrum occupancy, and corrects the decision-making errors caused by inaccurate or untimely calculation. The specific steps are to obtain the integrated data from the data set to determine the frequency band, count the signal occupancy time to calculate the occupancy rate, and compare it with the preset value to screen, such as calculating the occupancy rate within a certain time when analyzing a certain frequency band and judging whether it meets the requirements.

[0093] Subsequently, the number and distribution of spectrum holes in the second frequency band data set are obtained to calculate the spectrum data idle degree (S5). Its beneficial effect is to quantify the idle degree and intuitively reflect the spectrum idle condition to provide key indicators for spectrum sharing, such as preferentially allocating idle frequency bands according to the idle degree; it can also understand the fragmentation degree and potential of spectrum resources to guide resource integration, such as judging whether the distribution of spectrum holes is suitable for large-scale data transmission business. This solves the problem of unclear spectrum idle condition and difficult effective utilization, and overcomes the difficulty of being unable to plan and manage due to lack of understanding of the characteristics of holes. The solution steps are to obtain frequency band data, count the number, distribution and bandwidth of holes, and calculate the idle degree, such as analyzing the hole distribution of a certain frequency band and calculating its idle degree, and then filtering the second frequency band data set to obtain the idle frequency band set and map it to the high-dimensional feature space (S6). This can focus on high-idle-degree resources, reduce processing amount and improve efficiency, such as preferentially processing the most potential frequency band; after mapping, deep feature relationships can be mined to provide more information for optimization, such as discovering hidden relationships between hole distribution and interference. It solves the problem of complex spectrum data that is difficult to analyze directly for potential value and breaks through the limitations of low-dimensional space analysis. The solution steps are to set the idle degree threshold to filter the data to obtain the idle frequency band set, select the kernel function mapping, calculate the mapping value and feature vector, and then fit the data, such as setting the threshold to screen the idle frequency band and mapping it to the high-dimensional space;

[0094] Finally, the frequency band resource features are obtained through inverse mapping technology and the idle frequency band set is processed (S7). Its beneficial effect is to realize the association of high-dimensional data and original spectrum features, which is convenient for understanding and application, such as converting high-dimensional clustering results into actual spectrum hole density features; based on this, spectrum resources can be accurately classified, labeled and configured, such as allocating business according to hole density. This solves the problem of high-dimensional data being difficult to explain and apply to actual management, and improves the extensive management situation. The solution steps are to obtain the original spectrum features by inverse mapping, calculate the correlation strength, classify and label the resource features, and then process the idle frequency band set, such as obtaining the spectrum features by inverse mapping of the fitted frequency band and classifying and processing them accordingly.

[0095] In one embodiment, the step of obtaining the data tightness of each initial data, classifying the initial data based on the data tightness, and obtaining a plurality of data classification libraries comprises:

[0096] S201, obtaining frequency band feature information of each initial data, wherein the frequency band feature information comprises signal strength and spectrum occupation amount;

[0097] S202, obtaining a correlation value corresponding to each initial data according to the signal strength and the spectrum occupation amount;

[0098] S203, taking the correlation value as the data tightness corresponding to each initial data;

[0099] S204, obtaining a plurality of preset threshold intervals according to the correlation value;

[0100] S205, establishing a data classification library according to the preset threshold intervals;

[0101] S206, distinguishing all the correlation values based on each preset threshold interval, and classifying the initial data corresponding to each correlation value into the corresponding data classification library.

[0102] As described in steps S201-S206, the present application is of great significance by obtaining the frequency band feature information (including signal strength and spectrum occupation amount) of each initial data. Its beneficial effects are that the signal strength and the spectrum occupation amount are key attributes of spectrum data, and accurate acquisition of them can directly reflect the use state and quality of spectrum resources. For example, in the wireless network optimization scenario, accurate grasp of signal strength helps to determine the base station transmission power and antenna adjustment direction to ensure good signal coverage, reduce signal blind area, and improve user network experience; the statistics of spectrum occupation amount can determine the busy degree of each frequency band, and in spectrum resource allocation, priority is given to frequency bands with high spectrum occupation amount and critical business such as emergency communication to ensure smooth important communication. This step solves the problem of lack of key feature description of initial spectrum data state, avoiding inaccurate and incomplete analysis due to lack of necessary information in the subsequent analysis. Its solution steps are to first determine the relevant data fields or parameters in the data source, such as obtaining signal strength and spectrum occupation data from the base station monitoring system, and then pre-process the original data, including unifying the format, correcting the error value, and calibrating, so that the data from different sources are comparable.

[0103] Then, through step S202, the correlation value corresponding to each initial data is obtained according to the signal strength and spectrum occupancy. The beneficial effect of this is that the correlation value calculated by combining the two can more comprehensively and accurately characterize the internal relationship of the data and highlight the difference in characteristics. For example, when analyzing the spectrum data of a certain area, the correlation value can reflect the coordinated changes in signal strength and spectrum occupancy. Special cases of high signal strength accompanied by high spectrum occupancy or vice versa can be quantified, providing a basis for accurate classification. This step solves the problem that considering a single feature alone cannot accurately measure the overall characteristics of the spectrum data and that classification is difficult due to the lack of comprehensive indicators. The solution is to select a suitable calculation method, such as setting weights for the two according to business needs to calculate the weighted sum as the correlation value, and normalizing the signal strength and spectrum occupancy before calculation to prevent a certain feature from dominating the correlation value calculation.

[0104] Then, through step S203, the correlation value is used as the data closeness. Its beneficial effect is to provide a new comprehensive feature description method for spectrum data, which reflects the data correlation closeness from the overall perspective, and based on this, it can better capture the internal structure and pattern of spectrum data. For example, the data collected by two base stations in the same area at the same time period have strong synergy between signal strength and spectrum occupancy, and the data closeness is high. They are classified into one category to facilitate the analysis of the spectrum resource usage pattern in the area. This solves the problem of lack of a unified and effective feature measurement standard for spectrum data and the deviation of data analysis results due to improper use of indicators. The solution is to clarify that the correlation value is the data closeness measurement standard, record the closeness value of each initial data, and explain its value range and meaning, such as stipulating a value of 0-1, where 0 indicates weak correlation, 1 indicates strong correlation, and intermediate values ​​indicate different degrees of correlation.

[0105] Subsequently, through step S204, multiple preset threshold intervals are obtained according to the data density. The beneficial effect of this step is that the spectrum data can be reasonably divided, and different threshold intervals correspond to different density subsets, which helps to understand the distribution pattern of spectrum resources under different usage conditions, and provides clear and objective standards for classification to ensure the consistency of classification results. For example, in long-term spectrum monitoring, setting multiple threshold intervals can divide the data into different density categories, which is convenient for comparing the usage trends of spectrum resources in different periods. This step solves the problem of subjective arbitrariness and inaccurate and imprecise classification caused by the lack of clear standards for classification. The solution is to analyze the distribution of data density, such as drawing histograms and calculating statistical indicators, and then determine the number and range of threshold intervals in combination with spectrum management goals and needs. For example, when focusing on efficient utilization, narrow intervals can be set in high density areas.

[0106] Afterwards, a data classification library is established according to the preset threshold interval through S205. The beneficial effect is to realize the structured organization of spectrum data, improve the data management efficiency, and facilitate the retrieval and access. For example, data with high data closeness is classified into a specific classification library, representing an efficient utilization mode of spectrum resources, and when researching efficient utilization cases, data can be quickly located and obtained. This solves the problems of disordered spectrum data storage and management and low data analysis efficiency. The solution step is to create a classification library and mark it according to the threshold interval, define the storage structure and format, and ensure effective data storage and management, such as using numbers or names to identify the classification library, and determining the data field type, length, and storage method.

[0107] Finally, the initial data is classified based on the preset threshold interval through S206. The beneficial effect is to realize accurate classification, facilitate in-depth analysis and management, and discover hidden patterns and rules. For example, data of the same business type is classified into a category, which can be used to study the impact of the business on spectrum resource demand and optimize allocation strategies; independent analysis in the classification library can discover rules such as periodic changes in signal strength and spectrum occupancy, providing a basis for dynamic management. This step solves the problem of inaccurate classification of spectrum data, which leads to management decision errors. The solution step is to traverse the initial data, determine the interval to which the data closeness belongs to determine the classification library, accurately store the data, and update the classification library statistical information such as the number of data and the sum.

[0108] In one embodiment, the step of performing data division on each data classification library to obtain a plurality of data computing nodes based on the frequency band distributed nodes, and performing data processing on each data computing node based on machine learning parallel data processing to obtain a first frequency band data set, includes:

[0109] S301, obtaining a plurality of frequency ranges according to the frequency band distributed nodes;

[0110] S302, obtaining the starting node and the ending node of each frequency range;

[0111] S303, performing node allocation on each data classification library according to the starting node and the ending node corresponding to a plurality of frequency ranges to obtain a plurality of data computing nodes;

[0112] S304, performing data cleaning on each data computing node to obtain a cleaned data computing node;

[0113] S305, performing feature extraction on the data computing node to obtain key node feature information, wherein the key node feature information includes time slot distribution feature information, signal intermittent feature information, and spectrum occupancy feature information;

[0114] S306, performing normalization processing on the time slot distribution feature information to obtain a first feature;

[0115] S307, normalizing the signal intermittence feature information to obtain a second feature;

[0116] S308, normalizing the spectrum occupation feature information to obtain a third feature;

[0117] S309, integrating the first feature, the second feature and the third feature to obtain integrated data of the corresponding data computing node;

[0118] S310, obtaining a first frequency band data set according to the integrated data corresponding to all data computing nodes.

[0119] As described in steps S301-S310, the application obtains multiple frequency ranges according to frequency band distributed nodes. Its beneficial effects are that different frequency ranges have unique spectrum use characteristics and propagation laws, such as low frequency bands suitable for wide-area coverage of Internet of Things device communication, and high frequency bands conducive to high-speed data transmission services. By specifying the frequency range, accurate basis can be provided for spectrum resource management, improving resource utilization efficiency and avoiding service interference. For example, in the city intelligent transportation system, vehicle networking communication can be allocated in a specific low frequency band to ensure stable communication between vehicles. This step solves the problem of chaotic management of spectrum resources in the frequency dimension, avoiding the inability to take advantage of resource advantages due to lack of understanding of frequency characteristics. The solution step is to comprehensively scan and monitor the frequency band, collect signal strength, spectrum occupation and other information, and then divide the frequency range according to the business demand and spectrum management regulations, such as referring to the ITU standard to divide mobile communication, broadcast television and other business frequency bands.

[0120] Then, through step S302, the starting node and the ending node of each frequency range are obtained. This helps to accurately locate the frequency range and improve data processing accuracy, providing convenience for data management in distributed computing. For example, when calculating the spectrum occupation rate of a certain frequency band, clear boundaries can ensure accurate statistical data and provide a reliable basis for resource allocation. At the same time, it is convenient for computing nodes to quickly locate the data range and improve processing efficiency. This step solves the problem of data processing confusion and task allocation caused by unclear frequency range, avoiding data processing delay and resource waste caused by unclear range. The solution step is to determine the accurate frequency value in combination with the spectrum monitoring device accuracy and record and store related information, establish index or identification, such as recording the starting and ending frequencies and the metadata of the business type in the database.

[0121] Then, the data is allocated to the nodes according to the start and end node allocation data through S303. This can achieve reasonable allocation of spectrum data, take advantage of distributed computing, improve processing speed, and enhance system flexibility and scalability. For example, when processing massive spectrum data, different frequency band data is allocated to multiple nodes for parallel processing, shortening the processing time; when the business changes, it is convenient to adjust the node allocation strategy. This step solves the problem of uneven task allocation in distributed computing, avoiding data processing delay and resource waste caused by unreasonable allocation. The solution step is to develop an allocation strategy based on node performance and frequency range characteristics, such as allocating complex high-frequency band data to high-performance nodes, and then using a distributed computing framework to allocate data according to the strategy to ensure accuracy and integrity.

[0122] Subsequently, through S304, data cleaning can remove noise, outliers and error data, improve data quality, reduce computational complexity and storage requirements. For example, electromagnetic interference may cause errors in spectrum monitoring data, which can ensure the reliability of subsequent analysis after cleaning, while reducing storage and computing resource consumption. This step solves the problem of noise and error data affecting the accuracy of analysis results, avoiding resource waste and low processing efficiency caused by invalid data. The solution step is to set reasonable cleaning rules and thresholds, such as removing outliers according to the normal range of signal strength, using outlier detection and data smoothing algorithms, such as using the 3σ principle to detect signal strength outliers and using the moving average method to smooth the spectrum occupancy rate data. After S305, feature extraction can mine key information, reduce data dimensionality, and facilitate analysis and discovery of rules to provide decision support for management. For example, time slot distribution features can reflect the frequency spectrum time use rules, signal intermittence features can help understand signal stability, and spectrum occupancy features can reflect resource occupancy. By analyzing the time slot distribution, it can provide a basis for dynamic spectrum allocation. This step solves the problem of high-dimensional and complex spectrum data that is difficult to directly analyze and manage decision-making without targetedness. The solution step is to determine key features, such as determining the importance of time slot distribution and other features according to business needs, and then using appropriate algorithms to extract, such as using time domain analysis methods to extract time slot distribution features, using signal processing techniques to detect signal intermittence features, and using spectrum analysis algorithms to obtain spectrum occupancy features.

[0123] Next, through S306, the normalized time slot distribution feature information is obtained to make different node data comparable and improve the application effect of data analysis and machine learning algorithms. For example, when comparing the time slot usage of different base stations, normalization can accurately evaluate the relative efficiency, and help neural networks and other algorithms converge faster and improve accuracy. This step solves the problem of comparing time slot distribution features with different dimensions and ranges and affecting algorithm performance. The solution step is to choose the appropriate normalization method, such as min-max normalization or standardization, according to the characteristics of the features and the analysis needs, and then apply the method to process each node data and record the relevant parameters.

[0124] Then, in step S307, the second feature is obtained by normalizing the signal intermittence characteristic information. The effect is similar to the time slot distribution characteristic normalization, which can improve the signal stability evaluation accuracy and data mining efficiency. For example, when analyzing the signal quality of a large wireless network, the signal stability in different areas can be accurately compared after normalization, which helps the clustering algorithm to identify signal intermittence anomalies. This step solves the problem that signal intermittence characteristics cannot be compared and affect management decisions due to equipment and environmental factors. The solution step is to select a suitable normalization method according to the signal intermittence characteristics, such as selecting minimum-maximum normalization according to the signal intermittence proportion relationship, then processing the data and recording the parameters.

[0125] Then, in step S308, the third feature is obtained by normalizing the spectrum occupation characteristic information, which facilitates the comparison of spectrum occupation distribution and provides a basis for resource allocation and prediction, improving the accuracy of related models. For example, when comparing the spectrum occupation needs of different services, the differences can be clearly displayed after normalization, helping to predict spectrum demand. This step solves the problem that spectrum occupation characteristics cannot be compared and affect prediction and decision-making models due to differences in frequency bands and traffic volume. The solution step is to select a suitable normalization method, such as using minimum-maximum normalization to highlight the proportion difference of spectrum occupation, then processing the data and recording the parameters.

[0126] Finally, in step S310, the first, second, and third features are integrated to obtain corresponding node integrated data, and then the first frequency band data set is obtained by summarizing. This can effectively integrate the feature data processed by each node, form a data set that comprehensively reflects the spectrum status, and provide a complete data basis for subsequent spectrum resource evaluation, allocation, etc., such as providing data support for spectrum hole analysis. This step solves the problem of scattered data and lack of overall spectrum status data set. The solution step is to integrate the normalized feature data of each node according to the predetermined format and structure, ensuring data integrity and consistency, and forming a first frequency band data set that can be used for further analysis.

[0127] In one embodiment, the step of obtaining a plurality of spectrum occupation rates from the first frequency band data set, and screening the first frequency band data set based on the spectrum occupation rates to obtain a second frequency band data set, includes:

[0128] S401, obtaining a plurality of integrated data from the first frequency band data set;

[0129] S402, obtaining a plurality of first frequency bands from the integrated data based on frequency band distributed nodes;

[0130] S403, obtaining the node duration of signal occupation in a specific time window for each first frequency band;

[0131] S404, obtaining the total duration of signal occupation in a specific time window for all first frequency bands;

[0132] S405, obtaining a spectrum occupancy according to the node duration and the total duration;

[0133] S406, obtaining a preset spectrum occupancy of the first frequency band data set;

[0134] S407, judging whether the spectrum occupancy conforms to the preset spectrum occupancy;

[0135] If not, the first frequency band is excluded;

[0136] If yes, the first frequency band corresponding to the spectrum occupancy is integrated to obtain a second frequency band data set.

[0137] As described in steps S401-S407, the application obtains multiple integrated data from the first frequency band data set. The beneficial effect is that the integrated data summarizes the spectrum key information after the previous processing, which lays a solid foundation for comprehensive analysis of spectrum resource conditions and helps to mine potential correlations and rules. For example, when analyzing the spectrum use in urban areas, the integrated data covers the signal characteristics of different frequency bands in each time period and area. Through analysis, it can be found that the commercial area has high demand for high frequency bands during the day, and the residential area is more active in low frequency bands at night. This provides data support for spectrum management. This step solves the problems of scattered, disordered spectrum data, which is difficult to analyze comprehensively and cannot effectively utilize potential information. The solution steps are to analyze the data set structure and storage method first, to clarify the data correlation, and then to extract and integrate data according to the demand using SQL or data processing framework API, such as selecting related field combinations from the database table storing frequency band data to form the required format for analysis.

[0138] Then, according to step S402, multiple first frequency bands are obtained from the integrated data based on the frequency band distributed nodes. This is helpful for frequency spectrum resource management and planning at the frequency band level, allocating services according to frequency band characteristics, improving efficiency and reducing interference, and also adapting to dynamic changes in spectrum. For example, according to the frequency band distributed node configuration, the spectrum is divided into first frequency bands suitable for mobile communication, broadcast television and other services. When 5G services develop, suitable high frequency band first frequency bands can be quickly found from the integrated data and resources can be allocated. This step solves the problems of lack of effective division and management of frequency bands and inability to respond to changes in spectrum in a timely manner. The solution steps are to extract first frequency band data from the integrated data using data filtering conditions (such as frequency range, service type identifier) according to node configuration and division rules, and then verify and arrange the data, add metadata to ensure data completeness and accuracy, such as adding bandwidth, number and other information to mobile communication frequency bands.

[0139] Then, through the S403 step, the node duration of each first frequency band in a specific time window is obtained. This can accurately quantify the frequency band usage, evaluate the utilization and busy degree, provide basis for resource allocation optimization, and also analyze the usage regularity trend. Taking the traffic monitoring frequency band as an example, through long-term monitoring of node duration in different time periods, it is found that the early peak duration is longer, the late peak is second, and the night is shorter, which provides a reference for dynamic spectrum allocation, such as preferentially guaranteeing monitoring data transmission in peak period. This step solves the problems of being unable to accurately grasp the time use of the spectrum and the lack of timeliness and pertinence of management decision-making. The solution step is to first set a time window according to the business requirements (such as 1 minute for traffic monitoring), then filter the occupation records in the window from the first frequency band data, calculate the node duration by calculating the difference between the start and end time of the occupation or obtaining the occupation duration field, and calculate the signal interruption according to the rules, such as accumulating the continuous occupation time period.

[0140] Subsequently, through the S404 step, the total duration of signal occupation of all first frequency bands in a specific time window is obtained. This can grasp the spectrum use intensity and busy degree from the whole, serve as a basis for evaluating the utilization, and also facilitate comprehensive evaluation and comparison of multiple frequency bands. For example, by comparing the total duration in different time periods, it is found that the total duration of spectrum use on holidays is lower than that on weekdays, indicating that the business requirements change and resource allocation needs to be adjusted. The total duration of different frequency bands can also be compared to analyze their position in spectrum resource use, such as the total duration of high frequency bands being short, which needs to study methods to improve utilization. This step solves the problem of being unable to globally evaluate the spectrum use intensity and the lack of global view in allocation decision-making. The solution step is to traverse the first frequency band occupation records, and accumulate the node duration to obtain the total duration by using a loop or a database aggregation function. The total duration and related metadata are recorded, such as being stored in a database table associated with the time window and the calculation time, for easy query and analysis.

[0141] After that, through the S405 step, the spectrum occupation rate is obtained according to the node duration and the total duration. The spectrum occupation rate reflects the frequency band occupation degree in the form of intuitive percentage, which is convenient for comparing different frequency bands and providing quantitative basis for resource management. It can also establish a dynamic monitoring mechanism. For example, if the spectrum occupation rate of a mobile communication frequency band is high, the quality of service needs to be guaranteed or sharing needs to be considered; if the spectrum occupation rate of a broadcast television frequency band is low, it can be used as a resource for new business expansion. By regularly calculating the occupation rate in different time windows, abnormal fluctuations can be found in time, such as a sudden increase in the occupation rate of a certain frequency band, which may be caused by a new business or an interference source. This step solves the problem of being unable to intuitively compare the occupation degree of frequency bands and the lack of scientific accuracy in management decision-making. The solution step is to calculate the occupation rate of each first frequency band using a formula to ensure consistency, check and correct abnormal values, and record the occupation rate associated with the frequency band and the time window, such as storing it in a database table for easy query and processing.

[0142] Next, through the S406 step, the preset spectrum occupancy rate of the first frequency band data set is obtained. The preset value provides a clear standard for spectrum screening management, which is determined according to management goals, business needs and historical conditions, and guarantees fair and reasonable management. For example, for medical communication services with high real-time requirements, the preset value is low; for non-critical Internet of Things services, the preset value can be higher. This avoids the problems of arbitrariness and blindness, uneven distribution and low efficiency caused by the lack of standards in spectrum management. The solution step is to determine the preset principle by analyzing the management goals and business needs, such as ensuring the quality of service of critical services or improving overall utilization, and then determine the specific value by combining historical data statistical analysis, such as determining the appropriate value by referring to historical occupancy rate statistical indicators.

[0143] Finally, through the S407 step, it is judged whether the spectrum occupancy rate meets the preset spectrum occupancy rate and the frequency band is screened. The ones that do not meet the requirements are excluded, which can exclude unreasonable frequency bands and improve management efficiency; the ones that meet the requirements are integrated into the second frequency band data set, which provides a basis for subsequent decision-making. For example, frequency bands with excessively high spectrum occupancy rates that cause congestion or excessively low spectrum occupancy rates that waste resources are excluded, and after the integration of frequency bands that meet the requirements, different services can be reasonably allocated frequency bands according to their characteristics, optimizing resource allocation. This step solves the problems of unreasonable occupation of frequency bands in spectrum resources affecting management efficiency and unfair resource allocation. The solution step is to compare each first frequency band occupancy rate with the preset value, and if it does not meet the requirements, it is excluded, and if it meets the requirements, it is integrated according to the rules, such as frequency band characteristics, occupancy rate, etc., to form a second frequency band data set.

[0144] In one embodiment, the step of obtaining the number and distribution of spectrum holes in each frequency band data in the second frequency band data set, and obtaining the spectrum data idle degree based on the number and distribution of spectrum holes, comprises:

[0145] S501, obtaining a plurality of frequency band data from the second frequency band data set and taking it as a second frequency band;

[0146] S502, obtaining the number of spectrum holes according to the second frequency band;

[0147] S503, obtaining the distribution of each spectrum hole, and obtaining the start frequency and end frequency of each spectrum hole based on the distribution of the spectrum hole;

[0148] S504, obtaining the spectrum bandwidth corresponding to the spectrum hole according to the start frequency and end frequency;

[0149] S505, obtaining the total spectrum bandwidth of each frequency band data;

[0150] S506, calculating the spectrum data idle degree according to the spectrum bandwidth and the total spectrum bandwidth, wherein the calculation formula is:

[0151] ;

[0152] wherein, represents the spectrum data idle degree, represents the spectrum bandwidth of the th spectrum hole, represents the total spectrum bandwidth, and n represents the number of spectrum holes.

[0153] As described in steps S501-S506, the application obtains the frequency band data from the second frequency band data set as the second frequency band. Its beneficial effect lies in focusing on the high-potential frequency band after the preliminary screening, improving the pertinence and effectiveness of spectrum resource management, and facilitating hierarchical management and fine analysis. For example, when expanding new 5G millimeter wave business, by focusing on the second frequency band data, the frequency band suitable for millimeter wave communication can be quickly located, the search range is reduced, the resource allocation efficiency is improved, and special spectrum sharing strategies can be formulated for the second frequency band. This step solves the problems of lack of focus in spectrum resource analysis and lack of pertinence in management decision-making. The solution step is to analyze the data set structure format, clarify the frequency band data storage method and field information, use data reading and processing tools to extract frequency band data according to the rules, and arrange them into a format convenient for subsequent analysis, such as reading the data into a list, each element containing frequency band number, frequency range, etc.

[0154] Next, through step S502, the number of spectrum holes of the second frequency band is obtained. This can quantify the size of idle resources in the frequency band, providing a key basis for spectrum resource allocation and sharing. For example, when planning the deployment of large-scale Internet of Things devices, clearly defining the number of spectrum holes can reasonably arrange the number and layout of devices, ensure stable communication of devices, and avoid waste of resources. This step solves the problems of being unable to grasp the size of idle spectrum resources and unreasonable allocation. The solution step is to use spectrum hole detection technology, such as setting a signal strength threshold (-110 dBm), which is considered a hole if it is lower than this value, and using a counter to count the number of holes and store them for subsequent idle degree calculation.

[0155] Then, through step S503, the spectrum hole distribution and the start and end frequencies are obtained. This can accurately describe the range of the hole position, and have a good understanding of the idle structure of the spectrum resource, which helps to find the fragmentation problem. For example, when designing wireless communication frequency planning, according to the hole distribution, the communication frequency can be accurately placed to avoid interference, and if the holes are scattered, integration strategies can be developed. This step solves the problem of inaccurate understanding of the range of the hole position and the inability to solve the fragmentation problem. The solution step is to determine the start and end frequencies based on hole detection using spectrum monitoring equipment with high accuracy, and store them in a suitable data structure (such as a dictionary), with the frequency band number as the key and the hole frequency list as the value.

[0156] Subsequently, the spectral hole bandwidth is calculated through S504. This can quantify the actual size of the hole, which is of great significance for evaluating the idle capacity and helps to optimize resource matching. For example, high-definition video transmission services require large bandwidth, while sensor services require small bandwidth. According to the allocation of hole bandwidth, the efficiency of resource utilization can be improved. This step solves the problem of being unable to measure the available capacity of the hole and wasting resources or abnormal services. The solution step is to calculate the formula "end frequency - start frequency" to ensure consistent frequency units, associate the bandwidth with the hole information storage, and facilitate query retrieval, such as storing in a database table and building an index. Then, in S505, the total spectral bandwidth of the frequency band is obtained. This can overall grasp the size of the frequency band resources, which is an important basis for evaluating the utilization and planning of services. For example, when comparing the utilization efficiency of 5G communication resources in different frequency bands, the total spectral bandwidth combined with the hole information can determine the degree of utilization, providing a basis for reasonable allocation of services. This step solves the problem of lacking overall understanding of the size of the frequency band resources and being unable to accurately evaluate the utilization rate. The solution step is to directly obtain the total spectral bandwidth if the data directly provides it, otherwise to calculate it according to the frequency range, associate the total spectral bandwidth with other information of the frequency band, and build a frequency band information dataset.

[0157] Finally, through S506, the spectral data idle degree is calculated. This indicator comprehensively reflects the idle degree of spectral resources, providing a basis for dynamic management and optimization. For example, in spectrum trading, the idle degree can measure the value of the spectrum, assisting in pricing, and it can also dynamically adjust resource allocation according to the idle degree. This step solves the problem of lacking comprehensive indicators for evaluating the idle degree and decision-making errors. The solution step is to calculate the idle degree according to the formula, ensure the accuracy of the relevant parameters, associate the idle degree with the frequency band information storage, and form an idle degree dataset for query analysis and decision-making, such as sorting and filtering frequency band management allocation according to the idle degree.

[0158] In one embodiment, the step of filtering the second frequency band dataset to obtain an idle frequency band set and mapping the idle frequency band set to a high-dimensional feature space to obtain fitted data comprises:

[0159] S601, obtaining a preset spectral data idle degree;

[0160] S602, judging whether the spectral data idle degree corresponding to each second frequency band in the second frequency band dataset is greater than the preset spectral data idle degree;

[0161] If less, the corresponding second frequency band is removed;

[0162] If greater than or equal to, the corresponding second frequency band is taken as an idle frequency band, and an idle frequency band set is generated;

[0163] S603, obtaining a frequency band mapping value of the idle frequency band in the high-dimensional feature space based on a preset kernel function;

[0164] S604, obtaining a feature vector of each idle frequency band in a high-dimensional space according to the frequency band mapping value;

[0165] S605, calculating a relationship value of the idle frequency band and other frequency band data by a kernel function according to the feature vector;

[0166] S606, fitting similar idle frequency bands based on the relationship value to obtain fitting data.

[0167] As described in steps S601-S606, the present application obtains the idle degree of the preset frequency spectrum data. Its beneficial effect lies in providing clear standards and basis for frequency spectrum resource screening, which can be determined according to management goals, business needs and historical conditions, and flexible resource allocation can be realized. For example, when the spectrum is tight and the key business needs to be guaranteed, a low preset value is set to ensure high utilization rate; when the demand is large at the initial stage of new business development, the preset value is increased to quickly obtain idle resources. This solves the problems of no unified standard for screening and inability to adapt to different needs. The solution step is to first analyze the management goals to determine the principles, such as taking the utilization rate as the target to refer to the occupancy rate distribution, and then combining historical data statistics (mean, median, etc.) to determine specific values, such as setting the idle degree standard according to the historical occupancy rate range.

[0168] Next, through step S602, the relationship between the idle degree of the second frequency band and the preset value is judged to screen the frequency band. Accurate screening of idle frequency bands improves accuracy and effectiveness, reduces processing amount, and discovers potential value. For example, when processing massive data, low idle degree frequency bands are removed to reduce complexity; when exploring 5G millimeter wave applications, high idle degree frequency bands are screened to assist technology development. This step solves the problems of inaccurate identification of idle valuable frequency bands and inaccurate screening. The solution step is to obtain the idle degree (query or calculation) of the second frequency band, compare it with the preset value, mark the frequency band state, and ensure the correctness of data types and comparison logic.

[0169] Then, through step S603, the high-dimensional mapping value of the idle frequency band is obtained based on the preset kernel function. The frequency spectrum data is converted to a high-dimensional space to mine potential feature relationships, provide a new perspective method, and improve classification recognition accuracy. For example, the low-dimensional chaotic frequency spectrum hole distribution may present a regularity in high-dimensional space, which helps to distinguish different business frequency band characteristics. This solves the problems of insufficient low-dimensional features and lack of depth accuracy in decision-making. The solution step is to select a suitable kernel function (such as linear or Gaussian kernel function), set parameters (such as Gaussian kernel bandwidth) according to frequency spectrum characteristics and purposes, and calculate high-dimensional mapping values with idle frequency band data. The parameters can be determined by cross-validation.

[0170] Then, the high-dimensional feature vector is obtained from the frequency band mapping value through S604. Providing a structured representation for high-dimensional spectrum data helps to discover the internal structural pattern and facilitate subsequent analysis and processing. For example, in cluster analysis, the feature vector is input to improve the efficiency and accuracy of the algorithm, and the analysis of the vector relationship can discover the clustering structure of the spectrum resource. This step solves the problem of lack of structured representation of high-dimensional data and inability to accurately identify structural patterns. The solution step is to arrange the mapping values into a feature vector in a certain order, determine the order according to the spectrum characteristics, ensure the same dimension and meaning, and perform standardization or normalization processing if necessary.

[0171] After that, the frequency band relationship value is calculated by using the kernel function through S605. In-depth understanding of the mutual relationship of spectrum resources provides the basis for collaborative management and optimization, and discovers potential collaborative and complementary relationships. For example, when spectrum sharing, the idle frequency bands with low correlation are selected to reduce interference; the possibility of collaborative work between frequency bands can also be discovered to provide ideas for business expansion. This solves the problem of inability to accurately measure the relationship of spectrum resources and lack of collaborative optimization in decision-making. The solution step is to select an appropriate kernel function (which can be the same as or different from the mapping kernel function), calculate the relationship value using the feature vector, ensure the correct format and reasonable parameters, and appropriately process the relationship value.

[0172] Finally, similar idle frequency bands are fitted based on the relationship value through S606 to obtain fitting data. Integrating abstract similar frequency bands, discovering common patterns and rules, simplifying data and retaining key information, and improving management efficiency and decision-making scientificity. For example, when analyzing the distribution of spectrum holes, the model representing the commonness obtained by fitting helps to grasp the overall characteristics; when allocating spectrum, fitting data is used for quick decision-making, such as batch allocation after determining that the frequency band is suitable for the business. This step solves the problem of too complex spectrum resource data for management and decision-making and lack of overall understanding in analysis. The solution step is to select a fitting method (such as linear fitting or clustering algorithm) according to the relationship value, determine it according to the spectrum characteristics and purpose, set the parameters, perform fitting, evaluate and verify the results, and adjust the method or parameters for re-fitting if the results are not ideal.

[0173] In one embodiment, the step of obtaining the frequency band resource characteristics from the fitting data by inverse mapping technology comprises:

[0174] S701, obtaining a plurality of fitting frequency bands from the fitting data;

[0175] S702, obtaining the spectrum characteristics of each fitting frequency band in the original space by inverse mapping technology;

[0176] S703, obtaining the number of reduced-dimensional spectrum holes and the reduced-dimensional spectrum occupancy rate according to the spectrum characteristics;

[0177] S704, obtaining the correlation strength between each fitting frequency band according to the number of reduced-dimensional spectrum holes and the reduced-dimensional spectrum occupancy rate;

[0178] S705, classifying each fitting frequency band based on the correlation strength, and marking the classified fitting frequency bands to obtain a fitting frequency band set;

[0179] S706, obtaining a frequency band resource feature according to the fitting frequency band set, wherein the frequency band resource feature comprises a frequency band hole density and a frequency band occupancy rate.

[0180] As described in the above steps S701-S706, the present application obtains a plurality of fitting frequency bands from fitting data. The beneficial effect is that representative spectrum resource units can be extracted, which facilitates hierarchical management and fine analysis, and strategies are formulated according to characteristics, such as fitting frequency bands with strong signals and regular holes being preferentially allocated to high-quality services. This solves the problems of difficult management of fitting data and lack of decision-making pertinence. The solution steps are to first determine the fitting frequency band range characteristics in the clustering results, such as according to the clustering center, then extract by selecting a suitable algorithm (such as a distance or threshold method), and finally organize records, such as dividing fitting frequency bands according to the center and threshold after clustering of spectrum data.

[0181] Then, through the S702 step, the original space spectrum features of the fitting frequency bands are obtained by inverse mapping. The restored information is combined with the actual situation, which is helpful for verifying the high-dimensional analysis results. For example, the frequency range of the frequency band is determined according to the spectrum features in the original space, which is used for spectrum allocation and interference analysis; the special clustering mode in high dimension can be explained by inverse mapping. This step solves the problems of difficult management of high-dimensional data and unreliable decision-making. The solution steps are to first determine the inverse mapping technology and inverse operation method (such as the inverse operation corresponding to the Gaussian kernel function), then substitute the fitting frequency band data for calculation, check the reasonableness of the results, and store them, such as converting the fitting frequency band data into original space spectrum features by using the inverse operation of the Gaussian kernel.

[0182] Then, through the S703 step, the dimension-reduced spectrum hole number and occupancy rate are obtained according to the spectrum features. Reducing the dimension of data simplifies the description, which is beneficial for quickly evaluating the spectrum state, such as understanding the regional spectrum utilization when patrolling; it is also helpful for improving the analysis efficiency, such as sorting and screening according to the number of holes when searching for idle spectrum. This step solves the problems of high dimension of original spectrum features and complex calculation and difficult fast evaluation. The solution steps are to select a suitable dimension reduction method (such as PCA or LDA), substitute the original spectrum features for calculation (assuming the number of PCA principal components), verify the accuracy of the results, and check the accuracy by using PCA dimension reduction and comparing with the original data.

[0183] Then, the fitting frequency band association strength is fitted by S704, so as to understand the relationship between frequency bands and provide a basis for collaborative management, such as allocating frequency bands with small interference when spectrum sharing; and to find collaborative complementary relationship. This step solves the problem of difficult to measure frequency band relationship and lack of collaborative optimization in decision-making. The association strength is obtained by substituting the dimensionality reduction features, and the format is correct, and the normalized results are explained, such as calculating by using the correlation coefficient method and mapping the results to the [0, 1] interval.

[0184] Then, the fitting frequency band set is obtained by S705, and the fitting frequency band is classified and labeled according to the association strength. The spectrum resource classification management is realized, and the efficiency is increased and targeted, such as adopting a unified management strategy for similar frequency bands; and it is also helpful to find hidden mode rules, such as the periodic change of the spectrum hole of a certain type of frequency band, which provides a basis for dynamic management. This step solves the problem of lack of classification method and unscientific decision-making. The solution step is to select a suitable classification algorithm (such as K-means or hierarchical clustering), set parameters according to the algorithm requirements, evaluate and verify the effect (use indicators such as the silhouette coefficient), and adjust and reclassify if the effect is not ideal.

[0185] Then, the frequency band resource features are obtained according to the fitting frequency band set by S706. The key features of the spectrum are grasped as a whole, and a basis is provided for macro management decision-making, such as allocating according to the hole density and occupancy rate when planning spectrum resources; and it is also helpful to evaluate and compare spectrum resources at different levels, such as comparing regional hole density to help new business site selection. This step solves the problem of lack of overall feature description and lack of comprehensive long-term decision-making. The solution step is to extract the fitting frequency band information to calculate the hole density (hole number to bandwidth ratio) and occupancy rate (occupied time to total time ratio), and to summarize and arrange the feature data set, such as storing it in a database table for query and analysis after calculation. Finally, the idle frequency band is identified based on the resource features. The type and characteristics of the idle frequency band are accurately identified to ensure precise allocation and efficient use, such as allocating high-potential idle frequency bands to high-definition video transmission businesses; and to optimize the management process and improve efficiency, such as reducing monitoring investment for low-quality idle frequency bands. This step solves the problem of being unable to distinguish between idle frequency bands and low management efficiency. The solution step is to determine the type of idle frequency band according to the resource features (hole density, occupancy rate, etc.).

[0186] The application also provides a data processing system based on machine learning, comprising:

[0187] A first acquisition module 1 is configured to acquire a plurality of initial data based on a plurality of data sources;

[0188] A second acquisition module 2 is configured to acquire the data tightness of each initial data, classify the initial data based on the data tightness, and obtain a plurality of data classification libraries;

[0189] The third obtaining module 3 is configured to obtain a plurality of data computing nodes by dividing each of the data classification libraries based on the frequency band distributed nodes, and perform data processing on each of the data computing nodes based on the machine learning parallel data processing to obtain a first frequency band data set;

[0190] The fourth obtaining module 4 is configured to obtain a plurality of spectrum occupancy rates based on the first frequency band data set, and perform screening on the first frequency band data set based on the spectrum occupancy rates to obtain a second frequency band data set;

[0191] The fifth obtaining module 5 is configured to obtain a spectrum hole quantity and a distribution of each frequency band data in the second frequency band data set, and obtain a spectrum data idle degree based on the spectrum hole quantity and the distribution;

[0192] The sixth obtaining module 6 is configured to filter the second frequency band data set to obtain an idle frequency band set, and map the idle frequency band set to a high-dimensional feature space to obtain fitting data;

[0193] The processing module 7 is configured to obtain frequency band resource features based on the fitting data by an inverse mapping technology, and perform frequency band identification processing on the idle frequency band set based on the resource features.

[0194] It should be noted that each module and unit in the machine learning-based data processing system corresponds to each step in the machine learning-based data processing method.

[0195] As shown in Figure 3 , the present application also provides a computer device, which can be a server, and the internal structure thereof can be as shown in Figure 3 . The computer device comprises a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide calculation and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store all data required by the process of the machine learning-based data processing method. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the machine learning-based data processing method.

[0196] Those skilled in the art can understand that Figure 3 the structure shown in the description is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied.

[0197] An embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement any one of the above data processing methods based on machine learning.

[0198] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the computer program can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium provided by the present application and used in the embodiments can include non-volatile and / or volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.

[0199] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that processes, devices, articles or methods including a list of elements do not necessarily include those elements only, but can include other elements not expressly listed or inherent to such processes, devices, articles or methods. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, device, article or method including the element.

[0200] The above description is only the preferred embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.

Claims

1. A data processing method based on machine learning, characterized in that: include: Acquire multiple initial data based on multiple data sources; Obtaining the data compactness of each initial data, and classifying the initial data based on the data compactness to obtain multiple data classification libraries; Performing data division on each of the data classification libraries based on frequency band distributed nodes to obtain a plurality of data computing nodes, and performing data processing on each of the data computing nodes based on machine learning parallel data processing to obtain a first frequency band data set; Acquire multiple spectrum occupancies according to the first frequency band dataset, and filter the first frequency band dataset based on the spectrum occupancies to obtain a second frequency band dataset; Obtaining the number and distribution of spectrum holes in each frequency band data in the second frequency band data set, and obtaining spectrum data idleness based on the number and distribution of spectrum holes; filtering the second frequency band data set to obtain an idle frequency band set, and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data; The frequency band resource characteristics are obtained according to the fitting data through an inverse mapping technology, and frequency band identification processing is performed on the idle frequency band set based on the resource characteristics.

2. The data processing method based on machine learning according to claim 1, characterized in that: The step of obtaining the data compactness of each initial data, classifying the initial data based on the data compactness, and obtaining multiple data classification libraries includes: Acquire frequency band characteristic information of each of the initial data, wherein the frequency band characteristic information includes signal strength and spectrum occupancy; Acquire a correlation value corresponding to each of the initial data according to the signal strength and spectrum occupancy; Using the correlation value as the data compactness corresponding to each of the initial data; Acquire multiple preset threshold intervals according to the correlation value; Establishing a data classification library according to the preset threshold interval; All the correlation values ​​are distinguished based on each preset threshold interval, and the initial data corresponding to each of the correlation values ​​is classified into a corresponding data classification library.

3. The data processing method based on machine learning according to claim 1, characterized in that: The step of dividing each of the data classification libraries based on frequency band distributed nodes to obtain a plurality of data computing nodes, and performing data processing on each data computing node based on machine learning parallel data processing to obtain a first frequency band data set includes: Acquire starting nodes and ending nodes of multiple frequency ranges according to the frequency band distributed nodes; Allocating nodes to each of the data classification libraries according to the starting nodes and the ending nodes corresponding to the multiple frequency ranges to obtain multiple data calculation nodes; Performing data cleaning on each of the data computing nodes to obtain a cleaned data computing node; Performing feature extraction on the data computing node to obtain key node feature information, wherein the key node feature information includes time slot distribution feature information, signal intermittent feature information, and spectrum occupancy feature information; Normalizing the time slot distribution feature information to obtain a first feature; Normalizing the signal intermittent feature information to obtain the second feature; Normalizing the spectrum occupancy feature information to obtain the third feature; Data integration is performed based on the first feature, the second feature, and the third feature to obtain a first frequency band data set.

4. The data processing method based on machine learning according to claim 1, characterized in that: The step of obtaining multiple spectrum occupancies according to the first frequency band data set, and filtering the first frequency band data set based on the spectrum occupancies to obtain the second frequency band data set includes: Obtaining a plurality of integrated data according to the first frequency band data set; Acquiring a plurality of first frequency bands based on the integrated data at the frequency band distributed node; Obtaining the node duration occupied by the signal in a specific time window for each first frequency band; Obtaining the total duration of signal occupancy of all first frequency bands within a specific time window; Obtaining spectrum occupancy according to the node duration and the total duration; Obtaining a preset spectrum occupancy rate of the first frequency band data set; Determining whether the spectrum occupancy rate meets a preset spectrum occupancy rate; If it does not meet the requirements, the first frequency band will be eliminated; If it is consistent, the first frequency band corresponding to the spectrum occupancy rate is integrated to obtain a second frequency band data set.

5. The data processing method based on machine learning according to claim 1, characterized in that: The step of obtaining the number and distribution of spectrum holes in each frequency band data in the second frequency band data set, and obtaining the spectrum data idleness based on the number and distribution of spectrum holes includes: Acquire multiple frequency band data according to the second frequency band data set, and use the data as the second frequency band; Obtaining the number of spectrum holes according to the second frequency band; Obtaining a distribution of each of the spectrum holes, and obtaining a start frequency and an end frequency of each spectrum hole based on the distribution of the spectrum holes; Acquire a spectrum bandwidth corresponding to the spectrum hole according to the starting frequency and the ending frequency; Obtaining the total spectrum bandwidth of each frequency band data; The spectrum data idleness is calculated according to the spectrum bandwidth and the total spectrum bandwidth.

6. The data processing method based on machine learning according to claim 1, characterized in that: The step of filtering the second frequency band data set to obtain an idle frequency band set, and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data includes: Obtaining the preset spectrum data idleness; determining whether a spectrum data idleness corresponding to each second frequency band in the second frequency band data set is greater than a preset spectrum data idleness; If it is less than, the corresponding second frequency band will be eliminated; If it is greater than or equal to, the corresponding second frequency band is used as an idle frequency band, and an idle frequency band set is generated; Obtaining the frequency band mapping value of the idle frequency band in the high-dimensional feature space based on a preset kernel function; Obtaining a feature vector of each idle frequency band in a high-dimensional space according to the frequency band mapping value; Calculating the relationship value between the idle frequency band and other frequency band data by using a kernel function according to the feature vector; Similar idle frequency bands are fitted based on the relationship values ​​to obtain fitting data.

7. The data processing method based on machine learning according to claim 1, characterized in that: The step of obtaining frequency band resource characteristics according to the fitting data using an inverse mapping technology includes: Obtaining a plurality of fitting frequency bands according to the fitting data; The inverse operation of the fitting frequency band is performed through the inverse mapping technology to obtain the spectrum characteristics of each fitting frequency band in the original space; Obtaining the number of holes in the dimension-reduced spectrum and the occupancy rate of the dimension-reduced spectrum according to the spectrum characteristics; Obtaining the correlation strength between each fitting frequency band according to the number of holes in the dimension-reduced spectrum and the occupancy rate of the dimension-reduced spectrum; Classify each fitting frequency band based on the correlation strength, and mark the classified fitting frequency bands to obtain a fitting frequency band set; Frequency band resource characteristics are acquired according to the fitted frequency band set, wherein the frequency band resource characteristics include frequency band hole density and frequency band occupancy.

8. A data processing system based on machine learning, characterized in that: include: A first acquisition module is used to acquire multiple initial data based on multiple data sources; A second acquisition module is configured to acquire a data compactness of each initial data, and classify the initial data based on the data compactness to obtain a plurality of data classification libraries; A third acquisition module is configured to divide each of the data classification libraries into multiple data computing nodes based on the frequency band distributed nodes, and perform data processing on each data computing node based on machine learning parallel data processing to obtain a first frequency band data set; a fourth acquisition module, configured to acquire multiple spectrum occupancies according to the first frequency band dataset, and filter the first frequency band dataset based on the spectrum occupancies to obtain a second frequency band dataset; a fifth acquisition module, configured to acquire the number and distribution of spectrum holes in each frequency band data in the second frequency band data set, and acquire spectrum data idleness based on the number and distribution of spectrum holes; a sixth acquisition module, configured to filter the second frequency band data set to obtain an idle frequency band set, and map the idle frequency band set to a high-dimensional feature space to obtain fitting data; The processing module is used to obtain frequency band resource characteristics according to the fitting data through an inverse mapping technology, and perform frequency band identification processing on the idle frequency band set based on the resource characteristics.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for determining spectrum resource utilization degree

    CN105992215A

  • Spectrum channel clustering method based on extreme learning machine

    CN111353530A