Data processing method and system based on machine learning and computer equipment

Through the data processing method based on machine learning, the data tightness classification and frequency band distributed node division are used to solve the problem of poor data source quality control and dynamic adaptability in spectrum resource management, and the refined management and efficient utilization of spectrum resources are realized.

CN120337040AActive Publication Date: 2025-07-18HANGZHOU IDEACOME INTERNET FINANCIAL CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510813440.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing machine learning algorithms have insufficient quality control and poor dynamic adaptability in spectrum resource management, resulting in unreasonable data division and affecting the accuracy and flexibility of analysis results.

Method used

By acquiring initial data based on multiple data sources, classification according to the degree of data tightness, using frequency band distributed nodes to divide and process in parallel, obtaining spectrum occupancy and number of holes, filtering idle frequency bands and mapping to high-dimensional feature space, and inversely mapping to obtain frequency band resource features.

Benefits of technology

It realizes refined management of spectrum resources, improves data processing speed and analysis accuracy, ensures the reliability and integrity of analysis results, and optimizes spectrum resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337040A_ABST
    Figure CN120337040A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data processing method and system based on machine learning and computer equipment. Comprising the steps of obtaining multiple pieces of initial data based on multiple data sources; obtaining the data closeness degree of each piece of initial data, and classifying the initial data based on the data closeness degrees to obtain a plurality of data classification libraries; and performing data division on each data classification library based on frequency band distributed nodes to obtain a plurality of data calculation nodes, and performing data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set. Initial data are obtained from multiple sources, information of different channels is integrated, limitation of a single data source is avoided, a comprehensive and rich data basis is provided for spectrum resource analysis, reliability and integrity of an analysis result are guaranteed, and through data closeness degree classification, frequency band distributed node division, parallel processing and the like, the spectrum resource analysis efficiency is improved. And the data processing speed and the spectral analysis accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a data processing method, system and computer device based on machine learning. Background Art

[0002] Machine learning algorithms can automatically learn and mine potential rules and patterns from a large amount of spectrum data, realizing dynamic monitoring, intelligent analysis and optimized management of spectrum resources. By learning historical spectrum data, machine learning models can predict the usage trends of spectrum resources and allocate spectrum resources in advance for new services or sudden demands.

[0003] Existing machine learning algorithms have deficiencies in the quality control and dynamic adaptability of data sources. The data formats, precisions and reliabilities of different data sources may vary. Without effective data preprocessing and quality assessment mechanisms, the initial data may have noise, biases or missing values. For different types of spectrum data or application scenarios, parameters may need to be frequently adjusted, and the configuration of frequency band distributed nodes may not be flexible enough to dynamically adapt to changes in spectrum resource usage patterns, resulting in unreasonable data partitioning and affecting the accuracy of subsequent analysis results. Summary of the Invention

[0004] The main object of the present invention is to provide a data processing method based on machine learning, aiming to solve the technical problems in the prior art.

[0005] The present invention proposes a data processing method based on machine learning, including: Obtaining multiple initial data based on multiple data sources; Obtaining the data compactness of each initial data, classifying the initial data based on the data compactness to obtain multiple data classification libraries; Performing data partitioning on each data classification library based on frequency band distributed nodes to obtain multiple data calculation nodes, and performing data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set; Obtaining multiple spectrum occupancy rates according to the first frequency band data set, and screening the first frequency band data set based on the spectrum occupancy rates to obtain a second frequency band data set; Obtaining the number and distribution of spectrum holes of each frequency band data in the second frequency band data set, and obtaining the spectrum data idle degree based on the number and distribution of spectrum holes; Filtering the second frequency band data set to obtain an idle frequency band set, and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data; Obtaining frequency band resource characteristics according to the fitting data through inverse mapping technology, and performing frequency band identification processing on the idle frequency band set based on the resource characteristics.

[0006] Preferably, the steps of obtaining the data compactness of each initial data, classifying the initial data based on the data compactness, and obtaining a plurality of data classification libraries include: Obtain the frequency band characteristic information of each of the initial data, where the frequency band characteristic information includes signal strength and spectrum occupancy; Obtain the correlation value corresponding to each of the initial data according to the signal strength and spectrum occupancy; Use the correlation value as the data compactness corresponding to each of the initial data; Obtain a plurality of preset threshold intervals according to the correlation value; Establish a data classification library according to the preset threshold interval; Distinguish all the correlation values based on each preset threshold interval, and classify the initial data corresponding to each correlation value into the corresponding data classification library.

[0007] Preferably, the steps of dividing each of the data classification libraries based on the frequency band distributed nodes to obtain a plurality of data calculation nodes, and performing data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set include: Obtain a plurality of frequency ranges according to the frequency band distributed nodes; Obtain the start node and end node of each of the frequency ranges; Perform node allocation on each of the data classification libraries according to the start node and end node corresponding to the plurality of frequency ranges to obtain a plurality of data calculation nodes; Perform data cleaning on each of the data calculation nodes to obtain a cleaned data calculation node; Extract features from the data calculation nodes to obtain key node feature information, where the key node feature information includes time slot distribution feature information, signal intermittent feature information, and spectrum occupancy feature information; Perform normalization processing on the time slot distribution feature information to obtain a first feature; Perform normalization processing on the signal intermittent feature information to obtain a second feature; Perform normalization processing on the spectrum occupancy feature information to obtain a third feature; Integrate the first feature, the second feature, and the third feature to obtain integrated data corresponding to the data calculation nodes; Obtain a first frequency band data set according to the integrated data corresponding to all data calculation nodes.

[0008] Preferably, the steps of obtaining a plurality of spectrum occupancy rates according to the first frequency band data set, and screening the first frequency band data set based on the spectrum occupancy rates to obtain a second frequency band data set include: Obtain multiple integrated data according to the first frequency band data set; Based on the frequency band distributed nodes, obtain multiple first frequency bands according to the integrated data; Obtain the node duration occupied by the signal of each first frequency band within a specific time window; Obtain the total duration occupied by the signals of all the first frequency bands within a specific time window; Obtain the spectrum occupancy rate according to the node duration and the total duration; Obtain the preset spectrum occupancy rate of the first frequency band data set; Judge whether the spectrum occupancy rate conforms to the preset spectrum occupancy rate; If not, eliminate the first frequency band; If so, integrate the first frequency bands corresponding to the spectrum occupancy rate to obtain a second frequency band data set.

[0009] Preferably, the steps of obtaining the number and distribution of spectrum holes of each frequency band data in the second frequency band data set and obtaining the spectrum data idle degree based on the number and distribution of spectrum holes include: Obtain multiple frequency band data according to the second frequency band data set and use them as the second frequency bands; Obtain the number of spectrum holes according to the second frequency band; Obtain the distribution of each spectrum hole, and obtain the start frequency and end frequency of each spectrum hole based on the distribution of the spectrum hole; Obtain the spectrum bandwidth corresponding to the spectrum hole according to the start frequency and the end frequency; Obtain the total spectrum bandwidth of each frequency band data; Calculate the spectrum data idle degree according to the spectrum bandwidth and the total spectrum bandwidth, where the calculation formula is: ; Wherein, represents the spectrum data idle degree, represents the th spectrum bandwidth of the spectrum hole, represents the total spectrum bandwidth, and n represents the number of spectrum holes.

[0010] Preferably, the steps of filtering the second frequency band data set to obtain an idle frequency band set and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data include: Obtain the preset spectrum data idle degree; Judge whether the spectrum data idle degree corresponding to each second frequency band in the second frequency band data set is greater than the preset spectrum data idle degree; If less, eliminate the corresponding second frequency band; If it is greater than or equal to, use the corresponding second frequency band as the idle frequency band and generate an idle frequency band set; Obtain the frequency band mapping value of the idle frequency band in the high-dimensional feature space based on a preset kernel function; Obtain the feature vector of each idle frequency band in the high-dimensional space according to the frequency band mapping value; Calculate the relationship value between the idle frequency band and other frequency band data through the kernel function according to the feature vector; Fit the similar idle frequency bands based on the relationship value to obtain the fitting data.

[0011] Preferably, the step of obtaining the frequency band resource characteristics according to the fitting data by the inverse mapping technology includes: Obtain a plurality of fitting frequency bands according to the fitting data; Perform inverse operations on the fitting frequency bands through the inverse mapping technology to obtain the spectral characteristics of each fitting frequency band in the original space; Obtain the number of reduced-dimensional spectral holes and the reduced-dimensional spectral occupancy rate according to the spectral characteristics; Obtain the association strength between each fitting frequency band according to the number of reduced-dimensional spectral holes and the reduced-dimensional spectral occupancy rate; Classify each fitting frequency band based on the association strength and mark the classified fitting frequency bands to obtain a fitting frequency band set; Obtain the frequency band resource characteristics according to the fitting frequency band set, where the frequency band resource characteristics include the frequency band hole density and the frequency band occupancy rate.

[0012] The present application also provides a data processing system based on machine learning, including: The first acquisition module is used to acquire a plurality of initial data based on a plurality of data sources; The second acquisition module is used to obtain the data tightness of each initial data, classify the initial data based on the data tightness, and obtain a plurality of data classification libraries; The third acquisition module is used to perform data partitioning on each data classification library based on frequency band distributed nodes to obtain a plurality of data calculation nodes, and perform data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set; The fourth acquisition module is used to obtain a plurality of spectral occupancy rates according to the first frequency band data set, and filter the first frequency band data set based on the spectral occupancy rate to obtain a second frequency band data set; The fifth acquisition module is used to obtain the number and distribution of spectral holes of each frequency band data in the second frequency band data set, and obtain the spectral data idle degree based on the number and distribution of spectral holes; A sixth acquisition module, configured to filter the second frequency band dataset to obtain an idle frequency band set, and map the idle frequency band set to a high-dimensional feature space to obtain fitting data; A processing module, configured to obtain frequency band resource features according to the fitting data through an inverse mapping technique, and perform frequency band identification processing on the idle frequency band set based on the resource features.

[0013] The present invention further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above data processing method based on machine learning are implemented.

[0014] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above data processing method based on machine learning are implemented.

[0015] The beneficial effects of the present invention are as follows: initial data is obtained from multiple sources, information from different channels is integrated, the limitation of a single data source is avoided, a comprehensive and rich data basis is provided for spectrum resource analysis, the reliability and integrity of the analysis results are guaranteed, and through data tightness classification, frequency band distributed node division, parallel processing, etc., the data processing speed and spectrum analysis accuracy are improved. For example, accurately calculating the spectrum occupancy rate, the number of spectrum holes, etc., realizing the refined management of spectrum resources, and improving the utilization rate of spectrum resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0017] Figure 2 It is a schematic structural diagram of the device according to an embodiment of the present invention.

[0018] Figure 3 It is a schematic internal structure diagram of a computer device according to an embodiment of the present application.

[0019] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0021] As Figures 1 - 3 shown, the present application provides a data processing method based on machine learning, including: S1. Obtaining multiple initial data based on multiple data sources; S2. Obtaining the data tightness of each initial data, classifying the initial data based on the data tightness to obtain multiple data classification libraries; S3. Based on the frequency band distributed nodes, each of the data classification libraries is divided into multiple data calculation nodes, and based on machine learning parallel data processing, data processing is performed on each data calculation node to obtain a first frequency band data set; S4. According to the first frequency band data set, multiple spectrum occupancy rates are obtained, and based on the spectrum occupancy rates, the first frequency band data set is filtered to obtain a second frequency band data set; S5. Obtain the spectrum hole quantity and distribution of each frequency band data in the second frequency band data set, and obtain the spectrum data idle degree based on the spectrum hole quantity and distribution; S6. Filter the second frequency band data set to obtain an idle frequency band set, and map the idle frequency band set to a high-dimensional feature space to obtain fitting data; S7. Through the inverse mapping technology, obtain the frequency band resource characteristics according to the fitting data, and perform frequency band identification processing on the idle frequency band set based on the resource characteristics.

[0022] As described in the above steps S1 - S7, multiple initial data are obtained based on multiple data sources (S1). The beneficial effect is that it can comprehensively integrate spectrum data from different channels, ensuring the richness and diversity of the data. Just like collecting data from numerous sensors, base stations, and various wireless devices, it lays a solid foundation for subsequent accurate analysis. Then, obtain the data tightness of each initial data, and classify based on this to obtain multiple data classification libraries (S2), which enables spectrum data with similar characteristics to gather together, facilitating targeted analysis and improving processing efficiency. Just like classifying data with similar signal strengths and spectrum occupancy patterns into one category, it can more efficiently analyze the spectrum characteristics of this category. It can solve the problem that spectrum data is chaotic and disorderly, lacking organization and making it difficult to conduct in-depth analysis. For example, chaotic data will make it extremely difficult to find spectrum rules, and after classification, it can be effectively improved. The specific solution steps are to first extract the frequency band feature information in the initial data, such as signal strength and spectrum occupancy, calculate the relevant values as the measurement standard for tightness, divide the threshold interval to establish a classification library, and classify the data into the corresponding library. For example, different classification libraries are divided according to the signal strength range; Then, divide the data based on the frequency-band distributed nodes and process them in parallel to obtain the first frequency-band data set (S3). Its beneficial effects are to accelerate the processing of spectrum data, respond promptly to spectrum dynamic changes, such as quickly processing data in real-time spectrum monitoring scenarios to support dynamic spectrum allocation decisions; it can also deeply analyze the characteristics of each frequency band to provide a basis for refined management, such as optimizing communication parameters for different frequency bands. This can solve the problem of slow processing of massive spectrum data and inability to meet real-time requirements, and overcome the performance bottleneck of traditional centralized processing in the face of large-scale data. The solution steps are to determine the frequency range and node allocation according to the frequency-band distributed nodes, clean the data to ensure quality, extract features, normalize them, and then integrate the data. For example, allocate data in different frequency-band ranges to the corresponding computing nodes for processing; After that, obtain the spectrum occupancy rate based on the first frequency-band data set and filter it to obtain the second frequency-band data set (S4). This helps to accurately grasp the spectrum occupancy situation and provide a basis for reasonable allocation. For example, filter out idle frequency bands for the launch of new services. At the same time, it can optimize the spectrum configuration and improve the utilization rate. It solves the problem of unreasonable allocation due to the inability to accurately know the spectrum occupancy situation, and corrects decision-making mistakes caused by inaccurate or untimely calculations. The specific steps are to obtain the integrated data from the data set to determine the frequency band, count the signal occupancy duration to calculate the occupancy rate, and compare it with the preset value for filtering. For example, when analyzing a certain frequency band, calculate the occupancy rate within a specific time and judge whether it meets the requirements.

[0023] Subsequently, obtain the number and distribution of spectrum holes in the second frequency-band data set to calculate the spectrum data idle degree (S5). Its beneficial effects are to quantify the idle degree, intuitively reflect the spectrum idle situation, and provide key indicators for spectrum sharing, etc. For example, preferentially allocate idle frequency bands according to the idle degree; it can also understand the fragmentation degree and potential of spectrum resources to guide resource integration. For example, judge whether the distribution of spectrum holes is suitable for large-scale data transmission services. This solves the problem of unclear spectrum idle situation and difficulty in effective utilization, and overcomes the difficulty of unable to plan and manage due to the lack of understanding of the characteristics of the holes. The solution steps are to obtain the frequency-band data, count the number, distribution, and bandwidth of the holes, and calculate the idle degree. For example, analyze the hole distribution of a certain frequency band and calculate its idle degree. After that, filter the second frequency-band data set to obtain the idle frequency-band set and map it to the high-dimensional feature space (S6). This can focus on high-idle-degree resources, reduce the processing volume, and improve efficiency. For example, preferentially process the most potential frequency bands; after mapping, deep feature relationships can be mined to provide more information for optimization. For example, discover the hidden relationship between the hole distribution and interference. It solves the problem that spectrum data is complex and difficult to directly analyze potential value, and breaks through the limitation of low-dimensional space analysis. The solution steps are to set an idle degree threshold to filter the data to obtain the idle frequency-band set, select a kernel function for mapping, calculate the mapping value and feature vector, and then fit the data. For example, set a threshold to screen out idle frequency bands and map them to the high-dimensional space; Finally, the frequency band resource characteristics are obtained through the inverse mapping technology, and the idle frequency band set is processed (S7). The beneficial effect is to realize the association between high-dimensional data and the original spectrum characteristics, which is convenient for understanding and application. For example, converting the high-dimensional clustering results into the actual spectrum hole density characteristics; based on this, the spectrum resources can be accurately classified, marked, and configured. For example, services are allocated according to the hole density. This solves the problem that high-dimensional data is difficult to be interpreted and applied to actual management, and improves the extensive management situation. The solution steps are to use the inverse mapping to obtain the original spectrum characteristics, calculate the association strength, classify and mark to form the resource characteristics, and then process the idle frequency band set. For example, perform the inverse mapping on the fitted frequency band to obtain the spectrum characteristics and classify and process accordingly.

[0024] In one embodiment, the step of obtaining the data tightness of each initial data and classifying the initial data based on the data tightness to obtain a plurality of data classification libraries includes: S201. Obtain the frequency band characteristic information of each of the initial data, where the frequency band characteristic information includes signal strength and spectrum occupancy; S202. Obtain the corresponding correlation value for each of the initial data according to the signal strength and spectrum occupancy; S203. Use the correlation value as the data tightness corresponding to each of the initial data; S204. Obtain a plurality of preset threshold intervals according to the correlation value; S205. Establish a data classification library according to the preset threshold intervals; S206. Distinguish all the correlation values based on each preset threshold interval, and classify the initial data corresponding to each correlation value into the corresponding data classification library.

[0025] As described in the above steps S201-S206, it is of great significance for the present invention to obtain the frequency band characteristic information (including signal strength and spectrum occupancy) of each initial data. The beneficial effect is that, as the key attributes of spectrum data, the accurate acquisition of signal strength and spectrum occupancy can directly reflect the usage status and quality of spectrum resources. For example, in the scenario of wireless network optimization, the accurate mastery of signal strength helps to determine the transmitting power of the base station and the adjustment direction of the antenna to ensure good signal coverage, reduce signal blind spots, and improve the user network experience; the statistics of spectrum occupancy can clarify the busy degree of each frequency band. When allocating spectrum resources, priority is given to ensuring the frequency bands with high spectrum occupancy and key services such as emergency communication to ensure smooth important communication. This step solves the problem that the initial state of spectrum data lacks key feature descriptions, and avoids inaccuracy and incompleteness in subsequent analysis due to the lack of necessary information. The solution steps are to first determine the relevant data fields or parameters in the data source, such as obtaining signal strength and spectrum occupancy data from the base station monitoring system, and then preprocess the original data, including unifying the format, correcting error values, and calibration, so that the data from different sources can be comparable.

[0026] Next, in step S202, the correlation value corresponding to each initial data is obtained according to the signal strength and the spectrum occupancy. The beneficial effect of doing this is that by comprehensively calculating the correlation value from both, the internal relationship of the data can be described more comprehensively and accurately, highlighting the feature differences. For example, when analyzing the spectrum data of a certain area, the correlation value can reflect the co-variation of the signal strength and the spectrum occupancy. Special cases such as high signal strength accompanied by high spectrum occupancy or vice versa can be quantified, providing a basis for accurate classification. This step solves the problem that considering only a single feature alone cannot accurately measure the overall characteristics of the spectrum data and the difficulty in classification due to the lack of a comprehensive index. The solution steps are to select a suitable calculation method, such as setting weights for the two according to business requirements and calculating the weighted sum as the correlation value, and normalizing the signal strength and the spectrum occupancy before calculation to prevent a certain feature from dominating the calculation of the correlation value.

[0027] Then, in step S203, the correlation value is used as the data tightness degree. The beneficial effect of this is to provide a new comprehensive feature description method for the spectrum data, reflecting the degree of tightness of the data association as a whole. Based on this, the internal structure and pattern of the spectrum data can be better captured. For example, for the data collected by two base stations in the same area at the same time period, the signal strength and the spectrum occupancy have strong coordination, and the data tightness degree is high, which can be classified into one category to facilitate the analysis of the usage rules of the spectrum resources in this area. This solves the problems of the lack of a unified and effective feature measurement standard for the spectrum data and the deviation of the data analysis results due to the improper use of indicators. The solution steps are to clarify that the correlation value is the measurement standard for the data tightness degree, record the tightness degree value of each initial data, and explain its value range and meaning. For example, it is stipulated that the value ranges from 0 to 1, 0 indicates weak association, 1 indicates strong association, and intermediate values indicate different degrees of association.

[0028] Subsequently, in step S204, multiple preset threshold intervals are obtained according to the data tightness degree. The beneficial effect of this step is that it can reasonably divide the spectrum data. Different threshold intervals correspond to different subsets of tightness degrees, which helps to understand the distribution rules of the spectrum resources under different usage conditions, and provides a clear and objective standard for classification, ensuring the consistency of the classification results. For example, in long-term spectrum monitoring, setting multiple threshold intervals can divide the data into different tightness degree categories, facilitating the comparison of the usage trends of the spectrum resources in different periods. This step solves the problems of subjective randomness and inaccurate and non-refined classification due to the lack of a clear classification standard. The solution steps are to analyze the distribution of the data tightness degree, such as drawing a histogram and calculating statistical indicators, and then determine the number and range of the threshold intervals in combination with the spectrum management objectives and requirements. For example, when focusing on efficient utilization, narrow intervals are set in the high tightness degree area.

[0029] After that, through step S205, a data classification library is established according to a preset threshold range. Its beneficial effect is to achieve a structured organization of spectrum data, improve data management efficiency, and facilitate retrieval and access. For example, data with a high degree of data tightness is classified into a specific classification library, representing an efficient utilization mode of spectrum resources. When studying efficient utilization cases, data can be quickly located and obtained. This solves the problems of chaotic storage management of spectrum data and low data analysis efficiency. The solution steps are to create and identify a classification library according to the threshold range, define the storage structure and format, and ensure the effective storage and management of data. For example, the classification library is identified by numbers or names, and the data field type, length, and storage method are determined.

[0030] Finally, through step S206, the initial data is classified based on a preset threshold range. Its beneficial effect is to achieve accurate classification, facilitate in-depth analysis and management, and discover hidden pattern rules. For example, data of the same business type is classified into one category, and the impact of this business on spectrum resource requirements can be studied to optimize the allocation strategy; by analyzing independently in the classification library, rules such as the periodic changes of signal strength and spectrum occupancy can be discovered, providing a basis for dynamic management. This step solves the problems of ineffective classification of spectrum data and management decision-making errors caused by inaccurate classification. The solution steps are to traverse the initial data, determine the classification library by judging the interval to which the data tightness belongs, accurately store the data, and update the statistical information of the classification library, such as the number of data, the sum, etc.

[0031] In one embodiment, the step of dividing each of the data classification libraries by the band distributed nodes to obtain a plurality of data calculation nodes, and performing data processing on each of the data calculation nodes based on machine learning parallel data processing to obtain a first band data set includes: S301. Obtain a plurality of frequency ranges according to the band distributed nodes; S302. Obtain the start node and end node of each of the frequency ranges; S303. Perform node allocation on each of the data classification libraries according to the start nodes and end nodes corresponding to the plurality of frequency ranges to obtain a plurality of data calculation nodes; S304. Perform data cleaning on each of the data calculation nodes to obtain cleaned data calculation nodes; S305. Perform feature extraction on the data calculation nodes to obtain key node feature information, where the key node feature information includes time slot distribution feature information, signal intermittent feature information, and spectrum occupancy feature information; S306. Perform normalization processing on the time slot distribution feature information to obtain a first feature; S307. Perform normalization processing on the signal intermittent feature information to obtain a second feature; S308. Perform normalization processing on the spectrum occupancy feature information to obtain a third feature; S309. Integrate the first feature, the second feature, and the third feature to obtain integrated data corresponding to the data calculation node; S310. Obtain the first frequency band dataset based on the integrated data corresponding to all data calculation nodes.

[0032] As described in the above steps S301 - S310, the present invention obtains multiple frequency ranges based on frequency band distributed nodes. The beneficial effect is that different frequency ranges have unique spectrum usage characteristics and propagation laws. For example, the low frequency band is suitable for the communication of Internet of Things devices with wide - area coverage, and the high frequency band is conducive to high - speed data transmission services. By clarifying the frequency range, it can provide an accurate basis for spectrum resource management, improve resource utilization efficiency, and avoid service interference. For example, in an urban intelligent transportation system, vehicle - to - vehicle communication can be allocated to a specific low - frequency band to ensure stable communication between vehicles. This step solves the problem of chaotic management of spectrum resources in the frequency dimension and avoids the inability to give full play to the resource advantages due to the lack of understanding of frequency characteristics. The solution steps are to comprehensively scan and monitor the frequency band, collect information such as signal strength and spectrum occupancy, and then divide the frequency range according to service requirements and spectrum management regulations. For example, refer to the ITU standard to divide the frequency bands for services such as mobile communication and radio and television.

[0033] Next, through step S302, obtain the start node and end node of each frequency range. This helps to accurately locate the frequency range, improve the accuracy of data processing, and facilitate data management in distributed computing. For example, when calculating the spectrum occupancy rate of a certain frequency band, clear boundaries can ensure accurate statistical data and provide a reliable basis for resource allocation. At the same time, it is convenient for the calculation node to quickly locate the data range and improve the processing efficiency. This step solves the problems of chaotic data processing and task allocation caused by unclear definition of the frequency range, and avoids data processing delays and resource waste caused by unclear ranges. The solution steps are to determine accurate frequency values in combination with the accuracy of spectrum monitoring equipment, record and store relevant information, and establish indexes or identifiers. For example, record metadata such as start and end frequencies and the service type to which they belong in a database.

[0034] Then, through step S303, allocate data calculation nodes according to the start and end nodes. This can achieve reasonable allocation of spectrum data, give full play to the advantages of distributed computing, improve the processing speed, and enhance the flexibility and scalability of the system. For example, when processing a large amount of spectrum data, allocate data of different frequency bands to multiple nodes for parallel processing to shorten the processing time; when the service changes, it is convenient to adjust the node allocation strategy. This step solves the problem of uneven task allocation in distributed computing and avoids data processing delays and resource waste caused by unreasonable allocation. The solution steps are to formulate an allocation strategy based on the performance of the nodes and the characteristics of the frequency range. For example, allocate complex high - frequency band data to high - performance nodes, and then use the distributed computing framework to allocate data according to the strategy to ensure accuracy and integrity.

[0035] Subsequently, through step S304, data cleaning can remove noise, outliers, and incorrect data, improve data quality, reduce computational complexity, and storage requirements. For example, electromagnetic interference may cause errors in spectrum monitoring data. After cleaning, the reliability of subsequent analysis can be ensured, while reducing the consumption of storage and computing resources. This step solves the problem that noise and incorrect data affect the accuracy of analysis results and avoids resource waste and low processing efficiency caused by invalid data. The solution steps are to set reasonable cleaning rules and thresholds, such as removing outliers based on the normal range of signal strength, and adopting outlier detection and data smoothing algorithms, such as using the 3σ principle to detect signal strength outliers and using the moving average method to smooth spectrum occupancy data. Then, in step S305, feature extraction can mine key information, reduce the data dimension, facilitate analysis and discovery of patterns, and provide decision-making support for management. For example, the slot distribution feature can reflect the spectrum time usage pattern, the signal intermittency feature helps to understand signal stability, and the spectrum occupancy feature reflects the resource occupancy situation. By analyzing the slot distribution, a basis for dynamic spectrum allocation can be provided. This step solves the problems that spectrum data is highly dimensional and complex and difficult to directly analyze, and that management decisions lack pertinence. The solution steps are to determine key features, such as determining the importance of features such as slot distribution according to service requirements, and then adopting appropriate algorithms for extraction, such as using time-domain analysis methods to extract slot distribution features, using signal processing techniques to detect signal intermittency features, and using spectrum analysis algorithms to obtain spectrum occupancy features.

[0036] Next, through step S306, the slot distribution feature information is normalized to obtain the first feature, making the data of different nodes comparable and improving the application effect of data analysis and machine learning algorithms. For example, when comparing the slot usage of different base stations, after normalization, the relative efficiency can be accurately evaluated, and it also helps algorithms such as neural networks to converge faster and improve accuracy. This step solves the problem that slot distribution features cannot be compared due to different dimensions and ranges and affects algorithm performance. The solution steps are to select a suitable normalization method, such as min-max normalization or standardization, select according to the characteristics of the feature and analysis requirements, and then apply the method to process the data of each node and record relevant parameters.

[0037] Then, in step S307, the signal intermittency feature information is normalized to obtain the second feature, and the effect is similar to the normalization of the slot distribution feature, which can improve the accuracy of signal stability evaluation and data mining efficiency. For example, when analyzing the signal quality of a large-scale wireless network, after normalization, the signal stability in different regions can be accurately compared, which helps clustering algorithms to identify signal intermittency anomalies. This step solves the problem that signal intermittency features cannot be compared due to device and environmental factors and affects management decisions. The solution steps are to select a suitable normalization method according to the signal intermittency feature, such as selecting min-max normalization according to the proportional relationship of signal intermittency, and then processing the data and recording parameters.

[0038] After that, through step S308, the normalized spectrum occupancy feature information is used to obtain the third feature, which facilitates the comparison of spectrum occupancy distributions, provides a basis for resource allocation and prediction, and improves the accuracy of relevant models. For example, when comparing the spectrum occupancy requirements of different services, the differences can be clearly shown after normalization, which helps to predict the spectrum requirements. This step solves the problem that the spectrum occupancy features are incomparable due to differences in frequency bands and traffic volumes and affect the prediction and decision-making models. The solution steps are to select a suitable normalization method, such as using min-max normalization to highlight the differences in spectrum occupancy ratios, and then process the data and record the parameters.

[0039] Finally, through step S310, the first, second, and third features are integrated to obtain the integrated data corresponding to the nodes, and then the first frequency band dataset is summarized. This can effectively integrate the feature data processed by each node to form a dataset that comprehensively reflects the spectrum status, providing a complete data basis for subsequent spectrum resource evaluation, allocation, etc., such as providing data support for spectrum hole analysis. This step solves the problems of scattered data and the lack of a dataset for the overall spectrum status. The solution steps are to integrate the normalized feature data of each node according to a predetermined format and structure to ensure the integrity and consistency of the data, and form the first frequency band dataset that can be used for further analysis.

[0040] In one embodiment, the step of obtaining multiple spectrum occupancy rates according to the first frequency band dataset and screening the first frequency band dataset based on the spectrum occupancy rates to obtain the second frequency band dataset includes: S401. Obtain multiple pieces of integrated data according to the first frequency band dataset; S402. Based on the frequency band distribution nodes, obtain multiple first frequency bands according to the integrated data; S403. Obtain the node duration of the signal occupancy in each first frequency band within a specific time window; S404. Obtain the total duration of the signal occupancy in all the first frequency bands within a specific time window; S405. Obtain the spectrum occupancy rate according to the node duration and the total duration; S406. Obtain the preset spectrum occupancy rate of the first frequency band dataset; S407. Determine whether the spectrum occupancy rate meets the preset spectrum occupancy rate; If not, then eliminate this first frequency band; If so, then integrate the first frequency bands corresponding to this spectrum occupancy rate to obtain the second frequency band dataset.

[0041] As described in the above steps S401 - S407, the present invention obtains multiple integrated data from the first - frequency - band data set. Its beneficial effects are as follows: the integrated data summarizes the key spectrum information after the previous processing, laying a solid foundation for comprehensively analyzing the spectrum resource situation, and can also help to explore potential associations and rules. For example, when analyzing the spectrum usage in urban areas, the integrated data covers the signal characteristics of different frequency bands in each time period and region. Through analysis, it can be found that the demand for high - frequency bands is large during the day in commercial areas, and the low - frequency bands are more active at night in residential areas, providing data support for spectrum management. This step solves the problems that the spectrum data is scattered and disordered, making it difficult to comprehensively analyze and unable to effectively utilize potential information. The solution steps are to first analyze the data - set structure and storage method, clarify the data associations, and then use SQL or the API of the data - processing framework to extract the integrated data according to requirements. For example, relevant fields are selected from the database table storing frequency - band data and combined into the required format for analysis.

[0042] Next, through step S402, based on the frequency - band distributed nodes, multiple first - frequency bands are obtained according to the integrated data. This is helpful for spectrum - resource management and planning at the frequency - band level, allocating services according to the characteristics of the frequency bands, improving efficiency, reducing interference, and adapting to the dynamic changes of the spectrum. For example, according to the configuration of the frequency - band distributed nodes, the spectrum is divided into first - frequency bands suitable for services such as mobile communication and radio and television. When the 5G service develops, the appropriate high - frequency - band first - frequency bands can be quickly found from the integrated data and resources can be allocated. This step solves the problems of the lack of effective division and management of frequency bands and the inability to respond to spectrum changes in a timely manner. The solution steps are to extract the first - frequency - band data from the integrated data according to the node configuration and division rules using data - screening conditions (such as frequency range, service - type identifier), and then verify and sort it, adding metadata to ensure the integrity and accuracy of the data. For example, information such as bandwidth and number is added to the mobile - communication frequency band.

[0043] Then, through step S403, the node duration occupied by the signal of each first - frequency band within a specific time window is obtained. This can accurately quantify the frequency - band usage situation, evaluate the utilization rate and busyness, provide a basis for optimizing resource allocation, and can also analyze the usage rules and trends. Taking the traffic - monitoring frequency band as an example, by long - term monitoring of the node duration at different time periods, it is found that the morning - peak duration is longer, the evening - peak duration is the second, and the night duration is shorter, providing a reference for dynamic spectrum allocation. For example, priority is given to ensuring the transmission of monitoring data during peak hours. This step solves the problems of being unable to accurately grasp the time usage situation of the spectrum and the lack of timeliness and pertinence in management decisions. The solution steps are to first set the time window according to the service requirements (such as selecting 1 minute for traffic monitoring), then screen the occupancy records within this window from the first - frequency - band data, calculate the node duration by statistically calculating the difference between the start and end times of occupancy or obtaining the occupancy - duration field, and calculate according to the rules for signal - interruption situations, such as accumulating continuous occupancy time periods.

[0044] Subsequently, through step S404, obtain the total duration of signals occupied by all first frequency bands within a specific time window. This can comprehensively grasp the spectrum usage intensity and busyness, serving as a basis for evaluating the utilization rate and facilitating comprehensive evaluation and comparison of multiple frequency bands. For example, by comparing the total durations in different time periods, it is found that the total spectrum usage duration during holidays is lower than that on weekdays, indicating changes in business requirements and the need to adjust resource allocation. It is also possible to compare the total durations of different frequency bands to analyze their positions in spectrum resource usage. For example, the total duration of high-frequency bands is short, and methods to improve the utilization rate need to be studied. This step solves the problems of being unable to globally evaluate the spectrum usage intensity and the lack of a global perspective in allocation decisions. The solution steps are to traverse the occupancy records of the first frequency bands, accumulate the node durations using loops or database aggregation functions to obtain the total duration, and record and store the total duration and related metadata. For example, storing it in a database table associated with the time window and calculation time facilitates query and analysis.

[0045] After that, through step S405, obtain the spectrum occupancy rate based on the node duration and the total duration. The spectrum occupancy rate reflects the occupancy degree of the frequency band in an intuitive percentage form, facilitating comparison of different frequency bands, providing a quantitative basis for resource management, and enabling the establishment of a dynamic monitoring mechanism. For example, if the spectrum occupancy rate of a mobile communication frequency band is high, service quality needs to be guaranteed or sharing needs to be considered; if it is low for the radio and television frequency band, it can be used as a resource for new business expansion. By regularly calculating the occupancy rates in different time windows, abnormal fluctuations can be detected in a timely manner. For example, if the occupancy rate of a certain frequency band suddenly increases, there may be a new service or interference source. This step solves the problems of being unable to intuitively compare the occupancy degrees of frequency bands and the lack of scientific accuracy in management decisions. The solution steps are to calculate the occupancy rate of each first frequency band using a formula, ensure unit consistency, check and correct abnormal values, and record and store the occupancy rate associated with the frequency band and time window. For example, storing it in a database table facilitates query and processing.

[0046] Next, through step S406, obtain the preset spectrum occupancy rate of the first frequency band dataset. The preset value provides a clear standard for spectrum screening and management, which is determined based on management objectives, business requirements, and historical situations to ensure fair and reasonable management. For example, for medical communication services with high real-time requirements, the preset value is low; for non-critical Internet of Things services, the preset value can be higher. This avoids problems such as randomness, blindness, uneven distribution, and low efficiency caused by the lack of standards in spectrum management. The solution steps are to analyze management objectives and business requirements to determine the preset principle, such as guaranteeing the service quality of key services or improving the overall utilization rate, and then combine historical data statistical analysis to determine specific values. For example, determine a suitable value by referring to historical occupancy rate statistical indicators.

[0047] Finally, in step S407, it is determined whether the spectrum occupancy rate meets the preset spectrum occupancy rate and the frequency bands are screened. Those that do not meet the requirements are excluded, which can eliminate unreasonable occupied frequency bands and improve management efficiency; those that meet the requirements are integrated into the second frequency band dataset to provide a basis for subsequent decisions. For example, frequency bands with overly high spectrum occupancy rates causing congestion or overly low rates wasting resources are excluded. After integrating the compliant frequency bands, the frequency bands can be reasonably allocated to different services according to their characteristics to optimize resource allocation. This step solves the problem that unreasonable occupied frequency bands in spectrum resources affect management efficiency and unfair resource allocation. The solution steps are to compare the occupancy rate of each first frequency band with the preset value, exclude those that do not match, and integrate relevant data such as frequency band characteristics and occupancy rate according to the rules to form the second frequency band dataset.

[0048] In one embodiment, the step of obtaining the number and distribution of spectrum holes for each frequency band data in the second frequency band dataset and obtaining the spectrum data idle degree based on the number and distribution of spectrum holes includes: S501. Obtain a plurality of frequency band data according to the second frequency band dataset and use them as the second frequency band; S502. Obtain the number of spectrum holes according to the second frequency band; S503. Obtain the distribution of each of the spectrum holes, and obtain the start frequency and end frequency of each spectrum hole based on the distribution of the spectrum holes; S504. Obtain the spectrum bandwidth corresponding to the spectrum hole according to the start frequency and the end frequency; S505. Obtain the total spectrum bandwidth of each of the frequency band data; S506. Calculate the spectrum data idle degree according to the spectrum bandwidth and the total spectrum bandwidth, where the calculation formula is: ; Wherein, represents the spectrum data idle degree, represents the th spectrum bandwidth of the spectrum hole, represents the total spectrum bandwidth, and n represents the number of spectrum holes.

[0049] As described in the above steps S501 - S506, the present invention obtains frequency band data from the second - frequency - band data set as the second frequency band. Its beneficial effects are as follows: it can focus on the high - potential frequency bands after preliminary screening, improve the pertinence and effectiveness of spectrum resource management, and is conducive to hierarchical management and refined analysis. For example, when expanding new 5G millimeter - wave services, by focusing on the second - frequency - band data, the frequency bands suitable for millimeter - wave communication can be quickly located, the search range can be reduced, the resource allocation efficiency can be improved, and special spectrum sharing strategies can be formulated for the second frequency band. This step solves the problems of lack of focus in spectrum resource analysis and lack of pertinence in management decisions. The solution steps are to analyze the data - set structure format, clarify the storage method and field information of the frequency - band data, use data - reading and processing tools to extract the frequency - band data according to the rules, and organize it into a format convenient for subsequent analysis. For example, the data is read into a list, and each element contains information such as the frequency - band number and frequency range.

[0050] Next, through step S502, the number of spectrum holes in the second frequency band is obtained. This can quantify the scale of idle resources in the frequency band and provide a key basis for spectrum resource allocation and sharing. For example, when planning the deployment of large - scale Internet of Things devices, clarifying the number of spectrum holes can reasonably arrange the number and layout of devices, ensure stable communication of the devices, and avoid resource waste. This step solves the problems of inability to grasp the scale of idle spectrum resources and unreasonable allocation. The solution steps are to use spectrum - hole detection technology. For example, set a signal - strength threshold (-110 dBm), and values below this are regarded as holes. Use a counter to count the number of holes and store it for subsequent calculation of the idle degree.

[0051] Then, through step S503, the distribution of spectrum holes and the start and end frequencies are obtained. This can accurately describe the position range of the holes, have a clear understanding of the idle structure of spectrum resources, and help to discover fragmentation problems. For example, when designing a wireless communication frequency plan, the communication frequencies can be accurately placed based on the hole distribution to avoid interference. If it is found that the holes are scattered, an integration strategy can be formulated. This step solves the problems of inaccurate understanding of the position range of the holes and inability to solve fragmentation problems. The solution steps are based on hole detection, use the accuracy of spectrum monitoring equipment to determine the start and end frequencies, and store them in a suitable data structure (such as a dictionary), where the key is the frequency - band number and the value is a list of hole frequencies.

[0052] Subsequently, in step S504, the bandwidth of the spectrum hole is calculated. This can quantify the actual size of the hole, which is of great significance for evaluating the idle capacity and helps optimize resource matching. For example, high-definition video transmission services require a large bandwidth, while sensor services have a small bandwidth requirement. Allocating resources according to the hole bandwidth can improve the resource utilization efficiency. This step solves the problem of being unable to measure the available capacity of the hole and the waste of resources or abnormal services. The solution is to calculate using the formula "end frequency - start frequency", ensuring that the frequency units are consistent, associating the bandwidth with the hole information storage for easy query and retrieval, such as storing in a database table and creating an index. Then in step S505, the total spectrum bandwidth of the frequency band is obtained. This can provide an overall understanding of the scale of the frequency band resources and is an important basis for evaluating utilization and business planning. For example, when comparing the utilization efficiency of 5G communication resources in different frequency bands, the total spectrum bandwidth combined with the hole information can determine the utilization degree and provide a basis for reasonable business allocation. This step solves the problem of lacking an understanding of the overall resource scale of the frequency band and being unable to accurately evaluate the utilization rate. The solution is to directly obtain the total spectrum bandwidth if it is directly provided in the data, otherwise calculate it according to the frequency range, associate the total spectrum bandwidth with other information of the frequency band for storage, and establish a data set of frequency band information.

[0053] Finally, in step S506, the idle degree of the spectrum data is calculated. This indicator comprehensively reflects the idle degree of the spectrum resources and provides a basis for dynamic management and optimization. For example, in spectrum trading, the idle degree can measure the value of the spectrum, assist in pricing, and can also dynamically adjust resource allocation according to the idle degree. This step solves the problem of lacking a comprehensive indicator for evaluating the idle degree and decision-making errors. The solution is to calculate the idle degree according to the formula, ensure the accuracy of relevant parameters, associate the idle degree with the frequency band information for storage, form an idle degree data set for easy query, analysis, and decision-making, such as sorting and screening frequency bands for management and allocation according to the idle degree.

[0054] In one embodiment, the step of filtering the second frequency band data set to obtain an idle frequency band set and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data includes: S601. Obtain a preset spectrum data idle degree; S602. Determine whether the spectrum data idle degree corresponding to each second frequency band in the second frequency band data set is greater than the preset spectrum data idle degree; If it is less, eliminate the corresponding second frequency band; If it is greater than or equal to, use the corresponding second frequency band as an idle frequency band and generate an idle frequency band set; S603. Obtain the frequency band mapping value of the idle frequency band in the high-dimensional feature space based on a preset kernel function; S604. Obtain the feature vector of each idle frequency band in the high-dimensional space according to the frequency band mapping value; S605. Calculate the relationship value between the idle frequency band and the data of other frequency bands according to the feature vector through the kernel function; S606. Fit similar idle frequency bands based on relationship values to obtain fitting data.

[0055] As described in the above steps S601 - S606, the present invention obtains the preset spectrum data idle degree. Its beneficial effect is to provide a clear standard and basis for spectrum resource screening, which is determined according to management objectives, service requirements, and historical situations, and flexible resource allocation can be achieved. For example, when ensuring critical services in spectrum - tense areas, a low preset value is set to ensure high utilization; when the demand is large at the initial stage of new service development, the preset value is increased to quickly obtain idle resources. This solves the problems of no unified standard for screening and inability to adapt to different requirements. The solution steps are to first analyze the management objectives to determine the principle, such as referring to the occupancy rate distribution with the utilization rate as the goal, and then combine historical data statistics (mean, median, etc.) to determine the specific value, such as setting the idle degree standard according to the historical occupancy rate range.

[0056] Next, through step S602, judge the relationship between the second - band idle degree and the preset value to screen the frequency bands. Accurately screen idle frequency bands, improve accuracy and effectiveness, reduce the processing volume, and discover potential value. For example, when processing a large amount of data, low - idle - degree frequency bands are eliminated to reduce complexity; when exploring the application of 5G millimeter - wave, high - idle - degree frequency bands are screened to facilitate technological development. This step solves the problems of inability to accurately identify idle and valuable frequency bands and inaccurate screening. The solution steps are to traverse the second - band to obtain the idle degree (query or calculation), compare it with the preset value, and mark the frequency - band status to ensure that the data type and comparison logic are correct.

[0057] Then, through step S603, obtain the high - dimensional mapping values of the idle frequency bands based on the preset kernel function. Convert the spectrum data into a high - dimensional space, mine potential feature relationships, provide a new perspective method, and improve the accuracy of classification and recognition. For example, the disordered spectrum holes in low - dimension may show regularity in high - dimension, which helps to distinguish the characteristics of different service frequency bands. This solves the problems of insufficient low - dimensional features and lack of depth and accuracy in decision - making. The solution steps are to select a suitable kernel function (such as linear or Gaussian kernel function), which depends on the spectrum characteristics and purposes, set parameters (such as the Gaussian kernel bandwidth), calculate the high - dimensional mapping values using the idle - frequency - band data, and cross - validation can be used to determine the parameters.

[0058] Subsequently, through step S604, obtain high - dimensional feature vectors from the frequency - band mapping values. Provide a structured representation for high - dimensional spectrum data, which helps to discover the internal structural patterns and is convenient for subsequent analysis and processing. For example, in clustering analysis, the feature vectors are used as inputs to improve the efficiency and accuracy of the algorithm, and analyzing the vector relationships can discover the clustering structure of spectrum resources. This step solves the problems of lack of structured representation of high - dimensional data and inability to accurately identify structural patterns. The solution steps are to arrange the mapping values in a certain order to form feature vectors, determine the order according to the spectrum characteristics, ensure that the dimensions are the same and meaningful, and perform standardization or normalization processing if necessary.

[0059] After that, in step S605, the kernel function is used to calculate the frequency band relationship value. By deeply understanding the mutual relationship of spectrum resources, it provides a basis for collaborative management optimization and discovers potential collaborative complementary relationships. For example, when sharing spectrum, idle frequency bands with low correlation are selected according to the relationship value to reduce interference; it is also possible to discover the possibility of collaborative work between frequency bands, providing ideas for business expansion. This solves the problem of being unable to accurately measure the relationship of spectrum resources and the lack of collaborative optimization in decision-making. The solution steps are to select an appropriate kernel function (which can be the same as or different from the mapping kernel function), calculate the relationship value using the eigenvector, ensure the correct format and reasonable parameters, and appropriately process the relationship value.

[0060] Finally, in step S606, based on the relationship value, similar idle frequency bands are fitted to obtain fitting data. Integrate and abstract similar frequency bands, discover common pattern rules, simplify the data and retain key information, improving management efficiency and decision-making scientificity. For example, when analyzing the distribution of spectrum holes, a model representing the commonality is obtained by fitting, which helps to grasp the overall characteristics; when allocating spectrum, quick decisions are made based on the fitting data, such as batch allocation after determining the frequency band suitable for the service. This step solves the problems that the amount of spectrum resource data is too complex to manage and make decisions and the lack of overall grasp in analysis. The solution steps are to select a fitting method according to the relationship value (such as linear fitting or clustering algorithm), determine it according to the characteristics and purpose of the spectrum, set parameters and then perform fitting, evaluate and verify the results, and if not ideal, adjust the method or parameters and refit.

[0061] In one embodiment, the step of obtaining the frequency band resource characteristics according to the fitting data by the inverse mapping technology includes: S701. Obtain a plurality of fitted frequency bands according to the fitting data; S702. Perform an inverse operation on the fitted frequency bands through the inverse mapping technology to obtain the spectrum characteristics of each fitted frequency band in the original space; S703. Obtain the reduced-dimensional spectrum hole number and reduced-dimensional spectrum occupancy rate according to the spectrum characteristics; S704. Obtain the association strength between each fitted frequency band according to the reduced-dimensional spectrum hole number and reduced-dimensional spectrum occupancy rate; S705. Classify each fitted frequency band based on the association strength, and mark the classified fitted frequency bands to obtain a set of fitted frequency bands; S706. Obtain the frequency band resource characteristics according to the set of fitted frequency bands, where the frequency band resource characteristics include frequency band hole density and frequency band occupancy rate.

[0062] As described in the above steps S701 - S706, the present invention obtains multiple fitting frequency bands from the fitting data. Its beneficial effect lies in being able to extract representative spectrum resource units, facilitating hierarchical management and refined analysis, and formulating strategies according to characteristics. For example, the fitting frequency bands with strong signals and regular holes are preferentially allocated to high - quality services. This solves the problems of difficult management of fitting data and lack of pertinence in decision - making. The solution steps are to first clarify the structure of the fitting data, such as determining the characteristics of the fitting frequency band range according to the clustering center, etc. in the clustering results, then select a suitable algorithm (such as distance - based or threshold - based method) for extraction, and finally organize and record, such as dividing the fitting frequency bands according to the center and threshold after spectral data clustering.

[0063] Next, through step S702, the original - space spectral characteristics of the fitting frequency bands are obtained through inverse mapping. The combination of restored information and reality helps to verify the results of high - dimensional analysis. For example, in the original space, the frequency range of the frequency band is determined according to the spectral characteristics for spectral allocation and interference analysis; the high - dimensional special clustering pattern can be interpreted in practical significance through inverse mapping. This step solves the problems of difficult application of high - dimensional data in actual management and unreliable decision - making. The solution steps are to first determine the inverse mapping technology and inverse operation method (such as the inverse operation corresponding to the Gaussian kernel function), then substitute the fitting frequency band data for calculation, and store it after checking the rationality of the results. For example, use the Gaussian kernel inverse operation to convert the fitting frequency band data into the original - space spectral characteristics.

[0064] Then, through step S703, the number and occupancy rate of the reduced - dimension spectral holes are obtained according to the spectral characteristics. Reducing the data dimension simplifies the description, which is conducive to quickly evaluating the spectral state, such as understanding the spectral utilization situation of the area during inspection; it also helps to improve the analysis efficiency, such as sorting and screening according to the number of holes when searching for idle spectra. This step solves the problems of high - dimensionality and complex calculation of the original spectral characteristics and difficult rapid evaluation. The solution steps are to select a suitable dimension - reduction method (such as PCA or LDA), substitute the original spectral characteristics for calculation (assuming the number of principal components of PCA), and verify the accuracy of the results. For example, use PCA for dimension reduction and compare with the original data to check the accuracy.

[0065] Subsequently, through step S704, the association strength of the fitting frequency bands is obtained according to the reduced - dimension characteristics. It can understand the relationship between frequency bands and provide a basis for collaborative management. For example, when sharing spectra, allocate frequency bands with less interference according to this; it can also discover collaborative and complementary relationships. This step solves the problems of difficult measurement of the relationship between frequency bands and lack of collaborative optimization in decision - making. Obtain the association strength by substituting the reduced - dimension characteristics, ensure the correct format and interpret the normalized results. For example, calculate using the correlation coefficient method and map the results to the interval [0,1].

[0066] After that, through step S705, the fitting frequency bands are classified and marked according to the association strength to obtain a set of fitting frequency bands, realizing the classified management of spectrum resources, improving efficiency and enhancing pertinence. For example, the same type of frequency bands adopt a unified management strategy; it also helps to discover hidden pattern rules. For example, the periodic change of spectrum holes in a certain type of frequency band provides a basis for dynamic management. This step solves the problems of lack of classification methods and means and unscientific decision-making. The solution steps are to select a suitable classification algorithm (such as K-means or hierarchical clustering), set parameters for classification and marking according to the requirements of the algorithm, evaluate and verify the effect (using indicators such as silhouette coefficient). If it is not ideal, adjust and reclassify. For example, use K-means clustering and evaluate the effect according to the indicators.

[0067] Next, through step S706, the characteristics of frequency band resources are obtained according to the set of fitting frequency bands. Overall, grasp the key characteristics of the spectrum, providing a basis for macroscopic management decisions. For example, when planning spectrum resources, allocate according to the hole density and occupancy rate; it is also beneficial to evaluate and compare spectrum resources at different levels. For example, comparing the hole density in different regions helps in the location selection of new services. This step solves the problems of lack of overall feature description and lack of comprehensiveness and long-term perspective in decision-making. The solution steps are to extract the information of the fitting frequency bands, calculate the hole density (the ratio of the number of holes to the bandwidth) and the occupancy rate (the ratio of the occupied duration to the total duration), and summarize and organize them into a feature data set. For example, after calculation, store it in a database table for query and analysis. Finally, based on the resource characteristics, the idle frequency bands are identified. Accurately identify the type characteristics of the idle frequency bands, ensuring precise allocation and efficient utilization. For example, allocate high-potential idle frequency bands to high-definition video transmission services; it also optimizes the management process and improves efficiency. For example, reduce the monitoring input for low-quality idle frequency bands. This step solves the problems of inability to distinguish and utilize idle frequency bands and low management efficiency. The solution steps are to judge the type of idle frequency bands according to the resource characteristics (hole density, occupancy rate, etc.).

[0068] This application also provides a data processing system based on machine learning, including: The first acquisition module 1 is used to acquire a plurality of initial data based on a plurality of data sources; The second acquisition module 2 is used to acquire the data tightness of each initial data, classify the initial data based on the data tightness, and obtain a plurality of data classification libraries; The third acquisition module 3 is used to divide each data classification library based on the frequency band distributed nodes to obtain a plurality of data calculation nodes, and perform data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set; The fourth acquisition module 4 is used to obtain a plurality of spectrum occupancy rates according to the first frequency band data set, and screen the first frequency band data set based on the spectrum occupancy rates to obtain a second frequency band data set; The fifth acquisition module 5 is used to obtain the number and distribution of spectrum holes of each frequency band data in the second frequency band data set, and obtain the idle degree of spectrum data based on the number and distribution of spectrum holes; The sixth acquisition module 6 is configured to filter the second frequency band data set to obtain an idle frequency band set, and map the idle frequency band set to a high-dimensional feature space to obtain fitted data; The processing module 7 is configured to obtain frequency band resource features according to the fitted data through an inverse mapping technique, and perform frequency band identification processing on the idle frequency band set based on the resource features.

[0069] It should be noted that each module and unit in the data processing system based on machine learning corresponds one-to-one to the steps in the data processing method based on machine learning.

[0070] As Figure 3 shown, the present application further provides a computer device, which may be a server, and its internal structure may be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all data required for the process of the data processing method based on machine learning. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the data processing method based on machine learning.

[0071] Those skilled in the art can understand that Figure 3 the structure shown in

[0072] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.

[0073] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM can be obtained in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0074] It should be noted that in this document, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article, or method including that element.

[0075] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A data processing method based on machine learning, characterized in that, Including: Obtaining a plurality of initial data based on multiple data sources; Obtaining the data tightness of each initial data, classifying the initial data based on the data tightness to obtain a plurality of data classification libraries; Dividing each of the data classification libraries based on frequency band distributed nodes to obtain a plurality of data calculation nodes, and performing data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set; Obtaining a plurality of spectrum occupancy rates according to the first frequency band data set, and screening the first frequency band data set based on the spectrum occupancy rate to obtain a second frequency band data set; Obtaining the number and distribution of spectrum holes of each frequency band data in the second frequency band data set, and obtaining the spectrum data idle degree based on the number and distribution of spectrum holes; Filtering the second frequency band data set to obtain an idle frequency band set, and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data; Obtaining frequency band resource characteristics according to the fitting data through an inverse mapping technique, and performing frequency band identification processing on the idle frequency band set based on the resource characteristics.

2. The data processing method based on machine learning according to claim 1, wherein The step of obtaining the data tightness of each initial data, classifying the initial data based on the data tightness to obtain a plurality of data classification libraries includes: Obtaining the frequency band feature information of each initial data, where the frequency band feature information includes signal strength and spectrum occupancy; Obtaining a corresponding correlation value for each initial data according to the signal strength and spectrum occupancy; Taking the correlation value as the data tightness corresponding to each initial data; Obtaining a plurality of preset threshold intervals according to the correlation value; Establishing a data classification library according to the preset threshold interval; Distinguishing all the correlation values based on each preset threshold interval, and classifying the initial data corresponding to each correlation value into the corresponding data classification library.

3. The data processing method based on machine learning according to claim 1, wherein The step of dividing each of the data classification libraries based on frequency band distributed nodes to obtain a plurality of data calculation nodes, and performing data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set includes: Obtaining the start node and end node of a plurality of frequency ranges according to the frequency band distributed nodes; Performing node allocation on each data classification library according to the start node and end node corresponding to a plurality of the frequency ranges to obtain a plurality of data calculation nodes; Performing data cleaning on each data calculation node to obtain a cleaned data calculation node; Performing feature extraction on the data calculation node to obtain key node feature information, where the key node feature information includes time slot distribution feature information, signal intermittent feature information, and spectrum occupancy feature information; Performing normalization processing on the time slot distribution feature information to obtain a first feature; Performing normalization processing on the signal intermittent feature information to obtain a second feature; Performing normalization processing on the spectrum occupancy feature information to obtain a third feature; Obtaining a first frequency band data set according to the first feature, the second feature, and the third feature.

4. The data processing method based on machine learning according to claim 1, characterized in that The step of obtaining a plurality of spectrum occupancy rates according to the first frequency band data set, and screening the first frequency band data set based on the spectrum occupancy rate to obtain a second frequency band data set includes: Obtain multiple integrated data according to the first frequency band dataset; Based on the frequency band distributed nodes, obtain multiple first frequency bands according to the integrated data; Obtain the node duration occupied by the signal of each first frequency band within a specific time window; Obtain the total duration occupied by the signals of all the first frequency bands within a specific time window; Obtain the spectrum occupancy rate according to the node duration and the total duration; Obtain the preset spectrum occupancy rate of the first frequency band dataset; Judge whether the spectrum occupancy rate conforms to the preset spectrum occupancy rate; If not, eliminate the first frequency band; If so, integrate the first frequency bands corresponding to the spectrum occupancy rate to obtain a second frequency band dataset.

5. The data processing method based on machine learning according to claim 1, wherein The steps of obtaining the spectrum hole quantity and distribution of each frequency band data in the second frequency band dataset and obtaining the spectrum data idle degree based on the spectrum hole quantity and distribution include: Obtain multiple frequency band data according to the second frequency band dataset and use them as the second frequency band; Obtain the spectrum hole quantity according to the second frequency band; Obtain the distribution of each spectrum hole, and obtain the starting frequency and ending frequency of each spectrum hole based on the distribution of the spectrum hole; Obtain the spectrum bandwidth corresponding to the spectrum hole according to the starting frequency and the ending frequency; Obtain the total spectrum bandwidth of each frequency band data; Calculate the spectrum data idle degree according to the spectrum bandwidth and the total spectrum bandwidth.

6. The data processing method based on machine learning according to claim 1, characterized in that The steps of filtering the second frequency band dataset to obtain an idle frequency band set and mapping the idle frequency band set to a high-dimensional feature space to obtain fitting data include: Obtain the preset spectrum data idle degree; Judge whether the spectrum data idle degree corresponding to each second frequency band in the second frequency band dataset is greater than the preset spectrum data idle degree; If less, eliminate the corresponding second frequency band; If greater than or equal to, use the corresponding second frequency band as an idle frequency band and generate an idle frequency band set; Obtain the frequency band mapping value of the idle frequency band in the high-dimensional feature space based on the preset kernel function; Obtain the feature vector of each idle frequency band in the high-dimensional space according to the frequency band mapping value; Calculate the relationship value between the idle frequency band and other frequency band data through the kernel function according to the feature vector; Fit the similar idle frequency bands based on the relationship value to obtain fitting data.

7. The data processing method based on machine learning according to claim 1, wherein The steps of obtaining the frequency band resource characteristics according to the fitting data through the inverse mapping technology include: Obtain multiple fitting frequency bands according to the fitting data; Perform inverse operation on the fitting frequency bands through the inverse mapping technology to obtain the spectrum characteristics of each fitting frequency band in the original space; Obtain the reduced-dimensional spectrum hole quantity and reduced-dimensional spectrum occupancy rate according to the spectrum characteristics; Obtain the association strength between each fitting frequency band according to the reduced-dimensional spectrum hole quantity and the reduced-dimensional spectrum occupancy rate; Classify each fitting frequency band based on the association strength and mark the classified fitting frequency bands to obtain a fitting frequency band set; Obtain the frequency band resource characteristics according to the fitting frequency band set, where the frequency band resource characteristics include the frequency band hole density and the frequency band occupancy rate.

8. A data processing system based on machine learning, characterized in that, Include: A first acquisition module for acquiring multiple initial data based on multiple data sources; A second acquisition module, configured to acquire the data compactness of each piece of initial data, classify the initial data based on the data compactness to obtain a plurality of data classification libraries; A third acquisition module, configured to perform data partitioning on each of the data classification libraries based on frequency band distributed nodes to obtain a plurality of data calculation nodes, and perform data processing on each data calculation node based on machine learning parallel data processing to obtain a first frequency band data set; A fourth acquisition module, configured to obtain a plurality of spectrum occupancy rates according to the first frequency band data set, and filter the first frequency band data set based on the spectrum occupancy rates to obtain a second frequency band data set; A fifth acquisition module, configured to acquire the number and distribution of spectrum holes of each frequency band data in the second frequency band data set, and obtain the spectrum data idle degree based on the number and distribution of spectrum holes; A sixth acquisition module, configured to filter the second frequency band data set to obtain an idle frequency band set, and map the idle frequency band set to a high-dimensional feature space to obtain fitting data; A processing module, configured to obtain frequency band resource characteristics according to the fitting data through an inverse mapping technique, and perform frequency band identification processing on the idle frequency band set based on the resource characteristics.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for determining spectrum resource utilization degree

    CN105992215A

  • Spectrum channel clustering method based on extreme learning machine

    CN111353530A

  • Wireless spectrum intelligent allocation and edge computing cooperation method

    CN119562364A

  • Digital information transmission sharing operation and maintenance platform based on AI technology

    CN119602900A

  • Channel selection method and transmit end

    US20190069299A1