Data analysis report generation method and device, storage medium and electronic equipment
By converting multi-source data into natural language text and dividing it into segment analysis tasks, and using standard protocol knowledge graphs and hybrid prediction models to generate data analysis reports, the problems of high difficulty in integrating multi-source data and high error rate in manual processing are solved, achieving efficient and accurate automatic generation of data analysis reports and improving network operation and maintenance efficiency.
Patent Information
- Application Number
- CN202511248298.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-19
AI Technical Summary
In the planning, construction, maintenance and operation of wireless communication networks, existing technologies suffer from difficulties in integrating multi-source data, high error rates in manual processing, and long processing times, resulting in low efficiency in generating data analysis reports, which affects network operation and maintenance efficiency. Furthermore, they lack automated predictive analysis of real-time network data and future trends.
By acquiring target multi-source data and converting it into first natural language text, the analysis report is divided into multiple segment analysis tasks based on the elements of the analysis report. A task dependency graph is constructed, dynamic prompt words are generated using a pre-built standard protocol knowledge graph, and then input into a pre-trained hybrid prediction model to generate a data analysis report.
It enables automated generation of data analysis reports, improving report generation efficiency, ensuring compliance and readability, supporting the reuse of historical analysis logic, processing real-time data and predicting future trends, and improving network operation and maintenance efficiency.
Smart Images

Figure CN121168433A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a data analysis report generation method, a data analysis report generation device, a computer storage medium and an electronic device. BACKGROUND
[0002] In the process of planning, construction, maintenance, optimization and operation of a wireless communication network, an operation team needs to generate a large number of professional data analysis reports, such as network performance reports, resource planning reports, fault diagnosis reports, etc. on a regular basis. The current mainstream data analysis report generation scheme is a manual writing mode supplemented by basic automated tool support.
[0003] However, in the above scheme, the data is scattered in isolated systems such as network management systems, road testing tools and customer platforms, and the formats are not uniform, which leads to great difficulty in integrating multi-source data, high error rate of manual processing, long time consumption, low report generation efficiency, and thus affects the network operation efficiency and the dynamic response capability of the system is insufficient. SUMMARY
[0004] The present disclosure provides a data analysis report generation method, a data analysis report generation device, a computer storage medium and an electronic device, thereby automatically generating data analysis reports, reducing the difficulty of integrating multi-source data and the error rate caused by manual processing, improving the report generation efficiency, thereby improving the network operation efficiency, and improving the dynamic response capability of the system.
[0005] In a first aspect, an embodiment of the present disclosure provides a data analysis report generation method, the method comprising: obtaining target multi-source data, and converting the target multi-source data into a first natural language text; performing task division according to analysis report elements associated with a preselected report template, obtaining a plurality of segment analysis tasks, and constructing a task dependency graph for the plurality of segment analysis tasks; determining a target standard protocol matched with a second natural language text associated with each segment analysis task from a preconstructed standard protocol knowledge graph according to the task dependency graph and the second natural language text, to generate a dynamic prompt word; wherein the second natural language text is a natural language text associated with each segment analysis task selected from the first natural language text; inputting the target multi-source data into a pre-trained hybrid prediction model, to output a prediction result based on the hybrid prediction model; and generating a data analysis report based on the prediction result and the dynamic prompt word.
[0006] In a second aspect, an embodiment of the present disclosure provides a data analysis report generation apparatus, which comprises: a text conversion module configured to obtain target multi-source data and convert the target multi-source data into first natural language text; a dependency graph construction module configured to divide tasks according to analysis report elements associated with a preselected report template, obtain a plurality of segment analysis tasks, and construct a task dependency graph for the plurality of segment analysis tasks; a prompt word generation module configured to determine, from a preconstructed standard protocol knowledge graph, a target standard protocol matching a second natural language text associated with each segment analysis task according to the task dependency graph and the second natural language text, to generate dynamic prompt words; wherein the second natural language text is a natural language text associated with each segment analysis task selected from the first natural language text; a prediction result generation module configured to input the target multi-source data into a pre-trained hybrid prediction model to output a prediction result based on the hybrid prediction model; and an analysis report module configured to generate a data analysis report based on the prediction result and the dynamic prompt words.
[0007] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the data analysis report generation method described above.
[0008] In a fourth aspect, an embodiment of the present disclosure provides an electronic device, which comprises: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the executable instructions to implement the data analysis report generation method described above.
[0009] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the data analysis report generation method described above.
[0010] The technical solution of the present disclosure has the following beneficial effects:
[0011] The data analysis report generation method described above comprises the following steps: obtaining target multi-source data and converting the target multi-source data into first natural language text; dividing tasks according to analysis report elements associated with a preselected report template, obtaining a plurality of segment analysis tasks, and constructing a task dependency graph for the plurality of segment analysis tasks; determining, from a preconstructed standard protocol knowledge graph, a target standard protocol matching a second natural language text associated with each segment analysis task according to the task dependency graph and the second natural language text, to generate dynamic prompt words; wherein the second natural language text is a natural language text associated with each segment analysis task selected from the first natural language text; inputting the target multi-source data into a pre-trained hybrid prediction model to output a prediction result based on the hybrid prediction model; and generating a data analysis report based on the prediction result and the dynamic prompt words.
[0012] The method converts the collected target multi-source data into first natural language text, so as to directly use the natural language text to generate a data analysis report, and improve the readability of the data analysis report. The method is divided into multiple segment analysis tasks according to analysis report elements, and second natural language text matched with the first natural language text is collected from the first natural language text based on the dependency relationship between the segment analysis tasks, and a dynamic prompt word generated for real-time data flow is generated, so that the data analysis report automatic generation process is realized, thereby solving the technical problem of the existing technology that the format is not unified in the manual data report generation scheme, resulting in great difficulty in multi-source data integration, high error rate of manual processing, long time consumption, low report generation efficiency, and low network operation efficiency. The technical effects of automatically generating a data analysis report, improving the report generation efficiency, and improving the network operation efficiency are achieved. That is, the method provides a data analysis report automatic generation process with high efficiency, accuracy, and decision support.
[0013] In addition, based on the analysis report elements (abstract / method / result, etc.), the method is automatically disassembled into segment analysis tasks, and a vector database is embedded with 3GPP protocols, ITU protocols, industry white papers, and other knowledge bases to dynamically associate technical standards for each task. Each analysis segment has professional compliance, and the reuse of historical analysis logic is supported, which greatly improves the generation efficiency of similar reports. The method can also dynamically predict based on the collected target multi-source data, thereby overcoming the technical problems that traditional report generation tools cannot process real-time network data and lack automatic prediction analysis of future trends. The finally generated data analysis report not only contains the analysis report of the current data, but also contains the prediction analysis report, which facilitates subsequent layout for risk problems and improves network operation efficiency.
[0014] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings, which are incorporated into and form part of the specification, illustrate exemplary embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.
[0016] Figure 1 An application architecture diagram of one data analysis report generation system in the present exemplary embodiment is schematically shown;
[0017] Figure 2FIG. 1 shows a flow chart of a method for generating a data analysis report according to an example embodiment of the present disclosure;
[0018] Figure 3 FIG. 4 shows a flow chart of a method for selecting a hybrid prediction model according to an example embodiment of the present disclosure;
[0019] Figure 4 FIG. 5 shows a flow chart of a method for generating a data analysis report according to an example embodiment of the present disclosure;
[0020] Figure 5 FIG. 6 shows a schematic diagram of a data analysis report generation device according to an example embodiment of the present disclosure;
[0021] Figure 6 FIG. 7 shows a schematic diagram of an electronic device according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout. The terminology used in the description presented herein is not intended to be interpreted in any specific and / or particular manner. The terminology utilized in the present disclosure is for the purpose of describing exemplary embodiments only and is not intended to be limiting. The use of the terms "example" and / or "exemplary implementation" in the description is intended to illustrate at least one example embodiment. Such terms should not be construed to foreclose other embodiments that can fall within the scope of the present disclosure. Rather, such terms are used in conjunction with the phrase "in one example embodiment" and / or "in an example embodiment" to indicate that the feature, structure, and / or characteristic being described is included in at least one embodiment.
[0023] In addition, the drawings of the present disclosure are to be used only as illustrative tools and not as a definition of the scope of the present disclosure. They are provided solely for illustration of certain examples of the present disclosure and should not be taken as a definition of the scope of the present disclosure. Thus, the scope of the present disclosure is to be determined solely by the appended claims and their legal equivalents, whereas the examples shown in drawings are for illustrative purposes only and should not be construed as limiting the scope of the present disclosure. The use herein of "including," "comprising," "having," "containing," "involving," "decomposing," "carrying" or variations thereof is meant to encompass the items listed thereafter and any other items not specifically listed. Subject matter including "comprising," "having," "containing," "including," "carrying," "decomposing," "carrying" or variations thereof does not, without more constraints, preclude the inclusion of additional items regardless of whether more constraints are listed in the claims. The terms "a" and "an" and "the" and "at least one" and "one or more" are used interchangeably and mean "one or more than one" or "at least one (item)." The terms "plurality" and "a plurality" mean "two or more" or "at least two" (of the referenced item). The terms "comprises", "comprising", "including", "includes", "contains", "containing", "has", "having" or variants thereof are used synonymously with each other in this disclosure and mean it is "comprised of" or "including", "contains" or "has" the specified features or steps but not excluding the presence of other features or steps. The term "coupled" means the direct or indirect coupling between elements, which can be physical or logical.
[0024] The flow charts shown in the drawings are only illustrative and do not necessarily include all the steps. For example, some steps can be further decomposed, and some steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.
[0025] In the related art, in the planning, construction, maintenance, optimization and operation process of a wireless communication network, an operation team needs to generate a large amount of professional data analysis reports regularly, thereby facilitating subsequent network operation and maintenance of an operator. For example, network performance reports (for example, Key Performance Indicator (KPI) statistical reports, network coverage analysis reports, interference positioning analysis reports), resource planning reports (for example, capacity prediction reports, base station deployment suggestion reports), fault diagnosis reports (for example, root cause analysis reports, solution suggestion reports).
[0026] The mainstream solution in the current industry mainly relies on manual writing supplemented by a small amount of automatic tools to complete the generation of data analysis reports. However, the above method at least has the following technical problems:
[0027] 1) Difficulty in integrating multi-source heterogeneous data: Data is often scattered in different network management systems, road testing tools, customer platforms and other isolated systems, and the data formats are not uniform. The time consumption of manual data collection and cleaning accounts for more than 60%, and it is also prone to errors, resulting in poor reliability of the generated data analysis report.
[0028] 2) Highly dependent on manual experience for report writing: Report writing requires engineers to have both wireless communication professional knowledge (for example, 3rd Generation Partnership Project (3GPP) protocol) and writing ability, and the repetitive workload of similar reports is large (for example, full network KPI analysis reports need to be generated every month), but the reuse rate of analysis logic is low, resulting in high difficulty in generating data analysis reports.
[0029] 3) Insufficient dynamic data response capability: Traditional report generation tools cannot process real-time network data, and lack automatic prediction analysis of future trends.
[0030] Embodiments of the present disclosure consider the above problems and propose a data analysis report generation method which can be applied to any application scenario that needs data analysis, especially the generation of data analysis reports. The method converts the target multi-source data into a first natural language text after collecting it, so that the subsequent data analysis report can be generated directly using natural language text, improving the readability of the data analysis report. Moreover, the method is divided into multiple segment analysis tasks according to the analysis report elements, and the second natural language text matching the first natural language text is collected from the first natural language text based on the dependency relationship between them, and then the dynamic prompt words generated for real-time data flow are generated. The process realizes the automatic generation of the data analysis report, thereby solving the technical problems of the existing technology using artificial data report generation scheme, such as inconsistent format, difficulty in integrating multi-source data, high error rate of manual processing, long time consumption, low report generation efficiency, and impact on network operation efficiency. The technical effects of automatically generating data analysis reports, improving report generation efficiency, and improving network operation efficiency are achieved. That is, the method provides an efficient and accurate data analysis report automatic generation process that provides decision support.
[0031] In addition, based on the analysis report elements (abstract / method / result, etc.), the method is automatically disassembled into a segmented analysis task, and the vector database is embedded with 3GPP protocols, ITU protocols, industry white papers, and other knowledge bases to dynamically associate technical standards for each task. Ensure that each analysis segment has professional compliance, while supporting the reuse of historical analysis logic to significantly improve the efficiency of similar report generation. Moreover, the method can also dynamically predict based on the collected target multi-source data, thereby overcoming the technical problems of traditional report generation tools that cannot handle real-time network data and lack automated prediction analysis of future trends. The final generated data analysis report not only contains the analysis report of the current data, but also contains the prediction analysis report, thereby facilitating subsequent layout for risk issues and improving network operation efficiency.
[0032] The present disclosure proposes a data analysis report generation method and device which can be applied to Figure 1 the system architecture of the exemplary application environment shown.
[0033] As Figure 1 shown, the system architecture 100 can include a terminal device 101, a server 102. The terminal device 101 communicates with the server 102 through a network 103, and a data storage system 104 can store data required by the server 102 for processing. The data storage system 104 can be integrated on the server 104, or placed on the cloud or other servers.
[0034] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, there can be any number of terminal devices, networks, and servers. For example, server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0035] The terminal device 101 may include, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices, while portable wearable devices may include smartwatches, smart bracelets, and head-mounted devices. It should be noted that the data analysis report generation method provided in this embodiment can be executed independently on the server 102; correspondingly, the data analysis report generation device is generally located in the server 102. The data analysis report generation method provided in this embodiment can also be executed independently on the terminal device 101; correspondingly, the data analysis report generation device can also be located in the terminal device 101. The data analysis report generation method provided in this embodiment can also be executed collaboratively by the terminal device 101 and the server 102, for example, partially executed on the server 102 and partially executed on the terminal device 101. Correspondingly, some modules of the data analysis report generation device can be located in the server 102, and some modules can be located in the terminal device 101.
[0036] For example, in one exemplary embodiment, server 102 may receive data from terminal device 102, base station (…). Figure 1 The process involves: acquiring target multi-source data (not shown in the diagram) and converting it into first natural language text; dividing the data into multiple segment analysis tasks based on the analysis report elements associated with a pre-selected report template; constructing a task dependency graph for each segment analysis task; determining the target standard protocol matching the second natural language text associated with each segment analysis task from a pre-constructed standard protocol knowledge graph based on the task dependency graph and the second natural language text; generating dynamic prompts; wherein the second natural language text is selected from the first natural language text and associated with each segment analysis task; inputting the target multi-source data into a pre-trained hybrid prediction model to output prediction results based on the hybrid prediction model; and generating a data analysis report based on the prediction results and dynamic prompts.
[0037] However, it is easy for those skilled in the art to understand that the above application scenarios are only for example, and the present example embodiment is not limited thereto.
[0038] Next, the data analysis report generation method will be described below by taking the server 102 as an example. Figure 2 The flowchart of the data analysis report generation method in the present example embodiment is schematically shown in FIG. 2. Figure 2 The data analysis report generation method provided by the present embodiment includes the following steps S201-S205.
[0039] In step S201, the target multi-source data is obtained and converted into a first natural language text.
[0040] In step S202, the analysis report elements associated with the pre-selected report template are used to divide the tasks, and a plurality of segment analysis tasks are obtained. A task dependency graph is constructed for the plurality of segment analysis tasks.
[0041] In step S203, the task dependency graph and the second natural language text are used to determine the target standard protocol matched with the second natural language text associated with each segment analysis task from the pre-constructed standard protocol knowledge graph, to generate dynamic prompt words; wherein the second natural language text is the natural language text associated with each segment analysis task selected from the first natural language text.
[0042] In step S204, the target multi-source data is input into the pre-trained hybrid prediction model, to output a prediction result based on the hybrid prediction model.
[0043] In step S205, the data analysis report is generated based on the prediction result and the dynamic prompt words.
[0044] In the present example embodiment, the data analysis report generation method is applied to the server 102. Figure 2In the technical solution provided, the target multi-source data is converted into first natural language text after being collected by the method, so that the natural language text is directly used for generating a data analysis report subsequently, improving the readability of the data analysis report. Moreover, the method is divided into multiple segment analysis tasks according to analysis report elements, so as to collect second natural language text matched with the segment analysis tasks from the first natural language text based on the dependency relationship therebetween, and then generate dynamic prompt words for real-time data stream, which realizes the automatic generation process of the data analysis report, thereby solving the technical problems of the prior art, such as the non-uniform format, the difficulty in multi-source data integration, the high error rate of manual processing, the long time consumption, the low report generation efficiency, and the impact on network operation efficiency. The technical effects of automatically generating the data analysis report, improving the report generation efficiency, and improving the network operation efficiency are achieved. That is, the method provides a data analysis report automatic generation process with high efficiency, accuracy, and decision support.
[0045] In addition, the analysis report elements (abstract / method / result, etc.) are automatically disassembled into segment analysis tasks, and the vector database is embedded with knowledge bases such as 3GPP protocols, ITU protocols, and industry white papers, so as to dynamically associate technical standards for each task. Each analysis segment has professional compliance, and the reuse of historical analysis logic is supported, so that the generation efficiency of similar reports is greatly improved. Moreover, the method can dynamically predict based on the collected target multi-source data, thereby overcoming the technical problems of the conventional report generation tools, such as the inability to process real-time network data and the lack of automatic prediction analysis of future trends. The finally generated data analysis report not only contains the analysis report of the current data, but also contains the prediction analysis report, thereby facilitating the layout of risk problems in advance subsequently and improving the network operation efficiency.
[0046] The specific implementation of each step in the embodiments will be described in detail below in combination with specific embodiments. Figure 2 The specific implementation of each step in the embodiments will be described in detail below in combination with specific embodiments.
[0047] In step S201, target multi-source data is acquired, and the target multi-source data is converted into first natural language text.
[0048] The target multi-source data is collected from multiple input sources. In the application scenario of a communication network, the input sources may be, for example, a network management system, a road testing platform, a user complaint database, a base station, or a manually input operation and maintenance work order text, etc., and the present disclosure does not make any special limitation or enumeration.
[0049] For example, the target multi-source data may be data collected directly from multiple input sources, or data preprocessed from data collected from multiple input sources.
[0050] Next, the process of collecting data from multiple input sources and preprocessing the collected initial source data to obtain target multi-source data will be exemplarily described in combination with the following embodiments.
[0051] In an optional embodiment of the present disclosure, data is collected from multiple network data source systems based on a standardized interface to obtain initial multi-source data; the initial multi-source data is preprocessed to obtain target multi-source data.
[0052] Among them, the multiple network data sources are the multiple input sources shown in the above embodiments.
[0053] Exemplarily, data can be collected from multiple network data source systems through a standardized Application Programming Interface (API interface), i.e., a standardized interface, to obtain raw and unprocessed initial multi-source data.
[0054] Among them, the step of collecting data from multiple network data source systems can include at least two of the following:
[0055] Collecting real-time data streams for network parameters from at least one network management system.
[0056] Obtaining historical log texts of base stations from a distributed file system.
[0057] Obtaining operation and maintenance work order texts for communication networks.
[0058] Exemplarily, from the at least one network management system, such as a network management system, a road testing platform, a customer platform (e.g., a user complaint database), etc., key performance indicators (KPIs) for network performance of each generation network (e.g., 4G / 5G / 6G network, etc.) can be collected, such as reference signal receiving power (RSRP), signal to interference plus noise ratio (SINR), throughput, latency, etc.
[0059] For example, real-time stream data, such as KPI indicators, can be collected from at least one network management system every 3 minutes based on a Kafka interface.
[0060] Optionally, historical log text of the base station can also be obtained from a distributed file system. For example, hourly traffic statistics, hardware status code fields, and other historical logs of the base station can be obtained from HDFS (Hadoop Distributed File System).
[0061] It should be noted that, in order to ensure data availability, continuous data that meets or exceeds a preset time period can be obtained, for example, continuous data that meets or exceeds 6 months.
[0062] Optionally, manual input of maintenance work order text can also be obtained. Maintenance work order text is usually required to conform to the predefined XML Schema (XSD) specification, which is used to define the structure, data types and constraint rules of XML documents.
[0063] Accordingly, in an optional embodiment of this disclosure, the initial multi-source data undergoes data preprocessing, including one or more of the following:
[0064] A. In response to the initial source data containing a real-time data stream collected from at least one network management system, perform anomaly detection on the actual data stream according to the dynamic sliding window algorithm and / or the isolated forest algorithm to obtain the first target data.
[0065] The sliding window size in the dynamic sliding window algorithm is determined based on the network type corresponding to the network management system. Taking 4G and 5G networks as examples, the calculation process of the sliding window size can be referred to the following formula (1):
[0066] w size =K*(α*BW) 5G +(1-α)*BW 4G ) Formula (1)
[0067] In formula (1), w size The sliding window size; α is the proportion of 5G traffic; K is the baseline coefficient; BW 5G For 5G bandwidth parameters; BW 4G This refers to 4G bandwidth parameters. In this embodiment, dynamically adjusting the sliding window size can improve the efficiency and accuracy of anomaly detection.
[0068] Simultaneously, the isolated forest algorithm can be used to check for mutations in real-time data streams and to verify the rationality of the real-time physical layer as a key performance indicator of network performance. For example, the effective range of RSRP is limited to [-140, -44] dBm. If the RSRP value collected in the real-time data stream exceeds the above-defined effective range, it is determined to be abnormal / mutational data and is then removed.
[0069] B, in response to the initial source data contains from the distributed file system to obtain the historical log text, according to the conditional random field CRF model, the mapping relationship between each log code and network semantic event in the historical log text is established, and the second target data is obtained.
[0070] Among them, when the mapping relationship between each log code and network semantic event is established by the CRF model, in order to match the network semantics more, the precision of network semantic event can be improved by fusing 3GPP terminology library.
[0071] Exemplarily, in the case of obtaining the historical log text of the base station from the distributed file system, the conditional random field (CRF) model can be used to structure the historical log text, and the mapping relationship between each log code and network semantic event in the historical log text is established by the feature engineering of part of speech tagging, regular matching and matching 3GPP terminology library. It should be explained that the core of feature engineering is to establish the association between the structure of the above log code and the network semantic event by designing feature function.
[0072] For example, assuming that a log code in the historical log text is: "ERR_CODE=0x3A", and the network semantic event mapped with it is "power amplifier overload". Similarly, when determining the network semantic event, the technical compliance of the network semantic event can be checked and standardized by combining the communication standard protocol in the database.
[0073] C, the first target data and / or the second target data and / or the operation and maintenance order text are aligned in time and space dimensions, and the aligned target multi-source data is obtained.
[0074] Exemplarily, in order to realize the structured storage and unification of cross-system data, the cross-system data contained in the initial source data can be aligned in time and space dimensions. Time and space dimension alignment, namely, alignment from time dimension and space dimension respectively.
[0075] Assuming that the initial source data contains the above processed first target data, second target data and operation and maintenance order text, the above three kinds of cross-system data can be automatically aligned and stored in time and space dimensions.
[0076] In the execution of the above-mentioned spatio-temporal dimension alignment, in an optional embodiment of the present disclosure, the spatio-temporal dimension alignment is performed on the first target data and / or the second target data and / or the operation and maintenance work order text, including: based on the spatial dimension, using the GeoHash algorithm to encode the spatial position, and using the Voronoi diagram to determine the cell coverage boundary position; based on the time dimension, using cubic spline interpolation to align the first target data and / or the second target data and / or the operation and maintenance work order text with different sampling rates.
[0077] For example, the geographic space processing adopts GeoHash encoding base station position, and calculates the cell coverage boundary through the Voronoi diagram algorithm; the time dimension processing uses cubic spline interpolation to align the data streams with different sampling rates.
[0078] In this embodiment, by performing spatio-temporal dimension alignment processing on multi-source data, the challenges brought by data dispersion and heterogeneity can be solved, the synonymous fields in the data collected by different input sources are associated, the same entity in different input sources is identified, and different frequency time stamps are aligned, so that unified understanding, fusion and efficient utilization of multi-source data are facilitated, thereby eliminating the data island problem and solving the problems of low data collection efficiency and easy errors caused by cross-system collection and non-uniform format in the related technical solutions.
[0079] It should be emphasized that, through the above preprocessing operation, the output target multi-source data needs to meet the quality standards of: field missing rate < 0.1% (for example, the missing rate of the field can be determined by the 6σ criterion), and label accuracy ≥ 99.5% (for example, the accuracy of the label can be determined by the cross-validation method), so as to provide a reliable data basis for the generation of subsequent data analysis reports.
[0080] In addition to the above-mentioned embodiments, the initial source data can also be subjected to data cleaning, normalization, standardization and other preprocessing processes by using the ETL process. It should be explained that the ETL process refers to the key steps in data processing, including three stages of extraction (Extract), transformation (Transform) and loading (Load), which aims to process raw data into structured forms available for business. Extraction (Extract): Extracting raw data from source systems (such as databases, APIs or files), involving data source identification, connection establishment and extraction strategy formulation to ensure efficient acquisition of required data sets. Transformation (Transform): Cleaning (such as handling missing values or duplicates), calculation (such as generating new fields) and integration (such as associating multi-source data) of extracted data to address data quality issues and improve consistency. Loading (Load): Writing the transformed data into target systems (such as data warehouses or data marts) to support business analysis, reporting and decision-making applications. This process efficiently integrates scattered data to provide a reliable foundation for data analysis.
[0081] In this embodiment, by obtaining initial multi-source data from multiple input sources, and then respectively performing a series of preprocessing on the initial multi-source data, the data island problem is solved, the automatic alignment and structured storage of cross-system data are realized, high-quality input is provided for subsequent analysis, and the automatic collection and processing process reduces the time consumption of manual data arrangement by more than 70%, thereby improving the efficiency of generating data analysis reports.
[0082] Based on the above embodiment, after obtaining the preprocessed target multi-source data, the target multi-source data can be subjected to labeling / semantic processing, so that the data collected from network management systems, road testing platforms and other systems, historical log texts, and manually inputted operation and maintenance work order texts are all converted into first natural language texts to be semantically labeled.
[0083] For example, the RSRP measurement value is converted into a natural language description with semantics such as "cell A: RSRP = -85dBm (good)". It should be explained that, in the above semantic processing process, the technical compliance of the natural language description can also be checked and standardized by combining the communication standard protocols in the database. The communication standard protocols can be, for example, the 3rd Generation Partnership Project (3GPP protocol), the International Telecommunication Union (ITU protocol), industry white papers, etc. It should be noted that the ITU protocol and the 3GPP protocol are complementary: the ITU protocol defines the spectrum policy and global compatibility, and the 3GPP protocol implements the technical details. For example, the IMT-2020 standard of the ITU and the 3GPP 5G NR together constitute the complete 5G ecology.
[0084] In step S202, task division is performed according to the analysis report elements associated with the preselected report template, a plurality of segment analysis tasks are obtained, and a task dependency graph is constructed for the plurality of segment analysis tasks.
[0085] Among them, the analysis report element is one or more sub-elements corresponding to the overall report framework of the report template, for example, the analysis report element can be the elements such as abstract, introduction, background, method, result, discussion and conclusion that constitute the data analysis report; the task is the data analysis report generation task.
[0086] For example, a report template can be selected from a pre-defined report template library, and the overall structure of the selected report template is explicitly selected to perform framework analysis on the overall structure of the report template, and obtain the analysis report elements associated with the analysis report. The generation task of the data analysis report is divided into a plurality of segment analysis tasks according to the analysis report elements.
[0087] Optionally, the report template is described in JSON-LD format. JSON-LD (JSON for Linked Data) format is a JSON-based semantic data format designed to represent and publish Linked Data, enhancing the machine understanding of data by introducing context and vocabulary. Taking the JSON-LD format description corresponding to the network performance analysis report of interference analysis as an example, an exemplary fragment is as follows:
[0088]
[0089] The "data_requirements" in the above description represents data requirements, "cell_id" is an identifier, "interference_type" represents interference type, and "knowledge_dependencies" represents standard protocol basis. For example, "3GPP TS 36.213 Section 5.2" represents the conclusion obtained according to the analysis of Section 5.2 of TS 36.213 of the 3GPP protocol.
[0090] The task dependency graph is a visual model for representing the logical relationship between multiple segment analysis tasks, and its core is to describe the execution order and constraint conditions of the tasks through nodes and directed edges. The node is a plurality of segment analysis tasks, the directed edge represents the dependency relationship between the plurality of segment analysis tasks, and the arrow direction of the directed edge is used to define the execution order between the tasks. For example, the directed edge of "method task→conclusion task" indicates that the conclusion task needs to wait for the method task to complete before execution, i.e., the conclusion can be obtained after the method is written.
[0091] For example, based on each segment analysis task in the task dependency graph, a second natural language text matching each segment analysis task can be determined from the first natural language text. It can be understood that the second natural language text matching each segment analysis task can be a natural language text with repetition or a natural language text without repetition, and the embodiments of the present disclosure do not make any special limitation thereon, and the specific determination can be determined according to the actual situation.
[0092] In an optional embodiment, the improved C4.5 decision tree algorithm can be used to construct the task dependency graph. It should be explained that the C4.5 decision tree algorithm is a classic classification algorithm in machine learning, and its core feature is to select attributes through information gain rate and introduce pruning mechanism to improve generalization ability. Of course, other classification algorithms can also be used to construct the task dependency graph based on the plurality of segment analysis tasks, and the embodiments of the present disclosure do not make any special limitation and enumeration thereon.
[0093] In an optional embodiment, the edge weight of the directed edge of the task dependency graph constructed by the embodiments of the present disclosure can be determined by the data confidence of the target multi-source data and the business priority. For example, the data confidence can be determined by the integrity of the RSRP field, and the priority of each service can be determined based on the weighting coefficient of each VIP cell.
[0094] Optionally, when some parameters of the target multi-source data are missing, an intelligent completion strategy can be triggered, for example, a sliding window of the last 24 hours is used for smoothing to complete the missing parameters.
[0095] In step S203, a target standard protocol matched with the second natural language text associated with each segment analysis task is determined from a pre-constructed standard protocol knowledge graph according to the task dependency graph and the second natural language text, to generate a dynamic prompt word; wherein the second natural language text is a natural language text associated with each segment analysis task selected from the first natural language text.
[0096] The standard protocol knowledge graph is constructed by standardized entities in the field of communication networks, for example, the standard protocol knowledge graph can be composed of 120,000 standardized entities. For example, the standard protocol knowledge graph can be a 3GPP knowledge graph constructed based on the 3GPP protocol.
[0097] Exemplarily, the 3GPP knowledge graph can be constructed based on Neo4j. Neo4j stores data using a node-relation model, supports dynamic attribute key-value pairs, and has a built-in Cypher query language to implement efficient graph traversal operations.
[0098] In an optional embodiment of the present disclosure, in response to the update of the standard protocol, the standard protocol knowledge graph is updated based on the query encoder using the BERT model to obtain an updated standard protocol knowledge graph; and the target standard protocol matched with the second natural language text associated with each segment analysis task is determined based on the document encoder through the updated standard protocol knowledge graph.
[0099] Exemplarily, a double-encoder retrieval architecture can be used for the standard protocol knowledge graph. Specifically, the query encoder is based on BERT-base fine-tuning, and the document encoder outputs a 768-dimensional dense vector to realize double-channel update of the knowledge iteration layer.
[0100] Exemplarily, as the network is continuously updated and iterated, the 3GPP standard protocol is also continuously updated. When the system detects that the standard protocol is updated, the standard protocol knowledge graph can be updated by using the BERT model to query the encoder, and the updated standard protocol knowledge graph is obtained; and the document encoder determines the target standard protocol matched with the second natural language text associated with each segment analysis task through the updated standard protocol knowledge graph.
[0101] Exemplarily, the millisecond-level response can also be realized by the FAISS clustering index. The FAISS clustering index is used to realize the update and iteration of the standard protocol knowledge graph, that is, the BERT-base fine-tuning can be used to automatically monitor the 3GPP standard update and extract new constraint clauses through the BERT-QA model, so as to update and iterate the standard protocol knowledge graph according to the update and iteration of the 3GPP standard. At the same time, the excellent report segment that passes the artificial review is stored in the FAISS index library in the form of vector, so as to realize the millisecond-level response through the FAISS clustering index.
[0102] In an optional embodiment of the present disclosure, after the dynamic prompt word is generated, the method comprises: optimizing the dynamic prompt word based on a T5-3B model to obtain an optimized dynamic prompt word.
[0103] Exemplarily, in order to further optimize the generated dynamic prompt word, the dynamic prompt word can be optimized based on a T5-3B model.
[0104] It should be explained that the core design concept of the T5-3B (Text-to-Text Transfer Transformer 3B) model is to unify all natural language processing (NLP) tasks into a "text-to-text" conversion framework. Thus, the natural language polishing is performed through the T5-3B model.
[0105] In the above process of generating the dynamic prompt word, three quality inspection levels can be set: the first level verifies the integrity of the mandatory fields, the second level detects the contradictions through the knowledge graph reasoning technology (such as the power configuration violating TS38.104), and the third level performs the final check before outputting the dynamic prompt word (for example, the key indicator coverage rate is greater than or equal to 99%). Tests show that the above embodiments are significantly better than the baseline system in terms of task decomposition accuracy (98.2%) and knowledge retrieval recall rate (95.7% @Top5).
[0106] In step S204, the target multi-source data is input into the pre-trained hybrid prediction model to output a prediction result based on the hybrid prediction model.
[0107] Exemplarily, the real-time prediction process adopts a hybrid prediction framework to achieve performance analysis of multi-scale networks.
[0108] Before the target multi-source data is input into the pre-trained hybrid prediction model in step S204, the target multi-source data can also be further processed by a device. For example, historical KPI data (assuming more than 6 months and 15-minute granularity) is time-aligned by three times of spline interpolation, real-time data stream (5-minute granularity) is spatio-temporally associated with geographic information (GeoHash 150-meter grid) after being accessed via Kafka, and external factors such as weather API and holiday calendar are integrated. Feature extraction is performed, such as automatic construction of lag features (12-period flow backtracking) and composite indicators (such as flow-PRB utilization ratio), and the mutual information method is used to screen the most predictive features.
[0109] Further, in step S204, in an optional embodiment of the present disclosure, the TFT model can be used as the hybrid prediction model when the prediction length is less than 24 hours, and the Prophet+XGBoost combined model can be used as the hybrid prediction model when the prediction length is greater than or equal to 24 hours. Figure 3 The method can further include steps S301-S306.
[0110] Step S301, determining the prediction length.
[0111] Step S302, determining whether the prediction length is a first prediction length.
[0112] In response to the prediction length being the first prediction length, step S303 is performed, in which the TFT model is used as the hybrid prediction model, and step S304 is performed, in which the prediction result is output based on the TFT model.
[0113] Otherwise, in response to the prediction length being a second prediction length, step S305 is performed, in which the Prophet+XGBoost combined model is used as the hybrid prediction model, and step S306 is performed, in which the prediction result is output based on the Prophet+XGBoost combined model.
[0114] The first prediction length is less than the second prediction length.
[0115] Exemplarily, the prediction process can adopt a hybrid architecture design: for short-term prediction of the first prediction length (for example, the first prediction length is less than 24 hours), the TFT model can be used to process complex multi-element time series dependence. For medium and long-term prediction of the second prediction length (for example, the second prediction length is 1-3 months), the Prophet+XGBoost combined model can be used to capture macro trends.
[0116] It can be explained that the Prophet+XGBoost combined model is a prediction method that combines the advantages of time series decomposition and gradient boosting trees. Prophet decomposes the time trend, and then the result is used as a feature input to XGBoost for nonlinear modeling.
[0117] The Prophet function extracts explicit patterns from time series data using decomposable additive models (e.g., trend + seasonal + holiday components), supporting robustness against missing values and outliers. XGBoost utilizes the decomposition results from Prophet (e.g., trend residuals, seasonal components) and other external features (e.g., weather, promotional activities) to capture nonlinear relationships through gradient boosting of decision trees.
[0118] For example, in the training process of a hybrid prediction model, an incremental learning strategy can be implemented, such as weekly full training combined with daily online fine-tuning, and determining the optimal hyperparameters through Bayesian optimization (50 iterations). The loss function can be optimized using a combination of Pinball Loss and MAE.
[0119] In an optional embodiment, the prediction results can generate probability intervals using Monte Carlo Dropout and output a 5%–95% confidence range. An anomaly detection module uses an isolated forest to flag predictions exceeding a 3σ threshold. Finally, a structured prediction report is output, which typically includes point estimates and confidence intervals for key metrics (e.g., traffic prediction of 152.7 Mbps [138.2, 167.5]), trend analysis (increasing / decreasing), and an interpretation of impact factors based on SHAP values (e.g., user growth contribution +12.3%).
[0120] Furthermore, in an optional embodiment of this disclosure, the prediction result is output based on the hybrid prediction model, including: determining the risk object based on the hybrid prediction model, and generating early warning information with time stamps based on the risk object.
[0121] For example, when risks such as exceeding capacity limits are predicted, a time-stamped warning is automatically triggered (e.g., "Carrier expansion is required on June 28th"). In live network verification, this solution achieves a short-term prediction error of ≤5% (MAPE) and processing latency controlled within 30 seconds (P99), reducing prediction deviation by 42% compared to traditional methods.
[0122] In step S205, a data analysis report is generated based on the prediction results and dynamic prompts.
[0123] For example, when performing this step, a large language model can be used to generate a data analysis report based on the prediction results and dynamic prompts.
[0124] Professional-level report automation production can be implemented using a three-stage generation architecture. For example, in the first draft generation stage, a domain fine-tuned model based on the Deepseek architecture (continuously pre-trained on a 50 GB wireless communication corpus, extended with 2,143 professional terms) receives structured input data, including predicted indicators (with confidence intervals), 3GPP protocol references (accurate to section numbers), and task constraint conditions. The model controls output stability through a temperature parameter (temperature = 0.3) and dynamically embeds enhanced prompt word templates (such as "follow IEEE style, reference 3GPP TS 38.214 v16.4.0 power control requirements").
[0125] The following will refer to Figure 4 The entire data analysis report generation process of the data analysis report generation method of the exemplary embodiments of the present disclosure will be described in detail.
[0126] First, a multi-source database integration and preprocessing process is performed.
[0127] Exemplarily, heterogeneous data sources such as network management systems, road testing platforms, and user complaint databases can be connected through standardized API interfaces, and ETL processes can be used to clean, normalize, and labelize the original data, such as converting RSRP measurement values into "cell ARSRP = -85 dBm (good)" and other semantic natural language descriptions. The data island problem is solved, and automatic alignment and structured storage of cross-system data are achieved, providing high-quality input for subsequent analysis and reducing manual data arrangement time by more than 70%.
[0128] Then, task decomposition and knowledge enhancement are performed.
[0129] Exemplarily, the task decomposition and knowledge enhancement step can achieve intelligent arrangement of data analysis tasks, and its core architecture includes two subsystems that work together: a task decomposition subsystem and a knowledge enhancement subsystem. The task decomposition subsystem receives preprocessed standardized data and a pre-defined report template library. Exemplarily, an improved C4.5 decision tree algorithm can be used to construct a task dependency graph, and the edge weight is determined by data confidence (such as RSRP field integrity) and business priority (VIP cell weighting coefficient). When parameter missing is detected, an intelligent completion mechanism is triggered (such as using the last 24-hour time window by default). The knowledge enhancement subsystem relies on a 3GPP knowledge graph (containing 120,000 standardized entities) built by Neo4j, and uses a dual-encoder retrieval architecture: the query encoder is based on BERT-base fine-tuning, and the document encoder outputs a 768-dimensional dense vector, which is implemented through a FAISS clustering index to achieve millisecond-level response. The retrieval results are injected into the generation process through dynamic prompt word templates (such as "refer to 3GPP TS 38.331 section 6.2.3 for analysis of handover failure"), and are polished in natural language by a T5-3B model.
[0130] In this process, a three-level quality inspection checkpoint can be set up for verification. Specifically, the first level verifies the completeness of required fields, the second level detects technical inconsistencies through knowledge graph reasoning (such as power configuration violating TS 38.104), and the third level performs a final check before output (key indicator coverage ≥99%). Tests show that this module significantly outperforms the baseline system in both task decomposition accuracy (98.2%) and knowledge retrieval recall (95.7%@Top5).
[0131] Next, real-time predictive analysis is performed.
[0132] For example, in this real-time predictive analysis step, a hybrid predictive framework can be used to achieve multi-scale network performance analysis. The system first performs deep preprocessing on the input heterogeneous data: historical KPI data (requiring more than 6 months, 15-minute granularity) is time-aligned using cubic spline interpolation; real-time data streams (5-minute granularity) are spatiotemporally correlated with geospatial information (GeoHash 150-meter grid) after being accessed via Kafka, while also integrating external factors such as weather APIs and holiday calendars. During the feature engineering stage, lagging features (12-cycle traffic retrospective) and composite indicators (such as traffic-PRB utilization ratio) are automatically constructed, and the mutual information method is used to select the top 20 most predictive features. The prediction engine adopts a hybrid architecture design: short-term predictions (<24 hours) use Temporal Fusion Transformer to handle complex multivariate time-series dependencies, while medium- and long-term predictions (1-3 months) use a Prophet+XGBoost combined model to capture macro trends. The model training employs an incremental learning strategy, combining weekly full training with daily online fine-tuning. Optimal hyperparameters are determined through Bayesian optimization (50 iterations), and the loss function is optimized using a combination of Pinball Loss and MAE. Prediction results generate probability intervals using Monte Carlo Dropout and output a 5%-95% confidence range. The anomaly detection module uses isolated forests to mark predicted values exceeding the 3σ threshold. The system ultimately outputs a structured prediction report, including point estimates and confidence intervals for key indicators (e.g., traffic prediction of 152.7 Mbps [138.2, 167.5]), trend analysis (increasing / decreasing), and interpretation of impact factors based on SHAP values (e.g., user growth contribution +12.3%). When risks such as capacity overruns are predicted, a time-stamped warning is automatically triggered (e.g., "Carrier expansion required on June 28th"). In live network verification, this scheme achieves a short-term prediction error ≤5% (MAPE) and processing latency controlled within 30 seconds (P99), reducing prediction bias by 42% compared to traditional methods.
[0133] Furthermore, large language models are used for generation and optimization.
[0134] Exemplarily, in the large language model generation and optimization step, a three-stage generation architecture can be used to realize professional report automatic production. In the first draft generation stage, the field fine-tuning model based on the Deepseek architecture (continuously pre-trained on a 50GB wireless communication corpus, extended by 2,143 professional terms) receives structured input data, including prediction indicators (with confidence intervals), 3GPP protocol references (accurate to chapter number) and task constraint conditions. The model controls the output stability through the temperature parameter (temperature = 0.3), and dynamically embeds enhanced prompt word templates (such as "follow the IEEE style, quote 3GPP TS38.214v16.4.0 power control requirements"). The technical verification stage implements three verification: the rule engine ensures that the numerical indicator deviation is ≤1% (such as RSRP range verification), the knowledge graph checks the protocol clauses 100% accurately, and the isolation forest model detects logical contradictions. The style optimization stage reconstructs the paragraph logic through the T5-3B model (transition word density ≥ 3 / paragraph), and uses the PPO reinforcement learning algorithm to dynamically adjust the technical detail density (engineer version technical parameter completeness > 90%, management version key conclusion prominence > 80%). The quality control system monitors the BERTScore (≥ 0.85), coherence, and other technical-language dual-dimensional indicators in real time, and automatically triggers a three-level repair process (regeneration / constraint simplification / manual review) for low-confidence content. The scheme realizes a generation speed of 45 tokens per second on an A100 GPU, with a first-round generation qualified rate of 92.7%, a technical term accuracy rate improved by 58% compared with the general model, and all outputs strictly complying with the latest 3GPP standards.
[0135] Further, the visualization and interactive output.
[0136] Exemplarily, in the visualization and interactive output step, multi-modal presentation of analysis results can be realized. The system receives three types of core input: time series indicators (5-minute granularity), geographic spatial data (WGS84 coordinate system base station location + meter-level coverage radius) and natural language analysis conclusions (Markdown format), while carrying metadata such as visualization type, interaction level and security level.
[0137] Finally, closed-loop feedback learning.
[0138] Exemplarily, in the closed-loop feedback learning step, by constructing a complete report generation system self-evolution mechanism, continuous performance improvement is achieved through multi-dimensional data collection, deep analysis and strategy optimization. The system receives three types of core input data: user behavior data (including editing track, reading heat map and 5-star rating), system performance indicators (processing delay and resource consumption of each module) and abnormal log records. In the feedback analysis layer, the Diff algorithm is used to accurately compare the differences between the manually modified content and the original report, the reading path model based on LSTM is used to identify the key attention areas (such as chapters with chart stay time longer than 30 seconds), and the sentiment analysis technology is used to process the text evaluation. The reinforcement learning optimization layer designs a multi-dimensional state space (including prompt word templates, generation parameters and quality scores), constructs a composite reward function (technical accuracy weight 40%, language fluency 30%, generation efficiency 20%, resource consumption -10%), uses PPO algorithm for weekly strategy update, and at the same time retains the historical optimal strategy as a rollback benchmark. The knowledge iteration layer realizes double-channel update: automatically monitors 3GPP standard updates and extracts new constraint clauses through the BERT-QA model, and at the same time, the excellent report fragment passed by manual review is vectorized and stored in the FAISS index library.
[0139] To implement the above data analysis report generation method, in one embodiment of the present disclosure, a data analysis report generation device is provided. Figure 5 The schematic architecture of the data analysis report generation device is schematically shown.
[0140] The data analysis report generation device 500 includes a text conversion module 501, a dependency graph construction module 502, a prompt word generation module 503, a prediction result generation module 504, and an analysis report generation module 505.
[0141] The text conversion module 501 is configured to obtain target multi-source data and convert the target multi-source data into a first natural language text. The dependency graph construction module 502 is configured to perform task division according to analysis report elements associated with a report template selected in advance, obtain a plurality of segment analysis tasks, and construct a task dependency graph for the plurality of segment analysis tasks. The prompt word generation module 503 is configured to determine a target standard protocol matching a second natural language text associated with each segment analysis task from a pre-constructed standard protocol knowledge graph according to the task dependency graph and the second natural language text, to generate dynamic prompt words. The second natural language text is a natural language text associated with each segment analysis task selected from the first natural language text. The prediction result generation module 504 is configured to input the target multi-source data into a pre-trained hybrid prediction model, to output a prediction result based on the hybrid prediction model. The analysis report generation module 505 is configured to generate a data analysis report based on the prediction result and the dynamic prompt words.
[0142] In an optional embodiment of the present disclosure, the device further comprises a data acquisition module, the data acquisition module is configured to acquire data from a plurality of network data source systems based on a standardized interface to obtain initial multi-source data; and perform data preprocessing on the initial multi-source data to obtain target multi-source data; wherein the data preprocessing on the initial multi-source data comprises at least one of the following:
[0143] In response to the initial multi-source data containing real-time data streams acquired from at least one network management system, performing anomaly data detection on the actual data streams according to a dynamic sliding window algorithm and / or an isolation forest algorithm to obtain first target data; wherein the size of the sliding window of the dynamic sliding window algorithm is determined based on the network type corresponding to the network management system;
[0144] In response to the initial multi-source data containing historical log texts of base stations acquired from a distributed file system, establishing a mapping relationship between each log code in the historical log texts and a network semantic event according to a conditional random field (CRF) model to obtain second target data;
[0145] Performing spatio-temporal dimension alignment on the first target data and / or the second target data and / or the operation and maintenance order text contained in the initial multi-source data to obtain aligned target multi-source data,
[0146] In an optional embodiment of the present disclosure, the data acquisition module is configured to encode the base station position using a GeoHash algorithm based on a spatial dimension, and determine the cell coverage boundary position using a Voronoi diagram; and perform alignment on the first target data and / or the second target data and / or the operation and maintenance order text of different sampling rates using a cubic spline interpolation based on a time dimension.
[0147] In an optional embodiment of the present disclosure, the device further comprises an updating module and a protocol determination module, the updating module is configured to update a standard protocol knowledge graph based on a query encoder using a BERT model in response to a standard protocol update to obtain an updated standard protocol knowledge graph; and the protocol determination module is configured to determine a target standard protocol matching the second natural language text associated with each segment analysis task based on a document encoder through the updated standard protocol knowledge graph.
[0148] In an optional embodiment of the present disclosure, the device further comprises an optimization module, the optimization module is configured to optimize the dynamic prompt word according to a T5-3B model to obtain an optimized dynamic prompt word.
[0149] In an optional embodiment of the present disclosure, the prediction result generation module 504 is configured to determine a prediction length; in response to the prediction length being a first prediction length, adopt a time fusion transformer (TFT) model as a hybrid prediction model to output a prediction result based on the TFT model; and in response to the prediction length being a second prediction length, adopt a Prophet and XGBoost combined model as a hybrid prediction model to output a prediction result based on the Prophet and XGBoost combined model.
[0150] In an optional embodiment of the present disclosure, the prediction result generation module 504 is configured to determine a risk object based on the hybrid prediction model, and generate early warning information carrying a time label based on the risk object.
[0151] The data analysis report generation apparatus 500 provided by the embodiments of the present disclosure can execute the technical solutions of the data analysis report generation method in any of the above embodiments, and the implementation principles and beneficial effects thereof are similar to those of the data analysis report generation method. For details, refer to the implementation principles and beneficial effects of the data analysis report generation method, which will not be described here.
[0152] In the example embodiments of the present disclosure, a computer readable storage medium having a program product stored thereon is also provided, which can implement the above-mentioned method of the present disclosure. In some possible implementation manners, each aspect of the present disclosure can also be implemented in the form of a program product, which includes program codes for causing a terminal device to execute the steps described in the above-mentioned “example method” section according to various example embodiments of the present disclosure when the program product is running on the terminal device.
[0153] The program product for implementing the above-mentioned method according to the embodiments of the present disclosure can adopt a portable compact disc read-only memory (CD-ROM) and include program codes, and can run on a terminal device such as a personal computer. However, the program product of the present disclosure is not limited to this, and in this document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in conjunction with an instruction execution system, apparatus or device.
[0154] The program product can employ any combination of one or more computer-readable media or storage media. The computer-readable media or storage media can be readable by a computer. For example, the computer-readable media or storage media can include a non-transitory computer-readable storage medium. Examples of non-transitory computer-readable storage media include, but are not limited to, a magnetic storage media, an optical storage media, a solid state storage media, a hard disk, a floppy disk, a RAM, a ROM, a flash memory, a USB drive, a memory card, a DVD, a Blu-Ray disk, a CD, a media cartridge, and any other medium that can be used to store desired program code in a non-transitory fashion.
[0155] The computer-readable signal media can include a computer-readable storage medium that is non-transitory or a computer-readable transmission medium that can be any medium that can be used to carry program code for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable media can include any medium that is capable of storing or encoding a sequence of instructions for execution by the computer and that causes the computer to perform any one of the methodologies of the present disclosure. The computer-readable media can include, but are not limited to, solid-state memories, optical media, magnetic media, memory cards, memory sticks, and any other media that can be used to carry or store desired program code in a non-transitory fashion.
[0156] Program code embodied on the computer-readable media can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0157] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider. The present disclosure can also be implemented as a computer-readable storage medium having stored thereon a computer program that can direct a computer to operate as described and illustrated.
[0158] In the exemplary embodiments of the present disclosure, an electronic device capable of implementing the above-described method is also provided.
[0159] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method or a program product. Therefore, various aspects of the present application can be embodied in the form of entirely hardware embodiments, entirely software embodiments (including firmware, microcode, etc.), or embodiments combining software and hardware aspects, which can be generally referred to as "circuitry", "module" or "system".
[0160] The electronic device 600 according to this embodiment of the present application will be described below with reference to Figure 6 Figure 6 The electronic device 600 is merely an example and should not impose any limitation on the functions and usage range of the embodiments of the present application.
[0161] As shown in Figure 6 The components of the electronic device 600 can include, but are not limited to, the at least one processing unit 610, the at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), and a display unit 640.
[0162] The storage unit stores program codes which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present application described in the "Exemplary Method" section of the present specification. For example, the processing unit 610 can perform the steps S201 to S205 as shown in Figure 2
[0163] The storage unit 620 can include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 6201 and / or a cache memory 6202, and can further include a read-only memory (ROM) 6203.
[0164] The storage unit 620 can further include a program / utility 6204 having a set of program modules 6205, including but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which can include implementation of a network environment, alone or in combination.
[0165] The bus 630 can represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit bus, or a local bus using any of a variety of bus architectures.
[0166] The electronic device 600 can also communicate with one or more external devices 1000 such as a keyboard or pointing device, a Bluetooth device, or a device for enabling
[0167] Those skilled in the art will readily understand that the example embodiments described herein can be implemented by software and / or by hardware components. As such, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or on a network, and includes a number of instructions for enabling a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to perform the methods according to the embodiments of the present disclosure.
[0168] In addition, the above-described diagrams are merely schematic illustrations of the processes included in the method according to the example embodiments of the present application, and are not intended to be limiting. It will be readily understood that the processes shown in the above-described diagrams do not indicate or limit the time sequence of the processes. In addition, it will also be readily understood that the processes can be executed synchronously or asynchronously, for example, in a plurality of modules.
[0169] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. Indeed, according to the embodiments of the present disclosure, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into a plurality of modules or units.
[0170] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the disclosure being indicated by the following claims.
[0171] It should be understood that the present disclosure is not limited to the precise structures herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the claims that follow.
Claims
1. A method of generating a data analysis report, the method comprising: The method comprises: obtaining target multi-source data and converting the target multi-source data into first natural language text; dividing tasks according to analysis report elements associated with a preselected report template to obtain a plurality of segment analysis tasks, and constructing a task dependency graph for the plurality of segment analysis tasks; determining, from a pre-constructed standard protocol knowledge graph, a target standard protocol matching the second natural language text associated with each of the segment analysis tasks according to the task dependency graph and the second natural language text to generate dynamic prompt words; wherein the second natural language text is a natural language text selected from the first natural language text and associated with each of the segment analysis tasks; inputting the target multi-source data into a pre-trained hybrid prediction model to output a prediction result based on the hybrid prediction model; generating a data analysis report based on the prediction result and the dynamic prompt words.
2. The method of claim 1, wherein, The method comprises: collecting data from a plurality of network data source systems based on a standardized interface to obtain initial multi-source data; performing data preprocessing on the initial multi-source data to obtain the target multi-source data; wherein the data preprocessing on the initial multi-source data comprises at least one of the following: in response to the initial multi-source data containing real-time data streams collected from at least one network management system, performing abnormal data detection on the real-time data streams according to a dynamic sliding window algorithm and / or an isolation forest algorithm to obtain first target data; wherein the size of the sliding window of the dynamic sliding window algorithm is determined based on the network type corresponding to the network management system; in response to the initial multi-source data containing historical log text of a base station obtained from a distributed file system, establishing a mapping relationship between each log code in the historical log text and a network semantic event according to a conditional random field (CRF) model to obtain second target data; aligning the first target data and / or the second target data and / or operation and maintenance order text contained in the initial multi-source data in time and space dimensions to obtain the target multi-source data after alignment.
3. The method of claim 2, wherein, The method further comprises: in response to a standard protocol update, updating the standard protocol knowledge graph based on a query encoder using a BERT model to obtain an updated standard protocol knowledge graph; determining, based on a document encoder, a target standard protocol matching the second natural language text associated with each of the segment analysis tasks through the updated standard protocol knowledge graph.
4. The method of claim 1, wherein, After generating the dynamic prompt words, the method comprises: optimizing the dynamic prompt words based on a T5-3B model to obtain optimized dynamic prompt words. 5. The method according to claim 1 or 4, characterized in that, 6. The method of claim 1, wherein, inputting the target multi-source data into a pre-trained hybrid prediction model to output a prediction result based on the hybrid prediction model, comprising: determining a prediction length; in response to the prediction length being a first prediction length, adopting a Temporal Fusion Transformer (TFT) model as the hybrid prediction model to output the prediction result based on the TFT model; in response to the prediction length being a second prediction length, adopting a Prophet and XGBoost combined model as the hybrid prediction model to output the prediction result based on the Prophet and XGBoost combined model.
7. The method of claim 1, wherein, outputting a prediction result based on the hybrid prediction model, comprising: determining a risk object based on the hybrid prediction model, and generating early warning information carrying a time label based on the risk object.
8. A data analysis report generation apparatus characterized by comprising: The device comprises: a text conversion module configured to obtain target multi-source data and convert the target multi-source data into first natural language text; a dependency graph construction module configured to perform task division based on analysis report elements associated with a pre-selected report template, to obtain a plurality of segment analysis tasks, and to construct a task dependency graph for the plurality of segment analysis tasks; a prompt word generation module configured to determine, based on the task dependency graph and second natural language text, a target standard protocol matching the second natural language text associated with each of the segment analysis tasks from a pre-constructed standard protocol knowledge graph, to generate dynamic prompt words; wherein the second natural language text is natural language text associated with each of the segment analysis tasks selected from the first natural language text; a prediction result generation module configured to input the target multi-source data into a pre-trained hybrid prediction model to output a prediction result based on the hybrid prediction model; an analysis report generation module configured to generate a data analysis report based on the prediction result and the dynamic prompt words.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the data analysis report generation method of any one of claims 1 to 7.
10. An electronic device, comprising: comprises: a processor; and a memory configured to store executable instructions of the processor; wherein the processor is configured to execute the data analysis report generation method of any one of claims 1 to 7 by executing the executable instructions.
Citation Information
Cited By
Method and device for processing echocardiogram
CN122049552A