An internet of things security intelligent report generation method and system based on search enhancement generation

By using a retrieval-enhanced approach to collect, process, and optimize multi-source data from smart elevators, intelligent reports are generated, solving the problems of low report generation efficiency and insufficient insight, and achieving efficient and professional report generation.

CN122047206BActive Publication Date: 2026-06-26XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD
Filing Date
2026-04-01
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies suffer from low report generation efficiency and insufficient data correlation analysis when generating abnormal operation reports for smart elevators, resulting in long report generation cycles and insufficient insight.

Method used

By using a retrieval-enhanced approach, the system receives report generation requests, generates data collection strategies and report structures, collects multi-source data, performs vectorization processing and analysis space construction, generates a data point distribution set, performs feature analysis region division, optimizes data collection strategies, and generates intelligent reports by combining large language models.

Benefits of technology

It improves the efficiency and accuracy of report generation, enables refined data screening and optimization, forms a logically coherent and comprehensive contextual information body, has in-depth content analysis capabilities, and enhances the professionalism of reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122047206B_ABST
    Figure CN122047206B_ABST
Patent Text Reader

Abstract

The application provides a kind of based on search enhancement generation's internet of things security intelligent report generation method and system, it is related to data processing technical field, the method includes: data point is mapped to corresponding feature analysis area, generates collection optimization guidance quantity, the collection optimization guidance quantity includes supplementary collection data, increases collection frequency, prolongs collection time window;According to collection optimization guidance quantity, data collection strategy is adjusted, and optimization query instruction is generated;Optimized multi-source data set is collected from each data source by executing optimization query instruction;Optimized multi-source data set is carried out information fusion and structured organization, and comprehensive context information body is generated;Comprehensive context information body is combined with report structure, and prompt instruction combination is generated;Prompt instruction combination is input into pre-training large language model processing, and internet of things security analysis report is generated.The application effectively improves the efficiency, accuracy and professionalism of report generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for generating IoT security smart reports based on retrieval enhancement. Background Technology

[0002] In the operation and management of smart elevators, generating operational anomaly reports to investigate potential safety hazards often relies on real-time elevator operating data (such as car speed, door opening / closing response time, traction machine current, and sudden fault codes). However, current report generation methods face several technical limitations. Existing automated generation methods often have shortcomings in processing real-time dynamic data. Elevator data needs continuous real-time updates, such as operating parameters fed back every second and multiple door control anomaly signals triggered within a short period. Efficient processing of this dynamic data places high demands on the technical solutions, and existing solutions are prone to inefficiency, potentially leading to relatively long report generation cycles. Furthermore, while existing automated generation methods can initially statistically analyze basic information such as the number of faults and fault types, they often lack depth in data correlation analysis. For example, it is not easy to correlate increased door opening / closing delay frequency with recent guide rail lubrication warning data, thus hindering the report's insight into the root causes of anomalies. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a method and system for generating IoT security smart reports based on retrieval enhancement, which effectively improves the efficiency, accuracy and professionalism of report generation.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] Firstly, a method for generating IoT security smart reports based on retrieval enhancement is provided, the method comprising:

[0006] Step 1: Receive report generation request, and generate data collection strategy and report structure according to the report type in the request;

[0007] Step 2: Generate a unified query command based on the data acquisition strategy and report structure, and collect device status, behavioral characteristics, and security standard information from the IoT device status database, historical time series database, and security knowledge base to generate a multi-source data set;

[0008] Step 3: Vectorize the multi-source dataset to generate a data point distribution set and construct an analysis space. In the analysis space, determine the core data distribution area based on the data point distribution density and establish a benchmark analysis unit. Based on the distribution characteristics of the data points within the benchmark analysis unit, perform spatial partitioning on the benchmark analysis unit to generate multiple feature analysis regions.

[0009] Step 4: Map the data points to the corresponding feature analysis areas to generate acquisition optimization guidance. The acquisition optimization guidance includes supplementing acquisition data, increasing acquisition frequency, and extending the acquisition time window. Adjust the data acquisition strategy according to the acquisition optimization guidance to generate optimization query instructions.

[0010] Step 5: Execute the optimization query command to collect the optimized multi-source data set from various data sources; perform information fusion and structured organization on the optimized multi-source data set to generate a comprehensive contextual information body;

[0011] Step 6: Combine the comprehensive contextual information with the report structure to generate a set of prompt instructions; input the set of prompt instructions into a pre-trained large language model for processing to generate an IoT security analysis report.

[0012] Secondly, an IoT security intelligent report generation system based on retrieval enhancement includes:

[0013] The receiving module is used to receive report generation requests and generate data collection strategies and report structures based on the report type in the request.

[0014] The data acquisition module is used to generate unified query instructions based on data acquisition strategies and report structures, and to collect device status, behavioral characteristics, and security standard information from IoT device status database, historical time series database and security knowledge base to generate multi-source data sets;

[0015] The processing module is used to vectorize the multi-source dataset, generate a data point distribution set and construct an analysis space. In the analysis space, the core data distribution area is determined based on the data point distribution density, and a benchmark analysis unit is established. Based on the distribution characteristics of the data points within the benchmark analysis unit, the benchmark analysis unit is spatially partitioned to generate multiple feature analysis regions.

[0016] The optimization module maps data points to corresponding feature analysis regions and generates acquisition optimization guidance quantities, which include supplementing acquisition data, increasing acquisition frequency, and extending the acquisition time window. Based on the acquisition optimization guidance quantities, the data acquisition strategy is adjusted to generate optimization query instructions.

[0017] The fusion module is used to execute optimized query commands, collect optimized multi-source data sets from various data sources, perform information fusion and structured organization on the optimized multi-source data sets, and generate a comprehensive contextual information body.

[0018] The output module combines the comprehensive contextual information with the report structure to generate a set of prompt instructions; the set of prompt instructions is then input into a pre-trained large language model for processing to generate an IoT security analysis report.

[0019] Thirdly, a computing device, comprising:

[0020] One or more processors;

[0021] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0022] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0023] The above-described solution of the present invention has at least the following beneficial effects:

[0024] By vectorizing multi-source data into an analyzable set of data point distributions and constructing an analysis space containing feature analysis regions, and further filtering key data and generating optimization guidance quantities through data point mapping, the system achieves refined filtering and optimization of multi-source data. This solves the problems of insufficient data correlation and the burying of key information in data processing, thus improving the effectiveness of data processing. The optimized multi-source data is then fused and structured to form a logically coherent comprehensive contextual information body, avoiding the data fragmentation and lack of correlation in traditional reports. By leveraging the reasoning capabilities of large models, the system enables in-depth content such as trend analysis and risk warning, solving the problems of monotonous content and insufficient insight in automated reports, and bringing the report quality to the level of professional analysis. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of a method for generating IoT security smart reports based on retrieval enhancement, provided by an embodiment of the present invention.

[0026] Figure 2 This is a schematic diagram of an IoT security intelligent report generation system based on retrieval enhancement provided by an embodiment of the present invention. Detailed Implementation

[0027] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0028] like Figure 1 As shown, an embodiment of the present invention proposes a method for generating IoT security smart reports based on retrieval enhancement, the method comprising the following steps:

[0029] Step 1: Receive report generation request, and generate data collection strategy and report structure according to the report type in the request;

[0030] Step 2: Generate a unified query command based on the data acquisition strategy and report structure, and collect device status, behavioral characteristics, and security standard information from the IoT device status database, historical time series database, and security knowledge base to generate a multi-source data set;

[0031] Step 3: Vectorize the multi-source dataset to generate a data point distribution set and construct an analysis space. In the analysis space, determine the core data distribution area based on the data point distribution density and establish a benchmark analysis unit. Based on the distribution characteristics of the data points within the benchmark analysis unit, perform spatial partitioning on the benchmark analysis unit to generate multiple feature analysis regions.

[0032] Step 4: Map the data points to the corresponding feature analysis areas to generate acquisition optimization guidance. The acquisition optimization guidance includes supplementing acquisition data, increasing acquisition frequency, and extending the acquisition time window. Adjust the data acquisition strategy according to the acquisition optimization guidance to generate optimization query instructions.

[0033] Step 5: Execute the optimization query command to collect the optimized multi-source data set from various data sources; perform information fusion and structured organization on the optimized multi-source data set to generate a comprehensive contextual information body;

[0034] Step 6: Combine the comprehensive contextual information with the report structure to generate a set of prompt instructions; input the set of prompt instructions into a pre-trained large language model for processing to generate an IoT security analysis report.

[0035] In this embodiment of the invention, multi-source data is transformed into an analyzable set of data point distributions through vectorization processing, and an analysis space containing feature analysis regions is constructed. Furthermore, key data is filtered and optimized guidance quantities are generated through data point mapping, achieving refined filtering and optimization of multi-source data. This solves the problems of insufficient data correlation and the burying of key information in data processing, improving the effectiveness of data processing. The optimized multi-source data is then fused and structured to form a logically coherent comprehensive contextual information body, avoiding the fragmented and unrelated defects of traditional reports. Leveraging the reasoning capabilities of large models, in-depth content such as trend analysis and risk warnings is achieved, solving the problems of monotonous content and insufficient insight in automated reports, and ensuring that the report quality reaches the level of professional analysis.

[0036] In a preferred embodiment of the present invention, step 1, receiving a report generation request and generating a data acquisition strategy and report structure according to the report type in the request, specifically includes: receiving a report generation request from management personnel or a monitoring system, the request containing explicit report type information, such as a daily operation report, an anomaly fault analysis report, or a monthly safety assessment report for a smart elevator; retrieving the corresponding basic data acquisition strategy framework from the built-in strategy template library according to different report types, the framework corresponding to the daily operation report will focus on the periodic collection of routine operating parameters of the equipment, while the framework corresponding to the anomaly fault analysis report will strengthen the real-time and correlational collection of fault-related parameters; and determining the report structure according to the report type. The specific chapter composition of the report varies. Routine operation reports typically include chapters on equipment operation overview, key parameter statistics, and no abnormalities. Abnormal fault analysis reports, on the other hand, include chapters on fault phenomenon description, related parameter analysis, possible cause investigation, and handling suggestions. Based on the basic framework, the data collection strategy is refined according to the specific requirements of the report type. This includes clarifying the data categories, data sources, collection time range, and collection frequency. For example, abnormal fault analysis reports need to set the collection of equipment status data for one hour before and after the fault occurred, with the collection frequency increased to once per second. The chapter order and sub-items in each chapter are also adjusted accordingly to ensure that the data collection strategy corresponds to the report structure.

[0037] This embodiment combines report types to generate appropriate data collection strategies and report structures, ensuring that the data collection direction accurately matches the report requirements and avoiding invalid data collection.

[0038] In a preferred embodiment of the present invention, step 2, generating a unified query instruction based on the data acquisition strategy and report structure, and collecting device status, behavioral characteristics, and security standard information from the IoT device status database, historical time-series database, and security knowledge base to generate a multi-source data set, may include:

[0039] Step 201: Based on the data acquisition strategy corresponding to the report type, parse and generate a data acquisition requirement specification; based on the data acquisition requirement specification and the content framework determined by the report structure, generate a unified query instruction. Specifically, this includes: decomposing the generated data acquisition strategy for a specific report type, extracting the data acquisition objectives, data type requirements, data accuracy standards, and data association conditions, and organizing these contents into a structured data acquisition requirement specification; for example, for the data acquisition strategy of a smart elevator abnormal fault analysis report, it is parsed that the car running speed (accuracy to 0.1m / s, reference normal range 0.5 to 2.5m / s) and door opening / closing response time (accuracy to 0.1s, reference normal threshold 0.8 to 1.5s) need to be collected. The data acquisition requirements include data such as traction machine current (accuracy to 0.1A, reference rated range 10 to 15A), and these data need to be associated with the time of occurrence of fault codes (e.g., E1 represents door control sensor failure, E2 represents traction machine overload), forming a data acquisition requirement specification. Referring to the content framework of the report structure, determine the corresponding items of the required data in the requirement specification for each chapter. For example, the fault phenomenon description chapter in the report structure needs real-time status data when the fault occurs, such as the car position and door control status when fault code E1 occurs. The associated parameter analysis chapter needs historical behavioral characteristic data for the same period, such as statistics on door opening and closing response times for the same period in the past 30 days. Based on this correspondence, the data acquisition requirement specification is converted into a unified query command that can be recognized by various databases.

[0040] Step 202: Execute a unified query command to collect device status and operating parameters from the IoT device status database to obtain device status data; based on the device status data, extract corresponding historical behavior characteristics and trend change information from the historical time series database to obtain historical behavior characteristic data; combine the device status data and historical behavior characteristic data to obtain matching safety standard provisions and handling experience documents from the safety knowledge base to obtain safety knowledge documents. Specifically, this includes: sending the generated unified query command to the interface of the IoT device status database. After receiving the command, the database retrieves and extracts the real-time status data of the device within the corresponding time period according to the device ID (EL-001) and time range in the command, including the device's current operating mode (normal operating mode), the operating parameters fed back by each sensor (e.g., the smart elevator car position is 15th floor, position accuracy ±0.05m; door control status is not closed, corresponding signal low level 0V; traction machine temperature is 68℃, normal range 40 to 260℃), and whether there is a fault code (E1 door control). Information such as sensor malfunctions is collected and organized into equipment status data. Based on the equipment ID (EL-001) and time range (09:00 to 10:00 daily) in the equipment status data, a query request is sent to the historical time-series database. This database locates the historical records of the equipment based on the equipment ID, extracts data such as car speed and door opening / closing response time for the same time period within the past 30 days, and calculates the average value of these historical data (car speed 1.8 m / s, door opening / closing response time 1...). The historical behavioral characteristics of the equipment were analyzed by taking the following data: maximum value (car speed 2.2 m / s, door opening / closing response time 1.7 s), minimum value (car speed 1.5 m / s, door opening / closing response time 1.2 s), and frequency of change (door opening / closing delay occurred 6 times in the past 7 days). For example, the frequency and trend of door opening / closing delay during the daily period from 09:00 to 10:00, and the traction machine current gradually increased from 12.1A to 14.4A, an increase of about 0.5A per week.

[0041] Key parameter values ​​(car speed 1.9 m / s, door opening / closing time 1.6 s, traction machine temperature 68 ℃) and fault codes (E1) are extracted from equipment status data. Abnormal trend characteristics (maximum fluctuation of traction machine current 2.3 A, door opening / closing time exceeding historical average by 0.2 s) are extracted from historical behavior characteristic data. This information is used as search keywords to perform matching queries in the safety knowledge base. By matching keywords, the safety standard provisions and handling experience documents most relevant to the current equipment status and historical behavior are found and compiled into safety knowledge documents.

[0042] Step 203 involves integrating and processing equipment status data, historical behavior characteristic data, and safety knowledge documents to generate a multi-source data set. Specifically, this includes: standardizing the formats of the equipment status data, historical behavior characteristic data, and safety knowledge documents; converting numerical parameters in the equipment status data, such as car speed 1.9 m / s and door opening / closing time 1.6 s, and statistical results in the historical behavior characteristic data, such as 6 instances of delay and current fluctuation 2.3 A, into standardized numerical formats retaining two decimal places; and converting the text content in the safety knowledge documents into structured text fragments, each fragment containing keywords (e.g., E1 fault, door control sensor), a content summary (e.g., fault code E1 corresponds to a door control sensor fault, requiring inspection of wiring and cleaning of the probe), and a source identifier; and using the equipment ID (EL-001) and timestamp (October 20, 2024, 09:00 to 10:00) as the core associations to integrate the three types of data. The process involves matching and associating data, such as matching the equipment status data of elevator EL-001 from 09:00 to 10:00 on October 20, 2024 with the historical behavioral characteristic data of the same elevator over the same period in the past 30 days. Then, it associates the safety knowledge documents related to E1 fault handling that both data points to with these data, ensuring that various types of data from the same equipment within the same time period form a corresponding relationship. The associated data is then deduplicated, removing duplicate records (such as duplicate E1 fault code descriptions) or content with the same description (such as the same door control sensor inspection steps in different documents). The remaining data is then organized according to a hierarchical structure of equipment basic information (ID: EL-001, model: OTIS-3000), real-time status data (car speed 1.9m / s, door not closed), historical behavioral data (6 delays in the past 7 days), and associated knowledge documents (E1 fault handling clauses), forming a multi-source data set containing multiple types of information and logically related.

[0043] In this embodiment, a unified query command is generated by parsing the requirements, and data is collected and integrated from multiple source databases to ensure that the acquired data is comprehensive and closely related.

[0044] In a preferred embodiment of the present invention, step 3 involves vectorizing the multi-source data set to generate a data point distribution set and constructing an analysis space. Within the analysis space, a core data distribution region is determined based on the data point distribution density, and a benchmark analysis unit is established. Based on the distribution characteristics of data points within the benchmark analysis unit, the benchmark analysis unit is spatially partitioned to generate multiple feature analysis regions, which may include:

[0045] Step 301 involves feature extraction processing of equipment status data, historical behavior feature data, and safety knowledge documents. This process converts the semantic content and numerical attributes of various data types into feature value sets. Specifically, this includes: for equipment status data, extracting numerical attribute features, such as car running speed 1.9, door opening / closing response time 1.6, and traction machine current 13.2, directly using these values ​​as basic features; simultaneously extracting status features, such as equipment being in operation (converted to 1), presence of E1 fault code (converted to 1), and door not closed (converted to 0), converting these states into feature values ​​of 0 or 1; for historical behavior feature data, extracting trend features, such as the number of door opening / closing delays in the past 7 days (6) and the maximum fluctuation value of traction machine current (2.3), using these statistical results as feature values; extracting comparative features, such as the difference between the current car speed (1.9) and the historical average (1.8) (difference 0.1), and the difference between the current door opening / closing time (1.6) and the historical average (1.4) (difference 0.2), using these calculation results as feature values; and for safety knowledge documents, extracting semantic content features, extracting keywords from safety standard clauses and handling experience in the documents. For example, if a gate control malfunction occurs 8 times, insufficient lubrication occurs 5 times, or an E1 fault occurs 10 times, the frequency is used as the basic feature value. Simultaneously, weights are assigned based on the importance of the keywords, with the overall weight set between 0.1 and 1.0. The specific value depends on the closeness of the keyword's association with the current equipment malfunction type; the higher the association, the closer the weight is to 1.0, and the lower the association, the closer the weight is to 0.1. The E1 fault directly matches the existing E1 fault code in the current equipment, has the highest association, and is weighted 1.0. Gate control malfunctions directly match the current equipment's door not being closed. The correlation between the keywords is relatively high, so the weight is set to 0.8; the direct correlation between insufficient lubrication and the current anomaly is relatively low, so the weight is set to 0.6; the result of multiplying the frequency of keyword occurrences by the corresponding weights (gate anomaly 8×0.8=6.4, insufficient lubrication 5×0.6=3.0, E1 fault 10×1.0=10.0) is used as the feature value; all the feature values ​​extracted from the three types of data (1.9, 1.6, 13.2, 1, 1, 0, 6, 2.3, 0.1, 0.2, 6.4, 3.0, 10.0) are arranged in a uniform order to form a feature value set.

[0046] Step 302: Based on the feature value set, each data entity is converted into a corresponding multi-dimensional numerical vector through vectorization processing to generate a data point distribution set. According to the dimensional characteristics of each numerical vector in the data point distribution set, a corresponding analysis space is constructed. Specifically, this includes: clarifying the composition of the feature value set; each data entity's feature value set contains feature items in a fixed order, such as car speed, door opening / closing time, fault code status, historical delay count, keyword weight value, etc., and each feature item corresponds to an independent dimension, with the number of feature items completely matching the number of dimensions. Taking a certain data entity as an example, its feature value set includes car speed 1.9, door opening / closing time 1.6, fault code status 1, historical delay count 6, and keyword weight value 10.0. The feature items are then sequentially... Each value is used as an element of a vector, generating a multidimensional numerical vector (1.9, 1.6, 1, 6, 10.0) corresponding to that data entity. This process is repeated for all data entities. For example, if another data entity has a feature set of car speed 1.8, door opening / closing time 1.5, fault code status 0, historical delay count 5, and keyword weight value 8.5, the corresponding vector generated is (1.8, 1.5, 0, 5, 8.5). Another example is a feature set of car speed 2.0, door opening / closing time 1.7, fault code status 1, historical delay count 7, and keyword weight value 11.2, which corresponds to a vector generated as (2.0, 1.7, 1, 7, 11.2). All generated multidimensional vectors are categorized and summarized according to device ID and timestamp to form a data point distribution set.

[0047] Analyzing the dimensional composition of all vectors in the data point distribution set, and comparing multiple vectors within the set, we observe the elements of the first dimension. We find that they all originate from the car speed item in the feature value set, and the values ​​all reflect the real-time speed parameters of the equipment. Therefore, we determine that the first dimension represents the real-time operating parameter of the equipment (car speed). Observing the elements of the second dimension, we find that they all correspond to the door opening / closing duration item in the feature value set, and the values ​​reflect the real-time response time of the equipment's door control. We determine that the second dimension represents the real-time status (door opening / closing duration). Continuing to observe the elements of the third dimension, all of them are 0 or 1, corresponding to the conversion results of fault code status in the feature value set, reflecting whether the equipment has a fault. We determine that the third dimension represents the fault status (fault code status). The elements of the fourth dimension all come from the statistical results of historical delays, reflecting the abnormal frequency of the equipment's historical behavior. We determine that the fourth dimension represents historical behavior statistics (historical delay counts). The elements of the fifth dimension are all the product of the frequency of keyword occurrences and their weights, reflecting the closeness of the correlation between the data and safety knowledge. We determine that the fifth dimension represents the knowledge relevance (keyword weight value). In this way, we clarify the characteristic meaning of each dimension one by one.

[0048] A multidimensional analysis space is constructed based on the meaning of the features of each dimension. A dedicated coordinate axis is assigned to each feature item corresponding to each dimension, with the axis name consistent with the dimension's feature meaning. For example, the first axis is named "Car Speed" (m / s), the second "Door Opening / Closing Duration" (s), the third "Fault Code Status," the fourth "Historical Delay Count," and the fifth "Keyword Weight Value" (unitless). The scale range of each coordinate axis is determined based on the numerical range of all vectors under that dimension, while reserving a 5% to 10% redundancy interval to prevent data overflow. For instance, the numerical range of all vectors under the "Car Speed" dimension is 1.5 to 2.2 m / s, and the scale range is set to 0 to 3 m / s; the numerical range of the "Door Opening / Closing Duration" dimension is... The time interval is 1.2 to 1.7 seconds, with a scale range of 0 to 3 seconds; the fault code status dimension is only 0 and 1, with a scale range of 0 to 2; the historical delay count value range is 5 to 7 times, with a scale range of 0 to 10 times; the keyword weight value range is 8.5 to 11.2, with a scale range of 0 to 15; each coordinate axis is divided into scales at fixed intervals, such as car speed at 0.1 m / s intervals, door opening and closing time at 0.2 s intervals, and historical delay counts at 1 time intervals, to ensure that the scales are clearly distinguishable; after completing the coordinate axis settings, all vectors in the data point distribution set are mapped to the corresponding coordinate axis positions according to their respective dimension values, verifying that all vectors can find their corresponding positions in space, and finally forming a complete analysis space.

[0049] Step 303: In the analysis space, determine the core data distribution area based on the data point distribution density and establish a benchmark analysis unit; based on the distribution characteristics of data points within the benchmark analysis unit, perform spatial partitioning on the benchmark analysis unit to generate multiple feature analysis regions. Specifically, this includes: calculating the distribution density of all data points in the analysis space, and counting the number of data points contained within a certain range around each data point (centered on the point, with each dimension extended by a fixed value, such as car speed extended by 0.2 m / s, door opening / closing time extended by 0.3 s, fault code status extended by 0, historical delay count extended by 1, and keyword weight value extended by 1.0, forming a cubic region). The larger the number of data points, the more features are included in the region. The higher the density, for example, if a data point (1.9, 1.6, 1, 6, 10.0) is surrounded by a cube containing 15 data points, the density is higher than other areas. The minimum enclosing circle algorithm is used to determine the core data distribution area. First, find the highest density data points (e.g., 10). Use these points as the initial point set. Calculate the minimum circle that can enclose all points in this set. The calculation process involves randomly selecting two points, such as (1.9, 1.6) and (2.0, 1.7), as the endpoints of the circle's diameter. Calculate the circle's center ((1.9 + 2.0) / 2, (1.6 + 1.7) / 2) = (1.95, 1.65), with a radius equal to half the distance between the two points. ,radius 0.14, check if other points in the initial point set (such as (1.8, 1.5)) are inside the circle, and calculate the distance from that point to the center of the circle.

[0050] ( If the radius is greater than 0.14 and the point is not inside the circle, then add that point as a diameter endpoint. Reselect (1.8, 1.5) and (2.0, 1.7) as diameter endpoints, and calculate the new center ((1.8 + 2.0) / 2, (1.5 + 1.7) / 2) = (1.9, 1.6). The new radius is half the distance between the two points (the distance between the two points is...). ,radius Then check if other points are inside the circle again, and repeat this process until all points are inside the circle. The resulting circle is the smallest circle that surrounds the initial point set. Then gradually expand the range of the point set (e.g., increase to 20 or 30 data points) and repeat the calculation of the smallest circle until the circle contains more than half of the data points in the analysis space (e.g., if there are 100 total data points, the circle contains 51 or more). The area covered by the circle at this time is the core data distribution area. Based on the core data distribution area, the boundary range of this area (i.e., the range determined by the center (e.g., (1.9, 1.6)) and radius (e.g., 0.28) of the smallest enclosing circle) is used as the benchmark to establish the benchmark analysis unit. This unit contains the densest distribution of data points in the analysis space.

[0051] Within the benchmark analysis unit, the magnitude of the vector corresponding to each data point is calculated using a vector magnitude calculation algorithm. The calculation process involves taking the value of each dimension of the vector, multiplying the value of each dimension by itself to obtain the square value of that dimension. For example, for the vector (1.9, 1.6, 1, 6, 10.0), the square values ​​of each dimension are 1.9 × 1.9 = 3.61, 1.6 × 1.6 = 2.56, 1 × 1 = 1, 6 × 6 = 36, and 10.0 × 10.0 = 100. The square values ​​of all dimensions are then added together to obtain a sum (3.61 + 2.56 + 1 + 36 + 100 = 143.17). Finally, the square root of this sum is calculated. The result obtained is the magnitude of the vector (approximately 11.97). Based on the magnitude, the data points within the benchmark analysis unit are divided into different groups, such as a group with a magnitude of 9.0 to 10.5 (corresponding to regions where the data features are relatively close to the core features), a group with a magnitude of 10.5 to 12.0 (corresponding to regions where the data features are slightly deviated from the core features), and a group with a magnitude of 12.0 to 13.5 (corresponding to regions where the data features deviate from the core features). Based on the distribution range of these groups, the benchmark analysis unit is spatially divided, and the range of each group forms an independent feature analysis region, ultimately generating multiple feature analysis regions.

[0052] This embodiment combines the minimum enclosing circle algorithm and the vector magnitude calculation algorithm to make feature extraction more logical, thereby accurately identifying the core data distribution and different feature regions, improving the accuracy of data classification, providing a clear basis for subsequent data collection and optimization, and enhancing the depth and reliability of data analysis.

[0053] In a preferred embodiment of the present invention, step 4 involves mapping data points to corresponding feature analysis regions to generate acquisition optimization guidance, wherein the acquisition optimization guidance includes supplementing acquisition data, increasing acquisition frequency, and extending the acquisition time window; adjusting the data acquisition strategy according to the acquisition optimization guidance to generate optimization query instructions, which may include:

[0054] Step 401: Based on the feature similarity between the feature vectors of each data point in the data point distribution set and the feature analysis region center point, establish the mapping relationship between the data points and the feature analysis region. Specifically, this includes: clarifying the feature vector composition of each data point in the data point distribution set. The feature vector dimension is consistent with the multidimensional analysis space dimension constructed in step 302. Taking the smart elevator scenario as an example, the feature vector includes five dimensions: car speed, door opening and closing time, fault code status, historical delay count, and keyword weight value. For example, the feature vector of a certain data point is 1.9 and 1. 6, 1, 6, 10.0; Determine the center point feature vector for each feature analysis region. The value of each dimension of the center point vector is the average value of the corresponding dimension values ​​of all data points in that region. For example, there are 3 data points in the slightly abnormal data region, and their car speed dimension values ​​are 1.8, 1.9, and 2.0 respectively. The center point value of this dimension is (1.8 + 1.9 + 2.0) / 3, which is 1.9. Calculate all dimension values ​​of the center point vector in this way to obtain center point vectors such as 1.9, 1.6, 1, 5.8, and 9.9.

[0055] The feature similarity between the feature vectors of data points and the feature vectors of the center points of each region is calculated. The calculation process involves first summing the products of the corresponding dimensions of the two vectors. Taking the data point vectors 1.9, 1.6, 1, 6, 10.0 and the center point vectors 1.9, 1.6, 1, 5.8, 9.9 as an example, the sum of the products is 1.9 × 1.9 + 1.6 × 1.6 + 1 × 1 + 6 × 5.8 + 10.0 × 9.9. Then, the square root of the sum of the squares of each dimension of the two vectors is calculated. The sum of the squares of the data point vectors is 1.9 × 1.9 + 1.6 × 1.6 + ... The result of 1×1 + 6×6 + 10.0×10.0 gives the sum of squares of the center point vectors as 1.9×1.9 + 1.6×1.6 + 1×1 + 5.8×5.8 + 9.9×9.9. The square roots of these two sums are then taken. The sum of the products obtained in the first step is divided by the product of the two square roots obtained in the second step to obtain the feature similarity. After calculating the similarity between each data point and the center points of all feature analysis regions, the region with the highest similarity is selected as the corresponding region for that data point. This establishes the mapping relationship between each data point and the feature analysis region.

[0056] Step 402: Based on the mapping relationship, statistically analyze the data point distribution characteristics of each feature analysis region to generate data acquisition optimization guidance. Identify feature analysis regions with insufficient data quality based on the data acquisition optimization guidance and adjust the data acquisition strategy. Specifically, this includes: based on the mapping relationship, statistically analyzing the data point distribution characteristics of each feature analysis region. These distribution characteristics include the total number of data points within the region, the numerical fluctuation range of each dimension of data, and the temporal distribution density of the data points. Taking a moderately abnormal data region as an example, the total number of data points within the region is 4, the numerical fluctuation range of the car speed dimension is 2.1 to 2.5 m / s, and the temporal distribution density is only 1 data point per hour. If the total number of data points... If the number of data points is less than 5, the guidance includes supplementing the collection of data from the corresponding equipment in the area for the past 24 hours; if the fluctuation range of a certain dimension exceeds 50% of the normal range of that dimension, the guidance includes increasing the collection frequency of the parameter of that dimension; if the time distribution density is less than 2 data points per hour, the guidance includes extending the data collection time window for that area. For example, if the total number of data points in the moderately abnormal data area is 4 (less than 5) and the time distribution density is 1 (less than 2) per hour, the generated collection optimization guidance is to supplement the collection of data from the corresponding equipment in the area, such as the EL-001 elevator, for the past 24 hours, and increase the collection frequency from once per second to twice per second.

[0057] Based on the data collection optimization guidance, regions with insufficient data quality were identified. The identification criteria were regions containing guidance content such as supplementary data collection, increased frequency, and extended window. The aforementioned moderately abnormal data area was the region with insufficient data quality. The data collection strategy was adjusted for this region. The adjustments included expanding the collection time range from 1 hour before and after the original fault to nearly 24 hours, adding traction machine temperature to the original car speed and door opening and closing time types, and increasing the collection frequency from once per second to twice per second, to ensure that the supplemented regional data could cover more time nodes and parameter dimensions.

[0058] Step 403: Based on the adjusted data acquisition strategy, reconstruct the query conditions and data source priority configuration to generate optimized query instructions. Specifically, this includes: reconstructing the query conditions according to the adjusted acquisition strategy, specifying the equipment identifier (e.g., elevator EL-001), the acquisition time range (October 21, 2024, 00:00:00 to October 21, 2024, 24:00:00), the acquisition parameter types (car speed, door opening / closing time, traction machine temperature), the acquisition frequency requirement (twice per second), and the data accuracy standards (car speed accuracy 0.1 m / s, door opening / closing time accuracy 0.1 s, traction machine temperature accuracy 0.5℃); and configuring the data source priority. Priorities are set based on data timeliness and relevance: the IoT device status database stores real-time data and should be prioritized to supplement the latest data, with a priority of 1; the historical time series database stores historical data for comparative analysis, with a priority of 2; the security knowledge base stores standard documents and should be matched with security thresholds for supplementary parameters, with a priority of 3; query conditions and data source priority configurations are integrated to generate optimized query instructions. The instructions clearly indicate the query order of each data source, parameter extraction requirements, and data return format, such as JSON format including device ID, timestamp, parameter name, and parameter value, to ensure that each data source can return the required data in the order and according to the instructions.

[0059] This embodiment establishes a data mapping relationship by accurately calculating feature similarity, generates optimized guidance quantities by combining distribution characteristics, and adjusts the acquisition strategy to ensure that data acquisition can specifically supplement data in areas with insufficient quality, avoiding data loss or redundancy.

[0060] In a preferred embodiment of the present invention, step 5, executing an optimization query instruction to collect an optimized multi-source data set from various data sources; performing information fusion and structured organization on the optimized multi-source data set to generate a comprehensive contextual information body, may include:

[0061] Step 501: Execute the optimization query command to obtain supplementary collected device status data from the IoT device status database, supplementary historical behavior feature data from the historical time series database, and supplementary safety standard documents from the safety knowledge base, generating an optimized multi-source data set. Specifically, the optimization query command is sent to each database interface according to the data source priority. First, it is sent to the IoT device status database. Based on the device ID and time range in the command, this database extracts the car speed, door opening and closing duration, and traction machine temperature data of elevator EL-001 twice per second from 00:00 to 24:00 on October 21, 2024. During the process, invalid records with duplicate timestamps or parameter values ​​exceeding reasonable ranges are removed, resulting in 86,400 valid records. Each record contains a timestamp accurate to milliseconds and specific parameter values. These data are sorted in chronological order and organized into supplementary device status data.

[0062] A command is sent to the historical time-series database. Based on the device ID and parameter type, this database extracts the car speed, door opening / closing time, and traction machine temperature data for elevator EL-001 from 00:00:00 to 24:00:00 daily for the past 7 days. Outliers, such as negative temperatures, are first removed. Then, the average and maximum values ​​of each parameter, as well as the number of fluctuations exceeding the average by ±10% per hour, are calculated to form supplementary historical behavioral characteristic data, such as the average traction machine temperature of 62 degrees Celsius, maximum fluctuation of 5 degrees Celsius, and an average of 3 fluctuations per day for the past 7 days. A command is also sent to the safety knowledge base. Based on the supplemented traction machine temperature parameter, the knowledge base retrieves the corresponding safety standard documents through keyword matching, such as the relevant content in industry safety standards stipulating that the operating temperature of the traction machine should not exceed 70 degrees Celsius, as supplementary safety standard documents. The supplementary device status data, historical behavioral characteristic data, and safety standard documents are merged with the original data. Duplicate records are deleted by comparing timestamps and parameter types, generating an optimized multi-source data set.

[0063] Step 502: Based on the optimized multi-source dataset, perform time alignment processing to obtain multi-source data with unified time dimension; based on the multi-source data with unified time dimension, perform data association according to device identifier to obtain associated data at the device dimension; based on the associated data at the device dimension, perform fusion processing to obtain feature integration results, specifically including: performing time alignment processing on the optimized multi-source dataset, using the timestamp of the IoT device status database data as the benchmark, accurate to the millisecond level, such as October 21, 2024, 10:00:00.001, adjusting the timestamp of the historical time series database data, and mapping the historical data to the same time point of the same day according to the date, such as adjusting the historical data timestamp of October 20, 2024, 10:00:00.001. For comparison with 10:00:00.001 on October 21, 2024, the safety standard document has no time attribute and is marked as applicable for all time periods to ensure that all data is consistent in the time dimension, resulting in multi-source data with consistent time dimension. Data is associated according to device identifier, with device IDEL-001 as the association identifier. All device status data, historical behavior feature data, and safety standard documents belonging to EL-001 in the multi-source data with consistent time dimension are classified into the same group. For example, the real-time traction machine temperature of 68℃, the historical temperature of 62℃ and the safety standard of 70℃ at 10:00:00.001 on October 21, 2024 are grouped together to ensure that various types of data of the same device form a corresponding relationship, resulting in associated data at the device dimension.

[0064] The associated data at the equipment level are fused and processed by calculating the difference between real-time parameters and historical parameters, such as real-time traction machine temperature 68℃ - historical average 62℃ = 6℃; comparing the difference between real-time parameters and safety standards, such as real-time temperature 68℃ - standard 70℃ = -2℃; and counting the number of times parameters exceed the historical fluctuation range within a certain time period, such as the traction machine temperature exceeding the historical fluctuation range twice between 10:00 and 11:00. These calculation results are then integrated with the original parameters and standard clauses to obtain a feature integration result that includes real-time values, historical comparison values, standard comparison values, and the number of anomalies.

[0065] Step 503: Based on the feature integration results, organize and process the information according to the report structure requirements to generate a comprehensive contextual information body that includes data support, feature analysis, and knowledge references. Specifically, this includes: clarifying the report structure requirements. Taking the intelligent elevator abnormal fault analysis report as an example, the structure includes four chapters: real-time equipment status, historical behavior comparison, safety standard matching, and abnormal feature summary. Each chapter has sub-items ordered according to the degree of impact of parameters on equipment safety. Core parameters such as traction machine temperature are listed first, and secondary parameters such as door opening and closing time are listed later.

[0066] Next, organize the feature integration results according to the chapter requirements. In the "Real-time Equipment Status" chapter, fill in real-time parameter data, such as the average car speed of elevator EL-001 from 10:00 to 24:00 on October 21, 2024 (1.9 m / s) and the average traction machine temperature (68 degrees Celsius), as data support. In the "Historical Behavior Comparison" chapter, fill in the difference between real-time and historical parameters, such as a traction machine temperature difference of 6 degrees Celsius, exceeding the fluctuation range of the same period in the past 7 days by 1.5 degrees Celsius, as feature analysis. In the "Safety Standard Matching" chapter, fill in the comparison between real-time parameters and standards, such as the traction machine temperature of 68 degrees Celsius being lower than... The standard of 70 degrees Celsius has not been exceeded, and the corresponding standard content, such as the requirements for traction machine temperature in the elevator operation safety code, is cited as knowledge. The abnormal characteristics summary section summarizes the number of abnormalities, such as door opening and closing delays exceeding the historical range 3 times and traction machine temperature exceeding the standard for a short time once. In the process of organization, it is ensured that the content of each chapter includes data to support specific parameter values, characteristic analysis and comparison results and abnormal judgments, knowledge reference standard content or handling experience, and the content of each chapter is arranged in chronological order to avoid logical confusion, and finally form a comprehensive contextual information body with a clear structure and complete information.

[0067] This embodiment uses time alignment, device association, and fusion processing to form an organic whole from multiple sources, and integrates contextual information to encompass data, analysis, and knowledge, providing a basis for report generation.

[0068] In a preferred embodiment of the present invention, step 6, combining the comprehensive contextual information body with the report structure to generate a set of prompt instructions; inputting the set of prompt instructions into a pre-trained large language model for processing to generate an IoT security analysis report, may include:

[0069] Step 601: Based on the comprehensive contextual information body and report structure, generate an initial prompt framework. Based on the initial prompt framework, fill in the equipment status information, behavioral feature analysis results, and safety knowledge documents in the comprehensive contextual information body to obtain complete prompt content. Specifically, this includes: generating an initial prompt framework according to the report structure and the content composition of the comprehensive contextual information body. The framework must include the title, subtitle, and content guidance of each chapter of the report. For example, one is the real-time status of the equipment, including car operating parameters (fill in the real-time car speed data and acquisition range) and traction machine operating status (fill in the real-time traction machine temperature data and acquisition range); two is historical behavior comparison, including temperature comparison analysis (fill in the difference between real-time and historical temperatures and fluctuation judgment) and door control parameter comparison (fill in the comparison results of the entry switch duration); three is safety standard matching, including traction machine temperature standards (fill in the corresponding standard content and real-time value comparison); and four is anomaly feature summary (fill in the number of abnormalities and trends of each parameter).

[0070] Step 602a: Based on the complete prompt content, optimize the prompt content by performing paragraph structure optimization; based on the optimized prompt content, process it to obtain a combination of prompt instructions that includes content generation guidelines and format constraints; input the combination of prompt instructions into a pre-trained large language model, and use the semantic understanding ability of the pre-trained large language model to generate a preliminary report text. Specifically, this includes: optimizing the paragraph structure of the complete prompt content. The optimization method is to make each first-level heading (e.g., 1.1) and the real-time status of the equipment a separate paragraph and bold it, and the second-level heading (e.g., 1.1 Car Operation Parameters) is followed by the specific content. The content should be coherent and avoid piling up short sentences. For example, "1.1 Car Operation Parameters: From 00:00 to 24:00 on October 21, 2024, the average speed of the EL-001 elevator car was 1.9 m / s, and the collection frequency was every..." The previous section, "Data collected twice per second, covering 24 hours without missing data," was optimized to: "I. Real-time Equipment Status 1.1 Car Operating Parameters: From 00:00 to 24:00 on October 21, 2024, the average car speed of elevator EL-001 was 1.9 m / s. The data collection frequency was set to twice per second, covering 24 hours of data without any missing data or abnormal interruptions." All chapter content was optimized in this way to obtain the optimized prompts. These prompts were then processed to generate a set of prompt instructions. The content generation guidelines must clearly define the report generation requirements, such as analyzing the causes of equipment anomalies and proposing preliminary handling suggestions based on the provided parameter data and standard clauses. Format constraints must clearly define the report's font (SimSun), font size (title 12pt, body text 12pt), line spacing (1.5x), and chapter numbering rules.

[0071] The construction and training process of the pre-trained large language model involves collecting data from the IoT security field, including 500,000 elevator equipment parameter data, 100,000 fault handling cases, and 5,000 industry safety standard documents. The data is cleaned, removing duplicates and erroneous data, such as records with parameter values ​​exceeding reasonable ranges. Text-based data, such as standard clauses, is converted into structured text. A basic large language model, such as GPT-3.5, is selected as the initial model. Fine-tuning is performed using the cleaned IoT security data, with a learning rate of 0.0001 and 10 training rounds. After each training round, 20% of the validation data is used to evaluate the accuracy of the model's generated reports, such as whether parameter references are correct and standard matching is accurate. Model parameters are adjusted based on the evaluation results until the model's report accuracy on the validation set reaches over 90%, completing model training and construction. The prompt instructions are then input into the trained pre-trained large language model. The model uses semantic understanding to recognize the chapter structure, data content, and generation requirements in the prompts, organizes the language according to format constraints, and generates a preliminary report text containing equipment status, comparative analysis, standard matching, anomaly summary, and handling suggestions.

[0072] Step 602b: Based on the preliminary report text, perform logical coherence verification to obtain a logically complete draft report. Based on the draft report, perform format standardization adjustments to obtain a structurally complete IoT security analysis report. Specifically, this includes: performing logical coherence verification on the preliminary report text, including checking for logical gaps between chapters (e.g., whether the traction machine temperature of 71℃ mentioned in the real-time status of the equipment has a corresponding standard comparison in the safety standard matching), whether data references are consistent (e.g., whether the car speed values ​​for the same time period are consistent), and whether the anomaly analysis and handling suggestions correspond (e.g., whether the analysis that the traction machine temperature exceeds the standard is due to poor heat dissipation, and whether the suggestions include cleaning the heat sink). If logical problems are found (e.g., historical comparison data of door opening and closing delays are not mentioned), the corresponding feature analysis results are supplemented. If data references are inconsistent (e.g., one place writes a car speed of 1.9m / s, and another place writes 1.8m / s), the data is corrected by checking the integrated context information. After completing the verification, a logically complete draft report is obtained.

[0073] The draft report underwent formatting adjustments, including standardizing the font to SimSun, setting the title font size to 12pt and bold, and the body text font size to 12pt; adjusting the line spacing to 1.5; and standardizing the chapter numbering to follow the hierarchical format of I, 1.1, 1.1.1 to avoid confusion; setting borders (1pt solid lines) for data tables (such as parameter comparison tables) and bolding the table headers; and deleting redundant blank lines or inconsistent indentation to ensure overall format consistency. After the adjustments were completed, the formatting details were checked again (such as whether the page numbers are consecutive and whether the titles are centered), resulting in a structurally complete IoT security analysis report.

[0074] This embodiment optimizes the prompt structure, clarifies instruction requirements, and combines a specially trained pre-trained large language model to ensure that the generated report is logically coherent and formatted correctly. At the same time, it further improves the report quality through verification and adjustment to meet the professional needs of IoT security analysis.

[0075] like Figure 2 As shown, embodiments of the present invention also provide an IoT security intelligent report generation system based on retrieval enhancement, comprising:

[0076] The receiving module is used to receive report generation requests and generate data collection strategies and report structures based on the report type in the request.

[0077] The data acquisition module is used to generate unified query instructions based on data acquisition strategies and report structures, and to collect device status, behavioral characteristics, and security standard information from IoT device status database, historical time series database and security knowledge base to generate multi-source data sets;

[0078] The processing module is used to vectorize the multi-source dataset, generate a data point distribution set and construct an analysis space. In the analysis space, the core data distribution area is determined based on the data point distribution density, and a benchmark analysis unit is established. Based on the distribution characteristics of the data points within the benchmark analysis unit, the benchmark analysis unit is spatially partitioned to generate multiple feature analysis regions.

[0079] The optimization module maps data points to corresponding feature analysis regions and generates acquisition optimization guidance quantities, which include supplementing acquisition data, increasing acquisition frequency, and extending the acquisition time window. Based on the acquisition optimization guidance quantities, the data acquisition strategy is adjusted to generate optimization query instructions.

[0080] The fusion module is used to execute optimized query commands, collect optimized multi-source data sets from various data sources, perform information fusion and structured organization on the optimized multi-source data sets, and generate a comprehensive contextual information body.

[0081] The output module combines the comprehensive contextual information with the report structure to generate a set of prompt instructions; the set of prompt instructions is then input into a pre-trained large language model for processing to generate an IoT security analysis report.

[0082] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0083] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0084] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0085] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for generating IoT security intelligent reports based on retrieval enhancement, characterized in that, The method includes: Step 1: Receive report generation request, and generate data collection strategy and report structure according to the report type in the request; Step 2: Generate a unified query command based on the data acquisition strategy and report structure, and collect device status, behavioral characteristics, and security standard information from the IoT device status database, historical time series database, and security knowledge base to generate a multi-source data set; Step 3: Vectorize the multi-source dataset to generate a data point distribution set and construct an analysis space. In the analysis space, determine the core data distribution area based on the data point distribution density and establish a benchmark analysis unit. Based on the distribution characteristics of the data points within the benchmark analysis unit, perform spatial partitioning on the benchmark analysis unit to generate multiple feature analysis regions. Step 4: Map the data points to the corresponding feature analysis regions to generate data acquisition optimization guidance. This guidance includes supplementing data acquisition, increasing acquisition frequency, and extending the acquisition time window. Adjust the data acquisition strategy based on the optimization guidance to generate optimized query instructions. Specifically, this includes: establishing a mapping relationship between data points and feature analysis regions based on the feature similarity between the feature vectors of each data point in the data point distribution set and the feature similarity of the center point of the feature analysis region; statistically analyzing the data point distribution characteristics of each feature analysis region based on the mapping relationship to generate optimization guidance; identifying feature analysis regions with insufficient data quality based on the optimization guidance and adjusting the data acquisition strategy; and reconstructing the query conditions and data source priority configuration based on the adjusted data acquisition strategy to generate optimized query instructions. Step 5: Execute the optimization query command to collect optimized multi-source data sets from various data sources; perform information fusion and structured organization on the optimized multi-source data sets to generate a comprehensive contextual information body. Specifically, this includes: executing the optimization query command to obtain supplementary collected device status data from the IoT device status database, supplementary historical behavioral feature data from the historical time series database, and supplementary security standard documents from the security knowledge base, generating an optimized multi-source data set; performing time alignment processing on the optimized multi-source data set to obtain multi-source data with unified time dimensions; performing data association according to device identifiers based on the time-unified multi-source data to obtain associated data at the device dimension; performing fusion processing on the associated data at the device dimension to obtain feature integration results; and performing information organization processing according to report structure requirements based on the feature integration results to generate a comprehensive contextual information body containing data support, feature analysis, and knowledge references. Step 6: Combine the comprehensive contextual information with the report structure to generate a set of prompt instructions; input the set of prompt instructions into a pre-trained large language model for processing to generate an IoT security analysis report.

2. The IoT security intelligent report generation method based on retrieval enhancement generation according to claim 1, characterized in that, Based on data acquisition strategies and report structures, a unified query command is generated. Device status, behavioral characteristics, and security standard information are collected from IoT device status databases, historical time-series databases, and security knowledge bases, generating a multi-source dataset, including: Based on the data collection strategy corresponding to the report type, a data collection requirement specification is generated; based on the data collection requirement specification and the content framework determined by the report structure, a unified query instruction is generated. Execute a unified query command to collect device status and operating parameters from the IoT device status database to obtain device status data; based on the device status data, extract the corresponding historical behavior characteristics and trend change information of the device from the historical time series database to obtain historical behavior characteristic data; combine the device status data and historical behavior characteristic data to obtain matching security standard provisions and handling experience documents from the security knowledge base to obtain security knowledge documents; The system integrates and processes device status data, historical behavior data, and safety knowledge documents to generate a multi-source data set.

3. The IoT security intelligent report generation method based on retrieval enhancement as described in claim 2, characterized in that, The multi-source dataset is vectorized to generate a data point distribution set and construct an analysis space. In the analysis space, the core data distribution area is determined based on the data point distribution density, and a benchmark analysis unit is established. Based on the distribution characteristics of data points within the benchmark analysis unit, the benchmark analysis unit is spatially partitioned to generate multiple feature analysis regions, including: Feature extraction processing is performed on equipment status data, historical behavior feature data, and safety knowledge documents to convert the semantic content and numerical attributes of various types of data into feature value sets; Based on the feature numerical set, each data entity is converted into a corresponding multidimensional numerical vector through vectorization processing to generate a data point distribution set; according to the dimensional characteristics of each numerical vector in the data point distribution set, the corresponding analysis space is constructed. In the analysis space, the core data distribution area is determined based on the data point distribution density, and a benchmark analysis unit is established. Based on the distribution characteristics of data points within the benchmark analysis unit, the benchmark analysis unit is spatially divided to generate multiple feature analysis regions.

4. The IoT security intelligent report generation method based on retrieval enhancement generation according to claim 3, characterized in that, By combining the comprehensive contextual information with the report structure, a combination of prompts and instructions can be generated; The combined prompts are input into a pre-trained large language model for processing, generating an IoT security analysis report, including: Based on the integrated context information body and report structure, an initial prompt framework is generated; based on the initial prompt framework, the device status information, behavioral feature analysis results and security knowledge documents in the integrated context information body are populated to obtain the complete prompt content; Based on the complete prompt content, structural optimization is performed to obtain a combination of prompt instructions; this combination of prompt instructions is then input into a pre-trained large language model, and the text generation capability of the pre-trained large language model is used to obtain a structurally complete IoT security analysis report.

5. The IoT security intelligent report generation method based on retrieval enhancement generation according to claim 4, characterized in that, Based on the complete prompt content, structural optimization is performed to obtain a combination of prompt instructions. This combination of prompt instructions is then input into a pre-trained large language model. The pre-trained large language model's text generation capabilities are used to generate a structurally complete IoT security analysis report, including: Based on the complete prompt content, paragraph structure optimization is used to obtain optimized prompt content; based on the optimized prompt content, further processing is performed to obtain a combination of prompt instructions that includes content generation guidance and format constraints; the combination of prompt instructions is input into a pre-trained large language model, and the semantic understanding ability of the pre-trained large language model is used to generate a preliminary report text; Based on the initial report text, logical coherence verification is performed to obtain a draft report with complete content logic; based on the draft report, formatting adjustments are made to obtain a structurally complete IoT security analysis report.

6. A search-enhanced IoT security intelligent report generation system, wherein the system implements the method as described in any one of claims 1 to 5, characterized in that, include: The receiving module is used to receive report generation requests and generate data collection strategies and report structures based on the report type in the request. The data acquisition module is used to generate unified query instructions based on data acquisition strategies and report structures, and to collect device status, behavioral characteristics, and security standard information from IoT device status database, historical time series database and security knowledge base to generate multi-source data sets; The processing module is used to vectorize the multi-source dataset, generate a set of data point distributions and construct an analysis space. In the analysis space, the core data distribution area is determined based on the data point distribution density, and a benchmark analysis unit is established. Based on the distribution characteristics of data points within the benchmark analysis unit, the benchmark analysis unit is spatially divided to generate multiple feature analysis regions. The optimization module is used to map data points to corresponding feature analysis areas and generate acquisition optimization guidance quantities, which include supplementing acquisition data, increasing acquisition frequency, and extending acquisition time window. The data acquisition strategy is adjusted based on the acquisition optimization guidance, and optimized query instructions are generated. Specifically, this includes: establishing a mapping relationship between data points and feature analysis regions based on the feature similarity between the feature vectors of each data point in the data point distribution set and the feature of the center point of the feature analysis region; statistically analyzing the data point distribution characteristics of each feature analysis region based on the mapping relationship to generate acquisition optimization guidance; identifying feature analysis regions with insufficient data quality based on the acquisition optimization guidance and adjusting the data acquisition strategy; and reconstructing the query conditions and data source priority configuration based on the adjusted data acquisition strategy to generate optimized query instructions. The fusion module executes optimized query commands, collecting optimized multi-source data sets from various data sources. It then performs information fusion and structured organization on the optimized multi-source data sets to generate a comprehensive contextual information body. Specifically, this includes: executing optimized query commands to obtain supplementary device status data from the IoT device status database, supplementary historical behavioral feature data from the historical time-series database, and supplementary security standard documents from the security knowledge base, generating an optimized multi-source data set; performing time alignment processing on the optimized multi-source data set to obtain time-uniformed multi-source data; associating the time-uniformed multi-source data according to device identifiers to obtain device-dimensional associated data; performing fusion processing on the device-dimensional associated data to obtain feature integration results; and organizing information according to report structure requirements based on the feature integration results to generate a comprehensive contextual information body containing data support, feature analysis, and knowledge references. The output module combines the comprehensive contextual information with the report structure to generate a set of prompt instructions; the set of prompt instructions is then input into a pre-trained large language model for processing to generate an IoT security analysis report.

7. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Education data report content interaction method and system based on retrieval enhancement generation

    CN121168472A

  • Document interpretation and report generation method and device, equipment and medium

    CN121354154A