An industry risk knowledge graph construction method and system based on a large language model
By performing batch denoising and compression on video footage, dynamically adjusting the time window and verification frequency, and constructing a high-precision industry risk knowledge graph using a large language model, the problems of noise and modal alignment in video footage were solved, and the accurate and efficient extraction of risk information was achieved.
Patent Information
- Application Number
- CN202511316275.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-09-16
AI Technical Summary
In existing technologies, image noise, compression artifacts, and audio interference in video footage lead to misidentification and missed detection of entities and events. Multi-source modal semantic fragmentation and temporal asynchrony result in chaotic relationship extraction, making it difficult to achieve high-precision knowledge graph construction.
By performing batch denoising and compression on the original video footage, analyzing and processing error factors, optimizing the automated process, dynamically adjusting the time window and verification frequency, and using a large language model to construct structured risk information for multimodal data.
It improves the accuracy, completeness, and timeliness of risk entities and relationships in knowledge graphs, ensures video clarity and alignment accuracy, and enhances the accuracy, completeness, and efficiency of information extraction and knowledge graph construction.
Smart Images

Figure CN120822594B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of data processing, and particularly relates to an industry risk knowledge graph construction method and system based on a large language model. BACKGROUND
[0002] The knowledge graph construction technology automatically or semi-automatically extracts entities, attributes and semantic relationships therebetween from multi-source heterogeneous data, generates a structured knowledge representation taking a triple of "entity-relation-entity" as a basic unit, and gathers these triples into a graph database, so as to finally form a semantic network that can be queried, reasoned and dynamically updated. The industry risk knowledge graph construction system has achieved preliminary application in the fields of financial risk control, industrial safety and medical compliance, and meanwhile, the system also faces challenges such as high cost of large-scale multi-modal data preprocessing, insufficient cross-modal alignment accuracy and poor explainability of LLM reasoning. The industry is continuously improving the efficiency and accuracy of graph construction and accelerating the pace from pilot landing to large-scale commercialization through means such as dynamic parameter adjustment closed loop, knowledge enhancement prompt and fine-grained model fine-tuning.
[0003] For example, the patent with the publication number CN120104808A discloses a knowledge graph construction and application method based on the insurance industry, and belongs to the technical field of industry knowledge graph. The method comprises the following steps: S1, data collection and arrangement; collecting multi-source data of the insurance industry and arranging the data; S2, knowledge modeling: defining entities, determining relationships and setting attributes; S3, constructing a graph: entity recognition and disambiguation, relationship extraction and graph database storage; S4, graph optimization: adding semantic information, optimizing relationships and creating indexes; and S5, graph application.
[0004] For example, the patent with the publication number CN118069861A discloses a power grid risk knowledge graph construction method and system. The method comprises the following steps: obtaining power grid structured, semi-structured and unstructured data files; constructing a knowledge graph based on power safety event accidents according to the data files; formulating system boundary division and function module selection rules according to the power safety event accident knowledge graph, and constructing a FRAM model of power grid refined business processes; and fusing the FRAM model and the constructed power safety event accident knowledge graph to obtain a power grid risk knowledge graph based on the fused FRAM model.
[0005] The above-mentioned technology at least has the following technical problems:
[0006] Image noise, compression artifacts and audio interference commonly accompanied in the original video can cause a large number of misrecognition and missed detection in subsequent entity and event extraction, and the semantic discontinuity of multi-source modalities on different channels makes it difficult to align in the relationship extraction link. Secondly, the time asynchronization between multi-modal data can cause mismatch of voice, subtitles and picture information, leading to semantic confusion in entity and event extraction. SUMMARY
[0007] In view of the deficiencies of the prior art, the present application provides an industry risk knowledge graph construction method and system based on a large language model, which solves the problems involved in the above background technology.
[0008] To achieve the above object, the present application is implemented by the following technical scheme: an industry risk knowledge graph construction method based on a large language model, comprising batch-wise denoising and compression processing of a target industry original video material set and monitoring the processing process, analyzing each batch processing error factor, and optimizing the automation process of the original video material.
[0009] The automation process adjustment effect of the original video material is analyzed, thereby determining the secondary automation process adjustment execution requirement, and obtaining the multi-modal data of each batch of video expected material.
[0010] The multi-modal data of each batch of video expected material is input into a multi-modal analysis pipeline, the cross-modal matching factor of each batch is analyzed, and the multi-modal data time dislocation of each batch of video expected material is determined.
[0011] The time window of the multi-modal data of each batch of video expected material is dynamically updated, the multi-modal data of each batch of video expected material is audited, the verification trigger frequency is adjusted, and the audited multi-modal data of each batch of video expected material is input into a large language model to obtain the structured risk information of the graph construction of each batch of video expected material.
[0012] Further, the average processing time of each batch of video original video material, the average key frame content retention rate, the average video signal-to-noise ratio, and the average frame loss rate of the processing batch are obtained.
[0013] The average processing time of each batch of video original video material, the average key frame content retention rate, the average video signal-to-noise ratio, and the average frame loss rate of the processing batch are obtained.
[0014] Further, the automation process of the original video material is optimized, and the specific process is as follows:
[0015] The each batch processing error factor is extracted, the each batch processing error factor is difference processed with the average value of the each batch processing error factor, and the each batch processing average error is obtained.
[0016] If the average error of a batch is positive, the batch error factor is subtracted from the average of the batch error factor to obtain a first difference value factor of the batch error, the intensity and compression ratio of the noise filter corresponding to each interval of the first difference value factor of the batch error stored in the database are extracted, and the intensity and compression ratio of the noise filter corresponding to each interval of the first difference value factor of the batch error are mapped, denoted as the first adjustment value of the intensity and the first adjustment value of the compression ratio of the noise filter, and the automatic process of the original video material is optimized according to the first adjustment value of the intensity and the first adjustment value of the compression ratio of the noise filter.
[0017] If the average error of a batch is negative, the average of the batch error factor is subtracted from the batch error factor to obtain a second difference value factor of the batch error, the intensity and compression ratio of the noise filter corresponding to each interval of the second difference value factor of the batch error stored in the database are extracted, and the intensity and compression ratio of the noise filter corresponding to each interval of the second difference value factor of the batch error are mapped, denoted as the second adjustment value of the intensity and the second adjustment value of the compression ratio of the noise filter, and the automatic process of the original video material is optimized according to the second adjustment value of the intensity and the second adjustment value of the compression ratio of the noise filter.
[0018] Further, the adjustment effect of the automatic process of the original video material is analyzed, and thus the execution demand of the secondary automatic process adjustment is determined, and the specific process is as follows: a preset observation period is set, and the average error of each batch after optimization is obtained in the observation period, if the absolute value of the average error of each batch is higher than or equal to the preset error threshold, the automatic process of the original video material is re-optimized, if the absolute value of the average error of each batch is lower than the preset error threshold, the automatic process of the original video material does not need to be re-optimized, and thus the multi-modal data of each batch of video expected material is obtained.
[0019] Further, the cross-modal matching factor of each batch is analyzed, and the specific process is as follows: the average time delay deviation, time delay distribution variance and maximum deviation of each batch of video expected material are obtained, the average time delay deviation is compared with the defined time delay deviation, the time delay distribution variance is compared with the defined time delay distribution variance, and the maximum deviation is compared with the defined maximum deviation, and the proportion analysis and weight coefficient are introduced, and the cross-modal matching factor of each batch is obtained by comparing the batch error factor, which is used to evaluate and quantify the consistency and matching quality of different modal data in the time synchronization and fusion process.
[0020] Further, the multi-modal data time dislocation of each batch of video expected materials is determined, and the specific process is as follows: the cross-modal matching factor of each batch is extracted and compared with the set cross-modal matching factor threshold value, if the cross-modal matching factor of a batch is higher than or equal to the cross-modal matching factor threshold value, the time window of the multi-modal data of the video expected material of the batch is updated, if the cross-modal matching factor of a batch is lower than the cross-modal matching factor threshold value, the time window of the multi-modal data of the video expected material of the batch does not need to be updated.
[0021] Further, the time window of the multi-modal data of each batch of video expected materials is dynamically updated, and the specific process is as follows: a second observation period is preset after the time window is updated, the cross-modal matching factor of each batch after optimization is obtained in the second observation period, the number of batches whose cross-modal matching factor is higher than or equal to the preset cross-modal matching factor threshold value is counted, and is recorded as the time dislocation number of each batch, if the time dislocation number of each batch is higher than or equal to the preset time dislocation number threshold value, the time window of the multi-modal data of each batch of video expected materials is updated again, if the time dislocation number of each batch is lower than the preset time dislocation number threshold value, the time window of the multi-modal data of each batch of video expected materials does not need to be updated again.
[0022] Further, the multi-modal data of each batch of video expected materials is audited, and the specific process is as follows: in the preset time window, the pass rate of each batch of audit results of each batch of video expected materials is collected, and the average audit coherence factor is analyzed; the audit coherence factor threshold value stored in the database is extracted; if the average audit coherence factor is higher than or equal to the audit coherence factor threshold value, the multi-modal data of each batch of video expected materials after audit is imported into the large language model; if the average audit coherence factor is lower than the audit coherence factor threshold value, the verification trigger frequency is adjusted.
[0023] Further, the verification trigger frequency is adjusted, and the multi-modal data of each batch of video expected materials after audit is imported into the large language model to obtain the structured risk information of the graph construction of each batch of video expected materials, and the specific process is as follows: the average audit coherence factor and the audit coherence factor threshold value are extracted, the audit coherence factor threshold value is subtracted from the average audit coherence factor to obtain the verification trigger deviation value.
[0024] The verification trigger frequency corresponding to each interval of the verification trigger deviation value stored in the database is extracted, and the verification trigger frequency corresponding to each interval of the verification trigger deviation value is mapped and extracted, which is recorded as the verification trigger frequency adjustment value, and the verification trigger frequency is adjusted according to the verification trigger frequency adjustment value.
[0025] The prompt template stored in the database is extracted, the multi-modal data of the audited batches of video expected materials and the prompt template are imported into a large language model, and structured risk information of graph construction of the batches of video expected materials is output.
[0026] The second aspect of the application also provides an industry risk knowledge graph construction system based on a large language model, comprising: a preprocessing optimization module for batch-wise denoising and compression processing of a target industry original video material set and monitoring the processing process, analyzing each batch processing error factor, and optimizing the automatic process of the original video material.
[0027] A process adjustment evaluation module is used to analyze the automatic process adjustment effect of the original video material, thereby determining the secondary automatic process adjustment execution requirement, and obtaining multi-modal data of each batch of video expected materials.
[0028] A multi-modal analysis module is used to input the multi-modal data of each batch of video expected materials into a multi-modal analysis pipeline, analyze the cross-modal matching factor of each batch, and determine the multi-modal data time dislocation of each batch of video expected materials.
[0029] A time alignment and compliance audit module is used to dynamically update the time window of the multi-modal data of each batch of video expected materials, and to conduct content audit on the multi-modal data of each batch of video expected materials, and to adjust the verification trigger frequency, and to import the audited multi-modal data of each batch of video expected materials into a large language model to obtain structured risk information of graph construction of each batch of video expected materials.
[0030] The application has the following advantages:
[0031] (1) The industry risk knowledge graph construction method based on a large language model first performs batch-wise denoising and compression processing on the video material set, and determines whether secondary optimization is needed by evaluating the adjustment effect of the automatic process, and finally forms quality-controllable multi-modal data of each batch of video expected materials; then the multi-modal data is input into an analysis pipeline, the cross-modal matching factor is calculated and the time dislocation is determined, the time window is dynamically adjusted to realize millisecond-level alignment, the aligned data is audited for content and the verification frequency is intelligently adjusted to exclude non-compliant risks; finally, the audited multi-modal materials are sent to a large language model for named entity recognition, relationship and event extraction, structured risk information is generated and written into a graph, ensuring video clarity, alignment accuracy and compliance quality, effectively improving the accuracy, completeness and timeliness of risk entities and relationships in the knowledge graph.
[0032] (2) The application helps to dynamically adjust the filter strength and compression ratio by obtaining the processing error factor, avoids excessive information loss, and finally re-optimizes the automated preprocessing process using these fine-tuning parameters and outputs new batch expected materials, which not only balances the playback quality and processing speed between batches, but also provides consistent and high-fidelity input for subsequent multi-modal data analysis, greatly improving the accuracy, completeness and efficiency of information extraction and knowledge graph construction.
[0033] (3) The application realizes closed-loop control from video cleaning, dynamic adaptive optimization to multi-modal time alignment in the early stage by obtaining the cross-modal matching factor of each batch, not only ensuring the high fidelity and high availability of input data, but also providing accurate and stable cross-modal semantic support for subsequent entity and relationship extraction driven by large language models, greatly improving the accuracy and completeness of risk information extraction in the industry risk knowledge graph.
[0034] (4) The cross-modal matching factor of each batch helps to further tighten or relax the matching range by adjusting the time window parameter of all batch video expected materials, and when the number of time misalignment is below the threshold, it not only dynamically optimizes the multi-modal synchronization accuracy for different video content and recognition conditions, but also provides highly consistent timing context for subsequent large language models in entity recognition and relationship extraction, thereby significantly improving the accuracy and completeness of cross-modal risk information in the industry risk knowledge graph.
[0035] Of course, implementing any product of the present application does not necessarily require all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 is a schematic diagram of the method of the present application;
[0037] Figure 2 is a schematic diagram of the logic flow of an industry risk knowledge graph construction method based on a large language model of the present application;
[0038] Figure 3 is a schematic diagram of the system modules of the present application. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0040] In the description of the present application, it should be understood that the terms "opening", "upper", "lower", "thickness", "top", "middle", "length", "inner", "periphery" and the like indicate the orientation or positional relationship, only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the components or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore cannot be understood as a limitation on the present application.
[0041] Please refer to Figure 1 As shown in the drawings, the embodiment of the present application provides a method for constructing an industry risk knowledge graph based on a large language model, which comprises batch-wise denoising and compression processing of a target industry original video material set and monitoring the processing process, analyzing the error factors of each batch processing, and optimizing the automation process of the original video material.
[0042] The automation process adjustment effect of the original video material is analyzed, thereby determining the secondary automation process adjustment execution requirement, and obtaining the multi-modal data of the expected video material of each batch.
[0043] The multi-modal data of the expected video material of each batch is input into a multi-modal analysis pipeline, the cross-modal matching factor of each batch is analyzed, and the multi-modal data time dislocation of the expected video material of each batch is determined.
[0044] The time window of the multi-modal data of the expected video material of each batch is dynamically updated, the multi-modal data of the expected video material of each batch is audited, the verification trigger frequency is adjusted, the multi-modal data of the expected video material of each batch after the audit is imported into a large language model, and the structured risk information of the graph construction of the expected video material of each batch is obtained.
[0045] It should be noted that the target industry original video material set is preprocessed in batches, the material is divided into several controllable processing batches, so as to dynamically optimize in a large-scale scene. For each batch, the denoising and compression process is first executed to improve the overall picture quality and compressibility of the video, and algorithm-level correction is performed on the distortion area, excessive blur or color deviation that may be contained in the video. These denoising and compression parameters are adjusted by the multi-modal integrity indicators dynamically collected by the real-time quality monitoring system, for example, the key frame content retention rate, the recognition degree of OCR text, the target detection accuracy, etc. are taken as the core, the error with the expected value is calculated, if the error exceeds the preset threshold in this batch, the denoising filter parameter or compression ratio of the next batch is immediately adjusted slightly, so as to avoid quality jitter or continuous frame loss problem. This fine-tuning adopts a moving average strategy of a sliding window, so that the system keeps stable error sensitivity under multiple consecutive batches, preventing algorithm shock caused by extreme data fluctuations.
[0046] Specifically, the processing error factor of each batch is analyzed, and the specific process is as follows: the average processing time of the processing batch of the original video material of each batch of video, the average key frame content retention rate, the average video signal-to-noise ratio and the average frame loss rate are obtained.
[0047] It should be noted that after the denoising and compression processing of each batch is completed, the system automatically records the start and end time from video input to output, and the average processing time is obtained by dividing the total processing time of all video segments in the batch by the number of videos or the total number of frames; at the same time, in the key frame extraction link, the system compares the number of frames identified as key frames in each video before and after denoising with the total number of original key frames to calculate the key frame content retention rate; for the average video signal-to-noise ratio, the average value of all frames in the batch is obtained by taking the peak signal-to-noise ratio of each denoised image and the corresponding original image; and the average frame loss rate is obtained by dividing the difference between the actual output frame number and the original frame number before and after processing by the original frame number. All these indicators are collected synchronously after each batch processing is completed, which are the basic data for evaluating and optimizing the preprocessing process.
[0048] The average processing time of the original video material of each batch of video, the average key frame content retention rate, the average video signal-to-noise ratio and the average frame loss rate are compared with the defined processing time, the defined key frame content retention rate, the defined video signal-to-noise ratio and the defined frame loss rate, respectively, and the weight coefficient is introduced to obtain the processing error factor of each batch, which is used to evaluate the overall deviation between the integrity of the information and the processing efficiency.
[0049] It should be noted that the processing error factor of each batch is analyzed under the following conditions:
[0050] ;
[0051] In the formula, E j represents the jth batch processing error factor, t1 j represents the average processing time of the processing batch of the original video material of the jth batch of video, t1 represents the defined processing time set in the database, B1 j represents the average key frame content retention rate of the processing batch of the original video material of the jth batch of video, B1 represents the defined key frame content retention rate set in the database, X1 j represents the average video signal-to-noise ratio of the processing batch of the original video material of the jth batch of video, X1 represents the defined video signal-to-noise ratio set in the database, Z1 jZ1a+b+c+d, Z1 represents the threshold frame loss rate set in the database, a represents the weight coefficient corresponding to the average processing time set in the database, b represents the weight coefficient corresponding to the average key frame content retention rate set in the database, c represents the weight coefficient corresponding to the average video signal-to-noise ratio set in the database, d represents the weight coefficient corresponding to the average frame loss rate set in the database, j represents the number of each batch, j = 1, 2, 3, …, m, m is the total number of batches.
[0052] It should be noted that in video processing, the average processing time, the average key frame content retention rate, the average video signal-to-noise ratio and the average frame loss rate jointly determine the overall quality and efficiency of batch processing: first, the average processing time directly reflects the resource consumption and response speed of the system to the denoising filter and compression algorithm, but if the time is excessively reduced, the key frame retention rate or the signal-to-noise ratio will be sacrificed; second, the average key frame content retention rate reflects the detail fidelity of the picture after denoising and compression, if the value is continuously too low, it means that the filter or compression parameter is too aggressive, although the processing time can be shortened, it will lead to detail loss; the signal-to-noise ratio itself can be regarded as a barometer of denoising and compression quality, if it is continuously lower than expected, the processing pace needs to be slowed down (the time consumption is prolonged or the compression strength is reduced) to improve the picture quality, otherwise if the signal-to-noise ratio is much higher than the target, it means that there is redundancy in processing, which can be further compressed to speed up without seriously affecting the detail retention; finally, the average frame loss rate represents the frame compensation and bandwidth bottleneck problem in the process of denoising and compression, once the loss rate rises, even if the time consumption and picture quality indicators seem balanced, there is a serious risk of frame skipping or lag.
[0053] It should be noted that the weight coefficient corresponding to the average processing time, the weight coefficient corresponding to the average key frame content retention rate, the weight coefficient corresponding to the average video signal-to-noise ratio, and the weight coefficient corresponding to the average frame loss rate are stored in the database and the value range is usually set between 0 and 1. For example, by constructing a mapping table between the average processing time, the average key frame content retention rate, the average video signal-to-noise ratio, the average frame loss rate and the weight coefficient respectively, the real-time detected average processing time, the average key frame content retention rate, the average video signal-to-noise ratio and the average frame loss rate are respectively input into the corresponding mapping relationship table in the database, so as to quickly obtain the weight coefficient corresponding to the average processing time, the weight coefficient corresponding to the average key frame content retention rate, the weight coefficient corresponding to the average video signal-to-noise ratio, and the weight coefficient corresponding to the average frame loss rate.
[0054] Specifically, the automatic process of the original video material is optimized, and the specific process is as follows: extracting each batch processing error factor, performing difference processing on each batch processing error factor and the average value of each batch processing error factor to obtain the average error of each batch processing.
[0055] If the average error of a batch is positive, the batch error factor is subtracted from the average of the batch error factor to obtain a first difference value factor of the batch error, the intensity and compression ratio of the noise filter corresponding to each interval of the first difference value factor of the batch error stored in the database are extracted, and the intensity and compression ratio of the noise filter corresponding to each interval of the first difference value factor of the batch error are mapped, denoted as the first adjustment value of the intensity and the first adjustment value of the compression ratio of the noise filter, and the automatic process of the original video material is optimized according to the first adjustment value of the intensity and the first adjustment value of the compression ratio of the noise filter.
[0056] If the average error of a batch is negative, the average of the batch error factor is subtracted from the batch error factor to obtain a second difference value factor of the batch error, the intensity and compression ratio of the noise filter corresponding to each interval of the second difference value factor of the batch error stored in the database are extracted, and the intensity and compression ratio of the noise filter corresponding to each interval of the second difference value factor of the batch error are mapped, denoted as the second adjustment value of the intensity and the second adjustment value of the compression ratio of the noise filter, and the automatic process of the original video material is optimized according to the second adjustment value of the intensity and the second adjustment value of the compression ratio of the noise filter.
[0057] It should be noted that after obtaining the first adjustment value of the intensity and the first adjustment value of the compression ratio of the noise filter, the system will directly map these two adjustment values to the running parameters of the denoising and encoding module: specifically, the filter will automatically select the corresponding convolution kernel coefficient or filter threshold configuration file according to the first adjustment value, load a stronger high-pass filter template to improve the noise suppression effect, and the encoder will use a corresponding higher compression ratio gear (such as increasing the quantization parameter by 10% or reducing the average code rate target by 10%) when starting, both of which work together on the next batch of original video input to further speed up the processing speed on the premise of ensuring the information integrity margin; on the contrary, when the second adjustment value of the intensity and the second adjustment value of the compression ratio of the noise filter are used, the system will automatically switch to a more conservative set of preset parameters: the filter loads a smoother template with lower intensity to reduce the loss of picture details, and the encoder reduces the compression ratio (such as reducing the quantization parameter or increasing the average code rate target), thereby providing higher picture fidelity for the next batch of preprocessing, both optimization paths are checked by the parameter verification module before execution to ensure that the adjustment range does not exceed the specified threshold and takes effect immediately, to complete the closed-loop optimization of the automatic process of the original video material.
[0058] Specifically, the effect of the automatic process adjustment of the original video material is analyzed to determine the need for secondary automatic process adjustment, and the specific process is as follows: a preset observation period is set, the average error of each batch of processing after optimization is obtained in the observation period, if the absolute value of the average error of each batch of processing is higher than or equal to the preset error threshold, the automatic process of the original video material is re-optimized, if the absolute value of the average error of each batch of processing is lower than the preset error threshold, the automatic process of the original video material does not need to be re-optimized, thereby obtaining the multi-modal data of the expected video material of each batch.
[0059] It should be noted that the strength of the enhanced noise filter and the compression ratio are specifically adjusted by automatically fine-tuning the filter coefficients of the current batch, for example, increasing the compression ratio by about 5%-10%, increasing the compression ratio, i.e., increasing the target code rate by 5%-10%, and the system will slightly reduce the filter strength, for example, reducing the filter coefficients by about 5%-10% and reducing the compression ratio, for example, reducing the target code rate by 5%-10%, to retain more details. If the difference between the error factor of the current batch and the average value of the error factors of each batch in the monitoring period is zero, the current filter and compression parameters are kept unchanged.
[0060] It should be noted that when the absolute value of the average error of each batch exceeds the preset threshold, the weight or size parameter of the filter convolution kernel used in the current batch is slightly adjusted, i.e., the amplitude or weight coefficient of the convolution kernel is fine-tuned by 5%-10%. If the error is positive, it means that there is still tolerable redundant information in the picture, so the amplitude or kernel weight of the convolution kernel is increased to make the filter more effective in suppressing residual noise in the next batch; if the error is negative, it means that there is a risk of information loss, so the amplitude or kernel weight of the convolution kernel is reduced to make the filter more relaxed, thereby retaining as much detail and structural features as possible in the subsequent batches.
[0061] Specifically, the cross-modal matching factor of each batch is analyzed, and the specific process is as follows: the average time delay deviation, the time delay distribution variance, and the maximum deviation of the expected video material of each batch are obtained, the average time delay deviation is compared with the defined time delay deviation, the time delay distribution variance is compared with the defined time delay distribution variance, and the maximum deviation is compared with the defined maximum deviation, and the proportion analysis and the introduction of the weight coefficient and the error factor of each batch are performed to obtain the cross-modal matching factor of each batch. The cross-modal matching factor of each batch is used to evaluate and quantify the consistency and matching quality of different modal data in the time synchronization and fusion process.
[0062] It should be noted that after the multi-modal analysis of the expected video material in each batch, the system will first align the labels of the same voice, text OCR and key frame on the time axis, and pair them into several "event units". Then, for each event unit, the system will calculate three kinds of time deviation: first, the absolute value of the time difference between text and voice, text and key frame, and voice and key frame is calculated; the average time delay deviation is the total number of events divided by the sum of the absolute time differences of all event units, which is the overall deviation level; the time delay distribution variance is a measure of the dispersion of these deviation values, reflecting the stability of the deviation.
[0063] It should be noted that the maximum deviation refers to the maximum value among all time delay deviation values (i.e. the absolute time difference between OCR timestamp, ASR timestamp and key frame timestamp) in a batch of multi-modal alignment. It represents the most extreme degree of time misalignment in this batch, and is used to identify the maximum amplitude of time axis misalignment of two or more modal information in the most serious case. It can help the subsequent system to quickly find and locate the serious misalignment caused by sampling rate drift, recognition delay or frame loss, so as to trigger more targeted time correction or window adjustment.
[0064] It should be noted that the cross-modal matching factor of each batch is analyzed under the following conditions:
[0065] ;
[0066] In the formula, r j represents the cross-modal matching factor of the jth batch, t2 j represents the average time delay deviation of the jth batch of video expected material, t2 represents the set boundary time delay deviation in the database, Y1 j represents the time delay distribution variance of the jth batch of video expected material, Y1 represents the set boundary time delay distribution variance in the database, T3 j represents the maximum deviation of the jth batch of video expected material, T3 represents the set boundary maximum deviation in the database, E j represents the processing error factor of the jth batch, A represents the weight coefficient corresponding to the average time delay deviation set in the database, B represents the weight coefficient corresponding to the time delay distribution variance set in the database, L represents the weight coefficient corresponding to the maximum deviation set in the database, U represents the weight coefficient corresponding to the processing error factor set in the database, j represents the number of each batch, j = 1, 2, 3,..., m, m is the total number of batches.
[0067] It should be noted that in the cross-modal alignment processing of multi-modal video data, the three indexes of average time delay deviation, time delay distribution variance and maximum deviation together with the batch processing error factor determine the accuracy and stability of data synchronization. For example, the increase of average time delay deviation usually leads to the expansion of time delay distribution variance, because the larger the overall time synchronization offset is, the more unstable the time difference between different modalities is, thereby increasing the maximum deviation value. The maximum deviation reflects the synchronization failure risk in the extreme case. If the maximum deviation is too large, the processing error factor of the entire batch may increase significantly, indicating that the multi-modal fusion quality is decreasing. Conversely, the reduction of time delay distribution variance means that the time synchronization is more stable, and the maximum deviation is naturally limited. If the average time delay deviation is maintained at a low level, the overall error factor may decrease, and the system runs better.
[0068] It should be noted that the weight coefficient corresponding to the average time delay deviation, the weight coefficient corresponding to the time delay distribution variance, the weight coefficient corresponding to the maximum deviation, and the weight coefficient corresponding to the processing error factor are stored in the database and the value range is usually set between 0 and 1. For example, by constructing a mapping table between the average time delay deviation, the time delay distribution variance, the maximum deviation, the processing error factor and the weight coefficient respectively, the real-time detected average time delay deviation, the time delay distribution variance, the maximum deviation and the processing error factor are respectively input and transmitted to the corresponding mapping relationship table in the database, so as to quickly obtain the weight coefficient corresponding to the average time delay deviation, the weight coefficient corresponding to the time delay distribution variance, the weight coefficient corresponding to the maximum deviation, and the weight coefficient corresponding to the processing error factor.
[0069] It should be noted that if the average time delay deviation is too large (for example, significantly exceeding the historical mean), it means that the timestamp alignment between OCR / ASR and key frame is severely mismatched, and the time window needs to be increased to include more adjacent frames or text in the matching interval, thereby improving the fault tolerance of cross-modal alignment. If it is close to 0 or relatively small, it means that the time alignment effect is good, and the current window can be maintained to avoid unnecessarily increasing the window and introducing more noise. If the time delay distribution variance is too large, it means that the offset distribution of each batch is extremely unstable, and the fluctuation amplitude is high. The system can consider increasing the time window to cover more uncertainties. If it is too small and stable, it means that the alignment accuracy is very concentrated and reliable, and the time window can be appropriately narrowed to reduce redundant calculation and improve efficiency.
[0070] Specifically, the multi-modal data time dislocation determination of each batch of video expected materials is performed, and the specific process is that the cross-modal matching factor of each batch is extracted and compared with the set cross-modal matching factor threshold. If the cross-modal matching factor of a batch is higher than or equal to the cross-modal matching factor threshold, the time window of the multi-modal data of the batch of video expected materials is updated. If the cross-modal matching factor of a batch is lower than the cross-modal matching factor threshold, the time window of the multi-modal data of the batch of video expected materials does not need to be updated.
[0071] It should be noted that the time window of the multi-modal data of the batch of video expected materials is updated, and the specific process is that the cross-modal matching factor is significantly low, indicating that the multi-modal data (speech text, subtitle OCR, video frame label) is time dislocated, and the time window is too narrow to accommodate the normal delay between modalities. The time window needs to be enlarged, for example, adjusted by 10 to 20%, so as to cover more time context and increase the matching opportunity, thereby improving the alignment degree.
[0072] Specifically, the time window of the multi-modal data of each batch of video expected materials is dynamically updated, and the specific process is that a second observation period is preset after the time window is updated, the cross-modal matching factor of each batch after optimization is obtained in the second observation period, the number of batches whose cross-modal matching factor is higher than or equal to the preset cross-modal matching factor threshold is counted, and the number is recorded as the time dislocation number of each batch. If the time dislocation number of each batch is higher than or equal to the preset time dislocation number threshold, the time window of the multi-modal data of each batch of video expected materials is re-updated. If the time dislocation number of each batch is lower than the preset time dislocation number threshold, the time window of the multi-modal data of each batch of video expected materials does not need to be re-updated.
[0073] It should be noted that the time window of the multi-modal data of each batch of video expected materials is re-updated, and the specific process is that if the time dislocation number of each batch exceeds or reaches the preset time dislocation number threshold, it indicates that the current time window update strategy still cannot effectively control the time deviation of multi-modal alignment, and there is still a systematic dislocation risk. Therefore, the time window needs to be re-updated, and at this time the time window should be further adjusted by 5% to 10% on the basis of the last update step.
[0074] Specifically, the multi-modal data of each batch of video expected materials is audited, and the specific process is that the pass rate of each batch of audit results of each batch of video expected materials is collected within the preset time window, and the average audit coherence factor is analyzed.
[0075] It should be noted that, in a preset time window, the system will continuously record the audit pass rate of each batch to form a time-varying audit pass rate time sequence signal; then the short-time Fourier transform is applied to the signal, and the power spectrum density of each frequency point is obtained, and further the coherence analysis of the same signal in adjacent time periods is performed to obtain the coherence factor of each frequency; finally, the coherence factors of all frequency points are averaged to obtain an average audit coherence factor that comprehensively reflects the stability and periodicity of the audit pass rate value. The higher the factor, the smaller the fluctuation and the stronger the stability of the audit pass rate in the monitoring period, providing a quantitative evaluation basis for subsequent audit strategy and process optimization.
[0076] Extract the audit coherence factor threshold value stored in the database.
[0077] If the average audit coherence factor is higher than or equal to the audit coherence factor threshold value, the multi-modal data of the audited video expected material of each batch is imported into the large language model.
[0078] If the average audit coherence factor is lower than the audit coherence factor threshold value, the verification trigger frequency is adjusted.
[0079] Specifically, the verification trigger frequency is adjusted, and the multi-modal data of the audited video expected material of each batch is imported into the large language model to obtain the structured risk information of the graph construction of the video expected material of each batch. The specific process is: extract the average audit coherence factor and the audit coherence factor threshold value, subtract the average audit coherence factor from the audit coherence factor threshold value to obtain the verification trigger bias value.
[0080] Extract the verification trigger frequency corresponding to each interval of the verification trigger bias value stored in the database, and map the verification trigger frequency corresponding to each interval of the verification trigger bias value, denoted as the verification trigger frequency adjustment value, and adjust the verification trigger frequency according to the verification trigger frequency adjustment value.
[0081] Extract the prompt template stored in the database, and import the multi-modal data of the audited video expected material of each batch and the prompt template into the large language model to output the structured risk information of the graph construction of the video expected material of each batch.
[0082] It should be noted that the specific process of adjusting the verification trigger frequency according to the verification trigger frequency adjustment value is: multiplying the current verification trigger frequency by the verification trigger frequency adjustment value to obtain the verification trigger frequency to be executed this time, and configuring the frequency to the scheduling of content audit as the basis for the trigger period of subsequent automated content compliance and effectiveness verification, that is, in the subsequent video multi-modal audit process, the system will perform compliance check and consistency check on video frames, OCR subtitles, speech transcription text, and structured labels according to the updated verification trigger frequency.
[0083] It should be noted that the multimodal data of each batch (including descriptions of keyframe images, text extracted by OCR, speech text transcribed by ASR, and scene labels, etc.) are organized into the input carrier of the large language model in a unified JSON format and submitted to the large language model along with the prompt template. The model first performs named entity recognition under the guidance of the prompt, automatically locating entity types related to industry risks, such as "leakage event", "security warning", "violation", etc. Then, the model extracts relationships based on the semantic context of the entity and identifies the "causal", "parallel", or "control" relationships between entities. Next, the model further performs event extraction, extracting attributes such as time, location, and responsible party from the identified risky behaviors or events, and automatically labeling each entity-relationship-event triple with risk level and compliance status according to industry compliance rules. Finally, the big language model summarizes this structured risk information into a standard format that can be written to the graph (such as a JSON array with entity ID, relationship type, timestamp, and risk score), and returns the results to the pipeline. The subsequent graph writing component then injects nodes and edges into the graph database in batches based on these triples, thereby completing the generation of structured risk information corresponding to each batch of expected video materials.
[0084] like Figure 2 As shown, Figure 2 The diagram illustrates the logical flow of an industry risk knowledge graph construction method based on a large language model, as described in this invention. First, the collected original video materials are denoised, compressed, and error monitored in batches to determine if secondary optimization is needed in the automated process. This ensures the availability and accuracy of each batch of multimodal data. The optimized multimodal data is then input into a multimodal parsing module to extract temporal and semantic consistency. Finally, the multimodal data that has passed content review and completed time alignment is imported into the large language model to automatically generate structured risk information for the graph. Through multi-batch error closed-loop optimization and a dynamic adjustment mechanism for multimodal temporal misalignment, the robustness of multi-source video data is significantly enhanced. Furthermore, before extracting structured information, the temporal consistency and compliance of the multimodal data are dynamically updated multiple times, allowing the large language model to utilize more accurate input during the graph construction stage.
[0085] like Figure 3 As shown, the second aspect of the present invention also provides an industry risk knowledge graph construction system based on a large language model, including: a preprocessing optimization module, used to perform batch denoising and compression processing on the original video material set of the target industry and monitor the processing process, analyze the processing error factors of each batch, and optimize the automated process of the original video material.
[0086] A process adjustment evaluation module is configured to analyze the effect of the automatic process adjustment of the original video material, thereby determining the need for secondary automatic process adjustment execution, and obtaining multi-modal data of the video expected material of each batch.
[0087] A multi-modal analysis module is configured to input the multi-modal data of the video expected material of each batch into a multi-modal analysis pipeline, analyze the cross-modal matching factor of each batch, and perform multi-modal data time dislocation determination of the video expected material of each batch.
[0088] A time alignment and compliance audit module is configured to dynamically update the time window of the multi-modal data of the video expected material of each batch, perform content audit on the multi-modal data of the video expected material of each batch, adjust the verification trigger frequency, and input the audited multi-modal data of the video expected material of each batch into a large language model to obtain structured risk information of the graph construction of the video expected material of each batch.
[0089] It should be noted that, in this document, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0090] The preferred embodiments of the application disclosed above are only used to help explain the application. The preferred embodiments do not describe all the details of the application, nor limit the application to the specific embodiments described. It is obvious that many modifications and variations can be made according to the content of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the application, so that those skilled in the art can well understand and utilize the application. The application is limited only by the claims and their full scope and equivalents.
Claims
1. A method for constructing an industry risk knowledge graph based on a large language model, characterized in that, The method comprises the following steps: Batch denoising and compression processing of original video material sets of target industries are performed, and the processing process is monitored, the processing error factors of each batch are analyzed, and the automatic process of the original video material is optimized; The adjustment effect of the automatic process of the original video material is analyzed, thereby determining the execution demand of the secondary automatic process adjustment, and obtaining the multi-modal data of the expected video material of each batch; The multi-modal data of the expected video material of each batch is input into a multi-modal analysis pipeline, the cross-modal matching factor of each batch is analyzed, and the multi-modal data time dislocation of the expected video material of each batch is determined; The time window of the multi-modal data of the expected video material of each batch is dynamically updated, the multi-modal data of the expected video material of each batch is audited, the verification trigger frequency is adjusted, the multi-modal data of the expected video material of each batch after the audit is input into a large language model, and the structured risk information of the graph construction of the expected video material of each batch is obtained; The specific process of analyzing the processing error factors of each batch is as follows: The average processing time, average key frame content retention rate, average video signal-to-noise ratio and average frame loss rate of the processing batch of the original video material of each batch are obtained; The average processing time, average key frame content retention rate, average video signal-to-noise ratio and average frame loss rate of the processing batch of the original video material of each batch are obtained; The specific process of analyzing the cross-modal matching factor of each batch is as follows: The average time delay deviation, time delay distribution variance and maximum deviation of the expected video material of each batch are obtained, the average time delay deviation, time delay distribution variance and maximum deviation are analyzed by percentage, and weight coefficients are introduced to obtain the cross-modal matching factor of each batch.
2. The industry risk knowledge graph construction method based on a large language model according to claim 1, characterized in that: The specific process of optimizing the automatic process of the original video material is as follows: The processing error factors of each batch are extracted, the processing error factors of each batch are subtracted from the average value of the processing error factors of each batch, and the average processing error of each batch is obtained; If the average processing error of a batch is positive, the processing error factor of the batch is subtracted from the average value of the processing error factor of the batch, the first difference factor of the processing error of the batch is obtained, the strength of the noise filter and the compression ratio corresponding to each interval of the first difference factor of the batch processing error stored in the database are extracted, and the strength of the noise filter and the compression ratio corresponding to each interval of the first difference factor of the batch processing error are mapped, which are denoted as the first adjustment value of the strength of the noise filter and the first adjustment value of the compression ratio, and the automatic process of the original video material is optimized according to the first adjustment value of the strength of the noise filter and the first adjustment value of the compression ratio. If the average error of a batch is negative, the average value of the batch error factor is subtracted from the batch error factor to obtain a second difference value factor of the batch error, the intensity and compression ratio of the noise filter corresponding to each interval of the stored second difference value factor of the batch error are extracted, and the intensity and compression ratio of the noise filter corresponding to each interval of the second difference value factor of the batch error are mapped, denoted as the second adjustment value of the intensity and the second adjustment value of the compression ratio of the noise filter, and the original video material is optimized according to the second adjustment value of the intensity and the second adjustment value of the compression ratio of the noise filter.
3. The industry risk knowledge graph construction method based on a large language model according to claim 2, characterized in that: The effect of the automatic process of analyzing the original video material is adjusted to determine the need for secondary automatic process adjustment, and the specific process is as follows: A preset observation period is set, and the average error of each batch after optimization is obtained in the observation period. If the absolute value of the average error of each batch is higher than or equal to the preset error threshold, the automatic process of the original video material is re-optimized, and if the absolute value of the average error of each batch is lower than the preset error threshold, the automatic process of the original video material does not need to be re-optimized, thereby obtaining the multi-modal data of each batch of video expected material.
4. The industry risk knowledge graph construction method based on a large language model according to claim 1, characterized in that: The time dislocation of the multi-modal data of each batch of video expected material is determined, and the specific process is as follows: The cross-modal matching factor of each batch is extracted and compared with the set cross-modal matching factor threshold. If the cross-modal matching factor of a batch is higher than or equal to the cross-modal matching factor threshold, the time window of the multi-modal data of the batch of video expected material is updated, and if the cross-modal matching factor of a batch is lower than the cross-modal matching factor threshold, the time window of the multi-modal data of the batch of video expected material does not need to be updated.
5. The industry risk knowledge graph construction method based on a large language model according to claim 4, characterized in that: The time window of the multi-modal data of each batch of video expected material is dynamically updated, and the specific process is as follows: A second observation period is set after the time window is updated, and the cross-modal matching factor of each batch after optimization is obtained in the second observation period. The number of batches with cross-modal matching factors higher than or equal to the preset cross-modal matching factor threshold is counted and denoted as the time dislocation number of each batch. If the time dislocation number of each batch is higher than or equal to the preset time dislocation number threshold, the time window of the multi-modal data of each batch of video expected material is re-updated, and if the time dislocation number of each batch is lower than the preset time dislocation number threshold, the time window of the multi-modal data of each batch of video expected material does not need to be re-updated.
6. The industry risk knowledge graph construction method based on a large language model according to claim 1, characterized in that: The multi-modal data of each batch of video expected material is audited, and the specific process is as follows: In the preset time window, the pass rate of each batch of video expected material is collected, and the average audit coherence factor is analyzed. The audit coherence factor threshold stored in the database is extracted. If the average audit coherence factor is higher than or equal to the audit coherence factor threshold, the multi-modal data of each batch of video expected material after audit is imported into a large language model. If the average audit coherence factor is lower than the audit coherence factor threshold, the verification trigger frequency is adjusted.
7. The industry risk knowledge graph construction method based on a large language model according to claim 6, characterized in that: The verification trigger frequency is adjusted, the multi-modal data of the audited video expected material of each batch is introduced into the large language model to obtain the structured risk information of the graph construction of the video expected material of each batch, and the specific process is as follows: An average audit coherence factor and an audit coherence factor threshold are extracted, the audit coherence factor threshold is subtracted by the average audit coherence factor to obtain a verification trigger deviation value; The verification trigger frequency corresponding to each interval of the verification trigger deviation value stored in the database is extracted, and the verification trigger frequency corresponding to each interval of the verification trigger deviation value is mapped and extracted, which is recorded as a verification trigger frequency adjustment value, and the verification trigger frequency is adjusted according to the verification trigger frequency adjustment value; The prompt template stored in the database is extracted, and the multi-modal data of the audited video expected material of each batch and the prompt template are introduced into the large language model to output the structured risk information of the graph construction of the video expected material of each batch.
8. A system applying the industry risk knowledge graph construction method based on a large language model according to any one of claims 1-7, comprising: A preprocessing optimization module for batch-wise denoising and compression processing of the target industry original video material set and monitoring the processing process, analyzing the processing error factors of each batch, and optimizing the automatic process of the original video material; A process adjustment evaluation module for analyzing the adjustment effect of the automatic process of the original video material, thereby determining the secondary automatic process adjustment execution demand, and obtaining the multi-modal data of the video expected material of each batch; A multi-modal analysis module for inputting the multi-modal data of the video expected material of each batch into a multi-modal analysis pipeline, analyzing the cross-modal matching factor of each batch, and performing multi-modal data time dislocation judgment of the video expected material of each batch; A time alignment and compliance audit module for dynamically updating the time window of the multi-modal data of the video expected material of each batch, auditing the content of the multi-modal data of the video expected material of each batch, adjusting the verification trigger frequency, introducing the multi-modal data of the audited video expected material of each batch into the large language model to obtain the structured risk information of the graph construction of the video expected material of each batch.
Citation Information
Patent Citations
Power grid risk knowledge graph construction method and system
CN118069861A
Knowledge graph construction and application method based on insurance industry
CN120104808A
Multi-modal large-model-driven figure knowledge graph construction method and multi-modal large-model-driven figure knowledge graph construction system
CN120409657A
AIGC content generation method and system based on multi-modal fusion
CN120578796A