Intelligent preprocessing method and system based on multi-source heterogeneous data fusion

By determining the correlation analysis strategy of data capture nodes based on node mutation coefficients and complex coefficients, and using integrated and co-occurring correlation analysis, the problem of low data integration quality in the existing technology is solved, and the data fusion effect and decision analysis efficiency are improved.

CN120145320AActive Publication Date: 2025-06-13北京数字航宇科技有限公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510629789.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The prior art fails to determine targeted correlation analysis methods based on the actual situation of the regulatory area, resulting in poor integration quality of multi-source heterogeneous data, which in turn affects the data fusion effect and decision-making analysis efficiency.

Method used

By determining the association analysis strategy of each data capture node in the target service area based on the node mutation coefficient and the node complex coefficient, integrating association analysis or co-occurrence association analysis is adopted, and the optimization method of integrated data bundles is determined based on the integration analysis index and mutation association index.

Benefits of technology

It improves the fusion effect and integration quality of multi-source heterogeneous data, enhances the efficiency of transmission and analysis of regulatory data, and ensures the effectiveness of decision-making analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145320A_ABST
    Figure CN120145320A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, in particular to an intelligent preprocessing method and system based on multi-source heterogeneous data fusion. Determining a correlation analysis strategy of each data capture node in the target service area according to the node mutation coefficient and the node complex coefficient, and performing integration correlation analysis or co-occurrence correlation analysis on each fusion analysis data of the data capture node; during integration correlation analysis, determining an integration strategy of one type of data capture nodes according to an integration analysis index; during co-occurrence association analysis, determining an integrated data cluster of each second-class data capture node according to a mutation association index; under the condition that correlation analysis is completed, the data optimization mode of each corresponding integrated data cluster is determined according to the correlation analysis strategy of each data capture node, the fusion effect of the multi-source heterogeneous data is improved, and then the decision analysis efficiency based on the fusion result of the multi-source heterogeneous data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to an intelligent preprocessing method and system based on multi-source heterogeneous data fusion. Background Art

[0002] When conducting supervision for public areas, performing fusion processing on the obtained multi-source heterogeneous data can effectively improve the reliability and timeliness of the subsequent decision-making and analysis process. Among them, by effectively associating and preprocessing the content of multi-source heterogeneous data, the fusion result of multi-source heterogeneous data can be ensured. In the actual supervision process, the personnel flow and activity rules in different supervision areas are different. Using a single data association analysis method is likely to cause the quality of the obtained data integration not to meet the actual needs, and further lead to poor efficiency and effectiveness of the subsequent decision-making and analysis based on the data fusion result. Therefore, how to determine targeted data association analysis methods for different supervision areas in the actual supervision process to avoid low data integration quality is an urgent problem to be solved by those skilled in the art.

[0003] Chinese Patent Application Publication No. CN118839292A discloses a multi-source heterogeneous data fusion method applied to a smart city, including the following steps: S1, data collection and preprocessing; S2, data integration and matching, integrating and matching data from different data sources, and establishing a data model and semantic association; S3, data fusion and analysis, using data fusion technology to combine and integrate the integrated data to generate a more comprehensive and accurate description of the urban state; S4, modeling and prediction. However, the above solution has the following problems: It fails to determine a targeted association analysis method according to the actual situation of the supervision area to integrate and match different data, resulting in poor quality of the data integration result, and further leading to low fusion effect for multi-source heterogeneous data. Summary of the Invention

[0004] Therefore, the present invention provides an intelligent preprocessing method and system based on multi-source heterogeneous data fusion to overcome the problem in the prior art that a targeted association analysis method cannot be determined according to the actual situation of the supervision area to integrate and match different data, resulting in poor quality of the data integration result, and further leading to low fusion effect for multi-source heterogeneous data.

[0005] To achieve the above object, the present invention provides an intelligent preprocessing method based on multi-source heterogeneous data fusion, including: Determining the association analysis strategy for each data capture node in the target service area as integrated association analysis or co-occurrence association analysis for each fusion analysis data of the data capture node according to the node mutation coefficient and the node complexity coefficient; When performing integrated association analysis, the integration strategy for a type of data capture node is determined based on the integrated analysis index, which is an integrated data cluster determined according to the reference integration index and the reference flow association index, or the node conflict parameter and the reference difference stability parameter. When performing co-occurrence association analysis, the integrated data cluster of each type-II data capture node is determined according to the mutation association index. Under the condition that the association analysis is completed, the data optimization method for each corresponding integrated data cluster is determined according to the association analysis strategy of each data capture node. The method is to determine the preferred video segment according to each optimization analysis combination of the integrated data cluster, or to determine the key video frames according to the proportion of integrated nodes in the integrated data cluster.

[0006] Furthermore, periodically detect the node mutation coefficient and the node complexity coefficient of each data capture node in the target service area to determine the association analysis strategy of each data capture node. For a single data capture node, the node mutation coefficient is determined according to the regulatory mutation index and the mutation duration index of the data capture node, and the node complexity coefficient is determined according to the node connectivity index and the reference connectivity mutation index of the data capture node.

[0007] Furthermore, for any data capture node, if the node mutation coefficient of the data capture node is greater than the preset node mutation coefficient or the node complexity coefficient is greater than the preset node complexity coefficient, then perform integrated association analysis on the various fusion analysis data of the data capture node.

[0008] Furthermore, for any data capture node, if the node mutation coefficient of the data capture node is less than or equal to the preset node mutation coefficient and the node complexity coefficient is less than or equal to the preset node complexity coefficient, then perform co-occurrence association analysis on the various fusion analysis data of the data capture node.

[0009] Furthermore, under the condition of integrated analysis, determine the integrated analysis index of each type-I data capture node according to the node mutation richness and the connectivity mutation coefficient, and determine the integration strategy of each type-I data capture node according to the integrated analysis index, so as to divide the fusion analysis data of each type-I data capture node and its connected nodes. The integrated analysis index has a positive correlation with the node mutation richness and the connectivity mutation coefficient respectively. The integrated analysis condition is that there is fusion analysis data of data capture nodes that needs to be subjected to integrated association analysis. The data capture nodes to be subjected to integrated association analysis are denoted as type-I data capture nodes.

[0010] Furthermore, for a single type-I data capture node, If the integration analysis index of this type of data capture node is greater than the preset integration analysis index, determine the integrated data cluster of this type of data capture node according to the reference integration index and the reference flow correlation index; If the integration analysis index of this type of data capture node is less than or equal to the preset integration analysis index, determine the integrated data cluster of this type of data capture node according to the node conflict parameter and the reference difference stability parameter.

[0011] Further, under the co-occurrence analysis condition, divide the fusion analysis data of each type-II data capture node according to the mutation correlation index to determine the integrated data cluster of each type-II data capture node; The co-occurrence analysis condition is that there is fusion analysis data of data capture nodes that needs to be subjected to co-occurrence correlation analysis, and the data capture nodes subjected to co-occurrence correlation analysis are denoted as type-II data capture nodes.

[0012] Further, under the condition that the association analysis is completed, optimize the data for each integrated data cluster, and determine the data optimization method for each corresponding integrated data cluster according to the association analysis strategy of each data capture node; For a single integrated data cluster, If this integrated data cluster is determined through co-occurrence correlation analysis, determine the preferred video segment according to each optimization analysis combination of this integrated data cluster; If this integrated data cluster is determined through integration correlation analysis, determine the determination method of the key video frame according to the proportion of integrated nodes in this integrated data cluster; The condition for the completion of the association analysis is that there are data capture nodes in the target service area that have completed the determination of the integrated data cluster.

[0013] Further, for a single integrated data cluster determined through integration correlation analysis, If the proportion of integrated nodes is greater than the preset proportion of integrated nodes, determine the key video frame of this integrated data cluster according to the associated connected nodes; If the proportion of integrated nodes is less than or equal to the preset proportion of integrated nodes, determine the key video frame of this integrated data cluster according to the mutation dispersion index.

[0014] The present invention also provides an intelligent preprocessing system applying the intelligent preprocessing method based on multi-source heterogeneous data fusion as described above, including: A node evaluation module, configured to determine the association analysis strategy for each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient to perform integration correlation analysis or co-occurrence correlation analysis on each piece of fusion analysis data of the data capture node; An integration analysis module, which is connected to the node evaluation module, and is used to determine the integration strategy of a class of data capture nodes according to the integration analysis index as the integrated data cluster determined according to the reference integration index and the reference flow correlation index, or, the node conflict parameter and the reference difference stability parameter; A co-occurrence analysis module, which is connected to the node evaluation module, and is used to determine the integrated data cluster of each second-class data capture node according to the mutation correlation index; An optimization execution module, which is respectively connected to the integration analysis module and the co-occurrence analysis module, and is used to determine the data optimization method of the corresponding integrated data clusters according to the association analysis strategy of each data capture node as determining the preferred video segment according to each optimization analysis combination of the integrated data cluster, or, determining the key video frames according to the proportion of the integrated nodes of the integrated data cluster.

[0015] Compared with the prior art, the beneficial effects of the present invention are that the technical solution of the present invention determines the targeted association analysis strategy of each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient, so that the process of dividing the fusion analysis data conforms to the actual situation of the data obtained in the supervision process of each data capture node. In addition, the data in the obtained integrated data cluster is optimized specifically, while ensuring the improvement of the fusion effect of the obtained multi-source heterogeneous data, the transmission and analysis efficiency of the supervision data is improved.

[0016] Further, in the present invention, the node mutation coefficient and the node complexity coefficient are used to determine how to perform association analysis for different data capture nodes. Among them, the node mutation coefficient and the node complexity coefficient can characterize the frequency of mutation of the supervision data of the data capture node in a recent period and the interference probability of the surrounding area, so as to set a targeted method for determining the integrated data cluster, while ensuring that the obtained integrated data cluster can fully contain the associated fusion analysis data, and ensuring the division efficiency of the integrated data cluster.

[0017] Further, in the present invention, for a class of data capture nodes, the integration strategy is determined according to the integration analysis index. The supervision data of the sub-service area supervised by such data capture nodes has a large mutation probability or is greatly affected by the fluctuations in the surrounding area. Therefore, the integration processing of its fusion analysis data needs to consider the spatio-temporal association with the supervision data obtained by other data capture nodes.

[0018] Furthermore, the present invention further determines the situation of data mutation of a class of data capture nodes and their connected regions through the integrated analysis index. When the integrated analysis index is relatively large, it indicates that there is a relatively large data mutation in the connected nodes as a whole, and there is also a large overlap in the regulatory data corresponding to the mutation moments. By referring to the integrated index and the reference flow correlation index, the association between the various fusion analysis data within the determined integrated data cluster and the association degree between the mutation of the regulatory data and the personnel flow are ensured. When the integrated analysis index is relatively small, it indicates that there is less regulatory data with mutation in the connected region. At this time, through the node conflict parameter and the reference difference stability parameter, it is ensured that the sub-service regions corresponding to the various data integrated therein are within a relatively close range and the time interval between the moments when the data mutates is relatively stable, that is, it is ensured that the data it integrates can effectively represent the overall state of the data in its surrounding region, and the effectiveness of the data fusion result for the decision-making analysis process is ensured.

[0019] Furthermore, in the present invention, the data optimization method of each corresponding integrated data cluster is determined according to the association analysis strategy of each data capture node. For the data capture node determined through the co-occurrence association analysis, the dominant factor for the change in its regulatory data tends to be the change in the personnel activities within its corresponding sub-service region. By analyzing and combining some words to determine the interception moment of the preferred video segment, it effectively ensures that the obtained video information can assist in analyzing the cause of data mutation, and while ensuring the effectiveness of the fusion result of multi-source heterogeneous data, it improves the subsequent analysis and transmission efficiency of the fusion result.

[0020] Furthermore, for the data capture node determined through the integrated association analysis, the dominant factor for the change in its regulatory data mainly tends to be the personnel circulation situation between different regions. By determining the setting method of the time interval between frames during the extraction of key video frames through the integrated node ratio, it is ensured that the extraction of key video frames is more adaptable to the data change situation in the actual integrated data cluster. The present invention improves the subsequent analysis and transmission efficiency of the fusion result while ensuring the effectiveness of the fusion result of multi-source heterogeneous data. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the intelligent preprocessing method based on multi-source heterogeneous data fusion of the present invention; Figure 2 It is a flowchart of determining the association analysis strategy according to the node mutation coefficient and the node complexity coefficient of the present invention; Figure 3 It is a flowchart of determining the integration strategy of a class of data capture nodes according to the integrated analysis index of the present invention; Figure 4 It is a module connection diagram of the intelligent preprocessing system based on multi-source heterogeneous data fusion of the present invention. Detailed implementation manners

[0022] In order to make the objectives and advantages of the present invention more clear and understandable, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0023] The preferred implementation manners of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0024] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0025] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0026] Please refer to Figures 1 to 3 As shown, the present invention provides an intelligent preprocessing method for multi-source heterogeneous data fusion, including: Determining the correlation analysis strategy for each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient, which is to perform integrated correlation analysis or co-occurrence correlation analysis on each piece of fusion analysis data of the data capture node; During integrated correlation analysis, determining the integration strategy for a class of data capture nodes according to the integration analysis index, which is the integrated data cluster determined according to the correlation quality coefficient and the integration conflict coefficient, or the integrated data cluster of a class of data capture nodes determined according to the integration conflict coefficient; During co-occurrence correlation analysis, determining the integrated data cluster of each second-class data capture node according to the mutation correlation index and the correlation stability index; Under the condition that the correlation analysis is completed, determining the data optimization method according to the integration correlation ratio of the integrated data cluster for optimization processing, and the data optimization method is to determine the priority optimization coefficient of each video segment according to the correlation activity connection parameter and the stage matching index, or to determine the priority optimization coefficient of each video segment according to the correlation richness index and the stage matching index.

[0027] Among them, in the process of the present invention for monitoring public areas, data fusion processing is performed on the obtained monitoring data. The area to be monitored is denoted as the target service area. There are several data capture nodes in the present invention. Any one data capture node corresponds to only one sub-service area, and several monitoring devices for obtaining various monitoring data are provided. Each sub-service area is obtained by dividing the target service area, and users can perform adaptive division according to the actual working scenario. In the present invention, the types of monitoring devices and the monitoring data set for each data capture node are not specifically limited, and users can perform adaptive settings according to the actual working scenario. However, each data capture node is provided with a video acquisition device, and the specific installation location and device model are not limited, but it is necessary to ensure that the video data obtained by each data capture node can cover the corresponding sub-service area. The types of monitoring data in the present invention include, but are not limited to, the temperature, carbon dioxide concentration, noise decibel, and video monitoring data of the sub-service area corresponding to each data capture node.

[0028] In the present invention, several monitoring processing records are applied. Any one monitoring processing record records at least once the reference mutation degree, activity connectivity parameter, node mutation coefficient, node complexity coefficient, integration analysis index, integration quality coefficient, correlation quality coefficient, mutation correlation index, and integration node ratio during the data fusion processing of the various monitoring data obtained for the target service area. And each monitoring processing record corresponds to a qualified mark, and the qualified mark records whether the fusion quality of the multi-source heterogeneous data obtained during the monitoring process meets the user's requirements. It can be understood that users can determine whether the fusion quality of the multi-source heterogeneous data obtained during the monitoring process meets the requirements according to the self-set indicators. For example, the self-set indicators can be, but are not limited to, the decision response time, and the decision response time is the time used to output the decision result based on the fusion result of the obtained multi-source heterogeneous data.

[0029] Specifically, the node mutation coefficient and the node complexity coefficient of each data capture node in the target service area are periodically detected to determine the correlation analysis strategy of each data capture node; For a single data capture node, The node mutation coefficient is determined according to the monitoring mutation index and the mutation duration index of the data capture node, and the node complexity coefficient is determined according to the node connectivity index and the reference connectivity mutation index of the data capture node.

[0030] Among them, a cyclic node evaluation period is applied in the present invention. The duration of the node evaluation period can be determined by the user. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process, the shorter the duration of the node evaluation period. A duration of the node evaluation period is provided, and the duration of the node evaluation period is 1h. At the end of each node evaluation period, the node mutation coefficient and the node complexity coefficient of each data capture node are detected to determine the correlation analysis strategy of each data capture node. A cyclic data supervision period is also applied in the present invention. The duration of the data supervision period can be determined by the user. A duration of the data supervision period is provided, and the duration of the data supervision period is 10min. At the end of each data supervision period, the data that has completed data fusion processing during the current data supervision period is analyzed to output a decision analysis result. How to output a decision analysis result based on the data fusion results of each supervision data is easily understood by those skilled in the art and will not be elaborated here; For a single data capture node, the node mutation coefficient is the product of the supervision mutation index and the mutation duration index of the data capture node. The supervision mutation index is the number of mutation moments existing during the node evaluation stage. The mutation duration index is the maximum value of the mutation duration of each mutation moment during the node evaluation stage. The node complexity coefficient is the product of the node connectivity index and the reference connectivity mutation index of the data capture node. The node connectivity index is the number of connected nodes of the data capture node. The reference connectivity mutation index is the average value of the supervision mutation indices of each connected node of the data capture node. The end moment of the node evaluation stage is the end moment of the current node evaluation period. The duration of the node evaluation stage can be set by the user according to the actual working scenario. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process, the longer the duration of the node evaluation stage. A duration of the node evaluation stage is provided, and the duration of the node evaluation stage is 6 times the duration of the node evaluation period; For the end moment of a single data supervision cycle, if the reference mutation degree at this moment is greater than the preset reference mutation degree, then this moment is recorded as the mutation moment. The reference mutation degree is the average value of the mutation degrees of each supervision data obtained at this moment. For a single item of supervision data, the mutation degree corresponding to this moment = |the maximum value of this item of supervision data obtained within the current data supervision cycle - the maximum value of this item of supervision data obtained within the previous data supervision cycle of the current data supervision cycle| / the maximum value of this item of supervision data obtained within the previous data supervision cycle of the current data supervision cycle. For a single mutation moment, the maximum value of the mutation degrees of each supervision data at this mutation moment is recorded as the evaluation mutation degree. The value of the supervision data corresponding to the evaluation mutation degree obtained within the previous data supervision cycle during the determination process of this mutation moment is recorded as the continuous reference value. The data supervision cycle in which the maximum value of the value of the supervision data corresponding to the evaluation mutation degree after this mutation moment is less than the continuous reference value is recorded as the recovery moment. The mutation duration is the minimum value of the interval durations between this mutation moment and its respective recovery moments. The measurement unit of the mutation duration is hours; The value of the preset reference mutation degree can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision processing records. The higher the user's requirement for the fusion quality of multi-source heterogeneous data obtained during the supervision process, the smaller the value of the preset reference mutation degree. A method for obtaining the value of the preset reference mutation degree is provided. The average value of the reference mutation degrees of each mutation moment in the supervision processing records that meet the requirements for the fusion quality of multi-source heterogeneous data obtained during the supervision process is recorded as the preset reference mutation degree; For any two data capture nodes, if the activity connectivity parameter between the above two data capture nodes is greater than the preset activity connectivity parameter, then it is determined that the above two data capture nodes are mutually connected nodes. The activity connectivity parameter between the above two data capture nodes is determined according to the personnel flow correlation degree and the connectivity path index. The activity connectivity parameter is the product of the personnel flow correlation degree and the connectivity path index. The connectivity path index is the shortest duration required to go from one data capture node to the other data capture node. The measurement unit of the connectivity path index is minutes. The personnel flow correlation degree = the number of people who have simultaneously existed in the sub-service areas corresponding to the above two data capture nodes during the node evaluation stage / the personnel flow parameter of the above two data capture nodes during the node evaluation stage. For a single data capture node, the personnel flow parameter is the maximum value of the number of people present in this data capture node obtained each time during the node evaluation stage; The value of the preset activity connection parameter can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process, the smaller the value of the preset activity connection parameter. A method for obtaining the value of the preset activity connection parameter is provided. The average value of the activity connection parameters between the connected nodes in the supervision processing record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process is recorded as the preset activity connection parameter.

[0031] Specifically, for any data capture node, if the node mutation coefficient of the data capture node is greater than the preset node mutation coefficient or the node complexity coefficient is greater than the preset node complexity coefficient, then integrated correlation analysis is performed on the various fusion analysis data of the data capture node.

[0032] Specifically, for any data capture node, if the node mutation coefficient of the data capture node is less than or equal to the preset node mutation coefficient and the node complexity coefficient is less than or equal to the preset node complexity coefficient, then co-occurrence correlation analysis is performed on the various fusion analysis data of the data capture node.

[0033] Among them, for a single data capture node, the various supervision data obtained within the current node evaluation period are recorded as the fusion analysis data that the data capture node currently needs to perform fusion processing on. The correlation analysis strategy is determined according to the node mutation coefficient and the node complexity coefficient. The values of the preset node mutation coefficient and the preset node complexity coefficient can be determined by the user according to the actual working scenario. For example, the user can set them according to the supervision processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process, the smaller the value of the preset node mutation coefficient, and the smaller the value of the preset node complexity coefficient. A method for obtaining the value of the preset node mutation coefficient is provided. The supervision processing record that performs integrated correlation analysis on the various fusion analysis data of the data capture node is recorded as the correlation reference record. The average value of the node mutation coefficients of each data capture node in the correlation reference record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process is recorded as the preset node mutation coefficient. A method for obtaining the value of the preset node complexity coefficient is provided. The average value of the node complexity coefficients of each data capture node in the correlation reference record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process is recorded as the preset node complexity coefficient.

[0034] Specifically, under the condition of integrated analysis, the integrated analysis index of each type I data capture node is determined according to the node mutation richness and the connected mutation coefficient, and the integration strategy of each type I data capture node is determined based on the integrated analysis index, so as to divide the fusion analysis data of each type I data capture node and its connected nodes; The integrated analysis index is positively correlated with the node mutation richness and the connected mutation coefficient respectively; The integrated analysis condition is that the integrated correlation analysis needs to be carried out on the integrated analysis data with data capture nodes. The data capture nodes for which the integrated correlation analysis is carried out are denoted as a type of data capture nodes.

[0035] Specifically, for a single type of data capture node, If the integrated analysis index of this type of data capture node is greater than the preset integrated analysis index, the integrated data cluster of this type of data capture node is determined according to the reference integration index and the reference flow correlation index; If the integrated analysis index of this type of data capture node is less than or equal to the preset integrated analysis index, the integrated data cluster of this type of data capture node is determined according to the node conflict parameter and the reference difference stability parameter.

[0036] Among them, because the node mutation coefficient or the node complexity coefficient of a type of data capture node is large, it indicates that the supervision data of this type of data capture node is not continuously stable, and most data have relatively frequent fluctuations or there is a greater risk of interference in its corresponding sub-service area. By integrating various data that may have spatio-temporal correlations for this type of data capture node and ensuring that the correlation analysis process between data is more adaptable to the actual scenario, the data integration effect is more sufficient, and the quality of the correlation situation between different integrated analysis data in the obtained integrated data cluster is ensured; For a single type of data capture node, the integrated analysis index = ln(node mutation coincidence degree × connected mutation coefficient), the node mutation richness = the number of categories of coincident mutation supervision data determined by this type of data capture node during the node evaluation stage / the number of categories of supervision data at each mutation moment determined by this type of data capture node during the node evaluation stage. If a supervision data has mutation moments determined for this type of data capture node and any of its connected nodes during the node evaluation stage, then this supervision data is denoted as coincident mutation supervision data. The connected mutation coefficient is the average value of the node mutation coefficients of each connected node of this type of data capture node; The value of the preset integrated analysis index can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision processing record. The higher the user's requirement for the fusion quality of multi-source heterogeneous data obtained during the supervision process, the larger the value of the preset integrated analysis index. A method for determining the value of the preset integrated analysis index is provided. The supervision processing record for determining the integrated data cluster according to the reference integration index and the reference flow correlation index is denoted as a type of integrated reference record. The minimum value of the integrated analysis index of the type of data capture node in a type of integrated reference record that meets the user's requirement for the fusion quality of multi-source heterogeneous data obtained during the supervision process is denoted as the preset integrated analysis index; For a single type-1 data capture node, partition the various integrated analysis data of this type-1 data capture node and its connected nodes to obtain an integrated data cluster. The integration quality coefficient of any integrated data cluster determined by this type-1 data capture node is greater than the preset integration quality coefficient. If the integrated analysis index is greater than the preset integrated analysis index, the integration quality coefficient is determined according to the reference integration index and the reference flow correlation index. For a single integrated data cluster, the integration quality coefficient is the product of the reference integration index and the reference flow correlation index. Detect the number of times that each group of integrated analysis data within this integrated data cluster exists in the same integrated data cluster during the node evaluation stage, and record it as the integration index between the above-mentioned group of integrated analysis data. Each group of integrated analysis data only contains two pieces of integrated analysis data, and the integrated analysis data contained in each group is not completely the same. The reference integration index is the average value of the integration indexes between each group of integrated analysis data within this integrated data cluster. Detect the flow correlation index of each flow analysis combination within this integrated data cluster. Any flow analysis combination only contains two pieces of integrated analysis data belonging to this type-1 data capture node and its connected nodes respectively. For a single flow analysis combination, the flow correlation index = the absolute value of the difference between the mutation interval duration and the connection path index / the mutation interval duration. The mutation interval duration is the interval duration between the nearest mutation moments corresponding to the two pieces of integrated analysis within this flow analysis combination. The reference flow correlation index is the average value of the flow correlation indexes of each flow analysis combination within this integrated data cluster. If the integrated analysis index is less than or equal to the preset integrated analysis index, the integration quality coefficient is determined according to the node conflict parameter and the reference difference stability parameter. For a single integrated data cluster, the integration quality coefficient = ln(reference difference stability parameter / node conflict parameter). The node conflict parameter is the minimum value of the active connection parameter between the data capture node corresponding to the integrated analysis data within this integrated data cluster and this type-1 data capture node. The reference difference stability parameter is the average value of the difference stability parameters between the data capture nodes corresponding to each integrated analysis data within this integrated data cluster and this type-1 data capture node. For any two data capture nodes, the difference stability parameter = the minimum value of the interval time between the respective mutation moments corresponding to the above two data capture nodes / the maximum value of the interval time between the respective mutation moments corresponding to the above two data capture nodes; The value of the preset integration quality coefficient can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision processing records. The higher the user's requirement for the integration quality of the multi-source heterogeneous data obtained during the supervision process, the larger the value of the preset integration quality coefficient. A method for obtaining the value of the preset integration quality coefficient is provided. The supervision processing records for the integrated correlation analysis of the various integrated analysis data of the data capture nodes are recorded as integration reference records. The average value of the integration quality coefficients of the integrated data clusters in the integration reference records that meet the user's requirements for the integration quality of the multi-source heterogeneous data obtained during the supervision process is recorded as the preset integration quality coefficient.

[0037] Specifically, under the co-occurrence analysis condition, the integrated analysis data of each type-two data capture node are divided according to the mutation correlation index to determine the integrated data clusters of each type-two data capture node. The co-occurrence analysis condition is that there is integrated analysis data of a data capture node that needs to be subjected to co-occurrence correlation analysis. The data capture node for which co-occurrence correlation analysis is performed is recorded as a type-two data capture node.

[0038] Among them, for a type-two data capture node, since its node mutation coefficient and node complexity coefficient are both small, it indicates that the supervision data of this type of data capture node is relatively stable and its connected nodes are also relatively stable. Therefore, only the integrated analysis data of itself is divided, which not only ensures the correlation quality between the data in the divided integrated data clusters, but also ensures the effectiveness of the correlation analysis process. For a single type-two data capture node, the correlation quality coefficient of any determined integrated data cluster is greater than the preset correlation quality coefficient. For a single integrated data cluster, the correlation quality coefficient is the average value of the mutation correlation indices of each correlation analysis combination. Each correlation analysis combination only includes two integrated analysis data that exist at the mutation moment within the current node evaluation period. For a single correlation analysis combination, the mutation correlation index = the duration of the node evaluation period / the interval duration between the mutation moments of the two integrated analysis data within the current node evaluation period. The value of the preset correlation quality coefficient can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision processing records. The higher the user's requirement for the integration quality of the multi-source heterogeneous data obtained during the supervision process, the larger the value of the preset correlation quality coefficient. A method for obtaining the value of the preset correlation quality coefficient is provided. The supervision processing records for the co-occurrence correlation analysis of the various integrated analysis data of the data capture nodes are recorded as co-occurrence reference records. The average value of the correlation quality coefficients of the integrated data clusters in the co-occurrence reference records that meet the user's requirements for the integration quality of the multi-source heterogeneous data obtained during the supervision process is recorded as the preset correlation quality coefficient.

[0039] Specifically, under the condition that the association analysis is completed, data optimization is performed for each integrated data cluster, and the data optimization method for each integrated data cluster is determined according to the association analysis strategy of each data capture node. For a single integrated data cluster, if the integrated data cluster is determined through co-occurrence association analysis, the preferred video segments are determined according to each optimization analysis combination of the integrated data cluster. if the integrated data cluster is determined through integrated association analysis, the determination method of the key video frames is determined according to the proportion of the integrated nodes in the integrated data cluster. The condition for the completion of the association analysis is that there is a data capture node in the target service area that has completed the determination of the integrated data cluster.

[0040] Among them, for a single integrated data cluster determined through co-occurrence association analysis, the optimization analysis combination is the association analysis combination with a mutation association index greater than the preset mutation association index. The first mutation moment in the determined association analysis combination is recorded as the start moment of the preferred video segment, and the last mutation moment in the determined optimization analysis combination is recorded as the end moment of the preferred video segment. The value of the preset mutation association index can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision and processing records. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process, the larger the value of the preset mutation association index. A method for determining the value of the preset mutation association index is provided. The minimum value of the mutation association index of each optimization analysis combination in the supervision and processing records that meet the user's requirement for the fusion quality of the multi-source heterogeneous data obtained during the supervision process is recorded as the preset mutation association index.

[0041] Specifically, for a single integrated data cluster determined through integrated association analysis, if the proportion of the integrated nodes is greater than the preset proportion of the integrated nodes, the key video frames of the integrated data cluster are determined according to the associated connected nodes. if the proportion of the integrated nodes is less than or equal to the preset proportion of the integrated nodes, the key video frames of the integrated data cluster are determined according to the mutation dispersion index.

[0042] Among them, for a single integrated data cluster determined by integrated association analysis, the proportion of integration nodes = the number of a certain type of data capture nodes among the connected nodes of a certain type of data capture nodes corresponding to this integrated data cluster / the number of connected nodes of a certain type of data capture nodes corresponding to this integrated data cluster. The value of the preset proportion of integration nodes can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision processing records, providing a method for determining the value of the preset proportion of integration nodes. The supervision processing records of the key video frames of this integrated data cluster determined according to the associated connected nodes are recorded as frame extraction reference records. The minimum value of the proportion of integration nodes in the frame extraction reference records that meet the user's requirements for the fusion quality of multi-source heterogeneous data obtained during the supervision process is recorded as the preset proportion of integration nodes; For the integrated data cluster of a single certain type of data capture node, if the proportion of integration nodes is greater than the preset proportion of integration nodes, the connected nodes of this certain type of data capture node that all have mutation moments within the current node evaluation period are recorded as associated connected nodes. The mutation moments of each associated connected node and this certain type of data capture node are recorded as frame extraction reference moments. For any two frame extraction reference moments, if there are no other frame extraction reference moments between the above two frame extraction reference moments, then the above two frame extraction reference moments are recorded as a group of adjacent frame extraction reference moments. The interval duration between the frame extraction reference moments within each group of adjacent frame extraction reference moments is detected and recorded as the frame extraction reference duration. The key frame extraction duration is determined according to the average value of the frame extraction reference durations of each group of adjacent frame extraction reference moments. The key frame extraction duration has a positive correlation with the average value of the obtained frame extraction reference durations. The key video frames are extracted according to the key frame extraction duration. The key frame extraction duration is the interval duration between adjacent key video frames. If the proportion of integration nodes is less than or equal to the preset proportion of integration nodes, the key frame extraction duration has a positive correlation with the mutation dispersion index. For any two mutation moments, if there are no other mutation moments between the above two mutation moments, then the above two mutation moments are recorded as a group of adjacent mutation moments. The interval duration between the mutation moments within each group of mutation moments is detected and recorded as the dispersion duration. The average value of the dispersion durations of each group of adjacent mutation moments determined is recorded as the mutation dispersion index.

[0043] Please refer to Figure 4 As shown in the figure, it is the module connection diagram of the intelligent preprocessing system based on multi-source heterogeneous data fusion of the present invention. The present invention also provides an intelligent preprocessing system based on multi-source heterogeneous data fusion, including: A node evaluation module, used to determine the association analysis strategy of each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient, which is to perform integrated association analysis or co-occurrence association analysis on each fusion analysis data of the data capture node; An integration analysis module, which is connected to the node evaluation module and is used to determine the integration strategy of a type of data capture node according to the integration analysis index as an integrated data cluster determined according to the reference integration index and the reference flow association index, or, the node conflict parameter and the reference difference stability parameter; A co-occurrence analysis module, which is connected to the node evaluation module and is used to determine the integrated data cluster of each type-two data capture node according to the mutation association index; An optimization execution module, which is respectively connected to the integration analysis module and the co-occurrence analysis module, and is used to determine the data optimization method of the corresponding integrated data clusters according to the association analysis strategy of each data capture node as determining a preferred video segment according to each optimization analysis combination of the integrated data cluster, or, determining a key video frame according to the proportion of integrated nodes of the integrated data cluster.

[0044] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

[0045] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent preprocessing method based on multi-source heterogeneous data fusion, characterized in that: include: The correlation analysis strategy of each data capture node in the target service area is determined according to the node mutation coefficient and the node complexity coefficient, which is to perform integrated correlation analysis or co-occurrence correlation analysis on each fusion analysis data of the data capture node; During the integration association analysis, the integration strategy of a type of data capture node is determined according to the integration analysis index as an integrated data cluster determined according to the reference integration index and the reference flow association index, or the node conflict parameter and the reference difference stability parameter; During the co-occurrence association analysis, the integrated data cluster of each second-class data capture node is determined according to the mutation association index; When the correlation analysis is completed, the data optimization method for each corresponding integrated data cluster is determined according to the correlation analysis strategy of each data capture node, that is, the preferred video segment is determined according to each optimization analysis combination of the integrated data cluster, or the key video frame is determined according to the proportion of integrated nodes of the integrated data cluster.

2. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 1 is characterized in that: Periodically test the node mutation coefficient and node complexity coefficient of each data capture node in the target service area to determine the correlation analysis strategy of each data capture node; For a single data capture node, the node mutation coefficient is determined according to the supervision mutation index and the mutation persistence index of the data capture node, and the node complexity coefficient is determined according to the node connectivity index and the reference connectivity mutation index of the data capture node.

3. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 1 is characterized in that: For any data capture node, if the node mutation coefficient of the data capture node is greater than the preset node mutation coefficient or the node complexity coefficient is greater than the preset node complexity coefficient, then integrated correlation analysis is performed on the various fusion analysis data of the data capture node.

4. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 1 is characterized in that: For any data capture node, if the node mutation coefficient of the data capture node is less than or equal to the preset node mutation coefficient and the node complexity coefficient is less than or equal to the preset node complexity coefficient, co-occurrence association analysis is performed on each fusion analysis data of the data capture node.

5. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 3 is characterized in that: Under the integrated analysis conditions, the integrated analysis index of each type of data capture node is determined according to the node mutation richness and the connectivity mutation coefficient, and the integration strategy of each type of data capture node is determined according to the integrated analysis index, so as to divide the fusion analysis data of each type of data capture node and its connectivity nodes; The integration analysis index is positively correlated with the node mutation richness and the connectivity mutation coefficient respectively; The integration analysis condition is that the fusion analysis data of the data capture node needs to be integrated and associated; The data capture nodes that perform integrated association analysis are recorded as a type of data capture nodes.

6. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 5 is characterized in that: For a single Class I data capture node, If the integration analysis index of the data capture node of the type is greater than the preset integration analysis index, the integrated data cluster of the data capture node of the type is determined according to the reference integration index and the reference flow correlation index; If the integration analysis index of the data capture node of this type is less than or equal to the preset integration analysis index, the integrated data cluster of the data capture node of this type is determined according to the node conflict parameter and the reference difference stability parameter.

7. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 4 is characterized in that: Under the co-occurrence analysis condition, the fusion analysis data of each type II data capture node is divided according to the mutation association index to determine the integrated data cluster of each type II data capture node; The co-occurrence analysis condition is that the fusion analysis data of the data capture node needs to be subjected to co-occurrence association analysis; The data capture nodes for co-occurrence association analysis are recorded as the second type of data capture nodes.

8. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 7 is characterized in that: When the correlation analysis is completed, data optimization is performed for each integrated data cluster, and the corresponding data optimization method for each integrated data cluster is determined according to the correlation analysis strategy of each data capture node; For a single integrated data set, If the integrated data cluster is determined through co-occurrence association analysis, the preferred video segment is determined according to each optimized analysis combination of the integrated data cluster; If the integrated data cluster is determined through integrated association analysis, a method for determining the key video frame is determined according to the proportion of integrated nodes of the integrated data cluster; The completion condition of the association analysis is that there is a data capture node in the target service area to complete the determination of the integrated data cluster.

9. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 8 is characterized in that: For a single integrated data set identified through integrated association analysis, If the integration node ratio is greater than the preset integration node ratio, determining the key video frame of the integrated data cluster according to the associated connected nodes; If the integration node ratio is less than or equal to the preset integration node ratio, the key video frame of the integrated data cluster is determined according to the mutation dispersion index.

10. An intelligent preprocessing system based on multi-source heterogeneous data fusion using the intelligent preprocessing method according to any one of claims 1 to 9, characterized in that: include: A node evaluation module is used to determine the association analysis strategy of each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient, and to perform integrated association analysis or co-occurrence association analysis on various fusion analysis data of the data capture node; An integration analysis module, which is connected to the node evaluation module, and is used to determine, according to the integration analysis index, an integration strategy of a type of data capture node as an integrated data cluster determined according to a reference integration index and a reference flow association index, or a node conflict parameter and a reference difference stability parameter; A co-occurrence analysis module, which is connected to the node evaluation module and is used to determine the integrated data cluster of each second-class data capture node according to the mutation association index; An optimization execution module is respectively connected to the integration analysis module and the co-occurrence analysis module, and is used to determine the data optimization method of each corresponding integrated data cluster according to the association analysis strategy of each data capture node, that is, to determine the preferred video segment according to each optimization analysis combination of the integrated data cluster, or to determine the key video frame according to the proportion of integrated nodes of the integrated data cluster.

Citation Information

Patent Citations

  • Multi-source heterogeneous data fusion method applied to smart city

    CN118839292A

  • Data management system of simulation process

    CN118278188A

  • Multi-source heterogeneous data fusion processing and intelligent analysis method, device and equipment

    CN119557845A

  • Debris flow monitoring method based on multimode data communication

    CN119783004A

  • Grain informatization data security management system

    CN119808179A