Intelligent preprocessing method and system based on multi-source heterogeneous data fusion
By determining the correlation analysis strategy of data capture nodes based on node characteristics within the regulatory area, and optimizing the data through integration and co-occurrence analysis, the problem of low fusion effect of multi-source heterogeneous data was solved, and the data integration quality and analysis efficiency were improved.
Patent Information
- Application Number
- CN202510629789.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Existing technologies have failed to determine targeted correlation analysis methods based on the actual situation of the regulated area, resulting in poor fusion effect of multi-source heterogeneous data, poor data integration quality, and affecting the efficiency and effectiveness of subsequent decision analysis.
By determining the correlation analysis strategy for each data capture node within the target service area based on the node mutation coefficient and node complexity coefficient, and by employing integrated correlation analysis and co-occurrence correlation analysis, combined with the integrated analysis index and mutation correlation index, the data optimization method for data capture nodes is optimized to identify key video frames and video segments.
It improves the fusion effect of multi-source heterogeneous data, enhances the transmission and analysis efficiency of regulatory data, ensures that the data integration results meet actual needs, and supports efficient decision analysis.
Smart Images

Figure CN120145320B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, and in particular to an intelligent preprocessing method and system based on multi-source heterogeneous data fusion. BACKGROUND
[0002] When monitoring public areas, the multi-source heterogeneous data obtained is fused and processed, which can effectively improve the reliability and timeliness of the subsequent decision analysis process. Through effective association and content preprocessing of multi-source heterogeneous data, the fusion result of multi-source heterogeneous data can be ensured. In actual monitoring processes, the personnel flow and activity rules of different monitoring areas are different. Using a single data association analysis method can easily lead to the fact that the data integration quality cannot meet the actual needs, and thus the efficiency and effectiveness of subsequent decision analysis based on data fusion results are poor. Therefore, how to determine a targeted data association analysis method for different monitoring areas in actual monitoring processes to avoid low data integration quality is a problem that needs to be solved by those skilled in the art.
[0003] Chinese Patent Application Publication No. CN118839292A discloses a multi-source heterogeneous data fusion method applied to smart cities, including the following steps: S1, data collection and preprocessing; S2, data integration and matching, integrating and matching data from different data sources, establishing a data model and semantic association; S3, data fusion and analysis, using data fusion technology to combine and integrate the integrated data to produce a more comprehensive and accurate city state description; S4, modeling and prediction. However, the above-mentioned scheme has the following problems: it fails to determine a targeted association analysis method according to the actual situation of the monitoring area, to integrate and match different data, resulting in poor quality of data integration results, and thus low fusion effect of multi-source heterogeneous data. SUMMARY
[0004] Therefore, the present application provides an intelligent preprocessing method and system based on multi-source heterogeneous data fusion to overcome the problem in the prior art that a targeted association analysis method cannot be determined according to the actual situation of the monitoring area, to integrate and match different data, resulting in poor quality of data integration results, and thus low fusion effect of multi-source heterogeneous data.
[0005] To achieve the above-mentioned purpose, the present application provides an intelligent preprocessing method based on multi-source heterogeneous data fusion, comprising:
[0006] According to the node mutation coefficient and the node complexity coefficient, the association analysis strategy of each data capture node in the target service area is determined as integration association analysis or co-occurrence association analysis of each fusion analysis data of the data capture node.
[0007] In the integration correlation analysis, the integration strategy of a data capture node in the first category is determined according to the integration analysis index, which is determined according to the reference integration index and the reference flow correlation index, or the integration data cluster determined according to the node conflict parameter and the reference difference stability parameter;
[0008] In the co-occurrence correlation analysis, the integration data cluster of each data capture node in the second category is determined according to the mutation correlation index;
[0009] Under the condition of completing the correlation analysis, the data optimization mode of each integration data cluster corresponding to each data capture node is determined according to the correlation analysis strategy of the data capture node, which is determined according to the preferred video segment of each optimization analysis combination of the integration data cluster, or the determination mode of the key video frame according to the integration node proportion of the integration data cluster.
[0010] Further, the node mutation coefficient and the node complexity coefficient of each data capture node in the target service area are periodically detected to determine the correlation analysis strategy of each data capture node;
[0011] For a single data capture node, the node mutation coefficient is determined according to the supervision mutation index and the mutation duration index of the data capture node, and the node complexity coefficient is determined according to the node connectivity index and the reference connectivity mutation index of the data capture node.
[0012] Further, for any data capture node, if the node mutation coefficient of the data capture node is greater than the preset node mutation coefficient or the node complexity coefficient is greater than the preset node complexity coefficient, the integration correlation analysis is performed on the fusion analysis data of the data capture node.
[0013] Further, for any data capture node, if the node mutation coefficient of the data capture node is less than or equal to the preset node mutation coefficient and the node complexity coefficient is less than or equal to the preset node complexity coefficient, the co-occurrence correlation analysis is performed on the fusion analysis data of the data capture node.
[0014] Further, in the integration analysis condition, the integration analysis index of each data capture node in the first category is determined according to the node mutation richness and the connectivity mutation coefficient, and the integration strategy of each data capture node in the first category is determined according to the integration analysis index, so as to divide the fusion analysis data of each data capture node in the first category and its connected nodes;
[0015] The integration analysis index is positively correlated with the node mutation richness and the connectivity mutation coefficient, respectively;
[0016] The integration analysis condition is that the fusion analysis data of the data capture node needs to be integrated and correlated, and the data capture node to be integrated and correlated is recorded as a data capture node in the first category.
[0017] Further, for a single one-class data capture node,
[0018] If the integration analysis index of the one-class data capture node is greater than the preset integration analysis index, the integrated data cluster of the one-class data capture node is determined according to the reference integration index and the reference flow correlation index;
[0019] If the integration analysis index of the one-class data capture node is less than or equal to the preset integration analysis index, the integrated data cluster of the one-class data capture node is determined according to the node conflict parameter and the reference difference stability parameter.
[0020] Further, under the co-occurrence analysis condition, the fusion analysis data of each two-class data capture node is divided according to the mutation correlation index to determine the integrated data cluster of each two-class data capture node;
[0021] The co-occurrence analysis condition is that the fusion analysis data of the data capture node needs to be analyzed by co-occurrence correlation analysis, and the data capture node to be analyzed by co-occurrence correlation analysis is recorded as a two-class data capture node.
[0022] Further, under the correlation analysis completion condition, data optimization is performed on each integrated data cluster, and the data optimization mode of each integrated data cluster is determined according to the correlation analysis strategy of each data capture node;
[0023] For a single integrated data cluster,
[0024] If the integrated data cluster is determined by co-occurrence correlation analysis, the preferred video segment is determined according to each optimization analysis combination of the integrated data cluster;
[0025] If the integrated data cluster is determined by integration correlation analysis, the determination mode of the key video frame is determined according to the integration node proportion of the integrated data cluster;
[0026] The correlation analysis completion condition is that the data capture node in the target service area completes the determination of the integrated data cluster.
[0027] Further, for a single integrated data cluster determined by integration correlation analysis,
[0028] If the integration node proportion is greater than the preset integration node proportion, the key video frame of the integrated data cluster is determined according to the correlation connected node;
[0029] If the integration node proportion is less than or equal to the preset integration node proportion, the key video frame of the integrated data cluster is determined according to the mutation dispersion index.
[0030] The application also provides an intelligent preprocessing system applying the intelligent preprocessing method based on multi-source heterogeneous data fusion, comprising:
[0031] a node evaluation module configured to determine, according to the node mutation coefficient and the node complexity coefficient, the correlation analysis strategy of each data capture node in the target service area as integrated correlation analysis or co-occurrence correlation analysis of each fusion analysis data of the data capture node;
[0032] an integrated analysis module connected with the node evaluation module and configured to determine, according to the integrated analysis index, the integrated strategy of a type of data capture node as the integrated data cluster determined according to the reference integration index and the reference flow correlation index, or the node conflict parameter and the reference difference stability parameter;
[0033] a co-occurrence analysis module connected with the node evaluation module and configured to determine, according to the mutation correlation index, the integrated data cluster of each type of data capture node;
[0034] an optimization execution module connected with the integrated analysis module and the co-occurrence analysis module respectively and configured to determine, according to the correlation analysis strategy of each data capture node, the data optimization mode of the corresponding integrated data cluster as the preferred video paragraph determined according to each optimization analysis combination of the integrated data cluster, or the determination mode of the key video frame determined according to the integrated node proportion of the integrated data cluster.
[0035] Compared with the prior art, the beneficial effects of the present application lie in that the present application determines the targeted correlation analysis strategy of each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient, so that the division process of the fusion analysis data conforms to the actual situation of the data obtained in the supervision process of each data capture node, and in addition, the data in the obtained integrated data cluster is optimized in a targeted manner, which not only improves the fusion effect of the obtained multi-source heterogeneous data, but also improves the transmission and analysis efficiency of the supervision data.
[0036] Further, in the present application, the node mutation coefficient and the node complexity coefficient are used to determine how to perform correlation analysis on different data capture nodes, wherein the node mutation coefficient and the node complexity coefficient can represent the frequency of the supervision data of the data capture node in a period of time and the interference probability of the surrounding area, so as to set a targeted method for determining the integrated data cluster, which not only ensures that the integrated data cluster obtained by division can sufficiently contain the fusion analysis data with correlation, but also ensures the division efficiency of the integrated data cluster.
[0037] Further, in the present application, for a type of data capture node, the integrated strategy is determined according to the integrated analysis index, and the supervision data of the sub-service area supervised by this type of data capture node has a large mutation probability or is greatly affected by the fluctuation of the surrounding area, so that the integrated processing of the fusion analysis data needs to consider the spatio-temporal correlation of the supervision data obtained by other data capture nodes.
[0038] Further, the present application further determines the existence data mutation of a class of data capture nodes and their connected regions by the integrated analysis index. When the integrated analysis index is larger, it indicates that the connected nodes have larger data mutation in the whole and the corresponding supervision data also has larger coincidence when the mutation occurs. By referring to the integrated index and the reference flow correlation index, the correlation between the integrated data clusters and the correlation between the supervision data mutation and the personnel flow are ensured. When the integrated analysis index is smaller, it indicates that there are fewer supervision data mutations in the connected region. At this time, the node conflict parameter and the reference difference stability parameter are used to ensure that the data corresponding to the sub-service area has a closer range and the interval time between the data mutation time is more stable, that is, to ensure that the fused data can effectively represent the overall state of the surrounding data, and to ensure the effectiveness of the data fusion result for the decision analysis process.
[0039] Further, the present application determines the data optimization mode of each integrated data cluster according to the correlation analysis strategy of each data capture node. The data capture node determined by the co-occurrence correlation analysis has a dominant factor of supervision data change, which tends to be the change of personnel activity in the corresponding sub-service area. The preferred video paragraph interception time is determined by some speech analysis combination, which effectively ensures that the obtained video information can assist in analyzing the mutation reason of the data. At the same time, the effectiveness of the fusion result of multi-source heterogeneous data is ensured, and the analysis and transmission efficiency of the fusion result is improved.
[0040] Further, the data capture node determined by the integrated correlation analysis has a dominant factor of supervision data change, which mainly tends to the personnel flow between different regions. The integrated node proportion is used to determine the setting mode of the interval time between frames in the key video frame extraction process, which ensures that the extraction of key video frames is more suitable for the change of data in the actual integrated data cluster. The present application ensures the effectiveness of the fusion result of multi-source heterogeneous data, and improves the analysis and transmission efficiency of the fusion result. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 The schematic diagram of the intelligent preprocessing method based on multi-source heterogeneous data fusion of the present application;
[0042] Figure 2 The flow chart of determining the correlation analysis strategy according to the node mutation coefficient and the node complexity coefficient of the present application;
[0043] Figure 3 The flow chart of determining the integration strategy of a class of data capture nodes according to the integrated analysis index of the present application;
[0044] Figure 4 The module connection diagram of the intelligent preprocessing system based on multi-source heterogeneous data fusion of the application. DETAILED DESCRIPTION
[0045] In order to make the objects and advantages of the present application clearer, the present application will be further described below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0046] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. It should be understood by those skilled in the art that the embodiments are only used to explain the technical principles of the present application and are not used to limit the protection scope of the present application.
[0047] It should be noted that, in the description of the present application, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer" and the like are based on the direction or positional relationship shown in the drawings, which is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application.
[0048] In addition, it should also be noted that, in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connection" and "connection" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0049] Please refer to Figures 1 to 3 The present application provides a multi-source heterogeneous data fusion intelligent preprocessing method, which comprises the following steps:
[0050] According to the node mutation coefficient and the node complexity coefficient, the correlation analysis strategy of each data capture node in the target service area is determined as the integrated correlation analysis or co-occurrence correlation analysis of each fusion analysis data of the data capture node;
[0051] When the integrated correlation analysis is performed, the integrated strategy of a type of data capture node is determined according to the integrated analysis index, which is an integrated data cluster determined according to the correlation quality coefficient and the integrated conflict coefficient, or an integrated data cluster of a type of data capture node determined according to the integrated conflict coefficient;
[0052] When the co-occurrence correlation analysis is performed, the integrated data cluster of each type of data capture node is determined according to the mutation correlation index and the correlation stability index;
[0053] Under the condition of the correlation analysis completion, the data optimization mode is determined according to the integrated correlation proportion of the integrated data cluster, so as to perform optimization processing, the data optimization mode is to determine the priority optimization coefficient of each video paragraph according to the correlation activity connectivity parameter and the stage matching index, or to determine the priority optimization coefficient of each video paragraph according to the correlation richness index and the stage matching index.
[0054] In the process of monitoring in the public area, the data fusion processing is performed on the obtained monitoring data, and the area to be monitored is recorded as a target service area, there are several data capture nodes in the application, any data capture node corresponds to only a sub-service area, and there are several monitoring devices for obtaining various monitoring data, each sub-service area is divided from the target service area, and the user can adaptively divide according to the actual working scene, the application does not specifically limit the monitoring devices set by each data capture node and the type of monitoring data obtained, and the user can adaptively set according to the actual working scene, but each data capture node is provided with a video acquisition device, and the specific setting position and device type are not limited, but the video data obtained by each data capture node can cover the corresponding sub-service area, the type of monitoring data in the application includes but is not limited to the temperature, carbon dioxide concentration, noise decibel and video monitoring data of the sub-service area corresponding to each data capture node.
[0055] In the application, there are several supervision processing records, any supervision processing record records the reference mutation degree, activity connectivity parameter, node mutation coefficient, node complexity coefficient, integrated analysis index, integrated quality coefficient, correlation quality coefficient, mutation correlation index and integrated node proportion in the process of data fusion processing of the obtained monitoring data of the target service area, and each supervision processing record corresponds to a qualified mark, the qualified mark records whether the fusion quality of the multi-source heterogeneous data obtained in the supervision process meets the user's demand, it can be understood that the user can determine whether the fusion quality of the multi-source heterogeneous data obtained in the supervision process meets the demand according to the self-set index, for example, the self-set index can be but is not limited to the decision response time, the decision response time is the time length used for outputting a decision result based on the fusion result of the obtained multi-source heterogeneous data.
[0056] Specifically, the node mutation coefficient and the node complexity coefficient of each data capture node in the target service area are periodically detected to determine the correlation analysis strategy of each data capture node.
[0057] For a single data capture node,
[0058] The node mutation coefficient is determined according to the supervision mutation index and the mutation duration index of the data capture node, and the node complexity coefficient is determined according to the node connectivity index and the reference connectivity mutation index of the data capture node.
[0059] In the application, a circulating node evaluation period is applied, the length of the node evaluation period can be determined by the user, the higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process, the shorter the length of the node evaluation period, and a length of the node evaluation period is provided, the length of the node evaluation period is 1 hour, at the end of each node evaluation period, the node mutation coefficient and the node complexity coefficient of each data capture node are detected to determine the correlation analysis strategy of each data capture node, and a circulating data supervision period is also applied in the application, the length of the data supervision period can be determined by the user, a length of the data supervision period is provided, the length of the data supervision period is 10 minutes, at the end of each data supervision period, the data fusion processing data in the current data supervision period are analyzed to output a decision analysis result, and how to output the decision analysis result based on the data fusion result of each supervision data is easy to understand for those skilled in the art, and will not be described here.
[0060] For a single data capture node, the node mutation coefficient is the product of the supervision mutation index and the mutation duration index of the data capture node, the supervision mutation index is the number of mutation moments existing in the node evaluation stage, and the mutation duration index is the maximum value of the mutation duration of each mutation moment in the node evaluation stage, the node complexity coefficient is the product of the node connectivity index and the reference connectivity mutation index of the data capture node, the node connectivity index is the number of connected nodes of the data capture node, and the reference connectivity mutation index is the average value of the supervision mutation indexes of each connected node of the data capture node, the end time of the node evaluation stage is the end time of the current node evaluation period, the length of the node evaluation stage can be set by the user according to the actual working scene, the higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process, the longer the length of the node evaluation stage, and a length of the node evaluation stage is provided, the length of the node evaluation stage is 6 times the length of the node evaluation period.
[0061] If the reference mutation degree at the end time of a single data supervision period is greater than a preset reference mutation degree, the time is recorded as a mutation time, the reference mutation degree is the average of the mutation degrees of each item of supervision data obtained at the time, for a single item of supervision data, the mutation degree corresponding to the time = | the maximum value of the item of supervision data obtained in the current data supervision period - the maximum value of the item of supervision data obtained in the last data supervision period of the current data supervision period | / the maximum value of the item of supervision data obtained in the last data supervision period of the current data supervision period, for a single mutation time, the maximum value of the mutation degrees of each item of supervision data at the mutation time is recorded as an evaluation mutation degree, the value of the supervision data corresponding to the evaluation mutation degree obtained in the previous data supervision period in the mutation time determination process is recorded as a continuous reference value, the data supervision period in which the maximum value of the supervision data corresponding to the evaluation mutation degree after the mutation time is less than the continuous reference value is recorded as a recovery time, the mutation duration is the minimum value of the interval duration between the mutation time and each recovery time thereof, and the unit of the mutation duration is hour;
[0062] The value of the preset reference mutation degree can be determined by the user according to the actual working scene, for example, the user can set it according to the supervision processing record, the higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process, the smaller the value of the preset reference mutation degree, and a method for determining the value of the preset reference mutation degree is provided, the average of the reference mutation degrees of each mutation time in the supervision processing record meeting the requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is recorded as the value of the preset reference mutation degree;
[0063] If the active connectivity parameter between any two data capture nodes is greater than a preset active connectivity parameter, the two data capture nodes are determined to be connected nodes, the active connectivity parameter between the two data capture nodes is determined according to the personnel flow correlation degree and the connectivity path index, the active connectivity parameter is the product of the personnel flow correlation degree and the connectivity path index, the connectivity path index is the shortest duration required from one data capture node to another data capture node, the unit of the connectivity path index is minute, and the personnel flow correlation degree = the number of personnel existing in the sub-service areas corresponding to the two data capture nodes at the same time in the node evaluation stage / the personnel flow parameter of the two data capture nodes in the node evaluation stage, for a single data capture node, the personnel flow parameter is the maximum value of the number of personnel existing in the data capture node obtained in the node evaluation stage;
[0064] The value of the preset activity connectivity parameter can be determined by the user according to the actual working scenario. For example, the user can set it according to the supervision processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is, the smaller the value of the preset activity connectivity parameter is. A method for determining the value of the preset activity connectivity parameter is provided. The average value of the activity connectivity parameter between the internal connectivity nodes in the supervision processing record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is recorded as the preset activity connectivity parameter.
[0065] Specifically, for any data capture node, if the node mutation coefficient of the data capture node is greater than the preset node mutation coefficient or the node complexity coefficient is greater than the preset node complexity coefficient, the integrated correlation analysis is performed on the fusion analysis data of the data capture node.
[0066] Specifically, for any data capture node, if the node mutation coefficient of the data capture node is less than or equal to the preset node mutation coefficient and the node complexity coefficient is less than or equal to the preset node complexity coefficient, the co-occurrence correlation analysis is performed on the fusion analysis data of the data capture node.
[0067] In the method, for a single data capture node, the supervision data obtained in the current node evaluation period is recorded as the fusion analysis data of the data capture node that needs to be fused at present. The correlation analysis strategy is determined according to the node mutation coefficient and the node complexity coefficient. The values of the preset node mutation coefficient and the preset node complexity coefficient can be determined by the user according to the actual working scenario. For example, the user can set them according to the supervision processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is, the smaller the value of the preset node mutation coefficient is, and the smaller the value of the preset node complexity coefficient is. A method for determining the value of the preset node mutation coefficient is provided. The supervision processing record in which the integrated correlation analysis is performed on the fusion analysis data of the data capture node is recorded as the correlation reference record. The average value of the node mutation coefficient of each data capture node in the correlation reference record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is recorded as the preset node mutation coefficient. A method for determining the value of the preset node complexity coefficient is provided. The average value of the node complexity coefficient of each data capture node in the correlation reference record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is recorded as the preset node complexity coefficient.
[0068] Specifically, under the integrated analysis condition, the integrated analysis index of each type of data capture node is determined according to the node mutation richness and the connectivity mutation coefficient, and the integrated strategy of each type of data capture node is determined according to the integrated analysis index, so as to divide the fusion analysis data of each type of data capture node and its connected nodes.
[0069] The integrated analysis index is positively correlated with the node mutation richness and the connected mutation coefficient, respectively.
[0070] The integrated analysis condition is that the fusion analysis data of the data capture node needs to be integrated and correlated. The data capture node that will be subjected to integrated correlation analysis is recorded as a type of data capture node.
[0071] Specifically, for a single type of data capture node,
[0072] If the integrated analysis index of the type of data capture node is greater than the preset integrated analysis index, the integrated data cluster of the type of data capture node is determined according to the reference integration index and the reference flow correlation index.
[0073] If the integrated analysis index of the type of data capture node is less than or equal to the preset integrated analysis index, the integrated data cluster of the type of data capture node is determined according to the node conflict parameter and the reference difference stability parameter.
[0074] Since the node mutation coefficient or the node complexity coefficient of the type of data capture node is large, it indicates that the supervision data of this type of data capture node is not continuously stable, and more data exist frequent fluctuations or the corresponding sub-service area exists a large risk of being interfered. By integrating various types of data that may exist spatio-temporal correlation for this type of data capture node, and ensuring that the correlation analysis process between the data is more suitable for the actual scene, the data integration effect is more sufficient, and the quality of the correlation between the different fusion analysis data in the obtained integrated data cluster is ensured.
[0075] For a single type of data capture node, the integrated analysis index = ln(node mutation coincidence degree x connected mutation coefficient), the node mutation richness = the number of coincident mutation supervision data categories determined by the type of data capture node in the node evaluation stage / the number of categories of supervision data determined by the type of data capture node at each mutation moment in the node evaluation stage. If a supervision data faces the type of data capture node and any connected node, and the supervision data is determined to have a mutation moment in the node evaluation stage, the supervision data is recorded as coincident mutation supervision data. The connected mutation coefficient is the average of the node mutation coefficients of each connected node of the type of data capture node.
[0076] The preset integration analysis index value can be determined by the user according to the actual working scene. For example, the user can set it according to the regulatory processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the regulatory process, the larger the preset integration analysis index value. A method for determining the value of the preset integration analysis index is provided. The regulatory processing record of the integrated data cluster determined according to the reference integration index and the reference flow correlation index is recorded as a type of integrated reference record. The minimum value of the integration analysis index of a type of data capture node in the type of integrated reference record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the regulatory process is recorded as the preset integration analysis index.
[0077] The integration quality coefficient of any integrated data cluster determined by the one-type data capturing node is greater than a preset integration quality coefficient. If the integration analysis index is greater than a preset integration analysis index, the integration quality coefficient is determined according to a reference integration index and a reference flow correlation index. For a single integrated data cluster, the integration quality coefficient is the product of the reference integration index and the reference flow correlation index. The number of times that each group of fusion analysis data exists in the same integrated data cluster in the node evaluation stage is detected and recorded as the integration index between the two groups of fusion analysis data. Each group of fusion analysis data only contains two fusion analysis data, and the fusion analysis data contained in each group is not completely the same. The reference integration index is the average value of the integration index between each group of fusion analysis data in the integrated data cluster. The flow correlation index of each flow analysis combination in the integrated data cluster is detected. Any flow analysis combination only contains two fusion analysis data belonging to the one-type data capturing node and its connected nodes. For a single flow analysis combination, the flow correlation index = the absolute value of the difference between the mutation interval time length and the connected path index / the mutation interval time length. The mutation interval time length is the interval time length between the nearest mutation time corresponding to the two fusion analysis data in the flow analysis combination. The reference flow correlation index is the average value of the flow correlation index of each flow analysis combination in the integrated data cluster. If the integration analysis index is less than or equal to the preset integration analysis index, the integration quality coefficient is determined according to a node conflict parameter and a reference difference stability parameter. For a single integrated data cluster, the integration quality coefficient = ln (reference difference stability parameter / node conflict parameter). The node conflict parameter is the minimum value of the active connection parameter between the data capturing node corresponding to the fusion analysis data in the integrated data cluster and the one-type data capturing node. The reference difference stability parameter is the average value of the difference stability parameter between the data capturing node corresponding to each fusion analysis data in the integrated data cluster and the one-type data capturing node. For any two data capturing nodes, the difference stability parameter = the minimum value of the interval time between the mutation time corresponding to each of the two data capturing nodes / the maximum value of the interval time between the mutation time corresponding to each of the two data capturing nodes.
[0078] The user can determine the value of the preset integration quality coefficient according to the actual working scene. For example, the user can set the value according to the supervision processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is, the greater the value of the preset integration quality coefficient is. A method for determining the value of the preset integration quality coefficient is provided. The supervision processing record of the integration correlation analysis of the fusion analysis data of each data capture node is recorded as an integration reference record. The average value of the integration quality coefficients of each integrated data cluster in the integration reference record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is recorded as the preset integration quality coefficient.
[0079] Specifically, under the co-occurrence analysis condition, the fusion analysis data of each two-type data capture node is divided according to the mutation correlation index to determine the integrated data cluster of each two-type data capture node.
[0080] The co-occurrence analysis condition is that the fusion analysis data of the existing data capture node needs to be analyzed by co-occurrence correlation. The data capture node that needs to be analyzed by co-occurrence correlation is recorded as a two-type data capture node.
[0081] Among them, for the two-type data capture node, since its node mutation coefficient and node complexity coefficient are both small, it indicates that the supervision data of this type of data capture node is relatively stable and its connected nodes are also relatively stable. Therefore, only the fusion analysis data of itself is divided, which not only ensures the correlation quality between the data in the integrated data cluster obtained by division, but also ensures the effectiveness of the correlation analysis process.
[0082] For a single two-type data capture node, the correlation quality coefficient of any integrated data cluster determined is greater than the preset correlation quality coefficient. For a single integrated data cluster, the correlation quality coefficient is the average value of the mutation correlation indices of each correlation analysis combination. Each correlation analysis combination only contains two fusion analysis data existing in the current node evaluation period at the mutation moment. For a single correlation analysis combination, the mutation correlation index = the duration of the node evaluation period / the interval duration between the mutation moments of the two fusion analysis data in the current node evaluation period in the correlation analysis combination.
[0083] The preset correlation quality coefficient value can be determined by the user according to the actual working scene. For example, the user can set it according to the supervision processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process, the larger the preset correlation quality coefficient value. A method for determining the value of the preset correlation quality coefficient is provided. The supervision processing record of the co-occurrence correlation analysis of the fusion analysis data of each data capture node is recorded as a co-occurrence reference record. The average value of the correlation quality coefficients of each integrated data cluster in the co-occurrence reference record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is recorded as the preset correlation quality coefficient.
[0084] Specifically, under the correlation analysis completion condition, data optimization is performed on each integrated data cluster, and the data optimization mode of each integrated data cluster is determined according to the correlation analysis strategy of each data capture node;
[0085] For a single integrated data cluster,
[0086] If the integrated data cluster is determined by co-occurrence correlation analysis, the preferred video paragraph is determined according to the optimization analysis combination of the integrated data cluster;
[0087] If the integrated data cluster is determined by integration correlation analysis, the determination mode of the key video frame is determined according to the integration node proportion of the integrated data cluster;
[0088] The correlation analysis completion condition is that there is a data capture node in the target service area to complete the determination of the integrated data cluster.
[0089] Among them, for a single integrated data cluster determined by co-occurrence correlation analysis, the optimization analysis combination is an association analysis combination with a mutation correlation index greater than a preset mutation correlation index. The first mutation time in the determined association analysis combination is recorded as the starting time of the preferred video paragraph, and the last mutation time in the determined optimization analysis combination is recorded as the end time of the preferred video paragraph.
[0090] The value of the preset mutation correlation index can be determined by the user according to the actual working scene. For example, the user can set it according to the supervision processing record. The higher the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process, the larger the value of the preset mutation correlation index. A method for determining the value of the preset mutation correlation index is provided. The minimum value of the mutation correlation index of each optimization analysis combination in the supervision processing record that meets the user's requirement for the fusion quality of the multi-source heterogeneous data obtained in the supervision process is recorded as the preset mutation correlation index.
[0091] Specifically, for a single integrated data cluster determined by integration correlation analysis,
[0092] If the integrated node proportion is greater than the preset integrated node proportion, the key video frames of the integrated data cluster are determined according to the associated connected nodes;
[0093] If the integrated node proportion is less than or equal to the preset integrated node proportion, the key video frames of the integrated data cluster are determined according to the mutation dispersion index.
[0094] Wherein, for a single integrated data cluster determined by the integrated association analysis, the integrated node proportion = the number of one type of data capture node in the connected nodes of the one type of data capture node corresponding to the integrated data cluster / the number of the connected nodes of the one type of data capture node corresponding to the integrated data cluster, the value of the preset integrated node proportion can be determined by the user according to the actual working scene, for example, the user can set it according to the supervision processing record, and a method for providing a value of the preset integrated node proportion is provided, the supervision processing record for determining the key video frames according to the associated connected nodes is recorded as the frame extraction reference record, and the minimum value of the integrated node proportion in the frame extraction reference record meeting the user's fusion quality requirements for the multi-source heterogeneous data obtained in the supervision process is recorded as the preset integrated node proportion;
[0095] For the integrated data cluster of a single one type of data capture node, if the integrated node proportion is greater than the preset integrated node proportion, the connected nodes of the one type of data capture node existing at the mutation moment in the current node evaluation period are recorded as the associated connected nodes, the mutation moment of each associated connected node and the one type of data capture node are recorded as the frame extraction reference moment, for any two frame extraction reference moments, if there is no other frame extraction reference moment between the two frame extraction reference moments, the two frame extraction reference moments are recorded as a group of adjacent frame extraction reference moments, the interval time length between the frame extraction reference moments in each group of adjacent frame extraction reference moments is detected and recorded as the frame extraction reference time length, the average value of the frame extraction reference time length of each group of adjacent frame extraction reference moments is determined as the key frame extraction time length, the key frame extraction time length is positively correlated with the average value of the obtained frame extraction reference time length, the key video frames are extracted according to the key frame extraction time length, the key frame extraction time length is the interval time length between adjacent key video frames, if the integrated node proportion is less than or equal to the preset integrated node proportion, the key frame extraction time length is positively correlated with the mutation dispersion index, for any two mutation moments, if there is no other mutation moment between the two mutation moments, the two mutation moments are recorded as a group of adjacent mutation moments, the interval time length between the mutation moments in each group of mutation moments is detected and recorded as the dispersion time length, and the average value of the dispersion time length of each group of adjacent mutation moments is recorded as the mutation dispersion index.
[0096] Please refer to Figure 4As shown, it is the module connection diagram of the intelligent preprocessing system based on multi-source heterogeneous data fusion, and the application also provides an intelligent preprocessing system based on multi-source heterogeneous data fusion, which comprises:
[0097] The node evaluation module is used to determine the correlation analysis strategy of each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient, and the correlation analysis strategy is integration correlation analysis or co-occurrence correlation analysis on each fusion analysis data of the data capture node.
[0098] The integration analysis module is connected with the node evaluation module, and is used to determine the integration strategy of a type of data capture node according to the integration analysis index, and the integration strategy is the integrated data cluster determined according to the reference integration index and the reference flow correlation index, or the node conflict parameter and the reference difference stability parameter.
[0099] The co-occurrence analysis module is connected with the node evaluation module, and is used to determine the integrated data cluster of each two types of data capture nodes according to the mutation correlation index.
[0100] The optimization execution module is connected with the integration analysis module and the co-occurrence analysis module respectively, and is used to determine the data optimization mode of each integrated data cluster according to the correlation analysis strategy of each data capture node, and the data optimization mode is the preferred video paragraph determined according to each optimization analysis combination of the integrated data cluster, or the determination mode of the key video frame determined according to the integration node proportion of the integrated data cluster.
[0101] So far, the technical scheme of the application has been described in combination with the preferred embodiments shown in the drawings, but those skilled in the art can easily understand that the protection scope of the application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the application, and the technical scheme after the changes or replacements will fall within the protection scope of the application.
[0102] The above description is only the preferred embodiments of the application and is not used to limit the application; for those skilled in the art, the application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application shall be included in the protection scope of the application.
Claims
1. An intelligent preprocessing method based on multi-source heterogeneous data fusion, characterized in that, The method comprises the following steps: According to the node mutation coefficient and the node complexity coefficient, the association analysis strategy of each data capture node in the target service area is determined as the integrated association analysis or co-occurrence association analysis of the fusion analysis data of the data capture node; When the integrated association analysis is performed, the integrated strategy of the data capture node of a type is determined according to the integrated analysis index, that is, the integrated data cluster determined according to the reference integration index and the reference flow association index, or the node conflict parameter and the reference difference stability parameter; When the co-occurrence association analysis is performed, the integrated data cluster of each data capture node of a type is determined according to the mutation association index; Under the completion condition of the association analysis, the data optimization mode of the corresponding integrated data cluster is determined according to the association analysis strategy of each data capture node, that is, the optimal video paragraph is determined according to each optimization analysis combination of the integrated data cluster, or the determination mode of the key video frame is determined according to the integrated node proportion of the integrated data cluster; For a single data capture node, the node mutation coefficient is determined according to the supervision mutation index and the mutation duration index of the data capture node, the node complexity coefficient is determined according to the node connectivity index and the reference connectivity mutation index of the data capture node, the supervision mutation index is the number of mutation moments existing in the node evaluation stage, the mutation duration index is the maximum value of the mutation duration of each mutation moment in the node evaluation stage, the node connectivity index is the number of connected nodes of the data capture node, and the reference connectivity mutation index is the average value of the supervision mutation index of each connected node of the data capture node; Under the integrated analysis condition, the integrated analysis index of each data capture node of a type is determined according to the node mutation richness and the connectivity mutation coefficient, and the integrated strategy of each data capture node of a type is determined according to the integrated analysis index, so as to divide the fusion analysis data of each data capture node of a type and its connected nodes; The integrated analysis index is positively correlated with the node mutation richness and the connectivity mutation coefficient; The integrated analysis condition is that the fusion analysis data of the data capture node needs to be integrated and associated; The data capture node to be integrated and associated is recorded as a data capture node of a type; Under the co-occurrence analysis condition, the fusion analysis data of each data capture node of a type is divided according to the mutation association index, so as to determine the integrated data cluster of each data capture node of a type; The co-occurrence analysis condition is that the fusion analysis data of the data capture node needs to be co-occurrence associated; The data capture node to be co-occurrence associated is recorded as a data capture node of a type. 2.The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Periodically, the node mutation coefficient and the node complexity coefficient of each data capture node in the target service area are detected to determine the association analysis strategy of each data capture node. 3.The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, For any data capture node, if the node mutation coefficient of the data capture node is greater than the preset node mutation coefficient or the node complexity coefficient is greater than the preset node complexity coefficient, the integrated association analysis is performed on the fusion analysis data of the data capture node. 4.The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, For any data capture node, if the node mutation coefficient of the data capture node is less than or equal to the preset node mutation coefficient and the node complexity coefficient is less than or equal to the preset node complexity coefficient, the co-occurrence correlation analysis is performed on the fusion analysis data of the data capture node. 5.The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 4, characterized in that, For a single one-type data capture node, If the integration analysis index of the one-type data capture node is greater than the preset integration analysis index, the integrated data cluster of the one-type data capture node is determined according to the reference integration index and the reference flow correlation index; If the integration analysis index of the one-type data capture node is less than or equal to the preset integration analysis index, the integrated data cluster of the one-type data capture node is determined according to the node conflict parameter and the reference difference stability parameter. 6.The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, Under the completion condition of the correlation analysis, the data optimization is performed on each integrated data cluster, and the data optimization mode of each integrated data cluster is determined according to the correlation analysis strategy of each data capture node; For a single integrated data cluster, If the integrated data cluster is determined by the co-occurrence correlation analysis, the preferred video segment is determined according to each optimization analysis combination of the integrated data cluster; If the integrated data cluster is determined by the integration correlation analysis, the determination mode of the key video frame is determined according to the integration node proportion of the integrated data cluster. The completion condition of the correlation analysis is that there is a data capture node in the target service area to complete the determination of the integrated data cluster.
7. The intelligent preprocessing method based on multi-source heterogeneous data fusion according to claim 6, characterized in that, For a single integrated data cluster determined by the integration correlation analysis, If the integration node proportion is greater than the preset integration node proportion, the key video frame of the integrated data cluster is determined according to the correlation connected node; If the integration node proportion is less than or equal to the preset integration node proportion, the key video frame of the integrated data cluster is determined according to the mutation dispersion index.
8. An intelligent preprocessing system using the intelligent preprocessing method based on multi-source heterogeneous data fusion according to any one of claims 1 to 7, characterized in that, It includes: The node evaluation module is used to determine the correlation analysis strategy of each data capture node in the target service area according to the node mutation coefficient and the node complexity coefficient, that is, to perform integration correlation analysis or co-occurrence correlation analysis on the fusion analysis data of the data capture node; The integration analysis module is connected with the node evaluation module, and is used to determine the integration strategy of the one-type data capture node according to the integration analysis index, that is, to determine the integrated data cluster according to the reference integration index and the reference flow correlation index, or the node conflict parameter and the reference difference stability parameter; The co-occurrence analysis module is connected with the node evaluation module, and is used to determine the integrated data cluster of each two-type data capture node according to the mutation correlation index; The optimization execution module is connected with the integration analysis module and the co-occurrence analysis module, respectively, and is used to determine the data optimization mode of each integrated data cluster according to the correlation analysis strategy of each data capture node, that is, to determine the preferred video segment according to each optimization analysis combination of the integrated data cluster, or to determine the determination mode of the key video frame according to the integration node proportion of the integrated data cluster.
Citation Information
Patent Citations
Multi-source heterogeneous data fusion method applied to smart city
CN118839292A
Multi-source heterogeneous data fusion processing and intelligent analysis method, device and equipment
CN119557845A
Debris flow monitoring method based on multimode data communication
CN119783004A