Data association method and device, electronic equipment and storage medium

By employing a two-layer index architecture and a confidence factor evaluation-based association method, the problem of insufficient retrieval efficiency and reliability of existing multi-source data association methods under sudden operational conditions is solved. This enables efficient and reliable fusion of multi-source heterogeneous data at the launch site, improving the system's applicability and the real-time nature of decision-making.

CN122046277APending Publication Date: 2026-05-15CHINA ELECTRONICS CORP 6TH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA ELECTRONICS CORP 6TH RES INST
Filing Date
2026-02-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing multi-source data association methods struggle to balance retrieval efficiency and system scalability under sudden operational conditions. Fixed-weight assignment models fail to consider data source reliability and transmission link status, and qualitatively designed association algorithms lack differentiated strategies for different launch conditions, making it difficult to meet the high real-time and high reliability requirements of critical launch mission phases.

Method used

A two-layer index architecture (preset spatiotemporal data index and preset feature data index) is adopted for data matching. Standardized feature vectors are generated through spatiotemporal benchmark calibration and feature extraction standardization. Hash tables and differentiated feature data indexes are constructed. The confidence level and association strategy are dynamically adjusted by combining confidence factor evaluation and working condition adaptation.

Benefits of technology

Achieving efficient and reliable fusion of multi-source heterogeneous data in complex launch environments significantly improves system reliability and engineering applicability, ensuring efficient data retrieval and timely decision-making instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122046277A_ABST
    Figure CN122046277A_ABST
Patent Text Reader

Abstract

The invention provides a data association method and device, electronic equipment and a storage medium, and the method comprises the steps: when a data association request is detected, based on a preset spatio-temporal data index and a preset feature data index, extracting to-be-associated historical data from the historical data of each data monitoring device of a plurality of data monitoring devices; for each data monitoring device, determining a confidence coefficient corresponding to the to-be-associated historical data of the data monitoring device based on the current emission site working condition and the confidence factor corresponding to the to-be-associated historical data of the data monitoring device; and determining a data association strategy based on the current launching site working condition and the confidence coefficient corresponding to the to-be-associated historical data of each data monitoring device, and associating the to-be-associated historical data of all the data monitoring devices according to the data association strategy. By adopting the technical scheme provided by the invention, efficient and credible fusion of multi-source heterogeneous data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a data association method, apparatus, electronic device, and storage medium. Background Technology

[0002] Command and decision-making in the C3I (Command, Control, and Information) system of a smart launch site heavily relies on the collaborative support of multi-source heterogeneous data, including onboard sensor data, ground-based telemetry and control radar data, equipment status monitoring data, meteorological and environmental data, and command interaction data. This data exhibits characteristics of structural heterogeneity, spatiotemporal heterogeneity, and qualitative heterogeneity. This multi-dimensional heterogeneity presents significant technical challenges in unified modeling, rapid retrieval, and accurate correlation of launch site data. Existing methods for multi-source data correlation typically employ the following approaches: data indexing mechanisms using single hash indexes or tree indexes, fixed-weight assignment models, and qualitatively designed correlation algorithms.

[0003] However, methods using single hash indexes or tree indexes struggle to balance retrieval efficiency and system scalability under sudden operational conditions; fixed-weight assignment models fail to adequately consider data source reliability, transmission link status, and changes in launch conditions, making it easy for low-quality data to participate in association and amplify errors; qualitative design of association algorithms fails to design differentiated association strategies for different launch conditions, making it difficult to meet the actual needs for high real-time performance and high reliability in critical stages of launch missions. Summary of the Invention

[0004] In view of this, embodiments of this application provide a data association method, apparatus, electronic device, and storage medium to achieve efficient and reliable fusion of multi-source heterogeneous data, significantly improving the reliability and engineering applicability of the system in complex launch environments.

[0005] This application mainly includes the following aspects: In a first aspect, embodiments of this application provide a data association method, the association method comprising: When a data association request is detected, the historical data to be associated is extracted from the historical data of each of the multiple data monitoring devices based on the preset spatiotemporal data index and the preset feature data index. For each data monitoring device, based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated with that data monitoring device, the confidence level corresponding to the historical data to be associated with that data monitoring device is determined; Based on the current launch site conditions and the confidence level of the historical data to be associated for each data monitoring device, a data association strategy is determined, and the historical data to be associated for all data monitoring devices is associated according to the data association strategy.

[0006] Furthermore, based on a preset spatiotemporal data index and a preset feature data index, the historical data to be associated is extracted from the historical data of each of the multiple data monitoring devices, including: Based on a preset spatiotemporal data index, historical data that meets preset time and preset spatial coordinates is extracted from the historical data to be associated of each of the multiple data monitoring devices. For each data monitoring device, based on the preset feature data index corresponding to the data monitoring device, historical data that meets the preset features is extracted from the historical data of the data monitoring device that meets the preset time and preset spatial coordinates; The historical data of each data monitoring device after extraction is identified as the historical data to be associated.

[0007] Furthermore, the association method also includes: For each data monitoring device, based on the original data collected by the data monitoring device and the corresponding collection time, the original data collected by the data monitoring device is supplemented. Based on the reference time drift error, the acquisition time of the original data collected by the data monitoring device after the data completion process is time-calibrated. Based on the spatial coordinates of the launch site, the data monitoring device is adjusted from its current coordinate system to the coordinate system of the launch site in order to perform coordinate calibration on the spatial coordinates of the data monitoring device. Based on the structure type of the original data collected by the data monitoring device, feature extraction and standardization are performed on the original data collected by the data monitoring device after time calibration and the spatial coordinates of the data monitoring device after coordinate calibration to obtain the standardized feature vector corresponding to the data monitoring device. Spatiotemporal data indexes and feature data indexes are constructed based on the standardized feature vectors corresponding to the data monitoring equipment.

[0008] Furthermore, based on the structure type of the original data collected by the data monitoring device, feature extraction and standardization are performed on the time-calibrated original data collected by the data monitoring device and the coordinate-calibrated spatial coordinates of the data monitoring device to obtain the standardized feature vector corresponding to the data monitoring device, including: Based on the structure type of the original data collected by the data monitoring device, feature extraction is performed on the original data collected and the spatial coordinates of the data monitoring device after coordinate calibration to generate the feature vector corresponding to the data monitoring device. The feature vector corresponding to the data monitoring device is subjected to interference processing and normalization processing in sequence, and the processed feature vector corresponding to the data monitoring device is determined as the standardized feature vector corresponding to the data monitoring device.

[0009] Furthermore, the construction of spatiotemporal data indexes and feature data indexes for the standardized feature vectors corresponding to all data monitoring devices includes: Based on the metadata in the standardized feature vector corresponding to the data monitoring device, determine the hash key of the standardized feature vector corresponding to the data monitoring device; Create a hash table to store pointers to the standardized feature vectors corresponding to all data monitoring devices; Based on the hash key of the standardized feature vector corresponding to the data monitoring device, the pointer of the standardized feature vector corresponding to the data monitoring device is stored in the corresponding position in the hash table; Based on the structure type of the original data collected by the data monitoring device, a feature data index is constructed using the standardized feature vector corresponding to the data monitoring device.

[0010] Furthermore, determining the confidence level of the historical data to be associated by the data monitoring equipment based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated by the data monitoring equipment includes: Determine the weights of the confidence factors corresponding to the historical data to be associated with the data monitoring device; Based on the confidence factors corresponding to the historical data to be associated by the data monitoring device and the weights of the confidence factors corresponding to the historical data to be associated by the data monitoring device, the confidence level of the historical data to be associated by the data monitoring device is determined.

[0011] Furthermore, the process of determining the association strategy for the historical data to be associated with all data monitoring devices based on the current launch site conditions and the confidence level corresponding to the historical data to be associated with each data monitoring device includes: If the current launch site condition is detected as an equipment failure condition, then the highest confidence level is selected from the confidence levels corresponding to the historical data to be associated of all data monitoring devices; the historical data to be associated corresponding to the highest confidence level is determined as the association result of the historical data to be associated of all data monitoring devices. If the current launch site condition is detected as a routine monitoring condition, the historical data to be associated from all data monitoring devices is weighted and averaged with the corresponding confidence level, and the weighted average result is determined as the association result of the historical data to be associated from all data monitoring devices. If the current launch site condition is detected as a launch condition, the historical data to be associated from all data monitoring devices is weighted and averaged with the corresponding confidence level, and the weighted average result is verified by a preset rule constraint; if the verification is successful, the weighted average result of the successful verification is determined as the association result of the historical data to be associated from all data monitoring devices.

[0012] Secondly, embodiments of this application also provide a data association device, the association device comprising: The retrieval module is used to extract the historical data to be associated from the historical data of each of the multiple data monitoring devices based on the preset spatiotemporal data index and the preset feature data index when a data association request is detected. The confidence level determination module is used to determine the confidence level of the historical data to be associated for each data monitoring device based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated for that data monitoring device. The association module is used to determine the data association strategy based on the current launch site conditions and the confidence level of the historical data to be associated for each data monitoring device, and to associate the historical data to be associated for all data monitoring devices according to the data association strategy.

[0013] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory through the bus, and the machine-readable instructions are executed by the processor to perform the steps of the data association method described in the first aspect or any possible implementation of the first aspect.

[0014] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the data association method described in the first aspect or any possible implementation of the first aspect.

[0015] This application provides a data association method, apparatus, electronic device, and storage medium. When a data association request is detected, based on a preset spatiotemporal data index and a preset feature data index, historical data to be associated is extracted from the historical data of each of multiple data monitoring devices. For each data monitoring device, based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated of that data monitoring device, the confidence level corresponding to the historical data to be associated of that data monitoring device is determined. Based on the current launch site conditions and the confidence level corresponding to the historical data to be associated of each data monitoring device, a data association strategy is determined, and the historical data to be associated of all data monitoring devices is associated according to the data association strategy.

[0016] This enables efficient and reliable fusion of multi-source heterogeneous data, significantly improving the system's reliability and engineering applicability in complex launch environments.

[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This invention provides a flowchart of a data association method according to an embodiment of the present application. Figure 2 A second flowchart of a data association method provided in an embodiment of this application is shown; Figure 3 A flowchart of a data association method provided in an embodiment of this application is shown as third; Figure 4 A flowchart of a data association method provided in an embodiment of this application is shown as fourth; Figure 5 The fifth flowchart illustrates a data association method provided in an embodiment of this application; Figure 6 A flowchart of a data association method provided in an embodiment of this application is shown as sixth; Figure 7 This illustration shows a schematic diagram of the structure of a data association device provided in an embodiment of this application; Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0021] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0022] The methods, apparatus, electronic devices, or computer-readable storage media described in this application can be applied to any scenario that requires data association. This application does not limit the specific application scenario, and any scheme using the data association methods and apparatus provided in this application is within the protection scope of this application.

[0023] It is worth noting that the command and decision-making of the intelligent launch site C3I system highly relies on the collaborative support of multi-source heterogeneous data, including onboard sensor data, ground-based telemetry and control radar data, equipment status monitoring data, meteorological and environmental data, and command interaction data. This data exhibits characteristics of structural heterogeneity, spatiotemporal heterogeneity, and quality heterogeneity. This multi-dimensional heterogeneity presents significant technical challenges in unified modeling, rapid retrieval, and accurate correlation of launch site data. Existing multi-source data correlation methods typically employ the following approaches: data indexing mechanisms using single hash indexes or tree indexes, fixed-weight assignment models, and qualitatively designed correlation algorithms. However, single hash indexes or tree indexes struggle to balance retrieval efficiency and system scalability under unforeseen circumstances; fixed-weight assignment models fail to adequately consider data source reliability, transmission link status, and changes in launch conditions, making it easy for low-quality data to participate in correlation and amplify errors; and qualitatively designed correlation algorithms lack differentiated correlation strategies for different launch conditions, failing to meet the actual requirements for high real-time performance and high reliability during critical stages of launch missions.

[0024] To address the aforementioned problems, embodiments of this application propose... A data association method, apparatus, electronic device, and storage medium are disclosed to achieve efficient and reliable fusion of multi-source heterogeneous data, significantly improving the system's reliability and engineering applicability in complex launch environments.

[0025] To facilitate understanding of this application, the technical solutions provided in this application will be described in detail below with reference to specific embodiments.

[0026] In this embodiment, the C3I (Command, Control, Communication and Intelligence) system of the smart launch site relies on multi-source heterogeneous data for command and decision-making, including onboard sensor data, ground-based telemetry and control radar data, equipment status monitoring data, meteorological and environmental data, and command interaction data. This data exhibits significant heterogeneity in structure, spatiotemporal context, and quality. Specifically: data structure heterogeneity manifests as structured parameter tables, semi-structured XML messages, and unstructured waveform / image data; spatiotemporal heterogeneity manifests as sampling frequencies ranging from Hz to kHz, with inconsistent spatiotemporal benchmarks; and quality heterogeneity manifests as susceptibility to strong electromagnetic interference, equipment failures, and large fluctuations in data reliability. This presents challenges for data association and processing. Current multi-source data association algorithms mainly suffer from the following three core problems: 1. Insufficient adaptability of indexing algorithms: Currently, most algorithms use single hash indexes or tree indexes. These algorithms fail to effectively adapt to the needs of high-dimensional heterogeneous data, easily leading to index conflicts and retrieval delays when data volume increases significantly. This can cause delays in decision-making instructions, especially in emergency situations, increasing risks. 1. **Single Hash Index or Tree Index Algorithms:** Existing algorithms often employ single hash or tree index algorithms without constructing targeted quantitative mathematical models for high-dimensional heterogeneous data. In the face of massive amounts of data during sudden launch site emergencies, they often fail to effectively guarantee retrieval efficiency and system scalability, easily leading to index conflicts and retrieval delays, exceeding the system's millisecond-level response requirements. In critical scenarios such as emergency fault response, this delay may even lead to data retrieval timeouts and delayed decision-making instructions. 2. **Static Confidence Assessment Algorithms:** Existing algorithms are typically based on fixed weight assignment models, failing to consider dynamic factors such as data source reliability, transmission link status, and launch conditions. When low-quality data, such as radar waveform data affected by electromagnetic interference or abnormal parameters output by faulty sensors, are included in the analysis, the error rate of the output results may increase significantly. This high error rate can interfere with command decisions, especially during critical phases of the launch mission, such as propellant loading and ignition countdown, where such interference is particularly severe, making it difficult to meet the mission's stringent requirements for high reliability. 3. **Lack of Mathematical Support in Association Algorithms:** Most widely used association algorithms are based on qualitative design and lack compensation mechanisms for spatiotemporal calibration errors and quantitative models for confidence-weighted fusion. Furthermore, no specific differentiated correlation strategies were designed for different operating environments at the launch site (such as routine monitoring, emergency response, and special environments). During critical phases at the launch site, this deficiency in algorithm design made it difficult to meet the stringent requirements of the C3I system for high real-time performance and high reliability, affecting the system's adaptability and efficiency in actual operation.To address the aforementioned issues, it is necessary to develop indexing algorithms that are more adaptable to high-dimensional heterogeneous data, introduce dynamic quantization parameters to enhance the accuracy of confidence assessment, and establish mathematical models to support the quantitative design of association algorithms, thereby improving the performance and reliability of the C3I system under various operating conditions.

[0027] Please see Figure 1 , Figure 1 This is one of the flowcharts for a data association method provided in an embodiment of this application.

[0028] like Figure 1 As shown in the embodiments of this application, the data association method includes the following steps: Step S101: When a data association request is detected, based on the preset spatiotemporal data index and the preset feature data index, the historical data to be associated is extracted from the historical data of each of the multiple data monitoring devices.

[0029] Here, the historical data of the data monitoring equipment refers to data collected and stored by various data monitoring devices such as rocket-borne sensors and ground control stations within a preset time range in the past, and processed through standardization, indexing, etc., making it available for retrieval and correlation analysis. In this application, the spatiotemporal data index is a hash index, and the feature data index is a tree index.

[0030] In this embodiment, to meet the launch site's requirement for high real-time retrieval of massive amounts of data, a two-layer index architecture is adopted: first, coarse matching is achieved through a preset spatiotemporal data index; second, fine matching is achieved using a preset feature data index. This effectively improves retrieval efficiency and ensures fast response speed when processing large-scale data.

[0031] The following is combined with Figure 2 This section explains in detail how to extract the historical data to be associated from the historical data of each of multiple data monitoring devices based on preset spatiotemporal data indexes and preset feature data indexes.

[0032] Please see Figure 2 , Figure 2 This is a second flowchart of a data association method provided in an embodiment of this application.

[0033] like Figure 2 As shown, regarding step S101, in a specific implementation, as an example, the following steps may be included: Step S1011: Based on the preset spatiotemporal data index, extract historical data that meets the preset time and preset spatial coordinates from the historical data to be associated of each of the multiple data monitoring devices.

[0034] Here, by using a preset spatiotemporal data index (i.e., a global spatiotemporal hash index), the time range and spatial location in the request are used as keys to quickly filter out all historical data generated by data monitoring devices within the spatiotemporal window, thus completing the initial coarse matching.

[0035] Step S1012: For each data monitoring device, based on the preset feature data index corresponding to the data monitoring device, extract historical data that meets the preset features from the historical data that meets the preset time and preset spatial coordinates in the data monitoring device.

[0036] Here, using a preset feature data index (i.e., a tree index, such as a B+ tree, an inverted index, or a KD tree), based on the specific data features or content specified in the request, the historical data that meets the feature conditions is further precisely located and extracted from the coarsely matched historical data for each data monitoring device.

[0037] Step S1013: The historical data of each data monitoring device after extraction is identified as the historical data to be associated.

[0038] In one possible embodiment, such as Figure 3 As shown, the spatiotemporal data index and feature data index are constructed in the following manner: Step S11: For each data monitoring device, based on the original data collected by the data monitoring device and the collection time corresponding to the original data collected, complete the original data collected by the data monitoring device.

[0039] It should be noted that the core of the two-layer index architecture in this application lies in achieving rapid data matching using spatiotemporal attributes and feature attributes. However, the raw collected data faces the problem of inconsistent spatiotemporal benchmarks, such as device clock drift and coordinate system differences. Directly building an index under these conditions may lead to spatiotemporal association errors, such as misclassifying data from different times as data from the same period, or mistaking device data from different locations for data from the same region. To solve this problem, a unified correction of the spatiotemporal benchmark is required before index construction to ensure the consistency of the spatiotemporal attributes of the data, thereby improving the accuracy and reliability of the index.

[0040] In step S11, a nonlinear interpolation completion method is used for time calibration. Specifically, it is assumed that the standard discrete time series of the data monitoring equipment is... Where m is the length of the time series, corresponding to the original collected data as follows: When missing timestamps exist When the data is completed, the cubic spline interpolation algorithm is used. The specific calculation formula is shown in formula (1).

[0041] (1).

[0042] in, To complete the original collected data, , , and These are the coefficients of the cubic spline interpolation. The coefficients of the cubic spline interpolation must satisfy the boundary conditions shown in formula (2): (2).

[0043] Step S12: Based on the reference time drift error, the acquisition time of the original data collected by the data monitoring device after the data completion process is time-calibrated.

[0044] Here, regarding time drift compensation, the drift error between the device clock and the BeiDou time reference is... ,in, This represents the time drift error value. For the running time of data monitoring equipment, and If the drift coefficient is used, then the calibrated timestamp is: ,in, The calibrated acquisition time. The data acquisition time is specified. The drift coefficient can be obtained by least-squares fitting of historical calibration data.

[0045] Step S13: Based on the spatial coordinates of the launch site, adjust the data monitoring device from its current coordinate system to the coordinate system of the launch site to perform coordinate calibration on the spatial coordinates of the data monitoring device.

[0046] Here, the coordinate system of the data monitoring equipment can be transformed using a rotation matrix that characterizes the attitude angle deviation between the original coordinate system of the data monitoring equipment and the coordinate system where the launch site is located, and a rotation matrix that characterizes the offset of the installation position of the data monitoring equipment in the coordinate system where the launch site is located. As an example, the original coordinate system of the data monitoring equipment can be transformed into the coordinate system where the launch site is located using formula (3).

[0047] (3).

[0048] in, This is the original coordinate vector of the data monitoring equipment. Let be the coordinate vector of the launch site. It is a 3×3 rotation matrix. It is a translation vector. .here, It can be calculated using formula (4).

[0049] (4).

[0050] in, , and These are the spatial attitude angle parameters.

[0051] To improve the calibration accuracy of individual data monitoring devices, a multi-device collaborative calibration model is used. This model selects a launch site reference device (such as a BeiDou ground station or a laser rangefinder) as a calibration reference and uses the collected data from the data monitoring devices for mutual verification to optimize the rotation matrix and translation vector. The objective function for optimization is shown in formula (5).

[0052] (5).

[0053] Where K represents the number of data monitoring devices participating in the collaborative calibration. The 2-norm represents the calibration error.

[0054] In this embodiment, the calibration granularity is adaptively adjusted according to the launch conditions: during critical phases (such as 10 minutes before ignition), the time calibration accuracy is increased to the microsecond level and the spatial calibration accuracy is increased to the centimeter level to ensure the high reliability of critical decisions and controls; during normal operation, the calibration granularity is appropriately reduced to effectively balance calibration accuracy and computational overhead while meeting business requirements.

[0055] Step S14: Based on the structure type of the original data collected by the data monitoring device, feature extraction and standardization are performed on the original data collected by the data monitoring device after time calibration and the spatial coordinates of the data monitoring device after coordinate calibration to obtain the standardized feature vector corresponding to the data monitoring device.

[0056] Here, the data structure types include: structured data, semi-structured data, and unstructured data (i.e., unstructured waveform / image data). Structured data includes: database tables, CSV files, Excel spreadsheets, etc. Semi-structured data includes: XML, JSON, HTML messages, emails, etc. Unstructured data includes: text, images, videos, audio, waveforms, PDF / Word documents, etc.

[0057] The following is combined with Figure 4 This paper explains how, based on the structure type of the original data collected by the data monitoring device, features are extracted and standardized from the time-calibrated original data and the coordinate-calibrated spatial coordinates of the data monitoring device to obtain the standardized feature vector corresponding to the data monitoring device.

[0058] Please see Figure 4 , Figure 4This is the fourth flowchart of a data association method provided in an embodiment of this application.

[0059] like Figure 4 As shown, regarding step S14, in a specific implementation, as an example, the following steps may be included: Step S141: Based on the structure type of the original data collected by the data monitoring device, feature extraction is performed on the corresponding original data collected and the spatial coordinates of the data monitoring device after coordinate calibration to generate the feature vector corresponding to the data monitoring device.

[0060] Here, for feature extraction of structured data, a key field mapping algorithm is used to extract fields such as equipment number, sampling time, and core parameters, generating a structured feature vector, which is the feature vector corresponding to the structured data. Among these, the core parameters refer to data that directly affects the safety and success of the launch mission and provides the most critical basis for command and control decisions.

[0061] For feature extraction of semi-structured data, the XPath+label weight algorithm is used to parse the message, assign higher weights to the core instruction labels, and generate semi-structured feature vectors, that is, the feature vectors corresponding to the semi-structured data.

[0062] For feature extraction of unstructured data, a wavelet transform + lightweight CNN feature extraction algorithm is adopted to extract time-frequency domain features from measurement and control waveform data and visual features from equipment image data, generating unstructured feature vectors, i.e. feature vectors corresponding to unstructured data.

[0063] Step S142: The feature vector corresponding to the data monitoring device is subjected to interference processing and normalization processing in sequence, and the processed feature vector corresponding to the data monitoring device is determined as the standardized feature vector corresponding to the data monitoring device.

[0064] Here, for structured feature vectors, to improve their robustness under strong electromagnetic interference environments, the 3σ criterion is used to process abnormally abrupt parameters in the structured feature vectors affected by electromagnetic interference, i.e., parameters exceeding [μ] are removed. Outliers in the range of [3σ, μ+3σ], where μ is the mean of the structured feature vector and σ is the standard deviation of the structured feature vector. For semi-structured feature vectors, the label missing problem caused by link jitter during message transmission is repaired by using an LSTM-based label completion model, that is, interference processing is performed on the semi-structured feature vectors. Among them, the core instruction is the instruction used to directly intervene in or change the key safety process of the launch mission. For unstructured feature vectors, the waveform spikes and image noise caused by electromagnetic interference are processed by the wavelet threshold denoising algorithm. The quantization model can be shown in formula (6): (6).

[0065] in, These are the wavelet transform coefficients. An adaptive threshold (automatically calculated from noise energy). These are the wavelet coefficients after denoising.

[0066] In this embodiment of the application, the three types of feature vectors are mapped to the [0,1] interval by using the min-max normalization algorithm. The normalization calculation can be shown in formula (7): (7).

[0067] in, These are the eigenvalues ​​in the normalized eigenvectors. These are the eigenvalues ​​in the feature vector after interference processing. The minimum value under the feature dimension. This represents the maximum value within the feature dimension. The set of normalized feature values ​​is ultimately determined as the standardized feature vector, which contains metadata (such as the data source and spatiotemporal coordinates of the data collection).

[0068] To address the fundamental issue of the inability to directly correlate heterogeneous data, this application designs differentiated feature extraction sub-algorithms for data with different structural types. Furthermore, an electromagnetic interference robustness processing module is introduced to enhance the stability and reliability of data processing. Through these sub-algorithms, standardized feature vectors in a unified feature space are generated, and normalization processing eliminates dimensional differences between different data, thereby achieving effective correlation of heterogeneous data.

[0069] See again Figure 3 Step S15: Construct a spatiotemporal data index and a feature data index for the standardized feature vector corresponding to the data monitoring device.

[0070] The following is combined with Figure 5 This will illustrate how to construct a spatiotemporal data index and a feature data index based on the standardized feature vectors corresponding to the data monitoring device.

[0071] Please see Figure 5 , Figure 5 This is the fifth flowchart of a data association method provided in the embodiments of this application.

[0072] like Figure 5 As shown, regarding step S15, in a specific implementation, as an example, the following steps may be included: Step S151: Based on the metadata in the standardized feature vector corresponding to the data monitoring device, determine the hash key of the standardized feature vector. Here, the metadata includes: the calibrated acquisition time and the calibrated spatial coordinates. Spatiotemporal hash key. The mathematical model is shown in formula (8): (8).

[0073] in, For hash functions, For modulo operation, The weighting coefficients are the weighting coefficients for the spatiotemporal dimensions, and the weighting coefficients satisfy... and , The length of the hash table. The value is set to a prime number to reduce the probability of hash collisions.

[0074] Step S152: Create a hash table to store pointers to the standardized feature vectors corresponding to all data monitoring devices.

[0075] Step S153: Based on the hash key of the standardized feature vector corresponding to the data monitoring device, store the pointer of the standardized feature vector corresponding to the data monitoring device in the corresponding position of the hash table.

[0076] It should be noted that if the corresponding position in the hash table is full, i.e., a hash collision occurs, the chaining method combined with a rehashing combination strategy is used to resolve the issue. The rehashing function is shown in formula (9): (9) in, For rehashing functions, To generate random numbers between 1 and 10, =2 +1 to ensure efficient storage and retrieval of conflicting data.

[0077] Step S154: Based on the structure type of the original collected data corresponding to the data monitoring device, construct a feature data index for the standardized feature vector corresponding to the data monitoring device.

[0078] Here, we design quantized indexing models for three types of standardized feature vectors to improve the efficiency and accuracy of data retrieval.

[0079] For structured data, an optimized B+ tree indexing algorithm is used to construct the index, specifically based on sorting by key fields. As an example, assuming the extracted key fields are sampling timestamp, device number, and core parameter type, a third-order B+ tree index is constructed in ascending order of sampling timestamp → ascending order of device number → descending order of core parameters. The root node stores the sampling timestamp range, non-leaf nodes associate the device number with the core parameter combination, and leaf nodes store pointers to the standardized feature vectors corresponding to the structured data and their metadata through a doubly linked list, enabling fast single-point queries and range searches.

[0080] For semi-structured data, an inverted tree indexing algorithm is used to construct the index, specifically based on label feature mapping. As an example, the three-level structure of root label → parent label → child label in XML / JSON messages is parsed. Core instruction labels are assigned high weights (greater than or equal to 0.8), while ordinary labels are assigned basic weights (less than 0.2). Missing labels are then filled in using an LSTM model. A label tree + inverted index structure is constructed, with the root node corresponding to the top-level label. First and second-level nodes map to parent and child labels respectively, and child labels are associated with the inverted index, storing standardized feature vectors and pointers to them. During retrieval, the inverted index is quickly located via the label path and output in sorted order by initial confidence value, achieving precise matching of individual tagged data.

[0081] For unstructured data, the KD tree nearest neighbor retrieval algorithm is used to construct the index, and the weighted Manhattan distance metric function is defined as the basis for standardizing feature vector matching. The weighted Manhattan distance metric function is shown in formula (10).

[0082] (10).

[0083] in, Let be the distance function. , These represent the i-th and j-th eigenvectors, respectively. n The dimension representing the feature vector. The weights represent the k-th dimension features. , Represent , The k-th eigenvalue. Weights Calculated using the information gain method, and satisfying... .

[0084] One possible implementation involves splitting or merging index nodes in the dual-layer hybrid index structure when the data inflow rate exceeds a preset threshold or the transmission mode changes (e.g., from regular monitoring mode to fault emergency mode). This clears redundant indexes and ensures stable index retrieval efficiency. The data inflow rate is defined as the number of valid data entries entering the data processing center from all data sources (sensors, radar, cameras, command systems, etc.) per unit time. When abnormal data monitoring devices are identified (e.g., sensor malfunction or link interruption), the corresponding index node is marked as isolated, and an independent abnormal data index structure is constructed to prevent faulty data from interfering with the normal data retrieval process. As an example, the isolation trigger condition is... ,in, This refers to the abnormal data volume of the data monitoring equipment. This represents the total amount of data from the data monitoring equipment. The fault determination threshold is... It is obtained by training with historical fault data, and the value range is [0.2, 0.5].

[0085] See again Figure 1 In step S102, for each data monitoring device, based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated with the data monitoring device, the confidence level corresponding to the historical data to be associated with the data monitoring device is determined.

[0086] Here, the confidence factor includes four dimensions: data source reliability factor, transmission link stability factor, operating condition adaptability factor, and data quality factor.

[0087] The following is combined with Figure 6 This explains how, for each data monitoring device, based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated with that data monitoring device, the confidence level of the historical data to be associated with that data monitoring device is determined.

[0088] Please see Figure 6 , Figure 6 This is a flowchart of a data association method provided in an embodiment of this application.

[0089] like Figure 6 As shown, regarding step S102, in a specific implementation, as an example, the following steps may be included: Step S1021: Determine the weight of the confidence factor corresponding to the historical data to be associated by the data monitoring device.

[0090] Here, the entropy weighting method is used to calculate the weight coefficients. The information entropy of the confidence factor can be calculated using formula (11).

[0091] (11).

[0092] in, Indicates the first Information entropy of each dimension factor For the sample size, , Indicates the first The first dimension Factor values ​​for each sample.

[0093] (12).

[0094] in, The weights corresponding to the data source reliability factor. The weights corresponding to the transmission link stability factor. The weights corresponding to the operating condition adaptability factor These are the weights corresponding to the data quality factors. Information entropy is the reliability factor of the data source. Information entropy is the transmission link stability factor. Information entropy for the working condition adaptability factor. The information entropy of the data quality factor. Additionally, the following constraints are added: + + + =1, and , , , ∈[0.2, 0.6], to ensure that the weight ratio of each dimension factor is not excessively unbalanced.

[0095] Step S1022: Based on the confidence factor corresponding to the historical data to be associated by the data monitoring device and the weight of the confidence factor corresponding to the historical data to be associated by the data monitoring device, determine the confidence level of the historical data to be associated by the data monitoring device.

[0096] Here, the confidence level can be calculated using formula (13).

[0097] (13).

[0098] in, For confidence level, For data source reliability factor, For transmission link stability factor, Operating condition adaptability factor This is the data quality factor.

[0099] In this embodiment of the application, the reliability factor of the data source is determined based on the historical accuracy and data monitoring equipment level of the data monitoring equipment. Quantify, It can be calculated using formula (14).

[0100] (14).

[0101] in, This is the weighting coefficient, which can take a value of 0.6; The amount of historically correct data from the data monitoring equipment. This represents the total historical data volume of the device. and The ratio is used to characterize the historical accuracy of data monitoring equipment; This represents the equipment level coefficient, where core equipment is set to 1, and auxiliary equipment is set to... =0.5, common sensor value =0.3.

[0102] Based on packet loss rate during data transmission p With latency d l An exponential decay model is used to evaluate the stability factor of the transmission link. Quantify, It can be calculated using formula (15).

[0103] (15).

[0104] in, It is a natural constant. k p , k d The attenuation coefficient is... k p It can be set to 10. k d It can be set to 5, packet loss rate p The latency rate is the quotient of the total number of data packets sent and the number of packets lost. d l It is the quotient of average transmission delay and system delay threshold.

[0105] Based on operating condition type and data importance, operating condition adaptability factors are used. Quantify, It can be calculated using formula (16).

[0106] (16).

[0107] in, q For the number of working condition types, Let's take the state value of the i-th type of operating condition as an example. A key operating condition is a launch countdown. =1, under normal operating conditions =0.5, The importance weights of the corresponding data under this working condition are obtained by combining expert experience and machine learning models.

[0108] Based on data integrity, consistency, and noise level, data quality factors are evaluated. Quantify, It can be calculated using formula (17).

[0109] (17).

[0110] in, For data integrity, For the effective data length, (Total data length) For data consistency, For inconsistent data volume, For noise level, The standard deviation of the data noise. This is the noise threshold.

[0111] Step S103: Based on the current launch site conditions and the confidence level of the historical data to be associated for each data monitoring device, determine the data association strategy, and associate the historical data to be associated for all data monitoring devices according to the data association strategy.

[0112] Here, if the current launch site condition is detected as an equipment failure condition, the highest confidence level is selected from the historical data to be associated from all data monitoring devices. The historical data to be associated corresponding to the highest confidence level is then determined as the association result of the historical data to be associated from all data monitoring devices. This method reduces fusion latency and meets the real-time requirements of emergency decision-making.

[0113] If the current launch site operating conditions are detected as routine monitoring conditions, then a weighted average is calculated between the historical data to be associated from all data monitoring devices and their corresponding confidence levels. The result of this weighted average is then determined as the association result for the historical data to be associated from all data monitoring devices. This method takes into account the complementarity of information from multiple data sources, improving the stability of the association results.

[0114] If the current launch site condition is detected as a launch condition, the historical data to be associated from all data monitoring devices is weighted and averaged with the corresponding confidence levels. The weighted average result is then verified using preset rules. If the verification is successful, the weighted average result is determined as the association result for the historical data to be associated from all data monitoring devices. This method ensures that the association result meets launch safety requirements. Preset rules can include whether the thrust exceeds a safety threshold, whether the temperature exceeds a temperature threshold, etc.

[0115] It should be noted that the operating conditions are identified in the following way: based on the launch mission phase information (including fueling, transfer, countdown, ignition and orbit insertion, etc.), equipment operating status (normal, abnormal, fault, etc.) and environmental condition parameters (such as electromagnetic interference intensity and meteorological level, etc.), an operating condition feature vector is constructed, and the launch operating conditions are automatically identified based on a random forest classification model.

[0116] Prior to step S103, the association method includes: filtering out confidence levels greater than or equal to a confidence threshold from the confidence levels of all data monitoring devices. Here, the confidence threshold can be automatically adjusted according to the operating conditions. For example, in device failure conditions, the confidence threshold is 0.8; in normal monitoring conditions, the confidence threshold is 0.5; and in massive data conditions, the confidence threshold is increased according to a preset adjustment ratio.

[0117] When it is necessary to perform a weighted average of the historical data to be associated from the data monitoring equipment, the association result can be calculated using formula (18).

[0118] (18).

[0119] in, For the associated results, For the confidence value corresponding to the i-th data to be associated, This is the i-th piece of data to be associated.

[0120] Based on the standardized feature vectors of different data monitoring devices, an improved Apriori algorithm is used to mine association rules in the fused dataset in order to identify and output the implicit associations between the data.

[0121] This application overcomes the problems of high retrieval latency and easy conflict in traditional single indexes under complex launch conditions by introducing a hierarchical hybrid index and a condition-aware mechanism, ensuring stable system response under high real-time requirements. Simultaneously, the algorithm can effectively quantify the impact of factors such as electromagnetic interference, equipment failure, and changes in operating conditions on data quality, avoiding interference from low-quality data in the association results. The association process is driven by confidence level, and the accuracy of the association results is improved through threshold filtering and weighted fusion mechanisms, compensating for the shortcomings of traditional algorithms that prioritize matching over quality. Furthermore, a robust processing mechanism for strong electromagnetic interference environments is introduced to improve feature extraction and association stability of heterogeneous data. An adaptive adjustment strategy based on operating conditions enables automatic switching of algorithm parameters and association strategies, meeting the differentiated needs of the entire launch mission process and significantly enhancing the system's engineering applicability and reliability without manual intervention.

[0122] This application provides a data association method that enables efficient and reliable fusion of multi-source heterogeneous data, significantly improving the system's reliability and engineering applicability in complex launch environments.

[0123] Based on the same application concept, this application also provides a data association device corresponding to the data association method provided in the above embodiments. Since the principle of the device in this application to solve the problem is similar to the data association method in the above embodiments of this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0124] Please see Figure 7 , Figure 7 This is a schematic diagram of a data association device provided in an embodiment of this application.

[0125] like Figure 7 As shown in the illustration, the data association device 710 provided in this application embodiment includes: The retrieval module 711 is used to extract the historical data to be associated from the historical data of each of the multiple data monitoring devices based on the preset spatiotemporal data index and the preset feature data index when a data association request is detected. The confidence level determination module 712 is used to determine the confidence level of the historical data to be associated for each data monitoring device based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated for that data monitoring device. The association module 713 is used to determine the data association strategy based on the current launch site conditions and the confidence level of the historical data to be associated for each data monitoring device, and to associate the historical data to be associated for all data monitoring devices according to the data association strategy.

[0126] Furthermore, the retrieval module 711 is specifically used for: Based on a preset spatiotemporal data index, historical data that meets preset time and preset spatial coordinates is extracted from the historical data to be associated of each of the multiple data monitoring devices. For each data monitoring device, based on the preset feature data index corresponding to the data monitoring device, historical data that meets the preset features is extracted from the historical data of the data monitoring device that meets the preset time and preset spatial coordinates; The historical data of each data monitoring device after extraction is identified as the historical data to be associated.

[0127] Furthermore, the associated device 710 also includes: The completion module 714 is used to complete the original data collected by each data monitoring device based on the original data collected by the data monitoring device and the collection time corresponding to the original data collected. The time calibration module 715 is used to calibrate the acquisition time of the original data collected by the data monitoring device after the data monitoring device has been completed, based on the reference time drift error. The coordinate calibration module 716 is used to adjust the data monitoring device from its current coordinate system to the coordinate system of the launch site based on the spatial coordinates of the launch site, so as to perform coordinate calibration on the spatial coordinates of the data monitoring device. The processing module 717 is used to perform feature extraction and standardization processing on the original data collected by the data monitoring device after time calibration and the spatial coordinates of the data monitoring device after coordinate calibration, based on the structure type of the original data collected by the data monitoring device, to obtain the standardized feature vector corresponding to the data monitoring device. Module 718 is used to construct spatiotemporal data indexes and feature data indexes for the standardized feature vectors corresponding to the data monitoring equipment.

[0128] Furthermore, the processing module 717 is specifically used for: Based on the structure type of the original data collected by the data monitoring device, feature extraction is performed on the original data collected and the spatial coordinates of the data monitoring device after coordinate calibration to generate the feature vector corresponding to the data monitoring device. The feature vector corresponding to the data monitoring device is subjected to interference processing and normalization processing in sequence, and the processed feature vector corresponding to the data monitoring device is determined as the standardized feature vector corresponding to the data monitoring device.

[0129] Furthermore, the construction module 718 is specifically used for: Based on the metadata in the standardized feature vector corresponding to the data monitoring device, determine the hash key of the standardized feature vector corresponding to the data monitoring device; Create a hash table to store pointers to the standardized feature vectors corresponding to all data monitoring devices; Based on the hash key of the standardized feature vector corresponding to the data monitoring device, the pointer of the standardized feature vector corresponding to the data monitoring device is stored in the corresponding position in the hash table; Based on the structure type of the original data collected by the data monitoring device, a feature data index is constructed using the standardized feature vector corresponding to the data monitoring device.

[0130] Furthermore, when the confidence determination module 712 determines the confidence level of the historical data to be associated by the data monitoring device based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated by the data monitoring device, it is also specifically used for: Determine the weights of the confidence factors corresponding to the historical data to be associated with the data monitoring device; Based on the confidence factors corresponding to the historical data to be associated by the data monitoring device and the weights of the confidence factors corresponding to the historical data to be associated by the data monitoring device, the confidence level of the historical data to be associated by the data monitoring device is determined.

[0131] Furthermore, when determining the association strategy for the historical data to be associated of all data monitoring devices based on the current launch site conditions and the confidence level corresponding to the historical data to be associated of each data monitoring device, the association module 713 is also specifically used for: If the current launch site condition is detected as an equipment failure condition, then the highest confidence level is selected from the confidence levels corresponding to the historical data to be associated of all data monitoring devices; the historical data to be associated corresponding to the highest confidence level is determined as the association result of the historical data to be associated of all data monitoring devices. If the current launch site condition is detected as a routine monitoring condition, the historical data to be associated from all data monitoring devices is weighted and averaged with the corresponding confidence level, and the weighted average result is determined as the association result of the historical data to be associated from all data monitoring devices. If the current launch site condition is detected as a launch condition, the historical data to be associated from all data monitoring devices is weighted and averaged with the corresponding confidence level, and the weighted average result is verified by a preset rule constraint; if the verification is successful, the weighted average result of the successful verification is determined as the association result of the historical data to be associated from all data monitoring devices.

[0132] This application provides a data association device that enables efficient and reliable fusion of multi-source heterogeneous data, significantly improving the system's reliability and engineering applicability in complex launch environments.

[0133] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0134] like Figure 8 As shown, the electronic device 800 includes a processor 810, a memory 820, and a bus 830.

[0135] The memory 820 stores machine-readable instructions executable by the processor 810. When the electronic device 800 is running, the processor 810 and the memory 820 communicate via the bus 830. When the machine-readable instructions are executed by the processor 810, they can perform the operations described above. Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 The steps of the data association method in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.

[0136] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described actions. Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 The steps of the data association method in the method embodiment shown are described in detail in the method embodiment, and will not be repeated here.

[0137] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0138] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0139] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0140] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0141] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data association method, characterized in that, The association method includes: When a data association request is detected, the historical data to be associated is extracted from the historical data of each of the multiple data monitoring devices based on the preset spatiotemporal data index and the preset feature data index. For each data monitoring device, based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated with that data monitoring device, the confidence level corresponding to the historical data to be associated with that data monitoring device is determined; Based on the current launch site conditions and the confidence level of the historical data to be associated for each data monitoring device, a data association strategy is determined, and the historical data to be associated for all data monitoring devices is associated according to the data association strategy.

2. The data association method according to claim 1, characterized in that, The process involves extracting historical data to be associated from the historical data of each of the multiple data monitoring devices, based on a preset spatiotemporal data index and a preset feature data index. This includes: Based on a preset spatiotemporal data index, historical data that meets preset time and preset spatial coordinates is extracted from the historical data to be associated of each of the multiple data monitoring devices. For each data monitoring device, based on the preset feature data index corresponding to the data monitoring device, historical data that meets the preset features is extracted from the historical data of the data monitoring device that meets the preset time and preset spatial coordinates; The historical data of each data monitoring device after extraction is identified as the historical data to be associated.

3. The data association method according to claim 1, characterized in that, The association method further includes: For each data monitoring device, based on the original data collected by the data monitoring device and the corresponding collection time, the original data collected by the data monitoring device is supplemented. Based on the reference time drift error, the acquisition time of the original data collected by the data monitoring device after the data completion process is time-calibrated. Based on the spatial coordinates of the launch site, the data monitoring device is adjusted from its current coordinate system to the coordinate system of the launch site in order to perform coordinate calibration on the spatial coordinates of the data monitoring device. Based on the structure type of the original data collected by the data monitoring device, feature extraction and standardization are performed on the original data collected by the data monitoring device after time calibration and the spatial coordinates of the data monitoring device after coordinate calibration to obtain the standardized feature vector corresponding to the data monitoring device. Spatiotemporal data indexes and feature data indexes are constructed based on the standardized feature vectors corresponding to the data monitoring equipment.

4. The data association method according to claim 3, characterized in that, Based on the structure type of the original data collected by the data monitoring device, feature extraction and standardization are performed on the time-calibrated original data collected by the data monitoring device and the coordinate-calibrated spatial coordinates of the data monitoring device to obtain the standardized feature vector corresponding to the data monitoring device, including: Based on the structure type of the original data collected by the data monitoring device, feature extraction is performed on the original data collected and the spatial coordinates of the data monitoring device after coordinate calibration to generate the feature vector corresponding to the data monitoring device. The feature vector corresponding to the data monitoring device is subjected to interference processing and normalization processing in sequence, and the processed feature vector corresponding to the data monitoring device is determined as the standardized feature vector corresponding to the data monitoring device.

5. The data association method according to claim 3, characterized in that, The construction of spatiotemporal data indexes and feature data indexes for the standardized feature vectors corresponding to all data monitoring devices includes: Based on the metadata in the standardized feature vector corresponding to the data monitoring device, determine the hash key of the standardized feature vector corresponding to the data monitoring device; Create a hash table to store pointers to the standardized feature vectors corresponding to all data monitoring devices; Based on the hash key of the standardized feature vector corresponding to the data monitoring device, the pointer of the standardized feature vector corresponding to the data monitoring device is stored in the corresponding position in the hash table; Based on the structure type of the original data collected by the data monitoring device, a feature data index is constructed using the standardized feature vector corresponding to the data monitoring device.

6. The data association method according to claim 1, characterized in that, The process of determining the confidence level of the historical data to be associated by the data monitoring equipment based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated by the data monitoring equipment includes: Determine the weights of the confidence factors corresponding to the historical data to be associated with the data monitoring device; Based on the confidence factors corresponding to the historical data to be associated by the data monitoring device and the weights of the confidence factors corresponding to the historical data to be associated by the data monitoring device, the confidence level of the historical data to be associated by the data monitoring device is determined.

7. The data association method according to claim 1, characterized in that, The process of determining the association strategy for the historical data to be associated with all data monitoring devices based on the current launch site conditions and the confidence level of the historical data to be associated with each data monitoring device includes: If the current launch site condition is detected as an equipment failure condition, then the highest confidence level is selected from the confidence levels corresponding to the historical data to be associated of all data monitoring devices; the historical data to be associated corresponding to the highest confidence level is determined as the association result of the historical data to be associated of all data monitoring devices. If the current launch site condition is detected as a routine monitoring condition, the historical data to be associated from all data monitoring devices is weighted and averaged with the corresponding confidence level, and the weighted average result is determined as the association result of the historical data to be associated from all data monitoring devices. If the current launch site condition is detected as a launch condition, the historical data to be associated from all data monitoring devices is weighted and averaged with the corresponding confidence level, and the weighted average result is verified by a preset rule constraint; if the verification is successful, the weighted average result of the successful verification is determined as the association result of the historical data to be associated from all data monitoring devices.

8. A data association device, characterized in that, The data association device includes: The retrieval module is used to extract the historical data to be associated from the historical data of each of the multiple data monitoring devices based on the preset spatiotemporal data index and the preset feature data index when a data association request is detected. The confidence level determination module is used to determine the confidence level of the historical data to be associated for each data monitoring device based on the current launch site conditions and the confidence factor corresponding to the historical data to be associated for that data monitoring device. The association module is used to determine the data association strategy based on the current launch site conditions and the confidence level of the historical data to be associated for each data monitoring device, and to associate the historical data to be associated for all data monitoring devices according to the data association strategy.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the machine-readable instructions are executed by the processor to perform the steps of the data association method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the data association method as described in any one of claims 1 to 7.