A method and apparatus for synchronous acquisition and processing of multi-source heterogeneous data

CN122470329BActive Publication Date: 2026-09-15XIAN SHENGXIN TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610954405.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-15
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

[0005]本申请的目的是提供一种用于多源异构数据的同步采集处理方法及装置,用于解决现有技术中的多设备独立采集方式,在多源异构数据处理中,由于设备时钟不统一和数据格式差异,存在时序难以对齐及处理效率低的技术问题

Benefits of technology

本申请实施例提供的方法通过现场可编程门阵列捕获外部授时信号并驯化高稳晶振得到高频基准时钟,并构建高精度时间戳;所述现场可编程门阵列同时向至少两类异构传感器发出采样触发指令,获取带所述高精度时间戳的带时标数据帧;将所述带时标数据帧按预设统一格式进行序列化封装,形成连续数据流;通过消息队列中间件对所述连续数据流中的任意数据进行配置,得到任意独立消息队列,并为所述任意独立消息队列配置任意消息消费者;所述任意消息消费者异步拉取所述连续数据流解析得到所述任意数据,并引入预设目标映射关系对所述任意数据进行映射转换。达到了通过构建高精度时间基准并结合消息队列中间件,实现多源异构数据同步采集、统一关联及高效处理的技术效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470329B_ABST
    Figure CN122470329B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for synchronous acquisition and processing of multi-source heterogeneous data, relating to the field of data processing technology. The method includes: capturing an external timing signal and training a high-stability crystal oscillator to obtain a high-frequency reference clock; constructing a high-precision timestamp; issuing a sampling trigger command; acquiring time-stamped data frames; serializing and encapsulating them according to a preset unified format to form a continuous data stream; configuring arbitrary data in the continuous data stream using a message queue middleware to obtain arbitrary independent message queues; configuring arbitrary message consumers; pulling arbitrary data; and performing mapping and conversion. This addresses the technical problems of timing alignment difficulties and low processing efficiency in multi-source heterogeneous data processing due to inconsistent device clocks and data format differences. By constructing a high-precision time reference and combining it with a message queue middleware, the technical effects of synchronous acquisition, unified association, and efficient processing of multi-source heterogeneous data are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a method and apparatus for synchronous acquisition and processing of multi-source heterogeneous data. Background Technology

[0002] With the development of intelligent vehicles, industrial automation, intelligent equipment testing and Internet of Things technologies, in order to achieve comprehensive perception and accurate evaluation of the operating status of target objects, it is necessary to simultaneously collect multi-source heterogeneous data such as video data, serial port sensor data, controller area network bus data and analog data.

[0003] In existing technologies, multiple independent acquisition devices are often used to acquire different types of data, followed by software-based correlation and fusion processing. However, since each acquisition device typically operates with an independent clock, the lack of a unified time base between different data sources easily leads to inconsistent data timestamps. This makes it difficult to accurately align and correlate multi-source data, affecting the accuracy of data fusion results. Furthermore, different data types differ significantly in interface types, communication protocols, and data formats, increasing the difficulty of unified data processing and collaborative analysis. This makes it difficult to meet the application requirements of real-time acquisition, dynamic processing, and flexible expansion of multi-source heterogeneous data.

[0004] In existing technologies, the independent acquisition method using multiple devices presents technical problems such as difficulty in aligning timing and low processing efficiency in processing multi-source heterogeneous data due to inconsistent device clocks and differences in data formats. Summary of the Invention

[0005] The purpose of this application is to provide a method and apparatus for synchronous acquisition and processing of multi-source heterogeneous data, which solves the technical problems of difficult timing alignment and low processing efficiency in the existing multi-device independent acquisition method for multi-source heterogeneous data processing due to inconsistent device clocks and data format differences.

[0006] In view of the above problems, this application provides a method and apparatus for synchronous acquisition and processing of multi-source heterogeneous data.

[0007] The first aspect of this application provides a method for synchronous acquisition and processing of multi-source heterogeneous data. This method includes: capturing an external timing signal using a field-programmable gate array (FPGA) and training a high-stability crystal oscillator to obtain a high-frequency reference clock, and constructing a high-precision timestamp; the FPGA simultaneously sending sampling trigger commands to at least two types of heterogeneous sensors to acquire time-stamped data frames with the high-precision timestamp; serializing and encapsulating the time-stamped data frames according to a preset unified format to form a continuous data stream; configuring arbitrary data in the continuous data stream using a message queue middleware to obtain arbitrary independent message queues, and configuring arbitrary message consumers for the arbitrary independent message queues; the arbitrary message consumers asynchronously pulling the continuous data stream, parsing the arbitrary data, and introducing a preset target mapping relationship to perform mapping transformation on the arbitrary data.

[0008] Optionally, when there is no external time signal, the field-programmable gate array (FPGA) automatically switches to hold mode. In hold mode, the time counter set inside the FPGA continues to accumulate based on the most recently calibrated high-stability crystal oscillator frequency to maintain the output of the high-precision timestamp. After the external time signal is detected again, the time deviation is calculated and the time counter is gradually adjusted using a moving average filter to smoothly transition the time to the external time signal.

[0009] Optionally, the second pulse signal of an external timer is captured by the field-programmable gate array; the output frequency of the high-stability crystal oscillator is conditioned and phase-locked multiplied based on the second pulse signal to generate the high-frequency reference clock; the time code data of the external timer is parsed, and the high-frequency reference clock is accumulated and counted using the rising edge of the second pulse signal as a trigger to construct the high-precision timestamp; wherein, the high-precision timestamp includes absolute date and time and relative microsecond-level count value.

[0010] Optionally, the heterogeneous sensor includes at least a video acquisition device, a serial communication device, a controller area network bus device, and an analog signal acquisition device.

[0011] Optionally, the arbitrary independent message queue and the arbitrary message consumer are independently allocated to any processing task corresponding to the arbitrary data; the lifecycle of the arbitrary independent message queue is bound to the arbitrary processing task, and the arbitrary independent message queue is automatically revoked and resources are reclaimed when the arbitrary processing task ends and the arbitrary message consumer finishes consuming the data.

[0012] Optionally, a queue deregistration request is sent to the message queue middleware to disconnect the data connection channel between the arbitrary message consumer and the arbitrary independent message queue; the memory buffer occupied by the arbitrary independent message queue is released based on the queue deregistration request, and the identifier of the arbitrary independent message queue is removed from the message queue middleware; a resource reclamation status code is generated and fed back to the task scheduling component, and the life cycle of the arbitrary processing task is confirmed by the task scheduling component.

[0013] Optionally, after the arbitrary message consumer parses the arbitrary data, for multiple pieces of arbitrary data with the same primary key, it retains only the final state in the order of arrival, discards all intermediate states, and outputs the final state.

[0014] Optionally, an independent data buffer is allocated for each primary key; when any message consumer receives any data belonging to the same primary key, the arbitrary data is overwritten into the data buffer; when a preset time threshold is reached or an output trigger signal is received, the arbitrary data in the data buffer is output as the final state, and the data buffer is cleared.

[0015] Optionally, it is determined whether the arbitrary data parsed by the arbitrary message consumer contains data definition language operations; if it does, a source table schema image is created and a logical sequence number is used as a version identifier, and added to the preset target mapping relationship; when processing the subsequent data definition language, the source table schema image corresponding to the target version of the version identifier is matched based on the preset target mapping relationship and a mapping conversion is performed.

[0016] A second aspect of this application provides a synchronous acquisition and processing device for multi-source heterogeneous data. The device comprises: a timestamp construction module, used to capture an external timing signal via a field-programmable gate array (FPGA) and train a high-stability crystal oscillator to obtain a high-frequency reference clock, and construct a high-precision timestamp; a sampling instruction issuing module, used by the FPGA to simultaneously issue sampling trigger instructions to at least two types of heterogeneous sensors to acquire time-stamped data frames with the high-precision timestamp; a data encapsulation module, used to serialize and encapsulate the time-stamped data frames according to a preset unified format to form a continuous data stream; a data configuration module, used to configure any data in the continuous data stream through a message queue middleware to obtain any independent message queue, and configure any message consumer for the arbitrary independent message queue; and a data conversion module, used by the arbitrary message consumer to asynchronously pull the continuous data stream, parse the arbitrary data, and introduce a preset target mapping relationship to perform mapping conversion on the arbitrary data.

[0017] One or more technical solutions provided in this application have at least the following technical effects or advantages: The method provided in this application captures external timing signals using a field-programmable gate array (FPGA) and trains a high-stability crystal oscillator to obtain a high-frequency reference clock, constructing a high-precision timestamp. The FPGA simultaneously sends sampling trigger commands to at least two types of heterogeneous sensors to acquire time-stamped data frames with the high-precision timestamp. The time-stamped data frames are serialized and encapsulated according to a preset unified format to form a continuous data stream. Any data in the continuous data stream is configured using a message queue middleware to obtain arbitrary independent message queues, and arbitrary message consumers are configured for these queues. The arbitrary message consumers asynchronously pull the continuous data stream, parse the arbitrary data, and introduce a preset target mapping relationship to perform mapping transformation on the arbitrary data. This achieves the technical effect of synchronously acquiring, uniformly associating, and efficiently processing multi-source heterogeneous data by constructing a high-precision time reference and combining it with a message queue middleware.

[0018] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating a method for synchronous acquisition and processing of multi-source heterogeneous data provided in this application.

[0021] Figure 2 The schematic diagram of the multi-source heterogeneous data acquisition hardware provided in this application.

[0022] Figure 3 This application provides a schematic diagram of the structure of a device for synchronous acquisition and processing of multi-source heterogeneous data.

[0023] Figure labeling: 11 timestamp construction module, 12 sampling instruction issuance module, 13 data encapsulation module, 14 data configuration module, 15 data conversion module. Detailed Implementation

[0024] This application provides a method and apparatus for synchronous acquisition and processing of multi-source heterogeneous data. It addresses the technical problems of inconsistent timing and low processing efficiency in existing multi-device independent acquisition methods for multi-source heterogeneous data processing due to inconsistent device clocks and data format differences. The method achieves the technical effect of synchronous acquisition, unified association, and efficient processing of multi-source heterogeneous data by constructing a high-precision time base and combining it with message queue middleware.

[0025] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. It should also be noted that, for ease of description, only the parts related to the present invention are shown in the accompanying drawings, not all of them.

[0026] Example 1, as Figure 1 As shown, this application provides a method for synchronous acquisition and processing of multi-source heterogeneous data, the method comprising: A high-frequency reference clock is obtained by capturing external timing signals using a field-programmable gate array and training a high-stability crystal oscillator, and a high-precision timestamp is constructed.

[0027] Furthermore, by capturing external timing signals using a field-programmable gate array (FPGA) and training a high-stability crystal oscillator to obtain a high-frequency reference clock, a high-precision timestamp is constructed. This process includes: when no external timing signal is detected, the FPGA automatically switches to hold mode; in hold mode, a timekeeping counter internally set in the FPGA continues to accumulate based on the most recently calibrated high-stability crystal oscillator frequency, maintaining the output of the high-precision timestamp; after the external timing signal is detected again, the timekeeping deviation is calculated, and the timekeeping counter is gradually adjusted using a moving average filter to smoothly transition the time to the external timing signal.

[0028] Specifically, a Field-Programmable Gate Array (FPGA) is a programmable device whose internal logic circuitry can be configured using a hardware description language. It possesses parallel processing capabilities and deterministic response with extremely low latency. Connecting the FPGA to an external timing unit, such as a GPS or BeiDou BDS time receiver, via input pins captures external timing signals. These signals include second pulse signals and time code data. The second pulse signal is a TTL level signal that generates one rising edge per second, with its rising edge aligned with the integer second of Coordinated Universal Time (UTC), providing a precise second-level time reference. The time code data conforms to the NMEA-0183 protocol or IRIG-B code format and contains absolute date and time information such as the current year, month, day, hour, minute, and second.

[0029] The Field Programmable Gate Array (FPGA) continuously receives the timing signal output from the external timer. When the FPGA does not detect a new second pulse signal within the preset monitoring period, it determines that the external timing signal is interrupted. At this time, it automatically switches to hold mode and uses the local high-stability clock source to continue to maintain the working state of the time base operation, thereby avoiding the timestamp generation process from stopping or jumping due to timing interruption.

[0030] After entering hold mode, the timekeeping counter set internally by the FPGA continues to increment based on the most recently calibrated high-stability crystal oscillator frequency. This timekeeping counter is a continuously incrementing high-precision counting module, and its counting clock originates from the output frequency of the high-stability crystal oscillator after the most recent time synchronization calibration. For example, during normal time synchronization, the actual count value between multiple consecutive second pulses is compared with the theoretical count value to obtain the calibrated crystal oscillator's actual operating frequency, and this frequency parameter is stored in the FPGA's internal register. When the external time synchronization signal is lost, the timekeeping counter continues to increment according to this calibration frequency and, combined with the absolute time reference corresponding to the most recent valid time synchronization moment, calculates the current time, thereby continuously outputting a high-precision timestamp.

[0031] Once the external time synchronization signal is restored, the FPGA receives the second pulse signal and time code data again. It first reads the local time calculated by the current timekeeping counter and compares it with the newly received standard time synchronization time, calculating the time deviation between the two. This time deviation is the timekeeping deviation. For example, if the timekeeping counter calculates 12:00:10.000150, while the standard time given by the external timer is 12:00:10.000000, the timekeeping deviation is 150 microseconds. A moving average filtering method is used to average the timekeeping deviation over multiple consecutive time synchronization cycles, decomposing the total deviation into multiple smaller correction amounts, which are then gradually applied to the timekeeping counter. For example, the 150-microsecond deviation is distributed across the subsequent 10 time synchronization cycles for successive corrections, with only 15 microseconds corrected per cycle, thus gradually bringing the local time closer to the standard time and achieving a smooth regression of the time reference.

[0032] For example, a high-stability crystal oscillator, after time synchronization calibration, outputs at a frequency of 10MHz, corresponding to a counting period of 100ns. The last valid second pulse signal is received at 12:00:00 on June 1, 2025. Then, due to satellite signal obstruction causing a time synchronization interruption, the FPGA enters hold mode, and the timekeeping counter continues to run at a frequency of 10MHz, accumulating 50,000,000 counts during the interruption, and calculating the current time as 12:00:05. When the external time synchronization signal is received again at the 6th second, the FPGA calculates the timekeeping deviation to be +150 microseconds and sets the moving average window length to 10 second periods. Each period corrects by 15 microseconds. After 10 consecutive time synchronization period adjustments, the local time is fully aligned with the standard time again, ensuring that the timestamps do not roll back, jump, or interrupt, guaranteeing the consistency and reliability of the data timeline during long-term continuous acquisition.

[0033] When an external time source temporarily fails due to satellite obstruction, antenna malfunction, network anomaly, or environmental interference, the field-programmable gate array (FPGA) can continue to output a unified time reference by switching to hold mode. This maintains the time continuity and synchronization during the data acquisition process of multi-source heterogeneous sensors. After the time is restored, a smoothing correction mechanism is used to eliminate the time error accumulated during the timekeeping period, avoiding the impact of time abrupt changes on data sorting, time series analysis, and multi-source data fusion. This improves the reliability, stability, and long-term operation capability of the entire synchronous data acquisition.

[0034] Furthermore, the high-frequency reference clock is obtained by capturing an external timing signal and training a high-stability crystal oscillator using a field-programmable gate array (FPGA), and a high-precision timestamp is constructed. This includes: capturing the second pulse signal of an external timer using the FPGA; training and performing phase-locked loop frequency multiplication on the output frequency of the high-stability crystal oscillator based on the second pulse signal to generate the high-frequency reference clock; parsing the time code data of the external timer, and accumulating and counting the high-frequency reference clock using the rising edge of the second pulse signal as a trigger to construct the high-precision timestamp; wherein, the high-precision timestamp includes absolute date and time and a relative microsecond-level count value.

[0035] Specifically, a field-programmable gate array (FPGA) continuously captures the second pulse signal from an external timer. Simultaneously, the FPGA is equipped with a high-precision digital phase-locked loop (DPLL) module. Using the second pulse signal as a reference, the DPLL module "acclimates" the onboard high-stability crystal oscillator (OCXO or TCXO). Acclimation refers to fine-tuning the output frequency of the high-stability crystal oscillator using the external time signal, ensuring that the crystal's clock frequency remains consistently aligned with the standard time over a long period. The FPGA's DPLL module continuously compares the phase difference between the crystal's output frequency and the second pulse signal, dynamically adjusting the crystal's voltage-controlled voltage or frequency divider using a proportional-integral (PI) control algorithm to lock the crystal's output frequency to an integer multiple of the second pulse, such as 10MHz, thus ensuring long-term consistency between the crystal's frequency and the time source. The acclimated crystal oscillator signal is locked to the low-frequency stable signal output by the crystal oscillator through a digital phase-locked loop or an analog phase-locked loop, and the frequency is multiplied to generate a high-frequency reference clock, such as 100MHz to 200MHz, with a corresponding clock period of 5 nanoseconds to 10 nanoseconds, which serves as an internal high-frequency counter.

[0036] The FPGA parses the timecode data from the external timer in real time, retrieving the current absolute date and time. It then uses the rising edge of the second pulse signal as the absolute time reference to trigger an internal high-frequency counter, which accumulates a count in each high-frequency clock cycle. The accumulated value records the high-precision time interval since the last second pulse signal. This accumulated count is then combined with the absolute date and time to form a high-precision timestamp. This high-precision timestamp simultaneously contains the absolute date and time and a relative microsecond-level count value, such as 2025-06-01-12:00:00.0005234, thus ensuring the uniqueness and accuracy of the data's time.

[0037] By capturing external timing signals and constructing high-precision timestamps through FPGA, a unified, continuous, microsecond-level accurate time reference is provided for data collected by multi-source heterogeneous sensors. This ensures that even if the clocks of each sensor are different, the collected data can be aligned in the same time system, achieving accurate synchronous acquisition and avoiding data time misalignment caused by device clock drift or inconsistent timing among multiple acquisition devices.

[0038] The field-programmable gate array simultaneously sends sampling trigger commands to at least two types of heterogeneous sensors to acquire time-stamped data frames with the high-precision timestamp.

[0039] The time-stamped data frames are serialized and encapsulated according to a preset unified format to form a continuous data stream.

[0040] Furthermore, the heterogeneous sensor includes at least a video acquisition device, a serial communication device, a controller area network bus device, and an analog signal acquisition device.

[0041] Specifically, the information acquisition and processing unit uses the RK3588 processor as its control core, for example... Figure 2 As shown, the RK3588 processor is used as the control core of the information acquisition and processing unit. Through industrial standard interfaces such as GMAC, UART, CAN, SPI, IIC, and HDMI, combined with ADUM2582 isolated transceiver chip, YT8531SH / YT9215RB and switch chip, AD7606AD sampling chip, MS9282 / MS9292 / XS9950A video conversion chip, and LK1553C1M1 bus protocol chip, multiple heterogeneous data acquisition channels are formed, including 5 Gigabit Ethernet channels, 4 isolated RS422 channels, 2 isolated CAN bus channels, multi-channel analog signal acquisition, 1553B bus, XGA and PAL video acquisition and VGA output channels, respectively connecting to heterogeneous sensors such as video acquisition devices, serial communication devices, controller area network bus devices and analog signal acquisition devices.

[0042] The functions of each interface are as follows: Four hardware UARTs are expanded into four RS422 channels via isolation chips to realize serial data recording and master control communication; the CAN interface, in conjunction with an isolation chip, completes CAN bus data acquisition; two SPI interfaces are respectively connected to a 1553B protocol chip and an AD7606 sampling chip to realize 1553B bus and multi-channel analog data reading; two sets of RGMII (GMAC) interfaces respectively bring out LAN1 and LAN2~LAN5, a total of 5 network ports, for network data and multi-channel network video stream acquisition; HDMI_RX is connected to an MS9282 to realize XGA image acquisition, PAL video is decoded by XS9950A and then connected to the processor via SPI, and HDMI_TX, in conjunction with MS9292, realizes VGA video output; the IIC interface reads clock chip parameters on the one hand and completes register configuration of MS9282 and MS9292 video conversion chips on the other hand; onboard eMMC Flash is used for local storage of acquired data.

[0043] The FPGA acts as a unified clock synchronization and sampling trigger control unit. Based on a high-precision time base generated by an external timer, it sends synchronization sampling trigger commands to at least two types of heterogeneous sensors. Data collected by each heterogeneous sensor is transmitted to the RK3588 via its corresponding interface for unified reception, buffering, and encapsulation processing. The heterogeneous sensors include at least video acquisition devices, serial communication devices, controller area network (CAN) bus devices, and analog signal acquisition devices. The data collected by the heterogeneous sensors includes: image frames captured by video acquisition devices, such as industrial cameras or network cameras; serial data collected by serial communication devices, such as GPS modules, inertial measurement units, and environmental sensors; CAN bus device node status or control signals; and temperature, voltage, and current signal data collected by analog signal acquisition devices.

[0044] When the FPGA outputs a synchronous sampling trigger signal, the heterogeneous sensors corresponding to each interface collect data according to a unified time reference. After receiving the data returned by each interface, the RK3588 performs unified time stamp association processing on the data collected by the heterogeneous sensors based on the high-precision timestamp information provided by the FPGA, thereby forming a time-stamped data frame. The time-stamped data frame contains the data structure of the data source identifier, collection time information, and original business data, thereby ensuring the synchronization and traceability of multi-source heterogeneous data under a unified time coordinate.

[0045] The RK3588 then performs unified formatting on time-stamped data frames from different heterogeneous sensors. For example, each data frame is uniformly encapsulated as: frame header + data source identifier + timestamp + data type + data length + checksum field. The frame header identifies the start position of the data frame, the data source identifier distinguishes different sensor sources, the timestamp records the sampling time, the data type identifies video data, serial port data, CAN messages, or analog data, the data length identifies the data payload size, and the checksum field ensures transmission integrity. Through serialization encapsulation, a continuous data stream is formed, which refers to a chronological, uninterrupted sequence of bytes.

[0046] By using FPGA to uniformly trigger sampling of various sensors and attach high-precision timestamps, synchronization problems caused by clock drift and inconsistent sampling of individual sensors are avoided. Furthermore, the timestamped data frames are converted into a unified continuous data stream, eliminating data format differences between different acquisition devices, thereby improving the efficiency and accuracy of data synchronization processing.

[0047] By configuring any data in the continuous data stream using message queue middleware, arbitrary independent message queues can be obtained, and arbitrary message consumers can be configured for the arbitrary independent message queues.

[0048] Furthermore, arbitrary independent message queues are obtained by configuring arbitrary data in the continuous data stream through message queue middleware, and arbitrary message consumers are configured for the arbitrary independent message queues. This includes: independently allocating the arbitrary independent message queues and the arbitrary message consumers for any processing task corresponding to the arbitrary data; the lifecycle of the arbitrary independent message queues is bound to the arbitrary processing task, and the arbitrary independent message queues are automatically revoked and resources are reclaimed when the arbitrary processing task ends and the arbitrary message consumer finishes consuming the data.

[0049] Furthermore, automatically deregistering the arbitrary independent message queue and reclaiming resources includes: sending a queue deregistration request to the message queue middleware to disconnect the data connection channel between the arbitrary message consumer and the arbitrary independent message queue; releasing the memory buffer occupied by the arbitrary independent message queue based on the queue deregistration request, and removing the identifier of the arbitrary independent message queue from the message queue middleware; generating a resource reclamation status code and feeding it back to the task scheduling component, and confirming the end of the lifecycle of the arbitrary processing task through the task scheduling component.

[0050] Specifically, message queue middleware is a data transmission component located between the data producer and the data processor. It provides asynchronous message sending, caching, routing, and reliable transmission capabilities, thereby decoupling data acquisition and processing and enabling parallel multi-task processing. For example, message queue middleware can use ZeroMQ, RabbitMQ, DDS, or other data distribution frameworks that support publish-subscribe patterns. The message queue middleware mainly includes messages, queues, producers, and consumers. Messages are the data units to be transmitted; queues are temporary storage containers for messages, following a first-in, first-out (FIFO) principle; producers are components that send messages to queues; and consumers are components that pull messages from queues and process them.

[0051] When a continuous data stream enters the message queue middleware, corresponding processing task configuration parameters are generated according to the preset processing strategy and the processing task configuration requirements generated by the user's instructions. The configuration parameters include task identifier, data source type, data filtering conditions, processing logic, and resource requirement information. Among them, data filtering rules include specific sensor ID, data type, or timestamp range; processing logic type includes video analysis, CAN parsing, or analog filtering; and required resource information includes queue buffer size and number of consumer threads.

[0052] The message queue middleware dynamically filters the continuous data stream according to the task configuration parameters, routes data frames that meet the task filtering rules to the logical channel of the corresponding processing task, and dynamically creates an independent message queue for each processing task. Within the middleware, a unique queue identifier and independent cache space are allocated for temporarily storing data related to that task. Then, based on the processing logic type of the processing task, a message consumer corresponding to the processing task is instantiated. This message consumer can be an independent thread or process, corresponding to an algorithm module, analysis task, or storage module. For example, the video encoding module subscribes to the video frame queue, the CAN signal analysis module subscribes to the CAN message queue, and the analog data processing module subscribes to the analog queue. The message consumer is bound to this independent message queue and asynchronously pulls data from the queue according to preset rules, such as a polling mechanism or a blocking subscription mechanism, and performs processing operations according to the task logic. During the data inflow process, any data in the continuous data stream is routed and matched according to the data filtering rules in the task configuration parameters. Data that meets the conditions is written to the corresponding independent message queue, and is asynchronously pulled and processed by the bound message consumer, thereby achieving data isolation between different processing tasks, asynchronous parallel processing, and task-level resource management.

[0053] The lifecycle of any independent message queue is bound to its corresponding processing task. This lifecycle binding means that the message queue middleware maps the creation, operation, maintenance, and destruction of queue instances to the start, execution, and end states of processing tasks, allowing queue resources to dynamically adjust as task status changes. When the task scheduling component starts a processing task, the message queue middleware synchronously creates an independent message queue corresponding to that task and generates a task status control structure internally, including a task identifier, queue identifier, consumer identifier, message backlog counter, and task completion flag. The message backlog counter records the number of unconsumed messages in the current independent message queue, and the task completion flag indicates whether the processing task has been completed. During task execution, the message queue middleware continuously monitors the processing task status and the message backlog information of the corresponding independent message queue. When a processing task is marked as completed, the message queue middleware first sets the task completion flag to complete, but does not immediately destroy the queue. Instead, it continues to maintain the queue until the message backlog counter reaches zero, confirming that all messages in the independent message queue have been completely consumed by their corresponding message consumers. The message consumer continuously retrieves messages from the queue and performs processing operations through asynchronous retrieval until the queue is empty and no new messages are written.

[0054] When the task completion flag is in the completed state and the message backlog counter is zero, it indicates that any processing task has ended and any message consumer has finished consuming the data. At this time, the message queue middleware automatically cancels any independent message queue and reclaims resources to confirm that the processing task has completed all data processing and has officially terminated.

[0055] Message queue middleware enables flexible routing of any data in a continuous data stream to message queues and consumers configured independently for specific tasks, achieving efficient, flexible, and accurate processing of multi-source heterogeneous data, thereby ensuring data consistency and task isolation, while optimizing resource utilization.

[0056] Furthermore, automatically deregistering the arbitrary independent message queue and reclaiming resources includes: sending a queue deregistration request to the message queue middleware to disconnect the data connection channel between the arbitrary message consumer and the arbitrary independent message queue; releasing the memory buffer occupied by the arbitrary independent message queue based on the queue deregistration request, and removing the identifier of the arbitrary independent message queue from the message queue middleware; generating a resource reclamation status code and feeding it back to the task scheduling component, and confirming the end of the lifecycle of the arbitrary processing task through the task scheduling component.

[0057] Specifically, when the task scheduling component marks a task as completed, the message queue middleware first generates a queue deregistration request through the queue management interface. This request refers to an API call provided by the message queue middleware, used to identify a queue resource release request and trigger a recycling process. It includes the queue identifier to be deregistered, the task identifier, and current message backlog information. Upon receiving the request, the message queue middleware immediately disconnects the message transmission channel between the arbitrary message consumer and the target independent message queue, i.e., it stops producer writes and consumer pulls to ensure that no new data backlog is generated in the independent message queue, while also ensuring the integrity of existing messages in the independent message queue.

[0058] After disconnecting the data channel, the message queue middleware releases all memory buffers occupied by the independent message queue based on the queue deregistration request, including message cache, index table, and queue management metadata. Simultaneously, it deletes the identifier of any independent message queue. The queue cache release process employs a reference message backlog counter mechanism to ensure that all message consumers have completed reading data from the independent message queue and cleared the cache before releasing resources, thus preventing data loss.

[0059] After the queue resource is successfully released, the message queue middleware generates a resource reclamation status code, including the queue identifier, task identifier, reclamation timestamp, and reclamation result flag, and feeds it back to the task scheduling component through the task scheduling interface. Upon receiving the feedback, the task scheduling component confirms that the lifecycle of the corresponding processing task has officially ended, and can trigger subsequent task scheduling or resource reuse operations.

[0060] For example, consider a task named Task_01 with its own message queue, Queue_01, a queue buffer capacity of 1000 data frames, and a message consumer named Consumer_01. When Task_01 finishes processing and all remaining messages in Queue_01 are consumed by Consumer_01, the task scheduling component sends a deregistration request to the middleware. The data channel between Queue_01 and Consumer_01 is disconnected, the 1000-frame buffer space and index table of Queue_01 are released, the queue identifier Queue_01 is removed from the middleware registry, and a resource reclamation status code is generated: Task_01, Queue_01, reclamation successful. The timestamp 2025-06-01-10:12:45 is fed back to the task scheduling component, ultimately confirming the end of Task_01's lifecycle.

[0061] By automatically revoking independent message queues and reclaiming resources, consistent control over the lifecycle of processing tasks and message queues is achieved. This ensures that queue resources exist only during task execution and are automatically released after the task ends and all data is consumed, thereby avoiding long-term resource occupation and improving overall resource utilization efficiency.

[0062] The arbitrary message consumer asynchronously pulls the continuous data stream, parses it to obtain the arbitrary data, and introduces a preset target mapping relationship to perform mapping transformation on the arbitrary data.

[0063] Furthermore, a preset target mapping relationship is introduced to perform mapping transformation on the arbitrary data, including: determining whether the arbitrary data parsed by the arbitrary message consumer contains data definition language operations; if it does, creating a source table schema image and using a logical sequence number as a version identifier, and adding it to the preset target mapping relationship; when processing the subsequent data definition language, matching the source table schema image of the target version corresponding to the version identifier based on the preset target mapping relationship and performing mapping transformation.

[0064] Specifically, any message consumer asynchronously pulls continuous data streams. Based on its binding relationship with an independent message queue, the message consumer retrieves encapsulated, time-stamped data frames from the corresponding queue of the message queue middleware according to a preset retrieval strategy, such as a polling mechanism or a blocking subscription mechanism. Then, the retrieved time-stamped data frames are parsed and deserialized according to a preset unified data structure to extract data such as data source identifier, timestamp, data type, and data length, thereby restoring the original structured data that can be processed by business logic—that is, arbitrary data.

[0065] After obtaining any data, the message consumer uses a preset target mapping relationship to perform mapping transformation on the data. This preset target mapping relationship refers to a set of predefined or dynamically maintained data mapping rules used to describe the field correspondence, type conversion rules, and version dependencies between the source and target data structures. During the mapping transformation process, the message consumer first determines whether the parsed arbitrary data contains Data Definition Language (DDL) operation information. This DDL operation information represents instructions for changing the data source structure, such as adding fields, modifying field types, or deleting fields.

[0066] When any data is determined to contain Data Definition Language (DDL) operation information, a source table schema mirror is created based on the current parsing context. This source table schema mirror records the complete field definitions and hierarchical relationships of the current data structure and is marked with a logical sequence number as a version identifier. This logical sequence number is a monotonically increasing number generated according to the data arrival or change order, used to uniquely identify the data structure evolution version. Then, the source table schema mirror and its version identifier are written into a preset target mapping relationship to form a versioned mapping index.

[0067] In subsequent data processing, when the message consumer receives arbitrary data containing Data Definition Language (DDL) operations again, it matches the version identifier recorded in the preset target mapping relationship to locate the corresponding version of the source table schema mirror, and performs data transformation processing according to the mapping rules between the field structure defined in the mirror and the target data structure. This data transformation process includes field mapping, field type conversion, and structure reorganization, thereby ensuring compatibility and consistency between different versions of the data structure, enabling the data to be stably parsed and processed during structural evolution.

[0068] For example, a message consumer parses a data entry from a continuous data stream. The original data payload is a Controller Area Network (CAN) message structure and includes a Data Definition Language (DML) operation: adding a new field, Temperature. Upon detecting this DML operation, a source table schema mirror, Schema_V1, is generated, identified as schema version one. This source table schema mirror contains the fields {ID, Speed} (vehicle number and speed) and is assigned a logical sequence number (LSN=1001) as the version identifier, which is then written into the mapping table. Then, when subsequent data contains the new field Temperature, schema version one is matched based on LSN=1001, and the new field Temperature is mapped to the target structure schema version two, which includes {ID, Speed, Temperature} (vehicle number, speed, and temperature). This completes the field expansion and data conversion, ensuring compatibility between old and new versions of data within the same processing chain.

[0069] By introducing a versioned source table schema mirroring mechanism based on data definition language operation recognition, the structure version can be automatically recorded and a mapping relationship can be established when the data structure changes. This ensures the compatibility of data between different versions, avoids data parsing failures or processing anomalies caused by field changes, and enables continuous data processing and stable conversion under the condition of structural evolution.

[0070] Furthermore, after the arbitrary message consumer asynchronously pulls the continuous data stream and parses it to obtain the arbitrary data, the process further includes: after the arbitrary message consumer parses the arbitrary data, for multiple pieces of arbitrary data with the same primary key, only the final state is retained according to the arrival order, all intermediate states are discarded, and the final state is output.

[0071] Furthermore, for multiple arbitrary data with the same primary key, only the final state is retained in the order of arrival, and all intermediate states are discarded, including: allocating an independent data buffer for each primary key; when the arbitrary message consumer receives arbitrary data belonging to the same primary key, overwriting the arbitrary data into the data buffer; when a preset time threshold is reached or an output trigger signal is received, outputting the arbitrary data in the data buffer as the final state, and clearing the data buffer.

[0072] Specifically, after any message consumer asynchronously pulls and parses the continuous data stream to obtain arbitrary data, it groups and manages the data according to the primary key. An independent data cache is allocated to each unique primary key to store the latest data corresponding to that primary key. The data cache can adopt a hash table structure or a key-value pair mapping structure, where the primary key serves as the index. The cache stores the latest data frame for that primary key and its high-precision timestamp information. The primary key refers to a field that uniquely identifies the same data entity in the continuous data stream, such as a sensor ID or a business object ID.

[0073] When a message consumer receives any data belonging to the same primary key, it directly overwrites and writes the arbitrary data into the data cache of the corresponding primary key in the order of arrival, replacing the original arbitrary data. This ensures that the cache always saves the final state, which refers to the latest data of the data entity. This process smoothly eliminates intermediate update states during data processing, reduces memory usage, and reduces processing latency.

[0074] Upon reaching a preset time threshold or receiving an output trigger signal, any data in the data buffer is output as the final state, while the buffer is cleared to prepare for the next round of data reception and updates. The preset time threshold represents the maximum time interval the message consumer waits before outputting the final state, such as 10 milliseconds to 3 seconds, determined based on the data update frequency of the specific application. This prevents certain primary keys from never having their final state output due to a prolonged lack of new data. The output trigger signal refers to the active triggering of the final state output through external events, such as GPIO interrupts or timer expiration events.

[0075] By establishing an independent cache for each primary key and implementing an overwrite strategy, we can deduplicate frequently changing data and output the final state, effectively control the amount of data, reduce the load, and ensure that the business processing logic only processes the latest data state, thereby improving the overall efficiency and accuracy of data processing.

[0076] Example 2, based on the same inventive concept as the synchronous acquisition and processing method for multi-source heterogeneous data in the foregoing examples, such as... Figure 3 As shown, this application provides a synchronous acquisition and processing device for multi-source heterogeneous data, wherein the synchronous acquisition and processing device for multi-source heterogeneous data includes: The time stamp construction module 11 is used to capture external timing signals through a field-programmable gate array (FPGA) and train a high-stability crystal oscillator to obtain a high-frequency reference clock, and construct a high-precision time stamp; the sampling instruction issuing module 12 is used for the FPGA to simultaneously issue sampling trigger instructions to at least two types of heterogeneous sensors to obtain time-stamped data frames with the high-precision time stamp; the data encapsulation module 13 is used to serialize and encapsulate the time-stamped data frames according to a preset unified format to form a continuous data stream; the data configuration module 14 is used to configure any data in the continuous data stream through a message queue middleware to obtain any independent message queue, and configure any message consumer for the arbitrary independent message queue; the data conversion module 15 is used for the arbitrary message consumer to asynchronously pull the continuous data stream, parse the arbitrary data, and introduce a preset target mapping relationship to perform mapping conversion on the arbitrary data.

[0077] Furthermore, the timestamp construction module 11 is also used to: automatically switch the field-programmable gate array to hold mode when there is no external time signal; in the hold mode, the time counter set inside the field-programmable gate array continues to accumulate based on the most recently calibrated high-stability crystal oscillator frequency to maintain the output of the high-precision timestamp; after the external time signal is detected again, calculate the time deviation and gradually adjust the time counter using a moving average filtering method to make the time smoothly transition to the external time signal.

[0078] Furthermore, the timestamp construction module 11 is also used to: capture the second pulse signal of an external timer through the field-programmable gate array; refine and phase-locked multiply the output frequency of the high-stability crystal oscillator based on the second pulse signal to generate the high-frequency reference clock; parse the time code data of the external timer, and accumulate and count the high-frequency reference clock using the rising edge of the second pulse signal as a trigger to construct the high-precision timestamp; wherein, the high-precision timestamp includes absolute date and time and relative microsecond-level count value.

[0079] Furthermore, the sampling instruction issuing module 12 is also used to: include at least a video acquisition device, a serial communication device, a controller area network bus device, and an analog signal acquisition device.

[0080] Furthermore, the data configuration module 14 is also used to: independently allocate the arbitrary independent message queue and the arbitrary message consumer to any processing task corresponding to the arbitrary data; the life cycle of the arbitrary independent message queue is bound to the arbitrary processing task, and when the arbitrary processing task ends and the arbitrary message consumer finishes consumption, the arbitrary independent message queue is automatically revoked and resources are reclaimed.

[0081] Furthermore, the data configuration module 14 is also used to: send a queue deregistration request to the message queue middleware to disconnect the data connection channel between the arbitrary message consumer and the arbitrary independent message queue; release the memory buffer occupied by the arbitrary independent message queue based on the queue deregistration request, and remove the identifier of the arbitrary independent message queue from the message queue middleware; generate a resource reclamation status code and feed it back to the task scheduling component, and confirm the end of the life cycle of the arbitrary processing task through the task scheduling component.

[0082] Furthermore, the data conversion module 15 is also used to: after the arbitrary message consumer parses the arbitrary data, for multiple arbitrary data with the same primary key, retain only the final state according to the arrival order, discard all intermediate states, and output the final state.

[0083] Furthermore, the data conversion module 15 is also configured to: allocate an independent data buffer for each primary key; when the arbitrary message consumer receives arbitrary data belonging to the same primary key, overwrite the arbitrary data into the data buffer; when a preset time threshold is reached or an output trigger signal is received, output the arbitrary data in the data buffer as the final state and clear the data buffer.

[0084] Furthermore, the data conversion module 15 is also used to: determine whether the arbitrary data parsed by the arbitrary message consumer contains data definition language operations; if it does, create a source table pattern image and use the logical sequence number as a version identifier, and add it to the preset target mapping relationship; when processing the subsequent data definition language, match the source table pattern image of the target version corresponding to the version identifier based on the preset target mapping relationship and perform mapping conversion.

[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The method and specific example for synchronous acquisition and processing of multi-source heterogeneous data in the foregoing embodiment one are also applicable to the device for synchronous acquisition and processing of multi-source heterogeneous data in this embodiment. Through the foregoing detailed description of the method for synchronous acquisition and processing of multi-source heterogeneous data, those skilled in the art can clearly understand the device for synchronous acquisition and processing of multi-source heterogeneous data in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.

[0086] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0087] Obviously, those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for synchronous acquisition and processing of multi-source heterogeneous data, characterized in that, include: A high-frequency reference clock is obtained by capturing external timing signals using a field-programmable gate array and training a high-stability crystal oscillator, and a high-precision timestamp is constructed. The field-programmable gate array simultaneously sends sampling trigger commands to at least two types of heterogeneous sensors to acquire time-stamped data frames with the high-precision timestamp. The time-stamped data frames are serialized and encapsulated according to a preset unified format to form a continuous data stream; By configuring any data in the continuous data stream through the message queue middleware, any independent message queue can be obtained, and any message consumer can be configured for the arbitrary independent message queue. The arbitrary message consumer asynchronously pulls the continuous data stream, parses it to obtain the arbitrary data, and introduces a preset target mapping relationship to perform mapping transformation on the arbitrary data.

2. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 1, characterized in that, A high-frequency reference clock is obtained by capturing external timing signals using a field-programmable gate array and training a high-stability crystal oscillator, and a high-precision timestamp is constructed. This process includes: When no external timing signal is received, the field-programmable gate array automatically switches to hold mode; In the hold mode, the timekeeping counter set inside the field programmable gate array continues to accumulate based on the most recently calibrated high-stability crystal oscillator frequency, maintaining the output of the high-precision timestamp; After the external time signal is detected again, the timekeeping deviation is calculated and the timekeeping counter is gradually adjusted using a moving average filtering method to smoothly transition the time to the external time signal.

3. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 2, characterized in that, A high-frequency reference clock is obtained by capturing external timing signals using a field-programmable gate array and training a high-stability crystal oscillator, and a high-precision timestamp is constructed, including: The second pulse signal of an external timer is captured by the field-programmable gate array; The high-frequency reference clock is generated by training and phase-locked multiplication of the output frequency of the high-stability crystal oscillator based on the second pulse signal. The time code data of the external timer is parsed, and the high-frequency reference clock is accumulated and counted using the rising edge of the second pulse signal as a trigger to construct the high-precision timestamp; The high-precision timestamp includes absolute date and time and relative microsecond-level count value.

4. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 1, characterized in that, The heterogeneous sensor includes at least a video acquisition device, a serial communication device, a controller area network bus device, and an analog signal acquisition device.

5. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 1, characterized in that, By configuring arbitrary data in the continuous data stream using message queue middleware, arbitrary independent message queues can be obtained, and arbitrary message consumers can be configured for the arbitrary independent message queues, including: The arbitrary message queue and the arbitrary message consumer are independently allocated to any processing task corresponding to the arbitrary data. The lifecycle of any independent message queue is bound to the arbitrary processing task. When the arbitrary processing task ends and the arbitrary message consumer finishes consuming the message, the arbitrary independent message queue is automatically revoked and its resources are reclaimed.

6. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 5, characterized in that, Automatically revoke any independent message queue and reclaim resources, including: Send a queue deregistration request to the message queue middleware to disconnect the data connection channel between the arbitrary message consumer and the arbitrary independent message queue; Based on the queue cancellation request, the memory buffer occupied by the arbitrary independent message queue is released, and the identifier of the arbitrary independent message queue is removed from the message queue middleware; A resource recycling status code is generated and fed back to the task scheduling component, which then confirms the termination of the lifecycle of any processing task.

7. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 1, characterized in that, The arbitrary message consumer asynchronously pulls the continuous data stream and parses it to obtain the arbitrary data. Then, the arbitrary message consumer parses the arbitrary data and, for multiple arbitrary data with the same primary key, retains only the final state according to the arrival order, discards all intermediate states, and outputs the final state.

8. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 7, characterized in that, For multiple pieces of data with the same primary key, only the final state is retained in the order of arrival, and all intermediate states are discarded, including: Allocate a separate data cache for each primary key; When the arbitrary message consumer receives arbitrary data belonging to the same primary key, it overwrites the arbitrary data into the data cache area. When a preset time threshold is reached or an output trigger signal is received, any data in the data buffer is output as the final state, and the data buffer is cleared.

9. The method for synchronous acquisition and processing of multi-source heterogeneous data as described in claim 7, characterized in that, Introducing a preset target mapping relationship to perform mapping transformation on the arbitrary data includes: Determine whether the arbitrary data parsed by the arbitrary message consumer contains a Data Definition Language operation; If included, create a source table schema mirror and use the logical sequence number as the version identifier, and add it to the preset target mapping relationship; When processing the subsequent data definition language, the source table schema mirror of the target version corresponding to the version identifier is matched based on the preset target mapping relationship and the mapping conversion is performed.

10. A device for synchronous acquisition and processing of multi-source heterogeneous data, characterized in that, The step of implementing the synchronous acquisition and processing method for multi-source heterogeneous data according to any one of claims 1 to 9, wherein the synchronous acquisition and processing device for multi-source heterogeneous data comprises: The timestamp construction module is used to capture external timing signals through a field-programmable gate array and train a high-stability crystal oscillator to obtain a high-frequency reference clock, and to construct a high-precision timestamp. The sampling instruction issuing module is used for the field programmable gate array to simultaneously issue sampling trigger instructions to at least two types of heterogeneous sensors to obtain time-stamped data frames with the high-precision timestamp; The data encapsulation module is used to serialize and encapsulate the time-stamped data frames according to a preset unified format to form a continuous data stream; The data configuration module is used to configure any data in the continuous data stream through the message queue middleware to obtain any independent message queue, and to configure any message consumer for the arbitrary independent message queue. The data conversion module is used by the arbitrary message consumer to asynchronously pull the continuous data stream, parse the arbitrary data, and introduce a preset target mapping relationship to perform mapping conversion on the arbitrary data.

Citation Information

Patent Citations

  • Intelligent peripheral control method based on Android box and phase-locked loop synchronization circuit

    CN121461975A

  • High-precision soft tame synchronization and time service method and system

    CN122119627A