Multi-source sensing data processing method and system
By using FPGA and NPU to segment, screen and converge the perceived data in the multi-source sensing data processing system, the problem of excessive CPU burden in the prior art is solved, and more efficient data processing and target positioning are achieved.
Patent Information
- Application Number
- CN202510285752.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, multi-source aware data processing systems face challenges in real-time, data bandwidth and computing power, especially the CPU is overloaded in the data scheduling and caching process.
By segmenting and filtering multi-source aware data on the FPGA module and using NPU for preliminary screening and identification, the data processing burden of the CPU is reduced. The method includes dividing the first perceptual data into the second perceptual data, pushing it to the NPU through the DMA channel of the PCIE Switch for screening, and then coarsely positioning and fusing by the FPGA, and finally generating precisely positioned and perceptual data.
This method greatly reduces the workload of manual processing, reduces the misjudgment rate of target objects in perceived data, and improves computing efficiency by reducing the data storage burden of the CPU.
Smart Images

Figure CN120216433A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a multi-source perception data processing method and system. Background Art
[0002] With the geometric growth of perception information, especially in multi-source visual information processing tasks, the real-time performance, data bandwidth, and computing power of a single embedded processing platform are facing increasingly severe challenges. Effectively leveraging the capabilities of multiple processors such as CPUs, FPGAs, and NPUs, while optimizing the deployment of algorithm modules in each processor, has become the mainstream solution in the industry.
[0003] Generally, deep neural network algorithms that require high computing power but have relatively single and repetitive tasks are executed by NPU cards, such as screening algorithms, recognition algorithms, identification algorithms, etc.; while some complex and customized algorithms that require low computing power but high parallelism are executed by FPGA cards, such as image segmentation algorithms, rough positioning algorithms, fine positioning algorithms, multi-source fusion algorithms, etc.; the CPU executes various data scheduling instructions in the middle. However, the hardware system architecture of multi-source perception information is shown in Figure 1. Among them, the data processed by the NPU and FPGA both need to be scheduled and cached through the CPU memory, making the data processing burden of the CPU relatively heavy. Summary of the Invention
[0004] In view of this, one of the technical problems solved by the embodiments of the present application is to provide a multi-source perception data processing method and system to overcome or alleviate the above-mentioned defects in the prior art.
[0005] The technical solutions provided by the embodiments of the present application are as follows:
[0006] The present invention provides a multi-source perception data processing method, including:
[0007] Step 1: For each piece of first perception data from multiple sources, process it according to the following steps to determine the rough position of the target object on the first perception data:
[0008] Step 11: Based on the CPU, push the first perception data to the FPGA module;
[0009] Step 12: Based on the FPGA module, segment the pushed first perception data to obtain several pieces of second perception data and push them to the NPU through the DMA channel of the PCIE Switch. The resolution of the second perception data is less than that of the first perception data;
[0010] Step 13: Based on the NPU, screen the second perception data pushed, initially screen out the second perception data containing the target object and push it to the FPGA through the DMA channel of the PCIE Switch, and push the second perception data that does not contain the target object initially screened out to the recognition module inside the NPU to make a secondary judgment on the second perception data that does not contain the target object;
[0011] Step 14: Based on the FPGA module, perform rough positioning of the target object on the pushed second perception data containing the target object, obtain the rough position of the target object on the first perception data and obtain the second perception data after rough positioning;
[0012] Step 2: Based on the FPGA, for each of the multiple data sources, fuse the second perception data after rough positioning corresponding to the first perception data of other data sources onto the second perception data after rough positioning corresponding to the first perception data of each data source to obtain the third perception data;
[0013] Step 3: Based on the FPGA, perform precise positioning of the target object on the third perception data corresponding to each data source to obtain the precisely positioned perception data as the fourth perception data;
[0014] Step 4: Based on the FPGA, fuse the fourth perception data corresponding to all data sources to generate the output target object positioning perception data.
[0015] Optionally, the method is applicable to an intelligent acceleration card integrating two high-computing power calculation units, namely NPU and FPGA.
[0016] Optionally, the NPU and the FPGA are interconnected through a PCIE Switch, and the PCIE Switch is embedded with a DMA controller, enabling the NPU and the FPGA to access each other's memory spaces.
[0017] Optionally, the steps for the NPU to read the perception data in the FPGA memory include:
[0018] The NPU sends a request to the CPU to read the data in the FPGA memory;
[0019] The CPU sends a command to transfer data to the DMA engine of the PCIE Switch according to the request to read the data in the FPGA memory;
[0020] The FPGA directly writes data to the memory attached to the NPU through the DMA controller.
[0021] Optionally, in step 14, pushing the second perception data that does not contain the target object initially screened out to the recognition module inside the NPU to make a secondary judgment on the second perception data that does not contain the target object includes:
[0022] The recognition module inside the NPU performs target object recognition on the second perception data that has been preliminarily screened and does not contain the target object, to determine whether the second perception data that does not contain the target object truly does not contain the target object. If so, the second perception data that truly does not contain the target object is discarded. If the second perception data that does not contain the target object truly contains the target object, the perception data containing the target is pushed through the DMA engine channel of the PCIE Switch to the FPGA to re - execute rough positioning, recognition, and fine positioning.
[0023] Optionally, based on the FPGA module, rough positioning of the target object is performed on the pushed second perception data containing the target object, obtaining the rough position of the target object on the first perception data and obtaining the second perception data after rough positioning, including:
[0024] When the perception data is audio data, the target object is tracked through the direction and intensity characteristics of the sound to achieve rough positioning;
[0025] When the perception data is video data, the target object is tracked through various visual characteristics such as color, shape, and texture to achieve rough positioning;
[0026] When the perception data is image data, the physical position of the target object is determined to achieve rough positioning.
[0027] Optionally, when the perception data is image data, the physical position of the target object is determined to achieve rough positioning, specifically including:
[0028] First, coordinate conversion is performed on the target object in the image data among the earth coordinate system, geographical coordinate system, carrier coordinate system, airborne platform coordinate system, and camera coordinate system, and finally the conversion relationship between the geographical coordinate system and the earth coordinate system of the target object is obtained;
[0029] Based on the conversion relationship between the geographical coordinate system and the earth coordinate system of the target object, the longitude, latitude, and altitude of all target points of the target object in the geographical coordinate system are solved.
[0030] Optionally, before determining the physical position of the target object through the target object to achieve rough positioning when the perception data is image data, it also includes: the FPGA assigns a temporary mark to each pixel in each image data, and records the equivalence relationship of the temporary marks in an equivalence table;
[0031] All temporary marks with an equivalence relationship are made equivalent to the minimum value among them, and the connected regions are renumbered in natural number order;
[0032] Based on the connected regions numbered in natural number order, target object detection is performed.
[0033] Optionally, in step 3, based on the FPGA, the third sensing data corresponding to each data source is accurately positioned for the target object, and the accurately positioned sensing data is obtained as the fourth sensing data, including:
[0034] Input the data sources of different modalities into the deep neural network to perform a multi-modal image matching algorithm based on radiation change-insensitive feature transformation.
[0035] The present invention also provides a multi-data-source sensing data processing system for executing the multi-data-source sensing data processing method in the above method embodiment, including: a main board and a heterogeneous acceleration card;
[0036] The main board is mainly composed of a CPU and its corresponding memory;
[0037] The heterogeneous acceleration card is mainly composed of a PCIE Switch card that simultaneously integrates two high-computing-power computing units, namely an NPU and an FPGA, and the memories corresponding to the NPU and the FPGA respectively. The PCIE Switch card is embedded with a DMA controller. Among them, the NPU and the FPGA are interconnected through an internal PCIE bus, and the main board and the heterogeneous acceleration card are interconnected through an external PCIE bus.
[0038] Advantages of the present application:
[0039] In the technical solution of the present application, the multi-source sensing data is mainly segmented by the FPGA in advance, and the segmented multi-source sensing data is screened, and the segmented sensing data is marked and screened to screen out the sensing data that does not contain the target object, greatly reducing the workload of manual processing.
[0040] In addition, the sensing data that does not contain the target object preliminarily screened is further judged by the recognition module of the NPU, greatly reducing the large misjudgment of the target object in the sensing data.
[0041] Finally, the hardware system to which the multi-data-source sensing data processing method provided by the present application is applied is mainly composed of a PCIE Switch card that simultaneously integrates two high-computing-power computing units, namely an NPU and an FPGA, to complete the processing of multi-source sensing data. Among them, the NPU and the FPGA are interconnected through the PCIE Switch. The PCIE Switch is embedded with a DMA controller, enabling the NPU and the FPGA to access each other's memory spaces, realizing direct data interaction without the participation of the CPU's memory, greatly reducing the CPU data storage burden and improving the operation efficiency. Description of the Drawings
[0042] Figure 1A It is a schematic diagram of the steps of a multi-data-source sensing data processing method provided by the present application.
[0043] Figure 1B Schematic diagram of step 1 in a multi-source perception data processing method provided by this application.
[0044] Figure 2 Schematic diagram of an image pixel point in a geographic coordinate system and an aircraft coordinate system provided by this application.
[0045] Figure 3 Schematic diagram of the hardware structure of a multi-source perception data processing system provided by this application. Specific implementation manners
[0046] Implementing any technical solution of the embodiments of this application does not necessarily require achieving all the above advantages simultaneously.
[0047] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art shall fall within the scope protected by the embodiments of this application.
[0048] Some specific embodiments of the embodiments of this application will be described in detail hereinafter with reference to the accompanying drawings in an exemplary rather than restrictive manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale.
[0049] This invention provides a multi-source perception data processing method. It should be noted that this method is mainly used for processing multi-source perception data such as data segmentation, data screening, target tracking, data fusion, etc. Among them, the perception data can be audio data, video data, or image data. In combination with Figure 1A , 1B , taking the perception data as image data as an example, the multi-source perception data processing method provided by this application will be explained in detail. Specifically:
[0050] Step 1: For each piece of first perception data from multiple sources, process it according to the following steps to determine the rough position of the target object on the first perception data.
[0051] Step 11: Based on the CPU, push the first perception data to the FPGA module.
[0052] In one embodiment, the data sources 1 and 2 here are used as the first sensing data, which are the CCD and SAR image data with different imaging principles respectively. The CPU pushes the data of the two data sources to the FPGA module. Among them, CCD is a semiconductor device that converts optical images into electrical signals. It converts light into electrical signals through a charge-coupled device (CCD) and is mainly used to capture visible light images. It has high spatial resolution and rich color information and is suitable for application scenarios that require high details and color recognition, such as urban planning, land use assessment, and target recognition. However, CCD images are limited by weather and lighting conditions and cannot be used in harsh environments. SAR (Synthetic Aperture Radar) images are a technology that obtains images through radar signal reflection and can penetrate weather conditions such as clouds, rain, and snow to achieve all-weather imaging. SAR images mainly record the backscattering information of ground objects and are usually grayscale images without color information. SAR images have important applications in fields such as geological exploration, environmental monitoring, and military reconnaissance, but their spatial resolution is relatively low and data processing is complex. Generally speaking, CCD and SAR images each have their own advantages and limitations. CCD images have high spatial resolution and rich color information under visible light conditions and are suitable for tasks that require high details and color recognition; while SAR images are not affected by weather and can work all-weather, and are suitable for monitoring and reconnaissance tasks in harsh environments, but their spatial resolution is relatively low and data processing is complex.
[0053] Step 12: Based on the FPGA module, the pushed first sensing data is segmented to obtain several pieces of second sensing data and pushed to the NPU through the DMA channel of the PCIE Switch. The resolution of the second sensing data is smaller than that of the first sensing data.
[0054] In one embodiment, the first sensing data is a high-resolution image. For example, the resolutions of data sources 1 and 2 are 50000*10000. Among them, the resolutions of data sources 1 and 2 can be the same or different. The FPGA module cuts the high-resolution image into several small-resolution images for the next-level processing according to the rules set by the algorithm, such as 800*800, that is, several pieces of second sensing data with smaller resolutions are obtained.
[0055] Step 13: Based on the NPU, the pushed second sensing data is screened to initially screen out the second sensing data containing the target object and push it to the FPGA through the DMA channel of the PCIE Switch, and push the second sensing data initially screened out as not containing the target object to the recognition module inside the NPU to perform a secondary judgment on the second sensing data not containing the target object.
[0056] In one embodiment, the CPU pushes several small-resolution images to the NPU in the order set by the algorithm to execute the depth neural network algorithm for screening. The NPU screens out the second perception data containing the target object through the depth neural network algorithm. Among them, the depth neural network steps can be YoloV8 of the CNN class, or Picodet, Nanadet, YoloV3, YoloV5, but are not limited to these networks. New networks will be iterated irregularly following the latest open-source algorithms in the industry to improve the operation efficiency and accuracy. The second perception data that is initially screened out and contains the target object is pushed to the FPGA through the DMA channel of the PCIE Switch. It should be noted that in this process, only the CPU needs to schedule the data, and the FPGA directly writes data to the memory attached to the NPU through DMA, without pre-storing the data of the FPGA in the memory of the CPU and then reading the data into the memory of the NPU, which improves the data interaction efficiency. And the second perception data that is initially screened out and does not contain the target object is pushed to the recognition module inside the NPU to make a secondary judgment on the second perception data that does not contain the target object to determine whether it really does not contain the target object.
[0057] Optionally, in step 13, pushing the second perception data that is initially screened out and does not contain the target object to the recognition module inside the NPU to make a secondary judgment on the second perception data that does not contain the target object includes:
[0058] The recognition module inside the NPU performs target object recognition on the second perception data that is initially screened out and does not contain the target object to determine whether the second perception data that does not contain the target object really does not contain the target object. If so, the second perception data that really does not contain the target object is discarded. If the second perception data that does not contain the target object really contains the target object, the perception data containing the target is pushed to the FPGA through the DMA engine channel of the PCIE Switch to re-perform rough positioning, recognition, and fine positioning.
[0059] With the participation of this step in the recognition step, it reduces the large misjudgment of the target after the previous manual participation.
[0060] Step 14: Based on the FPGA module, perform rough positioning of the target object on the pushed second perception data containing the target object to obtain the rough position of the target object on the first perception data and obtain the second perception data after rough positioning.
[0061] In one embodiment, when the second perception data is image data, before determining the physical position of the target object through the target object to achieve rough positioning, it further includes: assigning a temporary label to each pixel in each image data through the FPGA, and recording the equivalence relationship of the temporary labels in an equivalence table; making all the temporary labels with an equivalence relationship equivalent to the minimum value among them, and renumbering the connected regions in the order of natural numbers; performing target object detection based on the connected regions numbered in the order of natural numbers.
[0062] It should be noted that a fixed constant is used as the binary segmentation threshold to segment the image into a foreground and a background, and a temporary label is assigned to each pixel to distinguish the background from the foreground. Further, the binary threshold at each pixel position is determined according to the local pixel value distribution of the image to achieve adaptive threshold segmentation, perform more detailed marking on different regions of the image, and implement image processing for uneven light or contrast changes. Set the pixel values of the regions in the same equivalence relationship to the minimum value within their regions, and sequentially number each region to determine the final connected region, and perform target object detection based on this connected region.
[0063] Optionally, based on the FPGA module, perform rough positioning of the target object on the second perception data containing the target object pushed, obtain the rough position of the target object on the first perception data, and obtain the second perception data after rough positioning, including:
[0064] When the perception data is image data, determine its physical position through the target object to achieve rough positioning.
[0065] Optionally, when the second perception data is image data, determine its physical position through the target object to achieve rough positioning, specifically including:
[0066] First, perform coordinate transformation between the earth coordinate system, geographic coordinate system, aircraft coordinate system, airborne platform coordinate system, and camera coordinate system on the target object in the image data, and finally obtain the transformation relationship between the target object in the geographic coordinate system and the earth coordinate system;
[0067] Based on the transformation relationship between the target object in the geographic coordinate system and the earth coordinate system, solve the longitude, latitude, and altitude of all target points of the target object in the geographic coordinate system.
[0068] In one embodiment, five basic coordinate systems are required during the calibration process of the collected image data, namely the earth coordinate system (ECEF, earth-centered earth-fixed), geographic coordinate system (NED, north-east-down), aircraft coordinate system (AC, aircraft), airborne platform coordinate system (P, platform), and camera coordinate system (S, sensor). Represents the transformation matrix from the A coordinate system to the B coordinate system.
[0069]
[0070] Among them, [X A Y A Z A T , [X B Y B Z B T Are the coordinates of the same point in the A coordinate system and the B coordinate system.
[0071] Specifically, first calculate the xyz in the aircraft coordinate system (AC, aircraft) (the y direction is the flight direction, for coordinate conversion, pixel size and focal length calculation angle, and then calculate the line-of-sight vector). The airborne platform is connected to the aircraft through shock absorbers. When there is no vibration, the platform coordinate system (P) P0 X P Y P Z P Coincides exactly with the aircraft coordinate system. When the aircraft vibrates, the shock absorbers can effectively reduce the vibration in the pitch and roll directions. If δ pitch And δ roll Are the error angles between the airborne optoelectronic platform and the aircraft in the pitch and roll directions caused by the shock absorbers respectively, then there are:
[0072]
[0073] The camera coordinate system (S) S O -X S Y S Z S The origin is located at the principal point of the imaging system. The external parameter relationship between the camera coordinate system and the airborne optoelectronic platform coordinate system:
[0074]
[0075] Among them, ξ is the Lie algebra representation of the external parameter transformation matrix.
[0076] Then, according to yaw pitch roll, calculate the xyz in the local north-east-up coordinate system (the coordinate system is towards the east, north, and vertically upward, with the aircraft as the origin), as Figure 2 Shown.
[0077] The origin of the aircraft coordinate system (AC) A0-X A Y A Z A Coincides with the origin of the geographic coordinate system, X A And Y A Pointing respectively to the nose and the right wing direction of the carrier aircraft, Z A It is perpendicular to the carrier aircraft and downward within the longitudinal symmetry plane of the carrier aircraft. Its relationship with the geographical coordinate system is shown in formula (4). If the three attitudes of the carrier aircraft are the heading angle ψ, the pitch angle θ, and the roll angle κ, then:
[0078]
[0079] Finally, calculate the blh of the target point based on the blh of the aircraft, that is, the longitude, latitude, and altitude of the target point in the geographical coordinate system (WGS-84 coordinate system).
[0080] The origin of the geographical coordinate system (NED) A0-NED is the position of the carrier aircraft. The N-axis and the E-axis point to the due north and the due east respectively, and the D-axis is perpendicular to the earth ellipsoid and downward. Its relationship with the earth coordinate system is expressed as follows:
[0081]
[0082] Among them, is the radius of curvature of the prime vertical corresponding to the carrier aircraft.
[0083] It should be noted that based on the FPGA module, the rough positioning of the target object is performed on the second perception data containing the target object pushed. Here, taking the second data as audio data and video data as an example, the rough positioning and fine positioning are performed on it, specifically including:
[0084] When the perception data is audio data, the target object is tracked through the direction and intensity characteristics of the sound to achieve rough positioning. Specifically, the audio data can be roughly and roughly positioned through characteristics such as the direction and intensity of the sound. For example, in a sound source localization system, by analyzing the sound signals received by the microphone array, the direction and position of the sound source can be roughly determined. Fine positioning: Filter the audio data to remove noise and retain the audio data corresponding to the frequency, and fine positioning can be achieved.
[0085] When the sensed data is video data, the target object is tracked through various visual features such as color, shape, and texture to achieve rough positioning. Specifically, the video data can be roughly positioned for the target through visual features such as color, shape, and texture. For example, in a video surveillance system, target objects of interest (such as pedestrians, vehicles, etc.) can be quickly identified through methods such as color segmentation and shape matching. The fine positioning of video data can be achieved through more advanced image processing techniques, such as target tracking, contour detection, feature point matching, etc. These techniques can accurately determine information such as the position, size, shape, and posture of the target object in the video frame. For example, based on the number of cameras required, the target positioning technology based on machine vision can be divided into monocular vision positioning, binocular vision positioning, and omnidirectional vision positioning. By using a machine vision system to replace the human eye for detection, identification, counting, etc., the problems existing in manual detection are solved, and the industrial production efficiency is greatly improved.
[0086] Step 2: Based on the FPGA, for each of the multiple data sources, fuse the second sensed data after rough positioning corresponding to the first sensed data of other data sources onto the second sensed data after rough positioning corresponding to the first sensed data of each data source to obtain the third sensed data.
[0087] In one embodiment, after the rough positioning of both Data Source 1 and Data Source 2 is completed respectively, the results are pushed to the fusion processing module inside the FPGA to perform the first-level judgment on the target classification, coordinates, and importance of the image containing the target. For example: if the positioning results of the CCD and the SAR differ by within 10%, the average value is taken; if it exceeds 20%, the positioning result of the SAR is used as the standard.
[0088] After the fusion processing is completed, the third sensed data is generated, and based on the third sensed data, the FPGA tags Data Source 1 and Data Source 2 according to the source, that is, marks them according to the order of receiving the pictures from the host computer and the current processing date and time, such as CCD_1_202410211200, SAR_1_202410211200. Finally, the third sensed data obtained after fusion is pushed to the NPU through the DMA channel of the PCIE Switch.
[0089] Step 3: Based on the FPGA, perform precise positioning of the target object on the third sensed data corresponding to each data source to obtain the precisely positioned sensed data as the fourth sensed data.
[0090] Optionally, in Step 3, based on the FPGA, performing precise positioning of the target object on the third sensed data corresponding to each data source to obtain the precisely positioned sensed data as the fourth sensed data includes:
[0091] The signal sources of different modes are input into the deep neural network to perform a multimodal image matching algorithm based on radiation change insensitive feature transformation.
[0092] In one embodiment, taking the third perception data as image data as an example, the RIFT algorithm aims to solve the problem of insensitivity to radiation changes in multimodal image matching. It uses phase congruence (PC) to detect feature points instead of relying on image intensity or gradient information. Phase congruence is a feature detection method that can capture significant structural information in an image and has good robustness to radiation changes. Specifically,
[0093] 1. Input image data into the deep neural network:
[0094] 1. Input: 1000×1000 image img
[0095] 2. Perform a two-dimensional Fourier transform on the input image to obtain the transformed image imagefft (complex result).
[0096] imagefft=FFT(img)
[0097] 2. Calculate the logGabor filter
[0098] The logGabor filter has two main parameters, one is the scale parameter and the other is the direction parameter. Usually 4 scales and 6 directions are used, so there are a total of 24 logGabor filters. These filters filter imagefft respectively to get the result
[0099] 1. Generate grids x and y. x and y are normalized two-dimensional coordinate matrices with values of (-0.5, 0.5) separated by image length and width. The specific grid data is shown in Table 1 (x coordinate matrix) and Table 2 (y coordinate matrix):
[0100] Table 1
[0101]
[0102] Table 2
[0103]
[0104] 2. Calculate radius
[0105] Use matrix dot multiplication, matrix square root, and two-dimensional inverse Fourier transform with zero frequency components moved to the center:
[0106] 3. Calculate sintheta and costheta
[0107] The arctan of the two-dimensional matrix is used, and the zero-frequency component is moved to the center of the two-dimensional inverse Fourier transform, and then
[0108]
[0109] radius = IFFTSHIFT(radius)
[0110] Perform sine and cosine trigonometric function calculations:
[0111]
[0112] theta = IFFTSHIFT(theta)
[0113] sintheta = sin(theta)
[0114] costheta = cos(theta)
[0115] 4. Calculate the low-pass filter lp
[0116] The two-dimensional inverse Fourier transform that moves the zero-frequency component to the center, matrix dot product, matrix square root, and 30th power are used (cutoff and n are input parameters, cutoff = 0.45, n = 15).
[0117] 5. Calculate the log Gabor filter (nscale and norient are given parameters, nscale = 4, norient = 6).
[0118] (1) f o f and sigmaOnf are calculated using exponential and logarithmic operations according to the given parameters.
[0119]
[0120] (2) Calculate spread using sintheta and costheta: cos, sin, and arctan are used.
[0121] (3) Perform matrix dot product on the two complex matrices of imagefft and filter, and then perform two-dimensional inverse Fourier transform.
[0122] u1 = IFFTSHIFT(x)
[0123] u2 = IFFTSHIFT(y)
[0124]
[0125] Step 4: Based on the FPGA, fuse the fourth perception data corresponding to all information sources to generate the output target object positioning perception data.
[0126] In one embodiment, after the precise positioning of the data sources 1 and 2 is completed respectively, the generated fourth perception data is pushed to the fusion processing module inside the FPGA to perform a second-level judgment on the classification, coordinates, and importance of the target in the image containing the target.
[0127] After the fusion processing is completed, the target object positioning perception data is generated and output.
[0128] Optionally, the method is applicable to an intelligent acceleration card integrating two high-computing-power computing units, namely NPU and FPGA.
[0129] Optionally, the NPU and the FPGA are interconnected through a PCIE Switch. The PCIE Switch is embedded with a DMA controller, enabling the NPU and the FPGA to access each other's memory spaces.
[0130] Optionally, the steps for the NPU to read the perception data in the FPGA memory include:
[0131] The NPU sends a request to the CPU to read the data in the FPGA memory.
[0132] The CPU sends a command to transfer the data to the DMA engine of the PCIE Switch according to the request to read the data in the FPGA memory.
[0133] The FPGA directly writes the data to the memory attached to the NPU through the DMA controller.
[0134] It should be noted that in this process, neither the data generated by the NPU nor the data generated by the FPGA needs to be uploaded to the CPU memory. Instead, only by the CPU issuing a data scheduling command, the FPGA directly writes the data to the memory attached to the NPU through the DMA controller, reducing the memory burden of the CPU and accelerating the operation efficiency.
[0135] The present invention also provides a multi-data-source perception data processing system for executing the multi-data-source perception data processing method in the above method embodiment, as Figure 3 shown, the system mainly includes: a main board and a heterogeneous acceleration card;
[0136] The main board is mainly composed of a CPU and its corresponding memory.
[0137] The heterogeneous acceleration card is mainly composed of a PCIE Switch card integrating two high-computing-power computing units, namely NPU and FPGA, and the memories corresponding to the NPU and the FPGA respectively. The PCIE Switch card is embedded with a DMA controller. Among them, the NPU and the FPGA are interconnected through an internal PCIE bus, and the main board and the heterogeneous acceleration card are interconnected through an external PCIE bus.
[0138] It should be noted that this system integrates two high-computing-power computing units, namely NPU and FPGA, in the same PCIE card; the NPU and FPGA are interconnected through an internal PCIE bus, and the motherboard and the heterogeneous acceleration card are interconnected through an external PCIE bus.
[0139] The NPU and FPGA are interconnected through a PCIE Switch. Among them, the PCIE Switch is embedded with a DMA controller, enabling the NPU and FPGA to access each other's memory space (DDR) through the internal PCIE bus, the upstream port, the downstream port #1, the downstream port 2, and the external PCIE bus, realizing direct data interaction without the participation of the CPU's memory.
[0140] Taking the example of the NPU reading data from the FPGA's memory, the following steps are required:
[0141] 1. The NPU sends a request to the CPU to read data from the FPGA's memory;
[0142] 2. The CPU sends a command to transfer data to the DMA engine of the PCIE Switch;
[0143] 3. The FPGA directly writes data to the memory attached to the NPU through the DMA.
[0144] In addition, the method executable by a multi-source perception data processing system provided in this embodiment is exactly the same as the multi-source perception data processing method provided in the above method embodiment, and will not be elaborated here.
[0145] The following further illustrates the specific implementation of the embodiments of the present application with reference to the accompanying drawings of the embodiments of the present application.
[0146] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for processing multi-source sensing data, characterized in that: These include; Step 1: for each first perception data from multiple sources, process it according to the following steps to determine the rough position of the target object on the first perception data: Step 11: Based on the CPU, push the first sensing data to the FPGA module; Step 12: Based on the FPGA module, the pushed first perception data is segmented to obtain a plurality of second perception data, and the second perception data is pushed to the NPU through the DMA channel of the PCIE Switch, wherein the resolution of the second perception data is smaller than the resolution of the first perception data; Step 13: Based on the NPU, the second perception data pushed is screened to preliminarily screen out the second perception data containing the target object and push it to the FPGA through the DMA channel of the PCIE Switch, and the second perception data that is preliminarily screened out and does not contain the target object is pushed to the recognition module inside the NPU, and a second judgment is made on the second perception data that does not contain the target object; Step 14: Based on the FPGA module, perform coarse positioning of the target object on the pushed second perception data containing the target object, obtain a rough position of the target object on the first perception data, and obtain the coarsely positioned second perception data; Step 2: Based on the FPGA, for each information source from multiple information sources, the second perception data after coarse positioning corresponding to the first perception data of other information sources is fused with the second perception data after coarse positioning corresponding to the first perception data of each information source to obtain third perception data; Step 3: Based on the FPGA, the third sensing data corresponding to each information source is used to accurately locate the target object, and precise positioning sensing data is obtained as fourth sensing data; Step 4: Based on the FPGA, the fourth perception data corresponding to all information sources are integrated to generate output target object positioning perception data.
2. A method for processing multi-source sensing data according to claim 1, characterized in that: The method is applicable to an intelligent acceleration card that integrates two high-computing-power computing units, NPU and FPGA.
3. A method for processing multi-source sensing data according to claim 2, characterized in that: The NPU and FPGA are interconnected through a PCIE Switch, and the PCIE Switch has a built-in DMA controller, so that the NPU and FPGA can access each other's memory space.
4. A method for processing multi-source sensing data according to any one of claims 2 or 3, characterized in that: The step of the NPU reading the perception data in the FPGA memory includes: The NPU sends a request to the CPU to read the FPGA memory data; The CPU sends a command to move data to the DMA engine of the PCIE Switch according to the request to read the FPGA memory data; The FPGA directly writes data to the memory under the NPU through the DMA controller.
5. The method for processing multi-source sensing data according to claim 1, characterized in that: In step 14, the second perception data that does not contain the target object is initially screened out and pushed to the recognition module inside the NPU, and a secondary judgment is performed on the second perception data that does not contain the target object, including: The identification module inside the NPU performs target object identification on the second perception data that is initially screened out and does not contain the target object, so as to determine whether the second perception data that does not contain the target object really does not contain the target object. If so, the second perception data that really does not contain the target object is discarded; if the second perception data that does not contain the target object really contains the target object, the perception data containing the target is pushed to the FPGA through the DMA engine channel of the PCIE Switch to re-execute coarse positioning, identification and fine positioning.
6. A method for processing multi-source sensing data according to claim 1, characterized in that: Based on the FPGA module, performing coarse positioning of the target object on the pushed second perception data containing the target object, obtaining a coarse position of the target object on the first perception data and obtaining the coarsely positioned second perception data, including: When the perception data is audio data, the target object is tracked by the direction and intensity characteristics of the sound to achieve rough positioning; When the perception data is video data, the target object is tracked through multiple visual features such as color, shape, and texture to achieve coarse positioning; When the perception data is image data, the physical position of the target object is determined to achieve coarse positioning.
7. A method for processing multi-source sensing data according to claim 6, characterized in that: When the perception data is image data, the physical position of the target object is determined to achieve rough positioning, which specifically includes: Firstly, coordinate system conversion is performed on the target object in the image data among the earth coordinate system, the geographic coordinate system, the carrier coordinate system, the airborne platform coordinate system and the camera coordinate system, and finally the conversion relationship between the geographic coordinate system and the earth coordinate system of the target object is obtained; Based on the conversion relationship between the geographic coordinate system and the earth coordinate system of the target object, the longitude, latitude and altitude of all target points of the target object in the geographic coordinate system are solved.
8. A method for processing multi-source sensing data according to claim 6, characterized in that: When the perception data is image data, the physical position of the target object is determined by the target object, and before the rough positioning is realized, the method further includes: assigning a temporary mark to each pixel in each image data through the FPGA, and recording the equivalent relationship of the temporary mark in an equivalent table; All temporary marks with equivalent relations are made equivalent to the minimum value among them, and the connected regions are renumbered in natural number order; Target object detection is performed based on connected regions numbered in natural number sequence.
9. The method for processing multi-source sensing data according to claim 1, characterized in that: In the step 3, based on the FPGA, the third perception data corresponding to each information source is used to accurately locate the target object, and the precise positioning perception data is obtained as the fourth perception data, including: The signal sources of different modes are input into the deep neural network to perform a multimodal image matching algorithm based on radiation change insensitive feature transformation.
10. A multi-source sensing data processing system, used to execute a multi-source sensing data processing method according to claims 1 to 9, characterized in that: Includes: motherboard and heterogeneous accelerator card; The mainboard is mainly composed of a CPU and its corresponding memory; The heterogeneous acceleration card is mainly composed of a PCIE Switch card that integrates two high-computing power computing units, NPU and FPGA, and the memories corresponding to the NPU and FPGA respectively, and the PCIE Switch card has an embedded DMA controller, wherein the NPU and FPGA are interconnected through an internal PCIE bus, and the mainboard and the heterogeneous acceleration card are interconnected through the external PCIE bus.