Data processing method and device based on vector processor and custom instruction accelerator

By combining a vector processor and a custom instruction accelerator, the problem of low computational efficiency of GNSS receivers is solved, and efficient processing of large-scale tracking channels of GNSS receivers is achieved, which improves processing efficiency and reduces resource waste.

CN120704738APending Publication Date: 2025-09-26INST OF MICROELECTRONICS CHINESE ACAD OF SCI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510825580.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When processing multiple satellite signals, the computational efficiency of general-purpose processors in existing GNSS receivers is low, resulting in serious waste of computing resources and making it difficult to meet the needs of efficiently processing large-scale tracking channels.

Method used

Using vector processors and custom instruction accelerators, by building target vector processors and custom instruction accelerators, using the tracking engine to generate multi-channel baseband signal correlation values, and using the intermediate frequency sampling clock to achieve data time alignment, the target vector processors and custom instruction accelerators are used to perform parallel and atomic-level accelerated processing of data blocks.

Benefits of technology

It significantly improves the processing efficiency of GNSS receivers, reduces instruction decoding and data memory access overhead, realizes operator acceleration of intra-channel tracking algorithms and parallel processing between channels, and adapts to flexible adjustments for different algorithms and requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704738A_ABST
    Figure CN120704738A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device based on a vector processor and a user-defined instruction accelerator, and relates to the technical field of satellite positioning and navigation. The method comprises the following steps of: extracting components with data-level parallel processing characteristics and components capable of being accelerated through a custom instruction based on calculation-intensive operation in a baseband tracking algorithm, and respectively mapping the components on a vector processor and a custom instruction accelerator to obtain a target vector processor and a target custom instruction accelerator; performing correlation operation on the baseband signals of the plurality of channels by using a tracking engine to generate correlation values of the plurality of channels; dumping the correlation values of the plurality of channels based on the timing period of the intermediate frequency sampling clock to obtain multi-channel correlation values with aligned time; recombining the multi-channel correlation values based on the variable names to obtain a plurality of continuous vector data blocks; and performing tracking acceleration on the plurality of vector data blocks by using a target vector processor and a target custom instruction accelerator to obtain a target tracking result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of satellite positioning and navigation technology, and more specifically, to a data processing method and device based on a vector processor and a custom instruction accelerator. Background Art

[0002] With the development of modern global navigation satellite systems (GNSS), the number of constellation satellites and the addition of navigation signal frequencies have continued to expand, significantly improving GNSS positioning accuracy, reliability, and service coverage. However, to simultaneously track and process more satellite signals, the number of baseband tracking channels required by receivers has exploded.

[0003] Currently, most GNSS receivers still rely on general-purpose CPUs to perform the computational tasks of each tracking loop, processing each channel sequentially. However, the computationally intensive operations in the tracking loop algorithms are difficult to efficiently support within the general-purpose CPU instruction set, resulting in low computational efficiency. Furthermore, the CPU serially processes each channel, requiring repeated instruction decoding and data memory access. These redundant operations waste resources and reduce computational efficiency.

[0004] Public content

[0005] In view of this, the present disclosure provides a data processing method and device based on a vector processor and a custom instruction accelerator.

[0006] One aspect of the present disclosure provides a data processing method based on a vector processor and a custom instruction accelerator, comprising: constructing a target vector processor and a target custom instruction accelerator based on computationally intensive operations included in a baseband tracking algorithm; utilizing a tracking engine to perform correlation operations on baseband signals of multiple channels respectively to generate correlation values ​​of the baseband signals of the multiple channels; based on a timing cycle of an intermediate frequency sampling clock, dumping the correlation values ​​corresponding to the baseband signals of the multiple channels to obtain time-aligned multi-channel correlation values; reorganizing the multi-channel correlation values ​​based on variable names to obtain multiple continuous vector data blocks; utilizing the target vector processor and the target custom instruction accelerator to perform tracking acceleration on the multiple vector data blocks to obtain target tracking results.

[0007] According to an embodiment of the present disclosure, the target vector processor and the target custom instruction accelerator are constructed based on the computationally intensive operations included in the baseband tracking algorithm, including: analyzing the computationally intensive operations included in the baseband tracking algorithm to obtain data-level parallel processing components and custom instruction acceleration components; and constructing a vector instruction set and a custom instruction set based on the data-level parallel processing components and the custom instruction acceleration components, respectively;

[0008] Based on the above-mentioned vector instruction set and the above-mentioned custom instruction set, the microarchitecture of the vector processor and the microarchitecture of the custom instruction accelerator are determined; based on the microarchitecture of the above-mentioned vector processor and the microarchitecture of the above-mentioned custom instruction accelerator, the above-mentioned vector instruction set and the above-mentioned custom instruction set are mapped to the vector processor and the custom instruction accelerator respectively to obtain the above-mentioned target vector processor and the above-mentioned target custom instruction accelerator.

[0009] According to an embodiment of the present disclosure, the above-mentioned analysis of the computationally intensive operations included in the above-mentioned baseband tracking algorithm to obtain data-level parallel processing components and custom instruction acceleration components includes: component extraction of at least one matrix and vector operation with data-level parallel processing characteristics included in the above-mentioned baseband tracking algorithm to obtain the above-mentioned data-level parallel processing components; component extraction of at least one mathematical operation requiring customized acceleration included in the above-mentioned baseband tracking algorithm to obtain the above-mentioned custom instruction acceleration components.

[0010] According to an embodiment of the present disclosure, based on the above-mentioned data-level parallel processing components and the above-mentioned custom instruction acceleration components, a vector instruction set and a custom instruction set are respectively constructed, including: performing component analysis on the above-mentioned data-level parallel processing components to obtain the vector length parameters and element bit width parameters of the above-mentioned vector processor; based on the above-mentioned vector length parameters, the above-mentioned element bit width parameters and the data type and operation type of the above-mentioned baseband tracking algorithm, determining multiple vector instruction subsets that support multi-channel parallelism and matrix operations; based on multiple of the above-mentioned vector instruction subsets, obtaining the above-mentioned vector instruction set.

[0011] According to an embodiment of the present disclosure, the above method also includes: merging the same operations of different algorithms in the above custom instruction acceleration components to obtain a general acceleration instruction subset; performing instruction granularity splitting on the complex operations included in the above baseband tracking algorithm to obtain multiple basic operation combinations, and using them as fine-grained custom acceleration instruction subsets; obtaining the above custom instruction set based on the above general acceleration instruction subset and the above fine-grained custom acceleration instruction subset.

[0012] According to an embodiment of the present disclosure, the above-mentioned dumping of the correlation values ​​corresponding to the baseband signals of the above-mentioned multiple channels based on the timing cycle of the intermediate frequency sampling clock to obtain time-aligned multi-channel correlation values ​​includes: at the triggering moment of each timing cycle of the above-mentioned intermediate frequency sampling clock, the correlation values ​​corresponding to the baseband signals of the above-mentioned multiple channels are synchronously dumped to obtain multi-channel correlation values ​​generated at the same moment.

[0013] According to an embodiment of the present disclosure, a tracking engine is used to perform correlation operations on the baseband signals of multiple channels respectively to generate correlation values ​​of the baseband signals of the above multiple channels, including: performing phase synchronization on the baseband signal of each channel based on a local pseudo-random code to obtain a pseudo-random code synchronization signal of each channel; performing phase synchronization on the pseudo-random code synchronization signal of each channel based on an orthogonal carrier to obtain a baseband orthogonal signal of each channel; performing multiple coherent integrations on the baseband orthogonal signal of each channel within a preset integration period to obtain multiple coherent integration results for each channel; and performing modulus operations on the multiple coherent integration results to obtain correlation values ​​of the baseband signals of the above multiple channels.

[0014] According to an embodiment of the present disclosure, the target vector processor and the target custom instruction accelerator are used to track and accelerate the multiple vector data blocks to obtain target tracking results, including: obtaining at least one single-channel related value from the multiple vector data blocks; using the target custom instruction accelerator to perform atomic-level operations corresponding to the at least one single-channel related value to obtain single-channel tracking results; using the target vector processor to perform loop filtering and loop matrix parallel operations on the multiple vector data blocks respectively to obtain multi-channel tracking results; based on the single-channel tracking results and the multi-channel tracking results, the target tracking results are obtained.

[0015] Another aspect of the present disclosure provides a data processing device based on a vector processor and a custom instruction accelerator, wherein the device includes: a construction module for constructing a target vector processor and a target custom instruction accelerator based on the computationally intensive operations included in the baseband tracking algorithm; an operation module for using a tracking engine to perform correlation operations on the baseband signals of multiple channels respectively, and generate correlation values ​​of the baseband signals of the multiple channels; a synchronization module for dumping the correlation values ​​corresponding to the baseband signals of the multiple channels based on the timing cycle of the intermediate frequency sampling clock, and obtaining time-aligned multi-channel correlation values; a reorganization module for reorganizing the multi-channel correlation values ​​based on the variable names to obtain multiple continuous vector data blocks; a tracking module for using the target vector processor and the target custom instruction accelerator to track and accelerate the multiple vector data blocks to obtain target tracking results.

[0016] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described above.

[0017] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above method when executed.

[0018] Another aspect of the present disclosure provides a computer program product, which includes computer-executable instructions. When the instructions are executed, they are used to implement the method described above.

[0019] According to the embodiments of the present disclosure, a vector processor and a custom instruction accelerator are specifically designed based on the computationally intensive operations in the baseband tracking algorithm. The tracking engine is used to generate multi-channel baseband signal correlation values, and the time-aligned dump of each channel data is realized with the help of the intermediate frequency sampling clock. The data is then reorganized according to the variable name to form a vector data block. Finally, the constructed target vector processor and the custom instruction accelerator are used to perform dual acceleration processing on the data block, which effectively solves the problems of low instruction operation efficiency and slow serial processing speed of traditional general-purpose processors when processing GNSS tracking channels. The operator acceleration of the tracking algorithm within the channel and parallel processing between channels are realized, which greatly improves the processing efficiency of large-scale tracking channels of GNSS receivers, while reducing the instruction decoding and data memory access overhead. The instruction set and parallelism can be flexibly adjusted according to different algorithms and requirements, which significantly enhances the practicality and adaptability of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0021] Figure 1 The flowchart of the data processing method based on the vector processor and the custom instruction accelerator according to the embodiment of the present disclosure is schematically shown;

[0022] Figure 2 Schematically shows a schematic diagram of a data-level parallel processing component and a custom instruction acceleration component according to an embodiment of the present disclosure;

[0023] Figure 3 Schematically shows a schematic diagram of the hardware relationship between the target vector processor and the target custom instruction accelerator and the CPU and the trace engine according to an embodiment of the present disclosure;

[0024] Figure 4 Schematically illustrates a schematic diagram of a tracking channel software and hardware collaborative data flow conditioning mechanism according to an embodiment of the present disclosure;

[0025] Figure 5 A block diagram schematically illustrates a data processing device based on a vector processor and a custom instruction accelerator according to an embodiment of the present disclosure; and

[0026] Figure 6 A block diagram of an electronic device suitable for implementing a data processing method based on a vector processor and a custom instruction accelerator according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0031] With the construction and development of modern global navigation satellite systems (GNSS) such as the Global Positioning System (GPS), Beidou Navigation Satellite System (BDS), and Galileo Navigation Satellite System (Galileo), the number of constellation satellites and the addition of navigation signal frequencies have continued to expand, significantly improving the positioning accuracy, reliability, and service coverage of GNSS.

[0032] However, to simultaneously track and process more satellite signals, the number of baseband tracking channels required by receivers has exploded. In consumer applications, products like smart terminals and in-vehicle navigation devices typically require 96 to 192 channels to meet the demands of multi-system compatibility and multi-frequency reception. In professional applications like surveying and mapping, such as high-precision geodesy and drone aerial surveying, the number of tracking channels climbs to 384 to achieve sub-meter or even centimeter-level positioning accuracy, with some high-end devices even requiring 1,408 channels.

[0033] Currently, most GNSS receivers still rely on general-purpose processors (CPUs) to perform the computational tasks of each tracking loop, and they process each channel sequentially. This processing method has significant drawbacks. On the one hand, tracking loop algorithms such as phase-locked loops (PLLs), frequency-locked loops (FLLs), and Kalman filters (KFs) contain a large number of computationally intensive operations such as multiplication and division, square accumulation, trigonometric functions, and matrix-vector operations. General-purpose processor instruction sets struggle to efficiently support such operations, resulting in low computational efficiency. On the other hand, when the CPU processes each channel serially, it must repeatedly perform instruction decoding and data memory access operations. As the number of channels increases dramatically, the resource waste caused by these redundant operations becomes increasingly serious, resulting in an exponential decrease in computational efficiency when faced with large-scale tracking channel processing tasks.

[0034] To compensate for the lack of computing efficiency, existing solutions often rely on high-performance CPUs or integrated multi-core CPUs, which not only significantly increases hardware costs, but also leads to low processing efficiency, significantly increased device power consumption, and difficulty meeting the needs of portable, low-power application scenarios.

[0035] In view of this, the embodiments of the present disclosure respectively design the vector processor and the custom instruction accelerator in a targeted manner through the computationally intensive operations in the baseband tracking algorithm, use the tracking engine to generate multi-channel baseband signal correlation values, and use the intermediate frequency sampling clock to realize the time-aligned dump of each channel data, and then reorganize the data according to the variable name to form a vector data block. Finally, the constructed target vector processor and the custom instruction accelerator are used to perform dual acceleration processing on the data block, which effectively solves the problems of low instruction operation efficiency and slow serial processing speed of traditional general-purpose processors when processing GNSS tracking channels, realizes the operator acceleration of the tracking algorithm within the channel and parallel processing between channels, greatly improves the processing efficiency of large-scale tracking channels of GNSS receivers, and reduces the instruction decoding and data memory access overhead. In addition, the instruction set and parallelism can be flexibly adjusted according to different algorithms and requirements, which significantly enhances the practicality and adaptability of the method.

[0036] Specifically, an embodiment of the present disclosure provides a data processing method based on a vector processor and a custom instruction accelerator, the method including: constructing a target vector processor and a target custom instruction accelerator based on the computationally intensive operations included in the baseband tracking algorithm; using a tracking engine to perform correlation operations on the baseband signals of multiple channels respectively to generate correlation values ​​of the baseband signals of multiple channels; based on the timing cycle of the intermediate frequency sampling clock, dumping the correlation values ​​corresponding to the baseband signals of multiple channels to obtain time-aligned multi-channel correlation values; reorganizing the multi-channel correlation values ​​based on the variable names to obtain multiple continuous vector data blocks; using the target vector processor and the target custom instruction accelerator to track and accelerate the multiple vector data blocks to obtain target tracking results.

[0037] It should be noted that the data processing method and device based on the vector processor and the custom instruction accelerator determined in the embodiments of the present disclosure can be used in the field of satellite positioning and navigation technology. The data processing method and device based on the vector processor and the custom instruction accelerator determined in the embodiments of the present disclosure can also be used in any field other than the field of satellite positioning and navigation technology, such as the field of wireless communication technology. The application field of the data processing method and device based on the vector processor and the custom instruction accelerator determined in the embodiments of the present disclosure is not limited.

[0038] Figure 1 The flowchart of the data processing method based on the vector processor and the custom instruction accelerator according to the embodiment of the present disclosure is schematically shown.

[0039] like Figure 1 As shown, the method includes operations S101 to S105.

[0040] In operation S101 , a target vector processor and a target custom instruction accelerator are constructed based on computationally intensive operations included in a baseband tracking algorithm.

[0041] In operation S102 , a tracking engine is used to perform correlation operations on baseband signals of the multiple channels to generate correlation values ​​of the baseband signals of the multiple channels.

[0042] In operation S103 , correlation values ​​corresponding to baseband signals of multiple channels are dumped based on a timing cycle of the intermediate frequency sampling clock to obtain time-aligned multi-channel correlation values.

[0043] In operation S104 , the multi-channel related values ​​are reorganized based on the variable names to obtain a plurality of continuous vector data blocks.

[0044] In operation S105 , a target vector processor and a target custom instruction accelerator are used to perform tracking acceleration on the plurality of vector data blocks to obtain a target tracking result.

[0045] In this embodiment, the baseband tracking algorithm is the core processing module of the GNSS receiver and can be used to extract satellite navigation data from the received radio frequency signal. The baseband tracking algorithm may include a phase-locked loop (PLL) / frequency-locked loop (FLL) algorithm, a Kalman filter (KF) algorithm, etc. The phase-locked loop (PLL) / frequency-locked loop (FLL) algorithm can eliminate the effects of Doppler shift by tracking the phase and frequency of the carrier signal, ensuring the accuracy of signal demodulation. For example, the PLL compares the phase difference between the received signal and the local carrier through a phase detector, and then generates an error control signal through a loop filter. The Kalman filter (KF) algorithm can be used to fuse multi-channel tracking data, optimize the phase and frequency prediction of satellite signals, and improve tracking accuracy. For example, the Costas KF algorithm updates state variables in real time through matrix operations.

[0046] In this embodiment, the baseband tracking algorithm includes a large number of computationally intensive operations, where specific types of computationally intensive operations may include, for example, high-frequency mathematical operations, such as multiplication and division, square accumulation, trigonometric functions and other atomic operations, and may also include batch matrix and vector operations, such as state transfer matrix updates in Kalman filtering, multi-channel data parallel processing and other operations.

[0047] In this embodiment, a traditional general-purpose processor (CPU) is difficult to efficiently process such operations due to limitations of its instruction set architecture. Therefore, the embodiments of the present disclosure can customize a dedicated hardware acceleration unit based on the algorithm calculation characteristics.

[0048] Specifically, for operations with data parallelism, such as multi-channel matrix multiplication of the same type, batch data parallel processing can be achieved by designing a vector instruction set and hardware architecture. For atomic operations that occur repeatedly and have low efficiency for general instructions, such as ATAN2 trigonometric functions, single-cycle high-speed calculations can be achieved through a custom instruction set. Based on the vector instruction set and custom instruction set, the microarchitecture design of the vector processor and the target custom instruction accelerator is then carried out, and the target vector processor and target custom instruction accelerator are designed.

[0049] A vector processing unit (VPU) can represent a processor that supports data-level parallelism (DLP), using vector instructions to process multiple data elements simultaneously, making it suitable for batch operations such as matrix and vector operations. A custom instruction accelerator (CIA) can represent a hardware acceleration unit customized for specific algorithms such as trigonometric functions, multiplication and division, improving the efficiency of atomic-level operations through custom instructions.

[0050] In this embodiment, the tracking engine may represent a hardware module in the GNSS receiver that is responsible for baseband signal correlation operations to generate correlation values ​​required for tracking.

[0051] Specifically, the received baseband signals of multiple channels are matched with the locally generated pseudo-random code and carrier signal to determine the phase and frequency offset of the signals and obtain the correlation values ​​of the baseband signals of the multiple channels.

[0052] Among them, each channel can correspond to the signal processing of a satellite, and the correlation value can be expressed as (I, Q), among which the I component can represent the cumulative result of the in-phase carrier component, reflecting the signal amplitude; the Q component can represent the cumulative result of the orthogonal carrier component, reflecting the signal phase offset; the larger the modulus of the correlation value, the higher the matching degree between the received signal and the local signal.

[0053] In this embodiment, the timing of correlation value generation is randomly distributed across channels due to the satellite Doppler effect. Therefore, the disclosed embodiment utilizes a millisecond timing cycle of an intermediate frequency sampling clock, such as a 10 MHz clock, as a trigger signal. All channels simultaneously output correlation values ​​to ensure consistent data stream timing. For example, every 1 ms clock rising edge, all eight channels simultaneously output correlation values ​​to avoid timing confusion.

[0054] In this embodiment, variable names can represent state parameters in the baseband tracking algorithm, such as "PLL error," "FLL frequency," and "KF state vector." Each variable can correspond to a hardware register or memory address. A vector data block can represent a continuous collection of data organized by variable name and can be accessed in batches by vector instructions.

[0055] In a specific embodiment, multi-channel related values ​​are reorganized based on variable names. After reorganization, the I components of all channels are stored continuously, and the Q components of all channels are stored continuously to form a vector data block [I1, I2, ..., In, Q1, Q2, ..., Qn], which is convenient for batch reading of vector instructions.

[0056] In this embodiment, a target vector processor is used to perform inter-channel parallel acceleration on multiple vector data blocks to batch process the same type of operations on multiple channels, and a target custom instruction accelerator is used to perform intra-channel operator acceleration on multiple vector data blocks to perform atomic operation acceleration on single-channel related values, and merge the tracking results of each channel, such as carrier phase and code phase estimation values, to form a target tracking result.

[0057] Based on this, the embodiments of the present disclosure respectively design the vector processor and the custom instruction accelerator based on the computationally intensive operations in the baseband tracking algorithm, use the tracking engine to generate multi-channel baseband signal correlation values, and use the intermediate frequency sampling clock to realize the time-aligned dump of each channel data, and then reorganize the data according to the variable name to form a vector data block. Finally, the constructed target vector processor and the custom instruction accelerator are used to perform dual acceleration processing on the data block, which effectively solves the problems of low instruction operation efficiency and slow serial processing speed of traditional general-purpose processors when processing GNSS tracking channels, realizes the operator acceleration of the tracking algorithm within the channel and parallel processing between channels, greatly improves the processing efficiency of large-scale tracking channels of GNSS receivers, and reduces the instruction decoding and data memory access overhead. In addition, the instruction set and parallelism can be flexibly adjusted according to different algorithms and requirements, which significantly enhances the practicality and adaptability of the method.

[0058] According to an embodiment of the present disclosure, a target vector processor and a target custom instruction accelerator are constructed based on the computationally intensive operations included in the baseband tracking algorithm, including: analyzing the computationally intensive operations included in the baseband tracking algorithm to obtain data-level parallel processing components and custom instruction acceleration components; constructing a vector instruction set and a custom instruction set based on the data-level parallel processing components and the custom instruction acceleration components, respectively; determining the microarchitecture of the vector processor and the microarchitecture of the custom instruction accelerator based on the vector instruction set and the custom instruction set; mapping the vector instruction set and the custom instruction set to the vector processor and the custom instruction accelerator, respectively, based on the microarchitecture of the vector processor and the microarchitecture of the custom instruction accelerator, to obtain a target vector processor and a target custom instruction accelerator.

[0059] In this embodiment, the computationally intensive operations in the baseband tracking algorithm are disassembled and divided into two categories according to their computational characteristics.

[0060] Data-level parallel processing (DLP) refers to operations that can process multiple data elements simultaneously. For example, in multi-channel GNSS tracking, loop filter operations in each channel with identical logic, such as PLL error calculation, can be processed in parallel. In single-channel GNSS tracking, matrix-vector multiplication operations in the KF algorithm, such as the product of the state transfer matrix and the state vector, can be computed element-wise in parallel.

[0061] The Custom Instruction Acceleration (CIA) component can represent atomic operations that are inefficient for traditional general-purpose instruction processing and require customized hardware instruction acceleration. For example, PLL / FLL algorithms contain a large number of multiplication and division operations, square-accumulation, two-quadrant inverse tangent ATAN / four-quadrant inverse tangent ATAN2 trigonometric operations, and square-accumulation operations in related value modulus calculations.

[0062] Figure 2 The diagram schematically shows a data-level parallel processing component and a custom instruction acceleration component according to an embodiment of the present disclosure.

[0063] like Figure 2 As shown in the figure, the baseband tracking algorithm is first divided into two parts: single-channel tracking and multi-channel processing. The single-channel tracking algorithm has two main branches: phase-locked loop / frequency-locked loop (PLL / FLL) tracking and Kalman filter (KF) tracking. PLL / FLL tracking includes algorithms of different loop orders (corresponding to different loop filters) and multiple phase / frequency detector algorithms. KF tracking includes linear KF algorithms of different orders and nonlinear EKF algorithms.

[0064] In this embodiment, 1st to 3rd order PLL loop filters, 1st to 2nd order FLL loop filters, two main PLL phase detectors, four main Costas PLL phase detectors, and three main FLL frequency detector algorithms in the PLL / FLL tracking algorithm are analyzed and common features are extracted. These include the following computationally intensive components that need to be accelerated: multiplication and division involving multiple variables, square accumulation, and ATAN / ATAN2 trigonometric functions. These operations are all CIA components.

[0065] In this embodiment, the KF algorithm and the EKF algorithm in the KF tracking algorithm are analyzed and common features are extracted. The phase detector used by the KF is consistent with that of the CostasPLL and the operations are the same. In addition, the KF also includes the following computationally intensive components that need to be accelerated: matrix and vector multiplication, addition, transposition, and matrix inversion operations. These operations can use vector instructions to operate on multiple matrix elements simultaneously and belong to the DLP components.

[0066] In this embodiment, a multi-channel tracking algorithm, that is, the above-mentioned single-channel tracking processing is performed on each channel of the receiver's multiple tracking channels. The data streams of each channel are independent of each other and have no dependencies. Since the operations to be performed on each channel are the same, multiple channels can be processed simultaneously through vector instructions, which belongs to the DLP component.

[0067] In this embodiment, a data-level parallel processing component (DLP) is used as a basis to construct an instruction set that supports batch data processing to obtain a vector instruction set. For example, corresponding vector instructions can be selected or customized based on the parallelism parameters and instruction type of the DLP component.

[0068] In this embodiment, a custom instruction set is constructed using a custom instruction acceleration component (CIA) as a foundation to build a dedicated hardware instruction set. For example, atomic operations such as ATAN2 within the CIA component can be abstracted into a single hardware instruction. Based on the number and type of operands, the instruction opcode and operand encoding are designed as custom instructions.

[0069] In this embodiment, the constructed vector instruction set and custom instruction set can be mapped to corresponding hardware units respectively to obtain a target vector processor and a target custom instruction accelerator adapted to the GNSS tracking task.

[0070] For example, the target vector processor can be obtained by converting the vector instruction set into the microarchitecture of a vector processing unit (VPU). Specifically, according to the type of vector instruction, the vector arithmetic unit (VAU, to support addition, subtraction, multiplication, and division operations), vector load and load unit (VLSU, to support batch data reading and writing), vector slide unit (VSU, to support data permutation), etc. are integrated. A vector register stack (VRF) is built to store intermediate results during the operation process, and the register bit width matches the vector length (for example, 16 32-bit registers form a 64-bit vector register). The vector load and load unit (VLSU) is designed to perform data access operations; multiple vector arithmetic units (VAUs) are configured to implement vector arithmetic operations; a vector slide unit (VSU) is set up to execute vector permutation instructions; and modules such as an instruction decoder, a main sequencer, and a dedicated register stack are equipped to ensure that the vector processor can accurately and efficiently execute the vector instruction set.

[0071] For example, a target custom instruction accelerator (CIA) can be obtained by converting a custom instruction set into its microarchitecture. Specifically, a register stack is constructed to store intermediate computation results; an arithmetic unit (ALU) is configured to implement the computational logic of the custom instructions; a state machine controls the execution of the ALU operations; and an instruction decoder is configured to fully decode the custom instructions. Furthermore, the custom instruction accelerator is tightly coupled to the processor core (CPU), enabling it to reuse the CPU's decoding and execution modules. The CPU then pre-decodes instructions and forwards them to the custom instruction accelerator, achieving efficient hardware acceleration.

[0072] Figure 3 The figure schematically shows the hardware relationship between the target vector processor and the target custom instruction accelerator, the CPU and the trace engine according to an embodiment of the present disclosure.

[0073] like Figure 3 As shown, the target vector processor microarchitecture includes:

[0074] Vector Register File (VRF) is used by functional units to store and access intermediate calculation results during the calculation process.

[0075] Functional units include: Vector Load / Store Unit (VLSU), which executes various memory access instructions; Vector Arithmetic Unit (VAU), which operates on multiple vector elements simultaneously and executes various vector arithmetic instructions; and Vector Slide Unit (VSU), which executes various vector permutation instructions.

[0076] Main Sequencer: Responsible for tracking the status of executing vector instructions. It sets up a scoreboard to track the processing progress of each vector element in each instruction and distributes instructions to the VLSU, VPU, and VSU.

[0077] Decoder: responsible for fully decoding vector instructions.

[0078] Special register file: Contains all control and status registers (CSRs) required by the vector instruction set.

[0079] like Figure 3 As shown, the target custom instruction accelerator microarchitecture includes:

[0080] Register stack: stores intermediate calculation results.

[0081] Arithmetic unit (ALU): implements the operation of custom instructions.

[0082] State machine: controls the ALU to perform operations.

[0083] Decoder: responsible for fully decoding custom instructions.

[0084] In this specific embodiment, a universal accelerator connection interface is used between the VPU and the CPU, and a large-bit-width data memory interface is expanded to meet the vector memory access requirements of the VPU. The CIA and the CPU are interconnected through a tight pipeline coupling. The CIA is embedded in the CPU decoding and execution levels and can call the arithmetic logic unit in the CPU execution level. The CPU is responsible for pre-decoding all instructions, fully decoding and executing scalar instructions, and forwarding custom instructions and vector instructions. The CPU is interconnected with the program memory and data memory via a bus. The tracing engine completes the relevant operations of the hardware channel, provides the processor with channel-related values, and is interconnected with the processor system composed of the CPU, VPU, and CIA via a bus.

[0085] Based on this, the embodiments of the present disclosure separate the data-level parallel processing components and the custom instruction acceleration components through the analysis of the computationally intensive operations in the baseband tracking algorithm, and based on this, respectively construct targeted vector instruction sets and custom instruction sets, and accurately map them to the vector processor and custom instruction accelerator, converting the computing requirements into hardware acceleration features, overcoming the problems of low instruction efficiency and serious resource waste when traditional general-purpose processors process GNSS tracking channels. On the one hand, the vector processor is used to batch accelerate parallel operations, and on the other hand, the custom instruction accelerator is used to efficiently process atomic-level complex operations, greatly improving the execution efficiency of various calculations in the baseband tracking algorithm.

[0086] According to an embodiment of the present disclosure, the computationally intensive operations included in the baseband tracking algorithm are analyzed to obtain data-level parallel processing components and custom instruction acceleration components, including: extracting components of at least one matrix and vector operation with data-level parallel processing characteristics included in the baseband tracking algorithm to obtain data-level parallel processing components; extracting components of at least one mathematical operation that requires customized acceleration included in the baseband tracking algorithm to obtain custom instruction acceleration components.

[0087] In this embodiment, the baseband tracking algorithm involves numerous matrix and vector operations that exhibit data-level parallel processing characteristics. Taking the Kalman filter (KF) algorithm as an example, when estimating the state of satellite signals, numerous matrix and vector multiplication, addition, transposition, and matrix inversion operations are involved. In a multi-channel tracking scenario, each channel must perform similar KF operations, and these operations are independent of each other, with no data dependencies. For example, when processing satellite signals from 10 channels, the multiplication of the state transfer matrix and the state vector in the KF algorithm for each channel can use the same operational logic. Extracting these common operations creates data-level parallel processing components.

[0088] In this embodiment, the baseband tracking algorithm also includes many mathematical operations that are inefficient due to conventional general-purpose instruction processing and require custom acceleration. In the phase-locked loop (PLL) and frequency-locked loop (FLL) algorithms, a large number of trigonometric operations, such as multiplication and division, square-accumulation, and two-quadrant inverse tangent (ATAN) and four-quadrant inverse tangent (ATAN2), frequently occur during the processing of the loop filter and phase / frequency detector. For example, the four-quadrant inverse tangent (ATAN2) operation requires complex software iteration algorithms on general-purpose processors, resulting in a tedious and time-consuming calculation process. However, in GNSS tracking algorithms, the ATAN2 operation is a critical step in determining signal phase and occurs frequently. Therefore, these operations are isolated and used as custom instruction acceleration components.

[0089] Based on this, the embodiments of the present disclosure analyze the computationally intensive operations in the baseband tracking algorithm and divide them into data-level parallel processing components and custom instruction acceleration components. Matrix and vector operations with data-level parallel characteristics are extracted, so that multi-channel or multi-data element operations that were originally executed serially can be adapted to vector processors for parallel accelerated processing, significantly reducing computation time. The extraction of mathematical operations that require customized acceleration can achieve fast calculations at the hardware level, significantly improving the computing speed compared to the software calculation method of general-purpose processors, thereby accelerating the execution of the entire baseband tracking algorithm.

[0090] According to an embodiment of the present disclosure, a vector instruction set and a custom instruction set are constructed based on data-level parallel processing components and custom instruction acceleration components, respectively, including: performing component analysis on the data-level parallel processing components to obtain vector length parameters and element bit width parameters of the vector processor; determining multiple vector instruction subsets that support multi-channel parallelism and matrix operations based on the vector length parameters, element bit width parameters, and the data type and operation type of the baseband tracking algorithm; and obtaining a vector instruction set based on multiple vector instruction subsets.

[0091] In this embodiment, key parameters of the vector processor are determined based on the operational characteristics of the DLP components.

[0092] In this embodiment, the vector length parameter includes the vector length and the matrix dimension. The vector length parameter can be determined by analyzing the computational scale of the data-level parallel processing component and determining the number of data elements that the vector processor can process at one time. For example, in multi-channel GNSS tracking, if 16 channels of KF state vectors are processed simultaneously, the vector length is set to 16. In matrix operations, if the state transition matrix is ​​3×3 dimensional, the matrix dimension parameter is set to 3×3.

[0093] In this embodiment, the element bit width parameter includes an element bit width range. The element bit width parameter can be determined by determining the bit width range of the vector elements based on the algorithm's accuracy requirements. For example, error calculations in PLL / FLL algorithms require a 32-bit floating-point type (ELEN=32); correlation value modulus calculations can use a 16-bit unsigned integer type (ELEN=16).

[0094] In this embodiment, the data type may include floating point type, integer type, etc., and the operation type may include addition, multiplication, shift, etc.

[0095] In this embodiment, based on the vector length parameter, the element width parameter, and the algorithm characteristics, instructions that support multi-channel parallelism and matrix operations are selected or designed. For example, an instruction subset that supports multi-channel parallelism and matrix operations is selected from a general vector instruction set such as the RISC-V vector extension instruction set.

[0096] In addition, the instruction operand format is determined based on the data type and operation type in the baseband tracking algorithm, and the vector instruction subsets of each function are integrated according to the instruction format specification to form a vector instruction set suitable for GNSS tracking, enabling it to efficiently perform multi-channel parallel operations and matrix operations.

[0097] Among them, the vector instruction set includes, for example, vector arithmetic instructions (addition, subtraction, multiplication and division), vector access instructions (VLSU), matrix operation instructions, vector permutation instructions (VSU), etc., covering the parallel computing requirements of GNSS tracking algorithms.

[0098] Based on this, the disclosed embodiments extract parameters such as vector length, matrix dimension, and element width from data-level parallel processing components, constructing a vector instruction subset based on the data type and operation type of the baseband tracking algorithm, ultimately forming a vector instruction set adapted to GNSS tracking scenarios. This process achieves a precise match between the vector processor and the parallel computing requirements of the algorithm. The vector length parameter supports batch processing of multi-channel data, the element width parameter adapts to different precision requirements, and the instruction subset covers core operations such as matrix multiplication and vector addition, enabling the vector processor to efficiently perform multi-channel parallel acceleration.

[0099] According to an embodiment of the present disclosure, constructing a vector instruction set and a custom instruction set based on a data-level parallel processing component and a custom instruction acceleration component respectively also includes: merging the same operations of different algorithms in the custom instruction acceleration component to obtain a general acceleration instruction subset; performing instruction granularity splitting on the complex operations included in the baseband tracking algorithm to obtain multiple basic operation combinations, and using them as a fine-grained custom acceleration instruction subset; obtaining a custom instruction set based on the general acceleration instruction subset and the fine-grained custom acceleration instruction subset.

[0100] In this embodiment, common acceleration instruction subsets can be created by merging and integrating the common operations involved in different algorithms within the CIA component. This avoids duplication of design, improves instruction reusability, and reduces instruction redundancy. For example, the phase detectors in the PLL / FLL and KF algorithms both use the ATAN2 operation, which can be merged into a single instruction. The square-accumulation operation for multi-channel correlation values ​​is common to both the PLL / FLL and KF algorithms and can also be merged into a single instruction.

[0101] In this embodiment, the complex operations included in the baseband tracking algorithm can be finely divided into instruction granularity. Taking the ATAN2 function as an example, it is represented as a piecewise function, using basic operations such as ATAN and addition to complete it. The separated basic operation instructions are then designed as fine-grained custom acceleration instructions that can be reused for other complex operations.

[0102] In this embodiment, a subset of general acceleration instructions and a subset of fine-grained custom acceleration instructions are integrated to form a complete custom instruction set. General acceleration instructions can be used to process high-frequency atomic operations such as multiplication and accumulation and square root, while fine-grained instructions can be used to process basic operation units such as symbol determination and piecewise function calculation to support the combined invocation of complex operations.

[0103] Based on this, the embodiments of the present disclosure achieve efficient construction of a custom instruction set by optimizing the custom instruction acceleration components. On the one hand, the same operations in different algorithms are merged to avoid repeated instruction design, effectively reducing the complexity of the hardware logic circuit and reducing the chip implementation cost; on the other hand, the complex operations are split into multiple basic operation combinations based on instruction granularity, which improves the reusability and flexibility of the instructions and enables the custom instruction accelerator to respond to diverse computing needs more efficiently. The final custom instruction set can not only quickly execute atomic operations such as trigonometric functions and multiplication and accumulation, but also realize complex operations through a combination of basic instructions. Compared with the traditional general instruction processing method, it significantly improves the execution efficiency of computationally intensive operations in the baseband tracking algorithm, while enhancing the adaptability of the hardware architecture to different algorithms and application scenarios, and provides strong computational acceleration support for GNSS tracking channel processing.

[0104] In a specific embodiment, the data-level parallel processing component and the custom instruction acceleration component can also be directly called from a knowledge base, etc. In addition, the BOC modulation and demodulation instructions of the Galileo satellite navigation can also be directly called to avoid repeated analysis.

[0105] In another specific embodiment, the DLP component and CIA component of the tracking software may be reconfigured using programming methods including but not limited to assembly and Intrinsic programming.

[0106] According to an embodiment of the present disclosure, based on the timing cycle of the intermediate frequency sampling clock, the correlation values ​​corresponding to the baseband signals of multiple channels are dumped to obtain time-aligned multi-channel correlation values, including: at the timing cycle triggering moment of each intermediate frequency sampling clock, the correlation values ​​corresponding to the baseband signals of the multiple channels are synchronously dumped to obtain multi-channel correlation values ​​generated at the same moment.

[0107] In a GNSS receiver, the generation of correlation values ​​for multi-channel baseband signals originally relies on satellite code periods or random triggers, resulting in time offsets in the data of each channel due to the satellite Doppler effect, such as signal frequency offset. Therefore, in an embodiment of the present disclosure, the correlation value dumping times of all channels are forcibly aligned through unified timing using an intermediate frequency sampling clock, providing a unified time base for the subsequent parallel operations of the vector processor, thereby solving the problem of low processing efficiency caused by data asynchrony in traditional solutions.

[0108] In this embodiment, an intermediate frequency sampling clock, such as a 10 MHz crystal oscillator, serves as the reference clock for baseband signal digitization. Its period, such as 100 ns, is significantly shorter than the baseband tracking algorithm's control period, such as 1 ms, thus enabling high-precision synchronization. In one specific embodiment, a 1 ms timing cycle can be used, with each cycle containing 10,000 intermediate frequency clock pulses, to ensure that the multi-channel dump timing error is less than 100 ns, or one clock cycle.

[0109] In this embodiment, on the hardware side, the intermediate frequency clock pulses can be counted by a counter, and a global synchronization signal is generated every accumulated 10,000 pulses, i.e., 1 ms. This signal can simultaneously trigger the relevant value dumping operation of all tracking channels.

[0110] For example, the counter starts at 0 and increases by 1 each time it receives a rising edge of the intermediate frequency clock. When the counter value reaches 9999, a synchronous dump is triggered on the next rising edge of the clock (that is, the 10,000th pulse), ensuring that all channels execute operations on the same hardware clock edge.

[0111] In this embodiment, the correlation values ​​of each tracking channel are stored in hardware registers (such as I / Q component registers and modulus registers). A synchronization signal is broadcast via a bus to the register control modules of all channels, forcing each channel to output register data to the data bus at the same time. For example, when the synchronization signal arrives, the registers of channels 1 through N simultaneously open their tri-state gates and output the correlation values ​​to the shared data bus. All channels are simultaneously dumped at exactly 1.000ms, ensuring consistent timestamps for multi-channel correlation values ​​and forming a time-aligned data stream.

[0112] In this embodiment, the time-aligned multi-channel correlation values ​​can be directly combined into vector data blocks such as [I1, I2, ..., I16, Q1, Q2, ..., Q16]. The vector processor can read them in batches without additional timing adjustment, avoiding the software-level timing calibration operations (consuming CPU cycles) and hardware-level FIFO cache accumulation (increasing storage overhead) caused by data asynchrony in traditional solutions.

[0113] In addition, the synchronized correlation values ​​can be read all at once by the vector access unit of the vector processor, ensuring phase consistency in multi-channel operations. For example, during PLL loop filtering, the error calculation of each channel is based on the correlation values ​​at the same moment, avoiding a decrease in tracking accuracy caused by phase offset.

[0114] Figure 4 The figure schematically shows a tracking channel software and hardware collaborative data flow conditioning mechanism according to an embodiment of the present disclosure.

[0115] like Figure 4As shown, taking parallel processing of N channels as an example, it is assumed that each channel has P variables, including related value variables and various counting values, tracking lock indicators and other variables.

[0116] For the hardware side, such as Figure 4 The upper portion of the dotted line shows a break from the traditional physical channel control mechanism that uses the satellite code period as the dump timing for hardware-related values, resulting in random temporal distribution of data dumps on each channel affected by the Doppler effect. Instead, a new mechanism uses a millisecond timing period driven by the intermediate frequency data sampling clock as the hardware-related value dump timing, with data on each channel dumped simultaneously. This mechanism thus shifts the asynchronous timing of the data streams on each channel to synchronous timing.

[0117] At the same time, the variable registers of each physical channel are organized according to the variable name, and the addresses are continuously distributed. When the synchronous dump moment arrives, the variable data streams of multiple channels are generated synchronously, and the processor accesses a set of variable data in the form of a vector according to the variable name address.

[0118] For software side Figure 4 The lower half of the dotted line shown breaks with the traditional channel variable data flow organization format, which uses channels as structural units and stores all variable data within a channel in a single structure. Instead, it redesigns the data flow organization format, which uses variables as structural units and stores all channel variable data with the same name in a single structure. This reorganizes the same-name variables in each channel that were previously distributed across the storage space into a continuous distribution format. When performing channel operations, vector instructions are used to access the same-name variable data in each channel in vector form according to the variable name, and the same operation is performed for each channel simultaneously. A mechanism is designed to handle changes in channel status and channel quantity during operation, grouping multiple channels in parallel according to channel status to ensure continuous and stable multi-channel DLP processing.

[0119] Based on this, the embodiment of the present disclosure uses the intermediate frequency sampling clock timing period as a trigger reference to achieve synchronous dumping of multi-channel baseband signal correlation values, solving the problem of channel data timing fragmentation caused by the satellite Doppler effect in traditional solutions. This mechanism enables the correlation values ​​of each channel to be generated at the same time, forming a time-aligned data stream, laying the foundation for the vector processor to read multi-channel data in batches. Compared with the traditional asynchronous dump method that requires channel-by-channel timing calibration, this method does not require additional cache storage for asynchronous data, reducing hardware storage resource consumption. At the same time, the time-aligned correlation values ​​ensure the phase consistency of the multi-channel tracking algorithm operation, avoid the reduction of tracking accuracy due to timing deviation, so as to improve the processing efficiency of large-scale tracking channels and provide key timing guarantees for the energy-efficient parallel processing of GNSS receivers.

[0120] According to an embodiment of the present disclosure, a tracking engine is used to perform correlation operations on the baseband signals of multiple channels respectively to generate correlation values ​​of the baseband signals of the multiple channels, including: performing phase synchronization on the baseband signal of each channel based on a local pseudo-random code to obtain a pseudo-random code synchronization signal of each channel; performing phase synchronization on the pseudo-random code synchronization signal of each channel based on an orthogonal carrier to obtain a baseband orthogonal signal of each channel; performing multiple coherent integrations on the baseband orthogonal signal of each channel within a preset integration period to obtain multiple coherent integration results for each channel; and performing modulus operations on the multiple coherent integration results to obtain correlation values ​​of the baseband signals of the multiple channels.

[0121] In this embodiment, because the local pseudorandom code (PRN) has excellent autocorrelation characteristics, its autocorrelation function reaches a peak when phases are aligned. Therefore, in this specific embodiment of the present disclosure, each channel corresponds to a satellite and generates a local PRN code identical to that of the receiving satellite. A delay-locked loop (DLL) is used to adjust the phase of the local PRN code to align it with the PRN code in the received signal. When the local PRN code and the received signal's PRN code are aligned in phase, multiplying them eliminates the effects of spread spectrum, generating a narrowband signal that serves as a pseudorandom code synchronization signal, converting the wideband spread spectrum signal into a narrowband signal and improving the signal-to-noise ratio (SNR) in subsequent carrier processing.

[0122] In this specific embodiment, a local carrier is generated with the same frequency and phase as the received signal and split into an in-phase (I) and a quadrature (Q) carrier. A phase-locked loop (PLL) or frequency-locked loop (FLL) is used to adjust the frequency and phase of the local carrier to align it with the carrier of the received signal. The pseudo-random code synchronization signal for each channel is multiplied by the I and Q carriers, respectively, to achieve down-conversion, resulting in baseband quadrature signals for each channel, such as I / Q signals.

[0123] In this specific embodiment, within a preset integration period, such as 1 ms, the baseband quadrature signal of each channel is accumulated multiple times to obtain multiple coherent integration results of each channel.

[0124] Among them, the preset integration period can be set based on the signal-to-noise ratio and dynamic response capability. For example, a long integration period (such as 20ms) can improve the weak signal detection capability, but has poor adaptability to high-dynamic scenarios (such as high-speed movement); a short integration period (such as 1ms) is suitable for high-dynamic scenarios, but the signal-to-noise ratio gain is limited.

[0125] In this specific embodiment, a modulus calculation is performed on the multiple coherent integration results for each channel to obtain correlation values ​​for the baseband signals of each of the multiple channels. The modulus value indicates the degree of match between the received signal and the local replica. The peak position of the correlation value corresponds to the phase alignment point between the received signal and the local signal; the peak strength reflects the signal strength and is used to determine satellite signal quality.

[0126] Based on this, the embodiment of the present disclosure performs phase synchronization with the baseband signal through a local pseudo-random code, strips off the signal spread spectrum modulation, and converts the broadband signal into a narrowband signal, thereby reducing the complexity for subsequent processing and improving the signal-to-noise ratio. The synchronized signal is further synchronized with the help of an orthogonal carrier, and accurate down-conversion of the RF signal to the baseband signal is achieved, making the signal convenient for digital processing. Then, through multiple coherent integrations within a preset integration period, the signal energy is effectively accumulated, noise interference is suppressed, and the signal-to-noise ratio of the signal is significantly improved. Finally, a modular operation is performed on the integration result to accurately quantify the degree of matching between the signal and the local reference signal, and generate the correlation value of each channel, ensuring that each channel can independently and accurately complete the conversion from the original baseband signal to the key correlation value, providing a high-quality data foundation for subsequent tracking acceleration based on a vector processor and a custom instruction accelerator.

[0127] According to an embodiment of the present disclosure, a target vector processor and a target custom instruction accelerator are used to perform tracking acceleration on multiple vector data blocks to obtain a target tracking result, including: obtaining at least one single-channel related value from the multiple vector data blocks; using the target custom instruction accelerator to perform atomic-level operations corresponding to the at least one single-channel related value to obtain a single-channel tracking result; using the target vector processor to perform loop filtering and loop matrix parallel operations on the multiple vector data blocks respectively to obtain a multi-channel tracking result; and obtaining a target tracking result based on the single-channel tracking result and the multi-channel tracking result.

[0128] In this specific embodiment, the single-channel correlation values ​​extracted from the vector data block can be processed by the target custom instruction accelerator to obtain a single-channel tracking result. Processing methods, such as phase detection and modulus calculation, can offload high-frequency atomic operations such as trigonometric functions and square roots to dedicated hardware, avoiding inefficient loop calculations on general-purpose processors. The single-channel tracking result can represent parameters such as phase error and frequency error obtained by the custom instruction accelerator and be used for loop control of the single channel.

[0129] In this specific embodiment, a target vector processor performs parallel operations on vector data blocks from multiple channels to generate multi-channel tracking results. Parallel operations, such as loop filtering and matrix operations, leverage the SIMD (Simplified Meaning) (SIMD) capabilities of the vector processor to convert traditional serial multi-channel operations into parallel processing, reducing instruction execution cycles. The multi-channel tracking results can represent multi-channel state vectors calculated in parallel by the vector processor, which are used for global navigation solutions.

[0130] The fine control parameters of a single channel are combined with the global state estimation of multiple channels to form the final tracking result.

[0131] Based on this, the embodiments of the present disclosure achieve a precise match between the computing tasks of the baseband tracking algorithm and the hardware capabilities by constructing a heterogeneous hardware collaborative acceleration architecture. The target custom instruction accelerator is used to perform atomic-level operations on single-channel related values, shortening the multi-cycle operations of the traditional software iterative algorithm to a single cycle. At the same time, the target vector processor performs parallel loop filtering and matrix operations on multi-channel vector data blocks, converting the serial processing that originally needed to be performed channel by channel into batch parallel operations. The two work together, giving full play to the hardware advantages of the custom instruction accelerator in special function calculations, and utilizing the SIMD parallel capabilities of the vector processor, significantly enhancing the navigation and positioning performance in high-sensitivity and high-dynamic scenarios.

[0132] Figure 5 The block diagram schematically shows a data processing device based on a vector processor and a custom instruction accelerator according to an embodiment of the present disclosure.

[0133] like Figure 5 As shown, the data processing device 500 based on the vector processor and the custom instruction accelerator includes a construction module 510 , an operation module 520 , a synchronization module 530 , a reorganization module 540 , and a tracking module 550 .

[0134] The construction module 510 is configured to construct a target vector processor and a target custom instruction accelerator based on the computationally intensive operations included in the baseband tracking algorithm. The operation module 520 is configured to utilize the tracking engine to perform correlation operations on the baseband signals of multiple channels to generate correlation values ​​for the baseband signals of the multiple channels.

[0135] The synchronization module 530 is configured to dump the correlation values ​​corresponding to the baseband signals of the multiple channels based on the timing cycle of the intermediate frequency sampling clock to obtain time-aligned multi-channel correlation values.

[0136] The reorganization module 540 is used to reorganize the multi-channel related values ​​based on the variable names to obtain multiple continuous vector data blocks.

[0137] The tracking module 550 is used to track and accelerate multiple vector data blocks using a target vector processor and a target custom instruction accelerator to obtain a target tracking result.

[0138] According to an embodiment of the present disclosure, the construction module 510 includes an analysis submodule, a first construction submodule, and a first mapping submodule.

[0139] The analysis submodule is used to analyze the computationally intensive operations included in the baseband tracking algorithm to obtain data-level parallel processing components and custom instruction acceleration components.

[0140] The first construction submodule is used to construct a vector instruction set and a custom instruction set based on a data-level parallel processing component and a custom instruction acceleration component.

[0141] A micro-architecture determination submodule, for determining the micro-architecture of the vector processor and the micro-architecture of the custom instruction accelerator based on the vector instruction set and the custom instruction set;

[0142] The first mapping submodule is used to map the vector instruction set and the custom instruction set to the vector processor and the custom instruction accelerator respectively based on the microarchitecture of the vector processor and the microarchitecture of the custom instruction accelerator, so as to obtain a target vector processor and a target custom instruction accelerator.

[0143] According to an embodiment of the present disclosure, the operation module 520 includes a first phase synchronization submodule, a second phase synchronization submodule, a coherent integration submodule, and a modulus operation submodule.

[0144] The first phase synchronization submodule is used to perform phase synchronization on the baseband signal of each channel based on the local pseudo-random code to obtain the pseudo-random code synchronization signal of each channel.

[0145] The second phase synchronization submodule is used to perform phase synchronization on the pseudo-random code synchronization signal of each channel based on the orthogonal carrier to obtain a baseband orthogonal signal of each channel.

[0146] The coherent integration submodule is used to perform multiple coherent integrations on the baseband orthogonal signal of each channel within a preset integration period to obtain multiple coherent integration results for each channel.

[0147] The module operation submodule is used to perform module operation on multiple coherent integration results to obtain the correlation values ​​of the baseband signals of multiple channels.

[0148] According to an embodiment of the present disclosure, the synchronization module 530 includes a synchronization dump submodule.

[0149] The synchronous dump submodule is used to synchronously dump the correlation values ​​corresponding to the baseband signals of multiple channels at the timing cycle triggering moment of each intermediate frequency sampling clock, and obtain the multi-channel correlation values ​​generated at the same moment.

[0150] According to an embodiment of the present disclosure, the tracking module 550 includes: a correlation value acquisition submodule, a first operation submodule, a second operation submodule, and a determination submodule.

[0151] The correlation value acquisition submodule is used to obtain at least one single-channel correlation value from multiple vector data blocks.

[0152] The first operation submodule is used to use the target custom instruction accelerator to perform atomic-level operations corresponding to at least one single-channel related value to obtain a single-channel tracking result.

[0153] The second operation submodule is used to use the target vector processor to perform loop filtering and loop matrix parallel operation on multiple vector data blocks respectively to obtain multi-channel tracking results.

[0154] The tracking result determination submodule is used to obtain the target tracking result based on the single-channel tracking result and the multi-channel tracking result.

[0155] According to the embodiments of the present invention, any number of modules, sub-modules, units, and sub-units, or at least part of the functions of any number of them, can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a computer program module, which can perform the corresponding functions when the computer program module is executed.

[0156] For example, any number of the building module 510, the computing module 520, the synchronization module 530, the reassembly module 540, and the tracking module 550 may be combined into a single module / unit / sub-unit, or any one of these modules / units / sub-units may be split into multiple modules / units / sub-units. Alternatively, at least part of the functionality of one or more of these modules / units / sub-units may be combined with at least part of the functionality of other modules / units / sub-units and implemented in a single module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the building module 510, the computing module 520, the synchronization module 530, the reassembly module 540, and the tracking module 550 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of software, hardware, and firmware, or any suitable combination of any of these. Alternatively, at least one of the construction module 510 , the operation module 520 , the synchronization module 530 , the reorganization module 540 , and the tracking module 550 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0157] It should be noted that the data processing device part based on the vector processor and the custom instruction accelerator in the embodiment of the present disclosure corresponds to the data processing method part based on the vector processor and the custom instruction accelerator in the embodiment of the present disclosure. The description of the data processing device part based on the vector processor and the custom instruction accelerator specifically refers to the data processing method part based on the vector processor and the custom instruction accelerator, which will not be repeated here.

[0158] Figure 6 A block diagram of an electronic device suitable for implementing a data processing method based on a vector processor and a custom instruction accelerator according to an embodiment of the present disclosure is schematically shown. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0159] like Figure 6As shown, the electronic device according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory ROM 602 or a program loaded from a storage portion 608 into a random access memory RAM 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for cache purposes. The processor 601 may include a single processing unit or multiple processing units for executing different actions of the data processing method flow based on a vector processor and a custom instruction accelerator according to an embodiment of the present disclosure.

[0160] Various programs and data required for the operation of the electronic device are stored in RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the programs in ROM 602 and / or RAM 603 to perform various operations of the method flow according to the embodiment of the present disclosure. It should be noted that the programs may also be stored in one or more memories other than ROM 602 and RAM 603. The processor 601 may also execute the programs stored in the one or more memories to perform various operations of the method flow according to the embodiment of the present disclosure.

[0161] According to an embodiment of the present disclosure, the electronic device may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device may further include one or more of the following components connected to the I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in the drive 610 as needed, so that computer programs read from the removable media can be installed in the storage section 608 as needed.

[0162] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.

[0163] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently without being incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, implements the data processing method based on the vector processor and the custom instruction accelerator according to the embodiments of the present disclosure.

[0164] According to embodiments of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0165] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603 .

[0166] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the data processing method based on the vector processor and the custom instruction accelerator provided by the embodiment of the present disclosure.

[0167] When the computer program is executed by the processor 601, the above functions defined in the system / device of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0168] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0169] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, and all of these combinations and / or couplings fall within the scope of the present disclosure.

[0171] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A data processing method based on a vector processor and a custom instruction accelerator, wherein: The method comprises: Build a target vector processor and target custom instruction accelerator based on the computationally intensive operations included in the baseband tracking algorithm; performing correlation operations on the baseband signals of the multiple channels respectively using a tracking engine to generate correlation values ​​of the baseband signals of the multiple channels; Based on a timing cycle of an intermediate frequency sampling clock, dumping correlation values ​​corresponding to the baseband signals of the plurality of channels to obtain time-aligned multi-channel correlation values; Reorganizing the multi-channel related values ​​based on variable names to obtain a plurality of continuous vector data blocks; The target vector processor and the target custom instruction accelerator are used to track and accelerate the multiple vector data blocks to obtain a target tracking result.

2. The method according to claim 1, wherein The target vector processor and target custom instruction accelerator are constructed based on the computationally intensive operations included in the baseband tracking algorithm, including: Analyzing computationally intensive operations included in the baseband tracking algorithm to obtain data-level parallel processing components and custom instruction acceleration components; Based on the data-level parallel processing component and the custom instruction acceleration component, construct a vector instruction set and a custom instruction set respectively; Determining a microarchitecture of a vector processor and a microarchitecture of a custom instruction accelerator based on the vector instruction set and the custom instruction set; Based on the micro-architecture of the vector processor and the micro-architecture of the custom instruction accelerator, the vector instruction set and the custom instruction set are mapped to the vector processor and the custom instruction accelerator respectively to obtain the target vector processor and the target custom instruction accelerator.

3. The method according to claim 2, wherein: The analysis of computationally intensive operations included in the baseband tracking algorithm to obtain data-level parallel processing components and custom instruction acceleration components includes: Extracting components of at least one matrix and vector operation with data-level parallel processing characteristics included in the baseband tracking algorithm to obtain the data-level parallel processing components; Component extraction is performed on at least one mathematical operation that requires customized acceleration and is included in the baseband tracking algorithm to obtain the customized instruction acceleration component.

4. The method according to claim 2, wherein: Based on the data-level parallel processing component and the custom instruction acceleration component, a vector instruction set and a custom instruction set are constructed respectively, including: Performing component analysis on the data-level parallel processing components to obtain a vector length parameter and an element bit width parameter of the vector processor, wherein the vector length parameter includes a vector length and a matrix dimension, and the element bit width parameter includes an element bit width range; Determining a plurality of vector instruction subsets supporting multi-channel parallelism and matrix operations based on the vector length parameter, the element bit width parameter, and the data type and operation type of the baseband tracking algorithm; The vector instruction set is obtained based on a plurality of the vector instruction subsets.

5. The method according to claim 4, wherein The method further comprises: Merging the same operations of different algorithms in the custom instruction acceleration components to obtain a general acceleration instruction subset; Performing instruction granularity splitting on the complex operations included in the baseband tracking algorithm to obtain multiple basic operation combinations, and using them as fine-grained custom acceleration instruction subsets; The custom instruction set is obtained based on the general acceleration instruction subset and the fine-grained custom acceleration instruction subset.

6. The method according to claim 1, wherein The method of dumping the correlation values ​​corresponding to the baseband signals of the plurality of channels based on the timing cycle of the intermediate frequency sampling clock to obtain time-aligned multi-channel correlation values ​​includes: At each timing cycle triggering moment of the intermediate frequency sampling clock, the correlation values ​​corresponding to the baseband signals of the multiple channels are synchronously dumped to obtain multi-channel correlation values ​​generated at the same moment.

7. The method according to claim 1, wherein Performing correlation operations on the baseband signals of the multiple channels respectively using a tracking engine to generate correlation values ​​of the baseband signals of the multiple channels, including: Performing phase synchronization on the baseband signal of each channel based on a local pseudo-random code to obtain a pseudo-random code synchronization signal of each channel; Performing phase synchronization on the pseudo-random code synchronization signal of each channel based on an orthogonal carrier wave to obtain a baseband orthogonal signal of each channel; Performing multiple coherent integrations on the baseband quadrature signal of each channel within a preset integration period to obtain multiple coherent integration results for each channel; Modulo operation is performed on the multiple coherent integration results to obtain correlation values ​​of the baseband signals of the multiple channels.

8. The method according to claim 1, wherein Tracking and accelerating the plurality of vector data blocks using the target vector processor and the target custom instruction accelerator to obtain a target tracking result includes: Obtaining at least one single-channel correlation value from the plurality of vector data blocks; Utilizing the target custom instruction accelerator to execute an atomic operation corresponding to the at least one single-channel related value to obtain a single-channel tracking result; Using the target vector processor to perform loop filtering and loop matrix parallel operation on the multiple vector data blocks respectively to obtain a multi-channel tracking result; The target tracking result is obtained based on the single-channel tracking result and the multi-channel tracking result.

9. A data processing device based on a vector processor and a custom instruction accelerator, wherein: The device comprises: Building blocks for building a target vector processor and a target custom instruction accelerator based on computationally intensive operations included in a baseband tracking algorithm; an operation module, configured to perform correlation operations on the baseband signals of the multiple channels respectively using a tracking engine to generate correlation values ​​of the baseband signals of the multiple channels; A synchronization module, configured to dump the correlation values ​​corresponding to the baseband signals of the plurality of channels based on a timing cycle of an intermediate frequency sampling clock, to obtain time-aligned multi-channel correlation values; a reorganization module, configured to reorganize the multi-channel related values ​​based on variable names to obtain a plurality of continuous vector data blocks; The tracking module is used to use the target vector processor and the target custom instruction accelerator to track and accelerate the multiple vector data blocks to obtain a target tracking result.

10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.