Data processing device and method
By designing a data processing device in a digital signal processing chip and using control signals to determine the processing method of data to be processed, the problems of bandwidth bottleneck and timing control complexity in high-resolution display scenarios are solved, and data processing with high throughput and low latency are achieved.
Patent Information
- Application Number
- CN202510401643.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-13
AI Technical Summary
The post-processing module of the digital signal processing chip needs to solve the bandwidth bottleneck problem through fusion or splitting functions in high-resolution display scenarios, resulting in a significant increase in timing control complexity and increased latency.
A data processing device is designed, including a control module and a coding module, and the method of splitting or fusion processing of the to-process data is determined by generating a control signal. The encoding module receives the to-process data and control signals, converts the data into multiple compressed slices, and performs a number-fetch operation periodically based on the control signal to achieve processing.
While reducing power consumption and area overhead, high throughput and low latency data processing is achieved, avoiding the delay problems caused by splitting or fusion processing in the prior art.
Smart Images

Figure CN120151541A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of display processing, and in particular, to a data processing apparatus and method. Background Art
[0002] The post-processing module (outctrl) of a digital signal processing chip (Digital Signal Processing, DSP) needs to integrate or split multi-channel data through a merge or split function to solve the bandwidth bottleneck problem in high-resolution display scenarios.
[0003] However, whether it is the data fusion processing of the merge module or the data splitting processing of the split module, both rely on long-term data caching processing, resulting in a significant increase in the complexity of timing control. Therefore, optimizing the latency caused by merging or splitting has become a challenge for the post-processing of digital signal processing chips. Summary of the Invention
[0004] In view of this, embodiments of the present disclosure provide a data processing apparatus and method to achieve high throughput and low latency while reducing power consumption and area overhead.
[0005] The technical solution of the present invention is implemented as follows:
[0006] Embodiments of the present disclosure disclose a data processing apparatus, including: a control module and an encoding module; wherein, the control module is configured to generate a control signal; wherein, the control signal is used to determine whether to perform split processing or fusion processing on the data to be processed; the encoding module is connected to the control module and is configured to receive the data to be processed and the control signal, convert the data to be processed into multiple compressed slices and then cache them, perform a data fetch operation on the cached data in cycles based on the control signal, and output the result of the data fetch operation.
[0007] In the above solution, the control signal includes: a fusion control signal and a split control signal; the encoding module is configured to perform the data fetch operation once per cycle when the fusion control signal is in an enabled state; wherein, each data fetch operation fetches the cached data of multiple compressed slices; or, when the split control signal is in an enabled state, perform multiple data fetch operations in parallel per cycle; wherein, each data fetch operation fetches the cached data of corresponding partial compressed slices among multiple compressed slices.
[0008] In the above solution, the encoding module includes: an output unit; the output unit is configured to, when the fusion control signal is in an enabled state, select any one of all the slice multiplexers of the output unit to perform the data fetching operation, and output the result of the data fetching operation through the output port connected to the slice multiplexer; or, when the splitting control signal is in an enabled state, select a plurality of the slice multiplexers from all the slice multiplexers of the output unit to perform the data fetching operation in parallel, and output the results of the data fetching operation in parallel through a plurality of the output ports connected to the plurality of slice multiplexers.
[0009] In the above solution, each of the slice multiplexers is connected to the control module and is configured to receive the control signal, generate a request signal and adjust the assignment of the request signal based on the control signal, and perform the data fetching operation according to the block size based on the adjusted assignment of the request signal; wherein, the assignment of the request signal is used to select the compressed slice corresponding to each data fetching operation among a plurality of the compressed slices; the data capacity of the block size is smaller than the data capacity of the compressed slice.
[0010] In the above solution, the output unit further includes: a first routing subunit; wherein, the first routing subunit is respectively connected to the control module and all the slice multiplexers, and is configured to, when the fusion control signal is in an enabled state, transmit the data of a plurality of the compressed slices cached to any one of the slice multiplexers selected by the fusion control signal; or, when the splitting control signal is in an enabled state, transmit the data of a plurality of the compressed slices cached to a plurality of the slice multiplexers selected by the splitting control signal respectively.
[0011] In the above solution, the encoding module further includes: an input unit; the input unit is configured to, when the fusion control signal is in an enabled state, select a plurality of input ports from all the input ports of the input unit to respectively receive a plurality of sub-data of the data to be processed, and decompose the sub-data received correspondingly through the slice unit connected to each input port; or, when the splitting control signal is in an enabled state, select any one of all the input ports of the input unit to receive the data to be processed, and select a plurality of the slice units from all the slice units of the input unit to decompose the data to be processed.
[0012] In the above solution, the input unit includes: a second routing subunit; the second routing subunit is respectively connected to the control unit and all the input ports, and is configured to, when the fusion control signal is in an enabled state, respectively transmit the sub-data received by each input port to the slice unit connected to the input port; or, when the split control signal is in an enabled state, transmit the data to be processed to multiple slice units selected by the split control signal.
[0013] In the above solution, the encoding module further includes: an encoding and compression unit; the encoding and compression unit is respectively connected to the input unit and the output unit, and is configured to convert the data received by each slice unit into a corresponding compressed slice, and cache the data of multiple compressed slices through multiple bitrate buffer units in the encoding and compression unit.
[0014] In the above solution, the control module includes: a first register and a second register; wherein, the first register is configured to generate the fusion control signal in an enabled state when performing the fusion process on the data to be processed; the second register is configured to generate the split control signal in an enabled state when performing the split process on the data to be processed.
[0015] The present disclosure also provides a data processing method, including: obtaining data to be processed and a control signal; wherein, the control signal is used to determine whether to perform a split process or a fusion process on the data to be processed; the control signal includes: a fusion control signal and a split control signal; converting the data to be processed into multiple compressed slices and then caching them; performing a data fetch operation on the cached data in cycles based on the control signal, and outputting the result of the data fetch operation; wherein, when the fusion control signal is in an enabled state, a data fetch operation is performed once in each cycle; each data fetch operation fetches the cached data of multiple compressed slices; or, when the split control signal is in an enabled state, multiple data fetch operations are performed in parallel in each cycle; each data fetch operation fetches the cached data of corresponding partial compressed slices among multiple compressed slices.
[0016] In the above solution, performing the fetch operation on the cached data according to a period based on the control signal includes: converting the data of multiple compressed slices into a first array; in response to the fusion control signal, converting the first array into a second array; generating multiple request signals, and adjusting the assignments of the multiple request signals in response to the fusion control signal; wherein, the assignment of each bit of one of the multiple adjusted request signals is used to indicate fetching data from the second array, and different assignments of the request signal respectively represent multiple data in the same row of the corresponding second array; the assignment of each bit of the remaining request signals is used to indicate not fetching data; in response to the fusion control signal and the adjusted multiple request signals, performing a fetch operation on the second array.
[0017] In the above solution, performing the fetch operation on the cached data according to a period based on the control signal includes: converting the data of multiple compressed slices into a first array; in response to the split control signal, converting the first array into multiple second arrays; generating multiple request signals, and adjusting the assignments of the multiple request signals in response to the split control signal; wherein, the assignment of each bit of each adjusted request signal is used to indicate fetching data from a corresponding one of the second arrays, and different assignments of each request signal respectively represent multiple data in the same row of the corresponding second array; in response to the split control signal and the adjusted multiple request signals, performing a fetch operation on the second arrays. Description of the Drawings
[0018] Figure 1 Schematic structural diagram of a prior art DSP chip provided by an embodiment of the present disclosure;
[0019] Figure 2 Schematic structure of a data processing device provided by an embodiment of the present disclosure Figure 1 ;
[0020] Figure 3 Schematic diagram of the data flow of the fusion processing of the data to be processed provided by an embodiment of the present disclosure;
[0021] Figure 4 Schematic diagram of the data flow of the split processing of the data to be processed provided by an embodiment of the present disclosure;
[0022] Figure 5 Schematic structure of a data processing device provided by an embodiment of the present disclosure Figure 2 ;
[0023] Figure 6 Schematic structural diagram of an input port provided by an embodiment of the present disclosure;
[0024] Figure 7Structural schematic of the input unit provided by an embodiment of the present disclosure Figure 1 ;
[0025] Figure 8 Structural schematic of the input unit provided by an embodiment of the present disclosure Figure 2 ;
[0026] Figure 9 Structural schematic diagram of the slice data path provided by an embodiment of the present disclosure;
[0027] Figure 10 Structural schematic of the output unit provided by an embodiment of the present disclosure Figure 1 ;
[0028] Figure 11 Schematic of the data flow of the output unit provided by an embodiment of the present disclosure Figure 1 ;
[0029] Figure 12 Structural schematic of the output unit provided by an embodiment of the present disclosure Figure 2 ;
[0030] Figure 13 Schematic of the data flow of the output unit provided by an embodiment of the present disclosure Figure 2 ;
[0031] Figure 14 Schematic of the data flow of the output unit provided by an embodiment of the present disclosure Figure 3 ;
[0032] Figure 15 Schematic of the data flow of the output unit provided by an embodiment of the present disclosure Figure 4 ;
[0033] Figure 16 Structural schematic of the data processing device provided by an embodiment of the present disclosure Figure 3 ;
[0034] Figure 17 Schematic of the data fusion method provided by an embodiment of the present disclosure Figure 1 ;
[0035] Figure 18 Schematic of the data fusion method provided by an embodiment of the present disclosure Figure 2 ;
[0036] Figure 19 Schematic of the data fusion method provided by an embodiment of the present disclosure Figure 3 。 Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be construed as limitations on the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.
[0038] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it should be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and they may be combined with each other without conflict.
[0039] If similar descriptions such as "first / second" appear in the application documents, the following explanation is added. In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this disclosure belongs. The terms used herein are only for the purpose of describing the embodiments of the present disclosure and are not intended to limit the present disclosure.
[0041] Figure 1 FIG. is a schematic structural diagram of a digital signal processing chip (DSP) 200, which is an optional prior art provided by an embodiment of the present disclosure. It should be noted that Figure 1 Only the first channel 210 and the second channel 220 are illustrated. The digital signal processing chip (DSP) 200 may further include a greater number of channels, which is not limited herein. Each channel includes a data conversion module (Data Convert), a data splitting module (Split), a fusion module (Merge), an encoder (DSC_ENC), a cache module, and a frame control module (FrameTiming).
[0042] Refer to Figure 1, the data conversion modules 211 and 212 are used to maintain signal integrity in scenarios such as audio processing and wireless communication. For example, the data conversion modules 211 and 212 can convert analog signals into digital signals. The fusion modules 212 and 222 are used to perform data fusion operations, that is: integrating multiple sub-streams or processing results from different channels into a single data stream to ensure timing and logical consistency. The splitting modules 213 and 223 are used to split a continuous data stream or a high-bandwidth signal into multiple sub-streams according to specific rules (such as frequency bands, time slices) to facilitate parallel processing or data transmission through different channels. The encoders 214a, 214b, 224a, and 224b are used to perform compression encoding on video or image data to reduce the transmission bandwidth requirement and maintain visual quality. For example, the encoder can use algorithms such as predictive coding and transform domain quantization to reduce redundant data. The buffer modules 215 and 225 are used to buffer the data compressed by the encoder. The frame control modules 216a, 216b, 226a, and 226b are used to manage the timing synchronization and alignment of data frames to ensure clock consistency and data integrity between the transmitter and the receiver. For example, the frame control module can use a phase-locked loop (PLL) to generate a stable clock signal to eliminate the phase offset caused by transmission delay.
[0043] Continue to refer to Figure 1 , during the process of data fusion processing, the data of the first channel 210 and the second channel 220 are fused through the fusion module 211a of the first channel 210, and then input into the input port of the encoder 214a of the first channel 210. Finally, after being compressed by the encoder 214a, it is output after passing through the buffer module 215 and the frame control module 216a. That is to say, the data to be processed is fused through the fusion module and then transmitted to the encoder for compression.
[0044] Continue to refer to Figure 1, during the process of splitting the data to be processed, after the data to be processed passes through the data conversion unit 211, it is input to the splitting module 213 to implement the splitting of the data to be processed. Then, the splitting module 213 inputs the split data into the encoders 214a and 2014b of the first channel 220 for compression, and stores it in the buffer module 215. Finally, to ensure that the time intervals of the output of the split data remain within the allowable range, after the frame control modules 216a and 216c generate corresponding timings, the data cached in the buffer module 215 is output. For example, when the data to be processed is an image, the port for receiving the split data is the MIPI DSI (Display Serial Interface) interface, and the data skew allows a slight timing difference between the two signals. The frame control modules 216a and 216c need to ensure that the two outputs complete the transmission of adjacent rows within the horizontal front porch (HFP) cycle of the screen. That is to say, in the prior art, the data to be processed is split by the splitting module, and after the splitting process is completed, it is transmitted to the encoder for compression. At the same time, the prior art also needs to set up frame control modules and buffer modules to ensure that multiple data after splitting can be output synchronously.
[0045] Figure 2 is a schematic structural diagram of an optional data processing device 100 provided by an embodiment of the present disclosure. Refer to Figure 2 , it should be noted that Figure 2 the exemplified control module 10 can be set in the advanced peripheral bus (APB) or other devices, and there is no limitation here.
[0046] In an embodiment of the present disclosure, refer to Figure 2 , the data processing device 100 includes a control module 10. The control module 10 is used to generate a control signal. The control signal is used to determine whether to perform splitting processing or fusion processing on the data to be processed. The control module 10 may include a first register 10a and a second register 10b. The first register 10a can generate a fusion control signal in an enabled state when performing fusion processing on the data to be processed. The second register 10b can generate a splitting control signal in an enabled state when performing splitting processing on the data to be processed.
[0047] In an embodiment of the present disclosure, refer to Figure 2, the control signal can be implemented by methods such as encoding. For example, when the fusion control signal is in the enabled state, the level of the fusion control signal can be 1. Conversely, when the fusion control signal is in the disabled state, the level of the fusion control signal can be 0. When the splitting control signal is in the enabled state, the level of the splitting control signal can be 1. Conversely, when the splitting control signal is in the disabled state, the level of the splitting control signal can be 0.
[0048] In the embodiments of the present disclosure, referring to Figure 2 , the data processing device 100 includes an encoding module 20. The encoding module can be an encoder, and the encoder can encode the data to be processed based on the Display Stream Compression (DSC) technology. The encoding module 20 can convert the data to be processed into multiple compressed slices. For example, the encoding module 20 can use the encoder to divide the received data to be processed into multiple slices. Then, the encoding module 20 can use algorithms such as residual quantization, Color History Index (ICH), and entropy coding (VLC) to compress the decomposed multiple slices to generate multiple compressed slices. The data volume of each slice is smaller than that of the data to be processed, and the data volume of the compressed compressed slice is smaller than that of the corresponding slice.
[0049] In the embodiments of the present disclosure, referring to Figure 2 , the encoding module 20 can cache multiple compressed slices, and perform a fetch operation on the cached data in cycles based on the control signal, and output the result of the fetch operation. For example, when the fusion control signal is in the enabled state, the encoding module 20 can perform a fetch operation on multiple compressed slices in each cycle of the fetch operation. Then, after multiple cycles of fetch operations, the multiple compressed slices are fused into one fetch operation result. Another example is that when the splitting control signal is in the enabled state, the encoding module 20 can perform multiple parallel fetch operations on multiple compressed slices in each cycle of the fetch operation. Then, after multiple cycles of fetch operations, the multiple compressed slices are split into multiple fetch operation results. The number of parallel fetch operations can be equal to the number of predetermined splits of the data to be processed; for example, if the data to be processed is to be split into 2, the number of parallel fetch operations can be 2, and the number of fetch operation results can be 2. That is to say, the encoding module 20 can implement the splitting process and the fusion process of the data to be processed during the fetch operation on multiple compressed slices.
[0050] It can be understood that the encoding module 20 can perform splitting processing and fusion processing on multiple compressed slices by controlling the data fetching operation in each cycle. In this way, compared with the prior art, the embodiments of the present disclosure do not need to additionally set up hardware such as a splitting module and a fusion module for splitting processing and fusion processing. Therefore, the chip area can be further reduced, the power consumption can be lowered, and the integration degree of the chip can be improved. At the same time, the embodiments of the present disclosure can avoid the delay caused by splitting or fusion processing in the prior art. Therefore, the delay in the processing of the data to be processed can be reduced, and the data processing rate can be further improved.
[0051] In some embodiments of the present disclosure, with reference to Figure 2 , the encoding module 20 is configured to perform a data fetching operation once in each cycle when the fusion control signal is in the enabled state. Wherein, each data fetching operation fetches the cached data of multiple compressed slices.
[0052] Figure 3 FIG. is a data stream of an optional encoding module 20 provided by the embodiments of the present disclosure for performing fusion processing on the data to be processed. It should be noted that Figure 3 the data to be processed shown includes two sub-data, namely: image A1 and image B1. After image A1 and image B1 are compressed, they are fused to generate image C1, that is: image C1 is the result of the data fetching operation after the fusion of image A1 and image B1. The data to be processed can also be other types of data, and the data to be processed can also include a greater number of sub-data, which is not limited here.
[0053] In the embodiments of the present disclosure, in combination with Figure 2 and Figure 3 , the encoding module 20 can convert the data to be processed into multiple compressed slices. For example, when the fusion control signal is in the enabled state, the encoding module 20 can split image A1 into 2 slices, slice0 and slice1, and split image B1 into 2 slices, slice2 and slice3. Then, the encoding module 20 can compress the slices slice0, slice1, slice2, and slice3, generate corresponding 4 compressed slices and cache them.
[0054] Further, the encoding module 20 can, based on the control of the splitting control signal, perform a data fetching operation on multiple compressed slices in each cycle to generate a data fetching operation result. For example, as Figure 3As shown, each data fetch operation of the encoding module 20 can generate one line of data of the image C1. During the process of generating the first line of data of the image C1, the first line of data of the image A1 is evenly decomposed into the slices Slice0 and Slice1, and the first line of data of the image B1 is evenly decomposed into the slices Slice2 and Slice3. At this time, the encoding module 20 can, in the first cycle, sequentially perform data fetch operations on the compressed slices corresponding to the 4 slices Slice0, Slice1, Slice2, and Slice3 according to the chunk size, and obtain 4 fetched chunks chunk0, chunk1, chunk2, and chunk3. The data in the fetched chunks chunk0 and chunk1 is the first line of data of the compressed image A1, and the data in the fetched chunks chunk2 and chunk3 is the first line of data of the compressed image B1. Finally, the encoding module 20 sequentially splices the fetched chunks chunk0, chunk1, chunk2, and chunk3. The data of the spliced fetched chunks chunk0 and chunk1 is fused into the first first half line of the image C1, and the data of the spliced fetched chunks chunk2 and chunk3 is fused into the first second half line of the image C1. The data fetch operations for the remaining lines of the image C1 are performed in the manner shown in the first line. In this way, the encoding module 20 can perform fusion processing on multiple compressed slices through multiple cycles of data fetch operations, and thus fuse the compressed images A1 and B1 into the image C1.
[0055] In some other embodiments of the present disclosure, referring to Figure 2 , the encoding module 20 is configured to perform multiple data fetch operations in parallel in each cycle when the split control signal is in the enabled state. Wherein, each data fetch operation fetches the cached data of the corresponding partial compressed slices among the multiple compressed slices.
[0056] Figure 4 is an optional data stream for the encoding module 20 provided by an embodiment of the present disclosure to perform split processing on the data to be processed. It should be noted that Figure 4 the data to be processed shown includes the image A2. The images B2 and C2 are the results of the data fetch operations after splitting the image A2. The image A2 can also be split into a greater number, which is not limited here.
[0057] In the embodiments of the present disclosure, in combination with Figure 2 and Figure 4, the encoding module 20 can convert the data to be processed into multiple compressed slices. For example, when the splitting control signal is in the enabled state, the encoding module 20 can split Image A2 into 4 slices: slice0, slice1, slice2, and slice3. Then, the encoding module 20 can compress the slices slice0, slice1, slice2, and slice3, generate the corresponding 4 compressed slices, and cache them.
[0058] Furthermore, the encoding module 20 can, based on the control of the splitting control signal, perform a parallel data fetching operation on multiple compressed slices in cycles to generate multiple data fetching operation results. For example, as Figure 4 shown, slices Slice0, Slice1, Slice2, and Slice3 are the data after decomposition of Image A2. The first first half row of Image A2 can be evenly decomposed in slices Slice0 and Slice1. The first second half row of Image A2 can be evenly decomposed in slices Slice2 and Slice3. During the process of generating the first row data of Image B2, the encoding module 20 can, in the first cycle, fetch data from 2 slices Slice0 and Slice1 in sequence according to the block size, obtain 2 fetched blocks chunk0 and chunk1, and then splice chunk0 and chunk1 into the first row of Image B2. During the process of generating the first row data of Image C2, the encoding module 20 can, in the first cycle, perform a data fetching operation on the compressed slices corresponding to Slice2 and Slice3 in sequence according to the block size, obtain 2 fetched blocks chunk2 and chunk3, and then splice the fetched blocks chunk2 and chunk3 into the first row of Image C2. The data fetching operation process of the remaining rows of Image B2 and C2 can be understood by referring to the first row, which will not be elaborated here. In this way, the encoding module 20 can perform splitting processing on multiple compressed slices through parallel data fetching operations in multiple cycles, thereby splitting the compressed Image A2 into Image B2 and Image C2.
[0059] Figure 5 is a schematic structural diagram of an optional encoding module 20 provided in an embodiment of the present disclosure. It should be noted that Figure 5 the illustrated encoding module 20 includes an input unit 21, a compression encoding unit 22, and an output unit 23. The input unit 21 is configured to receive the data to be processed and decompose the data to be processed into multiple slices. The compression encoding unit 22 is configured to compress the slices into compressed slices and cache the compressed slices. The output unit 23 is configured to perform a data fetching operation on the cached compressed slices, pack, and output the result of the data fetching operation.
[0060] Figure 6 is provided in an embodiment of the present disclosure Figure 5Schematic diagram of the input port Port0. It should be noted that Figure 5 For the remaining input ports such as the input port Port1, reference can be made to Figure 6 the input port Port0 shown for understanding, which will not be elaborated here.
[0061] In the embodiments of the present disclosure, with reference to Figure 6 , the input port Port0 includes a preprocessing subunit (Pre logic) 205, a pixel padding subunit (Pixel Padding Logic) 206, and a pixel demultiplexing subunit (Pixel de-multiplexer Logic) 207. When the data to be processed is an image, the preprocessing subunit 205 is used to perform preprocessing such as format conversion and color space adjustment on the input pixel data. For example, the original image data is converted into the format required for DSC encoding, such as converting the RGB format to the YUV format, and ensuring that the data bit width meets the input requirements of the subsequent compression algorithm. The pixel padding subunit 206 is used to meet the boundary alignment requirements of block compression and pad the pixels of the input image. When the original resolution cannot be divided evenly by the minimum processing block (such as a macroblock) of the compression algorithm, the image size is extended by padding redundant pixels to ensure the integrity of subsequent block processing. The pixel demultiplexing subunit 207 is used to demultiplex or reorganize the input pixel data stream. For example, according to different color channels or parallel processing requirements, through the selectors mux1 and mux2, the serial input data is demultiplexed into multiple parallel data to adapt to the multi-channel processing architecture inside the encoder. The FIFO (First Input First Output) is used to ensure that the processing order of the data is the same as the input order.
[0062] It should also be noted that, with reference to Figure 5, the input unit 21 has two input ports, Port0 and Port1, and four slicing units 21a, 21b, 21c, and 21d. The encoding module 20 may also include other numbers of input ports, and each input port may also include other numbers of slicing units, which are not limited here. In the case where no control signal is received, each input port in the encoding module 20 is independently configured; for example, slicing units 21a and 21b are assigned to input port Port0, and slicing units 21a and 21b only decompose the data received by input port Port0; multiplexer 231 is assigned to input port Port0, and multiplexer 231 only fetches the compressed slices cached in the first slice data path 22a and the second slice data path 22b. Slicing units 21c and 21d are assigned to input port Port1, and slicing units 21c and 21d only decompose the data received by input port Port1; multiplexer 231 is assigned to input port Port1, and multiplexer 231 only fetches the compressed slices cached in the third slice data path 22c and the fourth slice data path 22d.
[0063] In some embodiments of the present disclosure, refer to Figure 5 , the input unit 21 is configured to, when the fusion control signal is in an enabled state, select multiple input ports from all the input ports of the input unit 21 to respectively receive multiple sub-data of the data to be processed, and decompose the correspondingly received sub-data through the slicing units connected to each input port.
[0064] In an embodiment of the present disclosure, refer to Figure 5 , when the fusion control signal is in an enabled state, the input unit 21 may select multiple input ports from all the input ports of the input unit 21 to respectively receive multiple sub-data of the data to be processed. For example, the data to be processed includes Figure 3 the images A1 and B1 shown. At this time, the input unit may select two input ports to receive images A1 and B1. Input port Port0 may receive Figure 3 the image A1 shown, and slicing units 21a and 21b assigned to input port Port0 decompose image A1 to generate Figure 3 the slices Slice0 and Slice1 shown. Input port Port1 may receive Figure 3 the image B1 shown, and slicing units 21c and 21d assigned to input port Port1 decompose image B1 to generate Figure 3 the slices Slice2 and Slice3 shown.
[0065] In some embodiments of the present disclosure, refer to Figure 5, the output unit 23 is configured to select any one of all the slice multiplexers of the output unit 23 to perform a data fetch operation when the fusion control signal is in the enabled state, and output the result of the data fetch operation through the output port connected to the slice multiplexer.
[0066] In the embodiments of the present disclosure, refer to Figure 5 , when the fusion control signal is in the enabled state, the output unit 23 can select any one of all the slice multiplexers of the output unit 23 to perform a data fetch operation. For example, the data to be processed is decomposed into four slices Slice0, Slice1, Slice2, and Slice3 as shown in Figure 3 . The compressed slices after compression of the 4 slices Slice0, Slice1, Slice2, and Slice3 can be cached in the first slice data path 22a, the second slice data path 22b, the third slice data path 22c, and the fourth slice data path 22d in one-to-one correspondence. The control module 10 can output a fusion control signal to the slice multiplexer 231. Then, based on the control of the fusion control signal, the slice multiplexer 231 performs a data fetch operation on the cached data of multiple compressed slices. The multiplexer 231 performs a data fetch operation on the first slice data path 22a and the second slice data path 22b, and generates data fetch blocks Chunk0 and Chunk1 as shown in Figure 3 . The multiplexer 231 performs a data fetch operation on the third slice data path 22c and the fourth slice data path 22d, and generates data fetch blocks Chunk2 and Chunk3 as shown in Figure 3 . Then, the multiplexer 231 can splice the data fetch blocks Chunk0, Chunk1, Chunk2, and Chunk3, that is: fuse the first row of image A1 and the first row of image B1 into the first row of image C1. Finally, the slice multiplexer 231 can generate image C1 through multiple data fetch operations and output image C1 from its output end. That is to say, compared with the method of independently configuring input ports in the prior art, the embodiments of the present disclosure can use the fusion control signal to control the slice multiplexer to perform a data fetch operation on the compressed slices of other input ports, so as to realize the fusion processing of multiple compressed slices during the data fetch operation.
[0067] In some other embodiments of the present disclosure, refer to Figure 5 , when the split control signal is in the enabled state, any one of all the input ports of the input unit 21 is selected to receive the data to be processed, and multiple slice units are selected from all the slice units of the input unit 21 to decompose the data to be processed.
[0068] In the embodiments of the present disclosure, refer to Figure 5, when the splitting control signal is in the enabled state, the input unit 21 can select any one of all the input ports of the input unit 21 to receive the data to be processed. For example, the input unit 21 can receive Figure 4 the image A2 shown.
[0069] In the embodiments of the present disclosure, referring to Figure 5 , when the splitting control signal is in the enabled state, the input unit 21 selects a plurality of slice units from all the slice units of the input unit 21 to decompose the data to be processed. That is to say, after receiving the splitting control signal, the input unit 21 can flexibly call the slice units assigned to multiple input ports. For example, the input port Port0 receives Figure 4 the image A2 shown; at this time, the input port Port1 does not receive data and is in an idle state. The slice units 21c and 21d assigned to the input port Port1 are also in an idle state. The input unit 21 can, under the control of the splitting control signal, call the slice units 21c and 21d assigned to the input port Port1 to decompose the data to be processed. The slice units 21a and 21b assigned to the input port Port0, and the slice units 21c and 21d assigned to the input port Port1 all decompose the image A2 to generate 4 slices Slice0, Slice1, Slice2 and Slice3 as shown in Figure 4 . That is to say, compared with the method of independently configuring input ports in the prior art, the embodiments of the present disclosure can decompose the data to be processed by calling the slice units of other idle input ports. In this way, the embodiments of the present disclosure can increase the number of slice units for decomposing the data to be processed, so as to improve the decomposition speed of the data to be processed, and further improve the data processing efficiency.
[0070] In other embodiments of the present disclosure, referring to Figure 5 , the output unit 23 is configured to, when the splitting control signal is in the enabled state, select a plurality of slice multiplexers from all the slice multiplexers of the output unit 23 to perform a fetching operation in parallel, and output the results of the fetching operation in parallel through a plurality of output ports connected to the plurality of slice multiplexers.
[0071] In the embodiments of the present disclosure, referring to Figure 5 , the input unit 23 can select a plurality of slice multiplexers from all the slice multiplexers of the output unit 23 to perform a fetching operation in parallel. The output end of the slice multiplexer is the output port of the output unit 23. For example, Figure 4 the image A2 shown needs to be split and processed into images B2 and C2. The input unit 23 needs to select two slice multiplexers to perform a fetching operation. The data to be processed is decomposed into as shown in Figure 4The four slices shown, Slice0, Slice1, Slice2, and Slice3, the compressed slices after compression of the 4 slices Slice0, Slice1, Slice2, and Slice3 can be correspondingly cached in the first slice data path 22a, the second slice data path 22b, the third slice data path 22c, and the fourth slice data path 22d. The control module 10 can output a split control signal to the slice multiplexer 231 and the slice multiplexer 232. Then, based on the control of the split control signal, the slice multiplexer 231 and the slice multiplexer 232 perform an operation of fetching data from the cached data of multiple compressed slices in parallel.
[0072] The slice multiplexer 231 can perform an operation of fetching data from the compressed slices cached in the first slice data path 22a and the second slice data path 22b, obtain 2 fetched chunks chunk0 and chunk1, and splice chunk0 and chunk1 into the first row of the image B2. The slice multiplexer 232 can perform an operation of fetching data from the compressed slices cached in the third slice data path 22c and the fourth slice data path 22d, obtain 2 fetched chunks chunk2 and chunk3, and splice the fetched chunks chunk2 and chunk3 into the first row of the image C2. Finally, the slice multiplexer 231 can generate the image B2 through multiple operations of fetching data based on the split control signal, and output the image B2 from its output terminal. The slice multiplexer 232 can generate the image C2 through multiple operations of fetching data, and output the image C2 from its output terminal. That is to say, in the embodiments of the present disclosure, multiple slice multiplexers perform operations of fetching data in parallel based on the control of the same split control signal. In this way, the embodiments of the present disclosure can ensure that the operations of fetching data in parallel in each cycle are synchronized. Furthermore, the embodiments of the present disclosure can synchronously output the results of multiple operations of fetching data after split processing. At the same time, the embodiments of the present disclosure do not need to further set a frame control unit, and can also ensure the synchronous output of multiple data after split processing, further improving the chip integration and reducing the power consumption.
[0073] Figure 7 It is a schematic structural diagram of an optional input unit 21 provided by the embodiments of the present disclosure. It should be noted that Figure 7 A more specific example shows the data stream in which the control signal is a fusion control signal. Figure 7The input unit 21 illustrated in the example includes a second combination subunit (Combination Logic) 220, a second router (X-BAR Logic) 230, a second separation subunit (Separation Logic) 240, pixel buffer subunits (Pixel buffer) 250a, 250b, 250c, and 250d, read subunits (Read Request Logic) 260a, 260b, 260c, and 260d, and unpacking component analysis subunits (Unpackcomponent logic) 270a, 270b, 270c, and 250d. The second combination subunit 220 is used to receive multiple slices and combine the arrays corresponding to the multiple slices to generate a multi-dimensional array. The second separation subunit 230 is configured to receive the one-dimensional array after the transformation of the multi-dimensional array and transmit the pixel data in the one-dimensional array to the corresponding slice data path in the encoding and compression unit. The read subunits (Read Request Logic) 250a, 250b, 250c, and 250d are used to generate read request signals and read the data in the slice unit based on the read request signals. The pixel buffer subunits 250a, 250b, 250c, and 250d are used to receive and cache the read pixel data. The unpacking component analysis subunits 270a, 270b, 270c, and 250d are used to extract the parameter components of the pixel data. The FIFO (First Input First Output) is used to ensure that the processing order of the data is consistent with the input order.
[0074] In some embodiments of the present disclosure, referring to Figure 7 , the second routing subunit 230 is respectively connected to the control unit 10 and all input ports. For example, Figure 7 in the second routing subunit 230 is connected to the input port Port0 and the input port Port1. The second routing subunit 230 can receive the control signal output by the control module.
[0075] In the embodiments of the present disclosure, referring to Figure 7 , the second routing subunit 230 is configured to, when the fusion control signal is in the enabled state, respectively transmit the sub-data received by each input port to the slice unit connected to the input port. For example, when the fusion control signal is in the enabled state, the data to be processed is in the manner exemplified by D1 in Figure 7 . The two sub-data included in the data to be processed are respectively input to the second routing subunit 230 along the input port Port0 and the input port Port1, and then the second routing subunit 230 respectively transmits the data received by the input port Port0 to the slice units 21a and 21b. The second routing subunit 230 respectively transmits the data received by the input port Port1 to the slice units 21c and 21d.
[0076] Figure 8 It is a schematic structural diagram of an optional input unit 21 provided by an embodiment of the present disclosure. It should be noted that Figure 8 A more specific example shows the data stream in which the control signal is a split control signal. Figure 8 The devices in Figure 7 can be understood with reference to
[0077] In some other embodiments of the present disclosure, with reference to Figure 8 , when the split control signal is in the enabled state, the second routing subunit 230 transmits the data to be processed to a plurality of slice units selected by the split control signal. For example, when the split control signal is in the enabled state, the data to be processed is input to the second routing subunit 230 along the input port Port0 in the manner exemplified by D2 in Figure 8 . Then, the second routing subunit 230 transmits the data received by the input port Port0 to the slice units 21a and 21b, 21c and 21d respectively. In this way, the embodiment of the present disclosure can increase the number of slice units for decomposing the data to be processed, so as to improve the decomposition speed of the data to be processed, and further improve the data processing efficiency.
[0078] In some embodiments of the present disclosure, with reference to Figure 5 , the encoding and compression unit 22 is respectively connected to the input unit 21 and the output unit 23. The data of multiple compressed slices are cached through multiple bitrate buffer units in the encoding and compression unit 22. For example, the slice slice0 is compressed through the first slice data path 22a to generate a compressed slice, and the compressed slice corresponding to the slice slice0 is stored in the bitrate buffer unit 201.
[0079] Figure 9 It is a schematic structural diagram of an optional slice data path provided by an embodiment of the present disclosure. It should be noted that Figure 9 Specifically exemplifies Figure 5 the first slice data path 22a in Figure 5 The remaining slice data paths in Figure 9 can be understood with reference to the first slice data path 22a shown in
[0080] In the embodiments of the present disclosure, in combination with Figure 5 and Figure 9, the encoding and compression unit 22 can convert the data received by each slice unit into a corresponding compressed slice. For example, the first slice data path 22a includes a color space conversion subunit 241, a prediction quantization and reconstruction subunit 242, a line buffer subunit 243, a rate control subunit 244, an entropy encoding subunit 245, a sub-stream multiplexing subunit 246, and a data adjustment subunit 247. The color space conversion subunit 241 can receive the corresponding slice Slice0 and perform format conversion on Slice0. The prediction quantization and reconstruction subunit 242 can be connected to the color space conversion unit 241 and is configured to calculate the reconstruction value of the pixel data in Slice0. The line buffer subunit 243 can be connected to the prediction quantization and reconstruction unit 242 and is configured to receive and cache the reconstruction value. The entropy encoding subunit 245 can be connected to the prediction quantization and reconstruction unit 242 and is configured to encode Slice0 to generate a compressed slice. The rate control subunit 244 can be respectively connected to the entropy encoding unit 245 and the prediction quantization and reconstruction unit 242 and is configured to control the output bit rate of the compressed slice. The data adjustment subunit 247, which is respectively connected to the entropy encoding unit 245 and the sub-stream multiplexing subunit 246, is configured to adjust the size of the compressed slice. The sub-stream multiplexing subunit 246 is configured to multiplex multiple sub-streams in the compressed slice into one data packet.
[0081] Figure 10 It is a schematic structural diagram of an optional output unit 23 provided by an embodiment of the present disclosure. It should be noted that Figure 10 A more specific example shows the data stream after the input unit 23 receives the fusion control signal. Figure 10 The output unit 23 shown in the figure includes a first merging subunit 233, a first separating subunit 235, and data packing subunits 236 and 237. The first merging subunit 233 is configured to receive multiple compressed slices and convert the multiple compressed slices into a three-dimensional array. The second separating subunit 235 is configured to receive a request signal and fetch data from the one-dimensional array after the compressed slice is converted based on the request signal. The data packing subunits 236 and 237 are configured to receive the result of the fetch operation and output the result of the fetch operation after packing.
[0082] In some embodiments of the present disclosure, with reference to Figure 10 , the first routing subunit 234 is configured to, when the fusion control signal is in the enabled state, transmit the data of the multiple cached compressed slices to any one of the slice multiplexers selected by the fusion control signal.
[0083] Figure 11This is the data stream for transmitting data of the first routing subunit 234 provided by an embodiment of the present disclosure. It should be noted that "Slice0 data" can be Figure 10 the compressed slices cached in the medium bitrate buffer unit 201, "Slice1 data" can be Figure 10 the compressed slices cached in the medium bitrate buffer unit 202, "Slice2 data" can be Figure 10 the compressed slices cached in the medium bitrate buffer unit 203, "Slice3 data" can be Figure 10 the compressed slices cached in the medium bitrate buffer unit 204.
[0084] In an embodiment of the present disclosure, in combination with Figure 10 and Figure 11 , when the fusion control signal is in the enabled state, multiple compressed slices to be converted from the data to be processed are input into the first combining subunit 233 in the manner exemplified by D3 in Figure 10 along the medium bitrate buffer units 201, 202, 203, and 204. Then, as shown in Figure 11 , the first combining subunit 233 can convert the compressed slices in the medium bitrate buffer units 201, 202, 203, and 204 into a three-dimensional array array0[3:0], and transmit the three-dimensional array array0[3:0] to the first routing subunit 234.
[0085] Furthermore, the first routing subunit 234 can convert the received three-dimensional array array0[3:0] into a corresponding number of one-dimensional arrays according to the number of slice units allocated to each input port and the enabled state of the control signal. For example, as shown in Figure 11 , both input ports Port0 and Port1 include 2 slice units, and the fusion control signal is in the enabled state. At this time, the first routing subunit 234 converts the received three-dimensional array array0[3:0] into a one-dimensional array array1[3:0], and then transmits the one-dimensional array array1[3:0] to the first separating subunit 235. After the multiplexer 231 receives the fusion control signal in the enabled state, it sends a request signal to the first separating subunit 235. Then, the first separating subunit 235 transmits the one-dimensional array array1[3:0] to the multiplexer 231. In this way, compared with the prior art in which the input ports are independently configured, the embodiment of the present disclosure can control the first routing subunit 234 to fuse multiple compressed slices from different input ports through the fusion control signal.
[0086] Figure 12 This is a schematic structural diagram of an optional output unit 23 provided by an embodiment of the present disclosure. It should be noted that Figure 12More specific examples show the data stream in which the control signal is a split control signal. Figure 12 The devices in Figure 10 can be understood with reference to
[0087] and will not be elaborated here. Figure 12 In some other embodiments of the present disclosure, with reference to
[0088] Figure 13 Figure 3 is the data stream for the first routing subunit 234 to transmit data provided by the embodiments of the present disclosure. It should be noted that "Slice0 data" can be Figure 12 the compressed slices cached in the bitrate buffer unit 201 in Figure 12 "Slice1 data" can be Figure 12 the compressed slices cached in the bitrate buffer unit 202 in Figure 12 "Slice2 data" can be
[0089] In the embodiments of the present disclosure, in combination with Figure 12 and Figure 13 when the split control signal is in the enabled state, a part of the multiple compressed slices can be input into the first combining subunit 233 along the bitrate buffer units 201 and 202 in the manner exemplified by D4 in Figure 11 A part of the multiple compressed slices can be in the manner exemplified by D4 in Figure 11 Another part of the multiple compressed slices can be in the manner exemplified by D5 in Figure 11 and input into the first combining subunit 233 along the bitrate buffer units 203 and 204. Then, as shown in Figure 13 the first combining subunit 233 can convert the compressed slices in the bitrate buffer units 201, 202, 203, and 204 into a three-dimensional array array0[3:0] and transmit the three-dimensional array array0[3:0] to the first routing subunit 234.
[0090] Furthermore, the first routing subunit 234 can convert the received three-dimensional array array0[3:0] into the corresponding number of one-dimensional arrays according to the number of slice units assigned to each input port and the enabled state of the control signal. For example, as shown in Figure 13As shown, both input ports Port0 and Port1 include two slicing units, and the control signal is a split control signal in the enabled state. At this time, the first routing subunit 234 converts the received three-dimensional array array0[3:0] into two one-dimensional arrays array1[3:0] and array2[3:0], and then transmits the one-dimensional arrays array1[3:0] and array2[3:0] to the first separation subunit 235. After receiving the fusion control signal in the enabled state, the multiplexers 231 and 232 send a request signal to the first separation subunit 235. Then, the first separation subunit 235 transmits the one-dimensional array array1[3:0] to the multiplexer 231 and the one-dimensional array array2[3:0] to the multiplexer 232.
[0091] That is to say, multiple slice multiplexers in the embodiments of the present disclosure transmit the results of the fetch operation based on the same first routing subunit 234, and the transmission process of the first routing subunit 234 is controlled by a split control signal. In this way, the embodiments of the present disclosure can synchronize the transmission of parallel fetch operations through the first routing subunit 234. Thus, the embodiments of the present disclosure can ensure that parallel fetch operations in each cycle are synchronized. Furthermore, the embodiments of the present disclosure can synchronously output the results of multiple fetch operations after splitting processing.
[0092] In some embodiments of the present disclosure, each slice multiplexer is configured to receive a control signal, generate a request signal, and adjust the assignment of the request signal based on the control signal. The assignment of the request signal is used to select the compressed slice corresponding to each fetch operation among multiple compressed slices.
[0093] Figure 14 and 15 are the data streams of the optional slice multiplexer for fetch operations provided by the embodiments of the present disclosure. It should be noted that Figure 15 specifically exemplifies the data stream when the control signal is a fusion control signal, Figure 15 and specifically exemplifies the data stream when the control signal is a split control signal.
[0094] In the embodiments of the present disclosure, in combination with Figure 10 and Figure 14 , when the fusion control signal is in the enabled state, the multiplexer 231 generates a request signal slice_req and fetches data from the one-dimensional array array1[3:0] based on the generated request signal slice_req. For example, the request signal slice_req can adopt bitmask encoding, and the binary assignment of the request signal slice_req can directly map the fetch block that needs to be operated currently. As Figure 3As shown, the process of fetching data for the first row of image C1 is to splice four data fetching blocks chunk0, 1, 2, and 3. The multiplexer 231 can first assign the request signal slice_req to 1, and the multiplexer 231 fetches data from the one-dimensional array array1[3:0] to generate the data fetching block chunk0. Then, in the next few cycles, the multiplexer 231 can sequentially assign the request signal slice_req to 2, 4, and 8, and the multiplexer 231 fetches data from the one-dimensional array array1[3:0] to generate the data fetching blocks chunk1, chunk2, and chunk3.
[0095] In the embodiments of the present disclosure, in combination with Figure 12 and Figure 15 , when the split control signal is in the enabled state, the multiplexer 231 generates the request signal slice_req, and fetches data from the one-dimensional array array1[3:0] based on the generated request signal slice_req. The multiplexer 232 generates the request signal slice_req, and fetches data from the one-dimensional array array2[3:0] based on the generated request signal slice_req. As Figure 4 shown, the process of fetching data for the first row of image B2 is to splice two data fetching blocks chunk0 and 1, and the process of fetching data for the first row of image C2 is to splice two data fetching blocks chunk2 and 3. In the first cycle, the multiplexer 231 can first assign the request signal slice_req to 1, the multiplexer 231 fetches data from the one-dimensional array array1[3:0] to generate the data fetching block chunk0; the multiplexer 232 fetches data from the one-dimensional array array2[3:0] to generate the data fetching block chunk1. Then, in the next cycle, the multiplexer 231 can assign the request signal slice_req to 2, the multiplexer 231 fetches data from the one-dimensional array array1[3:0] to generate the data fetching block chunk1; the multiplexer 232 can assign the request signal slice_req to 2, and the multiplexer 232 fetches data from the one-dimensional array array2[3:0] to generate the data fetching block chunk3.
[0096] Figure 16 FIG. is a schematic structural diagram of another optional data processing apparatus 100 provided by the embodiments of the present disclosure. It should be noted that, with reference to Figure 16, the data processing device 100 may further include a splitting module 30 and a fusion module 40. Both the splitting module 30 and the fusion module 40 may be connected to the control module 10. The output ends of the splitting module 30 and the fusion module 40 are both connected to the encoding module 20. Both the splitting module 30 and the fusion module 40 may be designed with a dedicated bypass control circuit (Bypass), and enter the bypass mode after receiving a control signal. For example, the splitting module 30 may receive the splitting control signal generated by the control module 10, and when the splitting control signal is in the enabled state, bypass the received data to be processed to the control module 10. The fusion module 40 may receive the splitting control signal generated by the control module 10, and when the fusion control signal is in the enabled state, bypass the received data to be processed to the control module 10. That is to say, in the bypass mode, the splitting module 30 and the fusion module 40 transmit the received data to be processed through the standby channel. In this way, the embodiments of the present disclosure can avoid the delay caused by splitting or fusion processing in the prior art, thereby reducing the delay in the process of processing the data to be processed and further improving the data processing rate.
[0097] Figure 17 is a schematic flowchart of an optional data processing method provided by the embodiments of the present disclosure. It should be noted that Figure 17 the exemplified data processing method can be implemented by Figure 2 or Figure 10 the data processing device 100 shown, and will be described in conjunction with each step.
[0098] S101. Obtain the data to be processed and a control signal; wherein, the control signal is used to determine whether to perform splitting processing or fusion processing on the data to be processed; the control signal includes: a fusion control signal and a splitting control signal.
[0099] S102. Convert the data to be processed into multiple compressed slices and then cache them.
[0100] S103. Perform a data fetch operation on the cached data in cycles based on the control signal, and output the result of the data fetch operation; wherein, when the fusion control signal is in the enabled state, perform a data fetch operation once per cycle; each data fetch operation fetches the cached data of multiple compressed slices; or, when the splitting control signal is in the enabled state, perform multiple data fetch operations in parallel per cycle; each data fetch operation fetches the cached data of the corresponding partial compressed slices among the multiple compressed slices.
[0101] In the embodiments of the present disclosure, refer to Figure 2, when the fusion control signal is in the enabled state, the encoding module 20 can perform a fetch operation on multiple compressed slices in each cycle of the fetch operation. Then, after multiple cycles of fetch operations, the multiple compressed slices are fused into one fetch operation result. That is to say, the encoding module 20 can perform a fusion process on multiple sub-data of the data to be processed based on the control of the fusion control signal.
[0102] In the embodiments of the present disclosure, with reference to Figure 2 , when the split control signal is in the enabled state, the encoding module 20 can perform multiple parallel fetch operations on multiple compressed slices in each cycle of the fetch operation. Then, after multiple cycles of fetch operations, the multiple compressed slices are split into multiple fetch operation results. The number of parallel fetch operations can be equal to the number of predetermined splits of the data to be processed. For example, if the data to be processed is to be split into 2, the number of parallel fetch operations can be 2, and the number of fetch operation results can be 2. That is to say, the encoding module 20 can perform a split process on the data to be processed based on the control of the split control signal.
[0103] It can be understood that the encoding module can perform a split process and a fusion process on multiple compressed slices by controlling the fetch operation in each cycle. In this way, compared with the prior art, the embodiments of the present disclosure do not need to additionally set up hardware such as a split module and a fusion module to perform the split process and the fusion process. Therefore, the chip area can be further reduced, and the integration degree of the chip can be improved. At the same time, the embodiments of the present disclosure can avoid the delay caused by the split or fusion process in the prior art. Therefore, the delay in the process of processing the data to be processed can be reduced, and the data processing rate can be further improved.
[0104] In some embodiments of the present disclosure, it can be implemented through Figure 18 shown in S201~S204 to implement Figure 17 S103 in , and will be described in combination with each step.
[0105] S201. Convert the data of multiple compressed slices into a first array.
[0106] S202. In response to the fusion control signal, convert the first array into a second array.
[0107] In the embodiments of the present disclosure, in combination with Figure 10 and Figure 11 , when the fusion control signal is in the enabled state, multiple compressed slices converted from the data to be processed are input into the first combining subunit 233 in the manner exemplified by D3 in Figure 10 . Then, as shown in Figure 11As shown, the first combined unit subunit 233 can convert the compressed slices in the bitrate buffer units 201, 202, 203, and 204 into a three-dimensional array array0[3:0] (the first array), and transmit the three-dimensional array array0[3:0] to the first routing subunit 234.
[0108] Further, the first routing subunit 234 can convert the received three-dimensional array array0[3:0] into a corresponding number of one-dimensional arrays (the second arrays) according to the number of slice units assigned to each input port and the enabling state of the control signal. For example, as Figure 11 shown, both input ports Port0 and Port1 include 2 slice units, and the fusion control signal is in the enabled state. At this time, the first routing subunit 234 converts the received three-dimensional array array0[3:0] into a one-dimensional array array1[3:0], and then transmits the one-dimensional array array1[3:0] to the first separation subunit 235. After receiving the enabled fusion control signal, the multiplexer 231 sends a request signal to the first separation subunit 235. Then, the first separation subunit 235 transmits the one-dimensional array array1[3:0] (the second array) to the multiplexer 231.
[0109] S203. Generate a plurality of request signals, and adjust the assignment of the plurality of request signals in response to the fusion control signal; wherein, the assignment of each bit of one of the adjusted plurality of request signals is used to represent fetching data from the second array, and different assignments of the request signal respectively represent a plurality of data in the same row of the corresponding second array; the assignment of each bit of the remaining request signals is used to represent not fetching data.
[0110] S204. Perform a fetching operation on the second array in response to the fusion control signal and the adjusted plurality of request signals.
[0111] In the embodiments of the present disclosure, in combination with Figure 10 and Figure 14 , when the fusion control signal is in the enabled state, the multiplexer 231 generates a request signal slice_req, and fetches data from the one-dimensional array array1[3:0] based on the generated request signal slice_req. For example, the request signal slice_req can adopt bitmask encoding, and the binary assignment of the request signal slice_req can directly map the fetching block that needs to be operated currently. As Figure 3As shown, the process of fetching data from the first row of image C1 is to splice four fetch blocks chunk0, 1, 2, and 3. The multiplexer 231 can first assign the request signal slice_req to 1, and the multiplexer 231 fetches data from the one-dimensional array array1[3:0] to generate the fetch block chunk0. Then, in the next few cycles, the multiplexer 231 can sequentially assign the request signal slice_req to 2, 4, and 8, and the multiplexer 231 fetches data from the one-dimensional array array1[3:0] to generate the fetch blocks chunk1, chunk2, and chunk3.
[0112] In some embodiments of the present disclosure, it can be implemented through Figure 19 the S301 to S304 shown Figure 17 in to implement S103 in, which will be described in combination with each step.
[0113] S301: Convert the data of multiple compressed slices into a first array.
[0114] S302: In response to the split control signal, convert the first array into multiple second arrays.
[0115] S303: Generate multiple request signals and adjust the assignments of the multiple request signals in response to the split control signal; wherein, the assignment of each bit in each adjusted request signal is used to indicate fetching data from a corresponding second array, and different assignments of each request signal respectively represent multiple data in the same row of the corresponding second array.
[0116] S304: In response to the split control signal and the adjusted multiple request signals, perform a fetch operation on the second array.
[0117] In the embodiments of the present disclosure, in combination with Figure 12 and Figure 13 , when the split control signal is in the enabled state, a part of the multiple compressed slices can be input to the first combining unit 233 along the rate buffer units 201 and 202 in the manner exemplified by D4 in Figure 11 . A part of the multiple compressed slices can be input to the first combining unit 233 in the manner exemplified by D4 in Figure 11 , and another part of the multiple compressed slices can be input to the first combining unit 233 in the manner exemplified by D5 in Figure 11 along the rate buffer units 203 and 204. Then, as Figure 13 shown, the first combining unit sub-unit 233 can convert the compressed slices in the rate buffer units 201, 202, 203, and 204 into a three-dimensional array array0[3:0] (the first array), and transmit the three-dimensional array array0[3:0] to the first routing unit 234.
[0118] Furthermore, the first routing subunit 234 can convert the received three-dimensional array array0[3:0] into a corresponding number of one-dimensional arrays according to the number of slice units allocated to each input port and the enabling state of the control signal. For example, as Figure 11 shown, both input ports Port0 and Port1 include 2 slice units, and the splitting control signal of the control signal is in the enabled state. At this time, the first routing subunit 234 converts the received three-dimensional array array0[3:0] into two one-dimensional arrays array1[3:0] and array2[3:0] (the second array), and then transmits the one-dimensional arrays array1[3:0] and array2[3:0] to the first separation subunit 235. After receiving the enabled state fusion control signal, the multiplexers 231 and 232 send a request signal to the first separation subunit 235. Then, the first separation subunit 235 transmits the one-dimensional array array1[3:0] to the multiplexer 231 and the one-dimensional array array2[3:0] to the multiplexer 232.
[0119] In the embodiments of the present disclosure, in combination with Figure 12 and Figure 15 , when the splitting control signal is in the enabled state, the multiplexer 231 generates a request signal slice_req and fetches data from the one-dimensional array array1[3:0] based on the generated request signal slice_req. The multiplexer 232 generates a request signal slice_req and fetches data from the one-dimensional array array2[3:0] based on the generated request signal slice_req. As Figure 4 shown, the process of fetching data for the first row of image B2 is to splice 2 fetch blocks chunk0 and 1, and the process of fetching data for the first row of image C2 is to splice 2 fetch blocks chunk2 and 3. In the first cycle, the multiplexer 231 can first assign the request signal slice_req to 1, the multiplexer 231 fetches data from the one-dimensional array array1[3:0], and generates a fetch block chunk0; the multiplexer 232 fetches data from the one-dimensional array array2[3:0], and generates a fetch block chunk1. Then, in the next cycle, the multiplexer 231 can assign the request signal slice_req to 2, the multiplexer 231 fetches data from the one-dimensional array array1[3:0], and generates a fetch block chunk1; the multiplexer 232 can assign the request signal slice_req to 2, and the multiplexer 232 fetches data from the one-dimensional array array2[3:0], and generates a fetch block chunk3.
[0120] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.
[0121] The serial numbers of the above-described embodiments of the present disclosure are for description only and do not represent the superiority or inferiority of the embodiments. The methods disclosed in several method embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments. The features disclosed in several product embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new product embodiments. The features disclosed in several method or device embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0122] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A data processing device, comprising: Control module and encoding module; wherein, The control module is configured to generate a control signal; wherein the control signal is used to determine whether to perform splitting processing or fusion processing on the data to be processed; The encoding module is connected to the control module and is configured to receive the data to be processed and the control signal, convert the data to be processed into multiple compressed slices and cache them, perform data acquisition operations on the cached data periodically based on the control signal, and output the results of the data acquisition operations.
2. The data processing device according to claim 1, wherein the control signal comprises: Fusion control signals and split control signals; The encoding module is configured to perform the data acquisition operation once in each cycle when the fusion control signal is in an enabled state; wherein each data acquisition operation acquires data from cache data of a plurality of compressed slices; or When the split control signal is in an enabled state, the data fetching operation is performed multiple times in parallel in each cycle; wherein each data fetching operation fetches cache data of a corresponding part of the compressed slices among the multiple compressed slices.
3. The data processing device according to claim 2, wherein: The encoding module includes: an output unit; The output unit is configured to select any one of all the slice multiplexers of the output unit to perform the data acquisition operation when the fusion control signal is in an enabled state, and output the result of the data acquisition operation through the output port connected to the slice multiplexer; or When the split control signal is in an enabled state, a plurality of slice multiplexers are selected from all slice multiplexers of the output unit to perform the data acquisition operation in parallel, and the results of the data acquisition operation are output in parallel through a plurality of output ports connected to the plurality of slice multiplexers.
4. The data processing device according to claim 3, wherein: Each of the slice multiplexers is connected to the control module and is configured to receive the control signal, and generating a request signal and adjusting the value of the request signal based on the control signal, and, Based on the adjusted value of the request signal, the data fetching operation is performed according to the block size; wherein, The assignment of the request signal is used to select the compressed slice corresponding to each data access operation from the multiple compressed slices; the data capacity of the block size is smaller than the data capacity of the compressed slice.
5. The data processing device according to claim 3, wherein: The output unit further includes: a first routing subunit; wherein, The first routing subunit is connected to the control module and all the slice multiplexers respectively, and is configured to transmit the cached data of the plurality of compressed slices to any one of the slice multiplexers selected by the fusion control signal when the fusion control signal is in an enabled state; or When the split control signal is in an enabled state, the cached data of the plurality of compressed slices are respectively transmitted to the plurality of slice multiplexers selected by the split control signal.
6. The data processing device according to claim 3, wherein: The encoding module further includes: an input unit; The input unit is configured to select a plurality of input ports from all input ports of the input unit to respectively receive a plurality of sub-data of the data to be processed, and decompose the corresponding received sub-data through a slicing unit connected to each input port when the fusion control signal is in an enabled state; or When the split control signal is in an enabled state, any one of the input ports of the input unit is selected to receive the data to be processed, and a plurality of the slice units are selected from all the slice units of the input unit to decompose the data to be processed.
7. The data processing device according to claim 6, wherein: The input unit includes: a second routing subunit; The second routing sub-unit is connected to the control module and all the input ports respectively, and is configured to transmit the sub-data received by each input port to the slice unit connected to the input port respectively when the fusion control signal is in an enabled state; or When the split control signal is in an enabled state, the to-be-processed data is transmitted to the plurality of slice units selected by the split control signal.
8. The data processing device according to claim 6, wherein: The encoding module also includes: an encoding compression unit; The encoding compression unit is respectively connected to the input unit and the output unit, and is configured to convert the data received by each slicing unit into a corresponding compressed slice, and cache the data of multiple compressed slices through multiple bit rate buffer units in the encoding compression unit.
9. The data processing device according to claim 2, wherein: The control module includes: a first register and a second register; wherein, The first register is configured to generate the fusion control signal in an enabled state when the fusion processing is performed on the data to be processed; the second register is configured to generate the split control signal in an enabled state when the split processing is performed on the data to be processed.
10. A data processing method, comprising: Acquire data to be processed and a control signal; wherein the control signal is used to determine whether to perform split processing or fusion processing on the data to be processed; the control signal includes: a fusion control signal and a split control signal; Convert the data to be processed into multiple compressed slices and cache them; Based on the control signal, a data fetch operation is performed on the cached data in a cycle, and the result of the data fetch operation is output; wherein, When the fusion control signal is in an enabled state, a fetch operation is performed in each cycle; each time the fetch operation is performed, cache data of multiple compressed slices are fetched; or, when the splitting control signal is in an enabled state, multiple fetch operations are performed in parallel in each cycle; each time the fetch operation is performed, cache data of a corresponding part of the compressed slices among the multiple compressed slices are fetched.
11. The data processing method according to claim 10, wherein: Performing the data fetching operation on the cached data periodically based on the control signal includes: Convert the data of the plurality of compressed slices into a first array; In response to the fusion control signal, converting the first array into a second array; Generate multiple request signals, and adjust the assignment of the multiple request signals in response to the fusion control signal; wherein the assignment of each bit of one of the multiple request signals after the adjustment is used to indicate that data is to be fetched from the second array, and different assignments of the request signal respectively represent multiple data corresponding to the same row in the second array; the assignment of each bit of the remaining request signals is used to indicate that data is not to be fetched; In response to the fusion control signal and the adjusted plurality of request signals, a data fetch operation is performed on the second array.
12. The data processing method according to claim 10, wherein: Performing the data fetching operation on the cached data periodically based on the control signal includes: Convert the data of the plurality of compressed slices into a first array; In response to the split control signal, converting the first array into a plurality of second arrays; Generate multiple request signals, and adjust the assignment of the multiple request signals in response to the split control signal; wherein the assignment of each bit in each of the adjusted request signals is used to represent fetching data from a corresponding second array, and different assignments of each of the request signals respectively represent multiple data in the same row in the corresponding second array; In response to the split control signal and the adjusted plurality of request signals, a data fetch operation is performed on the second array.