Method, device and equipment for processing low-bit-width data by high-bit-width memory and medium
By storing the pending array to the vector register, the data alignment problem of 32-bit microcontrollers when processing low-bit width data is solved, efficient data processing and simplified data management are achieved, which improves processing efficiency and reduces development difficulty.
Patent Information
- Application Number
- CN202510033882.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-30
AI Technical Summary
32-bit microcontrollers often encounter data alignment problems when processing 8-bit or 16-bit data, which leads to complex data reading and may cause data misalignment, increasing the difficulty of software development and processing burden.
Data management and access is simplified by receiving multiple pending arrays and storing them into preset multiple vector registers, each vector register's channel stores one low bit-width data, thereby storing one low bit-width data per address in the high bit-width memory.
It significantly improves the speed and efficiency of data processing, reduces the number of instructions required to process data, simplifies data management and access, and reduces the difficulty of software development and the processing burden during runtime.
Smart Images

Figure CN120066982A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data communication technologies, and particularly to a method, apparatus, device, and medium for a high-width memory to process low-width data. Background Art
[0002] With the rapid development of electronic technologies, high-performance processor devices such as microcontrollers / microcontrollers / central processing units play an increasingly important role in various embedded systems. Among them, 32-bit microcontrollers have gradually become the mainstream choice in the industry due to their excellent processing capabilities and low power consumption.
[0003] However, in practical applications, some traditional communication protocols are still based on 8-bit or 16-bit system architectures, which leads to data alignment problems when 32-bit microcontrollers process 8-bit or 16-bit data transmitted by such system architectures. Specifically, taking a 32-bit microcontroller processing multiple consecutive 16-bit data as an example, when 16-bit aligned data is mapped to the memory space of a 32-bit microcontroller, if the number of 16-bit data is odd, the last 16-bit data will be stored separately in the high 16 bits or low 16 bits of a 32-bit address space. This situation not only makes data reading complex but also may cause data misalignment, that is, the high 16 bits and low 16 bits stored at the same 32-bit address are data of different attributes. When the 32-bit microcontroller processes these data, it must perform additional high-low bit splitting operations and also need to accurately calculate the storage addresses of data of each attribute. This undoubtedly greatly increases the software development difficulty and runtime processing burden of 32-bit microcontrollers. Summary of the Invention
[0004] In view of the above problems, embodiments of this application provide a method, apparatus, device, and medium for a high-width memory to process low-width data to solve the above technical problems.
[0005] In a first aspect, an embodiment of this application provides a method for a high-width memory to process low-width data, including:
[0006] Receiving a plurality of arrays to be processed, and sequentially storing each of the arrays to be processed starting from a preset starting address of the high-width memory. Each array to be processed includes a plurality of low-width data, and one address of the high-width memory stores at least 2 of the low-width data;
[0007] Storing the arrays to be processed into a preset plurality of vector registers, so that each low-width data in the i-th array to be processed is sequentially stored into the i-th channel of each of the vector registers, where i is all integers between 1 and n, and n is the number of arrays to be processed;
[0008] Sequentially obtain each of the arrays to be processed from the respective vector registers, and sequentially store each of the arrays to be processed starting from a preset starting address of the high-width memory, where one address of the high-width memory stores one low-width data.
[0009] In a second aspect, an embodiment of the present application further provides a device for a high-width memory to process low-width data, including:
[0010] A first storage module, configured to receive a plurality of arrays to be processed, and sequentially store each of the arrays to be processed starting from a preset starting address of the high-width memory. Each array to be processed includes a plurality of low-width data, and one address of the high-width memory stores at least 2 of the low-width data; a second storage module, configured to store the arrays to be processed into a preset plurality of vector registers, so that each of the low-width data in the i-th array to be processed is sequentially stored into the i-th channel of each of the vector registers, where i is all integers between 1 and n, and n is the number of arrays to be processed;
[0011] A third storage module, configured to sequentially obtain each of the arrays to be processed from the respective vector registers, and sequentially store each of the arrays to be processed starting from a preset starting address of the high-width memory, where one address of the high-width memory stores one low-width data.
[0012] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory and a processor, where:
[0013] The memory is used to store a computer program;
[0014] The processor is configured to read the computer program in the memory and execute the steps of the method for a high-width memory to process low-width data as in the first aspect.
[0015] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a readable computer program is stored, and when the program is executed by a processor, it implements the steps of the method for a high-width memory to process low-width data as in any one of the first aspects.
[0016] The method for a high-width memory to process low-width data provided by the embodiments of the present application uses a vector matrix to process an array to be processed. Each vector register in the vector matrix can process multiple low-width data simultaneously, significantly improving the speed and efficiency of data processing. In addition, when using a vector matrix to process data, a single instruction can operate on multiple data, further reducing the number of instructions required to process the data and improving the data processing efficiency. Finally, the embodiments of the present application achieve storing a low-width data at each address of the high-width memory. In this way, when processing low-width data in the high-width memory, the corresponding data can be directly accessed and modified using an offset address, without complex address calculations and data alignment operations, simplifying the data management and access and improving the processing efficiency of the high-width memory for processing low-width data.
[0017] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 The flowchart of the method for a high-width memory to process low-width data provided by the embodiments of the present application is shown.
[0020] Figure 2 The schematic diagram of a high-width memory provided by the embodiments of the present application receiving multiple arrays to be processed is shown.
[0021] Figure 3 The schematic diagram of a vector register storing an array to be processed provided by the embodiments of the present application is shown.
[0022] Figure 4 The schematic diagram of a high-width memory storing an array to be processed provided by the embodiments of the present application is shown.
[0023] Figure 5 Another flowchart of the method for a high-width memory to process low-width data provided by the embodiments of the present application is shown.
[0024] Figure 6 Another flowchart of the method for a high-width memory to process low-width data provided by the embodiments of the present application is shown.
[0025] Figure 7 The schematic diagram of the device for a high-width memory to process low-width data provided by the embodiments of the present application is shown.
[0026] Figure 8 Shows a schematic diagram of an electronic device provided by an embodiment of the present application.
[0027] Figure 9 Shows a schematic diagram of a computer storage medium provided by an embodiment of the present application. Detailed implementation manners
[0028] The following details the implementation manners of the present application. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as a limitation to the present application.
[0029] In the embodiments of the present application, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0030] In the description of the embodiments of the present application, words such as "example" or "for example" are used to indicate an example, illustration or description. Any embodiment or design described as "example" or "for example" in the embodiments of the present application is not construed as being more preferred or having more advantages than another embodiment or design. The use of words such as "example" or "for example" is intended to present relative concepts in a clear manner.
[0031] In addition, "a plurality of" in the embodiments of the present application means two or more. In view of this, "a plurality of" in the embodiments of the present application can also be understood as "at least two". "At least one" can be understood as one or more, for example, understood as one, two or more. For example, including at least one means including one, two or more, and does not limit which ones are included. For example, including at least one of A, B, and C, then what can be included is A, B, C, A and B, A and C, B and C, or A and B and C.
[0032] It should be noted that in the embodiments of the present application, "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / ", unless otherwise specified, generally represents an "or" relationship between the front and back associated objects.
[0033] An embodiment of the present application provides a method for a high-width memory to process low-width data. Figure 1 The flowchart of the method for a high-width memory to process low-width data provided by the embodiment of the present application is shown. As Figure 1 shown, the method includes:
[0034] Step S100: Receive multiple arrays to be processed, and sequentially store each array to be processed starting from a preset starting address in the high-width memory. Each array to be processed includes multiple low-width data, and at least 2 low-width data are stored in one address of the high-width memory.
[0035] Exemplarily, Figure 2 The schematic diagram of the high-width memory provided by the embodiment of the present application receiving multiple arrays to be processed is shown. As Figure 2 shown, taking a 32-bit memory as an example, the 32-bit memory receives four arrays to be processed, namely A, B, C, and D. Each array to be processed includes 5 low-width data (i.e., A0, A1, A2, A3, A4). The 32-bit memory sequentially stores the low-width data of each array to be processed starting from the preset starting address 0X00, that is, A0 and A1 are stored at address 0X00, A2 and A3 are stored at address 0X04, A4 and B0 are stored at address 0X08, and so on. D3 and D4 are stored at address 0X24.
[0036] It can be understood that in the embodiment of the present application, the starting address of the high-width memory can be preset according to the actual situation. For example, when addresses 0X00 and 0X04 are occupied, data A0 and A1 can also be stored starting from address 0X08, and the remaining data are sequentially stored in the remaining addresses.
[0037] It can be understood that Figure 2 only an example of storing 2 low-width data in one address of the high-width memory is shown, but the embodiment of the present application does not limit the number of low-width data that can be stored in one address of the high-width memory. For example, when each low-width data is 8-bit data, at most 4 low-width data can be stored in one address of the 32-bit memory.
[0038] Step S200: Store the arrays to be processed into a preset multiple vector registers, so that the low-width data in the i-th array to be processed are sequentially stored into the i-th channel of each vector register, where i is all integers between 1 and n, and n is the number of arrays to be processed stored in the high-width memory.
[0039] Exemplarily, Figure 3 The schematic diagram of the vector register storing the arrays to be processed provided by the embodiment of the present application is shown. As Figure 3As shown, taking a 32-bit memory as an example, each low-width data A0, A1, A2, A3, A4 in the first array A to be processed is sequentially stored into the first channel of vector registers Q0, Q1, Q2, Q3, Q4 respectively. Each low-width data B0, B1, B2, B3, B4 in the second array B to be processed is sequentially stored into the second channel of vector registers Q0, Q1, Q2, Q3, Q4 respectively, and so on. Each low-width data C0, C1, C2, C3, C4 in the third array C to be processed is sequentially stored into the third channel of vector registers Q0, Q1, Q2, Q3, Q4 respectively. Each low-width data D0, D1, D2, D3, D4 in the fourth array D to be processed is sequentially stored into the fourth channel of vector registers Q0, Q1, Q2, Q3, Q4 respectively.
[0040] Step S300: Sequentially obtain each array to be processed from each vector register, and sequentially store each array to be processed starting from a preset starting address in the high-width memory. One low-width data is stored in one address of the high-width memory.
[0041] Exemplarily, Figure 4 shows a schematic diagram of storing an array to be processed in the high-width memory provided by an embodiment of the present application. As Figure 4 shown, taking a 32-bit memory as an example, the 32-bit memory sequentially stores the low-width data of each array to be processed starting from the preset starting address 0X00. That is, A0 is stored at address 0X00, A1 is stored at address 0X04, A2 is stored at address 0X08, A3 is stored at 0X0C, A4 is stored at 0X10... Thus, one low-width data is stored in each address of the high-width memory, so that when processing low-width data in the high-width memory, the corresponding low-width data can be directly accessed and modified using the offset address.
[0042] The method for processing low-width data by the high-width memory provided by the embodiment of the present application stores the array to be processed by using vector registers. Each vector register can process multiple low-width data simultaneously, significantly improving the speed and efficiency of data processing. In addition, when using a vector matrix to process data, a single instruction can operate on multiple data, further reducing the number of instructions required to process data and improving the data processing efficiency. Finally, the embodiment of the present application realizes storing one low-width data in each address of the high-width memory, so that when processing low-width data in the high-width memory, the corresponding data can be directly accessed and modified using the offset address, without complex address calculation and data alignment operations, simplifying the data management and access, and improving the processing efficiency of the high-width memory for processing low-width data.
[0043] It can be understood that the high-width memory is the memory of a processor device such as a high-width microcontroller / microcontroller / central processing unit, etc. That is, the method for the high-width memory to process low-width data provided in the embodiments of the present application is applied to a processor device such as a high-width microcontroller / microcontroller / central processing unit, etc., and the high-width microcontroller / microcontroller / central processing unit, etc. should also include multiple vector registers and support corresponding technical means to configure the vector registers to execute the method for the high-width memory to process low-width data. For example, when the high-width processor device supports SIMD (Single Instruction, Multiple Data) technology, it can be used to execute the method for the high-width memory to process low-width data provided in the embodiments of the present application. Another example is that when the high-width processor device supports ARM NEON (ARM Advanced SIMD Extension) technology, it can be used to execute the method for the high-width memory to process low-width data provided in the embodiments of the present application.
[0044] In some embodiments, in the method for the high-width memory to process low-width data provided in the embodiments of the present application, in step S100: when receiving multiple arrays to be processed and sequentially storing each array to be processed starting from a preset starting address in the high-width memory, in each address of the high-width memory, the low-width data is stored starting from the high-order bit of the high-width memory address.
[0045] Exemplarily, as Figure 2 shown, taking a 32-bit memory as an example, the 32-bit memory sequentially stores the low-width data of each array to be processed starting from the preset starting address 0X00. That is, A0 is stored in the high-order bit of address 0X00, A1 is stored in the low-order bit of address 0X00, A2 is stored in the high-order bit of address 0X04, A3 is stored in the low-order bit of address 0X04, and so on. D3 is stored in the high-order bit of address 0X24, and D4 is stored in the low-order bit of address 0X24. Optionally, when the low-width data is 16-bit data, the two 16-bit data in the same address are respectively stored in the high 16 bits and the low 16 bits of the address; when the low-width data is 8-bit data and it is set that each address of the 32-bit memory stores 4 8-bit data, the four 8-bit data in the same address will be stored starting from the high-order bit of the address.
[0046] For the method for the high-width memory to process low-width data provided in the embodiments of the present application, when the 32-bit memory receives multiple arrays to be processed, it stores the low-width data starting from the high-order bit of the high-width memory address, fixing the storage positions of the low-width data, so that the low-width data can be directly read when read later, simplifying the access and management of the data and reducing the complexity in processing the data.
[0047] In some embodiments, in the method for a high-width memory to process low-width data provided by the embodiments of the present application, in step S100: in the step of receiving a plurality of arrays to be processed and sequentially storing each array to be processed starting from a preset starting address in the high-width memory,
[0048] When the bit width of the low-width data is not an integer multiple of 8, any data is filled to make the bit width of the low-width data an integer multiple of 8. Optionally, in the embodiments of the present application, the filling can be performed after the data is stored in the high-width memory, or can be performed before the data is stored in the high-width memory. The embodiments of the present application do not limit this.
[0049] Exemplarily, as Figure 2 shown, taking a 32-bit memory as an example, assuming the low-width data is 9-bit data, then the 9-bit data is filled with binary numbers 0 or 1 to be 16-bit data, so that the address of the 32-bit memory stores data in units of 16 bits. Assuming the low-width data is 7-bit data, then the 9-bit data is filled with binary numbers 0 or 1 to be 8-bit data, so that the address of the 32-bit memory stores data in units of 8 bits.
[0050] In the method for a high-width memory to process low-width data provided by the embodiments of the present application, by filling invalid data to make the bit width of the low-width data an integer multiple of 8 when the bit width of the low-width data is not an integer multiple of 8, the problem of unaligned access in the subsequent process of processing low-width data is avoided.
[0051] In some embodiments, Figure 5 Another flowchart showing the method for a high-width memory to process low-width data provided by the embodiments of the present application is shown. As Figure 5 shown, in the method for a high-width memory to process low-width data provided by the embodiments of the present application, after step S100 and before step S200, it further includes:
[0052] Step S150: Configure a vector matrix according to the quantity of low-width data in the array to be processed and the bit width of each low-width data. The vector matrix includes the quantity of vector registers and the channel bit width of each vector register. Optionally, the quantity of vector registers is set to store the corresponding quantity of low-width data, and the channel bit width of each vector register is set to store one low-width data in each channel. Among them, the total bit width of the vector registers remains unchanged. When the channel bit width of the vector registers is determined, the channel quantity of the vector registers will also be determined.
[0053] It can be understood that the method for a high-width memory to process low-width data provided by the embodiments of the present application is applied to processor devices such as high-width microcontrollers / microcontrollers / central processing units, and such processor devices should also include multiple vector registers and support corresponding technical means to implement configuring a vector matrix to execute the method for a high-width memory to process low-width data. For example, when a high-width processor device supports SIMD (Single Instruction, Multiple Data) technology, it can be used to execute the method for a high-width memory to process low-width data provided by the embodiments of the present application. For another example, when a high-width processor device supports ARM NEON (ARM Advanced SIMD Extension) technology, it can be used to execute the method for a high-width memory to process low-width data provided by the embodiments of the present application.
[0054] In some embodiments, in the method for a high-width memory to process low-width data provided by the embodiments of the present application, in step S150: in the step of configuring a vector matrix according to the number of low-width data in the array to be processed and the bit width of each low-width data,
[0055] The number of vector registers is the same as the number of low-width data in an array to be processed, and the channel bit width of each vector register is the quotient of the bit width of the vector register and the bit width of the low-width data. Optionally, the number of vector registers is set to store the corresponding number of low-width data, and the channel bit width of each vector register is set to store one low-width data in each channel. Since the total bit width of the vector register remains unchanged, when the channel bit width of the vector register is determined, the number of channels of the vector register will also be determined. Specifically, the number of channels of the vector register is the quotient of the total bit width of the vector register and the channel bit width. For example, when the total bit width of the vector register is 128 bits and the channel bit width is 16 bits, the vector register includes 8 channels; when the channel bit width is 8 bits, the vector register includes 16 channels.
[0056] It can be understood that when the initial bit width of the low-width data is not an integer multiple of 8, the vector matrix should be configured according to the bit width of the filled low-width data.
[0057] In some embodiments, Figure 6 shows another flowchart of the method for a high-width memory to process low-width data provided by the embodiments of the present application. As Figure 6 shown, in step S300: sequentially obtaining each array to be processed from each vector register and sequentially storing each array to be processed starting from the preset starting address of the high-width memory, where one address of the high-width memory stores one low-width data, includes:
[0058] Step S310: Configure the array cache vector register and the channel bit width of the array cache vector register, and store the low-bit-width data stored in the j-th channel of each vector register in the vector matrix into each channel of the array buffer vector register in sequence, where the initial value of j is 1. At this time, the data stored in the array cache vector register is a to-be-processed array that has been aligned. Optionally, the array cache vector register has the same configuration as each vector register in the vector matrix, and its channel bit width is the same as the bit width of the low-bit-width data to store the low-bit-width data.
[0059] Step S320: Store the low-bit-width data stored in each channel of the array cache vector register into the high-bit-width memory in sequence, and one low-bit-width data is stored at one address of the high-bit-width memory.
[0060] Step S330: Increase the value of j by 1, and return to execute Step S410: Configure the array cache vector register and the channel bit width of the array cache vector register, and store the low-bit-width data stored in the j-th channel of each vector register in the vector matrix into each channel of the array buffer vector register in sequence until all the low-bit-width data stored in each vector register in the vector matrix are stored in the high-bit-width memory. Optionally, in the embodiments of the present application, any means can be used to detect that all the low-bit-width data stored in each vector register in the vector matrix are stored in the high-bit-width memory. For example, in the embodiments of the present application, it can be detected that the value of j is greater than the number of to-be-processed arrays to terminate the return to execute Step S310. For another example, in the embodiments of the present application, it can also be detected that there is no data stored in each vector register in the vector matrix to terminate the return to execute Step S310. For another example, in the embodiments of the present application, it can also be detected that the high-bit-width memory has stored each to-be-processed array to terminate the return to execute Step S310.
[0061] Exemplarily, taking a 32-bit memory as an example, the low-width data A0, A1, A2, A3, A4 stored in the first channel of each vector register in the vector matrix are sequentially stored in the first channel, the second channel, the third channel, the fourth channel, and the fifth channel of the array cache vector register respectively. Then, the low-width data A0, A1, A2, A3, A4 stored in each channel of the array cache vector register are respectively stored at addresses 0X00, 0X04, 0X08, 0X0C, 0X10 of the 32-bit memory. Thus, the first array A to be processed has been completely stored in the 32-bit memory. Next, continue to sequentially store the low-width data B0, B1, B2, B3, B4 stored in the second channel of each vector register in the vector matrix in the first channel, the second channel, the third channel, the fourth channel, and the fifth channel of the array cache vector register respectively. Then, the low-width data B0, B1, B2, B3, B4 stored in each channel of the array cache vector register are stored starting from address 0X14 of the 32-bit memory. By this method, finally, all the arrays A, B, C, D to be processed are stored in the 32-bit memory.
[0062] The method for processing low-width data by a high-width memory provided in the embodiments of the present application sets an array cache vector register to store each array to be processed, and further realizes the alignment processing of the arrays to be processed in the array cache vector register. Transmitting the arrays to be processed from the array cache vector register to the high-width memory realizes storing one low-width data at each address in the high-width memory. In this way, when processing low-width data in the high-width memory, the corresponding data can be directly accessed and modified using the offset address, without complex address calculation and data alignment operations, simplifying the management and access of data.
[0063] In some embodiments, in the method for processing low-width data by a high-width memory provided in the embodiments of the present application, in step S300: sequentially obtaining each array to be processed from each vector register and sequentially storing each array to be processed starting from a preset starting address in the high-width memory, the low-width data is stored in the high-order bits of each address, or the low-width data is stored in the low-order bits of each address.
[0064] Exemplarily, taking a 32-bit memory as an example, assuming the low-width data is 16-bit data, when step S300 is executed, the low-width data will be stored in the lower 16 bits or the upper 16 bits of the 32-bit memory address.
[0065] In some embodiments, in the method for processing low-width data by a high-width memory provided in the embodiments of the present application, the high-width memory is a 32-bit wide memory, and each low-width data is 8-bit data or 16-bit data. Optionally, the embodiments of the present application are applied to a 32-bit wide memory to process 8-bit or 16-bit low-width data.
[0066] The method for a high-width memory to process low-width data provided by the embodiments of the present application is applied to processor devices such as high-width microcontrollers / microcontrollers / central processing units. A vector matrix is used to process an array to be processed. Each vector register in the vector matrix can process multiple low-width data simultaneously, significantly improving the speed and efficiency of data processing. In addition, when using a vector matrix to process data, a single instruction can operate on multiple data, further reducing the number of instructions required to process data and improving data processing efficiency. Finally, the embodiments of the present application implement storing one low-width data at each address of the high-width memory. In this way, when processing low-width data in the high-width memory, the corresponding data can be directly accessed and modified using the offset address, without complex address calculations and data alignment operations, simplifying data management and access, and improving the processing efficiency of the high-width memory for processing low-width data.
[0067] In addition, compared with directly receiving multiple arrays to be processed, writing the low-width data of multiple numbers to be processed into one address of the high-width memory respectively, this technical means still requires shift operations and data alignment operations on the low-width data of the arrays to be processed. Since the process of the processor performing operations on data is relatively slow, the method for a high-width memory to process low-width data provided by the embodiments of the present application does not require any address calculations and data alignment operations, significantly improving the processing efficiency of the high-width memory for processing low-width data.
[0068] Based on the above method for a high-width memory to process low-width data, the embodiments of the present application provide a device for a high-width memory to process low-width data. Figure 7 The schematic diagram of the device for a high-width memory to process low-width data provided by the embodiments of the present application is shown. As Figure 7 shown, the device includes:
[0069] A first storage module, configured to receive multiple arrays to be processed and sequentially store each array to be processed starting from a preset start address of the high-width memory. Each array to be processed includes multiple low-width data, and at least 2 low-width data are stored in one address of the high-width memory.
[0070] A second storage module, configured to store the array to be processed into a preset multiple vector registers, so that the low-width data in the i-th array to be processed are sequentially stored into the i-th channel of each vector register, where i is all integers between 1 and n, and n is the number of arrays to be processed.
[0071] A third storage module, configured to sequentially obtain each array to be processed from each vector register and sequentially store each array to be processed starting from a preset start address of the high-width memory. One low-width data is stored in one address of the high-width memory.
[0072] For other details of how each module in the above device for processing low-bit-width data with a high-bit-width memory implements the above technical solution, reference can be made to the description in the method for processing low-bit-width data with a high-bit-width memory provided in the above invention embodiment, which will not be elaborated here.
[0073] Based on the above method for processing low-bit-width data with a high-bit-width memory, an embodiment of the present application further provides an electronic device. Figure 8 The schematic diagram of the electronic device provided by the embodiment of the present application is shown. As Figure 8 shown, the electronic device provided by the embodiment of the present application includes a processor 81 and a memory 82 coupled to the processor 81. The memory 82 stores a computer program. When the computer program is executed by the processor 81, the processor 81 is caused to execute the steps of the method for processing low-bit-width data with a high-bit-width memory in the above embodiment.
[0074] For other details of how the processor 81 in the above electronic device implements the above technical solution, reference can be made to the description in the method for processing low-bit-width data with a high-bit-width memory provided in the above invention embodiment, which will not be elaborated here.
[0075] Among them, the processor 81 can also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 81 may be an integrated circuit chip with signal processing capabilities; the processor 81 can also be a general-purpose processor, a DSP (Digital Signal Process, digital signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gata Array, field programmable gate array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor, or the processor 81 can also be any conventional processor, etc.
[0076] Based on the above method for processing low-bit-width data with a high-bit-width memory, an embodiment of the present application further provides a computer-readable storage medium. Figure 9 The schematic diagram of the computer storage medium provided by the embodiment of the present application is shown. As Figure 9As shown in the figure, the embodiment of the present application further provides a computer-readable storage medium, on which a readable computer program 91 is stored; wherein, the computer program 91 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, magnetic disks or optical discs, ROM (Read-Only Memory), RAM (Random Access Memory), etc., which can store program codes, or terminal devices such as computers, servers, mobile phones, and tablets.
[0077] The above is only a preferred embodiment of the present application and does not impose any formal restrictions on the present application. Although the present application has been disclosed above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some modifications or refinements to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present application. However, as long as it does not depart from the content of the technical solution of the present application, any brief modifications, equivalent changes and refinements made to the above embodiments based on the technical essence of the present application still fall within the scope of the technical solution of the present application.
Claims
1. A method for processing low bit width data with a high bit width memory, characterized in that: include: Receive multiple arrays to be processed, and sequentially store the arrays to be processed starting from a preset starting address of a high-bit-width memory, each array to be processed includes multiple low-bit-width data, and one address of the high-bit-width memory stores at least 2 of the low-bit-width data; The array to be processed is stored in a plurality of preset vector registers, so that each low-bit width data in the i-th array to be processed is sequentially stored in the i-th channel of each vector register, where i is all integers between 1 and n, and n is the number of the array to be processed; The arrays to be processed are sequentially obtained from the vector registers, and the arrays to be processed are sequentially stored starting from a preset start address of the high bit width memory, wherein one address of the high bit width memory stores one low bit width data.
2. The method for processing low bit width data by a high bit width memory according to claim 1, characterized in that: In the step of receiving a plurality of arrays to be processed and sequentially storing each of the arrays to be processed starting from a preset starting address of a high bit width memory, In each address of the high bit width memory, the low bit width data is stored starting from the high bit of the high bit width memory address.
3. The method for processing low bit width data by a high bit width memory according to claim 1, characterized in that: In the step of receiving a plurality of arrays to be processed and sequentially storing each of the arrays to be processed starting from a preset starting address of a high bit width memory, When the bit width of the low bit width data is not an integer multiple of 8, arbitrary data is padded to make the bit width of the low bit width data an integer multiple of 8.
4. The method for processing low bit width data with a high bit width memory according to claim 1, characterized in that: After the step of receiving a plurality of arrays to be processed and sequentially storing each of the arrays to be processed starting from a preset starting address of a high-bit-width memory, and before the step of storing the arrays to be processed into a plurality of preset vector registers so that each low-bit-width data in the i-th array to be processed is sequentially stored into the i-th channel of each of the vector registers, the method further includes: A vector matrix is configured according to the number of low-bit-width data in the array to be processed and the bit width of each of the low-bit-width data, wherein the vector matrix includes the number of the vector registers and the channel bit width of each of the vector registers.
5. The method for processing low bit width data by a high bit width memory as claimed in claim 4, characterized in that: In the step of configuring the vector matrix according to the number of low-bit-width data in the array to be processed and the bit width of each low-bit-width data, The number of the vector registers is the same as the number of low-bit-width data in one array to be processed, and the channel bit width of each vector register is the quotient of the bit width of the vector register and the bit width of the low-bit-width data.
6. The method for processing low bit width data with a high bit width memory according to claim 4, characterized in that: The step of sequentially acquiring each array to be processed from each vector register and sequentially storing each array to be processed starting from a preset starting address of the high bit width memory comprises: Configure the array cache vector register and the channel bit width of the array cache vector register, and store the low bit width data stored in the j-th channel of each vector register in the vector matrix into each channel of the array cache vector register in sequence, where the initial value of j is 1; The low-bit width data stored in each channel of the array cache vector register is stored in the high-bit width memory in sequence, and one address of the high-bit width memory stores one low-bit width data; Increase the value of j by 1, and return to the execution step: configure the array cache vector register and the channel bit width of the array cache vector register, and store the low-bit-width data stored in the j-th channel of each vector register in the vector matrix in sequence into each channel of the array buffer vector register until the low-bit-width data stored in each vector register in the vector matrix are all stored in the high-bit-width memory.
7. The method for processing low bit width data with a high bit width memory according to claim 1, characterized in that: In the step of sequentially acquiring each array to be processed from each vector register and sequentially storing each array to be processed starting from a preset starting address of the high bit width memory, The low bit width data is stored in the high bit of each address, or the low bit width data is stored in the low bit of each address.
8. A device for processing low bit width data with a high bit width memory, characterized in that: include: A first storage module is used to receive a plurality of arrays to be processed, and sequentially store the arrays to be processed starting from a preset starting address of a high-bitwidth memory, each array to be processed includes a plurality of low-bitwidth data, and one address of the high-bitwidth memory stores at least two of the low-bitwidth data; A second storage module is used to store the array to be processed into a plurality of preset vector registers, so that each low-bit width data in the i-th array to be processed is sequentially stored in the i-th channel of each vector register, where i is all integers between 1 and n, and n is the number of the array to be processed; The third storage module is used to sequentially obtain each of the arrays to be processed from each of the vector registers, and sequentially store each of the arrays to be processed starting from a preset starting address of the high-bit-width memory, wherein one address of the high-bit-width memory stores one low-bit-width data.
9. An electronic device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store computer programs; The processor is used to read the computer program in the memory and execute the steps of the method for processing low-bit-width data with a high-bit-width memory according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A readable computer program is stored thereon, and when the program is executed by a processor, the steps of the method for processing low-bit-width data with a high-bit-width memory as claimed in any one of claims 1 to 7 are implemented.