Data storage method and device, data loading method and device, electronic equipment and medium

By splicing the data contents in the same data channel identification in a single-instruction multi-threaded architecture processor, vector data is obtained and stored in memory, the problem of low storage efficiency of vector data is solved and more efficient storage operations are achieved.

CN119987661AActive Publication Date: 2025-05-13BEIJING X RING TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510045852.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-13
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In single-instruction multithreaded architecture processors, the storage/loading efficiency of vector data is low, especially when spanning multiple vector registers.

Method used

By splicing the channel to identify the data content in the same data channel, vector data is obtained and stored in memory, thereby reducing the number of memory accesses and improving storage efficiency.

Benefits of technology

With this method, it is possible to improve the storage efficiency of vector data without increasing the number of memory accesses, and reduce the data processing amount of storage operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987661A_ABST
    Figure CN119987661A_ABST
Patent Text Reader

Abstract

The invention provides a data storage method and device, a data loading method and device, electronic equipment and a medium, and relates to the technical field of data processing.The data storage method comprises the steps that a storage instruction is received, the storage instruction is used for indicating that data contents in at least two vector registers are stored in a storage, any vector register is provided with at least one data channel; in response to the storage instruction, splicing data contents in the data channels with the same channel identifier in the at least two vector registers to obtain vector data corresponding to each channel identifier; and storing the vector data to a memory. The vector data is obtained by splicing the data contents in the data channels with the same channel identifier, and then the vector data is stored in the memory, so that the access times of the memory can be reduced, and the storage efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data storage method, a data loading method, a device, an electronic device and a medium. Background Art

[0002] In a Single Instruction, Multiple Threads (SIMT) architecture processor, the Vector Load / Store Unit (VLSU) is mainly responsible for moving data between vector registers and memory. Specifically, VLSU supports storage (writing data from vector registers to memory) and loading (reading data from memory to vector registers) operations of vector data.

[0003] However, in practical applications, vector data storage / loading often occurs across multiple vector registers. The increase in the number of vector registers will affect the storage / loading efficiency of vector data. Therefore, how to improve the storage / loading efficiency of vector data has become an important technical issue. Summary of the invention

[0004] The present application aims to solve one of the technical problems in the related art at least to some extent.

[0005] To this end, the present application proposes a data storage method, a data loading method, an apparatus, an electronic device and a medium to obtain vector data by splicing the data content in a data channel with the same channel identifier, and storing the vector data in a memory, thereby reducing the number of memory accesses and improving storage efficiency.

[0006] In one aspect, an embodiment of the present application provides a data storage method, including:

[0007] receiving a storage instruction, wherein the storage instruction is used to instruct to store data contents in at least two vector registers into a memory, wherein any of the vector registers has at least one data channel;

[0008] In response to the storage instruction, concatenate the data contents in the data channels with the same channel identifiers in the at least two vector registers to obtain vector data corresponding to each of the channel identifiers;

[0009] The vector data is stored in the memory.

[0010] Another aspect of the present application provides a data loading method, including:

[0011] receiving a load instruction, wherein the load instruction is used to instruct to load vector data in a memory into at least two vector registers, any of the vector registers having multiple data channels;

[0012] In response to the load instruction, the vector data is content-segmented to obtain corresponding data content;

[0013] The data content corresponding to the vector data is loaded into the data channel with the same channel identifier in the at least two vector registers.

[0014] Another aspect of the present application provides a data storage device, including:

[0015] A receiving module, used for receiving a storage instruction, wherein the storage instruction is used for instructing to store data contents in at least two vector registers into a memory, wherein any of the vector registers has at least one data channel;

[0016] a splicing module, configured to splice the data contents in the data channels with the same channel identifiers in the at least two vector registers in response to the storage instruction, to obtain vector data corresponding to each of the channel identifiers;

[0017] A storage module is used to store the vector data in the memory.

[0018] Another aspect of the present application provides a data loading device, including:

[0019] A receiving module, used for receiving a loading instruction, wherein the loading instruction is used for instructing to load vector data in a memory into at least two vector registers, any of the vector registers having multiple data channels;

[0020] A segmentation module, configured to segment the vector data into content in response to the load instruction to obtain corresponding data content;

[0021] A loading module is used to load the data content corresponding to the vector data into the data channel with the same channel identifier in the at least two vector registers.

[0022] Another aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the data storage method described in the first aspect or the data loading method described in the other aspect is implemented.

[0023] Another aspect of the present application provides a chip, including a processing circuit, wherein the processing circuit is used to implement the data storage method described in the first aspect or the data loading method described in the second aspect when executed.

[0024] Another aspect of the present application provides a non-temporary computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data storage method described in the first aspect or the data loading method described in the second aspect.

[0025] Another aspect of the present application provides a computer program product on which a computer program is stored. When the program is executed by a processor, the data storage method described in the first aspect or the data loading method described in the second aspect is implemented.

[0026] The data storage method, data loading method, device, electronic device and medium proposed in the present application receive a storage instruction, the storage instruction is used to instruct to store the data content in at least two vector registers into a memory, and any vector register has at least one data channel; in response to the storage instruction, the data content in the data channels with the same channel identifier in at least two vector registers is spliced ​​to obtain the vector data corresponding to each channel identifier; the vector data is stored in the memory. Among them, by splicing the data content in the data channels with the same channel identifier to obtain the vector data, and then storing the vector data in the memory, the number of memory accesses can be reduced and the storage efficiency can be improved.

[0027] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0029] Figure 1 A schematic diagram of a data storage method provided in an embodiment of the present application;

[0030] Figure 2 A schematic diagram of a flow chart of another data storage method provided in an embodiment of the present application;

[0031] Figure 3 A schematic diagram of a flow chart of another data storage method provided in an embodiment of the present application;

[0032] Figure 4 A schematic diagram of a data storage process provided in an embodiment of the present application;

[0033] Figure 5 Another schematic diagram of a data storage process according to an embodiment of the present application;

[0034] Figure 6 A flowchart of a data loading method provided in an embodiment of the present application;

[0035] Figure 7 A flowchart of another data loading method provided in an embodiment of the present application;

[0036] Figure 8 A flowchart of a data loading process provided in an embodiment of the present application;

[0037] Fig. 9 Another schematic diagram of the data loading process provided in the embodiment of the present application;

[0038] Fig.10 A schematic diagram of the structure of a data storage device provided in an embodiment of the present application;

[0039] Fig.11 A schematic diagram of the structure of a data loading device provided in an embodiment of the present application;

[0040] Fig.12 A block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0042] The storage layout of vector data in vector registers includes compact and fixed-length layouts. The fixed-length layout means that the vector register is divided into a fixed number of data channels (Lane), such as 32 / 64, and the available bit width of each Lane is also fixed. If the data type bit width does not exceed the allocated bit width of a single lane of the vector register, it is stored from the lowest bit, and the high bits are free; if the data type bit width is greater than the allocated bit width of a single lane of the vector register, multiple vector registers are used, and the Lanes with the same number in multiple vector registers are combined to store data. For example, the vector register bit width is 1024 bits (bits), which is fixedly divided into 32 Lanes, and 32 bits are allocated to each Lane: if the data type is 8 bits, each data only exists in the lower 8 bits of each lane of the vector register, and the upper 24 bits are free, so a total of 32 data are stored; if the data type is 64 bits, two vector registers are used, and the upper and lower 32 bits of the data are stored in the same numbered Lanes in the two vector registers.

[0043] Among them, for the case where the data type bit width is larger than the allocated bit width of a single lane of a vector register, in the related art, the instruction is split according to the number of vector registers. For example, for a 128-bit data type, a single instruction accesses 4 vector registers. Taking the storage process as an example, a single thread in the storage instruction needs to access the same numbered Lanes of 4 vector registers. The storage instruction is split into 4 sub-instructions, and the thread needs to access the memory multiple times in response to the 4 sub-instructions to complete the storage of the vector data. In summary, the above method has a low efficiency for the storage / loading of vector data.

[0044] In order to solve the above problems, the present application proposes a data storage method, a data loading method, an apparatus, an electronic device and a medium. The data storage method, data loading method, an apparatus, an electronic device and a medium of the embodiments of the present application are described below with reference to the accompanying drawings.

[0045] Figure 1 A flowchart of a data storage method provided in an embodiment of the present application. It should be noted that the data storage method of this embodiment can be applied to a data storage device. In some possible embodiments, the device can be configured in an electronic device or a chip so that the electronic device or the chip can perform a data storage function. In addition, in some possible embodiments, the data storage device can also be software in an electronic device, etc. Among them, the software is, for example, data storage software, etc. In addition, the data storage device can also be a VLSU in an electronic device. In the following embodiments, the execution subject is VLSU as an example for explanation.

[0046] like Figure 1 As shown, the method may include the following steps:

[0047] Step 101 : receiving a storage instruction, where the storage instruction is used to instruct to store data contents in at least two vector registers into a memory, and any vector register has at least one data channel.

[0048] The storage instruction may be used to indicate the source location of the data content in the vector register and the storage location of the data content in the memory. It should be noted that after receiving the storage instruction, the storage instruction may be decoded first.

[0049] The data channel has a channel identifier, which may include a channel number, a channel code, etc.; the channel identifiers corresponding to the data channels at the same position in at least two vector registers are the same.

[0050] In one embodiment of the present disclosure, data contents in at least two vector registers may indicate data of the same thread.

[0051] Step 102 , in response to a storage instruction, concatenate data contents in data channels with the same channel identifier in at least two vector registers to obtain vector data corresponding to each channel identifier.

[0052] The data contents in the data channels with the same channel identifiers in at least two vector registers constitute one vector data, and the data contents in at least two vector registers may correspond to at least one vector data.

[0053] It should be noted that the vector data is stored in at least two vector registers, from which it can be seen that the bit width of the vector data is greater than the allocated bit width of a single data channel in the vector register.

[0054] Step 103, storing the vector data into a memory.

[0055] The vector data and the storage location of the data content in the memory may be sent to the memory, so that the memory stores the vector data according to the storage location indicated by the storage instruction.

[0056] In the embodiment of the present application, a vector data only needs to access the memory once, and there is no need to access the memory multiple times for the data content included in the vector data, thereby improving the storage efficiency of the data.

[0057] In the data storage method of the embodiment of the present application, a storage instruction is received, and the storage instruction is used to instruct that the data content in at least two vector registers is stored in a memory, and any vector register has at least one data channel; in response to the storage instruction, the data content in the data channels with the same channel identifier in at least two vector registers is spliced ​​to obtain vector data corresponding to each channel identifier; and the vector data is stored in the memory. Among them, by splicing the data content in the data channels with the same channel identifier to obtain vector data, and then storing the vector data in the memory, the number of memory accesses can be reduced and storage efficiency can be improved.

[0058] Based on the above embodiments, Figure 2 A flowchart of another data storage method provided in an embodiment of the present application. It should be noted that the data storage method of this embodiment can be applied to a data storage device. In some possible embodiments, the device can be configured in an electronic device or a chip so that the electronic device or the chip can perform a data storage function. In addition, in some possible embodiments, the data storage device can also be software in an electronic device, etc. Among them, the software is, for example, data storage software, etc. In addition, the data storage device can also be a VLSU in an electronic device. In the following embodiments, the execution subject is VLSU as an example for explanation.

[0059] like Figure 2 As shown, the method comprises the following steps:

[0060] Step 201, receiving a storage instruction.

[0061] Step 202 , in response to a storage instruction, concatenate data contents in data channels with the same channel identifier in at least two vector registers to obtain vector data corresponding to each channel identifier.

[0062] The relevant explanations in the aforementioned embodiments are also applicable to step 201 and step 202, and the principles are the same, which will not be repeated here.

[0063] Step 203, concatenate the vector data corresponding to each channel identifier.

[0064] Wherein, when the data contents in at least two vector registers correspond to a plurality of vector data, the plurality of vector data are concatenated.

[0065] In one implementation of the embodiment of the present application, multiple vector data may be spliced ​​according to the arrangement order of data channels in any vector register.

[0066] Step 204 , dividing the spliced ​​vector data according to the set bit width to obtain at least one vector data group.

[0067] The set bit width may be a preset bit width; any vector data group includes at least one vector data.

[0068] It should be noted that a vector data group may correspond to a data group identifier, and the data group identifier may be used to indicate a storage batch of the vector data group.

[0069] Step 205 , storing the vector data in each vector data group into a memory in batches according to the vector data groups.

[0070] In the embodiment of the present application, the vector data in each vector data group may be stored in the memory in sequence. It should be noted that in the process of storing the vector data in any group of vector data groups in each batch, the VLSU may access the memory once or multiple times until all the vector data in the vector data group are stored in the memory.

[0071] In the data storage method of the embodiment of the present application, a storage instruction is received, and the storage instruction is used to instruct that the data contents in at least two vector registers are stored in a memory; in response to the storage instruction, the data contents in the data channels with the same channel identifiers in at least two vector registers are spliced ​​to obtain the vector data corresponding to each channel identifier; the vector data corresponding to each channel identifier is spliced; according to the set bit width, the spliced ​​vector data is divided to obtain at least one vector data group; according to the vector data group, the vector data in each vector data group is stored in the memory in batches. Among them, dividing the spliced ​​vector data into vector data groups and then storing the vector data in batches in the memory can reduce the data processing amount of each batch storage operation, improve storage efficiency, reduce the burden on the memory, and avoid storage congestion.

[0072] Based on the above embodiments, Figure 3 A flowchart of another data storage method provided in an embodiment of the present application. It should be noted that the data storage method of this embodiment can be applied to a data storage device. In some possible embodiments, the device can be configured in an electronic device or a chip so that the electronic device or the chip can perform a data storage function. In addition, in some possible embodiments, the data storage device can also be software in an electronic device, etc. Among them, the software is, for example, data storage software, etc. In addition, the data storage device can also be a VLSU in an electronic device. In the following embodiments, the execution subject is VLSU as an example for explanation.

[0073] like Figure 3 As shown, the method comprises the following steps:

[0074] Step 301: Receive a storage instruction, where the storage instruction is used to instruct to store data contents in at least two vector registers into a memory.

[0075] The memory includes at least one cache line, and the storage instruction may instruct to store data contents in at least two vector registers in the cache line.

[0076] Step 302 , in response to the storage instruction, concatenate the data contents in the data channels with the same channel identifier in at least two vector registers to obtain vector data corresponding to each channel identifier.

[0077] Step 303: splice the vector data corresponding to each channel identifier.

[0078] Step 304: segment the spliced ​​vector data according to the set bit width to obtain at least one vector data group.

[0079] Among them, steps 301 to 304 can refer to the relevant explanations in the aforementioned embodiments, and the principles are the same, so they will not be repeated here.

[0080] Step 305 : for any vector data group, determine the target vector data in the vector data group to be stored in the same cache line according to the cache line address of the vector data in the vector data group in the memory.

[0081] The cache line address is used to indicate a cache line for storing vector data; if the cache line addresses are the same, it can be indicated that the vector data will be stored in the same cache line.

[0082] The target vector data may refer to vector data in the vector data group that is to be stored in the same cache line. It should be noted that there are multiple target vector data.

[0083] It should be noted that the target vector data to be stored in the same cache line constitutes a data set, and any vector data group can correspond to one or more data sets. For example, assuming that the vector data group includes five vector data A, B, C, D, and E, A and B are to be stored in cache line a, and C, D, and E are to be stored in cache line b, then the vector data group corresponds to two data sets.

[0084] Step 306 , storing the target vector data to be stored in each cache line into the corresponding cache line in the memory.

[0085] Among them, for target vector data to be stored in the same cache line, multiple target vector data to be stored in the same cache line can be stored in the corresponding cache line through one memory access operation.

[0086] Since multiple target vector data are scattered and stored in corresponding cache lines, in order to make the target vector data the same as its offset in the accessed cache line and optimize the efficiency of data storage and access, in one implementation method of an embodiment of the present application, for any target vector data, the offset corresponding to the starting address is determined according to the starting address of the target vector data in the corresponding cache line; based on the offset, the target vector data belonging to the same cache line are rearranged through a crossbar matrix (Crossbar) to obtain the rearranged data of each cache line, and the rearranged position of each target vector data belonging to the same cache line in the rearranged data is determined; based on the rearranged position of each target vector data in the same cache line, a byte mask corresponding to the rearranged data of the corresponding cache line is generated; the rearranged data of each cache line and the corresponding byte mask are sent to the memory for storage.

[0087] Among them, the starting address of the target vector data in the corresponding cache line can be obtained according to the storage instruction. The starting address of the target vector data in the corresponding cache line is different, and the corresponding offset is different; the positions of each target vector data in the rearranged data may be non-adjacent, and the byte mask corresponding to the rearranged data is used by the memory to locate the position of the target vector data in the rearranged data. For example, the byte mask can be a 0-1 sequence, 1 represents the position of the target vector data (valid data bit), and 0 represents other positions (invalid data bit).

[0088] In an embodiment of the present application, target vector data belonging to the same cache line are rearranged based on the offset to obtain rearranged data, which can make the target vector data the same as its offset in the accessed cacheline, thereby improving data storage efficiency; the rearranged data and the corresponding byte mask are sent to the memory, which can enable the memory to quickly locate the target vector data in the rearranged data.

[0089] In one implementation of the embodiment of the present application, based on the starting address of the target vector data in the corresponding cache line, the starting addresses of the data contents corresponding to the target vector data in the cache line are determined; according to the starting addresses of the data contents corresponding to the target vector data in the cache line, the offset corresponding to the starting address of each data content in the cache line is determined; wherein, the offset corresponding to the starting address of the target vector data in the corresponding cache line includes the offset corresponding to the starting address of each data content in the cache line.

[0090] Among them, since the size of the data content is fixed, based on the starting address of the target vector data in the corresponding cache line, the starting address of each data content corresponding to the target vector data in the cache line can be calculated, and then based on the starting address of each data content corresponding to the target vector data in the cache line, the offset corresponding to the starting address of each data content in the cache line can be determined.

[0091] In an embodiment of the present application, the offset corresponding to the starting address of each data content corresponding to the target vector data in the cache line provides a more accurate positioning basis for data rearrangement. Therefore, the offset corresponding to the starting address of each data content corresponding to the target vector data in the cache line can further improve the data rearrangement effect.

[0092] In one implementation of the embodiment of the present application, the set bit width is determined based on the processable bit width of the cross-point matrix.

[0093] Based on the processable bit width of the cross-point matrix, the spliced ​​vector data group is divided, which can reduce the number of memory accesses and improve storage efficiency without changing the cross-point matrix.

[0094] like Figure 4As shown, Figure 4 The flowchart of the data storage process is shown in FIG. 4, where the thread ID in the figure indicates that the VLSU has T threads. Since the threads and data channel IDs in the VLSU correspond one to one, the thread ID in the figure can also refer to the channel ID. In the figure, 411 indicates four vector registers, each square represents a data channel, and the content in the square is the data content. The following is the specific storage process:

[0095] (1) The data contents in the data channels with the same channel identifier in the figure are spliced ​​to obtain the corresponding vector data; for example, the four D0s corresponding to the channel identifier 0 are spliced ​​to obtain the vector data corresponding to the channel identifier 0 (the data indicated by the arrow); the vector data corresponding to each channel identifier are spliced ​​to obtain the spliced ​​vector data 410.

[0096] (2) According to the processable bit width of Crossbar, the spliced ​​vector data is segmented to obtain at least one vector data group 401.

[0097] (3) For any vector data group, based on the cache line addresses of the vector data in the vector data group in the memory, determine the thread set 402 with the same corresponding cache line address; based on the thread set, determine the target vector data in the vector data group to be stored in the same cache line.

[0098] (4) For any target vector data, according to the starting address of the target vector data in the corresponding cache line, determine the offset corresponding to the starting address (located in 403, 403 includes the offset corresponding to each vector data of the vector data group); based on the offset, reorder the target vector data belonging to the same cache line through Crossbar404 to obtain the reordered data of each cache line, and determine the byte mask corresponding to the reordered data; send the reordered data of each cache line and the corresponding byte mask to the memory for storage, so as to store the reordered data in cacheline405.

[0099] In addition, the above embodiment describes the storage process of vector data when the vector data is stored in multiple vector registers (that is, the vector data bit width is greater than the allocated bit width of the data channel). It should be noted that the present application is also applicable to the case where the vector data bit width is less than or equal to the allocated bit width of the data channel, see Figure 5 , compared to Figure 4 , Figure 5 There is no need to splice the data content and split the vector data, that is, 501 is equivalent to a vector data group, and the content in each square in 501 is a vector data.

[0100] In the data storage method of the embodiment of the present application, a storage instruction is received, and in response to the storage instruction, the data contents in the data channels with the same channel identifier in at least two vector registers are spliced ​​to obtain the vector data corresponding to each channel identifier; the vector data corresponding to each channel identifier is spliced; according to the set bit width, the spliced ​​vector data is divided to obtain at least one vector data group; for any vector data group, according to the cache line address of the vector data in the vector data group in the memory, the target vector data to be stored in the same cache line in the vector data group is determined; the target vector data to be stored in each cache line is stored in the corresponding cache line in the memory. Among them, by accurately identifying and classifying the vector data to be stored in the same cache line, the goal of storing multiple target vector data at the same time in a single access is achieved, the number of memory accesses is reduced, and the storage efficiency is improved.

[0101] Based on the above embodiments, Figure 6 A flow chart of a data loading method provided in an embodiment of the present application. It should be noted that the data loading method of the present embodiment can be applied to a data loading device. In some possible embodiments, the device can be configured in an electronic device or a chip so that the electronic device or the chip can perform a data loading function. In addition, in some possible embodiments, the data loading device can also be software in an electronic device, etc. Among them, the software is, for example, data loading software, etc. In addition, the data loading device can also be a VLSU in an electronic device. Among them, the following embodiments are described by taking the execution subject as a VLSU as an example.

[0102] like Figure 6 As shown, the method comprises the following steps:

[0103] Step 601 : receiving a load instruction, where the load instruction is used to instruct to load vector data in a memory into at least two vector registers, and any vector register has multiple data channels.

[0104] The load instruction may be used to indicate a source position of the vector data in a memory, a loading position of the vector data in at least two vector registers, and a data amount of the vector data.

[0105] Step 602, in response to the load instruction, the vector data is content-segmented to obtain corresponding data content.

[0106] After the vector data is acquired, the vector data may be content segmented according to the allocated bit width of the data channel to obtain the corresponding data content.

[0107] Step 603: Load the data content corresponding to the vector data into the data channels with the same channel identifier in at least two vector registers.

[0108] Among them, the target channel identifier corresponding to the vector data can be determined according to the load instruction, and then the data content corresponding to the vector data can be loaded into the data channel corresponding to the target channel identifier in at least two vector registers.

[0109] In the data loading method of the embodiment of the present application, a loading instruction is received, and the loading instruction is used to instruct to load the vector data in the memory into at least two vector registers, and any vector register has multiple data channels; in response to the loading instruction, the vector data is content-segmented to obtain the corresponding data content; the data content corresponding to the vector data is loaded into the data channel with the same channel identifier in at least two vector registers. Compared with the solution of reading the data content of the vector data separately and then loading them separately in the related art, the present application directly reads the vector data, reduces the number of reads, and improves the loading efficiency.

[0110] Based on the above embodiments, Figure 7 A flow chart of a data loading method provided in an embodiment of the present application. It should be noted that the data loading method of this embodiment can be applied to a data loading device. In some possible embodiments, the device can be configured in an electronic device or a chip so that the electronic device or the chip can perform a data loading function. In addition, in some possible embodiments, the data loading device can also be software in an electronic device, etc. Among them, the software is, for example, data loading software, etc. In addition, the data loading device can also be a VLSU in an electronic device. Among them, the following embodiments are described by taking the execution subject as a VLSU as an example.

[0111] like Figure 7 As shown, the method comprises the following steps:

[0112] Step 701, receiving a loading instruction.

[0113] In one embodiment of the present disclosure, a load instruction may be used to indicate a source location of vector data in a memory, wherein the source location of the vector data in the memory includes a cache line address and a start address in the cache line.

[0114] Step 702 : According to any cache line address indicated by the load instruction, obtain from the memory the vector data stored in the cache line corresponding to the cache line address.

[0115] The same cache line address indicates that the vector data is stored in the same cache line.

[0116] In the embodiment of the present application, vector data stored in a cache line corresponding to a cache line address is obtained from a memory, so that a single read and simultaneous loading of multiple vector data can be achieved, thereby reducing the number of memory reads.

[0117] In one implementation of the embodiment of the present application, for any cache line, according to the starting address in the corresponding cache line address and the data amount information indicated by the load instruction, the storage data is read from the memory to obtain the storage data stored in the same cache line; based on the offset corresponding to the starting address, the storage data stored in the same cache line is rearranged through a cross point matrix to obtain the rearranged data of the corresponding cache line, wherein the rearranged data of any cache line includes at least one vector data.

[0118] The data volume information is used to indicate the data volume of the vector data, that is, the data bit width.

[0119] In the embodiment of the present application, by rearranging the storage data stored in the same cache line through a cross point matrix, the data layout can be optimized, thereby further improving the data loading efficiency.

[0120] Step 703, dividing the vector data into content to obtain corresponding data content.

[0121] Step 704, for any vector data in the rearranged data of any cache row, determine the target channel identifier corresponding to the vector data based on the target position of the vector data in the rearranged data and the target data group identifier of the vector data group to which the vector data belongs; wherein the target data group identifier is used to indicate the storage batch when the vector data is stored in the memory.

[0122] Wherein, in the vector data storage process before loading the vector data, the vector data has a vector data group to which it belongs, and the storage batch indicated by the data group identifier of the vector data group can be used to determine the channel identifier range corresponding to the vector data group. The target position of the vector data in the rearranged data can be used to determine the target channel identifier corresponding to the vector data in the channel identifier range. For example, the vector register has T data channels, and when the spliced ​​vector data is segmented during the vector data storage process, vector data group 1 includes vector data corresponding to channel identifiers 0-7, and vector data group 2 includes vector data corresponding to channel identifiers 8-15...; if the target data group identifier of the vector data group to which the vector data belongs is 2, and the target position of the vector data in the rearranged data is the second position from right to left, then the channel identifier range can be determined to be 8-15, and the target channel identifier is 9; if the target data group identifier of the vector data group to which the vector data belongs is 1, and the target position of the vector data in the rearranged data is the second position from right to left, then the channel identifier range can be determined to be 0-7, and the target channel identifier is 1.

[0123] Step 705: Load the data content corresponding to the vector data into the data channels corresponding to the target channel identifiers in at least two vector registers.

[0124] Among them, step 701 and step 703 can refer to the relevant explanations in the aforementioned embodiments, and the principles are the same, so they will not be repeated here.

[0125] like Figure 8 As shown, Figure 8 The following is a flowchart of the data loading process. The following is the specific storage process:

[0126] (1) According to any cache line address indicated by the load instruction, determine the thread set 802 that matches the cache line 805 corresponding to the cache line address; and read the vector data stored in the cache line 805 according to the thread set 802 .

[0127] (2) Based on the offset corresponding to the starting address of the vector data in the cache line 805 (located in 803, 803 includes the offset corresponding to each vector data in the cache line), the data is rearranged through Crossbar 804 to obtain 801.

[0128] (3) Determine the target channel identifier based on the target position of the vector data in the rearranged data 801 and the target data group identifier of the vector data group to which the vector data belongs.

[0129] (4) Content segmentation of the vector data in 801 is performed, and the segmented data content is loaded into the data channel corresponding to the target channel identifier in the vector register 811. It should be noted that 810 is for ease of understanding, and the actual loading process does not involve this process.

[0130] In addition, the above embodiment describes the loading process of vector data when the vector data bit width is greater than the allocated bit width of the data channel. It should be noted that the present application is also applicable to the case where the vector data bit width is less than or equal to the allocated bit width of the data channel. For details, see Fig. 9 , compared to Figure 8 , Fig. 9 There is no need to segment the vector data into content, that is, the content in each square in 901 is a vector data, rather than the data content of the vector data.

[0131] In the data loading method of the embodiment of the present application, a loading instruction is received; according to any cache line address indicated by the loading instruction, vector data stored in the cache line corresponding to the cache line address is obtained from the memory; the vector data is content-segmented to obtain the corresponding data content; for any vector data in the rearranged data of any cache line, the target channel identifier corresponding to the vector data is determined according to the target position of the vector data in the rearranged data and the target data group identifier of the vector data group to which the vector data belongs; wherein the target data group identifier is used to indicate the storage batch when the vector data is stored in the memory; the data content corresponding to the vector data is loaded into the data channel corresponding to the target channel identifier in at least two vector registers. Among them, obtaining the vector data stored in the cache line corresponding to the cache line address from the memory can realize a single read and load multiple vector data at the same time, thereby reducing the number of memory reads; data rearrangement of the storage data stored in the same cache line through the cross point matrix can optimize the data layout, thereby further improving the data loading efficiency.

[0132] Fig.10 A schematic diagram of the structure of a data storage device provided in an embodiment of the present application.

[0133] like Fig.10 As shown, the device may include:

[0134] A receiving module 1001 is used to receive a storage instruction, where the storage instruction is used to instruct to store data contents in at least two vector registers into a memory, and any vector register has at least one data channel;

[0135] A splicing module 1002 is used to splice data contents in data channels with the same channel identifier in at least two vector registers in response to a storage instruction to obtain vector data corresponding to each channel identifier;

[0136] The storage module 1003 is used to store the vector data into the memory.

[0137] Furthermore, in an implementation of the embodiment of the present application, the storage module 1003 is further configured to:

[0138] Splice the vector data corresponding to each channel identifier;

[0139] According to the set bit width, the spliced ​​vector data is segmented to obtain at least one vector data group;

[0140] According to the vector data groups, the vector data in each vector data group is stored in the memory in batches.

[0141] In an implementation of the embodiment of the present application, the memory includes a cache line, a storage module 1003, and is further used for:

[0142] For any vector data group, determine the target vector data in the vector data group to be stored in the same cache line according to the cache line address of the vector data in the vector data group in the memory;

[0143] The target vector data to be stored in each cache line is stored in the corresponding cache line in the memory.

[0144] In one implementation of the embodiment of the present application, the storage module 1003 is further configured to:

[0145] For any target vector data, according to the starting address of the target vector data in the corresponding cache line, determine the offset corresponding to the starting address;

[0146] Based on the offset, data rearrangement is performed on target vector data belonging to the same cache line through a cross point matrix to obtain rearranged data of each cache line, and a rearranged position of each target vector data belonging to the same cache line in the rearranged data is determined;

[0147] Based on the reordered positions of the target vector data of the same cache line, generating a byte mask corresponding to the reordered data of the corresponding cache line;

[0148] The reordered data and the corresponding byte mask of each cache line are sent to the memory for storage.

[0149] In one implementation of the embodiment of the present application, the storage module 1003 is further configured to:

[0150] Based on the starting address of the target vector data in the corresponding cache line, determine the starting address of each data content corresponding to the target vector data in the cache line;

[0151] According to the starting address of each data content corresponding to the target vector data in the cache line, determine the offset corresponding to the starting address of each data content in the cache line; wherein the offset corresponding to the starting address of the target vector data in the corresponding cache line includes the offset corresponding to the starting address of each data content in the cache line.

[0152] In an implementation of the embodiment of the present application, the set bit width is determined based on the processable bit width of the cross-point matrix.

[0153] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and will not be repeated here.

[0154] In the data storage device of the embodiment of the present application, a storage instruction is received, and the storage instruction is used to instruct that the data content in at least two vector registers is stored in the memory, and any vector register has at least one data channel; in response to the storage instruction, the data content in the data channels with the same channel identifier in at least two vector registers is spliced ​​to obtain the vector data corresponding to each channel identifier; and the vector data is stored in the memory. Among them, by splicing the data content in the data channels with the same channel identifier to obtain the vector data, and then storing the vector data in the memory, the number of memory accesses can be reduced and the storage efficiency can be improved.

[0155] Fig.11 A schematic diagram of the structure of a data loading device provided in an embodiment of the present application.

[0156] like Fig.11 As shown, the device may include:

[0157] A receiving module 1101 is used to receive a loading instruction, where the loading instruction is used to instruct to load vector data in a memory into at least two vector registers, and any vector register has multiple data channels;

[0158] A segmentation module 1102 is used to segment the vector data into content in response to a load instruction to obtain corresponding data content;

[0159] The loading module 1103 is used to load data content corresponding to the vector data into data channels with the same channel identifier in at least two vector registers.

[0160] Furthermore, in an implementation of the embodiment of the present application, the segmentation module 1102 is further configured to:

[0161] According to any cache line address indicated by the load instruction, obtain from the memory the vector data stored in the cache line corresponding to the cache line address;

[0162] Divide the vector data into content to obtain the corresponding data content.

[0163] In one implementation of the embodiment of the present application, the segmentation module 1102 is further configured to:

[0164] For any cache line, according to the start address in the corresponding cache line address and the data amount information indicated by the load instruction, the storage data is read from the memory to obtain the storage data stored in the same cache line;

[0165] Based on the offset corresponding to the start address, the storage data stored in the same cache line is rearranged through the cross point matrix to obtain the rearranged data of the corresponding cache line, wherein the rearranged data of any cache line includes at least one vector data.

[0166] In one implementation of the embodiment of the present application, the loading module 1103 is further used to:

[0167] For any vector data in the rearranged data of any cache line, determine the target channel identifier corresponding to the vector data according to the target position of the vector data in the rearranged data and the target data group identifier of the vector data group to which the vector data belongs; wherein the target data group identifier is used to indicate a storage batch when the vector data is stored in the memory;

[0168] The data content corresponding to the vector data is loaded into the data channel corresponding to the target channel identifier in at least two vector registers.

[0169] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and will not be repeated here.

[0170] The data loading device proposed in the embodiment of the present application receives a loading instruction, and the loading instruction is used to instruct to load the vector data in the memory into at least two vector registers, and any vector register has multiple data channels; in response to the loading instruction, the vector data is content-segmented to obtain the corresponding data content; the data content corresponding to the vector data is loaded into the data channel with the same channel identifier in at least two vector registers. Compared with the solution of reading the data content of the vector data separately and then loading them separately in the related art, the present application directly reads the vector data, reduces the number of reads, and improves the loading efficiency.

[0171] In order to implement the above embodiments, the present application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in the above method embodiments is implemented.

[0172] In order to implement the above-mentioned embodiments, the embodiments of the present application further provide a chip, including a processing circuit, wherein the processing circuit is used to implement the method described in the above-mentioned method embodiments when executed.

[0173] In order to implement the above embodiments, the present application also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in the above method embodiments is implemented.

[0174] In order to implement the above embodiments, the present application also proposes a computer program product on which a computer program is stored. When the computer program is executed by a processor, the method described in the above method embodiments is implemented.

[0175] Fig.12This is a block diagram of an electronic device provided in an embodiment of the present application. For example, the electronic device 1200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0176] Reference Fig.12 , the electronic device 1200 may include one or more of the following components: a processing component 1202 , a memory 1204 , a power component 1206 , a multimedia component 1208 , an audio component 1210 , an input / output (I / O) interface 1212 , a sensor component 1214 , and a communication component 1216 .

[0177] The processing component 1202 generally controls the overall operation of the electronic device 1200, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 1202 may include one or more processors 1220 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 1202 may include one or more modules to facilitate the interaction between the processing component 1202 and other components. For example, the processing component 1202 may include a multimedia module to facilitate the interaction between the multimedia component 1208 and the processing component 1202.

[0178] The memory 1204 is configured to store various types of data to support operations on the electronic device 1200. Examples of such data include instructions for any application or method operating on the electronic device 1200, contact data, phone book data, messages, pictures, videos, etc. The memory 1204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0179] The power component 1206 provides power to the various components of the electronic device 1200. The power component 1206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 1200.

[0180] The multimedia component 1208 includes a screen that provides an output interface between the electronic device 1200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1208 includes a front camera and / or a rear camera. When the electronic device 1200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0181] The audio component 1210 is configured to output and / or input audio signals. For example, the audio component 1210 includes a microphone (MIC), and when the electronic device 1200 is in an operation mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 1204 or sent via the communication component 1216. In some embodiments, the audio component 1210 also includes a speaker for outputting audio signals.

[0182] I / O interface 1212 provides an interface between processing component 1202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0183] The sensor assembly 1214 includes one or more sensors for providing various aspects of status assessment for the electronic device 1200. For example, the sensor assembly 1214 can detect the open / closed state of the electronic device 1200, the relative positioning of components, such as the display and keypad of the electronic device 1200, and the sensor assembly 1214 can also detect the position change of the electronic device 1200 or a component of the electronic device 1200, the presence or absence of user contact with the electronic device 1200, the orientation or acceleration / deceleration of the electronic device 1200, and the temperature change of the electronic device 1200. The sensor assembly 1214 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 1214 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 1214 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0184] The communication component 1216 is configured to facilitate wired or wireless communication between the electronic device 1200 and other devices. The electronic device 1200 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 1216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0185] In an exemplary embodiment, the electronic device 1200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0186] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1204 including instructions, and the instructions can be executed by a processor 1220 of the electronic device 1200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0187] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0188] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0189] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0190] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0191] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0192] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0193] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0194] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A data storage method, characterized in that: include: receiving a storage instruction, wherein the storage instruction is used to instruct to store data contents in at least two vector registers into a memory, wherein any of the vector registers has at least one data channel; In response to the storage instruction, concatenate the data contents in the data channels with the same channel identifiers in the at least two vector registers to obtain vector data corresponding to each of the channel identifiers; The vector data is stored in the memory.

2. The method according to claim 1, characterized in that The storing the vector data into the memory comprises: splicing the vector data corresponding to each of the channel identifiers; According to the set bit width, the spliced ​​vector data is segmented to obtain at least one vector data group; According to the vector data groups, the vector data in each of the vector data groups are stored in the memory in batches.

3. The method according to claim 2, characterized in that The memory includes a cache line, and storing the vector data in each of the vector data groups in batches into the memory according to the vector data groups includes: For any of the vector data groups, determining target vector data in the vector data group to be stored in the same cache line according to cache line addresses of the vector data in the vector data group in the memory; The target vector data to be stored in each of the cache lines is stored in the corresponding cache line in the memory.

4. The method according to claim 3, characterized in that The storing the target vector data to be stored in each of the cache lines into the corresponding cache line in the memory comprises: For any of the target vector data, determine an offset corresponding to the start address according to the start address of the target vector data in the corresponding cache line; Based on the offset, data rearrangement is performed on the target vector data belonging to the same cache line through a cross point matrix to obtain rearranged data of each cache line, and a rearranged position of each target vector data belonging to the same cache line in the rearranged data is determined; Based on the reordered positions of the target vector data of the same cache line, generating a byte mask corresponding to the reordered data of the corresponding cache line; The reordered data and the corresponding byte mask of each cache line are sent to the memory for storage.

5. The method according to claim 4, characterized in that The determining, according to the starting address of the target vector data in the corresponding cache line, an offset corresponding to the starting address comprises: Based on the starting address of the target vector data in the corresponding cache line, determining the starting address of each data content corresponding to the target vector data in the cache line; According to the starting addresses of the data contents corresponding to the target vector data in the cache line, determine the offsets corresponding to the starting addresses of the data contents in the cache line; wherein the offsets corresponding to the starting addresses of the target vector data in the corresponding cache line include the offsets corresponding to the starting addresses of the data contents in the cache line.

6. The method according to claim 4, characterized in that The set bit width is determined based on a processable bit width of the cross-point matrix.

7. A data loading method, characterized in that: include: receiving a load instruction, wherein the load instruction is used to instruct to load vector data in a memory into at least two vector registers, any of the vector registers having multiple data channels; In response to the load instruction, the vector data is content-segmented to obtain corresponding data content; The data content corresponding to the vector data is loaded into the data channel with the same channel identifier in the at least two vector registers.

8. The method according to claim 7, characterized in that In response to the load instruction, the vector data is content-segmented to obtain corresponding data content, including: According to any cache line address indicated by the load instruction, obtaining from the memory the vector data stored in the cache line corresponding to the cache line address; The vector data is content-segmented to obtain corresponding data content.

9. The method according to claim 8, characterized in that The acquiring, according to any cache line address indicated by the load instruction, vector data stored in a cache line corresponding to the cache line address from the memory comprises: For any cache line, according to the start address in the corresponding cache line address and the data amount information indicated by the load instruction, the memory is read to obtain the storage data stored in the same cache line; Based on the offset corresponding to the start address, the storage data stored in the same cache line is rearranged through a cross point matrix to obtain rearranged data of the corresponding cache line, wherein the rearranged data of any cache line includes at least one of the vector data.

10. The method according to claim 9, characterized in that The step of loading the data content corresponding to the vector data into the data channels with the same channel identifier in the at least two vector registers includes: For any of the vector data in the rearranged data of any cache line, determine the target channel identifier corresponding to the vector data according to the target position of the vector data in the rearranged data and the target data group identifier of the vector data group to which the vector data belongs; wherein the target data group identifier is used to indicate a storage batch when the vector data is stored in the memory; The data content corresponding to the vector data is loaded into the data channel corresponding to the target channel identifier in the at least two vector registers.

11. A data storage device, characterized in that: include: A receiving module, used for receiving a storage instruction, wherein the storage instruction is used for instructing to store data contents in at least two vector registers into a memory, wherein any of the vector registers has at least one data channel; a splicing module, configured to splice the data contents in the data channels with the same channel identifiers in the at least two vector registers in response to the storage instruction, to obtain vector data corresponding to each of the channel identifiers; A storage module is used to store the vector data in the memory.

12. A data loading device, characterized in that: include: A receiving module, used for receiving a loading instruction, wherein the loading instruction is used for instructing to load vector data in a memory into at least two vector registers, any of the vector registers having multiple data channels; A segmentation module, configured to segment the vector data into content in response to the load instruction to obtain corresponding data content; A loading module is used to load the data content corresponding to the vector data into the data channel with the same channel identifier in the at least two vector registers.

13. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1 to 6 or 7 to 10.

14. A chip, characterized in that: The method comprises a processing circuit, wherein the processing circuit is configured to implement the method according to any one of claims 1 to 6 or 7 to 10 when executed.

15. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform the steps of the method according to any one of claims 1 to 6 or 7 to 10.

16. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6 or 7 to 10.

Citation Information

Patent Citations

  • Data channel configuration method and device

    CN101425838A

  • Method and device for controlling transmission rate of wireless network, terminal equipment and storage medium

    CN108093444A

  • An apparatus and method for generating and processing a trace stream

    CN108345534A

  • Instructions and logic for lane-based strided scatter operations

    CN108369509A

  • Data splicing instruction processing method and data splicing instruction processing device

    CN111813447A