Data processing method of processor, electronic device, and storage medium
By performing block processing and broadcast operations on data sets, the problem of low processor data processing efficiency is solved and more efficient data operations are achieved.
Patent Information
- Application Number
- PCT/IB2024/062246
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-15
- Filing Date
- 2024-12-05
- Publication Date
- 2025-09-18
AI Technical Summary
Existing processors have problems with redundant memory access operations and low parallelism when performing data calculations, resulting in low computing efficiency.
By processing the data set in blocks and using the arrangement rule information to perform broadcast operations on multiple processing elements of the processor, the data set can be shared within rows and columns, avoiding the waste of writing intermediate results.
It improves the data processing efficiency of the processor, accelerates the calculation process through parallel processing, and reduces the time consumption of redundant memory access operations.
Smart Images

Figure IB2024062246_18092025_PF_FP_ABST
Abstract
Description
[0001] Processor Data Processing Method, Electronic Device, and Storage Medium TECHNICAL FIELD This application relates to large-scale model technology and the computer field, and more specifically, to a processor data processing method, electronic device, and storage medium. BACKGROUND In the related art, if a processor is required to perform calculations on two data sets, since both used data sets are stored in the computer's static random access memory (SRAM), the data sets need to be loaded from the SRAM into the computer's registers multiple times during the calculation process. This results in redundant memory access operations, thereby reducing computational efficiency. Furthermore, the low degree of parallelism between the multiple processing elements included in the processors of the related art also slows down the computation process. Therefore, the above analysis indicates that the related art still suffers from the technical problem of low processor data processing efficiency. Currently, no effective solution has been proposed to address this problem. SUMMARY OF THE INVENTION Embodiments of the present application provide a processor data processing method, electronic device, and storage medium to at least address the technical problem of low processor data processing efficiency. According to one aspect of the embodiments of the present application, a processor data processing method is provided. This method is applied to a data processing device, the data processing device including a processor comprised of multiple processing elements arranged according to permutation rule information. The method may include: obtaining a first data set and a second data set to be processed by the processor; performing block processing on the first data set and the second data set, respectively, to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; using the permutation rule information, dividing the first initial sub-data set along a first permutation direction and dividing the second initial sub-data set along a second permutation direction, respectively, to obtain multiple first target sub-data sets and multiple second target sub-data sets, wherein the first target sub-data sets are broadcasted to corresponding processing elements on the processor along the first permutation direction, and the second target sub-data sets are broadcasted to corresponding processing elements on the processor along the second permutation direction; performing operations on the first target sub-data sets and the second target sub-data sets in the corresponding processing elements to obtain output results of the corresponding processing elements; and determining an output result of the processor based on the output results of the processing elements. According to another aspect of an embodiment of the present application, a data processing method for a data stream processor is provided. The method is applied to a data processing device, the data processing device including a data stream processor comprised of multiple processing elements arranged according to the permutation rule information.The method may include: obtaining first and second data sets to be processed by a data stream processor from a large model; performing block processing on the first and second data sets, respectively, to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; utilizing permutation rule information, dividing the first initial sub-data set along a first permutation direction and dividing the second initial sub-data set along a second permutation direction, respectively, to obtain a first target sub-data set and multiple second target sub-data sets, wherein the first target sub-data set is broadcasted to corresponding processing elements on the processor along the first permutation direction, and the second target sub-data set is broadcasted to corresponding processing elements on the processor along the second permutation direction; performing operations on the first target sub-data set and the second target sub-data set in the corresponding processing elements to obtain output results of the corresponding processing elements; and determining an output result of the processor based on the output results of the processing elements. According to another aspect of an embodiment of the present application, another data processing method for a processor is provided. The method is applied to a data processing device, wherein the data processing device includes a processor, which may be composed of multiple processing elements arranged according to the permutation rule information. The method may include: obtaining a first data set and a second data set to be processed by a processor by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the first data set and the second data set; performing block processing on the first data set and the second data set, respectively, to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; using permutation rule information, dividing the first initial sub-data set along a first permutation direction and dividing the second initial sub-data set along a second permutation direction, respectively, to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets, wherein the first target sub-data set is broadcasted to corresponding processing elements on the processor along the first permutation direction, and the second target sub-data set is broadcasted to corresponding processing elements on the processor along the second permutation direction; performing operations on the first target sub-data set and the second target sub-data set in the corresponding processing elements to obtain output results of the corresponding processing elements; determining an output result of the processor based on the output results of the processing elements; and outputting the output result of the processor by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the output result of the processor. According to another aspect of an embodiment of the present application, a data processing system is provided.The system may include: a data acquisition terminal and a processor, the processor comprising multiple processing elements arranged according to arrangement rule information. The data acquisition terminal is configured to acquire a first data set and a second data set to be processed by the processor; the processor is configured to perform block processing on the first data set and the second data set, respectively, to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; using the arrangement rule information, the first initial sub-data set is divided along a first arrangement direction, and the second initial sub-data set is divided along a second arrangement direction, to obtain multiple first target sub-data sets and multiple second target sub-data sets; the first target sub-data sets are broadcasted to corresponding processing elements on the processor along the first arrangement direction, and the second target sub-data sets are broadcasted to corresponding processing elements on the processor along the second arrangement direction; the first target sub-data sets and the second target sub-data sets are input into corresponding processing elements for computation, obtaining output results of the corresponding processing elements; and an output result of the processor is determined based on the output results of the processing elements. According to another aspect of an embodiment of the present application, an electronic device is also provided. The electronic device may include a memory and a processor: the memory is configured to store computer-executable instructions, and the processor is configured to execute the computer-executable instructions. When the processor executes the computer-executable instructions, the processor implements the data processing method of the processor according to the embodiment of the present application. According to another embodiment of the present application, a processor is provided. The processor is configured to run a program, wherein when the program is executed, the processor executes the data processing method of the processor according to the embodiment of the present application. According to another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein when the program is executed, the device containing the storage medium is controlled to execute the data processing method of the processor according to the embodiment of the present application. According to another embodiment of the present application, a computer program product is provided. The computer program product includes a computer program, and when the processor executes the computer program, the computer program implements the data processing method of the processor according to the embodiment of the present application. In the embodiment of the present application, if a processor is required to process a first data set and a second data set to be multiplied, the first data set and the second data set can be obtained, and the first data set and the second data set can be processed in blocks to obtain corresponding first and second initial sub-data sets. The arrangement rule information can be used to divide the first initial sub-dataset in the first arrangement direction to obtain a plurality of divided first target sub-datasets.The second initial sub-dataset can be divided in the second arrangement direction to obtain multiple second target sub-datasets. The two target sub-datasets can be broadcast in the first or second arrangement direction to the corresponding processing elements. Operations are performed on the first and second target sub-datasets in the processing elements to obtain output results from each processing element. The output results from each processing element are integrated to obtain a final output result. In the embodiments of the present application, considering that operations on the two data sets can be processed in parallel by multiple processing elements in the processor, the purpose of accelerating operations is achieved. Furthermore, by broadcasting in the first and second arrangement directions, data in a row or column of the entire data set can be copied and broadcasted to other rows or columns, ensuring that the data in the rows or columns of the entire data set is the same. This allows for data sharing within the rows and columns, thereby avoiding the time-consuming operation of writing intermediate results, thereby improving the processor's data processing efficiency and resolving the technical problem of low processor data processing efficiency. It should be noted that the general description above and the detailed description below are merely examples and explanations of the present application and do not constitute limitations of the present application. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application.In the accompanying drawings: FIG1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method of a processor according to an embodiment of the present application; FIG2 is a structural block diagram of a computing environment of a data processing method of a processor according to an embodiment of the present application; FIG3 is a flow chart of a data processing method of a processor according to an embodiment of the present application; FIG4 is a flow chart of a data processing method of a data stream processor according to an embodiment of the present application; FIG5 is a flow chart of a data processing method of another processor according to an embodiment of the present application; FIG6 is a flow chart of a data processing system according to an embodiment of the present application; FIG7 is a schematic diagram of a computing unit structure according to an embodiment of the present application; FIG8 (a) is a schematic diagram of a matrix multiplication of a data stream processor in a related art; FIG8 (b) is a schematic diagram of a matrix multiplication of a data stream processor according to an embodiment of the present application; FIG9 (a) is a schematic diagram of a matrix cascade according to an embodiment of the present application; FIG9 (b) is a schematic diagram of another matrix cascade according to an embodiment of the present application; FIG10 is a schematic diagram of a matrix segmentation process according to an embodiment of the present application; FIG11 is a flow chart of a matrix multiplication optimization method for a data stream processor according to an embodiment of the present application; Figure 12 is a schematic diagram of a data processing device of a processor according to an embodiment of the present application; Figure 13 is a schematic diagram of a data processing device of a data stream processor according to an embodiment of the present application; Figure 14 is a schematic diagram of another data processing device of a processor according to an embodiment of the present application; Figure 15 is a structural block diagram of a computer terminal according to an embodiment of the present application; and Figure 16 is a block diagram of an electronic device for a data processing method of a processor according to an embodiment of the present application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS To help those skilled in the art better understand the present invention, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings. It should be understood that the described embodiments are merely some of the embodiments of the present application, and are not intended to be exhaustive. Based on the embodiments of the present application, all other embodiments derived by persons of ordinary skill in the art without inventive effort shall fall within the scope of protection of the present application. It should be noted that the terms "first," "second," and so on, in the specification and claims of the present application, and in the accompanying drawings, are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that the embodiments of the application described herein are capable of operation in sequences other than those illustrated or described herein.Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus. The technical solutions provided in this application are primarily implemented using large-scale model technology. Large-scale models herein refer to deep learning models with large-scale model parameters, typically including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. Large models, also known as foundational models, are pre-trained on large amounts of unlabeled corpora, producing pre-trained models with over 100 million parameters. These models are adaptable to a wide range of downstream tasks and exhibit good generalization capabilities. Examples include large-scale language models for information regression (LLMs) and multimodal pre-training models. It should be noted that in practical applications, large models can be fine-tuned using a small number of samples, allowing them to be applied to different tasks. For example, large models can be widely applied in fields such as natural language processing (NLP), computer vision, and speech processing. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image captioning (IC), and image generation. They can also be widely applied to natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. Therefore, the main application scenarios of large models include, but are not limited to, digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design. In the embodiments of this application, application generation is explained using the target language processing model obtained by training the generative large model in a conversational scenario as an example.First, some nouns or terms that appear in the description of the embodiments of this application are subject to the following interpretation: A data flow processor is a hardware architecture composed of numerous simple processing elements (PEs) arranged according to certain rules. The core concept is to allow data to flow within an array of execution units, reducing memory access times. Furthermore, the structure is regular and the wiring is uniform. Example 1 According to an embodiment of this application, a data processing method for a processor is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system, such as a set of computer-executable instructions. Furthermore, although the flowcharts illustrate a logical order, in some cases, the steps shown or described can be executed in a different order than that shown. The method embodiment provided in Example 1 of this application can be executed in a mobile terminal, a computer terminal, or a similar computing device. FIG1 is a hardware block diagram of a computer terminal (or mobile device) for implementing a data verification method according to an embodiment of the present application. As shown in FIG1 , the computer terminal 10 (or mobile device) may include one or more processors 102 (illustrated as 102a, 102b, ..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microcontroller unit (MCU) or a field programmable gate array (FPGA)), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, the computer terminal 100 (or mobile device) may include a display, an input / output (I / O) interface, a universal serial bus (USB) port (which may be included as one of the ports of a BUS), a network interface, a power supply, and / or a camera. Those skilled in the art will appreciate that the structure shown in FIG1 is merely illustrative and does not limit the structure of the electronic device. For example, the computer terminal 10 may include more or fewer components than those shown in FIG. 1 , or may have a configuration different from that shown in FIG. The hardware structure block diagram shown in FIG. 1 can serve not only as an exemplary block diagram of the aforementioned computer terminal 10 (or mobile device), but also as an exemplary block diagram of the aforementioned server. In an optional embodiment, FIG. 2 shows a block diagram of an embodiment using the computer terminal 10 (or mobile device) shown in FIG. 1 as a computing node in a computing environment 201.FIG2 is a block diagram of a computing environment for a data verification method according to an embodiment of the present application. As shown in FIG2 , computing environment 201 includes multiple computing nodes (e.g., servers) (illustrated as 210-1 and 210-2 in the figure) running on a distributed network. Each computing node includes local processing and memory resources. End users 202 can remotely run applications or store data in computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in computing environment 201, representing services "A," "D," "E," and "H," respectively. End users 202 can provision and access services through a web browser or other software application on a client. In some embodiments, end user 202's provision and / or request can be provided to an ingress gateway 230. Ingress gateway 230 can include a corresponding agent to handle provision and / or requests for services (one or more services provided in computing environment 201). Services are provided or deployed based on various virtualization technologies supported by computing environment 201. In some embodiments, services can be provided using virtual machine (VM)-based virtualization, container-based virtualization, and / or similar approaches. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While VMs virtualize machines, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance. In one embodiment based on container virtualization, several service containers can be assembled into a pod (e.g., a Kubernetes pod). For example, as shown in Figure 2, Pod 240-1, 240-2, 240-N (collectively referred to as Pod). A Pod may include an agent 245 and one or more containers.
[0002] 242-1, 242-2, and 242-M (collectively referred to as containers). One or more containers in a pod process requests related to one or more corresponding functions of a service. Proxy 245 typically controls network functions related to the service, such as routing and load balancing. Other services may also be equipped with pods similar to pods. During operation, executing a user request from end user 202 may require invoking one or more services in computing environment 201. Executing one or more functions of one service may require invoking one or more functions of another service. As shown in Figure 2, service "A" 220-1 receives a user request from end user 202 from ingress gateway 230. Service "A" 220-1 may invoke service "D" 220-2, and service "D" 220-2 may request service "E" 220-3 to execute one or more functions. The computing environment described above may be a cloud computing environment, where resource allocation is managed by the cloud service provider, allowing for feature development without having to worry about implementing, adjusting, or scaling servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Instead of expanding a single hardware device to handle potential load, services can be segmented to complete a set of functions that can be automatically and independently scaled. In the aforementioned operating environment, this application provides a data processing method for a processor, as shown in Figure 3, from the data processing device side. It should be noted that the data processing method for the processor in this embodiment can be executed by the mobile terminal in the embodiment shown in Figure 1. Figure 3 is a flow chart of a data processing method for a processor, according to an embodiment of this application. As shown in Figure 3, this method is applied to a data processing device, which includes a processor comprised of multiple processing elements arranged according to permutation rule information. The method may include the following steps: Step S302: Obtain a first data set and a second data set to be processed by the processor. In the technical solution provided in step S302 of this application, the processor may include multiple processing elements arranged according to the permutation rule information. The processor may be a data stream processor, capable of processing and analyzing data streams. It may receive real-time data streams from various sources, process, convert, filter, and analyze the data, and generate corresponding output results. In this embodiment of the application, the data processed by the processor may be matrices, and its function is to perform matrix multiplication to obtain corresponding output results. In this embodiment, the processing element may also be referred to as a PE unit in a processor. Alternatively, if the processing element is a PE unit used for matrix multiplication, it may be an arithmetic unit. A PE unit is a basic computing unit with independent computing capabilities in parallel computing and is a core component for implementing parallel computing.Permutation rule information can be used to arrange PE units so that data can flow within the PE unit array, and can also be referred to as array information. In this embodiment, a first data set and a second data set to be processed can be obtained. Optionally, the first data set and the second data set to be processed by the processor can be two matrices capable of matrix multiplication. That is, the first data set can be a first matrix, and the second data set can be a second matrix. Optionally, the rows and columns of the two matrices to be multiplied have a certain relationship, that is, the number of columns of the first matrix used for matrix multiplication is the same as the number of rows of the second matrix. The first matrix can be represented by X, XE RSXJ, where % can be used to represent the number of rows of the first matrix and % can be used to represent the number of columns of the first matrix. The second matrix can be represented by 4, AE RJxg, where % can be used to represent the number of rows of the second matrix and 3 can be used to represent the number of columns of the second matrix. It should be noted that the number of rows, columns, and formats of the first and second matrices described above are merely examples and are not specifically limited herein. Optionally, if the operation to be performed is matrix multiplication, the first and second data sets to be processed by the processor are the first and second matrices, respectively. The first and second matrices to be subjected to matrix multiplication can be obtained. Optionally, if matrix multiplication is required for the processor, the first and second matrices to be subjected to matrix multiplication can be obtained. Optionally, since the core calculation method in the large model can be matrix multiplication, the first and second matrices to be subjected to matrix multiplication in the large model can be obtained. It should be noted that the above-mentioned application scenarios for matrix multiplication are merely illustrative and are not specifically limited herein. Any application scenario that can utilize the matrix multiplication described in this application is within the scope of protection of the embodiments of this application. In step S304, the first and second data sets are respectively subjected to block processing to obtain multiple first initial sub-data sets and multiple second initial sub-data sets. In the technical solution provided in step S304 of this application, the first initial sub-data set can be a data set obtained by block processing of the first data set. If the first data set is a first matrix, the first initial sub-data set can be a block of the first matrix, or can also be referred to as a matrix block of the first matrix. The second initial sub-dataset may be a dataset obtained by performing block processing on the second dataset. If the second dataset is a second matrix, the second initial sub-dataset may be a block of the second matrix, which may also be referred to as a matrix block of the second matrix. This matrix block may also be referred to as a block matrix. Block processing may be the process of dividing the first dataset and the second dataset to obtain corresponding block matrices.In this embodiment, after obtaining the first and second data sets to be processed by the processor, the first data set can be processed in blocks to obtain multiple first initial sub-data sets corresponding to the block processing. The second data set can also be processed in blocks to obtain multiple second initial sub-data sets corresponding to the block processing. In this embodiment of the present application, since the first and second data sets to be processed by the processor may be large, direct computation without block processing would result in increased computational complexity and low efficiency. If the first and second data sets can be simplified, for example, by performing block processing on the two data sets to obtain multiple smaller initial sub-data sets, the above method can be used to divide a relatively difficult data set into multiple simpler and smaller initial sub-data sets, thereby simplifying the difficulty of processing the data sets by the processor and achieving the technical effect of improving the data processing efficiency of the processor. Optionally, the first matrix and the second matrix can be processed in blocks to obtain corresponding first initial block matrices and second initial block matrices. If the first data set is the first matrix, the first initial sub-data set can be the first initial block matrix. If the second data set is a second matrix, the second initial sub-data set can be a second initial block matrix. Optionally, because the data set is large, calculations cannot be performed on the entire data set. The first and second data sets can be divided accordingly according to the arrangement rule information of the processing elements arranged in the data stream processor to obtain first and second initial sub-data sets, thereby facilitating parallel operations on the larger data sets. For example, if the processing elements in the data stream processor are arranged in a 16*16 configuration and there are 256 PE units, matrix X and matrix 4 can be divided in a 16*16 configuration, with each block matrix having a size of 64*64. This configuration helps fully utilize the parallel computing capabilities of the PE units and decompose the large matrix multiplication task into multiple smaller blocks for parallel computation. In step S306, the first initial sub-data set is divided along the first arrangement direction, and the second initial sub-data set is divided along the second arrangement direction, using the arrangement rule information, to obtain multiple first target sub-data sets and multiple second target sub-data sets. In the technical solution provided in step S306 of the present application, the first arrangement direction may be used to divide the first initial sub-dataset after the block processing. If the first initial sub-dataset is a block of the first matrix, the first arrangement direction is the row direction. The second arrangement direction may be used to divide the second initial sub-dataset after the block processing. If the second initial sub-dataset is a block of the second matrix, the second arrangement direction is the column direction.The first target sub-dataset may be obtained by partitioning the first initial sub-dataset along a first arrangement direction. If the first initial sub-dataset is a first initial block matrix, the first target sub-dataset may be a first target block matrix. The second target sub-dataset may be obtained by partitioning the second initial sub-dataset along a second arrangement direction. If the second initial sub-dataset is a second initial block matrix, the second target sub-dataset may be a second target block matrix. The first target sub-dataset may be broadcasted to corresponding processing elements on the processor along the first arrangement direction. The second target sub-dataset may be broadcasted to corresponding processing elements on the processor along the second arrangement direction. Optionally, the above partitioning may also be referred to as slicing. In this embodiment, after the first dataset and the second dataset are partitioned to obtain a plurality of first initial sub-datasets and a plurality of second initial sub-datasets, the first initial sub-dataset may be partitioned along the first arrangement direction according to the arrangement rule information to obtain a plurality of first target sub-datasets. Alternatively, the second initial sub-dataset may be partitioned along the second arrangement direction according to the arrangement rule information to obtain a plurality of second target sub-datasets. The first target sub-dataset can be broadcast to the corresponding processing elements via the first arrangement direction, and the second target sub-dataset can also be broadcast to the corresponding processing elements via the second arrangement direction. Optionally, the first initial sub-dataset can be split in the first arrangement direction to obtain the split first target sub-dataset. The second initial sub-dataset can be split in the second arrangement direction to obtain the split second target sub-dataset. Optionally, the number of columns in the first matrix and the number of rows in the second matrix can be the same. In this case, the arrangement rule information can be used to partition the first initial block matrix in the row direction to obtain multiple first target block matrices. The second initial block matrix can also be partitioned in the column direction to obtain multiple second target block matrices. These two matrices can be broadcast in the row and column directions to the corresponding PE units. Because the processor cannot process the entire dataset in large quantities, the first and second initial sub-datasets can be further partitioned after block processing. The first initial sub-dataset can be partitioned by rows to obtain the first target sub-dataset, and the second initial sub-dataset can be partitioned by columns to obtain the second target sub-dataset. The first and second target sub-datasets can then be broadcasted in the row and column directions, respectively. In other words, the inner loop can be performed in both the row and column directions.The above method not only fully utilizes the computing and storage capabilities of the PE unit, but also meets certain storage restrictions, thereby achieving the purpose of improving computing efficiency and parallelism, and further achieving the technical effect of improving the processing efficiency of the data stream processor. In the technical solution provided in step S308 of the present application, the output results of the processing elements can be used to represent the processing results obtained by each processing element processing the target sub-dataset. In this embodiment, after using the permutation rule information to divide the first initial sub-dataset along the first permutation direction and the second initial sub-dataset along the second permutation direction to obtain corresponding multiple first target sub-datasets and second target sub-datasets, operations can be performed on the first target sub-datasets and the second target sub-datasets in the processing elements where the target sub-datasets are located to obtain the output results of each processing element. Optionally, a matrix multiplication operation can be performed on the first target block matrix and the second target block matrix to obtain a matrix multiplication result. Matrix multiplication operation can also be referred to as matrix multiplication. The output result of the processing element can be the matrix multiplication result obtained by performing matrix multiplication on the first target block matrix and the second target block matrix in the processing element. Step S310 determines the output result of the processor based on the output results of the processing element. In the technical solution provided in step S310 of the present application, the output result of the processor can be the final result obtained by performing operations on the first data set and the second data set. In this embodiment, after computing the first and second target sub-datasets in corresponding processing elements and obtaining the output results of each processing element, the output result of the entire processor can be determined based on the output results of each processing element. Optionally, after determining the output results of each processing element, the computation results of each block-partitioned and divided target sub-dataset can be obtained, i.e., the output results of each processing element. If the final computation result of computing the entire first and second data sets is desired, the output results of all processing elements are required to determine the processor output result. In this embodiment of the present application, considering that matrix multiplication can be processed in parallel by multiple processing elements in the processor, the matrix multiplication calculation process is accelerated. Furthermore, row and column broadcasting can be used to copy and broadcast data from a row or column of the entire matrix to other rows or columns, ensuring that the data in the rows or columns of the entire matrix are identical. This allows for data sharing within the rows and columns, thereby avoiding time-consuming operations such as writing intermediate results. This improves the processor's data processing efficiency and addresses the problem of low processor data processing efficiency.Through steps S302 to S310 of the present application, if a processor is required to process the first and second data sets, the first and second data sets to be multiplied can be obtained. The first and second data sets can be divided into blocks to obtain corresponding first and second initial sub-data sets. The first initial sub-data set can be divided along a first arrangement direction using the permutation rule information to obtain multiple first target sub-data sets. The second initial sub-data set can also be divided along a second arrangement direction to obtain multiple second target sub-data sets. The two target sub-data sets can then be broadcasted in the first or second arrangement direction to corresponding processing elements. Operations are then performed on the first and second target sub-data sets in the processing elements to obtain output results from each processing element. The output results from each processing element are then integrated to obtain a final output result. In the embodiment of the present application, considering that operations on two data sets can be processed in parallel by multiple processing elements in a processor, the purpose of accelerating operations is achieved. Furthermore, by broadcasting in the first and second arrangement directions, data in a row or column of the entire data set is copied and broadcasted to other rows or columns, ensuring that the data in the rows or columns of the entire data set is identical. This enables data sharing within the rows and columns, thereby avoiding the time-consuming operations of writing intermediate results. This further improves the processor's data processing efficiency and solves the technical problem of low processor data processing efficiency. The above-mentioned method of this embodiment is further described below. As an optional implementation, the arrangement rule information is used to indicate that the multiple processing elements are arranged according to a target number of rows and a target number of columns, with the first arrangement direction being the row direction and the second arrangement direction being the column direction. Step S306, using the arrangement rule information to divide the first initial sub-dataset in the first arrangement direction and the second initial sub-dataset in the second arrangement direction to obtain multiple first target sub-datasets and multiple second target sub-datasets, includes: dividing the first initial sub-dataset in the corresponding row direction using the target number of rows to obtain first target sub-datasets corresponding to the processing elements; and dividing the second initial sub-dataset in the corresponding column direction using the target number of columns to obtain second target sub-datasets corresponding to the processing elements.In this embodiment, when using the permutation rule information to partition the first initial sub-dataset in the first permutation direction and the second initial sub-dataset in the second permutation direction to obtain the first target sub-dataset and the second target sub-dataset, the first initial sub-dataset can be partitioned in the corresponding row direction using the target number of rows to obtain the first target sub-dataset corresponding to the processing element. The second initial sub-dataset can be partitioned in the corresponding column direction using the target number of columns to obtain the second target sub-dataset corresponding to the processing element. The permutation rule information can be used to indicate that the multiple processing elements are arranged according to the target number of rows and the target number of columns. The first permutation direction can be the row direction, and the second permutation direction can be the column direction. Optionally, since the permutation rule information indicates that the processing elements in the data stream processor are pre-arranged according to a certain number of rows and columns, that is, according to the target number of rows and the target number of columns, To ensure the processing efficiency of each processing element for a data set and the parallel computing capability between processing elements, the data sets required to be processed by each processing element in the data stream processor must also be arranged according to the aforementioned permutation rule information, so that each processing element can promptly and accurately process the data set received. Therefore, due to the particularity of matrix multiplication and the relationship between the rows and columns of the two matrices undergoing matrix multiplication, the rows and columns of these two matrices can be partitioned in the row and column directions to obtain corresponding target block matrices. Optionally, the first initial block matrix can be partitioned in the row direction according to the target number of rows in the permutation rule information to obtain multiple first target block matrices. Furthermore, the second initial block matrix can be partitioned in the column direction according to the target number of columns in the permutation rule information to obtain multiple second target block matrices. The first target block matrices can be broadcast in the row direction, and the second target block matrices can be broadcast in the column direction to the corresponding PE units. As an optional implementation, step S304, performing block processing on the first dataset and the second dataset to obtain a plurality of first initial sub-datasets and a plurality of second initial sub-datasets, includes: using the arrangement rule information to perform block processing on the first dataset and the second dataset to obtain the first initial sub-datasets and the second initial sub-datasets. In this embodiment, in the process of performing block processing on the first dataset and the second dataset to obtain the plurality of first initial sub-datasets and the plurality of second initial sub-datasets, the arrangement rule information may be used to perform block processing on the first dataset and the second dataset to obtain the first initial sub-datasets and the second initial sub-datasets.Optionally, the processing elements in the processor can be arranged in an array, and arrangement rule information for arranging the processing elements can be determined. During the initial block processing of the first and second data sets, the arrangement rule information can be used to block the first and second data sets, so that the arrangement of the first and second initial sub-data sets after block processing is the same as the arrangement of the current processing elements. For example, if the processing elements are pre-arranged so that the arrangement rule information of the processing elements is (Q, Q), that is, the processing elements can be arranged in the form of (Q, Q), to ensure that each processing element can participate in the calculation, each processing element can be assigned to a corresponding data set for processing. Therefore, the first and second data sets can be block-processed according to the arrangement rule information of the processing elements (Q, Q), so that each processing element can participate in the calculation, thereby reducing calculation time and achieving the technical effect of improving the data processing efficiency of the processor. It should be noted that the above-mentioned permutation rule information for permuting processing elements is merely illustrative and is not specifically limited herein. Any permutation rule information that enables each processing element in a processor to participate in the computation of a data set is within the scope of protection of the embodiments of this application. As an optional implementation, using the permutation rule information to partition the first and second data sets into blocks to obtain first and second initial sub-data sets includes: determining a first partition parameter of the processor using the permutation rule information; and partitioning the first and second data sets using the first partition parameter to obtain first and second initial sub-data sets. In this embodiment, when using the permutation rule information to perform block processing on the first and second data sets to obtain the first and second initial sub-data sets, the permutation rule information can be used to determine a first partitioning parameter of the processor. The first partitioning parameter can be used to partition the first and second data sets to obtain the first and second initial sub-data sets after the partitioning. The first partitioning parameter can be used to represent a parameter of an inner loop for processing the first and second data sets, and can also be referred to as the number of iterations of the inner loop or the loop variable of the inner loop. In this embodiment of the present application, the inner loop can be used to process the first and second data sets, as the inner loop can be used to process relatively complex data structures such as matrices, two-dimensional arrays, or multi-dimensional arrays.The inner loop can break down complex operations on the aforementioned data structure into multiple smaller steps, thereby simplifying the entire block processing operation for the first and second data sets and facilitating the maintenance of the code for the entire operation, thereby achieving the technical effect of improving the computational efficiency of the data stream processor. As an optional implementation, using a first partition parameter to partition the first and second data sets to obtain first and second initial sub-data sets arranged on the processor includes: partitioning the first data set in a second arrangement direction using the first partition parameter to obtain the first initial sub-data set; and partitioning the second data set in the first arrangement direction using the first partition parameter to obtain the second initial sub-data set. In this embodiment, when partitioning the first and second data sets in a second arrangement direction using the first partition parameter to obtain the first and second initial sub-data sets arranged on the processor, the first data set can be partitioned in the second arrangement direction using the first partition parameter to obtain the first initial sub-data set. The second data set can be partitioned in the first arrangement direction using the first partition parameter to obtain the second initial sub-data set. Optionally, an inner loop can be performed in the column direction of the first dataset and in the row direction of the second dataset. The number of loops can be represented by k, meaning that the first initial sub-dataset and the second initial sub-dataset processed by the inner loop can be subjected to inner loop calculations. Optionally, an inner loop can be performed in the column direction of the first matrix X and in the row direction of the second matrix X, with a number of loops k, thereby obtaining the first initial block matrix as e and the second initial block matrix as . eDuring the inner loop, each row of the first initial block matrix and each column of the second initial block matrix can be traversed, and the inner loop of matrix multiplication can be performed. In the inner loop, the first matrix and the second matrix are multiplied according to the number of loops, thereby implementing the matrix multiplication calculation process. The inner loop can ensure that every element in the matrix participates in the matrix multiplication calculation. It should be noted that the number of loops of the inner loop and the first initial block matrix and the second initial block matrix for matrix multiplication are merely examples and are not specifically limited here. Any process and method that can simplify the matrix multiplication of the first matrix and the second matrix using the inner loop is within the scope of protection of the embodiments of the present application. As an optional implementation, using a first partition parameter to partition the second data set in a first arrangement direction to obtain a second initial subset data set includes: if the data volume of the second data set is greater than a data volume threshold, partitioning the second data set using a second partition parameter of a processor, wherein the product of the second partition parameter and the first partition parameter is less than the product threshold; and partitioning the partitioned second data set in the first arrangement direction using the first partition parameter to obtain a second initial subset data set. In this embodiment, during the process of partitioning the second data set in the first arrangement direction using the first partition parameter to obtain the second initial subset data set, a relationship between the data volume of the second data set and the data volume threshold can be determined. If the data volume of the second data set is greater than the data volume threshold, the second data set can be partitioned using the second partition parameter of the processor, and the partitioned second data set can be partitioned in the first arrangement direction using the first partition parameter to obtain the second initial subset data set. Optionally, the second partition parameter can be the number of iterations of an outer loop (also referred to as a loop variable of the outer loop). The product of the second partition parameter and the first partition parameter is less than the product threshold. The product threshold can be a preset value that meets the requirements. The data volume can be used to represent the size of the second data set. If the second data set is a matrix, the data volume can be a parameter corresponding to the matrix size. The data volume threshold can be a preset value or a value set according to the actual situation of the processor. If the operation is a matrix, the data volume threshold can be a parameter used to determine whether the matrix size reaches a certain level of complexity. If the matrix size exceeds the data volume threshold, it indicates that the matrix size is large. It should be noted that the above-mentioned data volume threshold value and setting method are merely examples and are not specifically limited here. As long as the process and method can split the matrix when the matrix size reaches a certain level, it is within the scope of protection of the embodiments of this application.When the matrix size is large, directly performing full matrix calculations without further matrix processing can result in high computational complexity and reduced computational efficiency. However, in embodiments of the present application, considering the above issues, the matrix can be split again to reduce the size of the entire matrix, thereby reducing the complexity of matrix calculations and achieving the technical effect of improving the efficiency of matrix multiplication operations performed by the processor. Optionally, by determining that the data volume of the second matrix is greater than a data volume threshold, it can be indicated that the matrix size of the second matrix has increased, and the second matrix needs to be split again. The second matrix can be split in the column direction, that is, the second matrix can be split column-wise as an outer loop, and the number of loops can be eight. For example, if the matrix 4 to be calculated is large and the entire matrix cannot be calculated, the matrix 4 can be split again. Considering the characteristics of the data stream processor, the matrix 4 can be split column-wise. The column splitting of the matrix 4 as the outer loop has a number of loops of X € RSXj . A EBy performing column splitting on the matrix 4, the column-split matrix & can be obtained. It should be noted that the number of iterations of the outer loop and the column splitting process described above are merely illustrative and are not specifically limited herein. Any process and method capable of performing matrix splitting and outer looping when the matrix size is large is within the scope of the present invention. In the present invention, the outer loop can be used to control the number of executions or conditions of the inner loop, and multiple matrices can also be processed. Controlling the inner loop through the outer loop effectively manages complex operations. The outer loop processes multiple data sets, such as multiple arrays or matrices, enabling simultaneous processing of multiple data sets. The above analysis demonstrates that the outer loop plays an important role in data set operations, achieving the goal of increasing computational parallelism and, in turn, improving the efficiency of data stream processors in computing data sets. In related art, stream computing is used to perform matrix operations. However, because this computing method requires time for data flow, each computing unit cannot work simultaneously. For example, if Y = XA, the matrix X is divided into blocks (M, P), and the matrix M is divided into blocks (P, N). The total time required for data reading, calculation, and readout is (M + N + P - 2)7. Using stream computing can increase the time complexity of matrix operations. However, in the embodiments of the present application, considering the above issues, the PE units can be arranged in a (Q, Q) format. When all PE units in the data stream processor participate in the operation, the completion time is (2Q + P - 2)T. This method achieves the technical effect of reducing the time complexity of matrix multiplication operations performed by the data stream processor. Optionally, the first data set and the second data set can be divided into blocks. If the first data set and the second data set need to perform a matrix multiplication operation, assuming that the first data set is determined as the first matrix X and the second data set is determined as the second matrix, the first matrix and the second matrix can be divided into blocks and k*n blocks respectively. Then, the matrix multiplication result of the above two matrices is X4*. The calculation can be divided into the accumulation of k matrices. Since in the matrix multiplication, Alternatively, since the PE unit array can share data within rows and columns, as long as the input X block is column-wise broadcasted in the row direction and the 4 block is row-wise broadcasted in the column direction, the computation time can be reduced to (P + Q)T. It is worth noting that when multiplying multiple matrices, using methods in related art requires writing and reading intermediate results for each multiplication. However, the method in the embodiments of the present application avoids writing intermediate results, further saving time. In the embodiments of the present application, the above method enables data sharing and broadcasting within the PE unit, thereby optimizing the efficiency of matrix multiplication. Data can be shared within the PE unit, allowing the input X block to be column-wise broadcasted in the row direction and the 4 block to be row-wise broadcasted in the column direction. This operation helps accelerate matrix multiplication calculations, reduces data transmission and copying time, and improves computational efficiency. Through data sharing and broadcasting, the matrix multiplication computation time can be reduced to (P + Q)T. This indicates that the optimized computation time is significantly shortened, improving computational efficiency. Intermediate results do not need to be stored or read; they can be obtained by broadcasting them to the PE unit for the next calculation. This optimization strategy effectively reduces the storage and reading of intermediate results, accelerating the multi-layer matrix multiplication process. In the case of multi-layer matrix multiplication, the writing of intermediate results can be avoided, saving time. This optimization strategy helps reduce the time overhead of data transmission and read / write operations, thereby improving overall computational efficiency. In summary, the embodiments of the present application can effectively improve the efficiency of matrix multiplication by avoiding the writing of intermediate results and optimizing the storage and reading of intermediate results through data sharing and broadcasting operations. This optimization strategy helps reduce the time overhead of data transmission and read / write operations, achieving the technical effect of improving overall computational efficiency. For example, if the permutation rule information of the PE unit in the data stream processor is 16*16 and there are 256 PE units, assuming that the size of matrix X and matrix 4 are both 1024*1024, they can also be divided according to the permutation rule information of 16*16. In this case, the size of each block matrix is 64*64. In the related art, the computation time for completing the matrix multiplication of the first matrix and the second matrix is 46 T. Using the embodiment of the present application, the computation time is 327 T. This represents a 30% reduction in time compared to the related art, and the yield rate is positively correlated with the size of the PE unit array. In the embodiment of the present application, the stream processor is arranged in a 16*16 configuration, with a total of 256 PE units. The matrices X and A are partitioned in a 16*16 configuration, and each partition matrix is 64*64 in size.This configuration helps fully utilize the parallel computing capabilities of PE units, breaking down large matrix multiplication tasks into multiple smaller parallel computations. This reduces time consumption by 30% compared to related techniques. This means the new method reduces time overhead by 30% compared to related techniques, a significant improvement. This time savings is of great significance for large matrix multiplication calculations. The yield rate is positively correlated with array size. This indicates that as the array size increases, the advantages of the embodiments of the present application become more pronounced, enabling better utilization of large-scale PE unit arrays for efficient parallel computing. This further achieves the technical effect of improving the efficiency of matrix multiplication operations using data stream processors. As an optional implementation, the method further includes: obtaining the data volume of the first data set, the data volume of the second data set, and the storage space of the processing element; and determining the first partitioning parameter and the second partitioning parameter based on the permutation rule information, the data volume of the first data set, the data volume of the second data set, and the storage space. In this embodiment, the data volume of the first data set, the data volume of the second data set, and the storage space of the processing element can be obtained. Based on the arrangement rule information of the processing elements, the data volumes of the first and second data sets, and the storage space of the processing elements, a first partitioning parameter and a second partitioning parameter are determined. The data volumes of the first and second data sets represent the sizes of the first and second data sets. If the first and second data sets are matrices, the corresponding data volumes can be the matrix sizes of the corresponding matrices, i.e., the number of rows and columns. The storage space can represent the storage capacity of the corresponding processing element, i.e., the maximum storage capacity of the PE unit, also known as the storage limit. Optionally, if the first and second partitioning parameters need to be determined, the data volumes of the first and second data sets to be calculated, as well as the storage space of each processing element, can be obtained. The first and second partitioning parameters can be determined based on the aforementioned parameters. For example, assuming the array of PE units in a data stream processor is (q, q), the inputs of each PE unit can be determined to be q, ... The time for each outer loop can be determined as (k + q)T, and the total time is l(k + q)T. When k » q, the time is simplified to IkT. It should be noted that the storage of a single PE unit can satisfy I1 + I2 + O. j < M PE , where MPE can also The maximum storage capacity of the unit. The time required for a single PE unit to perform the two steps of data loading and calculation can be represented by T. In the embodiment of the present application, through the above operation, the computing power and storage capacity of the PE unit can be fully utilized. The technical effect of the rate. As an optional implementation, inputting the first target sub-dataset and the second target sub-dataset into a processing element for computation to obtain an output result of the corresponding processing element includes: determining a plurality of first identification information based on a first partitioning parameter, and determining a plurality of second identification information based on a second partitioning parameter, wherein the first identification information is a positive integer less than the first partitioning parameter, and the second identification information is a positive integer less than the second partitioning parameter; an input step of inputting the first target sub-dataset identified by current first identification information among the plurality of first identification information, and the second target sub-dataset identified by the first identification information and current second identification information among the plurality of second identification information, into a current processing element among the plurality of processing elements for computation, and merging the obtained computation result with an intermediate result of the first dataset and the second dataset on the current processing element to obtain an output result of the current processing element; and if the current first identification information has a next first identification information among the plurality of first identification information, and the current processing element has a next processing element among the plurality of processing elements, determining the next first identification information as the current first identification information, determining the output result of the current processing element as the intermediate result of the next processing element, determining the next processing element as the current processing element, and returning to the input step. In this embodiment, when inputting the first target sub-dataset and the second target sub-dataset into a processing element for computation and obtaining the output result of the corresponding processing element, multiple first identification information may be determined based on the first partitioning parameter, and multiple second identification information may be determined based on the second partitioning parameter. The first target sub-dataset identified by the current first identification information among the multiple first identification information, and the second target sub-dataset identified by the first identification information and the current second identification information among the multiple second identification information, may be input into the current processing element among the multiple processing elements for computation, and the resulting computation result may be combined with the intermediate result of the first dataset and the second dataset at the current processing element to obtain the output result of the current processing element. Optionally, it may be determined whether the current first identification information includes the next first identification information among the multiple first identification information, or whether the current processing element is the next processing element among the multiple processing elements. If so, the next first identification information may be determined as the current first identification information, the output result of the current processing element may be determined as the intermediate result of the next processing element, and the next processing element may be determined as the current processing element, and the above process may be returned to execution. Optionally, the first identification information may be a positive integer smaller than the first division parameter. The second identification information may be a positive integer smaller than the second division parameter.Optionally, the inputs of each PE unit are respectively nine =囹 X囹, ” 囹 X囹. Through the inputs of each of the above PE units, the corresponding intermediate result q = [^] X [剖 can be determined. Optionally, according to the first partitioning parameter k and the second partitioning parameter I, X and 4 can be partitioned to obtain
[0003] & = "2 X囹, = Tii X囹, A tj = [y X号]. Optionally, through the first partitioning parameter, the first identification information can be determined. That is, by determining the size of the first partitioning parameter k, the first identification information can be obtained as 1 < j < k. Optionally, through the second partitioning parameter, the second identification information can be determined. That is, by determining the size of the second partitioning parameter / , the second identification information can be obtained as 1 < i < L. Optionally, a single PE unit has two input terminals and two output terminals respectively. When the above nine and the upper are input into the register in the PE unit, the intermediate result can be obtained by performing an operation. The intermediate result or nine and & can be transmitted to the PE unit adjacent to the current PE unit. In the embodiment of the present application, with the structure that the PE unit has two input terminals and two output terminals, it can receive input data, perform operations, and pass the results to adjacent units. This structure reflects the characteristics of the data flow processor, that is, the flow of data between processing units, which helps to avoid redundant memory access operations and achieves the technical effect of improving computing efficiency. Optionally, in the inner loop, can be divided into q blocks in the row direction to obtain Xjq =捋, and can be broadcast in the row direction on the data flow processor. 4订 can be divided into q blocks in the column direction to obtain A ijq= preferably, and can be broadcast in the column direction on the data stream processor. Optionally, in the above process, the intermediate result can be updated, Oj = Oj + ** I; and the above intermediate result can be read, so that the final output result of the processor is Yj = Oj. Optionally, the above steps are repeatedly determined in a loop while i and j are within the corresponding value range, until i and j reach / and k, respectively, and then the loop can be terminated. As an optional implementation, multiple first identification information are sorted from small to large, and multiple second identification information are sorted from small to large. The method may also include: if the current first identification information is the largest first identification information among the multiple first identification information, and the current second identification information has the next second identification information in the multiple second identification information, then the next second identification information is determined as the current second identification information, the minimum first identification information among the multiple first identification information is determined as the current first identification information identifier, the next processing element is determined as the current processing element, and the output result of the current processing element is determined as the intermediate result on the current processing element, and the input step is returned to execute until the second identification information is the last second identification information among the multiple second identification information. In this embodiment, it is possible to determine whether the current first identification information is the largest first identification information among multiple first identification information, and whether the current second identification information has the next second identification information among the multiple second identification information. If so, the next second identification information can be determined as the current second identification information, and the smallest first identification information among the multiple first identification information can be determined as the current first identification information. The next processing element can also be determined as the current processing element, and the output result of the current processing element can be determined as the intermediate result of the current processing element. The input step is then returned to execution until the second identification information is the last second identification information among the multiple second identification information. The multiple first identification information are sorted from smallest to largest, and the multiple second identification information are sorted from smallest to largest. Optionally, in the outer loop, after traversing the first identification information under the current second identification information, that is, after the current first identification information is the largest identification information in the interval of the first identification information, the intermediate results under each first identification information under the next second identification information can be traversed. Optionally, after the current first identification information is the largest first identification information under the current second identification information, End the loop (end for). As an optional embodiment, step S310, determining the output result of the processor based on the output result of the processing element, includes: if the current processing element is the last processing element among the multiple processing elements, determining the output result of the current processing element as the output result of the processor. In this embodiment, during the process of determining the output result of the processor based on the output result of the processing element, it can be determined whether the processing element is the last processing element among the multiple processing elements. If so, the output result of the current processing element can be determined as the output result of the processor. Optionally, after a processing element performs an operation based on the two received data sets, the intermediate result of the operation can be transmitted to the adjacent processing element. The adjacent processing elements can perform corresponding operations based on the two received data sets and the intermediate result of the previous processing element. The operation continues until the last processing element among the processing elements is reached. The output result output by the last processing element is the output result obtained by the operation of the entire first and second data sets input into the processor. As an optional embodiment, when the first and second data sets are arranged in a matrix, the number of columns corresponding to the first data set is the same as the number of rows corresponding to the second data set. In this embodiment, if the first and second data sets are arranged in a matrix, the number of columns corresponding to the first data set is the same as the number of rows corresponding to the second data set. Alternatively, in this embodiment of the present application, if the two data sets to be operated on by the data stream processor are matrices and the operation to be performed is matrix multiplication, the first data set can be determined as the first matrix, and the second data set can be determined as the second matrix. The rows and columns of the first and second matrices to be multiplied have a certain correlation, that is, the number of columns of the first matrix is the same as the number of rows of the second matrix. This relationship between the rows and columns of the first and second matrices ensures that the first matrix can be multiplied with the second matrix. Alternatively, if matrix multiplication is required on the processor, the first and second matrices to be multiplied can be obtained and partitioned into blocks to obtain corresponding first and second initial block matrices. Since the number of columns of the first matrix and the number of rows of the second matrix are the same, the arrangement rule information can be used to partition the first initial block matrix in the row direction to obtain multiple first target block matrices. The second initial block matrix can be divided in the column direction to obtain multiple divided second target block matrices, and the above two matrices can be broadcast in the row direction or the column direction.Matrix multiplication can be performed on the first target block matrix and the second target block matrix to obtain a matrix multiplication result. The obtained matrix multiplication result can be merged with the intermediate result to obtain a final output result. As an optional implementation, step S302, obtaining the first and second data sets to be processed by the processor, includes: obtaining the first and second data sets to be processed by the processor from the large model. In this embodiment, during the process of obtaining the first and second data sets to be processed by the processor, the first and second data sets to be processed can be obtained from the large model. Optionally, because the core calculation method in the large model is matrix multiplication, if related technologies are used to perform operations on the two data sets in the large model, the various processing elements in the data stream processor cannot be effectively utilized. In other words, parallel processing of operations between the processing elements cannot be guaranteed. Therefore, the technical problem of low computational efficiency of the data stream processor still exists. However, in an embodiment of the present application, if a first data set and a second data set to be processed are present in a large model, the first data set and the second data set can be transferred to a data stream processor, effectively utilizing each processing element and performing operations in parallel across the processing elements, thereby achieving the technical effect of improving the computational efficiency of the data stream processor. An embodiment of the present application also provides a data processing method for a data stream processor from the data stream processor side. FIG4 is a flowchart of a data processing method for a data stream processor according to an embodiment of the present application. As shown in FIG4 , the method is applied to a data processing device including a data stream processor, which is composed of multiple processing elements arranged according to arrangement rule information. The method may include the following steps: Step S402: Obtaining the first data set and the second data set to be processed by the data stream processor from the large model. In the technical solution provided in step S402 above, if the first data set and the second data set to be processed are present in the large model, the first data set and the second data set in the large model can be transferred to the data stream processor for performing the corresponding operations. Optionally, since the core computation method in the large model is matrix multiplication, the first matrix and the second matrix required for matrix multiplication in the large model can be obtained. In step S404, the first dataset and the second dataset are respectively divided into blocks to obtain a plurality of first initial sub-datasets and a plurality of second initial sub-datasets. In the technical solution provided in step S404 of the present application, after obtaining the first and second datasets to be processed by the data stream processor from the large model, the first dataset can be divided into blocks to obtain a plurality of first initial sub-datasets corresponding to the block processing. Alternatively, the second dataset can be divided into blocks to obtain a plurality of second initial sub-datasets corresponding to the block processing.Optionally, the first matrix and the second matrix are respectively subjected to block processing to obtain corresponding first and second initial block matrices. If the first data set is the first matrix, the first initial sub-data set may be the first initial block matrix. If the second data set is the second matrix, the second initial sub-data set may be the second initial block matrix. Because the first and second data sets to be processed by the processor may be large, performing operations directly without block processing results in increased processor computational complexity and low efficiency. If the first and second data sets can be simplified, for example, by performing block processing on the two data sets to obtain multiple smaller initial sub-data sets, the above method can be used to divide a relatively difficult data set into multiple simpler and smaller initial sub-data sets, thereby simplifying the difficulty of processing the data sets by the processor and achieving the technical effect of improving the data processing efficiency of the processor. Optionally, due to the large size of the dataset, it may be impossible to perform calculations on the entire dataset. Therefore, the first dataset and the second dataset may be partitioned according to the permutation rule information of each processing element arranged in the data stream processor to obtain a first initial sub-dataset and a second initial sub-dataset, thereby facilitating parallel computing of the large dataset. In step S406, the first initial sub-dataset is partitioned along the first permutation direction and the second initial sub-dataset is partitioned along the second permutation direction using the permutation rule information to obtain a first target sub-dataset and multiple second target sub-datasets. The first target sub-dataset is broadcast to the corresponding processing elements on the processor along the first permutation direction, and the second target sub-dataset is broadcast to the corresponding processing elements on the processor along the second permutation direction. In the technical solution provided in step S406 of the present application, after the first dataset and the second dataset are partitioned to obtain the corresponding multiple first initial sub-datasets and multiple second initial sub-datasets, the first initial sub-dataset may be partitioned along the first permutation direction according to the permutation rule information to obtain multiple first target sub-datasets. Alternatively, the second initial sub-dataset may be divided along the second arrangement direction according to the arrangement rule information to obtain multiple second target sub-datasets. The first target sub-dataset may be broadcast to corresponding processing components along the first arrangement direction, and the second target sub-dataset may be broadcast to corresponding processing components along the second arrangement direction. Optionally, the first initial sub-dataset may be divided along the first arrangement direction to obtain divided first target sub-datasets. The second initial sub-dataset may be divided along the second arrangement direction to obtain divided second target sub-datasets.Optionally, since the number of columns in the first matrix and the number of rows in the second matrix are the same, the permutation rule information can be used to partition the first initial block matrix in the row direction to obtain multiple first target block matrices. The second initial block matrix can also be partitioned in the column direction to obtain multiple second target block matrices. These two matrices can be broadcasted in the row and column directions to corresponding PE units. In step S408, operations are performed on the first target sub-dataset and the second target sub-dataset in the corresponding processing element to obtain output results of the corresponding processing element. In the technical solution provided in step S408 of the present application, after the first initial sub-dataset is partitioned in the first permutation direction and the second initial sub-dataset is partitioned in the second permutation direction using the permutation rule information to obtain multiple first target sub-datasets and second target sub-datasets, operations can be performed on the first target sub-dataset and the second target sub-dataset in the processing element where the target sub-dataset resides to obtain output results of each processing element. Optionally, a matrix multiplication operation can be performed on the first target block matrix and the second target block matrix to obtain a matrix multiplication result. The output result of the processing element can be the matrix multiplication result obtained by performing the matrix multiplication on the first target block matrix and the second target block matrix in the processing element. In step S410, the output result of the processor is determined based on the output result of the processing element. In the technical solution provided in step S410 of this application, after performing operations on the first target sub-dataset and the second target sub-dataset in the corresponding processing element and obtaining the output results of each processing element, the output result of the entire processor can be determined based on the output results of the processing element. Optionally, after determining the output results of each processing element, the calculation results of each block-partitioned and divided target sub-dataset can be obtained, that is, the output results of each processing element. If the final calculation result of the entire first and second data sets is required, the output results of all processing elements are required to determine the output result of the processor.Through steps S402 to S410 of the present application, the first and second data sets to be processed by the data stream processor are obtained from the large model; the first and second data sets are respectively divided into blocks to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; the first initial sub-data sets are divided along a first arrangement direction and the second initial sub-data sets are divided along a second arrangement direction using the arrangement rule information to obtain a first target sub-data set and multiple second target sub-data sets; operations are performed on the first and second target sub-data sets in corresponding processing elements to obtain output results of the corresponding processing elements; and the output results of the processor are determined based on the output results of the processing elements. This achieves the technical effect of improving the data processing efficiency of the processor and solves the technical problem of low data processing efficiency of the processor. An embodiment of the present application also provides a data processing method for a processor. FIG5 is a flowchart of another data processing method for a processor according to an embodiment of the present application. As shown in FIG5 , the method is applied to a data processing device including a processor, which may be composed of multiple processing elements arranged according to arrangement rule information. The method may include the following steps: Step S502: Acquire a first data set and a second data set to be processed by the processor by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the first data set and the second data set. In the technical solution provided in step S502 of the present application, if a large model includes a first data set and a second data set to be processed by the processor, the first interface may be called to acquire the corresponding first data set and second data set, wherein the first interface may include a first parameter, and the parameter value of the first parameter may include the first data set and the second data set. Optionally, if the operation to be performed is matrix multiplication, the first data set and the second data set to be processed by the processor are the first matrix and the second matrix, respectively. The first matrix and the second matrix to be processed for the matrix multiplication may be acquired. Optionally, if matrix multiplication is required for the processor, the first matrix and the second matrix to be processed for the matrix multiplication may be acquired. Since the core calculation method in the large model is matrix multiplication, the first interface can be called to obtain the first matrix and the second matrix required for matrix multiplication in the large model. These matrices can then be transmitted to the processor for processing. In step S504, the first dataset and the second dataset are each processed in blocks to obtain multiple first initial sub-datasets and multiple second initial sub-datasets.In the technical solution provided in step S504 of the present application, after the processor obtains the first and second data sets to be processed by calling the first interface, the first data set can be partitioned to obtain multiple first initial sub-data sets corresponding to the partitioned data sets. The second data set can also be partitioned to obtain multiple second initial sub-data sets corresponding to the partitioned data sets. Optionally, the first matrix and the second matrix can be partitioned to obtain corresponding first initial partitioned matrices and second initial partitioned matrices. If the first data set is the first matrix, the first initial sub-data set can be the first initial partitioned matrix. If the second data set is the second matrix, the second initial sub-data set can be the second initial partitioned matrix. In this embodiment of the present application, by partitioning the first and second data sets, a relatively difficult data set can be partitioned into multiple simpler and smaller initial sub-data sets, thereby simplifying the difficulty of processing the data set by the processor and achieving the technical effect of improving the data processing efficiency of the processor. In step S506, the first initial sub-dataset is divided along the first arrangement direction, and the second initial sub-dataset is divided along the second arrangement direction, using the arrangement rule information, to obtain multiple first target sub-datasets and multiple second target sub-datasets. The first target sub-dataset is broadcast to corresponding processing elements on the processor along the first arrangement direction, and the second target sub-dataset is broadcast to corresponding processing elements on the processor along the second arrangement direction. In the technical solution provided in step S506 of the present application, the first initial sub-dataset can be divided along the first arrangement direction according to the arrangement rule information to obtain multiple first target sub-datasets. Alternatively, the second initial sub-dataset can be divided along the second arrangement direction according to the arrangement rule information to obtain multiple second target sub-datasets. The first target sub-datasets can be broadcast to corresponding processing elements along the first arrangement direction, and the second target sub-datasets can be broadcast to corresponding processing elements along the second arrangement direction. Optionally, the first initial sub-dataset can be segmented along the first arrangement direction to obtain segmented first target sub-datasets. The second initial sub-dataset can be segmented in the second arrangement direction to obtain a segmented second target sub-dataset. Optionally, since the number of columns in the first matrix and the number of rows in the second matrix are the same, the arrangement rule information can be used to segment the first initial block matrix in the row direction to obtain multiple first target block matrices. Furthermore, the second initial block matrix can be segmented in the column direction to obtain multiple second target block matrices. These two matrices can be broadcast in the row and column directions to the corresponding PE units.In an embodiment of the present application, because the data set is large and the entire data set cannot be computed within the processor, the first and second initial sub-datasets after block processing can be further partitioned. The first initial sub-dataset can be partitioned by rows to obtain a first target sub-dataset, and the second initial sub-dataset can be partitioned by columns to obtain a second target sub-dataset. The first target sub-dataset and the second target sub-dataset can be broadcasted in both the row and column directions. This method not only fully utilizes the computing and storage capabilities of the PE units but also meets certain storage constraints, thereby improving computing efficiency and parallelism, and further achieving the technical effect of improving the processing efficiency of the data stream processor. In step S508, the first and second target sub-datasets are computed in the corresponding processing elements to obtain the output results of the corresponding processing elements. In the technical solution provided in step S508 of the present application, the first and second target sub-datasets can be computed in the processing elements where the target sub-datasets are located to obtain the output results of each processing element. Optionally, a matrix multiplication operation can be performed on the first target block matrix and the second target block matrix to obtain a matrix multiplication result. The output result of the processing element can be the matrix multiplication result obtained by performing the matrix multiplication on the first target block matrix and the second target block matrix in the processing element. In step S510, the output result of the processor is determined based on the output result of the processing element. In the technical solution provided in step S510 of this application, after performing operations on the first target sub-dataset and the second target sub-dataset in the corresponding processing element to obtain the output results of each processing element, the output result of the entire processor can be determined based on the output results of the processing element. Optionally, after determining the output results of each processing element, the calculation results of each target sub-dataset after the blocks and divisions are calculated, that is, the output results of each processing element, can be obtained. If the final calculation result of the entire first data set and the second data set is required, the output results of all processing elements are required to determine the output result of the processor. In step S512, the output result of the processor is output by calling a second interface. The second interface includes a second parameter, and the parameter value of the second parameter includes the output result of the processor. In the technical solution provided in step S512 of the present application, after determining the output result of the processor, the second interface may be called to output the output result of the processor, wherein the second interface may include a second parameter. The parameter value of the second parameter may include the output result of the processor.Optionally, after the processor performs operations on the first and second datasets to obtain output results, the output results can be transmitted back to the large model via the second interface. Through steps S502 to S512 of the present application, the processor obtains the first and second datasets to be processed by calling the first interface; performs block processing on the first and second datasets, respectively, to obtain multiple first initial sub-datasets and multiple second initial sub-datasets; utilizes the permutation rule information to partition the first initial sub-dataset along a first permutation direction and the second initial sub-dataset along a second permutation direction, respectively, to obtain multiple first target sub-datasets and multiple second target sub-datasets; performs operations on the first target sub-datasets and the second target sub-datasets in corresponding processing elements to obtain output results of the corresponding processing elements; determines an output result of the processor based on the output results of the processing elements; and outputs the output result of the processor by calling the second interface. This achieves the technical effect of improving the processor's data processing efficiency and solves the technical problem of low processor data processing efficiency. Example 2 According to an embodiment of the present application, an embodiment of a data processing system is also provided. FIG6 is a schematic diagram of a data processing system according to an embodiment of the present application. As shown in FIG6 , data processing system 600 may include a data acquisition terminal 601 and a processor 602. The processor is composed of multiple processing elements arranged according to arrangement rule information. Data acquisition terminal 601 is configured to acquire a first data set and a second data set to be processed by the processor. In this embodiment, data acquisition terminal 601 may be used to acquire the first and second data sets to be processed in a large model. Optionally, if the large model contains the first and second data sets to be processed, the first and second data sets may be transmitted to data acquisition terminal 601. After receiving the first and second data sets, data acquisition terminal 601 may transmit the first and second data sets to processor 602, where processor 602 may perform operations on the first and second data sets. Optionally, if the first data set and the second data set that the large model needs to operate on are a first matrix and a second matrix, the first matrix and the second matrix to be subjected to the matrix multiplication operation can be obtained through the data acquisition terminal 601 and transmitted to the processor 602 to process the two matrices and perform the matrix multiplication operation.Processor 602 is configured to perform block processing on a first data set and a second data set, respectively, to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; use permutation rule information to divide the first initial sub-data set in a first permutation direction and the second initial sub-data set in a second permutation direction, respectively, to obtain multiple first target sub-data sets and multiple second target sub-data sets, wherein the first target sub-data set is broadcasted to corresponding processing elements on the processor in the first permutation direction, and the second target sub-data set is broadcasted to corresponding processing elements on the processor in the second permutation direction; input the first target sub-data set and the second target sub-data set into corresponding processing elements for computation to obtain output results of the corresponding processing elements; and determine an output result of the processor based on the output results of the processing elements. In this embodiment, processor 602 may perform block processing on the first data set to obtain multiple first initial sub-data sets corresponding to the block processing. Alternatively, processor 602 may perform block processing on the second data set to obtain multiple second initial sub-data sets corresponding to the block processing. The first initial sub-data set may be divided in the first permutation direction according to the permutation rule information to obtain multiple first target sub-data sets. Alternatively, the second initial sub-dataset may be divided along the second arrangement direction according to the arrangement rule information to obtain multiple second target sub-datasets. The first target sub-dataset may be broadcast to corresponding processing elements along the first arrangement direction, and the second target sub-dataset may be broadcast to corresponding processing elements along the second arrangement direction. Operations may be performed on the first target sub-dataset and the second target sub-dataset in the processing element where the target sub-dataset resides to obtain output results of each processing element. The output result of the entire processor may be determined based on the output results of the processing element. Optionally, block processing may be performed on the first matrix and the second matrix to obtain corresponding first initial block matrices and second initial block matrices. If the first dataset is the first matrix, the first initial sub-dataset may be the first initial block matrix. If the second dataset is the second matrix, the second initial sub-dataset may be the second initial block matrix. Alternatively, because the dataset is large, it may be impossible to perform calculations on the entire dataset. The first dataset and the second dataset may be partitioned accordingly according to the arrangement rule information of the processing elements arranged in the data stream processor to obtain a first initial sub-dataset and a second initial sub-dataset, thereby facilitating parallel computing of the large dataset. Alternatively, the first initial sub-dataset may be partitioned along the first arrangement direction to obtain a partitioned first target sub-dataset.The second initial sub-dataset can be segmented in the second arrangement direction to obtain a segmented second target sub-dataset. Optionally, since the number of columns in the first matrix and the number of rows in the second matrix are the same, the arrangement rule information can be used to segment the first initial block matrix in the row direction to obtain multiple first target block matrices. The second initial block matrix can also be segmented in the column direction to obtain multiple second target block matrices. These two matrices can be broadcasted in the row and column directions to corresponding PE units. Optionally, a matrix multiplication operation can be performed on the first target block matrix and the second target block matrix to obtain a matrix multiplication result. The output result of the processing element can be the matrix multiplication result obtained by performing the matrix multiplication of the first target block matrix and the second target block matrix in the processing element. Optionally, after determining the output result of each processing element, the calculation result of each target sub-dataset after segmentation and segmentation, i.e., the output result of each processing element, can be obtained. If the final result of the operation on the entire first and second data sets is desired, the output results of all processing elements are required to determine the processor's output result. In this embodiment, a data processing system is provided. A data acquisition terminal 601 acquires the first and second data sets to be processed by the processor. A processor 602 performs block processing on the first and second data sets, respectively, to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets. Using arrangement rule information, the first initial sub-data sets are divided along a first arrangement direction, and the second initial sub-data sets are divided along a second arrangement direction, to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets. The first target sub-data sets and the second target sub-data sets are input into corresponding processing elements for operation to obtain the output results of the corresponding processing elements. The processor's output result is determined based on the output results of the processing elements, thereby achieving the technical effect of improving the processor's data processing efficiency and resolving the technical problem of low processor data processing efficiency. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data, etc.) involved in this application, such as the data used for verification, are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse. Example 3 Currently, the core calculation of the large model LLM is matrix multiplication.In a matrix multiplication calculation, because data is stored in SRAM, the data needs to be loaded from SRAM into registers multiple times for calculation. This results in redundant memory access operations and reduces computational efficiency. This design avoids these redundant memory access operations, thereby improving computational efficiency. Furthermore, the low degree of parallelism among processing units slows the entire matrix multiplication process. Therefore, the technical issue of low processor data processing efficiency still exists. Currently, graphics processing units (GPUs), represented by Nvidia, are the mainstay of artificial intelligence (AI) computing hardware. However, when implementing matrix multiplication, the intermediate results of block-by-block matrix multiplication need to be written to SRAM and then read from SRAM again for the next use, resulting in a considerable amount of time waste. The data stream processor architecture, proposed in the 1980s, enables data flow within PE units, that is, data copying between PE units, thus avoiding this problem. When using a data stream processor, the streaming computing methods used in related technologies create data stream dependencies and fail to fully utilize the parallel capabilities of PE units. Therefore, the technical problem of low processor data processing efficiency still exists. Alternatively, when using a data stream processor for matrix multiplication, traditional streaming computing methods create data stream dependencies, meaning the input to the next PE is derived from a copy of the previous PE, thus failing to fully utilize the parallel capabilities of the PE units. Using the matrix multiplication solution proposed in the embodiments of the present application not only improves the parallelism of PE units but also provides additional benefits for cascaded matrix multiplication. Furthermore, the present application provides a matrix multiplication optimization method for data stream processors. If matrix multiplication is required for the processor, this method can obtain a first matrix and a second matrix to be multiplied, and can then perform block processing on each of the first and second matrices to obtain corresponding first and second initial block matrices. Since the first matrix has the same number of columns as the second matrix has the same number of rows, permutation rule information can be used to partition the first initial block matrix along the rows to obtain multiple first target block matrices. The second initial block matrix can be divided in the column direction to obtain multiple second target block matrices. The two matrices can be broadcast in the row direction or the column direction. Matrix multiplication can be performed on the first target block matrix and the second target block matrix to obtain a matrix multiplication result. The obtained matrix multiplication result can be combined with the intermediate result to obtain a final output result.In an embodiment of the present application, considering that matrix multiplication can be processed in parallel by multiple processing elements in a processor, the matrix multiplication calculation process is accelerated. Furthermore, data in a row or column of the entire matrix can be copied and broadcasted to other rows or columns through row and column broadcasting, ensuring that the data in the rows or columns of the entire matrix is identical. This allows for data sharing within rows and columns, thereby avoiding the time-consuming operation of writing intermediate results. This improves the processor's data processing efficiency and resolves the technical problem of low processor data processing efficiency. The above-mentioned method of this embodiment is further described below. In this embodiment, FIG7 is a schematic diagram of a computing unit structure according to an embodiment of the present application. As shown in FIG7 , a single PE unit has two input terminals and two output terminals. As shown in FIG7 , the two input terminals are input a1 and input a2, and the two output terminals are output b1 and output b2. Data can be transmitted to the computing unit via input a1, and data can be transmitted to the computing unit via input a2. Optionally, the data enters the register of the PE unit through two input terminals, where operations can be performed. At the same time, the input or intermediate results can be output and transmitted to adjacent units. Assuming that the maximum storage of the PE unit is , the sum of the input and intermediate results satisfies: / i + / 2 + 0 < M. O When a single PE performs a matrix multiplication of a certain size, it needs to perform two steps: data loading and operation. The time taken to complete the above process can be represented by T. Optionally, as shown in Figure 7, the PE unit has two input terminals and an output terminal. Data enters the register of the PE unit, the operation is performed, and the input or intermediate results are transmitted to the adjacent unit. The maximum storage of the PE unit is M, and The time required to complete the two steps of data loading and computation can be expressed as T. A PE unit has inputs and outputs that receive input data, perform computations, and pass results to adjacent units. This structure embodies the characteristics of a data flow processor: the flow of data between processing units, which helps avoid redundant memory accesses and improves computational efficiency. The maximum memory capacity of a PE unit is M, while the sum of inputs and intermediate results does not exceed M. This means that when performing matrix multiplication, data input and output must be effectively managed to ensure that memory capacity is within the limits. When performing matrix multiplication, a single PE performs the two steps of data loading and computation, and the time required to complete these steps can be expressed as T. This time complexity analysis helps evaluate data processing efficiency and performance. Figure 8(a) is a schematic diagram of matrix multiplication in a data stream processor in the related art. In stream computing in the related art, matrix multiplication can be performed using the data stream processor shown in Figure 8(a). The two diagonal grids in Figure 8(a) represent the matrix blocks after the two matrices undergoing matrix multiplication are divided, and the white blocks represent PE units. For example, if Y = XA, the matrix X is divided into blocks (M, P), and the matrix N is divided into blocks (P, N). Using stream computing, it takes (M + N + P - 2)T to complete data reading, calculation, and readout. When all PE units participate in the operation, the completion time is (2Q + P - 2)T. In the above stream computing, each PE unit cannot work simultaneously. Therefore, data flow requires time, resulting in the technical problem of low efficiency in completing matrix multiplication calculations using the above method. In an embodiment of the present application, FIG8(b) is a schematic diagram of a data stream processor matrix multiplication according to an embodiment of the present application. As shown in FIG8(b), the two diagonal grids represent the matrix blocks of the two matrices to be multiplied, and the white ones represent PE units. The PE units can be arranged in the form of (Q, Q). When all PE units in the data stream processor participate in the operation, the completion time is (2Q + P - 2)T. Through the above method, the technical effect of reducing the time complexity of the matrix multiplication operation performed by the data stream processor is achieved. Optionally, when performing the matrix multiplication operation, it is assumed that the matrix X and the matrix 4 are divided into blocks and k*n blocks. Then, the matrix multiplication result of the above two matrices X4 is *. The calculation can be split into the accumulation of k matrices. Since the matrix multiplication process involves data broadcasting, that is, the data of a row or column can be copied to In this embodiment, since the PE array can share data within rows and columns, as long as the input X block is column-broadcasted in the row direction and the 4 block is row-broadcasted in the column direction, the time can be reduced to (P + Q)T. It is worth noting that when multiple matrices are multiplied in succession, if the related art method is used, each multiplication calculation requires writing and reading the intermediate result. However, the method in the embodiment of the present application can avoid writing data of the intermediate result, further saving time. For example, the above method can implement a two-layer matrix multiplication Y = XAB, and Y = XAB can be split into 0 = X4 and Y = 0B. Figure 9 (a) is a schematic diagram of a matrix cascade according to an embodiment of the present application. As shown in Figure 9 (a), it shows the matrix cascade when the data stream processor processes 0 = X4. Figure 9 (b) is a schematic diagram of another matrix cascade according to an embodiment of the present application. As shown in Figure 9, it shows the matrix cascade when the data stream processor processes Y = 0B. Through the above operations, the intermediate result. No storage and read operations are required; the next computation can be obtained through broadcasting within the PE unit. Optionally, data can be shared within the PE array, so that the input block X is broadcasted column-wise in the row direction and the two blocks are broadcasted row-wise in the column direction. This operation helps accelerate matrix multiplication calculations, reduces data transmission and copying time, and improves computational efficiency. Optionally, through data sharing and broadcasting, the matrix multiplication calculation time can be reduced to (P + Q)T. This means that the optimized calculation process requires significantly shorter time, improving computational efficiency. Optionally, during the two-layer matrix multiplication Y = XAB, the intermediate results do not need to be stored or read; the next computation can be obtained through broadcasting within the PE array. This optimization strategy effectively reduces the storage and read operations of intermediate results, accelerating the multi-layer matrix multiplication process. Optionally, in the case of multi-layer matrix multiplication, by adopting the method in the embodiments of the present application, the data writing of intermediate results can be avoided, thereby saving time. This optimization strategy helps reduce the time overhead of data transmission and read and write operations, improving overall computational efficiency. In the embodiments of the present application, data sharing and broadcasting are used to avoid writing intermediate results and optimize the storage and reading of intermediate results, effectively improving the efficiency of matrix multiplication calculations. This optimization strategy helps reduce the time overhead of data transmission and read and write operations, thereby improving overall computing efficiency.For example, the stream processors are arranged in a 16*16 configuration, with a total of 256 PE units. The matrices X and A are divided into 16*16 sub-blocks, and each sub-block matrix has a size of 64*64. This setting helps to fully utilize the parallel computing capabilities of the PE units, breaking down the large matrix multiplication calculation task into multiple small parallel calculations. The time consumption is 30% less than that of related technologies. This means that the new method has a 30% reduction in time overhead compared to related technologies, which is a significant improvement. This saved time is of great significance for large matrix multiplication calculations. FIG. 10 is a schematic diagram of a matrix splitting process according to an embodiment of the present application. As shown in FIG. 10, when the scale of the matrix to be calculated is large, the complete matrix cannot be calculated directly, and the matrix can be split again. Considering the nature of the data stream processor, the matrices A and X can be split. The storage of a single PE needs to satisfy + / 2+ Oj < M PE , so in order to fully utilize the computing and storage capabilities of the PE array. Optionally, the data stream information can be (q, q). The first matrix is X f RJxs, and the second matrix is yl f RSx儿3. The first partitioning parameter can be determined by the following formula In this embodiment, array information of each computing unit included in the data stream processor, that is, the arrangement of each computing unit in the data stream processor, can be obtained. Step S1102: Determine whether the matrix needs to be partitioned. In this embodiment, after obtaining the matrix to be subjected to matrix multiplication, it can be determined whether the obtained matrix needs to be partitioned. If so, step S1104 can be executed; otherwise, step S1103 can be executed. Optionally, for some large matrices, performing matrix multiplication on the entire matrix may result in high complexity and may not be effectively applied to each computing unit. Therefore, for some large matrices, they can be further partitioned, and the matrix multiplication operation can be performed using the partitioned matrix. Optionally, for the obtained matrix, it can be determined whether the number of rows and columns of the matrix is greater than a preset threshold for the number of rows or columns. If so, the matrix size is large and it is a large matrix. The process jumps to step S1104 to partition the matrix. If both the number of rows and columns of the matrix are less than the preset threshold for the number of rows and columns, the matrix is small and the process jumps to step S1103 to perform a matrix multiplication operation. Step S1103 performs the matrix multiplication operation. In this embodiment, for some relatively simple matrices, that is, matrices with a small number of rows and columns, the matrix multiplication operation can be performed directly. Step S1104 calculates the size of the matrix partitions. In this embodiment, for some relatively complex matrices, that is, matrices with a large number of rows and / or columns, the matrix can be partitioned. The size of the matrix partitions can be determined based on the array information. Step S1105 loops through each matrix partition. In this embodiment, each matrix block can be cyclically executed through an inner loop and an outer loop. Step S1106: Obtain the execution result of the matrix multiplication operation. In this embodiment, the results determined by each PE unit are integrated to obtain the execution result of the matrix multiplication operation on two matrices by the entire data stream processor, that is, the operation result. In this embodiment of the present application, if matrix multiplication is required for the processor, a first matrix and a second matrix to be multiplied can be obtained. The first matrix and the second matrix can be separately partitioned to obtain corresponding first initial partitioned matrices and second initial partitioned matrices. Since the first matrix has the same number of columns as the second matrix has the same number of rows, the permutation rule information can be used to partition the first initial partitioned matrix in the row direction to obtain multiple first target partitioned matrices after partitioning.The second initial block matrix can be divided in the column direction to obtain multiple second target block matrices. These two matrices can be broadcast in the row or column direction. Matrix multiplication can be performed on the first target block matrix and the second target block matrix to obtain a matrix multiplication result. The obtained matrix multiplication result can be combined with the intermediate result to obtain a final output result. In this embodiment of the present application, considering that matrix multiplication can be processed in parallel by multiple processing elements in a processor, the matrix multiplication calculation process is accelerated. Furthermore, row and column broadcasting can be used to copy and broadcast data from a row or column of the entire matrix to other rows or columns, ensuring that the data in the rows or columns of the entire matrix is the same. This allows for data sharing within rows and columns, thereby avoiding the time-consuming operation of writing intermediate results. This improves the data processing efficiency of the processor and solves the technical problem of low data processing efficiency of the processor. Embodiment 4: According to this embodiment of the present application, a data processing device for a processor is provided for implementing the data processing method for the processor shown in FIG. 3 . FIG12 is a schematic diagram of a data processing device of a processor according to an embodiment of the present application. As shown in FIG12 , the data processing device 1200 of the processor may include: a first acquisition unit 1202, a first processing unit 1204, a first division unit 1206, a first operation unit 1208, and a first determination unit 1210. The first acquisition unit 1202 is configured to acquire a first data set and a second data set to be processed by the processor. The first processing unit 1204 is configured to perform block processing on the first data set and the second data set, respectively, to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets. The first division unit 1206 is configured to use permutation rule information to divide the first initial sub-data set along a first permutation direction and the second initial sub-data set along a second permutation direction, respectively, to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets. The first target sub-data sets are broadcast to corresponding processing elements on the processor along the first permutation direction, and the second target sub-data sets are broadcast to corresponding processing elements on the processor along the second permutation direction. The first operation unit 1208 is configured to operate on the first target sub-dataset and the second target sub-dataset in the corresponding processing element to obtain the output result of the corresponding processing element. The first determination unit 1210 is configured to determine the output result of the processor based on the output result of the processing element.The first acquisition unit 1202, first processing unit 1204, first division unit 1206, first operation unit 1208, and first determination unit 1210 described above correspond to steps S302 to S310 in Example 1. The examples and application scenarios implemented by these five units and the corresponding steps are the same, but are not limited to the content disclosed in Example 1. It should be noted that the above-mentioned units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned units may also be part of a device and run in the computer terminal 10 provided in Example 1. According to an embodiment of the present application, a data processing device for a data stream processor is also provided for implementing the data processing method of the data stream processor shown in FIG. 4 . FIG13 is a schematic diagram of a data processing device of a data stream processor according to an embodiment of the present application. As shown in FIG13 , the data processing device 1300 of the data stream processor may include: a second acquisition unit 1302, a second processing unit 1304, a second partitioning unit 1306, a second operation unit 1308, and a second determination unit 1310. The second acquisition unit 1302 is configured to acquire a first data set and a second data set to be processed by the data stream processor from a large model. The third processing unit 1304 is configured to perform block processing on the first data set and the second data set, respectively, to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets. The second partitioning unit 1306 is configured to use permutation rule information to partition the first initial sub-data set along a first permutation direction and the second initial sub-data set along a second permutation direction, respectively, to obtain a first target sub-data set and a plurality of second target sub-data sets. The first target sub-data set is broadcast to corresponding processing elements on the processor along the first permutation direction, and the second target sub-data set is broadcast to corresponding processing elements on the processor along the second permutation direction. The second operation unit 1308 is configured to operate on the first target sub-dataset and the second target sub-dataset in the corresponding processing element to obtain the output result of the corresponding processing element. The second determination unit 1310 is configured to determine the output result of the processor based on the output result of the processing element.It should be noted that the second acquisition unit 1302, second processing unit 1304, second division unit 1306, second operation unit 1308, and second determination unit 1310 described above correspond to steps S402 to S410 in Example 1. These five units and the corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Example 1. It should be noted that the above-mentioned units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). These units may also be part of an apparatus and run in the computer terminal 10 provided in Example 1. Figure 14 is a schematic diagram of a data processing apparatus of a processor according to an embodiment of the present application. As shown in Figure 14, the data processing apparatus 1400 of the processor may include a first calling unit 1402, a third processing unit 1404, a third division unit 1406, a third operation unit 1408, a third determination unit 1410, and a second calling unit 1412. The first calling unit 1402 is configured to obtain a first data set and a second data set to be processed by the processor by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the first data set and the second data set. The third processing unit 1404 is configured to perform block processing on the first data set and the second data set, respectively, to obtain multiple first initial sub-data sets and multiple second initial sub-data sets. The third partitioning unit 1406 is configured to use the permutation rule information to partition the first initial sub-data set along a first permutation direction and the second initial sub-data set along a second permutation direction, respectively, to obtain multiple first target sub-data sets and multiple second target sub-data sets. The first target sub-data sets are broadcast to corresponding processing elements on the processor along the first permutation direction, and the second target sub-data sets are broadcast to corresponding processing elements on the processor along the second permutation direction. The third computing unit 1408 is configured to perform computing on the first target sub-data set and the second target sub-data set in corresponding processing elements, to obtain output results of the corresponding processing elements. The third determining unit 1410 is configured to determine the output result of the processor based on the output results of the processing elements. The second calling unit 1412 is configured to output the output result of the processor by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the output result of the processor.It should be noted that the first calling unit 1402, the third processing unit 1404, the third dividing unit 1406, the third operation unit 1408, the third determining unit 1410, and the second calling unit 1412 correspond to steps S502 to S512 in Example 1. The six units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the above-mentioned units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, 102n). The above-mentioned units may also be part of an apparatus and run in the computer terminal 10 provided in Example 1. In the data processing apparatus of the processor, if the processor is required to process the first and second data sets, the first and second data sets to be multiplied may be obtained, and the first and second data sets may be divided into blocks for processing to obtain corresponding first and second initial sub-data sets. The permutation rule information can be used to partition the first initial sub-dataset in a first permutation direction to obtain multiple first target sub-datasets. The second initial sub-dataset can also be partitioned in a second permutation direction to obtain multiple second target sub-datasets. These two target sub-datasets can then be broadcast in the first or second permutation direction to corresponding processing elements. Operations are performed on the first and second target sub-datasets in the processing elements to obtain output results from each processing element. The output results from each processing element are then integrated to obtain a final output result. In this embodiment of the present application, considering that operations on two data sets can be processed in parallel by multiple processing elements in a processor, computational acceleration is achieved. Furthermore, by broadcasting in the first and second permutation directions, data in a row or column of the entire data set can be copied and broadcasted to other rows or columns, ensuring that the data in the rows or columns of the entire data set is identical. This allows for data sharing within rows and columns, thereby avoiding time-consuming operations such as writing intermediate results. This improves the processor's data processing efficiency and addresses the problem of low processor data processing efficiency. Embodiment 5 An embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may be replaced by a terminal device such as a mobile terminal.Optionally, in this embodiment, the computer terminal may be located in at least one of multiple network devices in a computer network. In this embodiment, the computer terminal may execute program code for the following steps in a data processing method of a processor: obtaining a first data set and a second data set to be processed by the processor; performing block processing on the first data set and the second data set, respectively, to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; utilizing permutation rule information, dividing the first initial sub-data set along a first permutation direction and dividing the second initial sub-data set along a second permutation direction, respectively, to obtain multiple first target sub-data sets and multiple second target sub-data sets, wherein the first target sub-data sets are broadcasted to corresponding processing elements on the processor along the first permutation direction, and the second target sub-data sets are broadcasted to corresponding processing elements on the processor along the second permutation direction; performing operations on the first target sub-data sets and the second target sub-data sets in the corresponding processing elements to obtain output results of the corresponding processing elements; and determining an output result of the processor based on the output results of the processing elements. Optionally, Figure 15 is a structural block diagram of a computer terminal according to an embodiment of the present application. As shown in Figure 15 , the computer terminal A may include one or more (only one shown) processors 1502, a memory 1504, and a transmission device 1506. The memory may be configured to store software programs and modules, such as the program instructions / modules corresponding to the data processing methods and apparatuses of the processors in the embodiments of the present application. The processor executes the software programs and modules stored in the memory to perform various functional applications and data processing, thereby implementing the aforementioned data processing methods of the processors. The memory may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory may further include memory located remotely from the processor, which may be connected to the terminal A via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. Optionally, the processor may further execute program code for the following steps: using the target number of rows, dividing the first initial sub-dataset in the corresponding row direction to obtain a first target sub-dataset corresponding to the processing element; and using the target number of columns, dividing the second initial sub-dataset in the corresponding column direction to obtain a second target sub-dataset corresponding to the processing element. Optionally, the processor may further execute program code for the following steps: using the permutation rule information, performing block processing on the first data set and the second data set, respectively, to obtain a first initial sub-dataset and a second initial sub-dataset.Optionally, the processor may further execute program code for the following steps: determining a first partition parameter of the processor using the arrangement rule information; partitioning the first data set and the second data set using the first partition parameter, respectively, to obtain a first initial sub-data set and a second initial sub-data set. Optionally, the processor may further execute program code for the following steps: partitioning the first data set in the second arrangement direction using the first partition parameter to obtain a first initial sub-data set; partitioning the second data set in the first arrangement direction using the first partition parameter to obtain a second initial sub-data set. Optionally, the processor may further execute program code for the following steps: if the data volume of the second data set is greater than a data volume threshold, partitioning the second data set using the processor's second partition parameter, wherein the product of the second partition parameter and the first partition parameter is less than the product threshold; and partitioning the partitioned second data set in the first arrangement direction using the first partition parameter to obtain a second initial sub-data set. Optionally, the processor may further execute program code for the following steps: obtaining the data volume of the first data set, the data volume of the second data set, and the storage space of the processing element; and determining the first partition parameter and the second partition parameter based on the arrangement rule information, the data volume of the first data set, the data volume of the second data set, and the storage space. Optionally, the processor may further execute program code of the following steps: determining a plurality of first identification information based on a first partitioning parameter, and determining a plurality of second identification information based on a second partitioning parameter, wherein the first identification information is a positive integer less than the first partitioning parameter, and the second identification information is a positive integer less than the second partitioning parameter; an input step of inputting a first target sub-dataset identified by a current first identification information among the plurality of first identification information, and a second target sub-dataset identified by the first identification information and the current second identification information among the plurality of second identification information, into a current processing element among the plurality of processing elements for operation, and merging the obtained operation result with an intermediate result of the first data set and the second data set on the current processing element to obtain an output result of the current processing element; if the current first identification information has a next first identification information among the plurality of first identification information, and the current processing element has a next processing element among the plurality of processing elements, determining the next first identification information as the current first identification information, determining the output result of the current processing element as the intermediate result of the next processing element, determining the next processing element as the current processing element, and returning to execute the input step.Optionally, the processor may further execute program code for the following steps: if the current first identification information is the largest first identification information among multiple first identification information, and the current second identification information has the next second identification information among the multiple second identification information, then determine the next second identification information as the current second identification information, determine the smallest first identification information among the multiple first identification information as the current first identification information identifier, determine the next processing element as the current processing element, and determine the output result of the current processing element as the intermediate result of the current processing element, and then return to the input step until the second identification information reaches the last second identification information among the multiple second identification information. Optionally, the processor may further execute program code for the following steps: if the current processing element is the last processing element among the multiple processing elements, then determine the output result of the current processing element as the output result of the processor. Optionally, the processor may further execute program code for the following steps: obtaining the first and second data sets to be processed by the processor from the large model. The processor can call information and applications stored in the memory through the transmission device to execute the following steps: obtain a first data set and a second data set to be processed by the data stream processor from the large model; block-process the first data set and the second data set respectively to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; use the arrangement rule information to divide the first initial sub-data set in a first arrangement direction and the second initial sub-data set in a second arrangement direction to obtain a first target sub-data set and multiple second target sub-data sets, wherein the first target sub-data set is broadcast to the corresponding processing element on the processor according to the first arrangement direction, and the second target sub-data set is broadcast to the corresponding processing element on the processor according to the second arrangement direction; perform operations on the first target sub-data set and the second target sub-data set in the corresponding processing element to obtain an output result of the corresponding processing element; and determine the output result of the processor based on the output result of the processing element.The processor can call information and applications stored in the memory through a transmission device to perform the following steps: obtaining a first data set and a second data set to be processed by the processor by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the first data set and the second data set; performing block processing on the first data set and the second data set, respectively, to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; using arrangement rule information, dividing the first initial sub-data set along a first arrangement direction and dividing the second initial sub-data set along a second arrangement direction, respectively, to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets, wherein the first target sub-data set is broadcasted to corresponding processing elements on the processor along the first arrangement direction, and the second target sub-data set is broadcasted to corresponding processing elements on the processor along the second arrangement direction; performing operations on the first target sub-data set and the second target sub-data set in the corresponding processing elements to obtain output results of the corresponding processing elements; determining an output result of the processor based on the output results of the processing elements; and outputting the output result of the processor by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the output result of the processor. Embodiments of the present application provide a data processing method for a processor. In an embodiment of the present application, if a processor is required to process the first and second data sets, the first and second data sets to be multiplied can be obtained. The first and second data sets can be divided into blocks to obtain corresponding first and second initial sub-data sets. The first initial sub-data set can be divided along a first arrangement direction using permutation rule information to obtain multiple first target sub-data sets. The second initial sub-data set can also be divided along a second arrangement direction to obtain multiple second target sub-data sets. These two target sub-data sets can then be broadcasted along the first or second arrangement direction to corresponding processing elements. Operations are then performed on the first and second target sub-data sets in the processing elements to obtain output results from each processing element. The output results from each processing element are then integrated to obtain a final output result.In the embodiments of the present application, considering that operations on two data sets can be processed in parallel by multiple processing elements in a processor, the purpose of accelerating operations is achieved. Furthermore, data in a row or column of the entire data set can be copied and broadcasted to other rows or columns by broadcasting in the first and second arrangement directions, making the data in the rows or columns of the entire data set identical. This allows for data sharing within rows and columns, thereby avoiding the time-consuming operations such as writing intermediate results. This further improves the processor's data processing efficiency and solves the technical problem of low processor data processing efficiency. Those skilled in the art will appreciate that the structure shown in FIG15 is merely illustrative, and computer terminal A may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG15 does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG15 , or have a configuration different from that shown in FIG15 . Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware associated with the terminal device through a program. The program can be stored in a computer-readable storage medium, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Example 6: The embodiments of the present application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the processor data processing method provided in Example 1. Optionally, in this embodiment, the computer-readable storage medium can be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining a first data set and a second data set to be processed by a processor; performing block processing on the first data set and the second data set respectively to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; using arrangement rule information, dividing the first initial sub-data set in a first arrangement direction and dividing the second initial sub-data set in a second arrangement direction to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets, wherein the first target sub-data set is broadcasted to corresponding processing elements on the processor according to the first arrangement direction, and the second target sub-data set is broadcasted to corresponding processing elements on the processor according to the second arrangement direction; performing operations on the first target sub-data set and the second target sub-data set in the corresponding processing elements to obtain output results of the corresponding processing elements; and determining an output result of the processor based on the output results of the processing elements. Optionally, the computer-readable storage medium may further execute program code for executing the following steps: using the target number of rows to partition the first initial sub-dataset in the corresponding row direction to obtain a first target sub-dataset corresponding to the processing element; and using the target number of columns to partition the second initial sub-dataset in the corresponding column direction to obtain a second target sub-dataset corresponding to the processing element. Optionally, the computer-readable storage medium may further execute program code for executing the following steps: using the permutation rule information to perform block processing on the first dataset and the second dataset, respectively, to obtain a first initial sub-dataset and a second initial sub-dataset. Optionally, the computer-readable storage medium may further execute program code for executing the following steps: using the permutation rule information to determine a first partition parameter for a processor; and using the first partition parameter to partition the first dataset and the second dataset, respectively, to obtain a first initial sub-dataset and a second initial sub-dataset. Optionally, the computer-readable storage medium may further execute program code for executing the following steps: using a first partition parameter, partitioning the first data set in the second arrangement direction to obtain a first initial subset data set; and using the first partition parameter, partitioning the second data set in the first arrangement direction to obtain a second initial subset data set. Optionally, the computer-readable storage medium may further execute program code for executing the following steps: if the data volume of the second data set is greater than a data volume threshold, partitioning the second data set using a second partition parameter of a processor, wherein the product of the second partition parameter and the first partition parameter is less than the product threshold; and partitioning the partitioned second data set in the first arrangement direction using the first partition parameter to obtain a second initial subset data set.Optionally, the computer-readable storage medium may further execute program code for the following steps: obtaining the data volume of the first data set, the data volume of the second data set, and the storage space of the processing element; and determining the first partition parameter and the second partition parameter based on the arrangement rule information, the data volume of the first data set, the data volume of the second data set, and the storage space. Optionally, the computer-readable storage medium may further execute program code for the following steps: determining a plurality of first identification information based on a first partitioning parameter, and determining a plurality of second identification information based on a second partitioning parameter, wherein the first identification information is a positive integer less than the first partitioning parameter, and the second identification information is a positive integer less than the second partitioning parameter; an input step of inputting a first target sub-dataset identified by current first identification information among the plurality of first identification information, and a second target sub-dataset identified by the first identification information and current second identification information among the plurality of second identification information, into a current processing element among the plurality of processing elements for operation, and merging the obtained operation result with an intermediate result of the first data set and the second data set on the current processing element to obtain an output result of the current processing element; and if the current first identification information has a next first identification information among the plurality of first identification information, and the current processing element has a next processing element among the plurality of processing elements, determining the next first identification information as the current first identification information, determining the output result of the current processing element as the intermediate result of the next processing element, determining the next processing element as the current processing element, and returning to the input step. Optionally, the computer-readable storage medium may further execute program code for the following steps: if the current first identification information is the largest first identification information among multiple first identification information, and the current second identification information has the next second identification information among the multiple second identification information, then determine the next second identification information as the current second identification information, determine the smallest first identification information among the multiple first identification information as the current first identification information identifier, determine the next processing element as the current processing element, and determine the output result of the current processing element as the intermediate result of the current processing element, and then return to the input step until the second identification information reaches the last second identification information among the multiple second identification information. Optionally, the computer-readable storage medium may further execute program code for the following steps: if the current processing element is the last processing element among the multiple processing elements, then determine the output result of the current processing element as the output result of the processor. Optionally, the computer-readable storage medium may further execute program code for the following steps: obtaining the first and second data sets to be processed by the processor from the large model.As an optional example, a computer-readable storage medium is configured to store program code for executing the following steps: obtaining a first data set and a second data set to be processed by a data stream processor from a large model; performing block processing on the first data set and the second data set respectively to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; using arrangement rule information, dividing the first initial sub-data set in a first arrangement direction and dividing the second initial sub-data set in a second arrangement direction to obtain a first target sub-data set and multiple second target sub-data sets, wherein the first target sub-data set is broadcast to corresponding processing elements on the processor according to the first arrangement direction, and the second target sub-data set is broadcast to corresponding processing elements on the processor according to the second arrangement direction; performing operations on the first target sub-data set and the second target sub-data set in the corresponding processing elements to obtain output results of the corresponding processing elements; and determining an output result of the processor based on the output results of the processing elements. As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: obtaining a first data set and a second data set to be processed by a processor by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the first data set and the second data set; performing block processing on the first data set and the second data set, respectively, to obtain multiple first initial sub-data sets and multiple second initial sub-data sets; using arrangement rule information, dividing the first initial sub-data set in a first arrangement direction and dividing the second initial sub-data set in a second arrangement direction, respectively, to obtain multiple first target sub-data sets and multiple second target sub-data sets, wherein the first target sub-data set is broadcasted to corresponding processing elements on the processor according to the first arrangement direction, and the second target sub-data set is broadcasted to corresponding processing elements on the processor according to the second arrangement direction; performing operations on the first target sub-data set and the second target sub-data set in the corresponding processing elements to obtain output results of the corresponding processing elements; determining an output result of the processor based on the output results of the processing elements; and outputting the output result of the processor by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the output result of the processor. Example 7 An embodiment of the present application may provide an electronic device, which may include a memory and a processor. FIG16 is a block diagram of an electronic device that implements a data processing method for a processor according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.An electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application as described and / or claimed herein. As shown in FIG16 , device 1600 includes a computing unit 1601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 1602 or loaded from a storage unit 1608 into a random access memory (RAM) 1603. RAM 1603 may also store various programs and data required for the operation of device 1600. Computing unit 1601, ROM 1602, and RAM 1603 are interconnected via a bus 1604. An input / output (I / O) interface 1605 is also connected to bus 1604. Multiple components in device 1600 are connected to I / O interface 1605, including: an input unit 1606, such as a keyboard and mouse; an output unit 1604, such as various types of displays and speakers; a storage unit 1608, such as a magnetic disk and optical disk; and a communication unit 1609, such as a network card, a modem, or a wireless communication transceiver. Communication unit 1609 allows device 1600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks. Computing unit 1601 can be any general-purpose or specialized processing component with processing and computing capabilities. Some examples of computing unit 1601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. Computing unit 1601 performs the various methods and processes described above, such as a processor's data processing method. For example, in some embodiments, the processor's data processing method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1600 via ROM 1602 and / or communication unit 1609.When the computer program is loaded into RAM 1603 and executed by computing unit 1601, one or more steps of the processor data processing method described above may be performed. Alternatively, in other embodiments, computing unit 1601 may be configured to execute the processor data processing method in any other appropriate manner (e.g., via firmware). Example 8: The present application also provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the processor data processing method described in the embodiments of the present application. Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be either a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, at least one input device, and at least one output device. Program code for implementing the methods of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on the remote machine or server.Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chips (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, at least one input device, and at least one output device. Program code for implementing the methods of the present application can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that, when executed by the processor or controller, the program code implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may execute entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server. In the context of this application, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.To provide user interaction, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD)) configured to display information to the user; a monitor; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be configured to provide user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input). The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., a data server), a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser). A user can interact with implementations of the systems and techniques described herein through the graphical user interface or web browser), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected through any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet. A computer system can include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship. The server can be a cloud server, a server in a distributed system, or a server integrated with a blockchain. It should be noted that the serial numbers of the above-mentioned embodiments of this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above-mentioned embodiments of this application, the description of each embodiment has its own focus. For portions not detailed in one embodiment, reference should be made to the relevant descriptions of other embodiments. In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways.The device embodiments described above are merely illustrative. For example, the division of units represents only one logical functional division. In actual implementation, other divisions may be employed. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be through interfaces, or indirect couplings or communication connections between units or modules, and may be electrical or other. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of these units may be selected to achieve the objectives of the present embodiments based on actual needs. Furthermore, the functional units in the various embodiments of the present application may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. These integrated units may be implemented in either hardware or software functional units. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for causing a computer device (such as a personal computer, server, or network device) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory, random access memory, a mobile hard drive, a magnetic disk, or an optical disk. The above are merely preferred embodiments of the present application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present application, and such improvements and modifications should also be considered within the scope of protection of the present application.
Claims
Claims 1. A data processing method for a processor, applied to a data processing device, wherein the data processing device includes a processor, the processor being composed of a plurality of processing elements arranged according to arrangement rule information, comprising: Obtaining a first data set and a second data set to be processed by the processor; The first data set and the second data set are respectively divided into blocks to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; the first initial sub-data sets are respectively divided in a first arrangement direction and the second initial sub-data sets are respectively divided in a second arrangement direction using the arrangement rule information to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets, wherein the first target sub-data sets are broadcasted on the processor to the corresponding processing elements in the first arrangement direction, and the second target sub-data sets are broadcasted on the processor to the corresponding processing elements in the second arrangement direction; in the corresponding processing elements, operations are performed on the first target sub-data sets and the second target sub-data sets to obtain output results of the corresponding processing elements; and an output result of the processor is determined based on the output results of the processing elements.
2. The method according to claim 1, wherein: The arrangement rule information is used to indicate that the multiple processing elements are arranged according to a target number of rows and a target number of columns, the first arrangement direction being a row direction, and the second arrangement direction being a column direction. Using the arrangement rule information, the first initial sub-dataset is divided in the first arrangement direction, and the second initial sub-dataset is divided in the second arrangement direction, to obtain a plurality of first target sub-datasets and a plurality of second target sub-datasets, including: using the target number of rows to divide the first initial sub-dataset in the corresponding row direction to obtain the first target sub-datasets corresponding to the processing elements; and using the target number of columns to divide the second initial sub-dataset in the corresponding column direction to obtain the second target sub-datasets corresponding to the processing elements.
3. The method according to claim 1, wherein: The first data set and the second data set are respectively divided into blocks to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets, including: using the arrangement rule information, the first data set and the second data set are respectively divided into blocks to obtain the first initial sub-data sets and the second initial sub-data sets.
4. The method according to claim 3, wherein: Using the arrangement rule information, respectively performing block processing on the first data set and the second data set to obtain the first initial sub-data set and the second initial sub-data set, includes: using the arrangement rule information, determining a first partitioning parameter of the processor; and using the first partitioning parameter, respectively partitioning the first data set and the second data set to obtain the first initial sub-data set and the second initial sub-data set. 45 5. The method according to claim 4, wherein: Using the first partition parameter to partition the first data set and the second data set respectively to obtain the first initial sub-data set and the second initial sub-data set arranged on the processor includes: using the first partition parameter to partition the first data set in the second arrangement direction to obtain the first initial sub-data set; and using the first partition parameter to partition the second data set in the first arrangement direction to obtain the second initial sub-data set.
6. The method according to claim 5, wherein: Using the first partition parameter, partitioning the second data set in the first arrangement direction to obtain the second initial sub-data set, including: if the data volume of the second data set is greater than a data volume threshold, partitioning the second data set using a second partition parameter of the processor, wherein a product of the second partition parameter and the first partition parameter is less than the product threshold; and partitioning the divided second data set in the first arrangement direction using the first partition parameter to obtain the second initial sub-data set.
7. The method according to claim 6, wherein: The method also includes: obtaining the data volume of the first data set, the data volume of the second data set, and the storage space of the processing element; and determining the first partition parameter and the second partition parameter based on the arrangement rule information, the data volume of the first data set, the data volume of the second data set, and the storage space.
8. The method according to claim 6, wherein: Inputting the first target sub-dataset and the second target sub-dataset into the processing element for operation to obtain an output result of the corresponding processing element includes: determining a plurality of first identification information based on the first partition parameter, and determining a plurality of second identification information based on the second partition parameter, wherein the first identification information is a positive integer less than the first partition parameter, and the second identification information is a positive integer less than the second partition parameter; an input step of inputting the first target sub-dataset identified by the current first identification information among the plurality of first identification information, and the second target sub-dataset identified by the first identification information and the current second identification information among the plurality of second identification information, into a current processing element among the plurality of processing elements for operation, and merging the obtained operation result with an intermediate result of the first dataset and the second dataset on the current processing element to obtain an output result of the current processing element; if the current first identification information has a next first identification information among the plurality of first identification information, and the current processing element has a next processing element among the plurality of processing elements, determining the next first identification information as the current first identification information, and determining the output result of the current processing element as the intermediate result of the next processing element. determining the next processing element as the current processing element, 46 Return to execute the input step.
9. The method according to claim 8, wherein The multiple first identification information are sorted from small to large, and the multiple second identification information are sorted from small to large. The method further includes: if the current first identification information is the largest first identification information among the multiple first identification information, and the current second identification information has the next second identification information among the multiple second identification information, then determining the next second identification information as the current second identification information, determining the smallest first identification information among the multiple first identification information as the current first identification information identifier, determining the next processing element as the current processing element, and determining the output result of the current processing element as the intermediate result on the current processing element, and returning to execute the input step until the second identification information is the last second identification information among the multiple second identification information.
10. The method according to claim 8, wherein: Determining the output result of the processor based on the output result of the processing element includes: if the current processing element is the last processing element among the multiple processing elements, determining the output result of the current processing element as the output result of the processor.
11. The method according to any one of claims 1 to 10, wherein: In the case where the first data set and the second data set are arranged in a matrix form, the number of columns corresponding to the first data set is the same as the number of rows corresponding to the second data set.
12. The method according to any one of claims 1 to 10, wherein: Acquiring a first data set and a second data set to be processed by the processor, comprising: acquiring the first data set and the second data set to be processed by the processor from a large model.
13. A data processing method for a data stream processor, applied to a data processing device, wherein the data processing device includes a data stream processor, the data stream processor being composed of a plurality of processing elements arranged according to arrangement rule information, comprising: Obtaining a first data set and a second data set to be processed by the data stream processor from the large model; The first data set and the second data set are respectively divided into blocks to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; the first initial sub-data set is respectively divided in a first arrangement direction and the second initial sub-data set is respectively divided in a second arrangement direction using the arrangement rule information to obtain a first target sub-data set and a plurality of second target sub-data sets, wherein the first target sub-data set is broadcasted on the processor to the corresponding processing elements in the first arrangement direction, and the second target sub-data set is broadcasted on the processor to the corresponding processing elements in the second arrangement direction; in the corresponding processing elements, operations are performed on the first target sub-data set and the second target sub-data set to obtain output results of the corresponding processing elements; and an output result of the processor is determined based on the output results of the processing elements. 47 14. A data processing method for a processor, applied to a data processing device, wherein the data processing device includes a processor, the processor being composed of a plurality of processing elements arranged according to arrangement rule information, comprising: The processor obtains a first data set and a second data set to be processed by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the first data set and the second data set; performs block processing on the first data set and the second data set respectively to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; uses the arrangement rule information to divide the first initial sub-data set in a first arrangement direction and the second initial sub-data set in a second arrangement direction to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets, wherein the first target sub-data set is broadcasted to the corresponding processing element on the processor in the first arrangement direction, and the second target sub-data set is broadcasted to the corresponding processing element on the processor in the second arrangement direction; performs operations on the first target sub-data set and the second target sub-data set in the corresponding processing element to obtain an output result of the corresponding processing element; determines an output result of the processor based on the output result of the processing element; and outputs the output result of the processor by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the output result of the processor.
15. A data processing system comprising a data acquisition terminal and a processor, wherein the processor comprises a plurality of processing elements arranged according to arrangement rule information, wherein: The data acquisition end is configured to acquire a first data set and a second data set to be processed by the processor; the processor is configured to perform block processing on the first data set and the second data set respectively to obtain a plurality of first initial sub-data sets and a plurality of second initial sub-data sets; using the arrangement rule information, the first initial sub-data set is divided in a first arrangement direction, and the second initial sub-data set is divided in a second arrangement direction to obtain a plurality of first target sub-data sets and a plurality of second target sub-data sets, wherein the first target sub-data sets are broadcasted on the processor to the corresponding processing elements in the first arrangement direction, and the second target sub-data sets are broadcasted on the processor to the corresponding processing elements in the second arrangement direction; the first target sub-data sets and the second target sub-data sets are input into the corresponding processing elements for calculation to obtain output results of the corresponding processing elements; and an output result of the processor is determined based on the output results of the processing elements.
16. An electronic device, comprising: a memory storing an executable program; A processor is configured to run the program, wherein the program executes claims 1 to The method according to any one of 14.
17. A computer-readable storage medium, comprising a stored executable program, wherein when the executable program is run, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 14.
18. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Tensor cache and access structure and method thereof
CN112925727A
Convolution operation circuit and operation method thereof
CN113869498A
Forward calculation method and system of multi-head attention mechanism based on super computer
CN115952393A
Calculation device, calculation method and related product
CN117235424A