Information processing method and device, equipment and storage medium
By receiving processing instructions associated with the logical data unit, determining the data slice and device serial number based on the identifier, and executing processing instructions at the corresponding computing device, the problems of low computing resource utilization and data synchronization consistency in traditional information processing are solved, and efficient distributed computing and a stable system are realized.
Patent Information
- Application Number
- CN202510400605.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
Traditional information processing methods cannot fully utilize the parallel processing capabilities of distributed computing resources, resulting in low utilization of computing resources, and there are data synchronization and consistency problems in distributed computing environments, affecting the accuracy of computing results and system stability.
By receiving processing instructions associated with the logical data unit, determining the data slice and device sequence number based on the identifier, and executing processing instructions at the corresponding computing device, managing the synchronization of communication and computing using data block primitives, supporting a dynamic mapping mechanism to adapt to complex models and dynamic data processing scenarios.
It improves the utilization rate of computing resources, reduces development costs, ensures the accuracy of calculation results and system stability, and adapts to the dynamic data processing needs of complex models.
Smart Images

Figure CN120336002A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for information processing. Background Art
[0002] In traditional information processing methods, data is usually processed as a whole, without fully utilizing the parallel processing capabilities of distributed computing resources. As the scale and complexity of models increase, the demand for computing resources also correspondingly increases. In this context, how to efficiently utilize multiple devices to process data in parallel has become an important research direction. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for information processing is provided. The method includes: during the operation of a model, receiving a processing instruction associated with a logical data unit, the processing instruction indicating an identifier of the logical data unit; based on the identifier, determining a data slice and a device number corresponding to the logical data unit; and at a computing device corresponding to the device number, executing the processing instruction corresponding to the data slice.
[0004] In a second aspect of the present disclosure, an apparatus for information processing is provided. The apparatus includes: a receiving module configured to receive, during the operation of a model, a processing instruction associated with a logical data unit, the processing instruction indicating an identifier of the logical data unit; a determining module configured to determine, based on the identifier, a data slice and a device number corresponding to the logical data unit; and an executing module configured to execute, at a computing device corresponding to the device number, the processing instruction corresponding to the data slice.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to execute the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium and can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this section of the specification is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Brief Description of the Drawings
[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, like or similar reference numerals denote like or similar elements, where:
[0009] Figure 1 FIG. shows a schematic diagram of an exemplary environment in which embodiments according to the present disclosure can be implemented;
[0010] Figure 2 FIG. shows a schematic diagram of different design spaces according to some embodiments of the present disclosure;
[0011] Figure 3 FIG. shows a flowchart of an exemplary process of information processing according to some embodiments of the present disclosure;
[0012] Figure 4 FIG. shows a schematic structural block diagram of an exemplary apparatus for information processing according to some embodiments of the present disclosure; and
[0013] Figure 5 FIG. shows a block diagram of an electronic device capable of implementing multiple embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0015] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Additionally, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different sections / subsections in any manner.
[0016] In the description of the embodiments of the present disclosure, the term "comprising" and its like shall be understood as an open inclusion, i.e., "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The term "some embodiments" shall be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0017] Embodiments of the present disclosure may involve user data, data acquisition, and / or data usage, etc. All of these aspects comply with corresponding laws, regulations, and related provisions. In the embodiments of the present disclosure, the collection, acquisition, processing, processing, forwarding, usage, etc. of all data are carried out on the premise that the user is aware of and confirms. Accordingly, when implementing the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained through appropriate means in accordance with relevant laws and regulations. The specific notification and / or authorization methods may vary according to the actual situation and application scenarios, and the scope of the present disclosure is not limited in this regard.
[0018] In the solutions of this specification and embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for basic functions, it will not affect the user's use of basic functions.
[0019] In a distributed computing environment, traditional data processing methods usually cannot fully utilize the parallel processing capabilities of multiple devices, resulting in low utilization of computing resources. In addition, as the model scale increases and the complexity improves, the demand for computing resources also increases accordingly, which further exacerbates the pressure on computing resources. In terms of communication, traditional data processing methods have data synchronization and consistency problems in a distributed computing environment. Data needs to be transmitted between different devices, and data competition and memory inconsistency problems during the communication process may affect the correctness of the calculation results and the stability of the system.
[0020] Embodiments of the present disclosure propose a solution for information processing. The solution includes: during the operation of the model, receiving a processing instruction associated with a logical data unit, where the processing instruction indicates an identifier of the logical data unit; based on the identifier, determining a data slice and a device number corresponding to the logical data unit; and at a computing device corresponding to the device number, executing the processing instruction corresponding to the data slice.
[0021] Thus, the embodiments of the present disclosure support processing instructions based on logical data units, which can hide low-level details, such as pointer management and barrier control, enabling developers to focus more on algorithm logic, thereby reducing development costs.
[0022] The following further describes various example implementations of this solution in detail with reference to the accompanying drawings.
[0023] Figure 1 FIG. 100 is a schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented. As Figure 1As shown, the model 110 can be deployed in a distributed system 120, which can include multiple computing nodes 130 for running the model 110 in parallel.
[0024] In some embodiments, the model 110 can include, for example, an appropriate machine learning model, such as a deep neural network. As an example, the model 110 can include an MLP (Multilayer Perceptron).
[0025] In some embodiments, the node 130 can run the kernel file of the model 110. In some embodiments, the kernel file can implement the fusion of computing and communication without independently building a computing kernel and a communication kernel.
[0026] The process of designing a fusion kernel based on logical data units (data blocks, such as tiles) will be introduced below. As an example, a logical data unit can be a data block obtained by slicing the target data to be processed. In some embodiments, the communication part and the computing part can be decoupled in multiple design spaces. As an example, such design spaces can include the size of the data block, the order of the data blocks, and the resource mapping logic. Figure 2 Schematic diagrams corresponding to different design spaces are shown.
[0027] As Figure 2 shown, the communication component and the computing component can independently formulate logic in three design spaces to optimize the component performance. For example, in the data block size subspace, the communication component and the computing component can select different data block sizes. For example, as Figure 2 shown, the communication part transmits a 128×128 data block each time, while the computing part consumes a 128×256 data block each time. This difference in data block size helps each component achieve optimal performance by aligning with the number of processing cores it uses.
[0028] For example, assume there is a task of data collection + matrix multiplication with a tensor size of M×N×K, where the data collection part binds the dimensions M and K to the processing cores, and the matrix multiplication part binds the dimensions M and N to the processing cores. If the communication component uses more cores, a smaller data block size can be used because all core resources can be fully utilized; conversely, if it uses fewer cores, a larger data block size can be used.
[0029] In addition, in the data block order subspace, the communication component can adopt a different data block order from the computing component. For example, communication can adopt various data transfer orders, such as ring order, full mesh all-to-all order, or other modes, while the computing component can start processing data blocks from any device number (e.g., rank). There are trade-offs in the selection of the data block order. If the computing component waits for data blocks from multiple device numbers, it may achieve better cache efficiency when processing larger blocks of data; however, this approach may result in longer waiting times. Conversely, if the computing component only waits for data blocks from one device number, it can start computing earlier, but this may lead to a reduction in overall computing efficiency. In Figure 2 the example shown, the communication component uses a ring order, and the computing component waits for data from two device numbers in each computing iteration.
[0030] In the resource binding subspace, the communication component and the computing component can be mapped to different units or the same unit, as Figure 2 shown. If the communication component uses a replication engine, it can avoid resource conflicts with the computing component. On the other hand, if the communication component uses a computing core for data replication, it may lead to resource conflicts with the computing component, but eliminates the computing overhead of the CPU.
[0031] The design space of decoupling communication and computing brings synchronization challenges. Since these two components use different data block sizes, data block orders, and resource mappings, synchronizing them requires complex low-level programming and communication instructions. Embodiments of the present disclosure further provide a set of block-based primitives.
[0032] In some embodiments, the primitives can include signal primitives and data primitives. The signal primitives can be used to transmit signals, and the data primitives can be used to transfer data. Table 1 further lists the primitives provided by the present disclosure. The first 6 in Table 1 are signal primitives, and the last 3 are data primitives.
[0033] Table 1 Primitives and Descriptions
[0034]
[0035]
[0036] As can be seen from Table 1, the following primitives can all implement signal processing or data processing based on tile_id: producer_tile_notify(tile_id,mode), consumer_tile_wait(tile_id), peer_tile_notify(tile_id,rank), peer_tile_wait(tile_id,rank), rank_notify(tile_id,rank), tile_push_data(tensors,tile_id,data), tile_pull_data(tensors,tile_id).
[0037] Furthermore, the primitive producer_tile_notify(tile_id,mode) can be used to trigger the generation of a first instruction, which can be the first instruction for notifying a first object (e.g., the corresponding consumer) that the logical data unit has completed the calculation. consumer_tile_wait(tile_id) can be used to trigger the generation of a second instruction, which is used to wait for a second object (e.g., the producer) to complete the calculation of the logical data unit. peer_tile_notify(tile_id,rank) can be used to trigger the generation of a third instruction, which is used to notify the peer device serial number that the logical data unit has completed the calculation;
[0038] peer_tile_wait(tile_id,rank) can be used to trigger the generation of a fourth instruction, which is used to wait for the peer device serial number to complete the calculation of the logical data unit. rank_notify(tile_id,rank) can be used to trigger the generation of a fifth instruction, which is used to notify the instruction device serial number that the logical data unit has completed the calculation.
[0039] In addition, the primitive tile_push_data(tensors,tile_id,data) can be used to trigger the generation of a sixth instruction for sending a data slice corresponding to the logical data unit to a third object (e.g., a remote tensor); tile_pull_data(tensors,tile_id) can be used to trigger the generation of a seventh instruction for obtaining a data slice corresponding to the logical data unit from a fourth object (e.g., a remote tensor).
[0040] Thus, the embodiments of the present disclosure can support using the above primitives to construct a kernel file for hybrid computing code and communication code.
[0041] The following shows a GEMM + ring ReduceScatter kernel constructed based on the above primitives.
[0042]
[0043] In the above example kernel file, the communication component and the computing component manage synchronization through the primitives producer_tile_notify and consumer_tile_wait. In addition, the kernel file also uses the primitives peer_tile_notify and peer_tile_wait to manage communication between peer device numbers.
[0044] It can be seen that by using primitives, embodiments of the present disclosure can significantly reduce the development workload. Developers only need to call the primitives to implement complex distributed computing tasks without writing a large amount of low-level code.
[0045] The following will further introduce the execution process of the hybrid kernel designed based on primitives.
[0046] Figure 3 A flowchart of an example process 300 of information processing according to some embodiments of the present disclosure is shown. Process 300 may be implemented at the distributed system 120.
[0047] As shown, at block 310, during the operation of the model, the distributed system 120 receives processing instructions associated with a logical data unit, and the processing instructions indicate an identifier of the logical data unit.
[0048] As an example, the distributed system 120 may obtain a kernel file of the model, and the kernel file may include communication code written based on the primitives described above. In addition, the kernel file may also include corresponding computing code.
[0049] Further, the kernel file may be compiled, and the code corresponding to the primitives may trigger the generation of corresponding processing instructions during the operation of the model. As mentioned above, such processing instructions may indicate an identifier (e.g., tile_id) of a logical data unit (e.g., a data block).
[0050] The following will further introduce the mapping process associated with the logical data unit.
[0051] As Figure 3 shown, at block 320, the distributed system 120 determines a data slice and a device number corresponding to the logical data unit based on the identifier.
[0052] In some embodiments, the mapping based on logical data units (e.g., data blocks) may include: shape mapping, device serial number mapping, and / or channel mapping. Specifically, shape mapping associates each data block identifier with a specific tensor shape slice. Device serial number mapping associates each data block identifier with a device serial number. Channel mapping assigns each data block identifier to a communication channel to execute the processing instruction based on the communication channel. As an example, f S , f R , f C respectively represent these three mappings.
[0053] In some embodiments, it may be supported to implement the mapping from the data block identifier to the above information based on static mapping or dynamic mapping. For static mapping, the distributed system 120 may map the identifier to the data slice and the device serial number based on the static mapping relationship, and the static mapping relationship is determined during the compilation of the model.
[0054] Specifically, static mapping refers to the mapping that can be determined at compile time. Static mapping is usually used in cases where the data sharding strategy is fixed, such as tensor parallel MLP and sequence parallel self-attention. As an example, affine operations can be used to process static mapping.
[0055] For example, for AllGather+GEMM (problem size is M×N×K) performed on R device serial numbers, each device serial number has C communication channels, the data block size used by the producer AllGather is Tmp×Tnp, and the input tensor is sharded along the M dimension. Given the tile_id p of the producer data block, its mapping process can be expressed as:
[0056]
[0057] range M =[tile_id p *Tm p , tile_id p *Tm p +Tm p ),
[0058]
[0059] In some embodiments, for dynamic mapping, the distributed system 120 may obtain the dynamic lookup table associated with the identifier, where the dynamic lookup table is constructed during the running of the model; and map the identifier to the data slice and the device serial number based on the dynamic lookup table.
[0060] Specifically, dynamic mapping refers to a mapping calculated at runtime, which is essential for workloads with dynamic data sharding requirements. For example, in the MoE data sharding strategy, dynamic routing determines the data distribution, and each data block may require tokens from any other rank. It is impossible to determine at compile time from which device number to collect data or on which channel to wait for a barrier. Therefore, the mapping must be calculated at runtime. To support dynamic mapping, the distributed system 120 can convert these mappings into lookup tables, whose values can be filled at runtime, while the access operations to these lookup tables are determined at compile time.
[0061] For example, the distributed system 120 can define the structure of the lookup tables at compile time, but not fill in the specific values. These lookup tables will be filled dynamically by the dynamic logic at runtime. Further, at runtime, the distributed system 120 can calculate the shape range, device number, and communication channel of each data block according to the dynamic logic (such as dynamic routing), and fill in the lookup tables. During subsequent execution, the distributed system 120 can access the lookup tables according to the data block identifier to obtain the shape range, device number, and communication channel of each data block.
[0062] Through dynamic mapping, the embodiments of the present disclosure can dynamically adjust data sharding and communication paths according to actual requirements at runtime, and are applicable to complex models and dynamic data processing scenarios.
[0063] In this way, the embodiments of the present disclosure can be mapped to the corresponding shape range, device number, and / or communication channel based on the identifier of the data block.
[0064] At block 330, the distributed system 120 executes the processing instructions corresponding to the data slice at the computing device corresponding to the device number.
[0065] Based on the mapping process described above, the distributed system 120 can execute the processing instructions corresponding to the data slice at the corresponding computing resources. For example, the computing resources to execute the processing instructions can be determined based on the device number, and the data slice to be processed can be determined based on the shape range. In addition, the processing instructions can also be executed based on the determined communication channel.
[0066] When compiling the kernel file, the distributed system 120 can also consider the memory consistency issue. In some embodiments, some primitives can be associated with preset consistency semantic information, and this semantic information can be used to control the data range corresponding to the access operations associated with the memory.
[0067] As an example, the primitives involving notify can carry release semantics, which can be used to ensure that all memory access operations before this primitive are not executed after this primitive. Additionally, the primitives involving wait carry acquire semantics, which can ensure that all memory access operations after this primitive are not executed before this primitive.
[0068] Furthermore, during the compilation process, the distributed system 120 can enforce strict data dependencies between signal primitives and their subsequent load / store operations. This ensures that signal primitives can be correctly reordered and expanded to accommodate pipeline optimizations. For multi-stage pipeline optimizations, the distributed system 120 can also ensure that memory access operations are not accidentally reordered in different stages of the pipeline.
[0069] As described above, embodiments of the present disclosure provide a set of block-based primitives that hide low-level details (such as pointer management and barrier control), enabling developers to focus more on algorithm logic rather than hardware details. For example, developers can write less code to achieve comparable performance.
[0070] Moreover, embodiments of the present disclosure decouple the design spaces of communication and computing, allowing both to be optimized independently, including different block sizes, block orders, and resource mappings. This flexibility enables developers to select the optimal strategy according to specific requirements.
[0071] On the other hand, based on the mapping mechanism described above, embodiments of the present disclosure are capable of effectively overlapping communication and computing, reducing communication overhead. Additionally, signal primitives provide strict memory consistency semantics, ensuring the correctness of memory operations in a multi-threaded / multi-process environment. This enables the correct transmission and processing of data to be guaranteed both at compile time and at runtime.
[0072] Furthermore, embodiments of the present disclosure support a dynamic mapping mechanism to allow the mapping strategy to be adjusted at runtime according to the actual data distribution, thereby better adapting to complex workloads.
[0073] Example devices and equipment
[0074] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an example apparatus 400 for information processing according to certain embodiments of the present disclosure is shown. The apparatus 400 can be implemented as or included in the distributed system 120. Each module / component in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.
[0075] As Figure 4As shown, the apparatus 400 includes a receiving module 410 configured to receive, during the operation of the model, processing instructions associated with a logical data unit, where the processing instructions indicate an identifier of the logical data unit; a determining module 420 configured to determine, based on the identifier, a data slice and a device serial number corresponding to the logical data unit; and an execution module 430 configured to execute, at a computing device corresponding to the device serial number, the processing instructions corresponding to the data slice.
[0076] In some embodiments, the determining module is further configured to: map the identifier to the data slice and the device serial number based on a static mapping relationship determined during the compilation of the model.
[0077] In some embodiments, the determining module is further configured to: obtain a dynamic lookup table associated with the identifier, where the dynamic lookup table is constructed during the operation of the model; and map the identifier to the data slice and the device serial number based on the dynamic lookup table.
[0078] In some embodiments, the apparatus 400 further includes a channel determining module configured to determine, based on the identifier, a communication channel corresponding to the logical data unit to execute the processing instructions based on the communication channel.
[0079] In some embodiments, the processing instructions include: a signal instruction for transmitting a signal; or a data instruction for transmitting data.
[0080] In some embodiments, the signal instruction includes one of the following: a first instruction for notifying a first object that the logical data unit has completed a calculation; a second instruction for waiting for a second object to complete the calculation of the logical data unit; a third instruction for notifying a peer device serial number that the logical data unit has completed a calculation; a fourth instruction for waiting for a peer device serial number to complete the calculation of the logical data unit; a fifth instruction for notifying an instruction device serial number that the logical data unit has completed a calculation.
[0081] In some embodiments, the data instruction includes one of the following: a sixth instruction for sending a data slice corresponding to the logical data unit to a third object; a seventh instruction for obtaining a data slice corresponding to the logical data unit from a fourth object.
[0082] In some embodiments, the processing instructions are associated with preset semantic information for controlling a data range corresponding to an access operation associated with a memory.
[0083] In some embodiments, the apparatus 400 further includes an instruction obtaining unit configured to: obtain a kernel file of the model, where the kernel file includes calculation code and communication code corresponding to the processing instructions; and obtain the processing instructions based on the kernel file.
[0084] In some embodiments, the logical data unit includes data blocks obtained by slicing the data to be processed.
[0085] Figure 5 FIG. shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 5 the illustrated electronic device 500 is merely exemplary and should not constitute any limitation on the functions and scope of the embodiments described herein. Figure 5 The illustrated electronic device 500 can be used to implement Figure 1 the distributed system 120.
[0086] As Figure 5 shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.
[0087] The electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.
[0088] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules that are configured to execute the various methods or actions of the various embodiments of the present disclosure.
[0089] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented by a single computing cluster or multiple computer machines that are capable of communicating via a communication link. Thus, the electronic device 500 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0090] The input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 560 can be one or more output devices, such as a display, speaker, printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed via the communication unit 540, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0091] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0092] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0093] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0094] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0095] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, and the module, segment of a program, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0096] The implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled persons in the art to understand the implementations disclosed herein.
Claims
1. A method for information processing, comprising: During the operation of a model, receiving a processing instruction associated with a logical data unit, the processing instruction indicating an identifier of the logical data unit; Based on the identifier, determining a data slice and a device serial number corresponding to the logical data unit; And At a computing device corresponding to the device serial number, executing the processing instruction corresponding to the data slice.
2. The method according to claim 1, wherein determining the data slice and the device serial number corresponding to the logical data unit based on the identifier comprises: Based on a static mapping relationship, mapping the identifier to the data slice and the device serial number, the static mapping relationship being determined during the compilation of the model.
3. The method according to claim 1, wherein determining the data slice and the device serial number corresponding to the logical data unit based on the identifier comprises: Obtaining a dynamic lookup table associated with the identifier, the dynamic lookup table being constructed during the operation of the model; And Based on the dynamic lookup table, mapping the identifier to the data slice and the device serial number.
4. The method according to claim 1, further comprising: Based on the identifier, determining a communication channel corresponding to the logical data unit to execute the processing instruction based on the communication channel.
5. The method according to claim 1, wherein the processing instruction comprises: A signal instruction for transmitting a signal; Or A data instruction for transmitting data.
6. The method according to claim 5, wherein the signal instruction comprises one of the following: A first instruction for notifying a first object that the logical data unit has completed a calculation; A second instruction for waiting for a second object to complete the calculation of the logical data unit; A third instruction for notifying a peer device serial number that the logical data unit has completed a calculation; A fourth instruction for waiting for a peer device serial number to complete the calculation of the logical data unit; A fifth instruction for notifying an instruction device serial number that the logical data unit has completed a calculation.
7. The method according to claim 5, wherein the data instruction comprises one of the following: A sixth instruction for sending the data slice corresponding to the logical data unit to a third object; A seventh instruction for obtaining the data slice corresponding to the logical data unit from a fourth object.
8. The method according to claim 1, wherein the processing instruction is associated with preset semantic information, the semantic information being used to control a data range corresponding to an access operation associated with a memory.
9. The method according to claim 1, further comprising: Obtaining a kernel file of the model, the kernel file including calculation code and communication code corresponding to the processing instruction; And Based on the kernel file, obtaining the processing instruction.
10. The method according to claim 1, wherein the logical data unit comprises a data block obtained by slicing data to be processed.
11. An apparatus for information processing, comprising: A receiving module, configured to receive, during the operation of the model, a processing instruction associated with a logical data unit, the processing instruction indicating an identifier of the logical data unit; A determining module, configured to determine, based on the identifier, a data slice and a device serial number corresponding to the logical data unit; And An execution module, configured to execute, at a computing device corresponding to the device serial number, the processing instruction corresponding to the data slice.
12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 10.