Data sharing method and device, electronic equipment and storage medium

By establishing a mapping relationship between storage paths and identification information in deep learning model inference, a second identification information representing the target data is generated, which solves the problem of redundant copying of data between different processors and achieves efficient data sharing and processing.

CN120723846BActive Publication Date: 2025-11-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511166091.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-25
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

During deep learning model inference, multiple copies of data between different processors lead to increased bandwidth and memory consumption, reduced system throughput, and low parallel efficiency.

Method used

By establishing a one-to-one mapping relationship between the storage path and the first identification information, a second identification information is generated to represent multiple target data. The second mapping relationship is then used to enable the second processor to automatically read the target data, simplifying the data sharing process and reducing redundant information transmission.

Benefits of technology

It improves data sharing efficiency, reduces redundant data transmission and delivery, and increases data processing speed and system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723846B_ABST
    Figure CN120723846B_ABST
Patent Text Reader

Abstract

The application provides a data sharing method and device, electronic equipment and storage medium, which can be applied to the technical field of computers. The method is applied to a first processor and includes the following steps: obtaining storage paths of multiple candidate data and first identification information, and a one-to-one first mapping relationship between the storage paths and the first identification information; determining, based on multiple target identification information in task information, multiple target data in which the first identification information is target identification information from the candidate data, and generating second identification information of the multiple target data; a one-to-many second mapping relationship exists between the second identification information and the multiple target identification information; and sending the second identification information to a second processor, so that the second processor obtains the multiple target data for processing based on multiple storage paths commonly specified by the second mapping relationship and the first mapping relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a data sharing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Deep learning model-based inference has wide applications in daily life. In actual inference processes, input data and the inference model may be deployed on different servers. This involves loading input data from disk, decoding it, preprocessing it, generating input tensors, and then copying the input tensors to another processor or acceleration device for model inference. This process involves multiple data copies, which may increase data access latency and consume bandwidth and memory resources. Summary of the Invention

[0003] In view of the above problems, this application provides a data sharing method, apparatus, electronic device and storage medium for reducing redundant data movement.

[0004] According to a first aspect of this application, a data sharing method is provided, which is applied to a first processor, comprising: acquiring storage paths and first identification information of multiple candidate data, wherein there is a one-to-one first mapping relationship between the storage paths and the first identification information; determining multiple target data from the candidate data whose first identification information is target identification information based on multiple target identification information in task information, and generating second identification information of the multiple target data; wherein there is a one-to-many second mapping relationship between the second identification information and the multiple target identification information; and sending the second identification information to a second processor so that the second processor can acquire and process the multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship.

[0005] According to a second aspect of this application, a data sharing apparatus is provided, comprising: an acquisition module, configured to acquire storage paths and first identification information of multiple candidate data, wherein a one-to-one first mapping relationship exists between the storage paths and the first identification information; a determination module, configured to determine multiple target data for which the first identification information is the target identification information from the candidate data based on multiple target identification information in task information, and generate second identification information of the multiple target data; wherein a one-to-many second mapping relationship exists between the second identification information and the multiple target identification information; and a sending module, configured to send the second identification information to a second processor, so that the second processor can acquire and process the multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship.

[0006] According to a third aspect of this application, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0007] According to a fourth aspect of this application, a computer-readable storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the steps of the above-described method.

[0008] In this embodiment, by establishing a one-to-one mapping relationship between the storage path and the first identification information, the storage location of the candidate data can be quickly and accurately located. Furthermore, a second identification information is used to represent multiple target data, enabling the second processor to automatically read multiple target data through the second mapping relationship. This eliminates the need for complex logic to read target data one by one, simplifying the data sharing process, reducing the transmission and transfer of redundant information, and improving data sharing efficiency. Attached Figure Description

[0009] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0010] Figure 1 This illustration schematically shows a hardware environment diagram of a data sharing method according to an embodiment of this application;

[0011] Figure 2 A flowchart illustrating a data sharing method according to an embodiment of this application is shown schematically;

[0012] Figure 3 This schematic diagram illustrates the storage path and first identification information for obtaining multiple candidate data according to an embodiment of this application.

[0013] Figure 4 This schematic diagram illustrates the principle of preprocessing raw data to obtain candidate data according to an embodiment of this application.

[0014] Figure 5 The flowchart illustrates a process according to an embodiment of this application, which involves generating second identification information based on different threads and invoking a second processor to process multiple target data.

[0015] Figure 6 This schematic diagram illustrates the principle of generating second identification information of multiple target data based on task information according to an embodiment of this application;

[0016] Figure 7This diagram illustrates the principle of dynamically binding the mapping relationship between the target model input tensor and the storage address through the tensor management interface according to an embodiment of this application;

[0017] Figure 8 This diagram illustrates the principle of dynamically binding the mapping relationship between the target model output tensor and the storage address through the tensor management interface according to an embodiment of this application;

[0018] Figure 9 A schematic block diagram of a computer system for an electronic device according to an embodiment of this application is shown. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] The methods and embodiments provided in this application can be executed on a server, mobile terminal, computer terminal, or similar computing device. Taking running on a server as an example, Figure 1 This diagram schematically illustrates the hardware environment of a method for creating a redundant independent disk array according to an embodiment of this application. Figure 1 As shown, a server may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1The structure shown is for illustrative purposes only and does not limit the structure of the server described above. For example, the server may also include components that are more complex than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0022] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the method for creating an independent redundant disk array in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0023] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0024] Optionally, the data sharing method in this embodiment can be executed by a server. Here, "server" refers to the entire server, including the relevant components within the server that need to execute the data sharing method, as well as the processor, etc.; the data sharing method in this embodiment can also be executed by the processor 102. In some examples of this embodiment, the data sharing method is described using the example of execution by the processor.

[0025] This application provides a data sharing method, apparatus, electronic device, and storage medium. Before introducing the technical solutions provided in this application, the relevant technologies involved in this application will be described first.

[0026] In applications that rely on deep learning for inference, a multi-processor collaborative inference approach is typically used to balance performance and cost. For example, the deep learning model is deployed on a second processor, which handles computationally intensive tasks (such as model computation), while the first processor performs logic-intensive tasks (such as data preprocessing and post-processing).

[0027] This process inevitably involves data interaction between the first and second processors, as well as the generation of multiple data copies. Data copies generated at different stages of deep learning inference may not be effectively reused, easily leading to duplicate copies (such as the same data being transferred multiple times between the first and second processors) and memory redundancy (such as multiple threads / processes independently storing the same data), resulting in problems such as increased memory usage and bandwidth consumption, reduced system throughput, and low parallel efficiency.

[0028] To address the aforementioned issues, this application provides a data sharing method.

[0029] Figure 2 A flowchart illustrating a data sharing method according to an embodiment of this application is shown schematically.

[0030] like Figure 2 As shown, the data sharing method of this embodiment includes operations S210 to S230.

[0031] In operation S210, the storage paths and first identification information of multiple candidate data are obtained, and there is a one-to-one first mapping relationship between the storage paths and the first identification information.

[0032] In this embodiment, multiple candidate data are pre-stored in the target storage area, and the storage path and first identification information of each candidate data are recorded to ensure that a one-to-one first mapping relationship is established between the storage path of the candidate data and the first identification information, that is, each first identification information uniquely corresponds to one storage path. The first mapping relationship can be stored in a database, configuration file or other suitable data structure.

[0033] For example, the target storage area can be a thread-safe cache container in memory. Thread-safe cache containers allow data to be safely accessed and manipulated by multiple threads simultaneously in a multi-threaded environment, avoiding concurrency issues such as data inconsistency and race conditions. For example, the thread-safe cache container can be a Map structure that supports multi-degree shared locks or atomic reference counting.

[0034] In operation S220, based on multiple target identification information in the task information, multiple target data whose first identification information is target identification information are determined from the candidate data, and second identification information of multiple target data is generated; there is a one-to-many second mapping relationship between the second identification information and the multiple target identification information.

[0035] In this embodiment of the application, when a new task arrives, the system receives task information containing multiple target identification information. This target identification information is used to specify the data that needs to be processed. The target identification information has a certain correlation with the first identification information of the candidate data. The target identification information may be the same as the first identification information, or it may be that the target identification information can be mapped to the first identification information through some rule.

[0036] Based on multiple target identifiers in the task information, the system can search for target data whose first identifier matches the target identifier from multiple candidate data. It can also generate corresponding second identifiers for the identified target data. The second identifier can be a randomly generated string, a combination of specific characters, etc., establishing a second mapping relationship between the second identifier and the first identifier corresponding to each of the multiple target data, so that the associated multiple target data can be found through the second identifier.

[0037] In operation S230, the second identification information is sent to the second processor so that the second processor can obtain multiple target data for processing based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship.

[0038] In this embodiment, the generated second identification information is sent to a second processor. Upon receiving the second identification information, the second processor finds the first identification information associated with it according to a second mapping relationship, and converts each first identification information into a corresponding storage path according to the first mapping relationship, thereby determining the specific storage location of multiple target data. The second processor reads target data from the target storage area according to the obtained multiple storage paths and processes the read target data. For example, the second processor can be another module within the system, or a remote server or device. The method of sending the second identification information can be determined according to the system architecture and communication protocol.

[0039] This application establishes a one-to-one mapping relationship between storage paths and first identification information, enabling rapid and accurate location of candidate data storage. Furthermore, second identification information represents multiple target data sets, allowing the second processor to automatically read these target data sets through the second mapping relationship. This eliminates the need for complex logic to read each target data set individually, simplifying the data sharing process, reducing the transmission and transfer of redundant information, and improving data sharing efficiency.

[0040] For example, the first processor mentioned in the embodiments of this application can be a central processing unit (CPU) responsible for processing non-computationally intensive tasks. The second processor can be a graphics processing unit (GPU) or other dedicated accelerators, such as a neural network processing unit (NPU) or a data processing unit (DPU). The specific choice of the second processor depends on the performance, power consumption and cost requirements of the application scenario.

[0041] In this embodiment of the application, before obtaining the storage path and first identification information of the candidate data, a target file can be created in advance so that the storage path and first identification information of the candidate data can be obtained directly from the target file, thereby improving access efficiency.

[0042] The following combination Figure 3 The process of obtaining the storage path and first identification information of multiple candidate data is described, along with specific embodiments.

[0043] Figure 3 The schematic diagram illustrates the storage path and first identification information for obtaining multiple candidate data according to an embodiment of this application.

[0044] like Figure 3 As shown, this embodiment of obtaining the storage paths and first identification information of multiple candidate data includes: writing the storage paths of the candidate data into a target file, whereby the target file is used to integrate the storage paths of multiple candidate data; and generating the first identification information of the candidate data based on the location of the candidate data in the target file.

[0045] In this embodiment of the application, a first file (dataset.txt file) containing the candidate data storage path can be read from an external data source at once to obtain the storage path of the candidate data.

[0046] For example, the dataset.txt file can be stored on disk as text, containing relative or absolute paths to the candidate data. After reading the storage paths of the candidate data from the dataset.txt file, these paths are written to a target file, which can be, for example, a list structure called image_list. The storage paths of the candidate data are written to the target file sequentially, allowing for quick identification of the candidate data and its storage path information. Compared to the dataset.txt file, the target file can be accessed directly through an index, improving access efficiency.

[0047] In this embodiment of the application, a first identifier (i.e., index) can be assigned to each candidate data according to the row number or column number in the target file based on the storage path of the candidate data.

[0048] For example, image_list[0] corresponds to candidate data with number 0, and image_list[1] corresponds to candidate data with number 1. The number of each candidate data in the target file is used as the first identification information of that candidate data. The contents of the target file can be found in Table 1.

[0049] Table 1 Contents of the target file

[0050]

[0051] When it is necessary to access a specific candidate data, the corresponding image path can be quickly found in image_list by specifying the candidate data number (i.e., the first identification information of the candidate data) based on the number, and then the actual image data can be accessed. For example, to access the candidate data with the number 1, the value of image_list[1] can be obtained directly from image_list, which is the storage path of the candidate data, and then the candidate data can be read according to this storage path.

[0052] This application's embodiments enable candidate data path pre-indexing and numbered addressing access. For example, only the first identifier information of the target data to be processed needs to be determined to obtain the storage address of the candidate data from the target file. There is no need to perform path concatenation or path resolution (such as concatenating directories, processing file names, etc.) on each access to candidate data; the storage path of the candidate data can be directly obtained from the target file based on the first identifier and accessed using the storage path, effectively improving data processing efficiency and reducing redundant calculations.

[0053] In this embodiment, the candidate data is data obtained after preprocessing the original image. Before performing operation S210, the data sharing method further includes: preprocessing the original data to obtain candidate data.

[0054] The following is combined Figure 4 The principles of raw data preprocessing are introduced, along with specific implementation examples.

[0055] Figure 4 The diagram illustrates the principle of preprocessing raw data to obtain candidate data according to an embodiment of this application.

[0056] like Figure 4As shown, the preprocessing of raw data to obtain candidate data in this embodiment includes: preprocessing multiple raw data sets to obtain multiple candidate data sets; the multiple candidate data sets are in a format that can be processed by a second processor. The candidate data are stored in a target storage area, and the storage path of the candidate data is determined. The first identification information of each of the multiple candidate data sets is associated with the storage path to obtain a first mapping relationship.

[0057] In this embodiment, the raw data may come from different sensors, databases, or file systems. These data may have the same or different formats. By preprocessing this raw data, it is converted into candidate data that meets the processing requirements of the second processor. Preprocessing may include operations such as data transformation and data normalization.

[0058] Taking a scenario where the original data is image data and the second processor processes the data by calling a neural network model as an example, neural network models typically require tensor inputs in a specific format (e.g., floating-point tensors with shape=(3,224,224)). Therefore, before processing the data based on the neural network model, the original data needs to be preprocessed. Preprocessing steps may include image decoding (JPEG→RGB), size scaling (Resize to the model input size), pixel normalization (Mean / Std normalization), channel arrangement (HWC→CHW), and data type conversion (uint8→float32). Through preprocessing operations, candidate data that can be recognized by the neural network model is obtained.

[0059] In this embodiment of the application, preprocessing is performed on all the original images to be processed at the beginning of the program. The preprocessing of all the original images is returned at once, which can effectively avoid repeating the preprocessing operation every time the second processor performs inference, reduce the amount of computation in the second processor's inference process, and significantly improve the inference speed.

[0060] In this embodiment, a suitable storage region can be selected as the target storage region to store candidate data based on factors such as data size and access frequency. The candidate data is stored in the target storage region according to a predetermined storage format, and a unique storage path is determined for each candidate region. The storage path can be generated, for example, based on data characteristics, timestamps, or other information.

[0061] In this embodiment, the first identifier information of each candidate data can be associated with its corresponding storage path to obtain a first mapping relationship. Candidate data can be stored in the target storage area as key-value pairs, with the first identifier information used as the key-value identifier to establish the first mapping relationship between the first identifier information of the candidate data and the storage path. When the result of obtaining candidate data is needed, only the first identifier information needs to be provided to retrieve the corresponding candidate data from the target storage area, improving data access efficiency.

[0062] This application embodiment preprocesses the raw data once and stores the preprocessing result (i.e., candidate data) in the target storage area, which effectively avoids repeatedly performing preprocessing operations before each inference. This effectively reduces the waste of computing time and memory copying, and improves the efficiency of data processing.

[0063] Furthermore, operations S220 and S230 in this embodiment can be implemented using different threads. Through the division of labor and cooperation between different threads, task scheduling and data processing are decoupled, thereby effectively improving resource utilization and avoiding serial blocking. The following is in conjunction with... Figure 5 The process of performing operations S220 and S230 based on multiple threads will be described with specific embodiments.

[0064] Figure 5 The schematic diagram illustrates the principle of generating second identification information based on different threads and calling a second processor to process multiple target data according to an embodiment of this application.

[0065] like Figure 5 As shown, the generation of second identification information based on different threads and the invocation of the second processor to process multiple target data in this embodiment include operations S510 to S520.

[0066] In operation S510, a first thread is started based on the received task information. The first thread is used to generate second identification information for multiple target data based on the task information and add the second identification information to the ready queue.

[0067] In this embodiment, upon receiving task information from an external source, a first thread is initiated. The task information may include a description of the target data to be processed, the processing type, and processing requirements. For example, in an image classification task, the task information may specify that a batch of images from a specific scene needs to be classified and the classification results provided.

[0068] In this embodiment, the first thread generates second identification information for multiple target data based on task information and adds the generated second identification information to the ready queue. The ready queue is a data structure used to store identification information of target data waiting to be processed, and can serve as a task scheduling and buffering mechanism. For example, if the task information involves processing a series of images, the second identification information can be a unique identifier for this series of images.

[0069] When operating S520, a second thread is started based on the second identification information. The second thread is used to call the second processor to obtain multiple target data for processing based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship.

[0070] In this embodiment, a second thread is started when identification information exists in the ready queue. The second thread can be triggered by polling the ready queue, event notification, or other methods. For example, the system can check the ready queue periodically to see if it is empty; if not, the second thread is started, and the second thread calls a second processor to process the target data.

[0071] In this embodiment of the application, the second thread can combine the first mapping relationship and the second mapping relationship to determine the accurate storage path of each target data. For example, it can determine the first identification information of multiple target data to be acquired according to the second mapping relationship, and locate the storage path corresponding to the first identification information according to the first mapping relationship, so that the second processor can directly read the target data from the storage path and reduce data copying in the intermediate process.

[0072] This embodiment of the application achieves simultaneous task scheduling and data processing through the division of labor and cooperation between a first thread and a second thread: the first thread is responsible for generating identification information to provide an orderly task flow for subsequent processing; the second thread focuses on the actual processing of the target data, making full use of the system's multi-core resources and improving the system's concurrent processing capabilities. The second identification information generated by the first thread based on the task information enables the second thread to read the corresponding target data at once, improving data reading efficiency.

[0073] In this embodiment of the application, generating the second identification information based on different threads and calling the second processor to process multiple target data may further include: triggering a target event in response to the first thread inputting the second identification information into the ready queue, the target event being used to notify the second thread to retrieve the second identification information from the ready queue.

[0074] In this embodiment, the ready queue can be, for example, a first-in, first-out (FIFO) data structure used to temporarily store second identification information awaiting processing. When the first thread successfully inputs one or more pieces of second identification information into the ready queue, a predefined target event is triggered. This target event notifies the second thread responsible for processing the target data that second identification information available for processing exists in the ready queue. The target event can be a signal, a message, or a callback function call.

[0075] In this embodiment, the second thread can register to listen for target events when the system starts. When the target event is triggered, the second thread will immediately receive a notification and enter a response state, retrieving one or more second identifiers from the ready queue according to the first-in-first-out principle. After retrieving the second identifiers, the second thread can call the second server to obtain the corresponding target data for processing.

[0076] This application embodiment achieves asynchronous communication through a target event notification mechanism. After triggering a target event, the first thread does not need to wait for the response from the second thread and can continue to execute other tasks, avoiding thread blocking. The second thread can respond to the target event and obtain the second identifier information at an appropriate time without needing to poll the ready queue in real time, reducing unnecessary overhead and improving response speed.

[0077] In this embodiment, a task queue may also be included, which is filled by a task generator thread and used to transmit task information. The task generator thread can generate target tasks according to preset rules or external input, and put the target tasks into the task queue. The first thread retrieves task information from the task queue.

[0078] In this embodiment, after filling the prepared second identifier information into the ready queue, the first thread acquires a global lock to update the task count and status, ensuring the atomicity of these operations in a multi-threaded concurrent environment. For example, updating the number of prepared second identifier information items and marking the current number of second identifier information items as ready. After completing data preparation and status updates, a target event is triggered to notify the second thread that the batch of data is ready and processing operations can begin. The global lock ensures the atomicity of task counting and status update operations in a multi-threaded concurrent environment. By acquiring the global lock when updating the task count and status, the first thread can prevent other threads from modifying this data simultaneously, thus avoiding data inconsistency issues.

[0079] In this embodiment, after the target data processing is completed, the second thread can write the task status (such as successful processing, failed processing, etc.) and processing result into a task completion table. The task completion table is a data structure used to record the task execution status. The task completion table can be queried periodically through scheduling information, and subsequent processing can be performed according to the task status, such as returning the processing result to the user or retrying failed tasks.

[0080] In this embodiment, the task generator thread, filling thread, and inference thread can work in parallel. By utilizing different threads to process tasks at different stages, idle waiting between threads can be effectively avoided, improving concurrent processing capabilities. The target event notification mechanism can effectively reduce the latency caused by each step, eliminating the need for frequent queries of the ready queue and effectively reducing communication overhead and waiting time between threads.

[0081] The data sharing method provided in this application embodiment can improve the efficiency of the second thread in data access and data processing by integrating the target identification information in the task information to obtain second identification information that can be accessed in one go after the first thread receives the task information.

[0082] The following is combined Figure 6 The principle of generating the second identification information is explained through specific embodiments.

[0083] Figure 6 The schematic diagram illustrates the principle of generating second identification information of multiple target data based on task information according to an embodiment of this application.

[0084] like Figure 6 As shown, this embodiment generates second identification information for multiple target data based on task information, including: determining target data and the storage path of the target data from candidate data based on the target identification information. The second identification information is generated for multiple target data based on the target identification information and / or storage paths, wherein the second identification information is a set of target identification information and / or storage paths for multiple target data.

[0085] In this embodiment, the first thread receives task information sent by the upstream task generator. The task information includes multiple target identifiers, such as [3, 8, 12, 20]. The target identifiers represent specific candidate data that needs to be processed. Based on the received target identifiers, the first thread determines the target data from the multiple candidate data in the target file and reads the storage path of the target data.

[0086] In this embodiment, the first thread can concatenate the read first identifier information and / or storage path into a second identifier information according to a specific function. Using the second identifier information, the parallel capabilities of the hardware can be utilized to perform inference calculations on multiple candidate data simultaneously. For example, candidate data in the same batch can be loaded into a second processor at once, and then processed simultaneously by parallel computing units, shortening inference time. Batch processing can optimize memory access patterns, reducing the number of memory accesses and latency. When processing batch data, the hardware can more effectively utilize caching mechanisms to improve data reading speed, thereby further enhancing inference performance.

[0087] In this embodiment, the positions of candidate data are recorded using an index list or pointer array, and multiple target data are logically concatenated to simulate batch processing. It should be noted that the actual candidate data is still scattered across the target storage area, but can be accessed all at once using the second identifier information. The second identifier information acts as an index list, allowing the second thread to indirectly access the target data without moving the data itself. After logical concatenation, the system can reduce communication overhead by making batch requests (e.g., telling the hardware at once, "I need data from these 3 addresses"), simulating the effect of "one-time access." For example, the second thread can submit multiple data access requests to the hardware (e.g., memory, GPU, disk) at once based on the second identifier information: "I need data from address A (image 3), address B (image 8), and address C (image 12)." The second identifier information establishes a mapping relationship between scattered data and overall access.

[0088] Furthermore, the data sharing method provided in this application embodiment, after the second thread obtains the second identification information from the scheduling queue, calls the second processor to acquire and process the target data.

[0089] In this embodiment of the application, calling the second processor to obtain multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship for processing includes: running the target model on the second processor through the target framework to process the multiple target data and obtain the processing result.

[0090] In this embodiment, the target framework can provide a tensor management interface, which is used to dynamically bind the association between the target model's input tensors and output tensors and their storage addresses. The input tensor is the carrier of data received by the target model, and the output tensor is the carrier of the processing result obtained by the target model after processing the target data. The association is used to clarify the specific storage locations of the input and output tensors. The storage address includes the storage address of the candidate data specified by the storage path and the storage address corresponding to the processing result. The association between the input tensor and the storage address can be determined based on both the second mapping relationship and the first mapping relationship.

[0091] In this embodiment, the second processor can process target data by running the target model. The target framework can be, for example, the ONNX Runtime. The ONNX Runtime runtime environment is started on the second processor to complete the inference of the target model based on the ONNX Runtime framework. In this process, the second processor, as the hardware device that actually executes the processing tasks, provides the physical basis for the operation of the target framework. The target framework, as a software framework, completes the inference of the target model by calling the processor's instruction set and hardware acceleration functions.

[0092] In this embodiment, the ONNX Runtime provides an interface called io_binding for tensor management. io_binding connects the input and output tensors of the target model to the storage addresses of the data. For the input tensors of the target model, the target storage path of the target data can be connected to the input tensor used by the target model to receive the target data, based on the second and first identification information, so that the target model knows where to obtain the data to be processed. For the output of the target model, a storage area for storing the processing results can be pre-allocated, and this storage area can be associated with the output tensors of the target model, so that the target model stores the processing results in this storage area after data processing is completed.

[0093] In this embodiment of the application, after the configuration of io_binding is completed, the target model is started to perform inference calculation. The target model will obtain target data according to the bound input tensor, process it according to the algorithm set inside the model, and store the processing result in the storage area corresponding to the pre-bound output tensor.

[0094] This application's embodiments establish a direct association between storage addresses and the model's input / output tensors through a tensor management interface, effectively avoiding the need for data to be copied back and forth between different storage areas as in traditional methods. This allows the target model to directly obtain target data through the association, reducing the data copying process and time, thereby effectively improving data processing speed.

[0095] Furthermore, in the data sharing method provided in this application embodiment, before calling the second processor to run the target model, the mapping relationship between the target model input tensor and the storage address can be dynamically bound through the tensor management interface, so that the target model can directly access the target data and process the target data quickly through the mapping relationship when performing inference.

[0096] The following is combined Figure 7 The principle of dynamically binding the input tensor and storage address of the target model is explained in detail with specific embodiments.

[0097] Figure 7 The diagram illustrates the principle of dynamically binding the mapping relationship between the target model input tensor and the storage address through the tensor management interface according to an embodiment of this application.

[0098] like Figure 7 As shown, this embodiment dynamically binds the mapping relationship between the target model input tensor and its storage address through a tensor management interface. This includes: generating first binding information based on second identification information, whereby the first binding information includes the storage address of the target data and parameter information of the target data. The first binding information is then input to the tensor management interface, which generates a metadata object matching the target data. This metadata object records a third mapping relationship between the input tensor and the storage address of the target data, enabling the target model to directly access the target data during inference through this third mapping relationship.

[0099] In this embodiment, the first binding information is the input parameter for generating the metadata object, which may include the storage address of the target data and parameter information (such as data type, shape, dimension, etc.). The storage address is used to inform the target model of the location of the data, and the parameter information is used to provide a "description" of the target data to the target model, so as to ensure that the target model can correctly read and parse the target data in the expected manner.

[0100] In this embodiment of the application, the storage address of the target data can be determined and the parameter information of the target data can be obtained based on the second identification information. The parameter information can be obtained through predefined configuration or through the target file corresponding to the first identification information.

[0101] In this application's embodiment set, the metadata object is dynamically created by the tensor management interface based on the first binding information, and is used to internally maintain the mapping relationship (i.e., the third mapping relationship) between the storage address of the input tensor and the target data. The metadata object can encapsulate metadata information such as the storage address, device type, data type, and shape of the target data, and provides a unified abstraction across devices. The tensor management interface achieves zero-copy binding of input and output data through the metadata object, thereby avoiding the data copying overhead in traditional methods.

[0102] In this embodiment, the metadata object can be obtained implicitly by providing the storage address of the target data and its parameter information (i.e., the first binding information) to the tensor management interface. Upon receiving the first binding information from IOBinding, the ONNX Runtime internally calls the constructor of OrtValue to generate a metadata object (OrtValue object) that matches the user data. The tensor binding interface can directly bind to the storage address of the target data through the metadata object, effectively avoiding repeated memory copies. Creating the metadata object marks the target data as potentially requiring future access by a second processor.

[0103] In this embodiment, in response to the creation of the metadata object, the CPU memory page containing the target data can be locked using CUDA's cudaHostRegister or a similar mechanism, making it pinned memory. This allows for asynchronous data transfer from the locked host memory to the GPU memory directly via DMA with the second processor (without additional copying).

[0104] It's important to note that the target data is not actually copied during this process. Instead, a "mapped view" is created, meaning the target data is already prepared and can be used directly on the second processor. No intermediate copying steps are required between host memory and GPU memory, thus avoiding the latency associated with traditional data copying.

[0105] In this embodiment, metadata objects can be automatically created by inputting the first binding information tensor management interface into the tensor management interface. This informs the ONNX Runtime that the input data is ready and located at the specified storage address, eliminating the need for data copying and allowing direct use of the data.

[0106] This application's embodiments utilize a model interface binding mechanism to directly map processed target data to a second processor, eliminating the need for intermediate data copying or relocation. This avoids implicit copying operations in explicit uploads and effectively reduces data transfer load between processors. This zero-copy memory binding method allows target data to be directly used in GPU memory, significantly reducing data transfer overhead between the host and GPU, and improving inference efficiency and throughput.

[0107] Furthermore, in the data sharing method provided in this application embodiment, before calling the second processor to run the target model, the mapping relationship between the target model output tensor and the storage address can be dynamically bound through the tensor management interface, so that the target model can directly and quickly output the processing result through the mapping relationship when performing inference.

[0108] The following is combined Figure 8 The document provides a detailed explanation of the principle of dynamically binding the output tensor and storage address of the target model, along with specific implementation examples.

[0109] Figure 8 The diagram illustrates the principle of dynamically binding the mapping relationship between the target model output tensor and the storage address through the tensor management interface according to an embodiment of this application.

[0110] like Figure 8 As shown, this embodiment dynamically binds the mapping relationship between the target model output tensor and the storage address through the tensor management interface, including: generating second binding information based on a pre-allocated storage area; the storage area is located in the second processor. The second binding information is input to the tensor management interface, which determines a fourth mapping relationship between the output tensor and the storage area, so that the target model outputs the processing result to the storage area specified by the fourth mapping relationship.

[0111] In this embodiment of the application, the second binding information may include, for example, the name of the target model output tensor, the device type (specifying whether the output is to the GPU or the CPU), the device ID, the target storage area pointer, etc.

[0112] In this embodiment, the second binding information is input to the tensor management interface. Based on the received second binding information, the tensor management interface determines a fourth mapping relationship between the output tensor and the storage area. After the target model is output, the ONNX Runtime directly writes the model output to a predefined storage area based on the fourth mapping relationship. The storage area can be a pre-allocated storage space, and the model output can be written to this storage area using the storage address corresponding to the processing result. For example, the pre-allocated storage area can be the storage area on a second processor (such as GPU memory).

[0113] This application embodiment directly outputs the processing results of the target model to the GPU memory based on the fourth mapping relationship, achieving zero-copy output. This allows the model output data to remain on the second processor, reducing the copying process between the second and first processors. If subsequent operations are still performed on the second processor, the processing results in the storage area can be used directly without any copying. If it is necessary to copy the processing results to the first processor, it can be triggered by user operations, avoiding unnecessary overhead and improving the flexibility of processing result storage.

[0114] In some embodiments, the target model can be invoked by multiple second threads to achieve parallel processing of data, reduce processing latency, and meet data processing needs in scenarios such as large-scale batch processing.

[0115] In scenarios where target processing is performed by calling the target model through multiple second threads, the data sharing method provided in this application embodiment further includes: processing multiple target data based on calling the target model through multiple second threads; wherein, each second thread independently binds its own input tensor storage address and output tensor storage address through an interface; the input tensor storage address is the storage address of the candidate data specified by the storage path, and the output tensor storage address is the storage address corresponding to the processing result.

[0116] In this embodiment, the target model can be called in parallel by multiple second threads to process different input data or different batches of the same data. An independent IOBinding interface is created for each second thread to independently bind the storage addresses of the input tensor and the output tensor, so as to realize data isolation between threads and reuse of model resources.

[0117] This application embodiment enables multi-threaded parallel calls to the target model by creating independent IOBinding interfaces for multiple second threads. This allows each second thread to bind its own input tensor and the storage address corresponding to its output tensor through an independent IOBinding interface, effectively achieving data isolation.

[0118] According to the data sharing method provided in the embodiments of this application, after preprocessing the original data to obtain candidate data, before storing the candidate data in the target storage area, the candidate data can be classified, and the target storage area corresponding to the candidate data can be determined based on the data type of the candidate data.

[0119] In this embodiment of the application, storing candidate data in the target storage area and determining the storage path of the candidate data may further include: storing the candidate data in the corresponding storage area according to the data type of the candidate data.

[0120] In this embodiment, the storage area includes a first storage area and a second storage area. The first storage area is a high-speed storage area, and the second storage area is a low-speed storage area; that is, the storage latency of the first storage area is lower than that of the second storage area. It should be noted that "high-speed storage" and "low-speed storage" are relative concepts here.

[0121] In this embodiment, the target storage area corresponding to the candidate data can be determined based on the data type of the candidate data. For example, the data type may include a first type of data and a second type of data, where the first type of data is accessed more frequently than the second type of data, the first type of data is stored in the first storage area, and the second type of data is stored in the second storage area.

[0122] In the embodiments of this application, determining the data type of candidate data includes at least one of the following: determining the data type based on the historical access frequency of candidate data; determining the data type based on the access information of candidate data in a specific time window; and determining the data type based on the association information between candidate data and target data.

[0123] For example, the number of times candidate data was accessed in the past 24 hours is counted, and the historical access frequency of each candidate data is calculated. Candidate data whose access frequency exceeds a preset threshold (e.g., 100 times / hour) is classified into one category. Candidate data with a significantly increased access frequency within a specific time period (e.g., 10:00-12:00) is classified into another category. Candidate data with a strong correlation (e.g., co-occurrence probability > 90%) with the target data (e.g., input data of the current target model inference task) is identified as another category. Candidate data of category one is allocated to the first storage area for storage. The remaining candidate data that does not meet the criteria for category one is allocated to the second storage area. The first storage area can be memory, and the second storage area can be disk.

[0124] This application's embodiments allocate different types of data to different storage areas, effectively reducing access latency while lowering storage costs and avoiding excessive memory expansion. Furthermore, candidate data is dynamically classified based on access frequency, time windows, and correlation, improving the flexibility of candidate data classification.

[0125] In embodiments of this application, the data sharing method may further include: storing verification information of candidate data in a target storage area, the verification information including preprocessing parameters and original data information; verifying the candidate data before determining the target data from the candidate data based on the target identification information; and changing the candidate data in the target storage area to invalid data if the verification information is inconsistent with the parameter information of the target data.

[0126] In this embodiment, verification information is generated during the preprocessing of the original data. This verification information may include preprocessing parameters (such as data hash value, timestamp, version number, etc.) and original data information (such as data size, data type, etc.). The verification information is bound to candidate data and stored in the target storage area. When the first thread determines the target data from the candidate data based on the target identifier information, the verification process is automatically triggered. The validity of the candidate data is determined by comparing the verification information with the parameter information of the candidate data. If the verification information matches the reference information of the target data, the candidate data is confirmed as the target data. If the verification information does not match the reference information of the target data, the candidate data in the target storage area is marked as invalid data.

[0127] In this embodiment of the application, when determining target data, the validity of candidate data is verified by the verification information bound to candidate data in the target storage area, so as to ensure data consistency and security, prevent data tampering, data loss and other problems, and ensure the correctness of data sharing.

[0128] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to perform the steps of any of the above method embodiments through the computer program.

[0129] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0130] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0131] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein the program executes the steps in any of the above method embodiments when it is run.

[0132] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0133] According to another aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program / instructions comprising program code for performing the method shown in the flowchart. In such an embodiment, reference is made to... Figure 9 The computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit 901, it performs various functions provided in the embodiments of this application. The sequence numbers of the embodiments of this application above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0134] Figure 9 A schematic block diagram of a computer system for an electronic device according to an embodiment of this application is shown. Figure 9As shown, the computer system 900 includes a central processing unit (CPU) 901, which performs various appropriate actions and processes based on programs stored in read-only memory (ROM) 902 or programs loaded from storage section 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for system operation. The CPU 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output interface 905 (I / O interface) is also connected to the bus 904.

[0135] The following components are connected to the input / output interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a local area network card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.

[0136] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 909, and / or installed from removable medium 911. When the computer program is executed by central processing unit 901, it performs various functions defined in the system of this application.

[0137] It should be noted that, Figure 9 The computer system 900 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0138] According to another aspect of the embodiments of this application, a data sharing apparatus is also provided, comprising: an acquisition module, configured to acquire storage paths and first identification information of multiple candidate data, wherein there is a one-to-one first mapping relationship between the storage paths and the first identification information; a determination module, configured to determine multiple target data from the candidate data whose first identification information is target identification information based on multiple target identification information in task information, and generate second identification information of the multiple target data; wherein there is a one-to-many second mapping relationship between the second identification information and the multiple target identification information; and a sending module, configured to send the second identification information to a second processor, so that the second processor acquires and processes the multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship.

[0139] It should be noted that the acquisition module in this embodiment can be used to perform the above operation S210, the determination module in this embodiment can be used to perform the above operation S220, and the sending module in this embodiment can be used to perform the above operation S230.

[0140] The embodiments provided in this application establish a one-to-one mapping relationship between the storage path and the first identification information, enabling rapid and accurate location of the candidate data's storage location. Furthermore, by using the second identification information to represent multiple target data, the second processor can automatically read multiple target data through the second mapping relationship, eliminating the need for complex logic to read each target data individually. This simplifies the data sharing process, reduces the transmission and transfer of redundant information, and improves data sharing efficiency.

[0141] Optionally, the acquisition module includes: a writing submodule, used to write the storage path of the candidate data to the target file, the target file being used to integrate the storage paths of multiple candidate data; and a generation submodule, used to generate the first identification information of the candidate data based on the location of the candidate data in the target file.

[0142] Optionally, the apparatus further includes: a preprocessing module for preprocessing multiple raw data to obtain multiple candidate data; the multiple candidate data being in a format that can be processed by a second processor; storing the candidate data in a target storage area and determining the storage path of the candidate data; associating the first identification information of each of the multiple candidate data with the storage path to obtain a first mapping relationship.

[0143] Optionally, the device further includes: a startup module, configured to start a first thread based on received task information, the first thread being configured to generate second identification information for multiple target data based on the task information and add the second identification information to a ready queue; and to start a second thread based on the second identification information, the second thread being configured to call a second processor to obtain multiple target data for processing based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship.

[0144] Optionally, the device further includes: a triggering module, configured to trigger a target event in response to the first thread inputting the second identification information into the ready queue, the target event being used to notify the second thread to retrieve the second identification information from the ready queue.

[0145] Optionally, the determining module includes: a determining submodule, used to determine target data and the storage path of the target data from the candidate data based on the target identification information; and to generate second identification information of multiple target data based on the target identification information and / or storage path, wherein the second identification information is a set of target identification information and / or storage paths of multiple target data.

[0146] Optionally, the sending module includes: a running submodule, used to run the target model on the second processor through the target framework to process multiple target data and obtain processing results; wherein, the target framework is used to provide a tensor management interface, which is used to dynamically bind the association between the target model's input tensors and output tensors and storage addresses, the association being determined based on a second mapping relationship and a first mapping relationship, and the storage address including the storage address of the candidate data specified by the storage path and the storage address corresponding to the processing results.

[0147] Optionally, the running submodule includes: a first running unit, used to generate first binding information based on second identification information, the first binding information including the storage address of the target data and parameter information of the target data; inputting the first binding information to the tensor management interface, the interface generating a metadata object matching the target data, the metadata object recording a third mapping relationship between the input tensor and the storage address of the target data, so that the target model can directly access the target data through the third mapping relationship during inference.

[0148] Optionally, the running submodule includes: a second running unit, used to store and generate second binding information based on a pre-allocated storage area; the storage area is located in the second processor; the second binding information is input to a tensor management interface, and the interface determines a fourth mapping relationship between the input tensor and the storage area, so that the target model outputs the inference results to the storage area specified by the fourth mapping relationship.

[0149] Optionally, the device further includes: a calling module, used to call the target model based on multiple second threads to process multiple target data; wherein each second thread independently binds its own input tensor storage address and output tensor storage address through an interface; the input tensor storage address is the storage address of the candidate data specified by the storage path, and the output tensor storage address is the storage address corresponding to the processing result.

[0150] Optionally, the device further includes: a storage module, used to store candidate data in corresponding storage areas according to the data type of the candidate data, the storage areas including a first storage area and a second storage area, the storage latency of the first storage area being less than that of the second storage area, the data types including first-class data and second-class data, the access frequency of first-class data being higher than that of second-class data, the first-class data being stored in the first storage area, and the second-class data being stored in the second storage area.

[0151] Optionally, the storage module may further include: a determination submodule, used to determine the data type based on the historical access frequency of the candidate data; to determine the data type based on the access information of the candidate data in a specific time window; and to determine the data type based on the association information between the candidate data and the target data.

[0152] Obviously, those skilled in the art should understand that the modules or steps of the embodiments of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of this application are not limited to any particular combination of hardware and software.

[0153] The above are merely preferred embodiments of this application and are not intended to limit the embodiments of this application. For those skilled in the art, various modifications and variations can be made to the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the embodiments of this application should be included within the protection scope of the embodiments of this application.

Claims

1. A data sharing method, characterized in that, The method is applied to a first processor and includes: The storage paths and first identification information of multiple candidate data are obtained, and there is a one-to-one first mapping relationship between the storage paths and the first identification information; the first identification information is determined based on the position of the candidate data in the target file, and the target file is used to integrate the storage paths of the multiple candidate data to be written. The first thread is started based on the task information. The first thread determines multiple target data whose first identification information is the target identification information and the storage path of the target data from the candidate data based on multiple target identification information in the task information, and generates second identification information of the multiple target data. There is a one-to-many second mapping relationship between the second identification information and the multiple target identification information. The second identification information is a set of target identification information and / or storage paths of the multiple target data. A second thread is started based on the second identification information, and the second thread calls the second processor to obtain the multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship for processing. The step of calling the second processor to obtain and process the multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship includes: running a target model on the second processor through a target framework to process the multiple target data and obtain processing results; wherein, the target framework is used to provide a tensor management interface, the tensor management interface is used to dynamically bind the association relationship between the target model input tensor and the storage address, the association relationship is determined based on the second mapping relationship and the first mapping relationship, and the storage address includes the storage address of the candidate data specified by the storage path.

2. The method according to claim 1, characterized in that, The acquisition of the storage path and first identifier information of multiple candidate data includes: Write the storage path of the candidate data into the target file; Based on the location of the candidate data in the target file, the first identification information of the candidate data is generated.

3. The method according to claim 2, characterized in that, The method further includes: Multiple raw data are preprocessed to obtain multiple candidate data; the multiple candidate data are in a format that can be processed by the second processor. Store the candidate data in the target storage area and determine the storage path of the candidate data; The first identification information of each of the multiple candidate data is associated with the storage path to obtain the first mapping relationship.

4. The method according to claim 1, characterized in that, Also includes: In response to the first thread inputting the second identification information into the ready queue, a target event is triggered, which is used to notify the second thread to retrieve the second identification information from the ready queue.

5. The method according to claim 1, characterized in that, The tensor management interface is also used to dynamically bind the relationship between the target model output tensor and the storage address, wherein the storage address includes the storage address corresponding to the processing result.

6. The method according to claim 5, characterized in that, The process of running the target model on the second processor via the target framework includes: First binding information is generated based on the second identification information, and the first binding information includes the storage address of the target data and the parameter information of the target data; The first binding information is input to the tensor management interface, which generates a metadata object that matches the target data. The metadata object records a third mapping relationship between the input tensor and the storage address of the target data, so that the target model can directly access the target data through the third mapping relationship during inference.

7. The method according to claim 5, characterized in that, Also includes: The second binding information is generated based on the pre-allocated storage area; The storage area is located in the second processor; The second binding information is input to the tensor management interface, which determines the fourth mapping relationship between the output tensor and the storage area, so that the target model outputs the inference result to the storage area specified by the fourth mapping relationship.

8. The method according to claim 5, characterized in that, Also includes: The target data is processed by calling the target model through multiple second threads; Each second thread independently binds its own input tensor storage address and output tensor storage address through the tensor management interface; the input tensor storage address is the storage address of the candidate data specified by the storage path, and the output tensor storage address is the storage address corresponding to the processing result.

9. The method according to claim 1, characterized in that, The method further includes: Based on the data type of the candidate data, the candidate data is stored in corresponding storage areas, including a first storage area and a second storage area. The storage latency of the first storage area is less than that of the second storage area. The data type includes a first type of data and a second type of data. The first type of data is accessed more frequently than the second type of data. The first type of data is stored in the first storage area, and the second type of data is stored in the second storage area.

10. The method according to claim 9, characterized in that, The determination of the data type of the candidate data includes at least one of the following: The data type is determined based on the historical access frequency of the candidate data; The data type is determined based on the access information of the candidate data within a specific time window; The data type is determined based on the correlation information between the candidate data and the target data.

11. A data sharing device, characterized in that, The device includes: An acquisition module is used to acquire the storage paths and first identification information of multiple candidate data, wherein there is a one-to-one first mapping relationship between the storage paths and the first identification information; the first identification information is determined based on the position of the candidate data in the target file; the target file is used to integrate the storage paths of the multiple candidate data to be written. The determination module is used to start a first thread based on task information. The first thread determines multiple target data whose first identification information is the target identification information and the storage path of the target data from the candidate data based on multiple target identification information in the task information, and generates second identification information of the multiple target data. There is a one-to-many second mapping relationship between the second identification information and the multiple target identification information. The second identification information is a set of target identification information and / or storage paths of the multiple target data. The sending module is used to start a second thread based on the second identification information, and the second thread calls the second processor to obtain the multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship for processing. The step of calling the second processor to obtain and process the multiple target data based on multiple storage paths jointly specified by the second mapping relationship and the first mapping relationship includes: The target model is run on the second processor through the target framework to process the multiple target data and obtain the processing result; wherein, the target framework is used to provide a tensor management interface, which is used to dynamically bind the association between the target model input tensor and the storage address, the association is determined based on the second mapping relationship and the first mapping relationship, and the storage address includes the storage address of the candidate data specified by the storage path.

12. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Business data processing method and device, computer equipment and storage medium

    CN113590304A

  • Business processing method and device, electronic equipment and storage medium

    CN119025101A