Data processing method and device, electronic equipment, storage medium and program product
By directly splicing and quantizing embedded vectors within the large model framework and combining different modal data for efficient inference, the problem of low efficiency in multimodal data processing is solved, and efficient multimodal data processing and memory optimization are achieved.
Patent Information
- Application Number
- CN202510660939.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-10-10
AI Technical Summary
The inference processing efficiency of multimodal data in existing technologies is low, and it is difficult to effectively combine different types of data for efficient inference.
Within the large model framework, the first inference result generated by the first model runtime engine is used as the input of the second model runtime engine. By directly splicing the embedding vectors of the first and second data, the quantization and dequantization parameters are used to execute in the same memory area, the memory area is shared to reduce memory usage, and the appropriate hardware carrier is selected for calculation.
It improves the efficiency and accuracy of multimodal data reasoning, reduces memory consumption and computing resource requirements, adapts to hardware changes without modifying model code, and realizes reasoning tasks of different paradigms.
Smart Images

Figure CN120764664A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, device, electronic device, storage medium, and program product. Background Art
[0002] With the rapid development of large model technology, more and more large models need to be deployed on edge devices, such as vehicles and mobile devices like mobile phones. Edge devices can process multiple types of data, including visual, text, and voice, and perform tasks such as environmental perception, path planning, and voice interaction.
[0003] In related technologies, multimodal data is classified in the model application layer, different types of data are transmitted to the corresponding models for inference, and the inference results of each model are returned to the application layer for merging to obtain the inference results of the multimodal data. Summary of the Invention
[0004] The present disclosure provides a data processing method, device, electronic device, storage medium and program product to solve the problem of low efficiency of reasoning processing in related technologies.
[0005] According to one aspect of an embodiment of the present disclosure, a data processing method is provided, which includes: receiving multimodal data, wherein the multimodal data includes first data and second data; calling a first interface to transmit the first data to a first model runtime engine, and receiving a first inference result output by the first model runtime engine; obtaining third data based on the first inference result and the second data; transmitting the third data to the second model runtime engine for processing; and outputting the inference result of the multimodal data.
[0006] Within the large model framework, the first inference result generated by the first model runtime engine is directly used as part of the input of the second model runtime engine, so that the second model runtime engine can perform inference based on the output of the first model runtime engine and the second data, thereby improving the efficiency of the inference processing.
[0007] In some exemplary embodiments of the present disclosure, the first inference result includes a first embedding vector, and the third data includes a third embedding vector; obtaining the third data according to the first inference result and the second data includes: obtaining a second embedding vector according to the second data; and concatenating the first embedding vector and the second embedding vector to obtain the third embedding vector.
[0008] The first inference result and the second data are directly concatenated in the embedding vector dimension to form a higher-dimensional joint vector. The embedding vector has been encoded into a high-dimensional semantic representation, which improves the efficiency of data reasoning.
[0009] In some exemplary embodiments of the present disclosure, the second data includes a prompt word; obtaining a second embedding vector based on the second data includes: segmenting the prompt word to obtain a Token (word unit) sequence; and embedding the Token sequence to obtain the second embedding vector.
[0010] In multimodal data processing tasks, prompt words are used to guide the model to understand data from other modalities. After converting the prompt words into embedding vectors, they are effectively fused or interacted with feature vectors from other modalities to improve data inference efficiency.
[0011] In some exemplary embodiments of the present disclosure, concatenating the first embedding vector and the second embedding vector to obtain a third embedding vector includes: quantizing the first embedding vector using a first parameter; dequantizing the quantized first embedding vector using a second parameter; and concatenating the dequantized first embedding vector and the second embedding vector to obtain the third embedding vector.
[0012] The quantized embedded vector can reduce data transmission overhead and memory usage. The dequantization parameter restores data accuracy locally and dynamically. By using different quantization and dequantization parameters, a balance between accuracy and efficiency is achieved.
[0013] In some exemplary embodiments of the present disclosure, quantization and inverse quantization are performed in the same memory area.
[0014] Quantization and dequantization operations can be performed directly in the same memory area without the need to allocate different memory spaces to store the intermediate results of quantization and dequantization respectively, thus reducing memory usage.
[0015] In some exemplary embodiments of the present disclosure, the first inference result and the second data share the same memory area.
[0016] The first inference result and the second data share the same memory area, avoiding reallocation of memory space for the merged third data, avoiding repeated storage of the same data, and reducing memory consumption.
[0017] In some exemplary embodiments of the present disclosure, calling the first interface, transmitting the first data to the first model runtime engine, and receiving the first inference result output by the first model runtime engine includes: transmitting the first data to the first model runtime engine; using the first model runtime engine to call the back-end capabilities encapsulated in the first executor to infer the first data and generate a first inference result; receiving the first inference result fed back by the first executor through the first model runtime engine, and receiving the first inference result output by the first model runtime engine.
[0018] The first executor allows the first model runtime engine to not care about the implementation details of the underlying hardware. If the underlying hardware changes, only the implementation of the first executor needs to be adjusted without modifying the code of the first model runtime engine.
[0019] In some exemplary embodiments of the present disclosure, transmitting the third data to the second model runtime engine for processing includes: transmitting the third data to the second model runtime engine; inferring the third data by calling the back-end capabilities encapsulated in the second executor through the second model runtime engine to generate inference results of multimodal data; and receiving the inference results of the multimodal data fed back by the second executor through the second model runtime engine.
[0020] The second executor frees the second model runtime engine from worrying about the implementation details of the underlying hardware. The second model runtime engine only needs to call unified interface methods. If the underlying hardware changes, only the second executor implementation needs to be adjusted, without modifying the second model runtime engine code.
[0021] In some exemplary embodiments of the present disclosure, the hardware carrier of the backend capability includes one or more of the following: CPU (Central Processing Unit), GPU (Graphics Processing Unit), and NPU (Neural-network Processing Unit).
[0022] By deploying various types of processors at the hardware platform layer, the most suitable hardware carrier can be selected for calculation according to task requirements, thereby improving data inference efficiency.
[0023] In some exemplary embodiments of the present disclosure, the method further includes: receiving image data; calling a second interface to transmit the image data to the first model runtime engine to generate an inference result of the image data; and calling a second interface to output the inference result of the image data.
[0024] The second interface is used to pass image data to the first model runtime engine. The first model runtime engine calls the first executor to perform inference tasks and generate inference results, enabling the image data to be efficiently matched to appropriate models and hardware resources, thereby improving inference efficiency and accuracy. Furthermore, the model inference system can not only perform inference on multimodal data, but also on image data, enabling inference tasks of different paradigms within the same model inference system.
[0025] In some exemplary embodiments of the present disclosure, the method further includes: receiving text data; calling a third interface to transmit the text data to a second model runtime engine to generate an inference result of the text data; and calling a third interface to output the inference result of the text data.
[0026] The third interface is responsible for passing the text data to the second model runtime engine. The second model runtime engine calls the second executor to perform the inference task and generate the inference results. This allows the text data to be efficiently matched to the appropriate model and hardware resources, thereby improving inference efficiency and accuracy. Furthermore, the model inference system can not only perform inference on multimodal data, but also on text data, enabling inference tasks of different paradigms within the same model inference system.
[0027] In some exemplary embodiments of the present disclosure, calling the first interface to transmit the first data to the first model runtime engine, and receiving the first inference result output by the first model runtime engine includes: calling the fourth interface to set the first LoRA (Low-Rank Adaptation) strength value; transmitting the first data and the first LoRA strength value to the first model runtime engine; loading the first model and the first LoRA matrix corresponding to the first model through the first model runtime engine; adjusting the first LoRA matrix using the first LoRA strength value; using the adjusted first LoRA matrix and the first model to infer the first data to generate a first inference result, and receiving the first inference result output by the first model runtime engine.
[0028] By introducing low-rank parameters on the basis of the existing model to update the model parameters, there is no need to adjust all the parameters of the entire model. Compared with full fine-tuning, the demand for computing resources is reduced.
[0029] In some exemplary embodiments of the present disclosure, transmitting the third data to the second model runtime engine for processing includes: calling the fourth interface to set the second LoRA strength value; transmitting the first data and the second LoRA strength value to the second model runtime engine; loading the second model and the second LoRA matrix corresponding to the second model through the second model runtime engine; adjusting the second LoRA matrix using the second LoRA strength value; and inferring the third data using the adjusted second LoRA matrix and the second model to generate an inference result of multimodal data.
[0030] According to another aspect of an embodiment of the present disclosure, a data processing device is provided, including: a receiving module for receiving multimodal data, wherein the multimodal data includes first data and second data; a first generating module for calling a first interface, transmitting the first data to a first model runtime engine, and receiving a first inference result output by the first model runtime engine; a data splicing module for obtaining third data based on the first inference result and the second data; a second generating module for transmitting the third data to the second model runtime engine for processing, and a first output module for outputting the inference result of the multimodal data.
[0031] In some exemplary embodiments of the present disclosure, the first inference result includes a first embedding vector, and the third data includes a third embedding vector; the data splicing module includes: a second embedding vector determination unit for obtaining a second embedding vector based on the second data; and a data splicing unit for splicing the first embedding vector and the second embedding vector to obtain a third embedding vector.
[0032] In some exemplary embodiments of the present disclosure, the second embedding vector determination unit includes: a word segmentation subunit, configured to segment the prompt word to obtain a Token sequence; and an embedding processing subunit, configured to embed the Token sequence to obtain a second embedding vector.
[0033] In some exemplary embodiments of the present disclosure, the splicing unit includes: a quantization subunit, used to quantize the first embedding vector using a first parameter; an inverse quantization subunit, used to inverse quantize the quantized first embedding vector using a second parameter; and a splicing subunit, used to splice the quantized first embedding vector and the second embedding vector to obtain a third embedding vector.
[0034] In some exemplary embodiments of the present disclosure, quantization and inverse quantization are performed in the same memory area.
[0035] In some exemplary embodiments of the present disclosure, the first inference result and the second data share the same memory area.
[0036] In some exemplary embodiments of the present disclosure, the first generation module includes: a first transmission unit for transmitting the first data to the first model runtime engine; a first inference unit for using the first model runtime engine to call the backend capabilities encapsulated in the first executor to infer the first data and generate a first inference result; and a first feedback unit for receiving the first inference result fed back by the first executor through the first model runtime engine.
[0037] In some exemplary embodiments of the present disclosure, the second generation module includes: a second transmission unit for transmitting the third data to the second model runtime engine; a second inference unit for inferring the third data by calling the back-end capabilities encapsulated in the second executor through the second model runtime engine to generate inference results of multimodal data; and a second feedback unit for receiving the inference results of the multimodal data fed back by the second executor through the second model runtime engine.
[0038] In some exemplary embodiments of the present disclosure, the hardware carrier of the backend capability includes one or more of the following: CPU, GPU, NPU.
[0039] In some exemplary embodiments of the present disclosure, the receiving module is also used to receive image data; the device also includes: a third generation module for calling the second interface to transmit the image data to the first model runtime engine to generate an inference result of the image data; and a second output module for calling the second interface to output the inference result of the image data.
[0040] In some exemplary embodiments of the present disclosure, the receiving module is also used to receive text data; the device also includes: a fourth generating module for calling the third interface to transmit the text data to the second model runtime engine to generate an inference result of the text data; and the second output module is used to call the third interface to output the inference result of the text data.
[0041] In some exemplary embodiments of the present disclosure, the first generation module is specifically used to call the fourth interface to set the first low-rank adaptive LoRA strength value; transmit the first data and the first LoRA strength value to the first model runtime engine; load the first model and the first LoRA matrix corresponding to the first model through the first model runtime engine; adjust the first LoRA matrix using the first LoRA strength value; and use the adjusted first LoRA matrix and the first model to infer the first data and generate a first inference result.
[0042] In some exemplary embodiments of the present disclosure, the second generation module is specifically used to call the fourth interface to set the second LoRA strength value; transmit the first data and the second LoRA strength value to the second model runtime engine; load the second model and the second LoRA matrix corresponding to the second model through the second model runtime engine; adjust the second LoRA matrix using the second LoRA strength value; use the adjusted second LoRA matrix and the second model to infer the third data to generate an inference result of multimodal data.
[0043] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor implements the steps of any one of the data processing methods described above.
[0044] According to another aspect of an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, which, when instructions in the storage medium are executed by a processor of a vehicle, enables an electronic device to perform the steps of any one of the data processing methods described above.
[0045] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including: a computer program or instructions, characterized in that when the computer program or instructions are executed by a processor, the steps of any one of the data processing methods described above are implemented.
[0046] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:
[0047] Within the large model framework, the first inference result generated by the first model runtime engine is used as part of the input of the second model runtime engine, so that the second model runtime engine can perform inference based on the output of the first model runtime engine and the second data, thereby improving the efficiency of the inference processing.
[0048] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0050] Figure 1 is a flowchart illustrating a data processing method according to some embodiments of the present disclosure.
[0051] Figure 2 It is a structural diagram of a model reasoning system according to some embodiments of the present disclosure.
[0052] Figure 3 It is a flowchart of an implementation process of generating third data according to some embodiments of the present disclosure.
[0053] Figure 4 is a flowchart illustrating another implementation process of generating third data according to some embodiments of the present disclosure.
[0054] Figure 5 is a flowchart illustrating another data processing method according to some embodiments of the present disclosure.
[0055] Figure 6 is a structural diagram of another model reasoning system 200 according to some embodiments of the present disclosure.
[0056] Figure 7 is a flowchart illustrating another data processing method according to some embodiments of the present disclosure.
[0057] Figure 8 2 is a structural diagram of yet another model reasoning system 200 according to some embodiments of the present disclosure.
[0058] Figure 9 is a flowchart illustrating another data processing method according to some embodiments of the present disclosure.
[0059] Figure 10 is a flowchart illustrating another data processing method according to some embodiments of the present disclosure.
[0060] Figure 11 FIG2 is a schematic diagram showing an example of a data processing method applied to a smart terminal according to an exemplary embodiment of the present disclosure.
[0061] Figure 12 FIG2 is a schematic diagram showing a data processing method applied to a vehicle according to an exemplary embodiment of the present disclosure.
[0062] Figure 13 It is a block diagram of a data processing device according to some embodiments of the present disclosure.
[0063] Figure 14 is a block diagram illustrating another data processing device according to some embodiments of the present disclosure.
[0064] Figure 15 is a block diagram illustrating yet another data processing device according to some embodiments of the present disclosure.
[0065] Figure 16 is a block diagram illustrating another data processing device according to some embodiments of the present disclosure.
[0066] Figure 17 It is a block diagram of an electronic device shown in an exemplary embodiment. DETAILED DESCRIPTION
[0067] Some embodiments of the present disclosure will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications and equivalents of the methods, devices and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but may be changed as becomes apparent after understanding the present disclosure, except for operations that must be performed in a specific order. In addition, for the sake of clarity and brevity, descriptions of features known in the art may be omitted.
[0068] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0069] Figure 1 is a flow chart showing a data processing method according to some embodiments of the present disclosure. Figure 1As shown, the data processing method provided by the present disclosure is used in a model reasoning system 200. The model reasoning system 200 can be deployed and run in a cloud server and / or end-side device. The end-side device can include various terminal devices such as vehicles and mobile phones. For example, in the field of smart cockpits, using the vehicle as an end-side device to perform various tasks, such as intelligent voice assistants and personalized services, can provide faster response speeds and a better user experience. Figure 1 The data processing method shown in includes the following steps.
[0070] In step S102 , multimodal data is received, where the multimodal data includes first data and second data.
[0071] Modality can be understood as the expression of different types of data or information. In the field of artificial intelligence, multimodal data can be understood as data in multiple forms, including images, text, audio, etc.
[0072] The multimodal data described above can be understood as data input into a model inference system that requires analysis and processing by the runtime engine and underlying hardware within the model inference system. In the field of vehicle control, receiving multimodal data can include receiving multimodal data transmitted by the vehicle.
[0073] Among them, the first data can be understood as data of a certain mode in the multimodal data. For example, the first data can be understood as other data in the multimodal data except text data. For example, the first data can be any one of audio data, image data, and video data. In this embodiment, the first data includes image data as an example for explanation.
[0074] The second data can be understood as data in other modalities in addition to the first data within the multimodal data. For example, the second data can be understood as text data within the multimodal data. Text data includes, but is not limited to, character strings, speech-to-text conversion results, descriptive text, and the like. For example, text data can include text data extracted from emails, social media posts, news articles, and navigation information.
[0075] Image data is collected by the vehicle's camera and then transmitted to the model inference system via a hardware interface. Camera acquisition devices include, but are not limited to, cameras inside the vehicle's cabin used to monitor the driver's status and cameras installed outside the vehicle for environmental awareness. Audio data is collected by the vehicle's microphone array and then transmitted to the model inference system via a hardware interface. Text information is collected by the vehicle's terminal device, either as user input or extracted from navigation, calendar, and other applications, and then transmitted to the model inference system.
[0076] In some example embodiments of the present disclosure, the application layer in the model inference system receives multi-modal data.
[0077] In step S104, the first interface is invoked to transmit the first data to the first model runtime engine and receive the first inference result output by the first model runtime engine.
[0078] The interface can be understood as an interface in the model inference system 200 for interacting with a specific model. The interface is used to transmit input data, invoke a runtime engine, and return an inference result. The interface is an API interface defined according to the invoked runtime engine. The interface is defined using any one of REST (Representational State Transfer), RPC (Remote Procedure Call), or other protocols.
[0079] The first interface is a bridge between the application layer and the multi-modal model, used to receive the first data in the multi-modal data and pass the first data to the first model runtime engine to perform an inference task.
[0080] The runtime engine can be understood as a software component responsible for loading, interpreting or compiling code during model execution, and managing its execution. The runtime engine provides a computing resource environment so that the model can run in it and perform memory allocation, garbage collection, exception handling, etc. Different programming languages and application fields have different runtime engines. The model runtime engine can be understood as a software component or service in the field of machine learning and artificial intelligence, which is used to load, interpret and execute a trained model. The functions of the model runtime engine include but are not limited to: loading a trained model file from a storage unit; calling an executor to perform inference operations according to input data; converting the inference result of the executor into a user-friendly format or a format required by other subsequent processing steps; effectively managing and allocating computing resources to ensure efficient operation of the model.
[0081] The first model runtime engine is a software component or service used to load, interpret and execute a trained traditional visual model. The above-mentioned visual model includes but is not limited to: CNN (Convolutional Neural Networks), YOLO (You Only Look Once) and the like.
[0082] The first inference result can be understood as the inference result obtained after the first model runtime engine infers the first data in the multimodal data. Exemplarily, the first inference result can be understood as the inference result obtained after the visual model runtime engine performs an inference task on the image data in the multimodal data. Based on the specific task type of the first model, the first inference result may include, but is not limited to, classification labels, detection boxes, segmentation maps, etc. Exemplarily, the first inference result includes a classification result, for example, the first inference result is "This is a cat."
[0083] The data representation form of the first reasoning result includes at least one of the following: descriptive text, embedded vector, structured information, time series, etc.
[0084] Methods for calling the first interface may include function calls, message queues, etc. Function calls can be understood as directly calling the interface function of the first interface when the first model runtime engine and the first interface are running in the same process. Message queues refer to asynchronous data transmission via message queues.
[0085] Transmitting the first data in the multimodal data to the first model runtime engine can be understood as transmitting the first data in the multimodal data to the execution environment where the first model runtime engine is located. The first model runtime engine receives the first data from the first interface and loads the first data into the computing environment.
[0086] The specific method for transmitting the first data from the first interface to the first model runtime engine depends on the architectural design of the model inference system. The transmission method includes memory sharing. Memory sharing can be understood as the first interface and the first model runtime engine being in the same process, and data can be directly transferred through memory sharing.
[0087] The first model runtime engine dynamically allocates computing resources based on task requirements and hardware resources. For example, for a vision model, the first model runtime engine calls the GPU interface to assign convolution operations to the GPU. After the GPU completes the image data inference task, it generates the first inference result, such as the category identification in an image classification task.
[0088] After the first model runtime engine receives the image data from the multimodal data, it calls the first executor to complete the inference task. The image data from the multimodal data is passed to the first executor, which loads the computation graph or instruction set of the first model and performs inference on the underlying hardware. The first executor returns the first inference result to the first model runtime engine, which then outputs the first inference result.
[0089] In step S106 , third data is obtained according to the first inference result and the second data.
[0090] In some possible implementations of the present disclosure, the first inference result and the second data are concatenated to obtain third data.
[0091] Splicing can be understood as combining two data sets in some way to form a single, integrated data set. Splicing can take different forms depending on the specific application scenario. For example, splicing the results of image recognition with related text descriptions.
[0092] Depending on different application scenarios, the methods of splicing the first inference result with the second data include but are not limited to: feature concatenation, attention mechanism splicing, multimodal embedding, etc.
[0093] Feature concatenation can be understood as directly concatenating the feature vector of the first inference result with the feature vector of the second data to form a new third feature vector. For example, if the feature vector of the first inference result is represented as [0, 1, 0, 0.9] and the feature vector of the second data is represented as [-0.2, 0.4, ..., 0.6], the third feature vector obtained by concatenating the two feature vectors is [0, 1, 0, 0.9, -0.2, 0.4, ..., 0.6].
[0094] The attention mechanism dynamically weights the importance of visual and textual features to generate new comprehensive features. For example, for the feature vector of the first inference result [0, 1, 0, 0.9] and the feature vector of the second data [-0.2, 0.4, ..., 0.6], an attention layer calculates the similarity or correlation between the first inference result and the second data, and adjusts the weight of each feature based on the calculation result to finally obtain the third feature vector.
[0095] The multimodal embedding concatenation process involves creating a multimodal embedding space in which the first inference result and the second data are mapped to vector representations of the same dimension. Using models trained on visual and textual data, the first inference result and the second data are each converted into vector representations of a specific dimension. Then, within the same embedding space, operations are performed on the feature vectors of the first inference result and the second data, such as averaging, addition, or more complex interactions, to generate a third data vector that captures information from both the first inference result and the second data.
[0096] In step S108 , the third data is transmitted to the second model runtime engine for inference.
[0097] The third data is transmitted to the second model runtime engine to generate an inference result of the multimodal data. The model inference system also includes a second model runtime engine. The second model runtime engine is used to load, interpret and execute the software components or services of the trained large language model. The above-mentioned large language models include but are not limited to: GPT (Generative Pre-trained Transformer), BERT (Bidirectional Encoder Representations from Transformers), T5 (Text-to-Text Transfer Transformer), etc.
[0098] The inference result of the multimodal data can be understood as the result of inference on the multimodal data received in step S102, or can refer to the inference result obtained after the second model runtime engine performs the inference task based on the third data.
[0099] Transmitting the third data to the second model runtime engine can be understood as transmitting the third data to the execution environment where the second model runtime engine is located, and the second model runtime engine loads the third data into the computing environment.
[0100] The specific method for transmitting the third data to the second model runtime engine 230 depends on the architectural design of the model inference system. The transmission method includes memory sharing. Memory sharing can be understood as the first model runtime engine and the second model runtime engine being in the same process, and data can be directly transferred through memory sharing.
[0101] In some exemplary embodiments of the present disclosure, the second model runtime engine invokes a second executor to perform inference tasks and generate inference results for third data. The second executor is a component used to perform inference tasks on text data. It interacts with the underlying hardware and invokes the forward propagation logic of the second model based on the third data to complete inference calculations; leverages the characteristics of the underlying hardware to accelerate the inference process; and decodes, formats, or otherwise processes the model output.
[0102] The second model runtime engine dynamically allocates computing resources based on task requirements and hardware resources. For example, for a large language model, the second model runtime engine may call the CPU interface to transfer the third data to the CPU for inference and generate inference results for the third data.
[0103] In step S110 , the inference result of the multimodal data is output.
[0104] In one possible implementation, after the second executor completes the inference task and generates the inference result of the multimodal data, the inference result of the multimodal data is transmitted to the first interface corresponding to the multimodal data through the second model runtime engine, so that the first interface corresponding to the multimodal data displays the inference result of the multimodal data to the user.
[0105] The first inference result generated by the first model runtime engine is directly used as part of the input of the second model runtime engine, so that the second model runtime engine can perform inference based on the output of the first model runtime engine and the second data, which can better capture the intrinsic connections and dependencies between different modalities, help maintain important correlations and semantic consistency between data, and improve data reasoning efficiency.
[0106] In some exemplary embodiments of the present disclosure, the data processing method provided in this embodiment is applied to the model reasoning system 200. Figure 2 is a structural diagram of a model reasoning system 200 according to some embodiments of the present disclosure, such as Figure 2 As shown, the model reasoning system 200 includes an interface layer 240, a runtime layer 250, and a hardware platform layer 280. Specifically, the interface layer 240 deploys a first interface 210 corresponding to multimodal data, the runtime layer 250 deploys a first model runtime engine 220 and a second model runtime engine 230. The hardware platform layer 280 deploys a first executor 260 and a second executor 270.
[0107] The interface layer 240 can be understood as an abstract layer located between the runtime layer 250 and the application layer. The interface layer 240 is used to define the communication method and message transmission method between the application layer and the runtime layer 250. The interface layer 240 can include multiple specific interfaces.
[0108] In one possible implementation, after the application layer receives the multimodal data, the application layer calls the first interface 210 corresponding to the multimodal data in the interface layer 240, and transmits the image data in the multimodal data to the first model runtime engine 220 of the runtime layer 250 through the first interface 210. The first model runtime engine 220 infers the image data in the multimodal data to obtain a first inference result.
[0109] In some exemplary embodiments of the present disclosure, the first model runtime engine invokes the first executor 260 to perform an image data inference task and generate a first inference result. The first executor 260 is configured to invoke the forward propagation logic of the first model based on the image data in the multimodal data to complete the inference calculation; utilize the characteristics of the underlying hardware to accelerate the inference process; and decode, format, or otherwise process the model output.
[0110] The first inference result and the second data are spliced in the same memory block to obtain third data.
[0111] When the second model runtime engine 230 receives the third data, the second executor 280 is called to complete the inference task. The third data is passed to the second executor 280, the second executor 280 loads the computation graph or instruction set of the second model, and executes the inference task of the third data on the underlying hardware. The second executor 280 returns the inference result of the third data to the second model runtime engine 230. The second model runtime engine 230 transmits the inference result of the multi-modal data to the corresponding first interface 210 of the multi-modal data, so that the corresponding first interface 210 of the multi-modal data feeds back the inference result of the multi-modal data to the application layer, and the user layer displays the inference result of the multi-modal data to the user.
[0112] As shown in FIG. 13, an implementation process flowchart for generating third data is provided for some exemplary embodiments of the present disclosure, including the following steps. Figure 3
[0113] In step S302, a second embedding vector is obtained according to the second data.
[0114] In the case where the second data includes text data, the text data is subjected to word segmentation processing to obtain a Token sequence, and the Token sequence is subjected to embedding processing to obtain a second embedding vector corresponding to the text data.
[0115] In the case where the second data includes image data, a pre-trained image model is used to extract image features of the image data, and the image features are subjected to embedding processing to obtain a second embedding vector corresponding to the image data.
[0116] In the case where the second data includes audio data, the audio signal is converted into a spectrum graph, a pre-trained audio model is used to extract audio features from the spectrum graph, and the audio features are subjected to embedding processing to obtain a second embedding vector corresponding to the audio data.
[0117] In the case where the second data includes video data, a pre-trained video model is used to extract video features from the video data, and the video features are subjected to embedding processing to obtain a second embedding vector corresponding to the video data.
[0118] In step S304, the first embedding vector and the second embedding vector are spliced to obtain a third embedding vector.
[0119] The first embedding vector can include an embedding vector of the first inference result in step S104, and the third data includes the third embedding vector.
[0120] In some exemplary embodiments of the present disclosure, after the first model runtime engine 220 performs inference on the first data, the generated first inference result is represented in the form of an embedding vector. In other words, after the first model runtime engine 220 performs inference on the first data, it outputs an embedding vector, namely, the first embedding vector.
[0121] In some exemplary embodiments of the present disclosure, after the first model runtime engine 220 infers the first data, the generated first inference result is represented in the form of a Token sequence, and the Token sequence is embedded to obtain a first embedding vector.
[0122] Concatenate the first and second embedding vectors to create a third embedding vector. This is done by concatenating the first and second embedding vectors end to end. For example, if the first embedding vector is [0.1, 0.2, 0.3] and the second embedding vector is [0.4, 0.5, 0.6], the resulting third embedding vector is [0.1, 0.2, 0.3, 0.4, 0.5, 0.6].
[0123] In this embodiment, the first inference result and the second data are concatenated in the embedding vector dimension to form a higher-dimensional vector. The embedding vector has been encoded into a high-dimensional semantic representation, thereby improving the inference efficiency.
[0124] like Figure 4 FIG. 1 is a flowchart of another implementation process for generating third data provided by some exemplary embodiments of the present disclosure. In this embodiment, the second data includes prompt words, and the method includes the following steps.
[0125] In step S402, the prompt word is segmented to obtain a Token (word unit) sequence.
[0126] The second data includes prompt words, which may include text data described in a continuous natural language. The prompt words can be understood as text instructions input by the user to guide the large language model to generate specific content.
[0127] Word segmentation can be understood as breaking a continuous natural language text into a meaningful sequence of tokens. A token sequence can be understood as the resulting list of tokens, which includes each individual word or token in the text data. For example, if the second data is "a brown puppy," the token sequence after word segmentation is: "a," "brown," "of," "puppy."
[0128] The method of segmenting the second data into words includes: segmenting the text data in the multimodal data using a natural language processing tool to obtain a Token sequence.
[0129] In step S404, the Token sequence is embedded to obtain a second embedding vector.
[0130] The embedding process can be understood as mapping a discrete Token sequence into a continuous vector space to generate a fixed-dimensional numerical vector. The second embedding vector is understood as a fixed-dimensional numerical vector obtained after the Token sequence in step S302 is embedded.
[0131] Embedding the Token sequence includes: for each Token in the Token sequence, using a pre-trained model to find an embedding vector corresponding to each Token. Then the embedding vectors of the Tokens are integrated to obtain the second embedding vector. In other words, the entire Token sequence is converted into a sequence of embedding vectors. The embedding vector sequence obtained after embedding is the "second embedding vector".
[0132] For example: the embedding vector corresponding to 'village' is: [0.1, 0.5, -0.3,]; the embedding vector corresponding to 'of' is: [0.0, 0.1, 0.2,..]; the embedding vector corresponding to 'autumn' is: [0.4, -0.2, 0.6,...].
[0133] In a multi-modal data processing task, the prompt word is used to guide the model to understand the data of other modalities. After converting the prompt word into an embedding vector, it can be effectively fused or interacted with the feature vector from other modalities, improving the inference efficiency.
[0134] In step S406, the first embedding vector is quantized using the first parameter.
[0135] Quantization can be understood as mapping a high-precision floating-point vector to a low-precision format representation, or quantization can also refer to the process of mapping a continuous range of values to a finite number of discrete values. The first parameter is used to quantize the parameter, and the first parameter can determine the conversion of the continuous embedding vector to a discrete value. The first embedding vector after quantization refers to the embedding vector obtained after the quantization step.
[0136] In one possible implementation, the first parameter is used to map the floating-point first embedding vector to the integer domain, and the low-precision quantized vector, i.e., the first embedding vector after quantization, is stored in a specified memory address.
[0137] In step S408, the first embedding vector after quantization is dequantized using the second parameter, wherein the quantization and dequantization are performed in the same memory area.
[0138] Dequantization can be understood as restoring low-precision data after quantization to high-precision floating-point representation, which is the inverse operation of quantization, and the second parameter is the parameter used for dequantization, which is different from the first parameter.
[0139] In a possible implementation, the quantized first embedded vector is converted back to a floating-point type using the second parameter to obtain a floating-point dequantized vector, that is, the dequantized first embedded vector, which is stored in a specified memory address.
[0140] Performing quantization and inverse quantization in the same memory area can be understood as allowing the quantization operation and the inverse quantization operation to be performed directly at the same memory location during the quantization and inverse quantization process, without the need to allocate additional memory space to store the intermediate results of quantization and inverse quantization respectively.
[0141] In other words, the quantized first embedded vector can directly overwrite the memory location where the first embedded vector before quantization is located. Similarly, when dequantization is performed, the dequantized first embedded vector can be directly written back to the memory location where the quantized first embedded vector is located, overwriting the quantized first embedded vector.
[0142] This avoids allocating additional memory space for intermediate results and reduces memory usage.
[0143] In step S410, the inversely quantized first embedding vector and the second embedding vector are concatenated to obtain a third embedding vector.
[0144] The inversely quantized first embedding vector and the second embedding vector are concatenated, that is, the first embedding vector and the second embedding vector are connected end to end to obtain a third embedding vector.
[0145] In this embodiment, quantization uses low-precision calculations to accelerate matrix operations, while inverse quantization restores high precision only before concatenation, potentially reducing the overall computational effort. The quantized embedded vectors reduce data transmission overhead, while inverse quantization parameter differences are dynamically recovered locally. By using different quantization and inverse quantization parameters, a balance between accuracy and efficiency is achieved.
[0146] like Figure 5 As shown, it is a flowchart of the implementation process of the data processing method provided by some other exemplary embodiments of the present disclosure, which includes the following steps.
[0147] In step S502 , multimodal data is received.
[0148] In step S504, the first interface is called to transmit the first data to the first model runtime engine.
[0149] In step S506, the first model runtime engine is used to call the backend capability encapsulated in the first executor to infer the first data and generate a first inference result, wherein the hardware carrier of the backend capability includes one or more of the following: CPU, GPU, NPU.
[0150] The functions provided by the underlying hardware are exposed to the runtime layer through encapsulated interfaces, enabling access to different types of underlying hardware in a unified manner without having to worry about the specific hardware implementation details.
[0151] In a computer system, CPU, GPU, and NPU are different types of hardware processors, each optimized for different computing tasks. In this embodiment, the hardware carriers of the backend capabilities include one or more of the following: CPU, GPU, and NPU.
[0152] Through interface encapsulation, different hardware capabilities can be accessed in a unified manner while shielding the complexity of the underlying hardware. This improves development efficiency, enhances cross-platform compatibility, and optimizes performance.
[0153] Figure 6 FIG is a structural diagram of a model reasoning system provided by some exemplary embodiments of the present disclosure. Figure 6 As shown, the above-mentioned hardware platform layer includes but is not limited to: CPU (Central Processing Unit) 610, GPU (Graphics Processing Unit) 620, QNN (Quantized Neural Network) 630, NPU (Neural-network Processing Unit) 640, etc. In addition, the hardware platform layer can also include other types of chips.
[0154] The CPU is a component of a computer that performs most computing tasks and controls various computer activities. GPUs are used for tasks requiring massively parallel processing, such as graphics rendering, scientific computing, and deep learning, and can also be used for general-purpose computing. NPUs are processors designed to accelerate the operation of machine learning algorithms, particularly neural networks.
[0155] The CPU defines a standardized set of operations for interacting with the runtime layer. Operators in the CPU are written using the instruction set architecture and a high-level programming language. Operators in the GPU are written using graphics APIs or parallel computing frameworks. Interfaces for the NPU are provided by deep learning frameworks or dedicated drivers.
[0156] An operator can be understood as the basic computational unit executed by the hardware. In deep learning, an operator can refer to a specific algorithmic operation, such as ReLU or convolutional layer. A programming language refers to the tool language used to write operators. Different hardware may be compatible with different languages. Programming languages include, but are not limited to, C / C++ and Python. The instruction set setting refers to the underlying instruction set supported by the CPU, which directly affects the types of operations that the hardware can perform.
[0157] OpenCL is an open parallel programming framework used to write cross-platform heterogeneous computing code, for example, for writing GPU-accelerated code.
[0158] The NPU API refers to the function library or toolkit provided by the NPU for developers to call NPU functions.
[0159] The first executor manages and executes tasks related to the visual model. It calls upon underlying hardware resources to complete inference or training of the visual model. It translates requests from the first model runtime engine into specific underlying operations. By encapsulating multiple abstract interfaces, the first executor hides the complexity of the underlying implementation and provides a unified external calling method. This supports dynamic switching between multiple backend hardware or software resources, such as CPUs, GPUs, NPUs, and other accelerators.
[0160] After executing the inference task, the first executor returns the inference result to the first model runtime engine 220 , and the first model runtime engine 220 receives the first inference result fed back by the first executor.
[0161] The first executor interface 610 allows the first model runtime engine 220 to not care about the implementation details of the underlying hardware. If the underlying hardware changes, only the implementation of the first executor 260 needs to be adjusted without modifying the code of the first model runtime engine 220.
[0162] In step S508 , the first inference result and the second data are concatenated to obtain third data, wherein the first inference result and the second data share the same memory area.
[0163] Sharing the same memory area means that the first inference result and the second data share the same storage area in the physical memory, rather than each occupying independent memory space.
[0164] In some exemplary embodiments of the present disclosure, memory sharing between the first inference result and the second data in the multimodal data may include: mapping the first inference result and the second data to the process address space through a file on a disk or a shared memory area, and calling the second interface 230 to directly read and write the same memory area.
[0165] In some exemplary embodiments of the present disclosure, in a multimodal data processing pipeline, a shared memory buffer is pre-allocated, and the first inference result and text data are directly written into the area.
[0166] Memory sharing between the first inference result and the second data in the multimodal data can avoid reallocating memory space for the merged third data, avoid repeated storage of the same data, and reduce memory consumption.
[0167] In step S510 , the third data is transmitted to the second model runtime engine.
[0168] In step S512, the second model runtime engine calls the backend capability encapsulated in the second executor to infer the third data to generate an inference result of the multimodal data.
[0169] The second executor shields the complexity of the underlying hardware implementation and provides a unified and standardized calling method for the runtime layer 250. For example, if the underlying hardware changes, only the implementation of the second executor needs to be adjusted without modifying the code logic of the runtime layer 250.
[0170] In the hardware platform layer 270, the executor is the component that directly interacts with the underlying hardware, responsible for executing inference tasks and generating final results. Within the entire model inference system 200, the executor is responsible for translating the computational logic of the model runtime engine into actual hardware operations. The second executor is also deployed in the hardware platform layer, but its design goals differ from those of the first executor. The second executor is specifically designed for large language models, responsible for efficiently executing inference tasks for these models and generating text, dialogue, or other natural language processing-related outputs.
[0171] The third data in the second model runtime engine 230 is transferred to the second executor 280 . After executing the inference task on the underlying hardware, the second executor 280 returns the inference result to the second model runtime engine 230 .
[0172] In step S514, the inference result of the multimodal data fed back by the second executor is received through the second model runtime engine.
[0173] After the second executor 280 executes the inference task, it returns the inference result of the third data to the second model runtime engine 230. The second model runtime engine 230 receives the inference result fed back by the second executor and sends the inference result to the interface layer 240. By calling the first interface 210 in the interface layer 240, the inference result of the multimodal data is sent to the application layer, so that the application layer can display the inference result of the multimodal data.
[0174] In step S516 , the inference result of the multimodal data is output.
[0175] The second executor interface frees the second model runtime engine 230 from the implementation details of the underlying hardware. The second model runtime engine 230 only needs to call unified interface methods without having to understand how the second executor interacts with the CPU. If the underlying hardware changes, only the second executor implementation needs to be adjusted, without modifying the second model runtime engine 230 code.
[0176] like Figure 7 , which is a flowchart of the implementation process of the data processing method provided by some exemplary embodiments of the present disclosure, includes the following steps.
[0177] In step S702 , image data is received.
[0178] Image data refers to visual information stored in digital form, including but not limited to: photos, charts, drawings, animations, video frames, etc.
[0179] After image data is collected by the camera device in the vehicle, the raw image data captured by the camera is transmitted to the model inference system 200 through a hardware interface. The camera collection device includes, but is not limited to: a camera inside the vehicle cabin for monitoring the driver's status, and a camera installed outside the vehicle for environmental perception.
[0180] In step S704, the second interface is called to transmit the image data to the first model runtime engine to generate an inference result of the image data.
[0181] Figure 8 FIG. 2 is a structural diagram of another model inference system 200 according to some embodiments of the present disclosure. Figure 8 As shown, the model inference system 200 also includes a second interface 810 corresponding to image data. The second interface 810 corresponding to image data is set in the interface layer 240 and is an interface for interacting with the first model runtime engine 220. The second interface 810 corresponding to image data is an interface of the visual model runtime engine.
[0182] The second interface 810 corresponding to the image data may be called by a function call, etc. Function call can be understood as the first model runtime engine 220 and the third interface corresponding to the image data running in the same process, and then directly calling the third interface corresponding to the image data.
[0183] Sending the image data to the first model runtime engine 220 can be understood as sending the image data to the execution environment where the first model runtime engine 220 is located. The first model runtime engine 220 receives the image data from the second interface 810 corresponding to the image data and loads the image data into the computing environment.
[0184] The specific method for transmitting image data from the second interface 810 corresponding to the image data to the first model runtime engine 220 depends on the system architecture design. In this embodiment, image data is transmitted from the second interface corresponding to the image data to the first model runtime engine 220 using a shared memory transmission method. Memory sharing means that the third interface corresponding to the image data and the first model runtime engine 220 are in the same process, and data can be directly transferred through shared memory.
[0185] In some exemplary embodiments of the present disclosure, the first model runtime engine 220 is used to call the first executor 260 to perform the inference task and generate the inference result of the image data.
[0186] After the image data is passed to the first executor 260 through the first model runtime engine 220, the first executor 260 will further process it, including: performing necessary conversions on the image data; calling the inference logic in the first model to complete the forward propagation calculation; decoding or formatting the output results of the first model to generate user-understandable inference results.
[0187] For example, for a visual model, convolution operations are assigned to the GPU. After the GPU completes the inference task, it generates inference results for the image data, such as category labels in an image classification task.
[0188] In step S706, the second interface is called to output the inference result of the image data.
[0189] In one possible implementation, after the first executor 260 executes the inference task, it generates an inference result of the image data, and transmits the inference result of the image data to the second interface 810 corresponding to the image data through the first model runtime engine 220, so that the second interface 810 corresponding to the image data transmits the inference result of the image data to the application layer, and then the application layer displays the inference result of the image data to the user.
[0190] In this embodiment, the second interface corresponding to the image data is called, and the second interface corresponding to the image data is used to pass the image data to the first model runtime engine 220. The first model runtime engine 220 calls the first executor to perform the reasoning task and generate the reasoning result, so that the image data can be efficiently matched to the appropriate model and hardware resources, thereby improving the reasoning efficiency and accuracy.
[0191] In addition, the model reasoning system 200 can not only realize the reasoning of multimodal data, but also realize the reasoning of image data, thereby realizing the reasoning tasks of different paradigms in the same model reasoning system 200.
[0192] like Figure 9As shown in FIG, it is a flowchart of the implementation process of the data processing method provided by some exemplary embodiments of the present disclosure, which includes the following steps.
[0193] In step S902 , text data is received.
[0194] Text data can come from a variety of sources, such as speech recognition results, text input from the user via the touchscreen, or log information from the vehicle system. For example, if a user says "lower the air conditioning temperature," the speech signal is converted into text data and then transmitted to the model inference system 200. Text data can be uploaded directly to the model inference system 200 through the API or transmitted as a file or streaming data.
[0195] In step S904, the third interface is called to transmit the text data to the second model runtime engine to generate an inference result of the text data.
[0196] like Figure 8 As shown, the model inference system 200 also includes a third interface 820 corresponding to text data. The third interface 820 corresponding to text data is provided in the interface layer 240 and is an interface for interacting with the second model runtime engine 230. The third interface 820 corresponding to text data is an interface for the large language model runtime engine.
[0197] The third interface 820 corresponding to text data serves as a bridge between the application layer and the second model runtime engine 230, and is used to receive text data and pass the text data to the second model runtime engine 230. The third interface 820 corresponding to text data is also used to perform preliminary formatting or preprocessing on the text data to ensure that the text data meets the requirements of the second model. The second model includes a large language model.
[0198] The specific method for transmitting text data from the third interface 820 corresponding to the text data to the second model runtime engine 230 depends on the system architecture design. In this embodiment, a memory-sharing transmission method is used to transmit image data from the third interface 820 corresponding to the text data to the second model runtime engine 230. Memory sharing means that the third interface 820 corresponding to the text data and the second model runtime engine 23024 are in the same process, and data can be directly transferred through memory sharing.
[0199] After receiving the text data, the second model runtime engine 230 calls the second executor to complete the inference task. This includes passing the text data to the second executor 280, which then loads the computation graph or instruction set of the second model and performs inference on the underlying hardware. The second executor 280 then returns the inference result of the text data to the second model runtime engine 230.
[0200] In a specific application scenario, the voice assistant system in the car collects the voice command issued by the user (for example, the user says "play music"), converts the voice command into text data, and transmits the text data to the model inference system 200. The model inference system 200 calls the third interface 820 corresponding to the text data to pass the text data to the large language model runtime engine. The large language runtime engine calls the executor of the large language model to load the large language model and executes Transformer reasoning on the TPU to generate the reasoning result. The result is returned to the third interface 820 corresponding to the text data and triggers the corresponding operation (such as turning on the car audio and playing music).
[0201] In some exemplary embodiments of the present disclosure, the second model runtime engine 230 is used to call the second executor interface, transmit the input data to the second executor 280, and the second executor 280 performs the reasoning task based on the text data to generate the reasoning result of the text data.
[0202] In step S906, the third interface is called to output the inference result of the text data.
[0203] In one possible implementation, after the second executor 280 executes the reasoning task, it generates an inference result of the text data, and transmits the inference result of the text data to the third interface 820 corresponding to the text data through the second model runtime engine 230, so that the third interface 820 corresponding to the text data transmits the inference result of the text data to the application layer, and then the application layer displays the inference result of the text data to the user.
[0204] In this embodiment, the fourth interface corresponding to the text data is called, and the fourth interface corresponding to the text data is responsible for passing the text data to the second model runtime engine 230. The second model runtime engine 230 calls the second executor to perform the reasoning task and generate the reasoning result, so that the text data can be efficiently matched to the appropriate model and hardware resources, thereby improving the reasoning efficiency and accuracy.
[0205] In addition, the model reasoning system 200 can not only realize the reasoning of multimodal data, but also realize the reasoning of text data, thereby realizing the reasoning tasks of different paradigms in the same model reasoning system 200.
[0206] like Figure 10 As shown in FIG, it is a flowchart of the implementation process of the data processing method provided by some exemplary embodiments of the present disclosure, which includes the following steps.
[0207] In step S1002 , multimodal data is received.
[0208] In step S1004, the fourth interface is called to set the first low-rank adaptive LoRA strength value.
[0209] LoRa is a technique used to adjust the weights of deep learning models by introducing a low-rank matrix to fine-tune the model without retraining the entire model. The LoRa strength value is a scalar parameter that controls the degree of adjustment to the original model weights. A larger LoRa strength value indicates a stronger influence of the LoRa matrix on the model; a smaller LoRA strength value indicates a weaker influence. The first LoRA strength value can be understood as the LoRA strength value used to adjust the first model.
[0210] like Figure 8 As shown, the model inference system 200 also includes a fourth interface. The fourth interface 830 is used to set the LoRA strength value of the user data to prepare for the subsequent model fine-tuning process.
[0211] In step S1006, the first interface is called to transmit the first data and the first LoRA intensity value to the first model runtime engine, and the first model and the first LoRA matrix corresponding to the first model are loaded through the first model runtime engine.
[0212] First data refers to the actual data input into the model and can be any form of data, such as text or images. The first model runtime engine is capable of loading and running deep learning models and processing attached LoRa matrices. The first LoRa matrix corresponding to the first model refers to the LoRa matrix that is bound to and adapted to the first model.
[0213] The first interface is called to establish a data channel, and the first data and the first LoRA strength value obtained are transmitted to the first model runtime engine. The first model runtime engine simultaneously loads the first model and its corresponding first LoRA matrix.
[0214] In step S1008, the first LoRA matrix is adjusted using the first LoRA strength value; and the first data is inferred using the adjusted first LoRA matrix and the first model to generate a first inference result.
[0215] The first LoRA matrix is adjusted using the first LoRA strength value, including: multiplying the first LoRA matrix by the first LoRA strength value to control the influence of the first LoRA matrix on the first model, combining the adjusted first LoRA matrix with the weight matrix of the first model to form a new first model, and using the new first model to infer the first data to generate a first inference result.
[0216] In step S1010 , a first inference result output by a first model runtime engine is received.
[0217] By introducing the first low-rank parameter based on the first model to update the model parameters, there is no need to adjust all the parameters of the first model. Compared with full fine-tuning, the demand for computing resources is reduced.
[0218] In step S1012 , third data is obtained according to the first inference result and the second data.
[0219] In step S1014, the fourth interface is called to set a second LoRA strength value.
[0220] The second LoRA strength value may be understood as a LoRA strength value used to adjust the second model.
[0221] In step S1016, the first data and the second LoRA intensity value are transmitted to a second model runtime engine, and the second model and a second LoRA matrix corresponding to the second model are loaded by the second model runtime engine.
[0222] The third data refers to the actual data input into the model, which can be any form of data such as text or images. The second model runtime engine is capable of loading and running the deep learning model and is capable of processing the attached LoRa matrix. The second LoRa matrix corresponding to the second model refers to the LoRa matrix that is bound to the second model and adapted to the second model.
[0223] The third data and the obtained second LoRA strength value are transmitted to the second model runtime engine, and the second model runtime engine simultaneously loads the second model and its corresponding second LoRA matrix.
[0224] The second LoRA matrix is adjusted using the second LoRA strength value; and the third data is inferred using the adjusted second LoRA matrix and the second model to generate an inference result of multimodal data.
[0225] The second LoRA matrix is adjusted using the second LoRA strength value, including: multiplying the second LoRA matrix by the second LoRA strength value, thereby controlling the influence of the second LoRA matrix on the first model, combining the adjusted second LoRA matrix with the weight matrix of the first model to form a new second model, and using the new second model to infer the third data to generate an inference result of the multimodal data.
[0226] In step S1018 , the inference result of the multimodal data is output.
[0227] By introducing a second low-rank parameter based on the second model to update the model parameters, there is no need to adjust all the parameters of the second model. Compared with full fine-tuning, the demand for computing resources is reduced.
[0228] In an exemplary application scenario, Figure 11 FIG. 1 is a schematic diagram showing a data processing method applied to a smart terminal according to an exemplary embodiment of the present disclosure. Figure 11 As shown, a dialogue page is displayed on the terminal interface. In response to user input into a text box, the dialogue page obtains multimodal data input by the user. The multimodal data includes image data 1110 and text data 1120. For example, the image data may be the cover of a book, and the text data may be some text information of the book, such as the title, author name, or a brief description of the book's contents. After receiving the multimodal data input by the user, the multimodal data is transmitted as input to the model inference system 200. The data inference method provided in this embodiment is executed to obtain inference results for the input data and display them on the dialogue page. The inference results may be a personalized recommendation list. For example, the inference results "Book Title 1, Author 1, Introduction 1; Book Title 2, Author 2, Introduction 2, etc." are displayed in text box 1130 on the dialogue page. This embodiment is merely an illustrative example of the application scenario of the data inference method, and is not intended to be limiting.
[0229] In an exemplary application scenario, Figure 12 FIG. 1 is a schematic diagram showing a data processing method applied to a vehicle according to an exemplary embodiment of the present disclosure. Figure 12 As shown, the camera 1210 installed in front of the vehicle continuously captures the road scene in front of the vehicle, including image data such as road signs, traffic light status, lane lines, and surrounding vehicles. The text information provided by the map navigation application in the vehicle terminal 1220 includes text information such as the destination address, recommended driving route, and estimated arrival time. The image data and text data are input into the model inference system 200 together, and the data inference method provided by this embodiment is executed to obtain the inference result of the input data. The inference result can be a voice prompt information. For example, when approaching a complex intersection, the microphone 1230 in the vehicle outputs a voice prompt of "Turn left in 500 meters ahead", and marks the specific turning point in the captured real-time video stream. The map navigation application in the vehicle terminal 1220 highlights the relevant road signs or traffic lights, helping the driver to make correct driving decisions more easily.
[0230] Figure 13 FIG1 is a block diagram of a data processing device according to some embodiments of the present disclosure. Figure 13 The data processing device includes: a receiving module 1310, a first generating module 1320, a data splicing module 1330, a second generating module 1340 and a first output module 1350.
[0231] Among them, the receiving module 1310 is used to receive multimodal data, wherein the multimodal data includes first data and second data; the first generation module 1320 is used to call the first interface, transmit the first data to the first model runtime engine, and receive the first inference result output by the first model runtime engine; the data splicing module 1330 is used to obtain third data based on the first inference result and the second data; the second generation module 1340 is used to transmit the third data to the second model runtime engine for processing, and the first output module 1350 is used to output the inference result of the multimodal data.
[0232] In some exemplary embodiments of the present disclosure, the first inference result includes a first embedding vector, and the third data includes a third embedding vector; Figure 14 As shown, the data splicing module 1330 includes: a second embedding vector determining unit 1410 for obtaining a second embedding vector according to the second data; and a data splicing unit 1420 for splicing the first embedding vector and the second embedding vector to obtain the third embedding vector.
[0233] In some exemplary embodiments of the present disclosure, the second embedding vector determining unit 1410 includes: a word segmentation subunit 1430 for segmenting the prompt word to obtain a token sequence; and an embedding processing subunit 1440 for embedding the token sequence to obtain a second embedding vector.
[0234] In some exemplary embodiments of the present disclosure, the data splicing unit 1420 includes: a quantization subunit 1450 for quantizing the first embedded vector using a first parameter; an inverse quantization subunit 1460 for inverse quantizing the quantized first embedded vector using a second parameter; and a splicing subunit 1470 for splicing the quantized first embedded vector and the second embedded vector to obtain a third embedded vector.
[0235] In some exemplary embodiments of the present disclosure, quantization and inverse quantization are performed in the same memory area.
[0236] In some exemplary embodiments of the present disclosure, the first inference result and the second data share the same memory area.
[0237] In some exemplary embodiments of the present disclosure, Figure 15 As shown, the first generation module 1320 includes: a first transmission unit 1510 for transmitting the first data to the first model runtime engine; a first reasoning unit 1520 for using the first model runtime engine to call the back-end capabilities encapsulated in the first executor to infer the first data and generate a first reasoning result; a first feedback unit 1530 for receiving the first reasoning result fed back by the first executor through the first model runtime engine, and receiving the first reasoning result output by the first model runtime engine.
[0238] In some exemplary embodiments of the present disclosure, Figure 15 As shown, the second generation module 1340 includes: a second transmission unit 1540 for transmitting the third data to the second model runtime engine; a second inference unit 1550 for inferring the third data by calling the back-end capabilities encapsulated in the second executor through the second model runtime engine to generate an inference result of the multimodal data; and a second feedback unit 1560 for receiving the inference result of the multimodal data fed back by the second executor through the second model runtime engine.
[0239] In some exemplary embodiments of the present disclosure, the hardware carrier of the backend capability includes one or more of the following: CPU, GPU, NPU.
[0240] In some exemplary embodiments of the present disclosure, the receiving module 1320 is further configured to receive image data; Figure 16 As shown, the device also includes: a third generation module 1610 for calling the second interface to transmit the image data to the first model runtime engine to generate an inference result of the image data; a second output module 1620 for calling the second interface to output the inference result of the image data.
[0241] In some exemplary embodiments of the present disclosure, the receiving module 1320 is further configured to receive text data; Figure 16 As shown, the device also includes: a fourth generation module 1630 for calling the third interface to transmit text data to the second model runtime engine to generate an inference result of the text data; a third output module 1640 for calling the third interface to output the inference result of the text data.
[0242] In some exemplary embodiments of the present disclosure, the first generation module 1220 is specifically used to call the fourth interface to set the first low-rank adaptive LoRA strength value; transmit the first data and the first LoRA strength value to the first model runtime engine; load the first model and the first LoRA matrix corresponding to the first model through the first model runtime engine; adjust the first LoRA matrix using the first LoRA strength value; use the adjusted first LoRA matrix and the first model to infer the first data to generate a first inference result, and receive the first inference result output by the first model runtime engine.
[0243] In some exemplary embodiments of the present disclosure, the second generation module 1240 is specifically used to call the fourth interface to set the second LoRA strength value; transmit the first data and the second LoRA strength value to the second model runtime engine; load the second model and the second LoRA matrix corresponding to the second model through the second model runtime engine; adjust the second LoRA matrix using the second LoRA strength value; use the adjusted second LoRA matrix and the second model to infer the third data to generate an inference result of multimodal data.
[0244] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0245] Figure 17 1 is a block diagram illustrating a data processing electronic device 1700 according to some embodiments of the present disclosure. For example, the electronic device 1700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0246] Reference Figure 17 , the electronic device 1700 may include one or more of the following components: a processing component 1702 , a memory 1704 , a power component 1706 , a multimedia component 1708 , an audio component 1710 , an input / output (I / O) interface 1712 , a sensor component 1714 , and a communication component 1716 .
[0247] Processing component 1702 controls the overall operation of electronic device 1700, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. Processing component 1702 may include one or more processors 1720 to execute instructions to perform all or part of the steps of the above-described method. In addition, processing component 1702 may include one or more modules to facilitate interaction between processing component 1702 and other components. For example, processing component 1702 may include a multimedia module to facilitate interaction between multimedia component 1708 and processing component 1702.
[0248] The memory 1704 is configured to store various types of data to support the operations of the device 1700. Examples of such data include instructions for any application or method operating on the electronic device 1700, contact data, phone book data, messages, pictures, videos, etc. The memory 1704 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0249] The power supply component 1706 provides power to the various components of the electronic device 1700. The power supply component 1706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 1700.
[0250] The multimedia component 1708 includes a screen that provides an output interface between the device 1700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 1708 includes a front camera and / or a rear camera. When the electronic device 1700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0251] The audio component 1710 is configured to output and / or input audio signals. For example, the audio component 1710 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 1700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 1704 or transmitted via the communication component 1716. In some embodiments, the audio component 1710 also includes a speaker for outputting audio signals.
[0252] I / O interface 1712 provides an interface between processing component 1702 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a start button, and a lock button.
[0253] Sensor assembly 1714 includes one or more sensors for providing various aspects of the status assessment of electronic device 1700. For example, sensor assembly 1714 can detect the open / closed state of electronic device 1700, the relative positioning of components, such as the display and keypad of device 1700. Sensor assembly 1714 can also detect changes in the position of electronic device 1700 or a component of electronic device 1700, the presence or absence of user contact with electronic device 1700, the orientation or acceleration / deceleration of electronic device 1700, and changes in the temperature of electronic device 1700. Sensor assembly 1714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1714 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1714 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0254] The communication component 1716 is configured to facilitate wired or wireless communication between the electronic device 1700 and other devices. The electronic device 1700 can access a wireless network based on a communication standard, such as WiFi, 3G, 4G, 5G, other communication standards, or a combination thereof. In some embodiments of the present disclosure, the communication component 1716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of the present disclosure, the communication component 1716 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0255] In some embodiments of the present disclosure, the electronic device 1700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0256] In some embodiments of the present disclosure, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory 1704 including instructions, and the instructions can be executed by a processor 1720 of an electronic device 1700 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0257] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of a mobile terminal, enables the mobile terminal to perform a data processing method, the method comprising: receiving multimodal data, wherein the multimodal data includes first data and second data; calling a first interface to transmit the first data to a first model runtime engine, and receiving a first inference result output by the first model runtime engine; obtaining third data based on the first inference result and the second data; transmitting the third data to a second model runtime engine for processing; and outputting the inference result of the multimodal data.
[0258] Based on the same inventive concept, an embodiment of the present disclosure further provides a vehicle, which includes the electronic device as in the above embodiment.
[0259] Based on the same inventive concept, the present disclosure also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the data processing method of any one of the above-mentioned method embodiments. Since the principles for solving the problems in this computer program product embodiment are similar to those in the above-mentioned method embodiment, the implementation of this computer program product embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated here.
[0260] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0261] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0262] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0263] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
[0264] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.
Claims
1. A data processing method, characterized in that: include: receiving multimodal data, wherein the multimodal data includes first data and second data; Calling a first interface to transmit the first data to a first model runtime engine, and receiving a first inference result output by the first model runtime engine; Obtain third data based on the first inference result and the second data; transmitting the third data to a second model runtime engine for processing; Outputting the inference result of the multimodal data.
2. The data processing method according to claim 1, wherein: The first inference result includes a first embedding vector, and the third data includes a third embedding vector; Obtaining third data according to the first inference result and the second data includes: Obtaining a second embedding vector according to the second data; The first embedding vector and the second embedding vector are concatenated to obtain the third embedding vector.
3. The data processing method according to claim 3, characterized in that: The second data includes a prompt word; Obtaining a second embedding vector according to the second data includes: Segment the prompt word to obtain a word-unit Token sequence; The Token sequence is embedded to obtain the second embedding vector.
4. The data processing method according to claim 2, wherein: The concatenating the first embedding vector and the second embedding vector to obtain the third embedding vector includes: quantizing the first embedding vector using a first parameter; Dequantizing the quantized first embedding vector using a second parameter; The inversely quantized first embedding vector and the second embedding vector are concatenated to obtain the third embedding vector.
5. The data processing method according to claim 4, characterized in that: The quantization and the inverse quantization are performed in the same memory area.
6. The data processing method according to any one of claim 1, characterized in that: The first inference result and the second data share the same memory area.
7. The data processing method according to claim 1, wherein: The calling of the first interface, transmitting the first data to the first model runtime engine, and receiving the first inference result output by the first model runtime engine includes: transmitting the first data to the first model runtime engine; Using the first model runtime engine to call the backend capability encapsulated in the first executor to infer the first data and generate a first inference result; receiving, through the first model runtime engine, a first inference result fed back by the first executor; Receive a first inference result output by the first model runtime engine.
8. The data processing method according to claim 1, wherein: The transmitting the third data to the second model runtime engine for processing comprises: transmitting the third data to the second model runtime engine; Inferring the third data by the second model runtime engine through the backend capability encapsulated in the second executor to generate an inference result of the multimodal data; The inference result of the multimodal data fed back by the second executor is received through the second model runtime engine.
9. The data processing method according to claim 7 or 8, characterized in that: The hardware carriers of the backend capabilities include one or more of the following: central processing unit CPU, graphics processing unit GPU, neural network processor NPU.
10. The data processing method according to claim 1, wherein: Also includes: receiving image data; Calling a second interface to transmit the image data to the first model runtime engine to generate an inference result of the image data; The second interface is called to output the inference result of the image data.
11. The data processing method according to claim 8, characterized in that: Also includes: Receive text data; Calling a third interface to transmit the text data to the second model runtime engine to generate an inference result of the text data; The third interface is called to output the inference result of the text data.
12. The data processing method according to claim 1, wherein: The calling of the first interface, transmitting the first data to the first model runtime engine, and receiving the first inference result output by the first model runtime engine includes: Call the fourth interface to set the first low-rank adaptation LoRA strength value; transmitting the first data and the first LoRA strength value to the first model runtime engine; Loading the first model and the first LoRA matrix corresponding to the first model through the first model runtime engine; adjusting the first LoRA matrix using the first LoRA strength value; Inferring the first data using the adjusted first LoRA matrix and the first model to generate a first inference result; Receive a first inference result output by a first model runtime engine.
13. The data processing method according to claim 1, wherein: The transmitting the third data to the second model runtime engine for processing comprises: Calling the fourth interface to set the second LoRA strength value; transmitting the first data and the second LoRA strength value to the second model runtime engine; Loading the second model and the second LoRA matrix corresponding to the second model through the second model runtime engine; adjusting the second LoRA matrix using the second LoRA strength value; The third data is inferred using the adjusted second LoRA matrix and the second model to generate an inference result of multimodal data.
14. A data processing device, characterized in that: include: The receiving module is used to receive multimodal data, wherein the multimodal data includes first data and second data; The first generating module is used to call the first interface, transmit the first data to the first model runtime engine, and receive the first inference result output by the first model runtime engine; A data splicing module is used to obtain third data based on the first inference result and the second data; The second generating module is used to transmit the third data to the second model runtime engine for processing; The first output module is used to output the inference result of the multimodal data.
15. The data processing device according to claim 14, characterized in that The first inference result includes a first embedding vector, and the third data includes a third embedding vector; The data splicing module includes: a second embedding vector determining unit for obtaining a second embedding vector according to the second data; and a data splicing unit for splicing the first embedding vector and the second embedding vector to obtain the third embedding vector.
16. The data processing device according to claim 15, characterized in that The second data includes a prompt word; The second embedding vector determination unit includes: a word segmentation subunit for segmenting the prompt word to obtain a word unit Token sequence; The embedding processing subunit is used to perform embedding processing on the Token sequence to obtain the second embedding vector.
17. The data processing device according to claim 15, characterized in that The data splicing unit includes: a quantization subunit for quantizing the first embedded vector using a first parameter; a dequantization subunit for dequantizing the quantized first embedded vector using a second parameter; and a splicing subunit for splicing the dequantized first embedded vector and the second embedded vector to obtain a third embedded vector.
18. The data processing device according to claim 14, characterized in that The quantization and the inverse quantization are performed in the same memory area.
19. The data processing device according to any one of claims 14 to 18, characterized in that: The first inference result and the second data share the same memory area.
20. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the steps of the data processing method according to any one of claims 1 to 13.
21. A non-transitory computer-readable storage medium, which, when instructions in the storage medium are executed by an electronic processor, enables the electronic device to perform the steps of the data processing method according to any one of claims 1 to 13.
22. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, the steps of the data processing method according to any one of claims 1 to 13 are implemented.