Data processing method and apparatus, and device and storage medium
Patent Information
- Application Number
- PCT/CN2026/084138
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-18
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026084138_01102026_PF_FP_ABST
Abstract
Description
A data processing method, apparatus, device, and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202510364702.6, filed with the State Intellectual Property Office of China on March 25, 2025, entitled "A Data Processing Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, specifically to a data processing method, apparatus, device, and storage medium. Background Technology
[0003] With the rapid development of artificial intelligence, large language models (LLMs) are widely used in natural language processing tasks, including question answering, summary generation, sentiment analysis, and other applications.
[0004] Large vision-language model (LVLM) is a model that combines vision and language processing capabilities. LVLM improves the processing performance of multimodal tasks by integrating LLM. However, the introduction of multimodal data such as images, videos, and audio data exacerbates the length of encoded sequences and data redundancy. Current sequence compression schemes suffer from high computational cost and high GPU memory usage, resulting in high inference overhead for large language models. Summary of the Invention
[0005] This application provides a data processing method, apparatus, device, and storage medium to achieve compression of multimodal sequences and reduce the inference overhead of large language models.
[0006] To achieve the above objectives, this application provides the following technical solution:
[0007] The first aspect of this application provides a data processing method applied to a multimodal reasoning system, the multimodal reasoning system including a preprocessing model and a large language model (LLM), the method including: acquiring a multimodal sequence; acquiring a target sequence of the multimodal sequence based on the preprocessing model, the preprocessing model being constructed based on the first N layers of the large language model (LLM), where N is greater than 1; and acquiring the reasoning result of the target sequence based on the large language model (LLM).
[0008] In this embodiment, the specific type of multimodal sequence is not limited. The multimodal sequence may include at least one of text modal sequences, image modal sequences, audio modal sequences, and video modal sequences. It is understood that the method of obtaining the multimodal sequence is not limited in this embodiment. For example, the multimodal sequence can be obtained by a multimodal encoder.
[0009] The multimodal reasoning system provided in this application embodiment can be applied to various multimodal understanding models, such as: Tongyi Qianwen Visual Language Multimodal Large Model Qwen2-VL, Shusheng Wanxiang Multimodal Large Model InternVL2, Large Language and Vision Assistant (LLaVA), etc., or general client models, and is not limited in this application embodiment.
[0010] It should be noted that the preprocessing model is built based on the first N layers of an LLM network, and the specific value of N is not limited in this embodiment. Optionally, N can be obtained through an automatic optimization algorithm based on a specific inference model and test dataset. Optionally, N can be selected from 2 to 28 layers. This embodiment does not impose any limitations.
[0011] This application provides a two-stage multimodal sequence inference scheme. In the first stage, a preprocessing model is used to compress the multimodal sequence. Since the preprocessing model is built based on the first N layers of a large language model, it can understand and process the multimodal sequence. Furthermore, the preprocessing model only needs to run once before the prefill stage of the LLM, eliminating the need for a key-value cache (KV cache) and effectively reducing computational load and GPU memory usage. In the second stage, the LLM can directly perform inference based on the target sequence. The LLM does not need to compress the sequence, reducing the impact of changes in sequence length within layers and lowering the inference overhead of the large language model.
[0012] In one possible implementation of the first aspect of this application, obtaining the target sequence of the multimodal sequence based on the preprocessing model includes: obtaining the attention score of each sequence unit in the multimodal sequence, the attention score including intramodal self-attention score and intermodal mutual attention score; determining the target sequence based on the intramodal self-attention score and the intermodal mutual attention score, wherein the number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence.
[0013] In this embodiment, the preprocessing model can compress the multimodal sequence based on the attention score of each sequence unit (token) in the multimodal sequence. The attention score of each sequence unit considers both its self-attention score within the modality and its mutual attention score between modalities. That is, when compressing the multimodal sequence, this embodiment not only considers the information density of the modality itself but also focuses on the information fusion between modalities, ensuring the reliability of multimodal sequence selection and compression, thereby guaranteeing the reliability of LLM inference.
[0014] In addition, in the embodiments of this application, the number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence. It can be seen that the sequence length of the target sequence obtained after compression is less than the sequence length of the multimodal sequence, which effectively reduces the sequence length while ensuring the model inference performance and reducing the inference overhead of large language models.
[0015] It should be noted that the calculation methods for the intramodal self-attention score and the intermodal mutual attention score of the sequence unit in this embodiment can refer to the existing attention score calculation methods, and will not be repeated in this embodiment.
[0016] In one possible implementation of the first aspect of this application, the compression ratio of the preprocessing model is R, where R is less than 1; determining the target sequence based on the intramodal self-attention score and the intermodal mutual attention score includes: determining the target sequence from the multimodal sequence according to the compression ratio based on the intramodal self-attention score and the intermodal mutual attention score.
[0017] In this embodiment, the preprocessing model is further configured with a compression ratio parameter R. By adjusting the compression ratio, the degree of compression of the multimodal sequence can be controlled, thereby controlling the sequence length of the target sequence. For example, if the multimodal sequence includes 100 sequence units and the compression ratio R is 50%, then the compressed target sequence will include 50 sequence units.
[0018] It should be noted that the compression ratio of the preprocessed model is R, and the specific value of R is not limited in this embodiment. Optionally, R can be obtained through an automatic optimization algorithm based on a specific model and test dataset. Optionally, R can be set to 50% to 90%. This embodiment does not impose any limitation.
[0019] In one possible implementation of the first aspect of this application, determining the target sequence from the multimodal sequence according to the compression ratio based on the intramodal self-attention score and the intermodal mutual attention score includes: determining the sequence to be compressed based on the intersection of the top K sequence units with the highest intramodal self-attention scores and the top K sequence units with the highest intermodal mutual attention scores; and determining the target sequence from the sequence to be compressed according to the compression ratio.
[0020] In this embodiment of the application, the intramodal self-attention score A of each sequence unit is first obtained. S And the intermodal mutual attention score A for each sequence unit. C Next, take the top K intramodal self-attention scores (TopK(A)) for each modality. S The TopK(A) intermodal mutual attention scores are used to determine the top K most important modalities. C TopK(A) intersection S )∧TopK(A C The sequence to be compressed is denoted as ; finally, the sequence to be compressed is compressed according to the compression ratio R to obtain the final target sequence. Therefore, this application provides a specific compression step for multimodal sequences. During the sequence compression process, both the self-attention score of sequence units within a mode and the mutual attention score of sequence units between modes are considered, ensuring the reliability of multimodal sequence selection and compression, and thus ensuring the reliability of LLM inference.
[0021] In one possible implementation of the first aspect of this application, the multimodal sequence includes a first modal sequence and a second modal sequence, wherein the first modal sequence includes a text modal sequence, and the second modal sequence includes at least one of an image modal sequence, an audio modal sequence, and a video modal sequence.
[0022] In this embodiment, the multimodal sequence can be further divided into two types of modal sequences: text modal sequences and audiovisual modal sequences, such as image modal sequences, audio modal sequences, and video modal sequences. Therefore, when calculating the intermodal mutual attention score of each sequence unit, it can be performed according to the two modal sequences. Compared to further dividing the audiovisual modal sequence into multiple types of modal sequences for intermodal mutual attention score calculation, this reduces the computational load of attention scores and the overhead of sequence compression.
[0023] In one possible implementation of the first aspect of this application, obtaining the inference result of the target sequence based on the Large Language Model (LLM) includes: inputting the sequence units in the target sequence and the sequence units in the context window of the preprocessing model into the Large Language Model (LLM) to obtain the inference result.
[0024] The context window, also known as the observation window, refers to the number of sequence unit tokens that a large language model can receive and consider when processing sequences. This window can be measured by a certain number of sequence unit tokens. In this embodiment, no limit is made on the number of sequence units contained in the context window. For example, the context window may include 50 sequence unit tokens.
[0025] In this embodiment, the sequence units in the target sequence obtained by compression and the sequence units in the context window of the large language model are input into the large language model for model inference to obtain the model inference result.
[0026] In one possible implementation of the first aspect of this application, the data processing method further includes: obtaining model configuration parameters, the model configuration parameters including the number of layers N and compression ratio R of the preprocessed model, the number of layers N and compression ratio R being determined based on an optimization algorithm; and constructing a preprocessed model based on the first N layers of the Large Language Model (LLM) and the compression ratio R.
[0027] In this embodiment, relevant model configuration parameters for configuring the preprocessing model can be obtained. The model configuration parameters include the number of layers N and the compression ratio R of the preprocessing model. These two parameters can be used to construct a preprocessing model that meets the requirements of multimodal sequence compression. The search space for the number of layers N and the compression ratio R of the preprocessing model can be defined based on a specific large language model and test dataset, and obtained by fitting through an automatic optimization algorithm to meet the user's model inference accuracy requirements.
[0028] It should be noted that the embodiments of this application do not limit the automatic optimization algorithm. For example, Bayesian optimization (BO), random search (RS), grid search (GS), etc. can be used.
[0029] Secondly, embodiments of this application also provide a data processing apparatus, which is applied to a multimodal reasoning system, the multimodal reasoning system including a preprocessing model and a Large Language Model (LLM), the apparatus comprising:
[0030] The sequence acquisition module is used to acquire multimodal sequences;
[0031] The sequence processing module is used to obtain the target sequence of the multimodal sequence based on the preprocessing model, wherein the preprocessing model is constructed based on the first N layers of the large language model LLM, and N is greater than 1;
[0032] The model inference module is used to obtain the inference results of the target sequence based on the large language model (LLM).
[0033] This application provides a data processing apparatus with a two-stage multimodal sequence inference scheme. In the first stage, a preprocessing model is used to compress the multimodal sequence. Since the preprocessing model is built based on the first N layers of a large language model, it can understand and process the multimodal sequence. Furthermore, the preprocessing model only needs to run once before the prefill stage of the LLM, eliminating the need for a key-value cache (KV cache) and effectively reducing computational load and GPU memory usage. In the second stage, the LLM can directly perform inference based on the target sequence. The LLM does not need to compress the sequence, reducing the impact of changes in sequence length within layers and lowering the inference overhead of the large language model.
[0034] In one possible implementation of the second aspect of this application, the sequence processing module includes:
[0035] The score acquisition submodule is used to acquire the attention score of each sequence unit in the multimodal sequence. The attention score includes the intramodal self-attention score and the intermodal mutual attention score.
[0036] The sequence compression submodule is used to determine the target sequence based on the intramodal self-attention score and the intermodal mutual attention score, wherein the number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence.
[0037] In one possible implementation of the second aspect of this application, the compression ratio of the preprocessing model is R, where R is less than 1;
[0038] The sequence compression submodule is specifically used to determine the target sequence from the multimodal sequence according to the compression ratio based on the intramodal self-attention score and the intermodal mutual attention score.
[0039] In one possible implementation of the second aspect of this application, the sequence compression submodule is specifically used for:
[0040] The sequence to be compressed is determined by the intersection of the top K sequence units with the highest intramodal self-attention scores and the top K sequence units with the highest intermodal mutual attention scores in the multimodal sequence.
[0041] The target sequence is determined from the sequence to be compressed according to the compression ratio.
[0042] In one possible implementation of the second aspect of this application, the multimodal sequence includes a first modal sequence and a second modal sequence, wherein the first modal sequence includes a text modal sequence, and the second modal sequence includes at least one of an image modal sequence, an audio modal sequence, and a video modal sequence.
[0043] In one possible implementation of the second aspect of this application, the model inference module is specifically used to input the sequence units in the target sequence and the sequence units in the context window of the preprocessing model into the Large Language Model (LLM) to obtain the inference result.
[0044] In one possible implementation of the second aspect of this application, the apparatus further includes:
[0045] The parameter acquisition module is used to acquire model configuration parameters, which include the number of layers N and the compression ratio R of the preprocessed model. The number of layers N and the compression ratio R are determined based on an optimization algorithm.
[0046] The model building module is used to build the preprocessed model based on the first N layers of the large language model LLM and the compression ratio R.
[0047] In the second aspect of this application, the constituent modules of the data processing apparatus can also perform the steps described in the first aspect and various possible implementations, as detailed in the foregoing description of the first aspect and various possible implementations.
[0048] Thirdly, embodiments of this application provide a server that may include a memory and a processor, wherein the memory is used to store computer programs or computer instructions, and the processor is used to execute the computer programs or computer instructions stored in the memory, so that the server performs the method of the first aspect of the embodiments of this application or any possible implementation of the first aspect.
[0049] Fourthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect above.
[0050] Fifthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first aspect above.
[0051] In a sixth aspect, embodiments of this application provide a communication device, which may include entities such as terminal devices or chips. The communication device includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, causing the communication device to perform the method as described in any one of the preceding first aspects.
[0052] In a seventh aspect, this application provides a chip system including a processor for supporting a data processing apparatus in implementing the functions involved in the above aspects, such as transmitting or processing data and / or information involved in the above methods.
[0053] In one possible design, the processor is coupled to the memory via an interface.
[0054] In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the data processing device. The chip system may consist of chips or may include chips and other discrete components.
[0055] Eighthly, embodiments of this application provide a chip including one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, it causes the electronic device to perform the method in the first aspect or any possible implementation of the first aspect.
[0056] Each of the third to eighth aspects corresponds to the first aspect and any one of the implementations of the first aspect, respectively. The technical effects corresponding to any one of the third and eighth aspects can be found in the technical effects corresponding to the first aspect and any one of the implementations of the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0057] Figure 1 is a schematic diagram of an artificial intelligence main framework applied in this application;
[0058] Figure 2 is a schematic diagram of a system architecture provided in this application;
[0059] Figure 3 is a schematic diagram of another system architecture provided in this application;
[0060] Figure 4 is a flowchart illustrating a data processing method provided in an embodiment of this application;
[0061] Figure 5 is a flowchart illustrating another data processing method provided in an embodiment of this application;
[0062] Figure 6 is a flowchart illustrating another data processing method provided in an embodiment of this application;
[0063] Figure 7 is a schematic diagram of a configuration method for a preprocessing model provided in an embodiment of this application;
[0064] Figure 8 is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;
[0065] Figure 9 is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0066] Figure 10 is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0067] This application provides a data processing method, apparatus, device, and storage medium, which aims to compress the input sequence of a multimodal understanding model and reduce the inference overhead of a large language model.
[0068] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0069] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0070] First, the overall workflow of an artificial intelligence system is described, as shown in Figure 1. Figure 1 is a structural diagram of the main framework of artificial intelligence. The framework is then elaborated on from two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.
[0071] (1) Infrastructure
[0072] Infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. This communication occurs through sensors; computing power is provided by intelligent chips, such as hardware acceleration chips (CPUs, NPUs, GPUs, ASICs, or FPGAs); the basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, which is then provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0073] (2) Data
[0074] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0075] (3) Data processing
[0076] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.
[0077] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.
[0078] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0079] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0080] (4) General ability
[0081] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0082] (5) Smart Products and Industry Applications
[0083] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0084] This application relates to applications of deep learning. To better understand the solutions in this application, the relevant terms and concepts that may be involved in this application are introduced below. It should be understood that the explanation of the relevant concepts may be limited due to the specific circumstances of this application, but it does not mean that this application is limited to that specific circumstance. The specific circumstances of different embodiments may also differ, and no specific limitation is made here.
[0085] Large language models (LLMs) are a class of deep learning-based artificial intelligence models designed to process and generate natural language text. Trained on large-scale text datasets, LLMs can understand and generate text similar to human language, performing various natural language processing tasks including text generation, translation, and sentiment analysis. The core technical architecture of LLMs is the Transformer model, a deep learning model based on a self-attention mechanism. The Transformer architecture captures dependencies in the sequence by calculating the attention score of each element in the input sequence to other elements, thereby achieving efficient processing of natural language.
[0086] Large vision-language model (LVLM): This is a model that combines vision and language processing capabilities to handle various types of tasks, including image description, question answering, and content creation. By integrating a large vision-language model (LVLM), it improves the processing performance of multimodal tasks, enabling it to understand image / audio content and generate relevant text descriptions, or generate corresponding images / audio based on text descriptions.
[0087] Neural network components: These refer to the basic units and structures that make up a neural network. These components work together to enable the neural network to learn and process complex data.
[0088] Referring to Figure 2, this application embodiment provides a system architecture 200. This system architecture includes a database 230 and a client device 240. A data acquisition device 260 is used to collect data and store it in the database 230. A training module 202 generates a target model / rule 201 based on the data maintained in the database 230. The following will describe in more detail how the training module 202 obtains the target model / rule 201 based on the data. The target model / rule 201 refers to the various models and neural network components mentioned in the following embodiments of this application; please refer to the relevant descriptions below for details.
[0089] The computation module may include a training module 202, and the target model / rules obtained by the training module 202 can be applied to different systems or devices. In Figure 2, the execution device 210 is configured with a transceiver 212, which may be a wireless transceiver, an optical transceiver, or a wired interface (such as an I / O interface), etc., to interact with external devices. The "user" can input data to the transceiver 212 through the client device 240. For example, in the following embodiments of this application, the client device 240 can send a target task to the execution device 210, requesting the execution device to build a neural network, and send a database for training to the execution device 210.
[0090] The execution device 210 can call data, code, etc. in the data storage system 250, and can also store data, instructions, etc. in the data storage system 250.
[0091] The calculation module 211 processes the input data using the target model / rule 201. Specifically, the calculation module 211 is used for:
[0092] Finally, transceiver 212 returns the constructed neural network to client device 240 for deployment in client device 240 or other devices.
[0093] At a deeper level, the training module 202 can obtain corresponding target models / rules 201 based on different data for different tasks, so as to provide users with better results.
[0094] In the scenario shown in Figure 2, the data input to the execution device 210 can be determined based on the user's input data. For example, the user can operate through the interface provided by the transceiver 212. Alternatively, the client device 240 can automatically input data to the transceiver 212 and obtain results. If the client device 240 needs user authorization to automatically input data, the user can set appropriate permissions in the client device 240. The user can view the results output by the execution device 210 on the client device 240; the specific presentation format can be display, sound, animation, etc. The client device 240 can also act as a data acquisition terminal, storing the acquired data associated with the target task into the database 230.
[0095] The training or update process mentioned in this application can be executed by the training module 202. It is understood that the training process of a neural network is learning how to transform the control space, more specifically, learning the weight matrix. The purpose of training a neural network is to make its output as close as possible to the expected value. Therefore, this can be achieved by comparing the current network's predicted value with the expected value, and then updating the weight vector of each layer of the neural network based on the difference between the two (of course, the weight vector can usually be initialized before the first update, i.e., pre-configured parameters for each layer in the deep neural network). For example, if the network's predicted value is too high, the values of the weights in the weight matrix are adjusted to lower the predicted value. This adjustment continues until the neural network's output value is close to or equal to the expected value. Specifically, the difference between the neural network's predicted value and the expected value can be measured using a loss function or an objective function. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, and the training of the neural network can be understood as a process of minimizing the loss as much as possible. The process of updating the weights of the starting network and training the serial network in the following embodiments of this application can be referred to in this process, and will not be repeated hereafter.
[0096] As shown in Figure 2, the target model / rule 201 is obtained by training module 202. The target model / rule 201 may be a network with Transformer architecture, deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN), residual network, or other neural networks.
[0097] During the training phase, database 230 stores a set of training samples. Training device 220 generates a target model / rule 201 for processing the samples and iteratively trains the target model / rule 201 using the sample set in the database to obtain a mature target model / rule 201, which is specifically represented as a neural network. The neural network obtained by training device 220 can be applied to different systems or devices.
[0098] During the inference phase, the execution device 210 can access data, code, etc., from the data storage system 250, or it can store data, instructions, etc., in the data storage system 250. The data storage system 250 can be located within the execution device 210, or it can be an external memory relative to the execution device 210. The computing module 211 can process the samples acquired by the execution device 210 through a neural network to obtain prediction results. The specific form of the prediction results is related to the function of the neural network.
[0099] It should be noted that Figure 2 is merely an exemplary schematic diagram of a system architecture provided in this application embodiment, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in Figure 2, the data storage system 250 is an external memory relative to the execution device 210. In other scenarios, the data storage system 250 can also be placed within the execution device 210.
[0100] The target model / rule 201 constructed based on the training module 202 can be applied to different systems or devices, such as mobile phones, tablets, laptops, augmented reality (AR) / virtual reality (VR), in-vehicle terminals, servers, or cloud devices.
[0101] The target model / rule 201 in this embodiment of the application may be a network of the large language model (LLM) in this application, or a large vision-language model (LVLM).
[0102] Referring to Figure 3, this application embodiment also provides a system architecture 300. The execution device 210 is implemented by one or more servers, optionally in conjunction with other computing devices, such as data storage, routers, load balancers, etc. The execution device 210 can be deployed on a single physical site or distributed across multiple physical sites. The execution device 210 can use data from the data storage system 350 or call program code from the data storage system 350 to implement the steps of the deep learning training method for computing devices corresponding to Figure 6 below.
[0103] Users can interact with execution device 210 by operating their respective user devices (e.g., local device 301 and local device 302). Each local device can represent any computing device, such as a personal computer, computer workstation, smartphone, tablet, smart camera, smart car or other type of cellular phone, media consumption device, wearable device, set-top box, game console, etc.
[0104] Each user's local device can interact with execution device 210 through a communication network using any communication mechanism / standard. The communication network can be a wide area network (WAN), a local area network (LAN), a point-to-point connection, or any combination thereof. Specifically, the communication network can include a wireless network, a wired network, or a combination of both. The wireless network includes, but is not limited to, any one or more combinations of: 5th-Generation (5G) systems, Long Term Evolution (LTE) systems, Global System for Mobile Communication (GSM) or Code Division Multiple Access (CDMA) networks, Wideband Code Division Multiple Access (WCDMA) networks, Wireless Fidelity (WiFi), Bluetooth, Zigbee, Radio Frequency Identification (RFID), Long Range (Lora) wireless communication, and Near Field Communication (NFC). The wired network can include fiber optic communication networks or networks composed of coaxial cables.
[0105] In another implementation, one or more aspects of the execution device 210 may be implemented by each local device. For example, local device 301 may provide local data or feedback calculation results to the execution device 210. This local device may also be referred to as a computing device.
[0106] It should be noted that all the functions of execution device 210 can also be implemented by a local device. For example, local device 301 implements the functions of execution device 210 and provides services to its own users, or provides services to users of local device 302.
[0107] With the rapid development of artificial intelligence, Large Language Models (LLMs) are widely used in natural language processing tasks, including question answering, summarization, and sentiment analysis. Multimodal Understanding Models (LVLMs) combine visual and language processing capabilities, improving performance in multimodal tasks by integrating LLMs. The introduction of multimodal data such as images, videos, and audio increases the length of encoded sequences (composed of sequence unit tokens) and data redundancy. Current sequence compression schemes suffer from high computational cost and high GPU memory usage, leading to significant inference overhead for large language models.
[0108] Therefore, this application provides a data processing method to compress multimodal sequences and reduce the inference overhead of large language models.
[0109] The data processing method provided in this application can be executed on a server or on a terminal device. The terminal device can be a mobile phone, tablet personal computer (TPC), media player, smart TV, laptop computer (LC), personal digital assistant (PDA), personal computer (PC), camera, camcorder, smartwatch, wearable device (WD), or autonomous vehicle, etc., and this application does not limit the specific device to this one.
[0110] This application provides a flowchart of a data processing method that can be used in a multimodal reasoning system. The multimodal reasoning system includes a pre-processing model (Pre-Model) and a large language model (LLM). The pre-processing model (Pre-Model) includes an N-layer network, and the large language model (LLM) includes an L-layer network. The N-layer network of the pre-processing model (Pre-Model) is constructed based on the first N layers of the L-layer network of the large language model (LLM).
[0111] In the framework of the multimodal inference system provided in this application embodiment, a pre-processing model (Pre-Model) is first used to perceive and compress multimodal sequences. Since the pre-processing model is built based on the first N layers of the large language model (LLM), it can understand and process multimodal sequences. Then, the compressed sequence output by the Pre-Model is input into the LLM for inference. Thus, during the model inference process, the LLM does not need to perform sequence compression, avoiding modifications to the LLM backbone. Furthermore, the pre-processing model only needs to run once before the prefill stage of the LLM, eliminating the need for key-value cache (KV cache) storage, effectively reducing computational load and GPU memory usage.
[0112] It should be noted that the Large Language Model (LLM) in this embodiment is a pre-trained inference model that can meet the user's inference needs. The training process of the LLM can be based on the user's inference needs, such as inference accuracy requirements, multimodal input data requirements, and inference time requirements. Therefore, the LLM in this embodiment is a large language model capable of meeting the user's requirements for multimodal input data and inference performance. This application does not limit the training process of the LLM.
[0113] The multimodal reasoning system provided in this application embodiment can be applied to various multimodal understanding models, such as: Tongyi Qianwen Visual Language Multimodal Large Model Qwen2-VL, Shusheng Wanxiang Multimodal Large Model InternVL2, Large Language and Vision Assistant (LLaVA), etc., or general client models, and is not limited in this application embodiment.
[0114] It should be noted that the preprocessing model is pre-built based on the first N layers of the LLM network, and the specific value of N is not limited in this embodiment. Optionally, N can be obtained through an automatic optimization algorithm based on a specific inference model and test dataset. Optionally, N can be selected from 2 to 28 layers. This embodiment does not impose any limitation.
[0115] Figure 4 shows a flowchart of a data processing method provided in an embodiment of this application. The method mainly includes the following steps:
[0116] Step S401: Obtain the multimodal sequence.
[0117] In this embodiment, the specific type of multimodal sequence is not limited. The multimodal sequence may include at least one of text modal sequence, image modal sequence, audio modal sequence and video modal sequence.
[0118] It is understood that the method of obtaining multimodal sequences is not limited in the embodiments of this application. For example, multimodal sequences can be obtained by a multimodal encoder.
[0119] Step S402: Obtain the target sequence of the multimodal sequence based on the preprocessing model. The preprocessing model is constructed based on the first N layers of the large language model LLM, where N is greater than 1.
[0120] Specifically, in this embodiment, the multimodal sequence can be processed based on a preprocessing model to obtain a target sequence. The number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence. Thus, the multimodal sequence can be compressed based on the preprocessing model, reducing the inference overhead of large language models.
[0121] Optionally, step S402 specifically includes: obtaining the attention score of each sequence unit in the multimodal sequence, the attention score including intramodal self-attention score and intermodal mutual attention score; and determining the target sequence based on the intramodal self-attention score and the intermodal mutual attention score.
[0122] In this embodiment, the preprocessing model can compress the multimodal sequence based on the attention score of each sequence unit (token) in the multimodal sequence. The attention score of each sequence unit considers both its self-attention score within the modality and its mutual attention score between modalities. That is, when compressing the multimodal sequence, this embodiment not only considers the information density of the modality itself but also focuses on the information fusion between modalities, ensuring the reliability of multimodal sequence selection and compression, and thus ensuring the reliability of LLM inference.
[0123] In addition, in the embodiments of this application, the number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence. It can be seen that the sequence length of the target sequence obtained after compression is less than the sequence length of the multimodal sequence, which effectively reduces the sequence length while ensuring the model inference performance and reducing the inference overhead of large language models.
[0124] It should be noted that the calculation methods for the intramodal self-attention score and the intermodal mutual attention score of the sequence unit in the embodiments of this application can refer to the existing attention score calculation methods, and are not limited in the embodiments of this application.
[0125] For example, the intramodal self-attention score of a sequence unit can be calculated by the dot product of the sequence unit in the modal sequence with all the sequence units in the modal sequence, as shown in the following formula:
[0126] First, assume there is a modal sequence X = [x1, x2, ..., xi], where xi represents the i-th element in the sequence (such as a word embedding vector or an image feature vector, which refers to different things in different modalities, and the element can be interchanged with the sequence unit in the embodiment).
[0127] Next, three weight matrices are defined: WQ (query matrix), WK (key matrix), and WV (value matrix). These matrices are parameters learned during the training of the large language model, and their dimensions typically match the embedding dimensions of the input elements.
[0128] Next, for each element xi in the sequence, its query vector qi, key vector ki, and value vector vi are calculated using the following formulas:
[0129] qi=xi*WQ, ki=xi*WK, vi=xi*WV;
[0130] Finally, for each element xi (or its query vector qi) in the sequence, calculate its dot product with all elements xj (or their key vector kj, where j can include i itself) in the sequence to obtain the intramodal self-attention score of the element. The formula is:
[0131] score(xi, xj) = qi * kj^T;
[0132] Where ^T denotes the transpose of a vector.
[0133] Understandably, to prevent the dot product result from being too large and causing the gradient of the softmax function to vanish, a scaling factor can be introduced to scale the dot product result, and the scaled dot product result can be normalized by the softmax function to obtain the final self-attention score.
[0134] For example, the intermodal mutual attention score of a sequence unit can be calculated by taking the dot product of the sequence unit in the modal sequence with all sequence units in the other modal sequences. The following example illustrates this using a multimodal sequence including a first modal sequence and a second modal sequence:
[0135] First, assume there exists a modality A sequence XA and a modality B sequence XB, XA = [xa1, xa2, ..., xai], XB = [xb1, xb2, ..., xbj], where xai represents the i-th element in modality A (such as a word embedding vector or an image feature vector, which has different meanings in different modalities, and the element can be interchanged with the sequence unit in the embodiment), and xbj represents the j-th element in modality B (such as a word embedding vector or an image feature vector, which has different meanings in different modalities, and the element can be interchanged with the sequence unit in the embodiment).
[0136] Next, for each element xai in mode A, generate the corresponding query vector qai, key vector kai, and value vector vai; for each element xbj in mode B, generate the corresponding query vector qbi, key vector kbi, and value vector vbi (see the introduction to autonomous force score calculation for specific generation steps).
[0137] Finally, for each element xai in mode A, calculate its dot product with all elements xbj (or its key vector kbj) in mode B to obtain the mutual attention score from mode A to mode B. The formula is: score(xai, xbj) = qai*kbj^T (T represents transpose).
[0138] Similarly, the mutual attention score from mode B to mode A can be calculated.
[0139] Similarly, in the process of calculating mutual attention scores, in order to prevent the gradient of the softmax function from vanishing due to an excessively large dot product result, a scaling factor can be introduced to scale the dot product result, and the scaled dot product result can be normalized by the softmax function to obtain the final self-attention score.
[0140] It should be noted that the above calculation methods for intramodal self-attention scores and intermodal mutual attention scores of sequence units are only illustrative examples. In practical applications, existing attention score calculation methods can be used for understanding and explanation. The above calculation methods do not constitute a limitation of the embodiments of this application.
[0141] Optionally, the multimodal sequence includes a first modal sequence and a second modal sequence, wherein the first modal sequence includes a text modal sequence, and the second modal sequence includes at least one of an image modal sequence, an audio modal sequence, and a video modal sequence.
[0142] In this embodiment, the multimodal sequence can be further divided into two types of modal sequences: text modal sequences and audiovisual modal sequences, such as image modal sequences, audio modal sequences, and video modal sequences. Therefore, when calculating the intermodal mutual attention score of each sequence unit, it can be performed according to the two modal sequences. Compared to further dividing the audiovisual modal sequence into multiple types of modal sequences for intermodal mutual attention score calculation, this reduces the computational load of attention scores and the overhead of sequence compression.
[0143] Optionally, the multimodal sequence can be further subdivided. For example, the multimodal sequence may include a first modal sequence, a second modal sequence, a third modal sequence, and a fourth modal sequence, where the first modal sequence is a text modal sequence, the second modal sequence is an image modal sequence, the third modal sequence is an audio modal sequence, and the fourth modal sequence is a video modal sequence. Therefore, when calculating the intermodal mutual attention score of each sequence unit, the scores can be calculated pairwise for each sequence unit. For example, the intermodal mutual attention score between sequence units in the first modal sequence and those in the second modal sequence (as the first score), the intermodal mutual attention score between sequence units in the first modal sequence and those in the third modal sequence (as the second score), and the intermodal mutual attention score between sequence units in the first modal sequence and those in the fourth modal sequence (as the third score) can be calculated separately. Finally, the first, second, and third scores are averaged to obtain the intermodal mutual attention score of the sequence units in the first modal sequence.
[0144] Optionally, in this embodiment, the preprocessing model is configured with a compression ratio parameter R, where R is less than 1; the above step S402 specifically includes: determining the target sequence from the multimodal sequence according to the compression ratio based on the intramodal self-attention score and the intermodal mutual attention score.
[0145] In this embodiment, the preprocessing model is further configured with a compression ratio parameter R. By adjusting the compression ratio, the degree of compression of the multimodal sequence can be controlled, thereby controlling the sequence length of the target sequence. For example, if the multimodal sequence includes 100 sequence units and the compression ratio R is 50%, then the compressed target sequence will include 50 sequence units.
[0146] It should be noted that the compression ratio of the preprocessed model is R, and the specific value of R is not limited in this embodiment. Optionally, R can be obtained through an automatic optimization algorithm based on a specific model and test dataset. Optionally, R can be set to 50% to 90%. This embodiment does not impose any limitation.
[0147] Optionally, the above steps involve determining the target sequence from the multimodal sequence based on the intramodal self-attention score and the intermodal mutual attention score, according to the compression ratio. Specifically, this includes: determining the sequence to be compressed based on the intersection of the top K sequence units with the highest intramodal self-attention scores and the top K sequence units with the highest intermodal mutual attention scores; and determining the target sequence from the sequence to be compressed according to the compression ratio.
[0148] The sequence compression process provided in this application embodiment may specifically include, firstly, obtaining the intramodal self-attention score A of each sequence unit. S And the intermodal mutual attention score A for each sequence unit. C Next, the top K intramodal self-attention scores (TopK(A)) of each sequence unit are taken. S The TopK(A) intermodal mutual attention scores are taken from the top K units in the sequence. C TopK(A) intersection S )∧TopK(A C The sequence to be compressed is denoted as ; finally, the sequence to be compressed is compressed according to the compression ratio R to obtain the final target sequence.
[0149] Optionally, the data processing method provided in this application embodiment further includes: obtaining model configuration parameters, including the number of layers N and compression ratio R of the preprocessed model, wherein the number of layers N and compression ratio R are determined based on an optimization algorithm; and constructing a preprocessed model based on the first N layers of the large language model LLM and the compression ratio R.
[0150] In this embodiment, relevant model configuration parameters for configuring the preprocessing model can be obtained. The model configuration parameters include the number of layers N and the compression ratio R of the preprocessing model. These two parameters can be used to construct a preprocessing model that meets the requirements of multimodal sequence compression. The search space for the number of layers N and the compression ratio R of the preprocessing model can be defined based on a specific large language model and test dataset, and obtained by fitting through an automatic optimization algorithm to meet the user's model inference accuracy requirements.
[0151] It should be noted that the embodiments of this application do not limit the automatic optimization algorithm. For example, Bayesian optimization (BO), random search (RS), grid search (GS), etc. can be used.
[0152] Step S403: Obtain the inference results of the target sequence based on the Large Language Model (LLM).
[0153] Specifically, the compressed results of the multimodal sequences are input into the large language model LLM for inference to obtain the inference results.
[0154] Optionally, step S403 specifically includes: inputting the sequence units in the target sequence and the sequence units in the context window of the preprocessing model into the large language model LLM to obtain the inference result.
[0155] The context window, also known as the observation window, refers to the number of sequence unit tokens that a large language model can receive and consider when processing sequences. This window can be measured by a certain number of sequence unit tokens. In this embodiment, no limit is made on the number of sequence units contained in the context window. For example, the context window may include 50 sequence unit tokens.
[0156] In this embodiment, the sequence units in the target sequence obtained by compression and the sequence units in the context window of the large language model are input into the large language model for model inference to obtain the model inference result.
[0157] The data processing method provided in this application embodiment will be described below with reference to the flowchart of another data processing method shown in Figure 5:
[0158] First, data from different modalities are encoded by a multimodal encoder to obtain multimodal sequences. In Figure 5, the data from different modalities are divided into visual modal data (including video, audio, and image data) and text modal data (including long text data). The multimodal encoder can use the Transformer model, which is a deep learning model architecture used for natural language processing (NLP) and other sequence-to-sequence tasks. It mainly consists of two parts: an encoder and a decoder, which can be composed of Transformer blocks.
[0159] Then, the multimodal sequence is input into a preprocessing model, which consists of an N-layer network architecture. This N-layer network architecture is constructed based on the first N layers of the Large Language Model (LLM). The preprocessing model will calculate the self-attention score A of each sequence unit in the multimodal sequence. S And the intermodal mutual attention score A for each sequence unit. C Based on the compression ratio R and the TopK(A) of the sequence to be compressed, S )∧TopK(A C The multimodal sequence is compressed to obtain the target sequence.
[0160] Finally, the target sequence is input into the Large Language Model (LLM) for model inference to obtain the model inference result. The Large Language Model (LLM) includes a complete L-layer network.
[0161] It should be noted that in Figure 5, This represents the intramodal self-attention score of the visual sequence. The intermodal mutual attention score represents the visual sequence. This represents the intramodal self-attention score of a text modal sequence. This represents the intermodal mutual attention score of a visual sequence.
[0162] Therefore, this application provides a two-stage multimodal sequence inference scheme. In the first stage, a preprocessing model is used to compress the multimodal sequence. Since the preprocessing model is built based on the first N layers of a large language model, it can understand and process the multimodal sequence. Furthermore, the preprocessing model only needs to run once before the prefill stage of the LLM, eliminating the need for a key-value cache (KV cache) and effectively reducing computational load and GPU memory usage. In the second stage, the LLM can directly perform inference based on the target sequence. The LLM does not need to compress the sequence, reducing the impact of changes in sequence length within layers and lowering the inference overhead of the large language model.
[0163] The data processing method provided in this application embodiment will be further described below with reference to the flowchart of another data processing method provided in Figure 6:
[0164] First, based on the multimodal test set and customer performance requirements, an automatic optimization algorithm is used to obtain the number of layers N and compression ratio R of the pre-processed model Pre-Model.
[0165] Then, a preprocessing model is constructed based on the number of layers N and the compression ratio R.
[0166] Next, the multimodal encoder converts the multimodal data into a multimodal sequence and inputs it into the preprocessing model. The preprocessing model will calculate the self-attention score A for each sequence unit in the multimodal sequence. S And the intermodal mutual attention score A for each sequence unit. C Based on the compression ratio R and the TopK(A) of the sequence to be compressed, S )∧TopK(A C The multimodal sequence is compressed (token filtering) to obtain the target sequence.
[0167] Finally, the target sequence is input into the Large Language Model (LLM) for model inference to obtain the model inference result.
[0168] Optionally, referring to the schematic diagram of a preprocessing model configuration method provided in Figure 7, this embodiment of the application provides a large model inference suite. For different client models (large language models), model configuration parameters are pre-optimized to obtain different client model configuration parameters. These parameters include the number of layers N and the compression ratio R of the preprocessing model. In actual use, users can select the corresponding model configuration parameters for different client models, construct a preprocessing model based on the model configuration parameters and the client model, and perform multimodal sequence compression and model inference. Therefore, users can selectively enable the preprocessing model as a plug-in feature before using the large language model for inference, thereby improving inference performance and reducing inference overhead.
[0169] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0170] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.
[0171] Please refer to Figure 8. An embodiment of this application provides a data processing device 800, which can be applied to a multimodal reasoning system. The multimodal reasoning system includes a preprocessing model and a Large Language Model (LLM). The device includes:
[0172] Sequence acquisition module 801 is used to acquire multimodal sequences;
[0173] Sequence processing module 802 is used to obtain the target sequence of the multimodal sequence based on a preprocessing model. The preprocessing model is constructed based on the first N layers of the large language model LLM, where N is greater than 1.
[0174] Model inference module 803 is used to obtain inference results of target sequences based on large language model LLM.
[0175] In one possible implementation of this application embodiment, the sequence processing module 801 includes:
[0176] The score acquisition submodule is used to acquire the attention score of each sequence unit in the multimodal sequence. The attention score includes the intramodal self-attention score and the intermodal mutual attention score.
[0177] The sequence compression submodule is used to determine the target sequence based on the intramodal self-attention score and the intermodal mutual attention score, wherein the number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence.
[0178] In one possible implementation of this application embodiment, the compression ratio of the preprocessed model is R, where R is less than 1;
[0179] The sequence compression submodule is specifically used to determine the target sequence from the multimodal sequence according to the compression ratio based on the intramodal self-attention score and the intermodal mutual attention score.
[0180] In one possible implementation of this application embodiment, the sequence compression submodule is specifically used to: determine the sequence to be compressed based on the intersection of the top K sequence units with the highest intramodal self-attention scores of each sequence unit in the multimodal sequence and the top K sequence units with the highest intermodal mutual attention scores of each sequence unit in the multimodal sequence; and determine the target sequence from the sequence to be compressed according to the compression ratio.
[0181] In one possible implementation of this application embodiment, the multimodal sequence includes: a first modal sequence and a second modal sequence, wherein the first modal sequence includes: a text modal sequence, and the second modal sequence includes: at least one of an image modal sequence, an audio modal sequence, and a video modal sequence.
[0182] In one possible implementation of this application embodiment, the model inference module 803 is specifically used to input the sequence units in the target sequence and the sequence units in the context window of the preprocessing model into the Large Language Model (LLM) to obtain the inference result.
[0183] In one possible implementation of this application embodiment, the data processing apparatus further includes:
[0184] The parameter acquisition module is used to acquire model configuration parameters, including the number of layers N and compression ratio R of the preprocessed model. The number of layers N and compression ratio R are determined based on an optimization algorithm.
[0185] The model building module is used to build a preprocessed model based on the first N layers of the large language model LLM and the compression ratio R.
[0186] According to the data processing apparatus provided in this application embodiment, before inputting the multimodal sequence into the large language model for inference, the multimodal sequence is first compressed by a preprocessing model. The preprocessing model is constructed based on the first N layers of the large language model, thus it can understand and process multimodal sequences. Furthermore, the preprocessing model only needs to run once before the prefill stage of the LLM, eliminating the need for key-value cache (KV cache), effectively reducing computational load and GPU memory usage. Moreover, the preprocessing model combines the self-attention score of sequence units within a modality and the mutual attention score of sequence units between modalities when compressing the multimodal sequence. That is, in this application embodiment, when compressing the multimodal sequence, not only is the information density of the modality itself considered, but also the information fusion between modalities, ensuring the reliability of multimodal sequence selection and compression, thereby ensuring the reliability of LLM inference.
[0187] In this embodiment of the application, the module is an example of a software functional unit, and the data processing device may include code running on a computing instance. The computing instance may be at least one of a physical host (computer device), a virtual machine, a container, or other computer devices.
[0188] Furthermore, the aforementioned computer equipment can be one or more devices. For example, the data processing device may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application can be distributed within the same region or in different regions. Similarly, the multiple hosts / virtual machines / containers used to run the code can be distributed within the same available zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0189] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a single region. Communication between two VPCs within the same region, and between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0190] As an example of a hardware functional unit, a data processing device may include at least one computer device, such as a server. Alternatively, the data processing device may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The aforementioned PLD may be implemented using a complex PLD (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0191] The data processing device comprises multiple computer devices that can be distributed within the same region or in different regions. Similarly, the multiple computer devices can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computer devices can be distributed within the same Virtual Private Cloud (VPC) or multiple VPCs. These multiple computer devices can be any combination of computer devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0192] This application also provides a computer device 130. As shown in FIG9, the computer device 130 includes a bus 132, a processor 134, a memory 136, and a communication interface 138. The processor 134, the memory 136, and the communication interface 138 communicate with each other via the bus 132. The computer device 130 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computer device 130.
[0193] Bus 132 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 9, but this does not imply that there is only one bus or one type of bus. Bus 134 can include pathways for transmitting information between various components of computer device 130 (e.g., memory 136, processor 134, communication interface 138).
[0194] The processor 134 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0195] Memory 136 may include volatile memory, such as random access memory (RAM). Processor 134 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0196] The memory 136 stores executable program code, and the processor 134 executes the executable program code to implement the functions of the aforementioned acquisition module and training module, thereby realizing the data processing method applied to the computer device cluster in the above embodiment. That is, the memory 136 stores instructions for executing the data processing method applied to the computer device cluster in the above embodiment.
[0197] The communication interface 138 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computer device 130 and other devices or communication networks.
[0198] This application also provides a computer program product that, when run on a computer, causes the computer to perform the steps executed by the data processing device in the method described in the foregoing embodiments.
[0199] The data processing apparatus provided in this application embodiment can be a chip, which includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in a storage unit to cause the chip within the server to execute the data processing method described in the above embodiments. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0200] Specifically, the aforementioned processing unit or processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0201] This application also provides a server. Referring to Figure 10, which is a schematic diagram of a server structure provided in this application embodiment, the server 10 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1522 (e.g., one or more processors) and a memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 can be temporary or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the figure), each module including a series of instruction operations on the server. Furthermore, the CPU 1522 may be configured to communicate with the storage media 1530 and execute the series of instruction operations in the storage media 1530 on the server 1500.
[0202] Server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0203] In this embodiment, the central processing unit 1522 is used to execute the data processing method in the embodiment corresponding to FIG4, FIG5, or FIG6.
[0204] It should be noted that the specific way in which the central processing unit 1522 executes the above steps is based on the same concept as the various method embodiments corresponding to Figures 4, 5, or 6 in this application, and the resulting technical effects are the same as those of the various method embodiments corresponding to Figures 4, 5, or 6 in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0205] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0206] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0207] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.
[0208] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0209] Finally, it should be noted that the above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data processing method, characterized in that, The method is applied to a multimodal reasoning system, which includes a preprocessing model and a large language model (LLM). The method includes: Obtain multimodal sequences; The target sequence of the multimodal sequence is obtained based on the preprocessing model, wherein the preprocessing model is constructed based on the first N layers of the large language model LLM, and N is greater than 1; The inference results of the target sequence are obtained based on the large language model LLM.
2. The method according to claim 1, characterized in that, The step of obtaining the target sequence of the multimodal sequence based on the preprocessing model includes: Obtain the attention score for each sequence unit in the multimodal sequence, wherein the attention score includes intramodal self-attention score and intermodal mutual attention score; The target sequence is determined based on the intramodal self-attention score and the intermodal mutual attention score, wherein the number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence.
3. The method according to claim 2, characterized in that, The compression ratio of the preprocessed model is R, where R is less than 1; Determining the target sequence based on the intramodal self-attention score and the intermodal mutual attention score includes: Based on the intramodal self-attention score and the intermodal mutual attention score, the target sequence is determined from the multimodal sequence according to the compression ratio.
4. The method according to claim 3, characterized in that, The step of determining the target sequence from the multimodal sequence according to the compression ratio based on the intramodal self-attention score and the intermodal mutual attention score includes: The sequence to be compressed is determined by the intersection of the top K sequence units with the highest intramodal self-attention scores and the top K sequence units with the highest intermodal mutual attention scores in the multimodal sequence. The target sequence is determined from the sequence to be compressed according to the compression ratio.
5. The method according to any one of claims 1 to 4, characterized in that, The multimodal sequence includes a first modal sequence and a second modal sequence, wherein the first modal sequence includes a text modal sequence, and the second modal sequence includes at least one of an image modal sequence, an audio modal sequence, and a video modal sequence.
6. The method according to any one of claims 1 to 5, characterized in that, The inference result obtained based on the large language model LLM for the target sequence includes: The sequence units in the target sequence and the sequence units in the context window of the preprocessing model are input into the Large Language Model (LLM) to obtain the inference result.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain model configuration parameters, including the number of layers N and the compression ratio R of the preprocessed model, wherein the number of layers N and the compression ratio R are determined based on an optimization algorithm; The preprocessing model is constructed based on the first N layers of the large language model LLM and the compression ratio R.
8. A data processing apparatus, characterized in that, The device is applied to a multimodal reasoning system, which includes a preprocessing model and a Large Language Model (LLM). The device includes: The sequence acquisition module is used to acquire multimodal sequences; The sequence processing module is used to obtain the target sequence of the multimodal sequence based on the preprocessing model, wherein the preprocessing model is constructed based on the first N layers of the large language model LLM, and N is greater than 1; The model inference module is used to obtain the inference results of the target sequence based on the large language model (LLM).
9. The apparatus according to claim 8, characterized in that, The sequence processing module includes: The score acquisition submodule is used to acquire the attention score of each sequence unit in the multimodal sequence. The attention score includes the intramodal self-attention score and the intermodal mutual attention score. The sequence compression submodule is used to determine the target sequence based on the intramodal self-attention score and the intermodal mutual attention score, wherein the number of sequence units in the target sequence is less than the number of sequence units in the multimodal sequence.
10. The apparatus according to claim 9, characterized in that, The compression ratio of the preprocessed model is R, where R is less than 1; The sequence compression submodule is specifically used to determine the target sequence from the multimodal sequence according to the compression ratio based on the intramodal self-attention score and the intermodal mutual attention score.
11. The apparatus according to claim 10, characterized in that, The sequence compression submodule is specifically used for: The sequence to be compressed is determined by the intersection of the top K sequence units with the highest intramodal self-attention scores and the top K sequence units with the highest intermodal mutual attention scores in the multimodal sequence. The target sequence is determined from the sequence to be compressed according to the compression ratio.
12. The apparatus according to any one of claims 8 to 11, characterized in that, The multimodal sequence includes a first modal sequence and a second modal sequence, wherein the first modal sequence includes a text modal sequence, and the second modal sequence includes at least one of an image modal sequence, an audio modal sequence, and a video modal sequence.
13. The apparatus according to any one of claims 8 to 12, characterized in that, The model inference module is specifically used to input the sequence units in the target sequence and the sequence units in the context window of the preprocessing model into the large language model LLM to obtain the inference result.
14. The apparatus according to any one of claims 8 to 13, characterized in that, The device further includes: The parameter acquisition module is used to acquire model configuration parameters, which include the number of layers N and the compression ratio R of the preprocessed model. The number of layers N and the compression ratio R are determined based on an optimization algorithm. The model building module is used to build the preprocessed model based on the first N layers of the large language model LLM and the compression ratio R.
15. A server, characterized in that, The server includes: Memory is used to store computer programs or computer instructions; A processor for executing a computer program or computer instructions stored in the memory, causing the server to perform the method as described in any one of claims 1 to 7.
16. A computer-readable storage medium for storing a computer program, which, when executed, performs the method according to any one of claims 1 to 7.
17. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1 to 7.
18. A chip, the chip comprising a processor and a data interface, the processor reading instructions stored in a memory through the data interface to execute the method as described in any one of claims 1 to 7.