Data processing method and apparatus based on large language model, and device
By converting the original data into text data format, using a large language model to quickly obtain answers, and converting the answers back to the original data format, the waste of computing resources and time when retrieving data from a large number of network data in the prior art is solved, and fast and efficient data retrieval is achieved.
Patent Information
- Application Number
- PCT/CN2023/139569
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-26
AI Technical Summary
The prior art requires a lot of computing resources and time to retrieve data required by users from a large amount of network data, resulting in waste of server resources and excessive waiting time for users.
The data processing method based on the large language model is adopted to convert the original data into text data format, input it to the trained large language model, output the answers in the corresponding text data format, and convert the answers back to the original data format, so as to quickly obtain the target answer.
Through this method, you can query the target answer from a large amount of data by simply consuming a small amount of computing resources, significantly saving server computing resources and time overhead, and quickly determining the corresponding answers to the input questions.
Smart Images

Figure CN2023139569_26062025_PF_FP_ABST
Abstract
Description
A data processing method, device and equipment based on large language model Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, and device based on a large language model (LLM). Background Art
[0002] With the rapid development of network technology, more and more network data is available. In order to retrieve the data required by users from a large amount of network data, users need to provide keywords, and the server retrieves the data required by users from a large amount of network data based on the keywords.
[0003] However, due to the large volume of network data, servers need to consume a large amount of computing resources to retrieve the data users need, and this process takes a long time. This results in significant server resource consumption and a long delay for users to obtain the data they need. For example, in real-world scenarios, numerous cameras are deployed, capturing a large number of images. Analyzing these images to obtain the data users need requires significant computing resources and time.
[0004] Summary of the Invention
[0005] In view of this, the present application provides a data processing method, apparatus and device based on a large language model, which can save server computing resources and time expenditure, and can quickly determine the target answer.
[0006] The present application provides a data processing method based on a large language model, the method comprising:
[0007] Obtaining a first input question in an original data format; if the original data format is not a text data format, converting the first input question into a second input question in a text data format;
[0008] Inputting the second input question to the trained target large language model, and having the target large language model output a first target answer in a text data format corresponding to the second input question;
[0009] The first target answer is converted into a second target answer in the original data format.
[0010] The present application provides a data processing device based on a large language model, the device comprising:
[0011] An acquisition module, configured to acquire a first input question in an original data format; if the original data format is not a text data format, convert the first input question into a second input question in a text data format;
[0012] a processing module, configured to input the second input question into a trained target large language model, and have the target large language model output a first target answer in a text data format corresponding to the second input question;
[0013] A conversion module is used to convert the first target answer into a second target answer in an original data format.
[0014] The present application provides an electronic device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the above-mentioned data processing method based on a large language model.
[0015] As can be seen from the above technical solution, in an embodiment of the present application, the first input question can be converted into a second input question in a text data format, the second input question is input into the target large language model, the target large language model outputs the first target answer in a text data format corresponding to the second input question, and the first target answer is converted into a second target answer in an original data format. Thus, the target answer is obtained with the help of the target large language model, and an accurate and reliable target answer is obtained by inputting the input question in a text data format. Only a small amount of computing resources is required to query the target answer corresponding to the input question from a large amount of data. The time to obtain the target answer is shorter, thereby saving the computing resources of the server and saving time overhead. It is possible to quickly determine the target answer corresponding to the input question, provide logical analysis and reasoning capabilities, and use the analysis capabilities of the target large language model to give an accurate and reliable target answer. Moreover, by converting multimodal data (such as pictures, videos, voice, etc.) into text data and inputting it into the target large language model, the target large language model supports multimodal data. The target large language model only needs to support the input and output of text data, thereby meeting the input and output requirements of multimodal data and supporting the functions of the multimodal large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments of the present application or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the embodiments of the present application.
[0017] FIG1 is a schematic flow chart of a data processing method based on a large language model in one embodiment;
[0018] FIG2 is a schematic structural diagram of a target large language model in one embodiment;
[0019] FIG3 is a schematic flow chart of a data processing method based on a large language model in one embodiment;
[0020] FIG4A is a schematic diagram of a BLIP2 algorithm in one embodiment;
[0021] FIG4B is a schematic diagram of a whisper algorithm in one embodiment;
[0022] FIG4C is a schematic diagram of an SD algorithm in one embodiment;
[0023] FIG5 is a schematic structural diagram of a data processing device based on a large language model in an embodiment;
[0024] FIG6 is a hardware structure diagram of an electronic device in an embodiment. DETAILED DESCRIPTION
[0025] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a," "the," and "the" used in this application and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more associated listed items.
[0026] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" used may also be interpreted as "at the time of" or "when" or "in response to determining".
[0027] In an embodiment of the present application, a data processing method based on a large language model is proposed. The method can be applied to any electronic device, such as a server. As shown in FIG1 , the method may include:
[0028] Step 101: Obtain a first input question in an original data format; if the original data format is not a text data format, convert the first input question into a second input question in a text data format.
[0029] Step 102: Input the second input question to the trained target large language model, and the target large language model outputs a first target answer in a text data format corresponding to the second input question.
[0030] Step 103: Convert the first target answer into a second target answer in the original data format.
[0031] In one example, converting the first input question into a second input question in a text data format may include, but is not limited to: if the original data format is an image data format, using an image-to-text conversion algorithm to convert the first input question in the image data format into a second input question in a text data format.
[0032] Alternatively, if the original data format is a voice data format, a voice-to-text conversion algorithm is used to convert the first input question in the voice data format into a second input question in the text data format.
[0033] Alternatively, if the original data format is a video data format, a video-to-text conversion algorithm is used to convert the first input question in the video data format into a second input question in the text data format.
[0034] In one example, converting a first target answer into a second target answer in the original data format may include, but is not limited to: if the original data format is an image data format, using a text-to-image conversion algorithm to convert the first target answer in text data format into a second target answer in image data format.
[0035] Alternatively, if the original data format is a voice data format, a text-to-voice conversion algorithm is used to convert the first target answer in text data format into a second target answer in voice data format.
[0036] Alternatively, if the original data format is a video data format, a text-to-video conversion algorithm is used to convert the first target answer in text data format into a second target answer in video data format.
[0037] In an example, the second input question is input into a trained target large language model, and the target large language model outputs a first target answer in a text data format corresponding to the second input question, which may include but is not limited to: obtaining a first token identifier and a second token identifier corresponding to the original data format.
[0038] The first token identifier is added before the second input question, and the second token identifier is added after the second input question; wherein the first token identifier is used to indicate the starting position of the second input question, and the second token identifier is used to indicate the ending position of the second input question.
[0039] The modified second input question is input to the target large language model, and the target large language model outputs a first target answer in a text data format corresponding to the second input question.
[0040] Among them, the first target answer has a first token identifier in front and a second token identifier behind, and the first token identifier is used to indicate the starting position of the first target answer, and the second token identifier is used to indicate the ending position of the first target answer.
[0041] In one example, converting a first target answer into a second target answer in a raw data format includes: determining, based on data content output by a target large language model, data content between a first token identifier and a second token identifier as the first target answer; determining a raw data format corresponding to the first token identifier or the second token identifier; and converting the first target answer into the second target answer in the raw data format.
[0042] In one example, after obtaining the first input question in the original data format, if the original data format is a text data format, the first input question is directly input into the trained target large language model, and the target large language model outputs the first target answer in the text data format corresponding to the first input question.
[0043] In an example, the training process for the target large language model may include but is not limited to: obtaining a sample dataset, where the sample dataset may include multiple sample questions and sample labels corresponding to each sample question, where the sample questions may be in text data format, and the sample labels may be in text data format.
[0044] The initial large language model to be trained is trained based on the sample questions (such as multiple sample questions) in the sample data set and the sample labels corresponding to the sample questions to obtain a trained candidate large language model.
[0045] The candidate large language model is used as the target large language model. Or,
[0046] Select some sample questions from the sample data set as reference sample questions, add a first token identifier in any data format in front of the reference sample question, and add a second token identifier in the data format after the reference sample question; train the candidate large language model based on the modified reference sample question and the sample labels corresponding to the reference sample question to obtain the target large language model.
[0047] In one example, the above execution order is only for the convenience of describing the examples given. In actual applications, the execution order between the steps can also be changed, and this execution order is not limited. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the steps included in the method may be more or less than those described in this specification. In addition, the single step described in this specification may be decomposed into multiple steps for description in other embodiments. The multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0048] As can be seen from the above technical solution, in an embodiment of the present application, the first input question can be converted into a second input question in a text data format, the second input question is input into the target large language model, the target large language model outputs the first target answer in a text data format corresponding to the second input question, and the first target answer is converted into a second target answer in an original data format. Thus, the target answer is obtained with the help of the target large language model, and an accurate and reliable target answer is obtained by inputting the input question in a text data format. Only a small amount of computing resources is required to query the target answer corresponding to the input question from a large amount of data. The time to obtain the target answer is shorter, thereby saving the computing resources of the server and saving time overhead. It is possible to quickly determine the target answer corresponding to the input question, provide logical analysis and reasoning capabilities, and use the analysis capabilities of the target large language model to give an accurate and reliable target answer. Moreover, by converting multimodal data (such as pictures, videos, voice, etc.) into text data and inputting it into the target large language model, the target large language model supports multimodal data. The target large language model only needs to support the input and output of text data, thereby meeting the input and output requirements of multimodal data and supporting the functions of the multimodal large language model.
[0049] The above technical solutions of the embodiments of the present application are described below in conjunction with specific application scenarios.
[0050] Due to the large volume of network data (such as text, video, voice, and image data), servers need to consume a large amount of computing resources to retrieve the data users need, which takes a long time. This leads to a large consumption of server resources and a long delay for users to obtain the required data. For example, in real-world scenarios, a large number of cameras are deployed, capturing a large number of images. Analyzing these images to obtain the data users need requires a large amount of computing resources and a significant amount of time.
[0051] To address these findings, a large language model (LLM) can be used to analyze data. For example, the question data can be fed into the LLM, which then outputs the target answer. This allows the target answer to be retrieved from a large amount of data, consuming only a small amount of computing resources. This shortens the time required to obtain the target answer, conserving server computing resources and time, while leveraging the analytical power of the LLM to deliver accurate and reliable answers.
[0052] For example, ChatGPT is a large language model based on natural language processing and machine learning that can simulate human language communication and enable intelligent conversations with users. It has been widely used in intelligent customer service, chatbots, voice assistants, and other fields, providing people with more convenient, fast, and intelligent services. Based on this, ChatGPT can be used to analyze data and obtain the target answer corresponding to the question data.
[0053] With the rapid development of artificial intelligence, multimodal large language models have been proposed and used. These models support input data such as video, voice, image, and text. Specifically, when a question is fed into the multimodal large language model using video, voice, image, or text data, it can output the target answer corresponding to the question data. This means that the multimodal large language model can process question data in various formats.
[0054] To achieve the above functions, it is necessary to use multimodal data to train the multimodal large language model to obtain a trained multimodal large language model. In other words, it is necessary to use video data, voice data, image data, and text data to train the multimodal large language model. In order to train the multimodal large language model, a large amount of high-quality data, such as video data, voice data, image data, and text data, is required, and the training cost is relatively high. The server needs to consume a large amount of computing resources to complete the training of the multimodal large language model, and it takes a lot of time to complete the training of the multimodal large language model.
[0055] When using a large multimodal language model, it goes through encoding, decoding, and linear adaptation processes. The input to the large multimodal language model is adapted embedding data, and this adaptation must be performed on both the input and output sides. Therefore, alignment training is required for the large multimodal language model and the linear projector (or adaptor, such as a fully connected layer). This alignment training requires a large amount of high-quality data and is costly.
[0056] In response to the above findings, the present application provides a data processing method based on a large language model, which converts multimodal data (such as pictures, videos, voice and other data) into text data and inputs it into a target large language model, so that the target large language model supports multimodal data. The target large language model only needs to support the input and output of text data, thereby meeting the input and output requirements of multimodal data and supporting multimodal functions. For example, a target large language model under a unified multimodal large model architecture is designed, and the input and output of the target large language model are both in text form, thereby avoiding alignment training on the input side and the input side.
[0057] FIG2 is a schematic diagram of the structure of a target large language model. The input data of the target large language model is in text format, and the output data of the target large language model is in text format. For example, the input data in text format is provided to the target large language model, which processes the input data and outputs the output data in text format.
[0058] As shown in FIG2 , if input data in an image data format is obtained, the image data format input data is converted into text data format. The text data format input data is provided to the target large language model, which outputs output data in a text data format. The text data format output data is then converted into image data format.
[0059] As shown in FIG2 , if input data in a speech data format is obtained, the input data in the speech data format is converted into input data in a text data format. The input data in the text data format is provided to the target large language model, which then outputs output data in a text data format. The output data in the text data format is then converted into output data in a speech data format.
[0060] As shown in FIG2 , if input data in a video data format is obtained, the video data format input data is converted into text data format. The text data format input data is provided to the target large language model, which then outputs output data in a text data format. The text data format output data is then converted into video data format.
[0061] As shown in Figure 2, if input data in text data format is obtained, the input data in text data format is directly provided to the target large language model, and the target large language model outputs output data in text data format. In this way, the output data in text data format can be directly obtained.
[0062] Based on the trained target large language model, an embodiment of the present application proposes a data processing method based on a large language model. FIG3 is a flow chart of the method, which may include:
[0063] Step 301: Obtain a first input question in original data format.
[0064] For example, a question in the original data format input by the user may be received and recorded as the first input question. Other methods may be used to obtain the first input question in the original data format, and there is no limitation to this.
[0065] Step 302: Determine the original data format of the first input question. If the original data format is an image data format, proceed to step 303. If the original data format is a voice data format, proceed to step 304. If the original data format is a video data format, proceed to step 305. If the original data format is a text data format, proceed to step 306. Of course, the above are just a few examples of original data formats and are not intended to be limiting.
[0066] In one example, if the first input question is in an image data format, the original data format is determined to be an image data format. Alternatively, if the first input question is in a voice data format, the original data format is determined to be a voice data format. Alternatively, if the first input question is in a video data format, the original data format is determined to be a video data format. Alternatively, if the first input question is in a text data format, the original data format is determined to be a text data format.
[0067] Step 303: If the original data format is a picture data format (ie, image data format), a picture-to-text conversion algorithm is used to convert the first input question in the picture data format into a second input question in the text data format. After step 303, step 306 may also be performed.
[0068] In one example, an image-to-text conversion algorithm may include, but is not limited to, a BLIP (Bootstrapping Language-Image Pre-training) algorithm and a BLIP2 algorithm. The BLIP or BLIP2 algorithm can be used to convert a first input question in image data format into a second input question in text data format. Of course, the BLIP or BLIP2 algorithm is merely an example and is not limiting. Any algorithm capable of converting a first input question in image data format into a second input question in text data format is sufficient.
[0069] For example, an example of the BLIP2 algorithm can be shown in Figure 4A. The BLIP2 algorithm may include an Image Encoder and a Q-Former. The Q-Former may include a first sub-network (multiple) and a second sub-network (multiple). The first sub-network may include a Self Attention network layer, a Cross Attention network layer, and a Feed Forward network layer. The second sub-network may include a Self Attention network layer and a Feed Forward network layer.
[0070] During training, a sample image can be provided to the Image Encoder to obtain multiple image features. These image features and the learned query vector are used as input features for the first sub-network, which then outputs image-text matching features. The text description features corresponding to the sample image can be used as input features for the second sub-network, which then outputs image-based text generation features. The Q-Former is then trained based on these features.
[0071] During use, the image to be detected (i.e., the first input question in the image data format, i.e., the first input question is an image) can be provided to the Image Encoder to obtain multiple image features, and the multiple image features and the learned queries are used as input features of the first sub-network. The first sub-network outputs a second input question in the text data format, i.e., the second input question is a text description of the first input question.
[0072] At this point, the BLIP2 algorithm can be used to convert the first input problem in the image data format into the second input problem in the text data format, where the second input problem describes the first input problem in text.
[0073] For example, the BLIP2 algorithm's Q-Former supports multiple modes, such as LM and ITM. If the LM mode is selected, the Q-Former can output text data. If the ITM mode is selected, the Q-Former can output embedding data.
[0074] Based on this, in order to output text data, the LM mode of Q-Former can be selected. In the LM mode, the first input question in the image data format can be provided to the Image Encoder to obtain multiple image features. The multiple image features and the learned queries are used as the input features of the first sub-network, and the first sub-network outputs the second input question in the text data format. This process will not be repeated here.
[0075] Step 304: If the original data format is a voice data format (i.e., a speech data format), a voice-to-text conversion algorithm is used to convert the first input question in the voice data format into a second input question in the text data format. After step 304, step 306 may also be performed.
[0076] In one example, a speech-to-text conversion algorithm may include, but is not limited to, a whisper algorithm, which can convert a first input question in speech data format into a second input question in text data format. Of course, the whisper algorithm is merely an example and is not limiting; any algorithm capable of converting a first input question in speech data format into a second input question in text data format will suffice.
[0077] For example, an example of the whisper algorithm can be shown in Figure 4B. The whisper algorithm may include a first subnetwork (such as the first subnetwork includes 2 convolutional layers and 1 GELU layer), a second subnetwork (such as the second subnetwork includes multiple Encoder Blocks), and a third subnetwork (such as the third subnetwork includes multiple Decoder Blocks), and a Cross Attention network layer is between the second subnetwork and the third subnetwork.
[0078] The speech to be detected (i.e., the first input question in the voice data format, the first input question is speech) can be provided to the whisper algorithm, which processes the speech to be detected and outputs the second input question in the text data format, i.e., the second input question is a description of the first input question using text.
[0079] At this point, the whisper algorithm can be used to convert the first input question in the voice data format into a second input question in the text data format, where the second input question describes the first input question in text.
[0080] Step 305: If the original data format is a video data format (ie, a video data format), a video-to-text conversion algorithm is used to convert the first input question in the video data format into a second input question in the text data format. After step 305, step 306 may also be performed.
[0081] In one example, a video-to-text conversion algorithm can be used to convert a first input question in video data format (i.e., the first input question is a video) into a second input question in text data format, i.e., the second input question is a textual description of the first input question. This conversion algorithm is not limited; it only needs to be able to convert the first input question in video data format into the second input question in text data format.
[0082] In one example, a first input problem (i.e., a video) in video data format can be decomposed into multiple images, with no restrictions on the decomposition process. For each image, an image-to-text conversion algorithm, such as the BLIP algorithm or the BLIP2 algorithm, can be used to convert the image into text in a text data format. The multiple texts corresponding to the multiple images are then combined to obtain a second input problem in text data format.
[0083] At this point, the first input question in the video data format can be converted into a second input question in the text data format, and the second input question describes the first input question in text.
[0084] Step 306: Input the input question in text data format to the target large language model.
[0085] In one example, for a first input question in an image data format, after converting the first input question in an image data format into a second input question in a text data format, a first token identifier and a second token identifier corresponding to the image data format can be obtained. The first token identifier is used to indicate the starting position of the second input question, and the second token identifier is used to indicate the ending position of the second input question.
[0086] The first token identifier is added before the second input question, and the second token identifier is added after the second input question to obtain the second input question in a modified text data format. After obtaining the modified second input question, the modified second input question is input into the target large language model.
[0087] For example, the first token identifier corresponding to the image data format can be , the second token identifier corresponding to the image data format can be <\IMG>, then the modified second input question can be: AAA<\IMG>, Indicates the starting position of the second input question, AAA indicates the second input question in text data format, and <\IMG> indicates the ending position of the second input question.
[0088] In one example, for a first input question in voice data format, after converting the first input question in voice data format to a second input question in text data format, a first token identifier and a second token identifier corresponding to the voice data format can be obtained. The first token identifier is used to indicate the starting position of the second input question, and the second token identifier is used to indicate the ending position of the second input question.
[0089] The first token identifier is added before the second input question, and the second token identifier is added after the second input question to obtain the second input question in a modified text data format. After obtaining the modified second input question, the modified second input question is input into the target large language model.
[0090] For example, the first token identifier corresponding to the voice data format can be <speech>, the second token identifier corresponding to the voice data format can be <\SPEECH>, then the modified second input question can be: <speech>AAA<\SPEECH>, <speech>Indicates the starting position of the second input question, AAA indicates the second input question in text data format, and <\SPEECH> indicates the ending position of the second input question.
[0091] In one example, for a first input question in a video data format, after converting the first input question in the video data format into a second input question in a text data format, a first token identifier and a second token identifier corresponding to the video data format can be obtained. The first token identifier is used to indicate the starting position of the second input question, and the second token identifier is used to indicate the ending position of the second input question.
[0092] The first token identifier is added before the second input question, and the second token identifier is added after the second input question to obtain the second input question in a modified text data format. After obtaining the modified second input question, the modified second input question is input into the target large language model.
[0093] For example, the first token corresponding to the video data format can be <vd>, the second token identifier corresponding to the video data format can be <\VD>, then the modified second input question can be: <vd>AAA<\VD>, <vd>Indicates the starting position of the second input question, AAA indicates the second input question in text data format, and <\VD> indicates the ending position of the second input question.
[0094] In one example, for the first input question in text data format, the first input question in text data format can be directly input into the target large language model, and there is no restriction on this process.
[0095] Step 307: Output a first target answer in text data format through the target large language model.
[0096] In an example, if the input to the target large language model is AAA<\IMG>, then the target large language model can The data content between <\IMG> is used as the input question (i.e., the second input question). Then, the target large language model obtains the first target answer in the text data format corresponding to the input question through analysis and reasoning, and outputs the first target answer in the text data format.
[0097] When the target large language model outputs the first target answer, the first target answer has a first token identifier in front and a second token identifier behind, and the first token identifier indicates the start position of the first target answer, and the second token identifier indicates the end position of the first target answer. For example, the target large language model outputs BBB<\IMG>, Indicates the starting position of the first target answer, <\IMG> indicates the ending position of the first target answer, and BBB indicates the first target answer in text data format.
[0098] In an example, if the input to the target large language model is <speech>AAA<\SPEECH>, then the target large language model can <speech>The data content between <\SPEECH> is used as the input question (i.e., the second input question). Then, the target large language model obtains the first target answer in the text data format corresponding to the input question through analysis and reasoning, and outputs the first target answer in the text data format.
[0099] When the target large language model outputs the first target answer, the first target answer has a first token identifier in front and a second token identifier behind. The first token identifier indicates the start position of the first target answer, and the second token identifier indicates the end position of the first target answer. <speech>BBB<\SPEECH>, <speech>Indicates the starting position of the first target answer, <\SPEECH> indicates the ending position of the first target answer, and BBB indicates the first target answer.
[0100] In an example, if the input to the target large language model is <vd>AAA<\VD>, then the target large language model can <vd>The data content between <\VD> is used as the input question (i.e., the second input question). Then, the target large language model obtains the first target answer in the text data format corresponding to the input question through analysis and reasoning, and outputs the first target answer in the text data format.
[0101] When the target large language model outputs the first target answer, the first target answer has a first token identifier in front and a second token identifier behind, and the first token identifier indicates the start position of the first target answer, and the second token identifier indicates the end position of the first target answer. For example, the target large language model outputs <vd>BBB<\VD>, <vd>Indicates the starting position of the first target answer, <\VD> indicates the ending position of the first target answer, and BBB indicates the first target answer in text data format.
[0102] In an example, if the input to the target large language model is AAA, the target large language model can use AAA as the input question (i.e., the first input question). The target large language model analyzes and infers the first target answer in the text data format corresponding to the input question and outputs the first target answer.
[0103] In one example, when the target large language model receives an input question in text data format (such as the first input question or the second input question), it can perform a Tokenizer conversion on the input question to obtain Embedding data, perform analysis and reasoning based on the Embedding data to obtain an inference result, and perform a DeTokenizer conversion on the inference result to obtain a first target answer in text data format.
[0104] Step 308: Convert the first target answer into a second target answer in the original data format.
[0105] In an example, if the target large language model outputs BBB<\IMG>, then based on the data content output by the target large language model, the data content (BBB) between the first token identifier and the second token identifier can be determined as the first target answer. Then, the original data format corresponding to the first token identifier or the second token identifier is determined, that is, Alternatively, <\IMG> corresponds to an image data format. Then, the first target answer can be converted into a second target answer in an image data format.
[0106] For example, since the original data format is an image data format, a text-to-image conversion algorithm is used to convert the first target answer in text data format into a second target answer in image data format.
[0107] For example, text-to-image conversion algorithms may include, but are not limited to, the SD (Stable Diffusion) algorithm, which can be used to convert a first target answer in text data format into a second target answer in image data format. Of course, the SD algorithm is merely an example and is not limiting. Any algorithm capable of converting a first target answer in text data format into a second target answer in image data format is sufficient.
[0108] For example, an example of the SD algorithm can be shown in FIG4C . The SD algorithm may include a Text Model, a UNet network, an autoencoder, and the like.
[0109] During use, the text to be detected (i.e., the first target answer in text data format, i.e., the first target answer is text) can be provided to the Text Model to obtain text embeddings (text embedding features). The text embeddings are input to the UNet network and autoencoder to obtain the second target answer in image data format, i.e., the second target answer uses an image to describe the first target answer.
[0110] In an example, if the target large language model outputs <speech>BBB<\SPEECH>, then based on the data content output by the target large language model, the data content (BBB) between the first token identifier and the second token identifier can be determined as the first target answer. Then, the original data format corresponding to the first token identifier or the second token identifier is determined, that is, <speech>Or <\SPEECH> corresponds to the voice data format. Then, the first target answer can be converted into a second target answer in the voice data format.
[0111] For example, since the original data format is a voice data format, a text-to-voice conversion algorithm is used to convert the first target answer in text data format into a second target answer in voice data format.
[0112] For example, text-to-speech conversion algorithms (also known as speech synthesis algorithms) may include, but are not limited to, TTS (Text-To-Speech) algorithms, Azure TTS algorithms, and Amazon Polly algorithms. These algorithms can be used to convert a first target answer in text data format into a second target answer in speech data format. Of course, these algorithms are merely examples and are not limiting. Any algorithm that can convert a first target answer in text data format into a second target answer in speech data format is sufficient.
[0113] In an example, if the target large language model outputs <vd>BBB<\VD>, then based on the data content output by the target large language model, the data content (BBB) between the first token identifier and the second token identifier can be determined as the first target answer. Then, the original data format corresponding to the first token identifier or the second token identifier can be determined, that is, <vd>Or <\VD> corresponds to the video data format. Then, the first target answer can be converted into a second target answer in the video data format.
[0114] For example, since the original data format is a video data format, a text-to-video conversion algorithm is used to convert the first target answer in text data format into a second target answer in video data format.
[0115] For example, a text-to-video conversion algorithm can be used to convert a first target answer in text data format into a second target answer in video data format, with no restrictions on this conversion process. Alternatively, a text-to-image conversion algorithm can be used to convert the first target answer in text data format into multiple frames of images, which can then be combined into the second target answer in video data format.
[0116] At this point, the second target answer in the original data format can be obtained, completing the analysis and reasoning process.
[0117] In an example, the training process for the target large language model may include but is not limited to:
[0118] A sample dataset is obtained, which may include multiple sample questions and sample labels corresponding to each sample question (i.e., sample answers corresponding to the sample questions). The sample questions may be in text data format, and the sample labels may be in text data format. An initial large language model to be trained is trained based on the sample questions (e.g., multiple sample questions) in the sample dataset and the sample labels corresponding to the sample questions to obtain a trained candidate large language model. This embodiment does not limit this training process.
[0119] After obtaining the candidate large language model, the candidate large language model can be used as the target large language model.
[0120] Alternatively, after obtaining the candidate large language model, the candidate large language model may be fine-tuned so that the candidate large language model supports token identification, thereby obtaining a fine-tuned target large language model.
[0121] During the fine-tuning process of the candidate large language model, some sample questions can be selected from the sample dataset as reference sample questions. For each reference sample question, a first token identifier of any data format (such as image data format, voice data format, video data format) is added to the front of the reference sample question, and a second token identifier of the data format is added to the back of the reference sample question.
[0122] For example, for some reference sample questions, add , add <\IMG> after these reference sample questions. For some reference sample questions, add <speech>, add <\SPEECH> after these reference sample questions. For some reference sample questions, add <vd>, add <\VD> after these reference sample questions. Based on the above processing, the modified reference sample questions can be obtained.
[0123] The candidate large language model is trained based on the modified reference sample question and the sample label corresponding to the reference sample question to obtain a target large language model. This embodiment does not limit this training process.
[0124] In one example, when the initial large language model is trained to obtain a candidate large language model, the input and output of the initial large language model are both text data, eliminating the Adaptor modules (linear mapping layers for alignment) on the input and output sides, thus avoiding tedious input-side alignment and output-side alignment training.
[0125] After training to obtain a candidate large language model, fine-tune the candidate large language model to obtain a target large language model, as long as the target large language model supports token identification.
[0126] It can be seen from the above technical solution that in the embodiment of the present application, the target answer is obtained with the help of the target large language model, and the accurate and reliable target answer is obtained by inputting the input question in the text data format. It only takes a small amount of computing resources to query the target answer corresponding to the input question from a large amount of data. The time to obtain the target answer is shorter, thereby saving the computing resources of the server and saving time overhead. It can quickly determine the target answer corresponding to the input question, provide logical analysis and reasoning capabilities, and use the analysis capabilities of the target large language model to give accurate and reliable target answers. Moreover, by converting multimodal data (such as pictures, videos, voice and other data) into text data and inputting it into the target large language model, the target large language model supports multimodal data. The target large language model only needs to support the input and output of text data, thereby meeting the input and output requirements of multimodal data and supporting the functions of the multimodal large language model.
[0127] The target large language model directly processes text data, eliminating the linear mapping layer on the input and output sides. There is no need for alignment operations between the input side and the target large language model, or between the output side and the target large language model. There is no need for tedious alignment training, which can greatly simplify the training process of the target large language model and meet more business needs.
[0128] Based on the same inventive concept, a data processing device and electronic device based on a large language model corresponding to the above-mentioned data processing method based on a large language model are also provided. Since the principles of solving problems by the data processing device and the electronic device are similar to those of the data processing method based on a large language model, the implementation of the data processing device and the electronic device can refer to the data processing method based on a large language model, and the repeated parts will not be repeated.
[0129] Based on the same application concept as the above method, an embodiment of the present application proposes a data processing device based on a large language model. FIG5 is a schematic diagram of the structure of the device, which may include:
[0130] An acquisition module 51 is configured to acquire a first input question in an original data format; if the original data format is not a text data format, convert the first input question into a second input question in a text data format;
[0131] A processing module 52 is configured to input the second input question into a trained target large language model, and have the target large language model output a first target answer in a text data format corresponding to the second input question;
[0132] The conversion module 53 is configured to convert the first target answer into a second target answer in an original data format.
[0133] In one example, when the acquisition module 51 converts the first input question into the second input question in text data format, it is specifically used to: if the original data format is an image data format, then use an image-to-text conversion algorithm to convert the first input question in image data format into the second input question in text data format. Alternatively, if the original data format is a voice data format, then use a voice-to-text conversion algorithm to convert the first input question in voice data format into the second input question in text data format. Alternatively, if the original data format is a video data format, then use a video-to-text conversion algorithm to convert the first input question in video data format into the second input question in text data format.
[0134] In one example, when the conversion module 53 converts the first target answer into the second target answer in the original data format, it is specifically used to: if the original data format is an image data format, then a text-to-image conversion algorithm is used to convert the first target answer in the text data format into the second target answer in the image data format. Alternatively, if the original data format is a voice data format, then a text-to-voice conversion algorithm is used to convert the first target answer in the text data format into the second target answer in the voice data format. Alternatively, if the original data format is a video data format, then a text-to-video conversion algorithm is used to convert the first target answer in the text data format into the second target answer in the video data format.
[0135] In one example, the processing module 52 inputs the second input question to the trained target large language model, and the target large language model outputs the first target answer in the text data format corresponding to the second input question, which is specifically used to: obtain the first token identifier and the second token identifier corresponding to the original data format. Add the first token identifier in front of the second input question, and add the second token identifier after the second input question. The first token identifier indicates the starting position of the second input question, and the second token identifier indicates the ending position of the second input question. Input the modified second input question to the target large language model, and the target large language model outputs the first target answer in the text data format corresponding to the second input question.
[0136] The first target answer is preceded by the first token identifier and followed by the second token identifier. The first token identifier indicates the starting position of the first target answer, and the second token identifier indicates the ending position of the first target answer.
[0137] In one example, when the conversion module 53 converts the first target answer into the second target answer in the original data format, it is specifically used to: determine the data content between the first token identifier and the second token identifier as the first target answer based on the data content output by the target large language model; determine the original data format corresponding to the first token identifier or the second token identifier; and convert the first target answer into the second target answer in the original data format.
[0138] In one example, the apparatus further includes: a training module for training to obtain the target large language model. Specifically, when the training module trains to obtain the target large language model, it is used to:
[0139] Obtain a sample data set, the sample data set including a plurality of sample questions and a sample label corresponding to each sample question, the sample question being in a text data format, and the sample label being in a text data format;
[0140] Training the initial large language model to be trained based on the sample questions in the sample data set and the sample labels corresponding to the sample questions to obtain a trained candidate large language model;
[0141] Using the candidate large language model as the target large language model; or,
[0142] Select some sample questions from the sample data set as reference sample questions, add a first token identifier in any data format in front of the reference sample question, and add a second token identifier in the data format after the reference sample question; train the candidate large language model based on the modified reference sample question and the sample label corresponding to the reference sample question to obtain the target large language model.
[0143] Based on the same concept as the above method, an electronic device is proposed in an example of the present application, as shown in Figure 6, the electronic device includes a processor 611 and a machine-readable storage medium 612, and the machine-readable storage medium 612 stores machine-executable instructions that can be executed by the processor 611; the processor 611 is used to execute the machine-executable instructions to implement the data processing method based on the large language model disclosed in the above example.
[0144] In one example, the processor 611 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 611 may be implemented in at least one hardware form selected from the group consisting of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array).
[0145] The processor 611 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in the standby state.
[0146] In some embodiments, the processor 611 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen.
[0147] In one example, the electronic device may optionally include a peripheral device interface 613 and at least one peripheral device. The processor 611 and the peripheral device interface 613 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 613 via a bus, signal lines, or circuit boards. The peripheral device may include at least one of a radio frequency circuit 614 and a power supply 615.
[0148] The RF circuit 614 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 614 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 614 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. The RF circuit 614 may optionally include an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a subscriber identity module card, and the like.
[0149] The radio frequency circuit 614 can communicate with the user equipment via at least one wireless communication protocol, including but not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network.
[0150] The power supply 615 is used to supply power to various components in the electronic device. The power supply 615 can be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0151] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the data processing method based on the large language model disclosed in the above example of the present application can be implemented.
[0152] The machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.
[0153] The systems, devices, modules, or units described in the above embodiments may be implemented by a computer entity or by a product having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.
[0154] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0155] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0156] The present application is described with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0157] Moreover, these computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0159] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.< / vd> < / speech> < / vd> < / vd> < / speech> < / speech> < / vd> < / vd> < / vd> < / vd> < / speech> < / speech> < / speech> < / speech> < / vd> < / vd> < / vd> < / speech> < / speech> < / speech>
Claims
1. A data processing method based on a large language model, characterized in that, The method includes: Obtain a first input question in an original data format; if the original data format is not a text data format, convert the first input question into a second input question in a text data format; Input the second input question into a trained target large language model, and the target large language model outputs a first target answer in the text data format corresponding to the second input question; Convert the first target answer into a second target answer in the original data format.
2. The method according to claim 1, wherein The conversion of the first input question into a second input question in a text data format includes: If the original data format is a picture data format, use an algorithm for converting pictures and text to convert the first input question in the picture data format into a second input question in the text data format; or, If the original data format is a voice data format, use an algorithm for converting voice and text to convert the first input question in the voice data format into a second input question in the text data format; or, If the original data format is a video data format, use an algorithm for converting video and text to convert the first input question in the video data format into a second input question in the text data format.
3. The method according to claim 1 or 2, wherein The conversion of the first target answer into a second target answer in the original data format includes: If the original data format is a picture data format, use an algorithm for converting text and pictures to convert the first target answer in the text data format into a second target answer in the picture data format; or, If the original data format is a voice data format, use an algorithm for converting text and voice to convert the first target answer in the text data format into a second target answer in the voice data format; or, If the original data format is a video data format, use an algorithm for converting text and video to convert the first target answer in the text data format into a second target answer in the video data format.
4. The method according to claim 1, wherein The inputting of the second input question into a trained target large language model, and the target large language model outputting a first target answer in the text data format corresponding to the second input question includes: Obtain a first token identifier and a second token identifier corresponding to the original data format; Add the first token identifier in front of the second input question and add the second token identifier behind the second input question; wherein, the first token identifier represents the start position of the second input question, and the second token identifier represents the end position of the second input question; Input the modified second input question into the target large language model, and the target large language model outputs a first target answer in the text data format corresponding to the second input question; Among them, the first token identifier is in front of the first target answer, and the second token identifier is behind the first target answer. The first token identifier represents the start position of the first target answer, and the second token identifier represents the end position of the first target answer.
5. The method according to claim 4, wherein the converting the first target answer into a second target answer in the original data format includes: Based on the data content output by the target large language model, determining the data content between the first token identifier and the second token identifier as the first target answer; determining the original data format corresponding to the first token identifier or the second token identifier; converting the first target answer into a second target answer in the original data format.
6. The method according to claim 1, wherein for the training process of the target large language model, it includes: obtaining a sample data set, the sample data set includes a plurality of sample questions and sample labels corresponding to each sample question, the sample questions are in text data format, and the sample labels are in text data format; training an initial large language model to be trained based on the sample questions in the sample data set and the sample labels corresponding to the sample questions to obtain a trained candidate large language model; using the candidate large language model as the target large language model; or, selecting some sample questions from the sample data set as reference sample questions, adding a first token identifier in any data format in front of the reference sample questions, and adding a second token identifier in this data format behind the reference sample questions; training the candidate large language model based on the modified reference sample questions and the sample labels corresponding to the reference sample questions to obtain the target large language model.
7. A data processing device based on a large language model, characterized in that, The device includes: an acquisition module, configured to acquire a first input question in the original data format; if the original data format is not in text data format, convert the first input question into a second input question in text data format; a processing module, configured to input the second input question into a trained target large language model, and the target large language model outputs a first target answer in text data format corresponding to the second input question; a conversion module, configured to convert the first target answer into a second target answer in the original data format.
8. The device according to claim 7, characterized in that When the acquisition module converts the first input question into a second input question in text data format, it specifically uses: if the original data format is in picture data format, then using an algorithm for converting pictures and text to convert the first input question in picture data format into a second input question in text data format; or, if the original data format is in voice data format, then using an algorithm for converting voice and text to convert the first input question in voice data format into a second input question in text data format; or, If the original data format is a video data format, a video-to-text conversion algorithm is adopted to convert the first input problem in the video data format into a second input problem in the text data format.
9. The device according to claim 7 or 8, characterized in that, When the conversion module converts the first target answer into a second target answer in the original data format, it specifically is used for: If the original data format is a picture data format, a text-to-picture conversion algorithm is adopted to convert the first target answer in the text data format into a second target answer in the picture data format; Or, If the original data format is a voice data format, a text-to-voice conversion algorithm is adopted to convert the first target answer in the text data format into a second target answer in the voice data format; Or, If the original data format is a video data format, a text-to-video conversion algorithm is adopted to convert the first target answer in the text data format into a second target answer in the video data format.
10. The device according to claim 7, characterized in that When the processing module inputs the second input problem to the trained target large language model, and the target large language model outputs the first target answer in the text data format corresponding to the second input problem, it specifically is used for: Obtaining a first token identifier and a second token identifier corresponding to the original data format; Adding the first token identifier in front of the second input problem and adding the second token identifier behind the second input problem; wherein, the first token identifier represents the start position of the second input problem position, and the second token identifier represents the end position of the second input problem; Inputting the modified second input problem to the target large language model, and the target large language model outputs the first target answer in the text data format corresponding to the second input problem; Wherein, the first token identifier is in front of the first target answer, the second token identifier is behind the first target answer, the first token identifier represents the start position of the first target answer, and the second token identifier represents the end position of the first target answer.
11. The device according to claim 10, characterized in that, When the conversion module converts the first target answer into a second target answer in the original data format, it specifically is used for: Based on the data content output by the target large language model, determining the data content between the first token identifier and the second token identifier as the first target answer; Determining the original data format corresponding to the first token identifier or the second token identifier; Converting the first target answer into a second target answer in the original data format.
12. The device according to claim 7, characterized in that, It further includes: A training module for training to obtain the target large language model; Wherein, when the training module trains to obtain the target large language model, it specifically is used for: Obtaining a sample data set, the sample data set includes a plurality of sample problems and sample labels corresponding to each sample problem, the sample problems are in the text data format, and the sample labels are in the text data format; Training an initial large language model to be trained based on the sample problems in the sample dataset and the sample labels corresponding to the sample problems to obtain a trained candidate large language model; Using the candidate large language model as the target large language model; or, Selecting some sample problems from the sample dataset as reference sample problems, adding a first token identifier in any data format before the reference sample problems, and adding a second token identifier in the data format after the reference sample problems; training the candidate large language model based on the modified reference sample problems and the sample labels corresponding to the reference sample problems to obtain the target large language model.
13. An electronic device, characterized in that, Including: A processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is configured to execute the machine-executable instructions to implement the method according to any one of claims 1-7.
Citation Information
Patent Citations
Data query method and device, computer equipment and storage medium
CN116842036A
Question and answer method and device, electronic equipment and storage medium
CN117056483A
Question text generation method and device of large language model, equipment and medium
CN117093696A
Monolingual conversion device
JP2022164001A