Artificial intelligence dialogue method in freight field, electronic equipment and readable storage medium
By using the source Q&A seeds and historical dialogue information in the freight field, and fine-tuning the dialogue model with user portrait data, the problem of the artificial intelligence dialogue model deviating from the real scene is solved, and accurate reply and efficient communication are achieved.
Patent Information
- Application Number
- CN202510408501.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
The current artificial intelligence dialogue model is prone to deviating from the real dialogue scenarios in the freight field, resulting in model hallucinations and poor user experience.
By obtaining multiple source Q&A seeds and historical dialogue information, using a large language pre-trained model to generate dialogue generation models, generate multiple dialogue information, and combine user portrait data to fine-tune the dialogue model to ensure the accuracy of the reply.
It improves the accuracy of the artificial intelligence dialogue model in the freight field, reduces model illusion, improves communication efficiency between drivers and shippers, and reduces the cost of manual labeling.
Smart Images

Figure CN120336472A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of intelligent freight transportation technology, and in particular, to an artificial intelligence dialogue method, an electronic device, and a readable storage medium in the field of freight transportation. Background Art
[0002] In the intelligent freight transportation scenario, for the transportation demand form proposed by the cargo owner, after the driver determines that they can undertake the transportation demand form, the cargo owner and the driver reach an order online and carry out subsequent transportation. The above process can be implemented through a freight application (app).
[0003] When the driver browses the transportation demand form, they may need to inquire about the source of goods information, and the cargo owner usually does not reply online in real time. Therefore, an artificial intelligence (AI) dialogue model can be used to communicate with the driver.
[0004] However, the current reply of the AI dialogue model is prone to deviate from the real dialogue scenario, resulting in the phenomenon of model hallucination and poor user experience. Summary of the Invention
[0005] The present disclosure provides an artificial intelligence dialogue method, an electronic device, and a readable storage medium in the field of freight transportation to solve the problem that the reply of the current artificial intelligence dialogue model deviates from the real dialogue scenario.
[0006] In a first aspect, the present disclosure provides an artificial intelligence dialogue method in the field of freight transportation, and the method includes:
[0007] Receiving a first message sent by a terminal device; the first message includes a source of goods problem raised for a first source of goods;
[0008] Inputting the source of goods problem, the information of the first source of goods, and the context information of the source of goods problem into an artificial intelligence dialogue model to obtain a source of goods reply corresponding to the source of goods problem;
[0009] Among them, the artificial intelligence dialogue model is obtained by adjusting the initial artificial intelligence dialogue model based on multiple sample freight conversations; the initial artificial intelligence dialogue model is established based on a large language pre-trained model; the multiple sample freight conversations are obtained by inputting the first source of goods Q&A information, the source of goods Q&A seeds, multiple dialogue process data, and the first prompt information into a dialogue generation model; the first prompt information is used to indicate generating multiple sample freight conversations according to the first source of goods Q&A information, the source of goods Q&A seeds, and the multiple dialogue process data; the first source of goods Q&A information is obtained by performing a Q&A generation operation according to the obtained multiple source of goods Q&A seeds and the intention information corresponding to each source of goods Q&A seed; each source of goods Q&A seed contains a source of goods question seed and a source of goods answer seed; the first source of goods Q&A information contains a first source of goods question and a first source of goods answer; the intention information corresponding to the source of goods Q&A seed is used to indicate the questioning intention of the source of goods question seed in the source of goods Q&A seed; the multiple dialogue process data are generated according to multiple source of goods historical conversation information, and each source of goods historical conversation information contains multi-round Q&A information.
[0010] Send the source of goods answer corresponding to the source of goods question to the terminal device.
[0011] In some embodiments, the method further includes:
[0012] Obtain multiple source of goods Q&A seeds; each source of goods Q&A seed contains a source of goods question seed and a source of goods answer seed;
[0013] Perform a Q&A generation operation according to the multiple source of goods Q&A seeds and the intention information corresponding to each source of goods Q&A seed to obtain the first source of goods Q&A information; the first source of goods Q&A information contains a first source of goods question and a first source of goods answer; the intention information corresponding to the source of goods Q&A seed is used to indicate the questioning intention of the source of goods question seed in the source of goods Q&A seed;
[0014] Generate multiple dialogue process data according to multiple source of goods historical conversation information; each source of goods historical conversation information contains multi-round Q&A information;
[0015] Input the first source of goods Q&A information, the source of goods Q&A seeds, the multiple dialogue process data, and the first prompt information into a dialogue generation model to obtain multiple dialogue information; the first prompt information is used to indicate generating multiple sample freight conversations according to the first source of goods Q&A information, the source of goods Q&A seeds, and the multiple dialogue process data; the dialogue generation model is established based on a large language pre-trained model.
[0016] In some embodiments, the performing a Q&A generation operation according to the multiple source of goods Q&A seeds and the intention information corresponding to each source of goods Q&A seed to obtain the first source of goods Q&A information includes:
[0017] Store the multiple source supply Q&A seeds in a task pool;
[0018] Obtain all source supply Q&A in the task pool as second source supply Q&A information; the second source supply Q&A information includes second source supply questions and second source supply answers;
[0019] Identify the questioning intentions of the second source supply questions respectively to obtain the intention information of each second source supply question;
[0020] Perform question generation operations according to each second source supply question and the intention information of each second source supply question to obtain first source supply questions;
[0021] Determine the first source supply responses corresponding to each first source supply question;
[0022] Take each first source supply question and the first source supply response corresponding to the first source supply question as third source supply Q&A information;
[0023] Store the third source supply question information in the task pool;
[0024] Input the first source supply Q&A information, the source supply Q&A seeds, the multiple conversation flow data, and the first prompt information into a conversation generation model to obtain multiple conversation information, including:
[0025] Obtain all source supply Q&A information in the task pool from the task pool as fourth source supply Q&A information;
[0026] Input the fourth source supply Q&A information, the multiple conversation flow data, and the first prompt information into a conversation generation model to obtain multiple conversation information.
[0027] In some embodiments, the step of taking each first source supply question and the first source supply response corresponding to the first source supply question as third source supply Q&A information includes:
[0028] Based on all first source supply questions, the first source supply responses corresponding to the first source supply questions, and the second source supply Q&A information in the task pool, perform duplicate removal processing on all first source supply questions and the first source supply responses corresponding to the first source supply questions to obtain third source supply Q&A information.
[0029] In some embodiments, after storing the third source supply question information in the task pool, it further includes:
[0030] Judge whether the quantity of the third source supply Q&A information is greater than or equal to a preset quantity value;
[0031] In the case where the number of the third source question-and-answer information is greater than or equal to the preset number value, return to execute obtaining all the source question-and-answer in the task pool as the second source question-and-answer information until the number of the third source question-and-answer information is less than the preset number value.
[0032] In some embodiments, inputting the first source question-and-answer information, the source question-and-answer seed, the multiple dialogue process data, and the first prompt information into a dialogue generation model to obtain multiple dialogue information, including:
[0033] Input the first source question-and-answer information, the source question-and-answer seed, the multiple dialogue process data, user portrait data, and the first prompt information into a dialogue generation model to obtain multiple dialogue information; the first prompt information is specifically used to indicate generating multiple sample freight conversations according to the first source question-and-answer information, the source question-and-answer seed, the user portrait data, and the multiple dialogue process data; the user portrait data at least includes at least one of the following: user basic information, user language style information, user personality characteristic information, user consumption habit information, and user social attribute information.
[0034] In some embodiments, after inputting the first source question-and-answer information, the source question-and-answer seed, the multiple dialogue process data, and the first prompt information into a dialogue generation model to obtain multiple dialogue information, it further includes:
[0035] Input the multiple dialogue information and a preset quality inspection standard into a dialogue quality inspection model to obtain a quality inspection result of the multiple dialogue information;
[0036] Use the multiple dialogue information indicated by the quality inspection result to pass the quality inspection to update the multiple dialogue information.
[0037] In a second aspect, an embodiment of the present disclosure provides a method for generating data in the freight field, and the method includes:
[0038] Obtain multiple source question-and-answer seeds; each source question-and-answer seed contains a source question seed and a source answer seed;
[0039] Perform a question-and-answer generation operation according to the multiple source question-and-answer seeds and the intention information corresponding to each source question-and-answer seed to obtain first source question-and-answer information; the first source question-and-answer information contains a first source question and a first source answer; the intention information corresponding to the source question-and-answer seed is used to indicate the questioning intention of the source question seed in the source question-and-answer seed.
[0040] Generate multiple dialogue process data according to multiple source historical dialogue information; each source historical dialogue information contains multi-round question-and-answer information;
[0041] Input the first source question-and-answer information, the source question-and-answer seeds, the multiple conversation flow data, and the first prompt information into a conversation generation model to obtain multiple conversation messages; the first prompt information is used to indicate generating multiple sample freight conversations according to the first source question-and-answer information, the source question-and-answer seeds, and the multiple conversation flow data; the conversation generation model is established based on a large language pre-trained model.
[0042] In some embodiments, the obtaining the first source question-and-answer information by performing a question-and-answer generation operation according to the multiple source question-and-answer seeds and the intent information corresponding to each source question-and-answer seed includes:
[0043] Store the multiple source question-and-answer seeds in a task pool;
[0044] Obtain all the source question-and-answer in the task pool as the second source question-and-answer information; the second source question-and-answer information includes second source questions and second source answers;
[0045] Identify the question intent of each second source question respectively to obtain the intent information of each second source question;
[0046] Perform a question generation operation according to each second source question and the intent information of each second source question to obtain the first source question;
[0047] Determine the first source reply corresponding to each first source question;
[0048] Take each first source question and the first source reply corresponding to the first source question as the third source question-and-answer information;
[0049] Store the third source question information in the task pool;
[0050] The inputting the first source question-and-answer information, the source question-and-answer seeds, the multiple conversation flow data, and the first prompt information into a conversation generation model to obtain multiple conversation messages includes:
[0051] Obtain all the source question-and-answer information in the task pool from the task pool as the fourth source question-and-answer information;
[0052] Input the fourth source question-and-answer information, the multiple conversation flow data, and the first prompt information into a conversation generation model to obtain multiple conversation messages.
[0053] In some embodiments, the taking each first source question and the first source reply corresponding to the first source question as the third source question-and-answer information includes:
[0054] Based on all the first source problems, the first source replies corresponding to the first source problems, and the second source Q&A information in the task pool, duplicate removal is performed on all the first source problems and the first source replies corresponding to the first source problems to obtain the third source Q&A information.
[0055] In some embodiments, after storing the third source problem information in the task pool, the following is further included:
[0056] Determine whether the quantity of the third source Q&A information is greater than or equal to a preset quantity value;
[0057] When the quantity of the third source Q&A information is greater than or equal to the preset quantity value, return to execute obtaining all the source Q&A in the task pool as the second source Q&A information until the quantity of the third source Q&A information is less than the preset quantity value.
[0058] In some embodiments, inputting the first source Q&A information, the source Q&A seeds, the multiple dialogue process data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information includes:
[0059] Input the first source Q&A information, the source Q&A seeds, the multiple dialogue process data, the user portrait data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information; the first prompt information is specifically used to indicate generating multiple sample freight conversations according to the first source Q&A information, the source Q&A seeds, the user portrait data, and the multiple dialogue process data; the user portrait data at least includes at least one of the following: user basic information, user language style information, user personality characteristic information, user consumption habit information, and user social attribute information.
[0060] In some embodiments, after inputting the first source Q&A information, the source Q&A seeds, the multiple dialogue process data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information, the following is further included:
[0061] Input the multiple dialogue information and the preset quality inspection standard into the dialogue quality inspection model to obtain the quality inspection results of the multiple dialogue information;
[0062] Use the quality inspection results to indicate the multiple dialogue information that passes the quality inspection to update the multiple dialogue information.
[0063] In a third aspect, the present disclosure provides an electronic device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and when the processor executes the program, the method described in the first aspect above is implemented.
[0064] Fourthly, the present disclosure provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, the method described in the second aspect above is implemented.
[0065] Fifthly, the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0066] Sixthly, the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the second aspect above is implemented.
[0067] The artificial intelligence dialogue method, electronic device, and readable storage medium provided by the embodiments of the present disclosure in the freight field. The server obtains a plurality of source-of-goods Q&A seeds, where each source-of-goods Q&A seed contains a source-of-goods question seed and a source-of-goods answer seed. According to the plurality of source-of-goods Q&A seeds and the intention information indicating the questioning intention of each source-of-goods Q&A seed, a Q&A generation operation is performed to obtain the first source-of-goods Q&A information. Thus, based on fewer source-of-goods Q&A seeds and referring to the questioning intention of the source-of-goods Q&A seeds, more first source-of-goods Q&A information in different scenarios is generated. The first source-of-goods Q&A information is in the form of question and answer. According to a plurality of source-of-goods historical dialogue information containing multi-round Q&A information, a plurality of dialogue process data are generated. Thus, according to the source-of-goods historical dialogue information in the real scenario, dialogue process data that conform to different scenarios are obtained. The first source-of-goods Q&A information, the source-of-goods Q&A seeds, the plurality of dialogue process data, and the first prompt information are input into a dialogue generation model established based on a large language pre-trained model to obtain a plurality of dialogue information. Thus, through a small number of source-of-goods Q&A seeds, a large number of diverse and authentic dialogue information can be quickly generated, reducing labor costs. In the freight field, a large amount of dialogue information is automatically obtained for subsequent fine-tuning of the initial artificial intelligence dialogue model, so as to obtain an artificial intelligence dialogue model that conforms to the freight field. In practical applications, the artificial intelligence dialogue model can give accurate answers to the inquiries of drivers, reduce the problem of model hallucinations, and improve the communication efficiency with drivers. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is an architecture diagram of an artificial intelligence dialogue system in the freight field provided by the embodiments of the present disclosure;
[0069] Figure 2 is a flowchart of a data generation method in the freight field provided by the embodiments of the present disclosure;
[0070] Figure 3 is a flowchart of a method for adjusting an artificial intelligence dialogue model provided by the embodiments of the present disclosure;
[0071] Figure 4 A flowchart of a method for adjusting an artificial intelligence dialogue model provided by an embodiment of the present disclosure;
[0072] Figure 5 A schematic diagram of the principle of a data generation method in the freight field provided by an embodiment of the present disclosure. Detailed implementation manners
[0073] In the present disclosure, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a alone, b alone, or c alone can represent: a alone, b alone, c alone, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b, and c, where a, b, and c can be single or multiple. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0074] The orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "upper", "lower", "left", "right", "front", "rear", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present disclosure.
[0075] The terms "connected" and "coupled" should be understood in a broad sense. For example, the "connection" or "coupling" of a circuit structure can refer not only to a physical connection but also to an electrical connection or a signal connection. For example, it can be a direct connection, that is, a physical connection, or it can be indirectly connected through at least one intermediate element, as long as the circuit is electrically connected. It can also be the internal connection of two elements; the signal connection can refer not only to the signal connection through a circuit but also to the signal connection through a media medium. For example, radio waves. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific situations.
[0076] In the intelligent transportation scenario, the cargo owner can fill in the freight request information on the freight application of the terminal device. The freight source information can include: loading location information, unloading location information, required vehicle type, required vehicle length, price, cargo information, etc. The freight application of the terminal device used by the cargo owner can be called the cargo owner terminal, and the cargo owner terminal sends the freight request information as a freight source to the corresponding server. The driver can also perform operations such as searching on the freight application of the terminal device. The freight application of the terminal device used by the driver can be called the driver terminal, and the server sends the current freight source to the driver terminal for display. After seeing the displayed freight source, the driver can inquire with the cargo owner about the freight source information. However, usually the cargo owner does not answer the driver's inquiries in real time online. Therefore, AI can be used to reply to the driver's inquiries based on the freight source information to improve communication efficiency. After the driver determines the freight source information, the driver can accept the order, and the cargo owner and the driver can sign the order to carry out subsequent transportation.
[0077] An artificial intelligence dialogue method in the freight field provided by this application can be applied to an artificial intelligence dialogue system in the freight field. Please refer to Figure 1 , Figure 1 FIG. is an architecture diagram of an artificial intelligence dialogue system in the freight field provided by an embodiment of the present disclosure. The artificial intelligence dialogue system in the freight field includes a server 102 and a terminal device 101. The server 102 and the terminal device 101 are connected through a network. Among them, the server 102 can be a single server or a server cluster composed of two or more servers. The terminal device 101 can include but is not limited to a terminal device, a terminal, a user equipment (UE), a mobile station, a mobile terminal, etc. The terminal device 101 can specifically be a mobile phone, a tablet computer, a computer with wireless transceiver function, a wearable device, a robot, a smart home device, etc.
[0078] A freight application can be installed in the terminal device 101. The driver user can log in to the freight application to view the current freight sources and inquire with the server 102 about the relevant freight source information for any current freight source. The server 102 obtains the corresponding reply through the artificial intelligence dialogue model and sends the corresponding reply to the terminal device 101. The above process of inquiring about the freight source can be carried out in text form, or in voice form, or in other forms. The present disclosure does not make any limitations on this.
[0079] The artificial intelligence dialogue model is pre-stored in the server 102. For the artificial intelligence dialogue model in the freight field, usually a large amount of high-quality data in the freight field is required to adjust the initial artificial intelligence dialogue model. Here, the adjustment can also be called fine-tuning.
[0080] Generally, data in the freight transportation field requires manual annotation of a large amount of data. However, there are many problems with manual annotation. The problems with manual annotation are introduced below by way of example.
[0081] Problem 1: Slow speed.
[0082] Annotation tasks usually involve a large amount of text data. Annotators need to browse, understand, and mark the data item by item, which not only requires a lot of time but also a high degree of concentration and a process of repeated verification. Therefore, when generating large-scale and high-quality data sets, the speed of manual annotation is slow.
[0083] Problem 2: High cost.
[0084] Annotation of large-scale data sets usually requires a large number of annotators. Especially in scenarios that require high precision and complex annotation tasks, it becomes even more necessary to hire professionals with relevant domain knowledge, and the corresponding training costs also need to be considered before the task.
[0085] Problem 3: Uneven quality.
[0086] In the case of limited concentration, human error is inevitable, and it is also affected by the subjective differences between different annotators and the differences caused by the uneven skill levels of different annotators. There are also certain fluctuations in quality inspection. Therefore, it is difficult to ensure high-quality output of the final data.
[0087] To solve the problems of slow speed and high labor cost caused by the above-mentioned manually labeled data, the present disclosure provides a method for generating data in the freight transportation field. The server obtains a plurality of source-of-goods question-and-answer seeds, where each source-of-goods question-and-answer seed contains a source-of-goods question seed and a source-of-goods answer seed. According to the plurality of source-of-goods question-and-answer seeds and the intention information indicating the questioning intention of each source-of-goods question-and-answer seed, a question-and-answer generation operation is performed to obtain the first source-of-goods question-and-answer information. Thus, based on fewer source-of-goods question-and-answer seeds and referring to the questioning intention of the source-of-goods question-and-answer seeds, more first source-of-goods question-and-answer information in different scenarios is generated. The first source-of-goods question-and-answer information is in the form of a question and an answer. According to a plurality of source-of-goods historical dialogue information including multi-round question-and-answer information, a plurality of dialogue process data are generated. Thus, according to the source-of-goods historical dialogue information in the real scenario, dialogue process data conforming to different scenarios are obtained. The first source-of-goods question-and-answer information, the source-of-goods question-and-answer seeds, the plurality of dialogue process data, and the first prompt information are input into a dialogue generation model established based on a large language pre-trained model to obtain a plurality of dialogue information. Thus, through a small number of source-of-goods question-and-answer seeds, a large number of diverse dialogue information can be quickly generated, reducing the labor cost. In the freight transportation field, a large amount of dialogue information is automatically obtained for subsequent fine-tuning of the initial artificial intelligence dialogue model, so as to obtain an artificial intelligence dialogue model conforming to the freight transportation field. In practical applications, the artificial intelligence dialogue model can give accurate answers to the inquiries of drivers, reduce the problem of model hallucinations, and improve the communication efficiency with drivers.
[0088] The following uses specific embodiments to elaborate in detail on the technical solutions provided by the present disclosure.
[0089] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for generating data in the freight transportation field provided by an embodiment of the present disclosure. As Figure 2 shown, the method provided in this embodiment is executed by one or more servers, and this server can be the same as or different from the server 102 in the above Figure 1 shown embodiment, and the present disclosure does not make any limitations in this regard. The method provided in this embodiment may include the following steps 201-step 204.
[0090] Step 201, obtain a plurality of source-of-goods question-and-answer seeds.
[0091] Among them, each source-of-goods question-and-answer seed contains a source-of-goods question seed and a source-of-goods answer seed. The source-of-goods question seed is a question raised for the source of goods, and the source-of-goods answer seed is an answer to the source-of-goods question seed. It can be understood that the source-of-goods question-and-answer seeds are in the form of a question and an answer, and the source-of-goods answer seed in each source-of-goods question-and-answer seed corresponds to the source-of-goods question seed. In addition, the source-of-goods question-and-answer seeds can be written by artificially simulating the questions and answers between a driver and an AI dialogue model. The source-of-goods question-and-answer seeds need to cover various types of questions about the source of goods.
[0092] Exemplarily, 1000 source-related Q&A seeds can be obtained.
[0093] Step 202: Perform a Q&A generation operation based on multiple source-related Q&A seeds and the intent information corresponding to each source-related Q&A seed to obtain the first source-related Q&A information.
[0094] Among them, the intent information corresponding to the source-related Q&A seed is used to indicate the questioning intent of the source question seed in the source-related Q&A seed. The intent information corresponding to the source-related Q&A seed can be directly obtained or classified by the server according to the source-related Q&A seed.
[0095] Exemplarily, the intent information corresponding to the source-related Q&A seed may include but is not limited to at least one of the following: asking about the name of the goods, asking about the loading and unloading address, asking about the vehicle specifications, asking for a price, and negotiating a price.
[0096] Perform a Q&A generation operation based on multiple source-related Q&A seeds and the intent information corresponding to each source-related Q&A seed to obtain the first source-related Q&A information. Among them, the first source-related Q&A information includes a first source-related question and a first source-related answer.
[0097] Step 203: Generate multiple dialogue process data based on multiple source-related historical dialogue information.
[0098] Among them, each source-related historical dialogue information contains multiple rounds of Q&A information. The source-related historical dialogue information can be the real dialogue between the user on the freight platform and the AI dialogue model. Further, a part of the dialogue can be obtained from the real dialogue by means of random sampling, etc., as the source-related historical dialogue information.
[0099] Among them, the dialogue process data is the data used to represent the dialogue Q&A process. For example, the dialogue process may be: in the first round of dialogue, the driver asks about the location of the unloading place, and the answer obtained is that the unloading place is in Logistics Park A; in the second round of dialogue, the driver asks whether his vehicle can transport this goods, and the answer obtained is yes.
[0100] Further, multiple source-related historical dialogue information can be input into the dialogue process generation model to obtain multiple dialogue process data. Among them, the dialogue process generation model is a pre-trained model. For example, the dialogue process generation model can be established based on a large model.
[0101] It should be noted that the execution of step 203 has no sequence with steps 201 - 202. Step 203 can be executed first, followed by steps 201 - 202; steps 201 - 202 can be executed first, followed by step 203; or step 203 and steps 201 - 202 can be executed simultaneously. The present disclosure does not limit the execution sequence of step 203 and steps 201 - 202.
[0102] Step 204: Input the first source - related Q&A information, source - related Q&A seeds, multiple dialogue process data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information.
[0103] Among them, the first prompt information is used to indicate generating multiple sample freight - related dialogues based on the first source - related Q&A information, source - related Q&A seeds, and multiple dialogue process data.
[0104] Among them, the dialogue generation model is established based on a large - language pre - trained model.
[0105] In this embodiment, the server obtains multiple source - related Q&A seeds. Each source - related Q&A seed contains a source - question seed and a source - answer seed. According to the multiple source - related Q&A seeds and the intention information indicating the questioning intention of each source - related Q&A seed, a Q&A generation operation is performed to obtain the first source - related Q&A information. Thus, based on fewer source - related Q&A seeds and referring to the questioning intention of the source - related Q&A seeds, more first source - related Q&A information in different scenarios is generated. The first source - related Q&A information is in the form of question - and - answer. Multiple dialogue process data are generated based on multiple source - related historical dialogue information containing multi - round Q&A information. Thus, according to the source - related historical dialogue information in the real scenario, dialogue process data that conform to different scenarios are obtained. The first source - related Q&A information, source - related Q&A seeds, multiple dialogue process data, and the first prompt information are input into the dialogue generation model established based on the large - language pre - trained model to obtain multiple dialogue information. Thus, through a small number of source - related Q&A seeds, a large number of diverse and authentic dialogue information can be quickly generated, reducing the labor cost. In the freight field, a large amount of dialogue information is automatically obtained for subsequent fine - tuning of the initial artificial - intelligence dialogue model, so as to obtain an artificial - intelligence dialogue model that conforms to the freight field. In practical applications, the artificial - intelligence dialogue model can give accurate answers to the drivers' inquiries, reduce the model hallucination problem, and improve the communication efficiency with the drivers.
[0106] In some embodiments, step 202 can be implemented through the following steps 2021 - 2027. Correspondingly, step 204 can be implemented through the following steps 2041 and 2042.
[0107] Step 2021: Store multiple source - related Q&A seeds in the task pool.
[0108] Among them, the initial state of the task pool can be empty, and the task pool is used to store source supply question and answer information.
[0109] Step 2022: Obtain all the source supply questions and answers in the task pool as the second source supply question and answer information.
[0110] Among them, the second source supply question and answer information includes second source supply questions and second source supply answers.
[0111] Obtain all the source supply questions and answers in the task pool from the current task pool as the second source supply question and answer information. It can be understood that the second source supply question and answer information obtained from the task pool for the first time is the source supply question and answer seed.
[0112] Step 2023: Identify the questioning intentions of the second source supply questions respectively to obtain the intention information of each second source supply question.
[0113] For the second source supply questions in each second source supply question and answer information, perform questioning intention identification respectively, that is, perform intention classification to obtain the intention information of each second source supply question.
[0114] Furthermore, the server can input the source supply question and answer seed into the intention classification model to obtain the intention information corresponding to the source supply question and answer seed. Among them, the intention classification model is a pre-trained model. For example, it can be a natural language understanding (NLU) model trained based on a large language model (LLM).
[0115] Furthermore, through the intention information corresponding to the source supply question and answer seed, the scenario to which the intention information corresponding to the source supply question and answer seed belongs can be obtained. For example, the intention information of asking about the loading and unloading address belongs to the loading and unloading scenario, asking about vehicle specifications belongs to the vehicle scenario, and inquiry and negotiation belong to the negotiation scenario. Furthermore, based on multiple source supply question and answer seeds and the scenarios to which each source supply question and answer seed belongs, a question generation operation can be performed to obtain the first source supply question and answer information.
[0116] Step 2024: Perform a question generation operation according to each second source supply question and the intention information of each second source supply question to obtain multiple first source supply questions.
[0117] The server generates new similar questions, that is, the first source supply questions, according to the second source supply questions and with reference to the intention information of the second source supply questions.
[0118] Furthermore, each second source supply question and the intention information of each second source supply question can be input into the question generation module to obtain the first source supply questions. Among them, the question generation module is a pre-trained model. For example, the question generation module can be established based on a large model.
[0119] Step 2025: Determine the first source replies corresponding to each first source problem.
[0120] The server generates the first source replies corresponding to the first source problems. Thus, a corresponding first source reply is generated for each first source problem.
[0121] Furthermore, each first source problem can be input into a reply generation module to obtain the first source replies corresponding to the first source problems. Exemplarily, the reply generation model can preset reply rules for reply generation. Among them, the reply rules can generate replies according to the scenarios of the current source problems.
[0122] Step 2026: Use each first source problem and the first source reply corresponding to the first source problem as the third source question-and-answer information.
[0123] Use each first source problem and the first reply corresponding to the first source problem as the third source question-and-answer information.
[0124] Step 2027: Store the third source problem information in the task pool.
[0125] Store the third source problem information in the task pool. At this time, the third source question-and-answer information can be cleared.
[0126] Step 2041: Obtain all the source question-and-answer information in the task pool from the task pool as the fourth source question-and-answer information.
[0127] Step 2042: Input the fourth source question-and-answer information, multiple dialogue process data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information.
[0128] In this embodiment, the source question-and-answer seeds have been stored in the task pool. Now, the generated third source question-and-answer information is stored in the task pool. Then, the current task pool contains the originally stored source question-and-answer seeds and the newly generated first source question-and-answer information. The source question-and-answer information in the task pool increases. Thus, more dialogue information can be generated according to the source question-and-answer information in the task pool. Therefore, with a small number of source question-and-answer seeds, a large number of diverse dialogue information can be quickly generated, reducing the labor cost.
[0129] In some embodiments, step 2026 can be implemented through the following step 20261.
[0130] Step 20261: Based on all the first source problems, the first source replies corresponding to the first source problems, and the second source question-and-answer information in the task pool, perform duplicate removal processing on all the first source problems and the first source replies corresponding to the first source problems to obtain the third source question-and-answer information.
[0131] Among them, for each first source question and its corresponding first source answer, the similarity values with the second source Q&A information can be respectively calculated. When there is a similarity value greater than the preset similarity threshold, the first source question and its corresponding first source answer are deleted. The similarity threshold is preset and can be dynamically set.
[0132] Furthermore, the first source questions and their corresponding first source answers can be de-duplicated by means of vectorization technology. Among them, all first source questions and their corresponding first source answers and all second source Q&A information can be converted into numerical vectors through a vectorization module, and the similarity values between the numerical vectors can be calculated. Exemplarily, similarity values can be calculated using cosine similarity, Jaccard similarity, etc.
[0133] In this embodiment, the repetition rate of source Q&A is controlled by dynamically setting the similarity threshold. If the similarity exceeds the similarity threshold, the two source Q&As are considered similar, and thus the duplicate source Q&As are removed. Through the de-duplication process, the first source questions with better diversity and their corresponding first source answers are retained and re-added to the task pool as examples for the next round of Q&A generation, further ensuring the diversity of the generated data.
[0134] In some embodiments, after step 2027, the following step 2028 may further be included.
[0135] Step 2028: Determine whether the number of third source Q&A information is greater than or equal to a preset quantity value.
[0136] Among them, the preset quantity value is preset.
[0137] If so, return to execute step 2022; if not, continue to execute step 204.
[0138] In this embodiment, during the current round of generation operation, if the number of the obtained third source problem information is large, the obtained third source question-and-answer information is stored in the task pool. During the process of cyclically generating new first source problems, the initial multi-round generated source question-and-answer information is also large after deduplication, that is, the number of the third source question-and-answer information is greater than or equal to the preset quantity value. Then, the third source question-and-answer information generated in this round is stored in the task pool, and the system returns to execute the step of obtaining all the source question-and-answer information in the task pool as the second source question-and-answer information, that is, generating new source question-and-answer information based on the source question-and-answer information in the current task pool again. After generating the third source question-and-answer information in multiple rounds, the number of the third source question-and-answer information may be small. Therefore, when the number of the third source question-and-answer information is less than the preset quantity value, it can be considered that fewer new source question-and-answer information is generated through the task pool. To save processing resources, the generated first source question-and-answer information can be stored in the task pool, and the generation of new source question-and-answer information can be stopped. Thus, while ensuring the generation of a large number of diverse third source question-and-answer information, processing resources are saved, and the efficiency of dialogue information generation is improved.
[0139] In some embodiments, step 204 may be implemented through the following step 2041.
[0140] Step 2041: Input the first source question-and-answer information, the source question-and-answer seed, multiple dialogue process data, the user portrait data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information.
[0141] The user portrait data refers to various information of the user. Generally, the user portrait data may include, but is not limited to, at least one of the following: user basic information, user language style information, user personality characteristic information, user consumption habit information, and user social attribute information, etc. Among them, the user basic information refers to the age, gender, region, and accent of the user, etc. The user language style information is used to indicate the language habit of the user during the conversation. The user personality characteristic information refers to the personality of the user usually in the conversation. For example, it can be irritable, calm, etc. The user consumption habit information refers to the consumption habit of the user in the freight scenario. For example, the user likes to share a ride, etc. The user social attribute information refers to the role information of the user in society.
[0142] The user portrait data can be directly obtained, or it can be automatically generated by summarizing and analyzing the data in the obtained user basic information database, user behavior information database, and user historical conversation database through a large model. The user basic information database refers to the set storing the basic information of the user. The user behavior information database refers to the set storing the behavior information of the user, which may include the bargaining habit, communication habit, inquiry habit, etc. of the user in history. The user historical conversation database refers to the set storing the data of the historical communication conversations of the user on the freight platform.
[0143] Among them, the first prompt information is specifically used to indicate that multiple sample freight conversations are generated according to the first source question-and-answer information, the source question-and-answer seeds, the user portrait data, and multiple conversation process data.
[0144] In this embodiment, by inputting the obtained user portrait data, the first source question-and-answer information, the source question-and-answer seeds, multiple conversation process data, and the first prompt information into the dialogue generation model together, it provides a basis for generating more personalized conversations, making the obtained multiple conversation information more in line with the habits of various users. The generated conversation information conforms to the settings of the user portrait in terms of style, and also has high diversity and authenticity.
[0145] In some embodiments, after step 204, the following steps 205 and 206 may further be included.
[0146] Step 205: Input the multiple conversation information and the preset quality inspection standard into the dialogue quality inspection model to obtain the quality inspection results of the multiple conversation information.
[0147] Among them, the dialogue quality inspection model is a model different from the dialogue generation model, so as to quality-inspect the dialogue information generated by the dialogue generation model through different models. The dialogue quality inspection model can also be established based on a large model. If the dialogue quality inspection model is established based on a large model, then the parameters or the model structure of the dialogue quality inspection model is different from that of the dialogue generation model. The dialogue quality inspection model is a large model with a larger scale than the dialogue generation model.
[0148] Step 206: Use the quality inspection results to indicate the multiple conversation information that passes the quality inspection and update the multiple conversation information.
[0149] In this embodiment, to ensure the quality of the generated dialogue information, the dialogue quality inspection model can be used to detect the quality of the dialogue information generated by the dialogue generation model. The dialogue quality inspection model can judge whether the dialogue information meets the expected requirements according to the preset quality inspection standard, and determine the dialogue information that meets the expected requirements as the final dialogue information. Thus, the accuracy of the generated dialogue information is ensured, and the quality of the dialogue information is improved.
[0150] In some scenarios, multiple conversation information is obtained through the above embodiments, and the initial artificial intelligence dialogue model can be adjusted using the multiple conversation information to obtain an artificial intelligence dialogue model. The following will be described in detail with specific embodiments.
[0151] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a method for adjusting an artificial intelligence dialogue model provided by an embodiment of the present disclosure. The method provided by this embodiment can be in Figure 2It can be implemented based on the embodiments shown, or can be implemented independently. The method provided in this embodiment includes the following steps 301 and 302.
[0152] Step 301: Obtain a plurality of conversation messages.
[0153] Among them, the plurality of conversation messages can be Figure 2 the conversation messages obtained by the method of the embodiments shown.
[0154] Step 302: Use the plurality of conversation messages to adjust the initial artificial intelligence dialogue model to obtain an artificial intelligence dialogue model.
[0155] Among them, the initial artificial intelligence dialogue model is established based on a large language pre-training model.
[0156] Among them, adjusting the initial artificial intelligence dialogue model can also be referred to as fine-tuning the initial artificial intelligence dialogue model.
[0157] For the method provided in this embodiment, its implementation principle and beneficial effects are similar to those of the above embodiments, and will not be elaborated here.
[0158] In some scenarios, the artificial intelligence dialogue model obtained through the above embodiments can answer questions raised by users in actual applications. The following will be described in detail with specific embodiments.
[0159] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for adjusting an artificial intelligence dialogue model provided by an embodiment of the present disclosure. The method provided in this embodiment can be implemented Figure 3 based on the embodiments shown, or can be implemented independently. The method provided in this embodiment includes the following steps 401 and 402.
[0160] Step 401: Receive a first message sent by a terminal device.
[0161] Among them, the first message contains a source problem raised for a first source of goods.
[0162] Step 402: Input the source problem, the information of the first source of goods, and the context information of the source problem into the artificial intelligence dialogue model to obtain a source reply corresponding to the source problem.
[0163] Among them, the information of the first source of goods refers to the transportation information related to the first source of goods.
[0164] Furthermore, the information of the first source of goods may include, but is not limited to, loading location information, unloading location information, cargo weight, cargo volume, required vehicle type, required vehicle length, etc.
[0165] Among them, the artificial intelligence dialogue model can be obtained through the aboveFigure 3 The artificial intelligence dialogue model obtained by the method of the illustrated embodiment.
[0166] Step 403: Send a source reply corresponding to the source problem to the terminal device.
[0167] The method provided in this embodiment has a similar implementation principle and beneficial effects to those of the above embodiments, which will not be elaborated here.
[0168] Next, in combination with Figure 5 Exemplarily illustrate the data generation method in the freight field provided by the embodiments of the present disclosure.
[0169] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the principle of a data generation method in the freight field provided by the embodiments of the present disclosure. As Figure 5 shown, the data generation method provided in this embodiment can be applied to a data generation device. The data generation device includes a pre-module and a main generation module. The pre-module includes a scenario dialogue generation module 510, a dialogue process generation module 520, and a user portrait generation module 530. The main generation module includes a multi-round dialogue generation module 540. The scenario dialogue generation module 510 is used to execute steps 201 and 202 in the above embodiments, and generate dialogue data in different scenarios through the scenario dialogue generation module 510. The dialogue process generation module 520 is used to execute step 203 in the above embodiments, and summarize the common dialogue processes existing online by using the dialogue process generation module 520. The user portrait generation module 530 is used to generate user portrait data by combining the data in the basic information library, the user behavior library, and the historical dialogue library. The multi-round dialogue generation module 540 is used to execute step 204 in the above embodiments, and can generate a large amount of high-quality, diverse, and highly fitting freight dialogue data online with reference to the text data generated by these pre-modules. The specific implementation of the above modules will be introduced below.
[0170] The scenario dialogue generation module 510 includes four parts. At the beginning of the task, multiple high-quality Q&A between the driver and the robot are manually written and stored as source Q&A seeds in the handwritten Q&A seed pool, and the source Q&A seeds in the handwritten Q&A seed pool are stored in the task pool.
[0171] The driver intention classification module 511 classifies the intentions of the source problems in the source Q&A seeds through a natural language understanding model trained based on a large model, and extracts the intention information therein. For example, asking about the name of the goods, asking about the loading and unloading address, asking about the vehicle specifications, asking for prices and bargaining, etc. These source problems are classified into different scenarios. For example, asking about the loading and unloading address belongs to the loading and unloading scenario, asking for prices and bargaining belongs to the bargaining scenario, and asking about the vehicle specifications belongs to the vehicle scenario.
[0172] The question generation module 512 generates new driver questions according to the set scenario requirements and example questions through a question generation model (Question LLM). The new driver questions correspond to the first source of goods questions in the above embodiments.
[0173] The answer generation module 513 replies to the new driver questions generated by the question generation module 512 through an answer generation model (Answer LLM) according to the preset reply rules to obtain the first source of goods reply, ensuring that the reply logic and structure of each scenario are clear. Among them, the question generation model and the answer generation model are two different large models.
[0174] The vectorization module 514 performs deduplication processing on the generated Q&A through vectorization technology to ensure data diversity. The vectorization module will convert the text dialogue into a numerical vector and identify duplicate texts by calculating the similarity between vectors. For example, similarity calculation methods such as cosine similarity and Jaccard similarity can be used. By dynamically setting a threshold to control the repetition rate of the text, if the similarity exceeds the set threshold, the two texts are considered similar, and the duplicate text is removed. Retain the texts with better diversity and re-add them to the task pool as examples for the next round of Q&A generation to further ensure the diversity of the generated data. The vectorization module 514 is used to execute step 20261 in the above embodiments.
[0175] The dialogue flow generation module 520 randomly samples the dialogues in the online dialogue pool and uses the LLM to summarize the dialogue flows of different dialogues, providing a dialogue reference for the LLM to automatically generate new dialogues later, and helping the model generate structured and practical dialogue data. Among them, the online dialogue pool is pre-constructed, and the dialogues in the online dialogue pool all come from the real dialogues between driver users and robots after the freight platform is launched.
[0176] The user portrait generation module 530 is to further enrich the diversity of the generated data. By depicting the real portraits of online users, it provides more dimensional reference bases for the generation of multi-round dialogue data. This module is based on the data in the user basic information library, behavior information library, and historical dialogue library, and uses a large model to summarize and analyze these data to automatically generate user portraits. The content of the user portrait not only includes language style and personality characteristics, but also covers multiple aspects such as consumption habits and social attributes, so as to provide a basis for generating more personalized data.
[0177] The multi-round dialogue generation module 540 combines the scenario dialogue, dialogue flow, and user portrait, refers to the previously generated content, and uses a large model to generate multi-round dialogue data. These generated multi-round dialogue data not only conform to the settings of the user portrait in terms of style, but also have high diversity and authenticity.
[0178] To ensure the quality of the generated data, after obtaining the multi-turn conversation data, the data quality inspection module 541 performs quality inspection on the generated multi-turn conversation data. The data quality inspection module 541 uses a large model with a larger scale than the conversation generation module or a large model different from the conversation generation model to perform quality inspection on the multi-turn conversation data. It judges whether the generated data meets the expected requirements according to the preset quality inspection standards, and only the data that meets the requirements will be output as the final result, ensuring the accuracy and high quality of the generated multi-turn conversation data.
[0179] In this embodiment, based on the data automatic generation framework of the large model, using the powerful data generation ability of the large model, large-scale and high-quality data is automatically generated. It improves the data generation method in the logistics domain, improves the speed, quality, and diversity of data generation, releases the annotation labor force, reduces the enterprise's employment cost, and improves the optimization and iteration of the subsequent logistics large model.
[0180] This disclosure provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the artificial intelligence conversation method in the freight domain in the above embodiment, which will not be elaborated here.
[0181] This disclosure provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the data generation method in the freight domain in the above embodiment, which will not be elaborated here.
[0182] This disclosure provides an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the adjustment method of the artificial intelligence conversation model in the above embodiment, which will not be elaborated here.
[0183] Based on the artificial intelligence conversation method in the freight domain described in any of the above embodiments, this disclosure also provides a computer-readable storage medium. For example, a non-temporary computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. Computer instructions are stored on this storage medium for executing the artificial intelligence conversation method in the freight domain described in any of the above embodiments, which will not be elaborated here.
[0184] Based on the data generation method in the freight field described in any of the above embodiments, an embodiment of the present disclosure further provides a computer-readable storage medium. For example, a non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. Computer instructions are stored on this storage medium for executing the data generation method in the freight field described in any of the above embodiments, which will not be elaborated here.
[0185] Based on the adjustment method of the artificial intelligence dialogue model described in any of the above embodiments, an embodiment of the present disclosure further provides a computer-readable storage medium. For example, a non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. Computer instructions are stored on this storage medium for executing the adjustment method of the artificial intelligence dialogue model described in any of the above embodiments, which will not be elaborated here.
[0186] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0187] After considering the specification and practicing the disclosure herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
Claims
1. An artificial intelligence dialogue method in the field of freight transportation, characterized in that, The method includes: Receiving a first message sent by a terminal device; the first message contains a source problem raised for a first source of goods; Inputting the source problem, the information of the first source of goods, and the context information of the source problem into an artificial intelligence dialogue model to obtain a source reply corresponding to the source problem; Wherein, the artificial intelligence dialogue model is obtained by adjusting an initial artificial intelligence dialogue model based on multiple sample freight conversations; the initial artificial intelligence dialogue model is established by adjusting a large language pre-training model; the multiple sample freight conversations are obtained by inputting first source question and answer information, source question and answer seeds, multiple dialogue process data, and first prompt information into a dialogue generation model; the first prompt information is used to indicate generating multiple sample freight conversations according to the first source question and answer information, the source question and answer seeds, and the multiple dialogue process data; the first source question and answer information is obtained by performing a question and answer generation operation according to the obtained multiple source question and answer seeds and the intention information corresponding to each source question and answer seed; each source question and answer seed contains a source problem seed and a source reply seed; the first source question and answer information contains a first source problem and a first source reply; the intention information corresponding to the source question and answer seed is used to indicate the questioning intention of the source problem seed in the source question and answer seed; the multiple dialogue process data are generated according to multiple source historical conversation information, and each source historical conversation information contains multi-round question and answer information; Sending the source reply corresponding to the source problem to the terminal device.
2. The method according to claim 1, wherein The method further includes: Obtaining multiple source question and answer seeds; each source question and answer seed contains a source problem seed and a source reply seed; Performing a question and answer generation operation according to the multiple source question and answer seeds and the intention information corresponding to each source question and answer seed to obtain first source question and answer information; the first source question and answer information contains a first source problem and a first source reply; the intention information corresponding to the source question and answer seed is used to indicate the questioning intention of the source problem seed in the source question and answer seed; Generating multiple dialogue process data according to multiple source historical conversation information; each source historical conversation information contains multi-round question and answer information; Inputting the first source question and answer information, the source question and answer seeds, the multiple dialogue process data, and first prompt information into a dialogue generation model to obtain multiple dialogue information; the first prompt information is used to indicate generating multiple sample freight conversations according to the first source question and answer information, the source question and answer seeds, and the multiple dialogue process data; the dialogue generation model is established based on a large language pre-training model.
3. A data generation method in the field of freight transportation, characterized in that, The method includes: Obtaining multiple source question and answer seeds; each source question and answer seed contains a source problem seed and a source reply seed; Perform a question-and-answer generation operation based on the multiple source question-and-answer seeds and the intent information corresponding to each source question-and-answer seed to obtain the first source question-and-answer information; the first source question-and-answer information includes a first source question and a first source answer; the intent information corresponding to the source question-and-answer seed is used to indicate the questioning intent of the source question seed in the source question-and-answer seed. Generate multiple dialogue process data based on multiple source historical dialogue information; each source historical dialogue information includes multiple rounds of question-and-answer information. Input the first source question-and-answer information, the source question-and-answer seeds, the multiple dialogue process data, and the first prompt information into a dialogue generation model to obtain multiple dialogue information; the first prompt information is used to indicate generating multiple sample freight conversations based on the first source question-and-answer information, the source question-and-answer seeds, and the multiple dialogue process data; the dialogue generation model is established based on a large language pre-trained model.
4. The method according to claim 2 or 3, characterized in that, The performing a question-and-answer generation operation based on the multiple source question-and-answer seeds and the intent information corresponding to each source question-and-answer seed to obtain the first source question-and-answer information includes: Store the multiple source question-and-answer seeds in a task pool. Obtain all the source questions and answers in the task pool as the second source question-and-answer information; the second source question-and-answer information includes a second source question and a second source answer. Identify the questioning intent of each second source question respectively to obtain the intent information of each second source question. Perform a question generation operation based on each second source question and the intent information of each second source question to obtain the first source question. Determine the first source answer corresponding to each first source question. Take each first source question and the first source answer corresponding to the first source question as the third source question-and-answer information. Store the third source question information in the task pool. The inputting the first source question-and-answer information, the source question-and-answer seeds, the multiple dialogue process data, and the first prompt information into a dialogue generation model to obtain multiple dialogue information includes: Obtain all the source question-and-answer information in the task pool from the task pool as the fourth source question-and-answer information. Input the fourth source question-and-answer information, the multiple dialogue process data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information.
5. The method according to claim 4, wherein The taking each first source question and the first source answer corresponding to the first source question as the third source question-and-answer information includes: Based on all the first source questions, the first source answers corresponding to the first source questions, and the second source question-and-answer information in the task pool, perform a deduplication process on all the first source questions and the first source answers corresponding to the first source questions to obtain the third source question-and-answer information.
6. The method according to claim 5, wherein After storing the third source question information in the task pool, it further includes: Judge whether the quantity of the third source question-and-answer information is greater than or equal to a preset quantity value. In the case that the quantity of the third source question-and-answer information is greater than or equal to the preset quantity value, return to execute obtaining all the source question-and-answer in the task pool as the second source question-and-answer information until the quantity of the third source question-and-answer information is less than the preset quantity value.
7. The method according to claim 2 or 3, characterized in that, Inputting the first source question-and-answer information, the source question-and-answer seed, the multiple dialogue process data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information, including: Inputting the first source question-and-answer information, the source question-and-answer seed, the multiple dialogue process data, the user portrait data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information; the first prompt information is specifically used to indicate generating multiple sample freight conversations according to the first source question-and-answer information, the source question-and-answer seed, the user portrait data, and the multiple dialogue process data; the user portrait data at least includes at least one of the following: user basic information, user language style information, user personality characteristic information, user consumption habit information, and user social attribute information.
8. The method according to claim 2 or 3, characterized in that, After inputting the first source question-and-answer information, the source question-and-answer seed, the multiple dialogue process data, and the first prompt information into the dialogue generation model to obtain multiple dialogue information, it further includes: Inputting the multiple dialogue information and the preset quality inspection standard into the dialogue quality inspection model to obtain the quality inspection result of the multiple dialogue information; Using the multiple dialogue information indicated by the quality inspection result to pass the quality inspection to update the multiple dialogue information.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.