Intelligent man-machine conversation method, device and equipment and storage medium
Through the intelligent human-computer dialogue method, users' intentions are analyzed and accurate computing power solutions are generated, which solves the problem of large language models generating inaccurate computing power solutions, and achieves efficient computing power resource utilization and business goals.
Patent Information
- Application Number
- CN202510091682.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
Large language models have hallucinations, data set quality problems and data timeliness when generating computing power solutions, resulting in inaccurate computing power solutions and affecting user experience.
It provides an intelligent human-computer dialogue method, which obtains user input information, analyzes user intentions, obtains target parameter values of preset task parameters based on multiple rounds of dialogue, retrieves target evaluation data in the evaluation database, and generates an accurate target computing power plan.
This method can accurately analyze user intentions, provide accurate and quantitative computing power solutions, help users make data-driven decisions, and ensure efficient utilization of computing power resources and smooth realization of business goals.
Smart Images

Figure CN120011541A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of intelligent computing technology, and in particular to an intelligent human-computer dialogue method, device, equipment and storage medium. Background Art
[0002] In the field of intelligent computing, the main technology for generating computing power solutions is Retrieval Augmented Generation (RAG) technology. The core of this technology is to retrieve local knowledge bases or Internet data, filter out information that is highly relevant to user queries, and load it into a general large-scale language model to infer and generate computing power solutions.
[0003] However, due to the hallucination problem of large language models, the data quality problem of large language model datasets, and the timeliness problem of local knowledge bases and Internet data, large language models may generate inaccurate computing power solutions, affecting user experience. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide an intelligent human-computer dialogue method, device, equipment and storage medium.
[0005] A first aspect of an embodiment of the present disclosure provides an intelligent human-computer dialogue method, the method comprising:
[0006] Get the user input information of the current round of dialogue;
[0007] Inputting the context information of the current round of dialogue into a pre-trained user intent classification model, and obtaining the user intent of the current round of dialogue output by the user intent classification model, wherein the context information includes user input information;
[0008] If the user intention of the current round of dialogue is to generate a computing power solution, target parameter values of multiple preset task parameters are obtained based on multiple rounds of dialogue, wherein the parameter values of the multiple preset task parameters do not represent different computing tasks at the same time;
[0009] Retrieving target evaluation data including target parameter values of the plurality of preset task parameters from a pre-established evaluation database, wherein the evaluation database includes a plurality of evaluation data, each of which includes quantitative data on performance and / or cost performance of a hardware component in executing a computing task under a computing power scheme;
[0010] Generate a target computing power plan based on the target evaluation data;
[0011] Output the target computing power solution.
[0012] A second aspect of the embodiments of the present disclosure provides an intelligent human-machine dialogue device, the device comprising:
[0013] The first acquisition module is used to obtain user input information of the current round of dialogue;
[0014] A first determination module is used to input the context information of the current round of dialogue into a pre-trained user intention classification model, and obtain the user intention of the current round of dialogue output by the user intention classification model, wherein the context information includes user input information;
[0015] A second acquisition module is used to acquire parameter values of multiple preset task parameters based on multiple rounds of dialogue if the user intention of the current round of dialogue is to generate a computing power solution;
[0016] A first retrieval module is used to retrieve target evaluation data including parameter values of the plurality of preset task parameters in a pre-established evaluation database, wherein the evaluation database includes a plurality of evaluation data, each of which includes quantitative data on performance and / or cost performance of a hardware component in a business scenario;
[0017] A first generating module, configured to generate a target computing power solution based on the target evaluation data;
[0018] The first output module is used to output the target computing power solution.
[0019] A third aspect of an embodiment of the present disclosure provides an electronic device, the server comprising: a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the method of the first aspect above.
[0020] A fourth aspect of an embodiment of the present disclosure provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect described above can be implemented.
[0021] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:
[0022] The disclosed embodiments can obtain user input information of the current round of dialogue; input context information of the current round of dialogue into a pre-trained user intent classification model, and obtain user intent of the current round of dialogue output by the user intent classification model, wherein the context information includes user input information; if the user intent of the current round of dialogue is to generate a computing power plan, obtain target parameter values of multiple preset task parameters based on multiple rounds of dialogue, wherein the parameter values of the multiple preset task parameters do not represent different computing tasks at the same time; retrieve target evaluation data including target parameter values of multiple preset task parameters from a pre-established evaluation database, wherein the evaluation database includes multiple evaluation data, each evaluation data including quantitative data on the performance and / or cost-effectiveness of a hardware component in executing a computing task under a computing power plan; generate a target computing power plan based on the target evaluation data; and output the target computing power plan. It can be seen that the above scheme can accurately analyze user intentions, and when the user intention is to generate a computing power plan, it can deeply explore and understand the user's specific computing tasks by obtaining the target parameter values of multiple preset task parameters, thereby providing an accurate and quantified target computing power plan for the computing task. This can help users make data-driven decisions in actual applications, ensure the efficient use of computing power resources and the smooth realization of business goals. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0025] Figure 1 is a flow chart of an intelligent human-computer dialogue method provided by an embodiment of the present disclosure;
[0026] Figure 2 is a logic diagram of an example of an intelligent human-computer dialogue method provided by an embodiment of the present disclosure;
[0027] Figure 3 It is a logic diagram of an example of generating a computing power solution provided by an embodiment of the present disclosure;
[0028] Figure 4 is a logic diagram of a computing power knowledge question and answer example provided in an embodiment of the present disclosure;
[0029] Figure 5It is a structural schematic diagram of an intelligent human-computer dialogue device provided by an embodiment of the present disclosure;
[0030] Figure 6 It is a structural schematic diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0032] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0033] In the field of intelligent computing, the main technology for generating computing power solutions and computing power knowledge question and answer results is RAG technology. The core of this technology is to retrieve local knowledge bases or Internet data, filter out information that is highly relevant to user queries, and load it into a general large language model to infer and generate computing power solutions and computing power knowledge question and answer results. However, this method has several significant defects: 1. The hallucination problem of large language models: Since the output of large language models is based on statistical probability, their accuracy cannot be fully guaranteed. Large language models may produce hallucinations or make non-factual statements, especially in scenarios with high understanding thresholds and strict accuracy requirements, which may cause users to receive misleading information and thus damage user trust. 2. The timeliness problem of local knowledge bases and Internet data: Both local knowledge bases and Internet data have timeliness problems. If the computing power knowledge is not effectively updated, the system may output outdated or erroneous computing power solutions and computing power knowledge question and answer results. This defect is particularly obvious in knowledge-intensive fields or those that require real-time updates, as it may cause users to rely on inaccurate or outdated information to make decisions. 3. Data quality issues of large language model datasets: Since many datasets may be outdated or unreliable, quality assurance is challenging. This not only affects the accuracy of the computing power plan and computing power knowledge question and answer results, but may also exacerbate the hallucination problem of large language models. As a result, large language models may generate inaccurate computing power plans and computing power knowledge question and answer results, affecting user experience.
[0034] In view of this, the embodiments of the present disclosure provide an intelligent human-computer dialogue method, device, equipment and storage medium. Figure 11 is a flow chart of an intelligent human-computer dialogue method provided by an embodiment of the present disclosure, and the method can be executed by an electronic device. The electronic device can be exemplarily understood as a device such as a mobile phone, a tablet, a computer, a server, etc. Figure 1 As shown, the method provided in this embodiment includes the following steps:
[0035] S110: Obtain user input information of the current round of dialogue.
[0036] In the disclosed embodiment, the electronic device has an intelligent human-computer dialogue function, and thus can receive user input information input by the user. The domain to which the user input information belongs may include the computing power domain (i.e., the user input information is related information about computing power) and / or the non-computing power domain; the form of the user input information may include text, image, voice and / or video, etc.; the input method of the user input information may include inputting the user input information through a keyboard, mouse, touch screen and / or voice module, etc., but is not limited thereto.
[0037] Specifically, a round of dialogue is a dialogue turn, which includes a user input information and the reply information generated by the electronic device in response to the user input information. For example, a round of dialogue includes "What is CPU? (i.e. user input information)" and "CPU (Central Processing Unit), called central processing unit in Chinese, is one of the core components of a computer system, responsible for executing instructions and processing data. It is usually called the "brain" of a computer because almost all computing business scenarios need to be completed through the CPU (i.e. reply information)".
[0038] Specifically, the current round of dialogue is an ongoing round of dialogue.
[0039] S120: Input the context information of the current round of dialogue into a pre-trained user intent classification model, and obtain the user intent of the current round of dialogue output by the user intent classification model, wherein the context information includes user input information.
[0040] In the disclosed embodiment, after obtaining the context information of the current round of conversation, the electronic device can call the user intent classification model to deeply analyze the user intent, so as to subsequently provide differentiated services for different user intents, thereby effectively improving user experience and service efficiency.
[0041] Specifically, context information refers to relevant background information required to analyze user intent.
[0042] In some embodiments, the context information of the current turn of the conversation includes user input information of the current turn of the conversation.
[0043] In other embodiments, if the current round of dialogue is the first round of dialogue, the context information of the current round of dialogue includes user input information; if the current round of dialogue is not the first round of dialogue, the context information of the current round of dialogue includes user input information, at least one round of historical dialogue and session embedding information.
[0044] Specifically, the first round of dialogue refers to the first round of dialogue.
[0045] Specifically, the historical round of dialogue is the dialogue that has been completed before the current round of dialogue. It should be noted that the context information of the current round of dialogue includes how many rounds of historical dialogue, and this disclosure does not limit this. For example, the context information of the current round of dialogue includes: the previous 2 rounds of dialogue (i.e., the previous round of dialogue and the previous round of dialogue), the previous 5 rounds of dialogue or the previous 10 rounds of dialogue, etc.; for another example, the context information of the current round of dialogue includes: dialogues that occurred within a preset time length (such as 3 minutes, 5 minutes or 10 minutes, etc.) from the current moment, but is not limited to this.
[0046] Specifically, session embedding information refers to specific tags embedded in the conversation to assist in analyzing user intent.
[0047] It is understandable that analyzing user intentions by combining user input information from the current round of conversation, at least one round of historical conversations, and session embedding information can enable electronic devices to more deeply understand the user's true intentions, and thus provide more accurate, efficient, and user-friendly services.
[0048] Specifically, user intent refers to a goal or purpose that a user hopes to achieve when engaging in human-computer dialogue with an electronic device.
[0049] In some embodiments, the types of user intent classified by the user intent classification model include computing power solution generation and computing power knowledge question and answer.
[0050] In other embodiments, the types of user intent classified by the user intent classification model include computing power solution generation, computing power knowledge question and answer, and non-computing power knowledge question and answer.
[0051] Specifically, the user intends to generate a computing power solution, indicating that the user wants to design a computing resource allocation plan for a specific business scenario (or a specific task). The computing power solution may include multiple aspects such as hardware model selection, software platform construction, and / or algorithm optimization, but is not limited thereto.
[0052] Specifically, the user's intention for a computing power knowledge question and answer indicates that the user wants the electronic device to answer questions related to computing power (such as questions about various concepts, technologies, and applications related to computing power) raised by the user.
[0053] Specifically, the user's intention for non-computing knowledge questions and answers indicates that the user wants the electronic device to answer the non-computing related questions raised by the user.
[0054] It should be noted that there are many specific training methods for the user intent classification model, which are not limited in this disclosure. For example, multiple context information is obtained as samples, and each sample is marked with a real user intent; multiple samples are divided into a training set and a test set; the following training is performed in a loop until the difference between the second loss values corresponding to two adjacent trainings is less than a preset threshold: the training set is input into the user intent classification model and the first predicted user intent output by the user intent classification model is obtained, the first loss value is calculated based on the real user intent corresponding to the training set and the first predicted user intent, the model parameters of the user intent classification model are updated based on the first loss value, the test set is input into the user intent classification model and the second predicted user intent output by the user intent classification model is obtained, the second loss value is calculated based on the real user intent corresponding to the test set and the second predicted user intent; the user intent classification model corresponding to the smallest second loss value is used as the trained user intent classification model. But it is not limited to this.
[0055] S130: If the user intention of the current round of dialogue is to generate a computing power solution, target parameter values of multiple preset task parameters are obtained based on multiple rounds of dialogue, wherein the parameter values of the multiple preset task parameters represent different computing tasks at the same time.
[0056] In the disclosed embodiment, if the user in the current round of conversation intends to generate a computing power plan, the electronic device can guide the user to provide target parameter values for multiple preset task parameters through multiple rounds of conversation to clarify the computing power plan for the computing task that the user wants to generate.
[0057] Specifically, the computing task may be represented by parameter values of a plurality of preset task parameters. For two different computing tasks, at least one parameter value of the preset task parameter is different between them.
[0058] Optionally, the computing tasks may include model computing tasks, physical simulation computing tasks, image processing computing tasks, etc., but are not limited thereto.
[0059] Further optionally, for a model calculation task, a plurality of preset task parameters may include: model application scenario, model name, parameter size, number of tokens, and expected completion time.
[0060] Specifically, the model application scenario refers to the application stage of the model. For example, the model application scenario includes training, pre-training, fine-tuning and / or inference.
[0061] Specifically, the model name refers to the type of model. For example, the model names include BERT (Bidirectional Encoder Representations from Transformers), ResNet (Residual Network), GPT (Generative Pre-trained Transformer), etc.
[0062] Specifically, parameter size refers to the total number of all learnable parameters in the model, including weights and bias terms.
[0063] Specifically, the number of tokens refers to the number of smallest units into which the input data is divided.
[0064] Specifically, the expected completion time refers to the expected time required to complete a certain application stage of the model.
[0065] For training, the expected completion time refers to the total time required from initializing the model to achieving satisfactory performance indicators.
[0066] For pre-training, the expected completion time refers to the time required to reach a certain level of feature extraction capability during the pre-training phase.
[0067] For fine-tuning, the expected completion time refers to the time required during the fine-tuning phase to adapt the model to the new business scenario and achieve the expected performance.
[0068] For inference, the expected completion time refers to the time required for a single or batch of input data to pass through the model to obtain the output result.
[0069] S140. Retrieve target evaluation data including target parameter values of multiple preset task parameters from a pre-established evaluation database, wherein the evaluation database includes multiple evaluation data, each evaluation data including quantitative data on the performance and / or cost-effectiveness of a hardware component in executing a computing task under a computing power scheme.
[0070] In the disclosed embodiment, in order to establish an evaluation database, multiple computing tasks, multiple models of hardware components and multiple computing power schemes may be pre-listed. The following evaluation is performed for each hardware component: for each computing task, the performance and / or cost performance (i.e., the relationship between performance and price) of the hardware component when the hardware component performs the computing task under each computing power scheme is tested, thereby obtaining quantitative data of the performance and / or quantitative data of the cost performance of the hardware component when performing the computing task under each computing power scheme.
[0071] Specifically, the hardware components described here refer to hardware with computing capabilities, such as CPU, GPU, TPU, and / or FPGA, but are not limited to these.
[0072] Specifically, the quantitative data of performance refers to the numerical indicators used to describe the working effect of a hardware component in executing a computing task under a computing power scheme. Exemplarily, when the computing task is a model computing task, the quantitative data of performance may include floating point operations per second (FLOPS), throughput, latency, GPU utilization, accuracy, precision, and / or recall, etc., but are not limited to this.
[0073] Specifically, the quantitative data of cost performance refers to an indicator used to measure the relationship between the performance provided by a hardware component and its cost. Exemplarily, the quantitative data of performance performance may include star ratings and / or scores, etc., but is not limited thereto.
[0074] Specifically, each piece of evaluation data includes device information of the hardware component (such as model or serial number, etc.), parameter values of multiple preset task parameters corresponding to the computing task, and quantitative data of performance and / or cost performance. In the evaluation database, evaluation data containing target parameter values of multiple preset task parameters are retrieved to obtain target evaluation data.
[0075] S150. Generate a target computing power plan based on the target evaluation data.
[0076] In the disclosed embodiment, the electronic device can perform data verification and formatting on the target evaluation data to generate a target computing power plan so as to output the target computing power plan in a structured manner, thereby ensuring that the user can receive accurate and clear computing power plan information.
[0077] In some embodiments, S150 includes: S151, detecting whether the target evaluation data is valid;
[0078] S152. If the target evaluation data is valid, organize the target evaluation data according to a preset format to obtain a target computing power solution.
[0079] Specifically, the validity of the target evaluation data can be detected from at least one of the following aspects: (1) Format verification: Check whether the parameters in the target evaluation data conform to the corresponding expected format, for example, the number of floating-point operations per second should be a numeric value. (2) Logical consistency: Verify whether the logical relationship between different parameters is reasonable. For example, if the model name is a lightweight model, then the size of its parameters should be small. (3) Range check: Ensure that the values in the target evaluation data are within a reasonable range. For example, the size of the parameters should not exceed billions. But it is not limited to this.
[0080] Specifically, the preset format includes a table and / or a chart, but is not limited thereto.
[0081] Of course, in some other embodiments, before S151, the following may be included: according to the quantitative data of the cost performance in all target evaluation data, all target evaluation data are sorted from high to low according to the cost performance, and the target evaluation data with a cost performance lower than a preset threshold are deleted. In this way, some more preferred computing power solutions can be selected for the user.
[0082] S160: Output the target computing power plan.
[0083] In the disclosed embodiment, the electronic device may output the target computing power plan to the user through display, voice broadcast, etc.
[0084] The disclosed embodiment adopts the above-mentioned scheme, which can accurately analyze the user's intention, and when the user's intention is to generate a computing power plan, it can deeply explore and understand the user's specific computing task by obtaining the target parameter values of multiple preset task parameters, thereby providing an accurate and quantified target computing power plan for the computing task. In this way, it can help users make data-driven decisions in actual applications, ensure the efficient use of computing power resources and the smooth realization of business goals.
[0085] In some other embodiments of the present disclosure, obtaining target parameter values of multiple preset task parameters based on multiple rounds of dialogues includes: extracting target parameter values of preset task parameters from user input information of the current round of dialogue;
[0086] Check whether the extracted target parameter value is valid;
[0087] If there are still unresolved task parameters among multiple preset task parameters, the first reply information of the current round of dialogue is generated based on the unresolved task parameters to start the next round of dialogue until the parameter values of multiple preset task parameters are extracted and valid, wherein the unresolved task parameters are preset task parameters for which valid parameter values have not been extracted, and the first reply information is used to guide the user to provide parameter values of the unresolved task parameters.
[0088] Specifically, when the user input information is not text, the voice, video, etc. input by the user can be converted into text, and then the target parameter value of the preset task parameter can be extracted from the text.
[0089] Specifically, whether the extracted target parameter value is valid can be detected from at least one of the following aspects: (1) Format verification: Check whether the extracted target parameter value conforms to the corresponding expected format, for example, the parameter value should be a numeric value. (2) Range check: Ensure that the extracted target parameter value is within a reasonable range. For example, the parameter value should not exceed billions. But it is not limited to this.
[0090] Specifically, unresolved task parameters refer to preset task parameters that have not yet extracted valid parameter values after previous conversations. Exemplarily, the multiple preset task parameters are model application scenarios, model names, parameter amounts, number of tokens, and expected completion time. If the target parameter values of the model application scenarios have been extracted through previous conversations, the model name, parameter amounts, number of tokens, and expected completion time are all unresolved task parameters.
[0091] Specifically, the first reply information refers to reply information in response to user input information, and is used to guide the user to provide a parameter value of at least one unresolved task parameter.
[0092] In some embodiments, generating first reply information of the current round of dialogue according to the unresolved task parameters includes: selecting at least one unresolved task parameter;
[0093] From the plurality of candidate reply messages, a first reply message matching the selected unresolved task parameters is screened.
[0094] Specifically, the candidate reply information refers to information used to guide the user to provide parameter values of the preset task parameters. One candidate reply information can be used to guide the user to provide one, two, or more parameter values of the preset task parameters, without limitation.
[0095] Specifically, for each preset task parameter, at least one piece of information for guiding the user to provide its parameter value may be preset, that is, at least one candidate reply information corresponding thereto may be preset. In this way, multiple candidate reply information may be obtained.
[0096] Specifically, one, two, or more unresolved task parameters may be selected, and this is not limited. Then, candidate reply information corresponding to the selected unresolved task parameters is screened from the multiple candidate reply information, and one of the screened candidate reply information is selected as the first reply information.
[0097] Exemplarily, the model name, parameter size, number of tokens and expected completion time are all unresolved task parameters, from which the parameter size is selected, and candidate reply information corresponding to the parameter size is filtered from multiple candidate reply information. For example, "What is the parameter size of the model you want to train?", then "What is the parameter size of the model you want to train?" can be used as the first reply information.
[0098] In some other embodiments, generating the first reply information of the current round of dialogue according to the unresolved task parameters includes: selecting at least one unresolved task parameter;
[0099] The first reply information is generated by combining the preset template and the selected preset task parameters.
[0100] Specifically, the preset template is a pre-set text message to be filled in. By filling the preset task parameters into the filling position, the first reply information can be obtained.
[0101] Exemplarily, the preset template is "Please provide the (fill in the blank) of the model to be processed", the model name, parameter size, token number and expected completion time are all unresolved task parameters, and the model name is selected from them. Combined with the preset template and model name, the first reply information "Please provide the model name of the model to be processed" can be obtained.
[0102] It can be understood that by setting up to screen the first reply information that matches the selected unresolved task parameters from multiple candidate reply information, or filling the selected unresolved task parameters into a preset template to obtain the first reply information, the method of obtaining the first reply information can be simple and quick, which is conducive to timely responding to user input information and improving user experience.
[0103] It is also understandable that through multiple rounds of dialogue, users can be guided to gradually provide complete parameter values of multiple preset task parameters, so as to obtain comprehensive computing task information, which is conducive to providing a more matching target computing power solution for the computing task that the user wants to perform. In addition, by setting the first reply information of the current round to be generated according to the unresolved task parameters, it is possible to avoid asking users again for the resolved task parameters (i.e. the preset task parameters whose parameter values have been extracted), which is conducive to improving efficiency.
[0104] Optionally, the method also includes: if there are still unresolved task parameters among the multiple preset task parameters, generating conversation embedding information corresponding to the current round of dialogue, wherein the conversation embedding information belongs to context information, which is used to instruct the user intention classification model to determine the user intention of the next round of dialogue as computing power plan generation.
[0105] Specifically, in the process of multiple rounds of conversation, after each round of conversation, if there are still unresolved task parameters among multiple preset task parameters, the electronic device will perform session embedding processing to obtain session embedding information, so that in the next round of conversation, the electronic device can also determine the user intention of the next round of conversation as a computing power solution generation based on the session embedding information, so as to continue to interact with the user about the parameter values of the preset task parameters in the next round of conversation, thereby ensuring the continuity of multiple rounds of conversation and the integrity of the parameter values of multiple preset task parameters.
[0106] It is understandable that in the next round of conversation, the electronic device will detect the conversation embedding information generated after the previous round of conversation, so that in the user intention recognition stage, it will give priority to continuing to execute multiple rounds of conversations based on the conversation embedding information, and then continue to collect the parameter values of the preset task parameters. It can be seen that the conversation embedding information can effectively avoid errors in user intention recognition, ensure the continuity of multiple rounds of conversations and the integrity of the parameter values of multiple preset task parameters, and thus improve the accuracy of the final target computing power solution.
[0107] In another embodiment of the present disclosure, the method further includes: if the user intention of the current round of dialogue is computing power knowledge question and answer, inputting the user input information of the current round of dialogue into a pre-trained keyword extraction model, and obtaining the keywords of the current round of dialogue output by the keyword extraction model;
[0108] A hybrid retrieval strategy is used to retrieve the knowledge fragments with the highest relevance to the retrieval information of the current round of dialogue from the knowledge base. The retrieval information includes user input information and keywords.
[0109] The knowledge fragments, user input information, reasoning prompt words and sample prompt words of the current round of dialogue are input into the pre-trained first answer generation model to obtain the second reply information of the current round of dialogue streamed output by the first answer generation model, and the second reply information of the current round of dialogue is output in real time.
[0110] Specifically, the electronic device can perform vectorization processing on the user input information, for example, by calling a pre-trained vectorization processing model to convert the user input information into a vector form, and calling a keyword extraction model to extract keywords from the user input information.
[0111] Specifically, the electronic device can adopt a hybrid search strategy, that is, combining vector similarity search and keyword search, comprehensively considering semantic similarity and keyword matching, and finding the knowledge fragment with the highest correlation with the search information (that is, user input information and keywords) from the knowledge base. It can be understood that vector similarity search can make up for the shortcomings of keyword search in processing fuzzy queries because it can understand semantic similarity; while keyword search ensures the precise positioning of specific terms, thereby providing more comprehensive and accurate search results.
[0112] Specifically, the reasoning prompt words are used to guide the first answer generation model on how to generate the second reply information in an expected manner. For example, the reasoning prompt words may instruct the first answer generation model: "Please make sure the answer is concise and clear."
[0113] Specifically, the sample prompt word refers to adding several similar examples to the input as a guide, so that the first answer generation model can learn how to imitate these examples to generate new, qualified second reply information. For example, if you want the first answer generation model to return an answer in a certain output format, you can include several correctly formatted examples in the input, and the first answer generation model will try to generate a similarly formatted second reply information based on these examples.
[0114] Specifically, streaming output means that the first answer generation model does not generate the entire answer at once, but gradually generates a part of the content and outputs it immediately. This can show the progress of the generation process in real time and allow users to interrupt or modify the query when necessary.
[0115] Specifically, the electronic device calls the first answer generation model, inputs the retrieved knowledge fragment and the user input information into the first answer generation model, sets the inference prompt words, and uses a small number of sample prompt words to provide a reference template for the output format. The first answer generation model is called through streaming output for inference, and finally generates the second reply information (i.e., the answer to the user input information).
[0116] It should be noted that there are many specific training methods for the vectorization processing model, the keyword extraction model, and the first answer generation model, which are not limited in this disclosure. For example, the training method of the user intent classification model can be understood with reference to the training method, but is not limited to this.
[0117] It is understandable that by calling the first answer generation model and combining the carefully designed reasoning prompt words and sample prompt words, the electronic device guides the first answer generation model to output the second reply information in a standardized output format. This structured output ensures the uniformity and readability of the answer. And calling the first answer generation model for reasoning through streaming output allows users to see the process of the answer gradually forming, while also providing flexibility so that users can give feedback or further clarify questions during the generation process.
[0118] In another embodiment of the present disclosure, the method further includes: if the user intention of the current round of dialogue is non-computing knowledge question and answer, inputting the user input information of the current round of dialogue into a pre-trained second answer generation model, and obtaining the third reply information of the current round of dialogue output by the second answer generation model;
[0119] Output the third reply information of the current round of dialogue.
[0120] Optionally, the electronic device calls the second answer generation model, inputs the user input information into the second answer generation model, sets reasoning prompt words, and uses a small number of sample prompt words to provide a reference template for the output format. The second answer generation model is called through streaming output for reasoning, and finally a third reply information (i.e., the answer to the user input information) is generated. It can be understood that by calling the second answer generation model and combining the carefully designed reasoning prompt words and sample prompt words, the electronic device guides the second answer generation model to output the third reply information in a standardized output format. This structured output ensures the uniformity and readability of the answer. And calling the second answer generation model for reasoning through streaming output allows users to see the process of the answer gradually forming, while also providing flexibility so that users can give feedback or further clarify questions during the generation process.
[0121] It should be noted that there are many specific training methods for the second answer generation model, which are not limited in this disclosure. For example, it can be understood by referring to the training method of the user intention classification model, but it is not limited to this.
[0122] It is understandable that by setting up the electronic device to be able to answer the user's non-computing knowledge questions, the electronic device can not only provide information on computing power, but also provide information on non-computing power, which is conducive to meeting the user's wider needs.
[0123] In another embodiment of the present disclosure, when the user intention is computing power knowledge Q&A and non-computing power knowledge Q&A, the method also includes: if the user input information of the current round of dialogue includes the target parameter value of the preset task parameter, generating session embedding information corresponding to the current round of dialogue, wherein the session embedding information belongs to context information, which is used to instruct the user intention classification model to determine the user intention of the next round of dialogue as computing power solution generation.
[0124] Specifically, it is detected whether the user input information of the current round of dialogue includes the target parameter value of the preset task parameter. If the user input information of the current round of dialogue includes the target parameter value of the preset task parameter, the electronic device will perform session embedding processing to obtain session embedding information, so that in the next round of dialogue, the electronic device will determine the user intention of the next round of dialogue as the generation of computing power plan according to the session embedding information, so as to guide the user to the interaction of computing power plan generation. Of course, in other embodiments, it is also possible to detect whether the user input information of multiple consecutive rounds of dialogues includes the target parameter value of the preset task parameter as of the current round of dialogue. If so, the electronic device will perform session embedding processing to obtain session embedding information, so that in the next round of dialogue, the electronic device will determine the user intention of the next round of dialogue as the generation of computing power plan according to the session embedding information.
[0125] It is understandable that in an actual conversation scenario, the user may want to generate a computing power plan, but is not sure how to ask, so at the beginning he may only ask some general questions related to the preset task parameters, such as "How long does it take to train a large model?" In the disclosed embodiment, by setting the conversation tracking information, the user can be slowly guided to the generation of the computing power plan, so as to guide the user to provide the target parameter values of multiple preset task parameters required to generate the computing power plan, thereby efficiently generating the computing power plan.
[0126] The following is a detailed description of the intelligent human-computer dialogue method provided by the embodiment of the present disclosure with reference to a specific example. Figure 2 is a logic diagram of an example of an intelligent human-computer dialogue method provided by an embodiment of the present disclosure; Figure 3 It is a logic diagram of an example of generating a computing power solution provided by an embodiment of the present disclosure; Figure 4 is a logic diagram of a computing power knowledge question and answer example provided by the embodiment of the present disclosure. Figure 2 As shown, the embodiment of the present disclosure introduces an intelligent dialogue system, which uses a large model, a multi-round dialogue mechanism, a knowledge base retrieval enhancement technology, and actual computing power data (i.e., an evaluation database) to provide users with accurate knowledge question and answer services and computing power solution recommendations. The functional architecture of the intelligent dialogue system is designed as three core branches: 1. Computing power knowledge question and answer branch: Focus on answering users' inquiries about general computing power knowledge, relying on RAG technology, prompt word enhancement technology, and calling large model APIs for deep reasoning to accurately output answers to questions. 2. Computing power solution generation branch: Dedicated to providing users with accurate and quantified computing power solutions for specific computing tasks. Mainly based on multi-round dialogue technology, conversation embedding strategy, and actual database data, it responds to users' personalized needs in a structured manner to ensure the practicality and accuracy of the computing power solution. 3. General dialogue question and answer branch: Responsible for responding to users' general questions related to non-computing power, so as to provide extensive dialogue support and cover diverse consulting needs. After receiving user input information, the intelligent dialogue system deeply analyzes which of the above three branches the user's intention is based on user input information, historical rounds of dialogue, and conversation embedding information. When identifying user intent, you can call the user intent classification model API and combine it with the prompt word template to guide the user intent classification model to output the results in a standardized JSON format (for example, output the number corresponding to the user intent). Figure 2 and Figure 3As shown, for the computing power knowledge question and answer branch, first, the intelligent dialogue system vectorizes the user input information, uses the vector processing model API to convert the user input information into a vector form, and calls the keyword extraction model API to extract the keywords in the user input information. Then, the intelligent dialogue system adopts a hybrid retrieval method, combining vector similarity retrieval and keyword retrieval to extract the knowledge fragment with the highest correlation with the user input information in the knowledge base. Finally, the intelligent dialogue system uses the retrieved knowledge fragment as context information and calls the first answer generation model API. At this time, the intelligent dialogue system inputs the context information together with the user input information, sets the reasoning prompt word, uses a small number of sample prompt words to provide a reference template for the output format, and calls the first answer generation model for reasoning through streaming output, and finally generates the answer (i.e., the second reply information). In addition, when the user input information includes the parameter values of the necessary parameters of the computing power plan, the intelligent dialogue system will also perform session embedding processing. In this way, when the user enters the user information again, the intelligent dialogue system will detect the current session embedding information, so that in the intention recognition stage, the intelligent dialogue system will prioritize the user intention to the computing power plan generation branch based on the session embedding information. For the computing power solution generation branch, the intelligent dialogue system will extract the parameter values of the necessary parameters (preset task parameters) of the computing power solution from the user input information and determine whether the extracted parameter values are valid. The necessary parameters include, for example, model application scenarios (training, pre-training, fine-tuning or reasoning), model name, parameter size, number of tokens, and expected completion time. If it is found that the necessary parameters are missing (that is, there are unresolved task parameters), the intelligent dialogue system will start a multi-round dialogue mechanism, through repeated iterations of intelligent dialogue system prompts, user feedback, parameter value extraction and parameter value verification processes, until the user enters the parameter values of all necessary parameters. In addition, during the multi-round dialogue process, the intelligent dialogue system will perform conversation point processing after each dialogue, so that when the user enters the user information again, the intelligent dialogue system will detect the current conversation point information, so that in the intention recognition stage, the intelligent dialogue system will give priority to continuing to perform multi-round dialogue information collection based on the conversation point information. Conversation point can effectively avoid errors in user intention recognition, ensure the coherence of multi-round dialogues, and ensure that the parameter values of all necessary parameters can be tracked and collected, which is conducive to improving the accuracy of the final computing power solution output. After collecting the parameter values of all necessary parameters, the intelligent dialogue system queries the actual evaluation database based on the collected parameter values of all necessary parameters. If it is to generate a training computing power plan, it queries the evaluation database; if it is to generate an inference computing power plan, it queries the inference evaluation database, so as to obtain the target evaluation data. The target evaluation data includes but is not limited to graphics card models, training / inference frameworks, performance evaluation indicators (i.e., quantitative data of performance) and cost-effectiveness information.Among them, the evaluation data in the evaluation database is obtained by performance evaluation based on multiple different models of computing power hardware servers and multiple different training / inference frameworks, and the computing power plan indicators are calculated and saved in combination with the actual computing power hardware price. Finally, the intelligent dialogue system verifies and formats the queried target evaluation data, and then returns the processed data to the front end for display to ensure that users can receive accurate and clear computing power plans. For non-computing power knowledge questions and answers, the intelligent dialogue system can call the second answer generation model API, set the reasoning prompt words, use a small number of sample prompt words to provide a reference template for the output format, call the second answer generation model through streaming output for reasoning, and finally generate the answer (i.e., the third reply information). In addition, when the user input information includes the parameter values of the necessary parameters of the computing power plan, the intelligent dialogue system will also perform session embedding processing, so that when the user enters the user information again, the intelligent dialogue system will detect the current session embedding information, so that in the intention recognition stage, the intelligent dialogue system will prioritize the user intention to the computing power plan generation branch based on the session embedding information.
[0127] The intelligent dialogue system provided by the disclosed embodiment can not only provide a wide range of computing power knowledge question and answer services to meet the user's needs for basic knowledge, but also deeply explore and understand the user's specific business scenarios, thereby providing accurate and quantified computing power solutions to meet the user's complex needs for computing power resource configuration and management. In this way, a comprehensive, efficient and user-friendly computing power question and answer platform is provided to users, which not only enhances users' understanding of computing power resources, but also helps them make data-driven decisions in practical applications, ensuring the efficient use of computing power resources and the smooth realization of business goals.
[0128] To sum up, in the field of intelligent computing, users often face a challenge: how to make wise choices among many computing power hardware models, training / inference frameworks, and hardware components with huge price differences. In order to solve this problem, the present disclosure provides users with an accurate, intuitive and structured computing power solution through intelligent dialogue, so that they can easily compare and make decisions to determine the most appropriate hardware and framework selection. The present disclosure not only simplifies the user's selection process for hardware and frameworks, but also takes into account the user's personalized needs. Users can specify a specific amount of training / inference data and set the expected completion time, etc. The technical solution disclosed in the present disclosure can generate accurate, quantitative and structured computing power solutions based on these specific computing tasks, which not only greatly improves the user experience, but also ensures that users can obtain tailored solutions. In this way, users no longer need to hesitate between complex technical parameters and prices, which not only saves users' time and resources, but also improves their work efficiency and project success rate. In addition, unlike the traditional RAG-based computing power question and answer scenario, the present disclosure integrates the intent recognition capability of the large model, which can not only support the RAG retrieval enhancement method to generate computing power knowledge questions and answers, but also combine measured data in the computing power solution question and answer scenario, and through the multi-round dialogue mechanism, provide accurate and structured computing power solutions for training / inference scenarios. In addition, considering that the technical solution of the present disclosure contains multiple functional branches, it may cause the illusion of intent recognition in the multi-round dialogue stage. In order to solve this challenge, the present disclosure adopts the conversation point technology, and presets the conversation point information at the end of each conversation. When the conversation is carried out again, these conversation point information will be used first to assist the intention recognition decision, thereby effectively reducing the illusion problem of intent recognition and ensuring the continuity and accuracy of the conversation. It can be seen that the technical solution of the present disclosure not only improves the efficiency and accuracy of computing power questions and answers, but also enhances the user experience, so that users can get more reliable and personalized support when facing complex computing power configuration problems.
[0129] Figure 5 1 is a schematic diagram of the structure of an intelligent human-computer dialogue device provided by an embodiment of the present disclosure. The intelligent human-computer dialogue device can be understood as the above electronic device or a part of the functional modules in the above electronic device. Figure 5 As shown, the intelligent human-machine dialogue device includes:
[0130] A first acquisition module 510 is used to acquire user input information of the current round of dialogue;
[0131] A first determination module 520 is configured to input the context information of the current round of dialogue into a pre-trained user intention classification model, and obtain the user intention of the current round of dialogue output by the user intention classification model, wherein the context information includes user input information;
[0132] A second acquisition module 530 is used to acquire parameter values of multiple preset task parameters based on multiple rounds of dialogue if the user intention of the current round of dialogue is to generate a computing power solution;
[0133] A first retrieval module 540 is used to retrieve target evaluation data including parameter values of the plurality of preset task parameters in a pre-established evaluation database, wherein the evaluation database includes a plurality of evaluation data, each of which includes quantitative data on performance and / or cost performance of a hardware component in a business scenario;
[0134] A first generating module 550, configured to generate a target computing power solution based on the target evaluation data;
[0135] The first output module 560 is used to output the target computing power solution.
[0136] Optionally, the multiple preset task parameters include model application scenarios, model names, parameter quantities, token quantities, and expected completion time.
[0137] Optionally, the second acquisition module 530 includes:
[0138] A first extraction submodule is used to extract a target parameter value of a preset task parameter from the user input information of the current round of dialogue if the user intention of the current round of dialogue is to generate a computing power solution;
[0139] The first detection submodule is used to detect whether the extracted target parameter value is valid;
[0140] The first reply submodule is used to generate the first reply information of the current round of dialogue according to the unresolved task parameters if there are still unresolved task parameters among the multiple preset task parameters, so as to start the next round of dialogue, until the parameter values of the multiple preset task parameters are extracted and valid, wherein the unresolved task parameters are preset task parameters for which valid parameter values have not been extracted, and the first reply information is used to guide the user to provide the parameter values of the unresolved task parameters.
[0141] Optionally, the device also includes: a first session embedding processing module, which is used to generate session embedding information corresponding to the current round of dialogue if there are still unresolved task parameters among the multiple preset task parameters, wherein the session embedding information belongs to the context information and is used to instruct the user intention classification model to determine the user intention of the next round of dialogue as computing power solution generation.
[0142] Optionally, the first generating module 550 is specifically used to detect whether the target evaluation data is valid;
[0143] If the target evaluation data is valid, the target evaluation data is organized according to a preset format to obtain the target computing power solution.
[0144] Optionally, the device further includes: a third acquisition module, configured to input the user input information of the current round of dialogue into a pre-trained keyword extraction model if the user intention of the current round of dialogue is computing power knowledge quiz, and obtain the keywords of the current round of dialogue output by the keyword extraction model;
[0145] A second retrieval module is used to use a hybrid retrieval strategy to retrieve from the knowledge base the knowledge fragment with the highest relevance to the retrieval information of the current round of dialogue, wherein the retrieval information includes user input information and keywords;
[0146] The second output module is used to input the knowledge fragments, user input information, reasoning prompt words and sample prompt words of the current round of dialogue into a pre-trained first answer generation model to obtain the second reply information of the current round of dialogue streamingly output by the first answer generation model, and output the second reply information of the current round of dialogue in real time.
[0147] Optionally, the device further includes: a fourth acquisition module, configured to input the user input information of the current round of dialogue into a pre-trained second answer generation model if the user intention of the current round of dialogue is non-computing knowledge question and answer, and obtain the third reply information of the current round of dialogue output by the second answer generation model;
[0148] The third output module is used to output the third reply information of the current round of dialogue.
[0149] Optionally, the device also includes: a second session embedding processing module, which is used to generate session embedding information corresponding to the current round of dialogue if the user input information of the current round of dialogue includes the target parameter value of the preset task parameter, wherein the session embedding information belongs to the context information and is used to instruct the user intention classification model to determine the user intention of the next round of dialogue as computing power solution generation.
[0150] Optionally, if the current round of dialogue is a first round of dialogue, the context information of the current round of dialogue includes user input information;
[0151] If the current round of dialogue is not the first round of dialogue, the context information of the current round of dialogue includes user input information, at least one round of historical dialogue and conversation embedding information.
[0152] The device provided in this embodiment can execute the method of any of the above embodiments, and its execution method and beneficial effects are similar, which will not be repeated here.
[0153] An embodiment of the present disclosure further provides an electronic device, which includes: a memory, in which a computer program is stored; and a processor, for executing the computer program. When the computer program is executed by the processor, the method of any of the above embodiments can be implemented.
[0154] For example, Figure 6 Schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 6 , which shows a schematic diagram of the structure of an electronic device 600 suitable for implementing the embodiment of the present disclosure. The electronic device 600 in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0155] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0156] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0157] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0158] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0159] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0160] The computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
[0161] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains user input information of the current round of dialogue;
[0162] Inputting the context information of the current round of dialogue into a pre-trained user intent classification model, and obtaining the user intent of the current round of dialogue output by the user intent classification model, wherein the context information includes user input information;
[0163] If the user intention of the current round of dialogue is to generate a computing power solution, target parameter values of multiple preset task parameters are obtained based on multiple rounds of dialogue, wherein the parameter values of the multiple preset task parameters do not represent different computing tasks at the same time;
[0164] Retrieving target evaluation data including target parameter values of the plurality of preset task parameters from a pre-established evaluation database, wherein the evaluation database includes a plurality of evaluation data, each of which includes quantitative data on performance and / or cost performance of a hardware component in executing a computing task under a computing power scheme;
[0165] Generate a target computing power plan based on the target evaluation data;
[0166] Output the target computing power solution.
[0167] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0168] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0169] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0170] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0171] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0172] The embodiments of the present disclosure further provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of any of the above embodiments can be implemented. The execution method and beneficial effects are similar and will not be repeated here.
[0173] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the information "includes a..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0174] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent human-computer dialogue method, characterized in that: include: Get the user input information of the current round of dialogue; Inputting the context information of the current round of dialogue into a pre-trained user intent classification model, and obtaining the user intent of the current round of dialogue output by the user intent classification model, wherein the context information includes user input information; If the user intention of the current round of dialogue is to generate a computing power solution, target parameter values of multiple preset task parameters are obtained based on multiple rounds of dialogue, wherein the parameter values of the multiple preset task parameters do not represent different computing tasks at the same time; Retrieving target evaluation data including target parameter values of the plurality of preset task parameters from a pre-established evaluation database, wherein the evaluation database includes a plurality of evaluation data, each of which includes quantitative data on performance and / or cost performance of a hardware component in executing a computing task under a computing power scheme; Generate a target computing power plan based on the target evaluation data; Output the target computing power solution.
2. The method according to claim 1, characterized in that The multiple preset task parameters include model application scenario, model name, parameter size, token number and expected completion time.
3. The method according to claim 1, characterized in that The step of obtaining target parameter values of multiple preset task parameters based on multiple rounds of dialogues includes: Extracting target parameter values of preset task parameters from user input information of the current round of dialogue; Check whether the extracted target parameter value is valid; If there are still unresolved task parameters among the multiple preset task parameters, a first reply message of the current round of dialogue is generated based on the unresolved task parameters to start the next round of dialogue until the parameter values of the multiple preset task parameters are extracted and valid, wherein the unresolved task parameters are preset task parameters for which valid parameter values have not been extracted, and the first reply message is used to guide the user to provide parameter values of the unresolved task parameters.
4. The method according to claim 3, characterized in that Also includes: If there are still unresolved task parameters among the multiple preset task parameters, generate session embedding information corresponding to the current round of dialogue, wherein the session embedding information belongs to the context information and is used to instruct the user intention classification model to determine the user intention of the next round of dialogue as computing power solution generation.
5. The method according to claim 1, characterized in that The generating a target computing power solution based on the target evaluation data includes: Detecting whether the target evaluation data is valid; If the target evaluation data is valid, the target evaluation data is organized according to a preset format to obtain the target computing power solution.
6. The method according to claim 1, characterized in that Also includes: If the user intention of the current round of dialogue is to answer questions about computing power knowledge, input the user input information of the current round of dialogue into a pre-trained keyword extraction model, and obtain the keywords of the current round of dialogue output by the keyword extraction model; Adopting a hybrid search strategy, searching from the knowledge base the knowledge fragment with the highest relevance to the search information of the current round of dialogue, wherein the search information includes user input information and keywords; The knowledge fragments, user input information, reasoning prompt words and sample prompt words of the current round of dialogue are input into a pre-trained first answer generation model to obtain the second reply information of the current round of dialogue streamed output by the first answer generation model, and the second reply information of the current round of dialogue is output in real time.
7. The method according to claim 1, characterized in that Also includes: If the user intention of the current round of dialogue is non-computing knowledge question and answer, input the user input information of the current round of dialogue into a pre-trained second answer generation model, and obtain the third reply information of the current round of dialogue output by the second answer generation model; Output the third reply information of the current round of dialogue.
8. The method according to claim 6 or 7, characterized in that: Also includes: If the user input information of the current round of dialogue includes the target parameter value of the preset task parameter, the session embedding information corresponding to the current round of dialogue is generated, wherein the session embedding information belongs to the context information and is used to instruct the user intention classification model to determine the user intention of the next round of dialogue as computing power solution generation.
9. The method according to claim 1, characterized in that: If the current round of dialogue is the first round of dialogue, the context information of the current round of dialogue includes user input information; If the current round of dialogue is not the first round of dialogue, the context information of the current round of dialogue includes user input information, at least one round of historical dialogue and conversation embedding information.
10. An intelligent human-computer dialogue device, characterized in that: include: The first acquisition module is used to obtain user input information of the current round of dialogue; A first determination module is used to input the context information of the current round of dialogue into a pre-trained user intention classification model, and obtain the user intention of the current round of dialogue output by the user intention classification model, wherein the context information includes user input information; A second acquisition module is used to acquire parameter values of multiple preset task parameters based on multiple rounds of dialogue if the user intention of the current round of dialogue is to generate a computing power solution; A first retrieval module is used to retrieve target evaluation data including parameter values of the plurality of preset task parameters in a pre-established evaluation database, wherein the evaluation database includes a plurality of evaluation data, each of which includes quantitative data on performance and / or cost performance of a hardware component in a business scenario; A first generating module, configured to generate a target computing power solution based on the target evaluation data; The first output module is used to output the target computing power solution.
11. An electronic device, characterized in that: include: A processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Intelligent dialogue method and system, electronic equipment, storage medium and program product
CN120849568A