Replay generation method and device, electronic equipment, storage medium and computer program product

Through pre-configured reply guidance and target matching, the problem of inefficient output control of language model is solved, more flexible and accurate output control is achieved, and the adaptability and efficiency of language model in complex scenarios is improved.

CN120258128APending Publication Date: 2025-07-04BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510147138.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently control the output results of the language model to meet the needs of specific scenarios. Especially in complex contexts and application scenarios, manually set constraint information is difficult to cover all possible scenarios, resulting in inefficient output.

Method used

By preconfiguring multiple replies, match the target replies related to the query content, and input them into the trained language model to constrain their output, including dynamically adjusting the replies based on the constraint rules of professional fields and question-and-answer scenarios.

Benefits of technology

It improves the control efficiency and flexibility of the output results of the language model, enhances the matching of the output results with the Q&A scenario, and makes up for the shortcomings of manual input constraint information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258128A_ABST
    Figure CN120258128A_ABST
Patent Text Reader

Abstract

The invention relates to a reply generation method and device, electronic equipment, a storage medium and a computer program product, relates to the technical field of artificial intelligence, and can effectively improve the control efficiency of a language model output result. The method comprises the following steps: acquiring inquiry content to be replied; determining a plurality of pre-configured reply guides used for constraining the reply mode of the inquiry content; the plurality of reply guides comprise reply guides for inquiry contents of various professional fields and / or question and answer scenes; in the multiple reply guides, matching a target reply guide related to the professional field and / or question and answer scene of the inquiry content; and inputting the target reply guide and the inquiry content into a trained language model, and outputting a first reply content for the inquiry content by the language model according to the target reply guide. By adopting the method, the language model output result control efficiency can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a reply generation method, apparatus, electronic device, storage medium, and computer program product. Background Art

[0002] With the rapid development of artificial intelligence technology, large-scale pre-trained language models have played an important role in the field of natural language processing and other fields. Such models can generate rich and diverse content by learning a large amount of data, and have demonstrated excellent processing performance in dialogue systems, content creation, and other aspects. While the model's creative ability is constantly improving, how to finely control the output of the model to ensure that the generated content meets the requirements and standards of specific scenarios has become an urgent problem to be solved.

[0003] In the related art, the generated content of the language model is mainly restricted by manually setting constraint information by the user. However, in the face of complex contexts and application scenarios, this method often fails to accurately control the output results of the language model, and often requires the user to try repeatedly many times.

[0004] It can be seen that the related art is difficult to efficiently control the language model to output information that meets the specific scenario requirements of the user. Summary of the Invention

[0005] The present disclosure provides a reply generation method, apparatus, electronic device, storage medium, and computer program product to at least solve the problem of low control efficiency of the output results of the language model in the related art. The technical solution of the present disclosure is as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, a reply generation method is provided, including:

[0007] Obtain the inquiry content to be replied;

[0008] Determine multiple reply guidelines pre-configured for restricting the reply manner of the inquiry content; the multiple reply guidelines include reply guidelines for inquiry content in various professional fields and / or question-and-answer scenarios;

[0009] Among the multiple reply guidelines, match the target reply guideline related to the professional field and / or the question-and-answer scenario of the inquiry content;

[0010] Input the target reply guideline and the inquiry content into a trained language model, and the language model outputs a first reply content for the inquiry content according to the target reply guideline.

[0011] In one exemplary embodiment, among the multiple reply guidelines, there is a first reply guideline determined based on the domain rules corresponding to the professional field involved in the question and answer, and the professional field includes at least one business field;

[0012] Before the step of obtaining the inquiry content to be replied, the method further includes:

[0013] Obtaining the business processing rules corresponding to the business field;

[0014] Configuring the first reply guideline according to the business processing rules; the first reply guideline is used to constrain the reply content to be output to conform to the business processing rules when the reply content to be output is related to the business field.

[0015] In one exemplary embodiment, the obtaining of the inquiry content to be replied includes:

[0016] Obtaining business behavior information for the account to be identified for whether there is an abnormal behavior; the business behavior information is behavior information in the business field;

[0017] Generating the inquiry content to be replied according to the business behavior information; the inquiry content is used to inquire whether there is an abnormality in the business behavior information.

[0018] In one exemplary embodiment, among the multiple reply guidelines, there is a second reply guideline determined based on the reply manner of the sample reply content in at least one of the question and answer scenarios;

[0019] Before the step of obtaining the inquiry content to be replied, the method further includes:

[0020] Collecting at least one sample question and answer pair in the question and answer scenario; the sample question and answer pair includes sample inquiry content, and positive sample reply content and negative sample reply content for the sample inquiry content;

[0021] Inputting the sample question and answer pair into the trained language model, and the language model determines the implicit reply rule for the question and answer scenario according to the difference in the reply manner between the positive sample reply content and the negative sample reply content;

[0022] Configuring the second reply guideline based on the implicit reply rule.

[0023] In one exemplary embodiment, the collecting of at least one sample question and answer pair in the question and answer scenario includes:

[0024] Obtaining the sample inquiry content based on the historical inquiry content in at least one of the question and answer scenarios;

[0025] Determine multiple historical reply contents in response to the historical inquiry content. Based on the target historical reply content selected by the account from the multiple historical reply contents, determine the positive sample reply content, and based on the historical reply contents other than the target historical reply content, determine the negative sample reply content.

[0026] In one exemplary embodiment, after matching out the target reply guidelines related to the professional field and / or the Q&A scenario of the inquiry content among the multiple reply guidelines, it further includes:

[0027] In the case where the number of the multiple target reply guidelines exceeds a preset threshold, determine the relevance of the multiple target reply guidelines to the professional field and / or the Q&A scenario of the inquiry content;

[0028] Screen the multiple target reply guidelines according to the relevance to obtain at least one target reply guideline whose quantity does not exceed the preset threshold and whose relevance meets the relevance condition; the preset threshold is determined according to the reply guideline processing ability of the trained language model;

[0029] The step of inputting the target reply guideline and the inquiry content into the trained language model includes:

[0030] Input the at least one target reply guideline and the inquiry content into the trained language model.

[0031] In one exemplary embodiment, after the step of the language model generating a first reply content for the inquiry content according to the target reply guideline, it further includes:

[0032] Input the first reply content and the target reply guideline into the language model, and when the first reply content does not conform to the target reply guideline, the language model adjusts the first reply content to obtain a second reply content that conforms to the target reply guideline;

[0033] Use the second reply content as the reply result of the inquiry content.

[0034] According to the second aspect of the embodiments of the present disclosure, there is provided a reply generation device, including:

[0035] An inquiry content receiving unit configured to execute obtaining the inquiry content to be replied;

[0036] A candidate guideline determining unit configured to execute determining multiple reply guidelines pre-configured for restricting the reply manner of the inquiry content; the multiple reply guidelines include reply guidelines for inquiry contents in various professional fields and / or Q&A scenarios;

[0037] A matching unit configured to perform matching, from multiple said reply guidelines, a target reply guideline related to the professional field and / or the Q&A scenario of the said inquiry content.

[0038] A first reply obtaining unit configured to perform inputting the said target reply guideline and the said inquiry content into a trained language model, and output, by the language model according to the said target reply guideline, a first reply content for the said inquiry content.

[0039] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0040] A processor;

[0041] A memory for storing instructions executable by the said processor;

[0042] Wherein, the said processor is configured to execute the said instructions to implement the reply generation method as described in any one of the above.

[0043] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, which, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the reply generation method as described in any one of the above.

[0044] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, which includes instructions that, when executed by a processor of an electronic device, enable the electronic device to execute the reply generation method as described in any one of the above.

[0045] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0046] On the one hand, by providing multiple configurable reply guidelines, inputting the target reply guideline and the inquiry content into the language model together, various reply constraint rules can be flexibly maintained, and the language model can flexibly output reply content related to the current Q&A scenario according to the currently dynamically input target reply guideline, improving the flexibility of output control. On the other hand, by providing multiple reply guidelines for inquiry content in various professional fields and Q&A scenarios and matching out the target reply guideline therefrom, rich relevant constraint information can be provided for various Q&As, effectively making up for the deficiency of manually input constraint information and improving the matching degree of the output result with the Q&A scenario. Thus, the control efficiency of the output result of the language model can be effectively improved.

[0047] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0048] The accompanying drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation on the present disclosure.

[0049] Figure 1 It is a flowchart of a reply generation method shown according to an exemplary embodiment.

[0050] Figure 2 It is a flowchart of steps for obtaining a second reply guidance shown according to an exemplary embodiment.

[0051] Figure 3 It is a flowchart of another reply generation method shown according to an exemplary embodiment.

[0052] Figure 4 It is a block diagram of a reply generation device shown according to an exemplary embodiment.

[0053] Figure 5 It is a block diagram of an electronic device shown according to an exemplary embodiment.

[0054] Figure 6 It is a block diagram of another electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0055] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0056] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above accompanying drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0057] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0058] In order to enable those skilled in the art to better understand the present disclosure, the related technologies are introduced below.

[0059] With the rapid development of artificial intelligence technology, large-scale pre-trained language models have played an important role in the field of natural language processing and other fields. By learning a vast amount of data, such models can generate rich and diverse content and have demonstrated excellent processing performance in machine translation, automatic summarization, dialogue systems, content creation, etc. While the model's creative ability is continuously improving, how to finely control the output of the model to ensure that the generated content meets the requirements and standards of specific scenarios has become an urgent problem to be solved.

[0060] In related technologies, it is mainly the user who manually sets constraint information to limit the generated content of the language model. In some examples, during the model training process, the user inputs keywords and style configuration information to tell the language model which style information should be focused on during the feature decoding stage, so that the model generates works with a specified style. For another example, manually design features for ranking, and combine additional context information, the topic distribution of the generated results, or the similarity score of the embedding to change the sampling distribution of the model, and then control the output content of the model.

[0061] However, in the face of complex contexts and application scenarios, the manually set constraint information often fails to cover all possible scenarios. At the same time, in the above methods, the modification or update of the constraint information requires retraining the model, which is time-consuming, laborious, and inefficient. When dealing with newly emerging abnormal problems, it is impossible to adjust the model output in a timely manner, which is too laggy and difficult to achieve rapid response. It can be seen that related technologies are difficult to efficiently control the information output by the language model to meet the specific scenario requirements of users.

[0062] Based on this, the present disclosure provides a reply generation method, device, electronic device, storage medium, and computer program product to at least solve the problem of low efficiency in controlling the output result of the language model in related technologies.

[0063] In an exemplary embodiment, as Figure 1 shown, a reply generation method is provided. In this embodiment, the method is exemplified by being applied to a server. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the server can be implemented by an independent server or a server cluster composed of multiple servers; the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc.

[0064] In this embodiment, the method includes the following steps:

[0065] In step S110, obtain the inquiry content to be replied to.

[0066] Among them, the inquiry content (query) is the content used to inquire about or confirm specified information. In some exemplary embodiments, the inquiry content may be information containing a question. Taking the content constructed by an interrogative sentence pattern or a rhetorical question sentence pattern as an example, the inquiry content may be "What is the area of A?" "Isn't B an animal?" It can be understood that the inquiry content may also be information without a question. Taking the content expressed by a declarative sentence pattern in a dialogue system as an example, the inquiry content may be "C is a part of D".

[0067] In specific implementation, the inquiry content to be replied to can be obtained. In some embodiments, the inquiry content may be the inquiry content input by a real user, or may be information automatically constructed according to the inquiry content construction rule.

[0068] In step S120, determine multiple reply guidelines pre-configured for constraining the reply manner; the multiple reply guidelines include reply guidelines for inquiry content in various professional fields and / or Q&A scenarios.

[0069] Regarding the situation in the related art where the constraint information needs to be modified by retraining the model, in some exemplary embodiments, a general language model may also be pre-trained. This language model can follow general instructions, so that when generating content, the user can set the constraint information and input it as an instruction into the language model to achieve controllable output of the language model. However, this method still requires the user to customize the constraint information, and the constraint information customized by the user is often only simple and obvious restrictions, and the effect is limited when facing complex contexts and application scenarios.

[0070] In view of this, in the embodiments of the present disclosure, multiple reply guidelines may be pre-configured. The reply guidelines may also be referred to as control rules, which can be used to constrain the reply manner of the language model to the inquiry content. In some exemplary embodiments, the multiple pre-configured reply guidelines may form an output control rule pool, and the output control rule pool may contain multiple reply guidelines; the multiple reply guidelines may be configured in one or more ways. For example, they may be manually added by the user, automatically crawled from the Internet through a crawler, or obtained by summarizing and generalizing information by a language model. Of course, they may also be obtained through other means.

[0071] In practical applications, the questions and answers may involve information related to specific professional fields. For example, the inquiry content may include information related to the telecommunications business field or legal knowledge field. In other cases, some question-and-answer information may not correspond to a specific professional field or field rules, but rather reply to the inquiry content in a specific scenario. In this regard, in the embodiments of the present disclosure, multiple reply guidelines may be preset. The multiple reply guidelines may be reply guidelines for inquiry content in various professional fields and / or question-and-answer scenarios. Among them, the professional field may be an information field divided according to a specified dimension. For example, the professional field may be a knowledge field or a business field, such as the chemical field, legal field, e-commerce field, logistics field, etc.

[0072] The question-and-answer scenario can be determined according to the specific business scenario where the question-and-answer interaction occurs and / or the theme around which the question-and-answer interaction revolves. For example, the question-and-answer scenario may include one or more of the following: intelligent customer service question-and-answer scenario, comment area automatic reply scenario, abnormal information review scenario, encyclopedia automatic question-and-answer scenario. Another example is that the question-and-answer scenario may include a question-and-answer scenario for the entertainment life theme or a question-and-answer scenario for the education and learning theme. In specific implementation, the automatic question-and-answer method based on artificial intelligence can be applied to one or more question-and-answer scenarios. The requirements for replying to the inquiry content may vary in different question-and-answer scenarios. For example, in the comment area automatic reply scenario, due to the word limit of the comment, it is often required that the provided reply content be limited within a specified word threshold as much as possible. For the encyclopedia automatic question-and-answer scenario, in order to improve the user's understanding efficiency of the information, more detailed explanations may be provided.

[0073] It can be understood that compared with the user manually inputting constraint information, the present application can provide rich relevant constraint information for various questions and answers by pre-providing reply guidelines for inquiry content in various professional fields and / or question-and-answer scenarios, effectively making up for the deficiency of manually inputting constraint information.

[0074] In step S130, among the multiple reply guidelines, match the target reply guideline related to the professional field and / or question-and-answer scenario of the inquiry content.

[0075] Specifically, in the present disclosure, by pre-configuring multiple reply guidelines for constraining the reply manner, it is possible to provide a variety of detailed and sufficient constraint information for different question-and-answer scenarios in the case where the user does not provide manually input constraint information. Correspondingly, among the multiple reply guidelines, there may be reply guidelines that are relatively relevant to the current question-and-answer scenario of the inquiry content, and there may also be reply guidelines that have a low relevance to the current question-and-answer scenario of the inquiry content.

[0076] In this regard, in this step, after obtaining multiple reply guidelines, the target reply guidelines related to the Q&A scenario or professional field of the inquiry content can be matched among the multiple reply guidelines. In some exemplary embodiments, the target reply guidelines related to the Q&A scenario of the inquiry content can be determined by at least one of keyword matching and feature matching.

[0077] Specifically, for example, the feature of the inquiry content corresponding to the inquiry content can be extracted, and the feature of the guideline content corresponding to each of the multiple reply guidelines is matched with the feature of the inquiry content. According to the reply guidelines corresponding to the successfully matched guideline content features, at least one target reply guideline can be obtained. Another example is that the query keyword in the inquiry content can be extracted, and the guideline keyword corresponding to each of the multiple reply guidelines is matched with the query keyword. According to the reply guidelines corresponding to the successfully matched guideline keywords, at least one target reply guideline can be obtained.

[0078] In step S140, the target reply guideline and the inquiry content are input into the trained language model, and the language model outputs the first reply content for the inquiry content according to the target reply guideline.

[0079] Among them, the language model can be a model pre-trained based on a large amount of text data. Based on a large amount of text data, the language model can learn rich language knowledge and context understanding ability. In this embodiment, the language model can have a generation function and can output corresponding text according to the provided prompt information. In some examples, the language model can be a self-developed model or a model obtained through other means.

[0080] Specifically, after determining the target reply guideline, in this step, the target reply guideline can be used as the prompt information and input into the trained language model together with the inquiry content. Thus, the language model can determine the relevant reply mode constraints according to the target reply guideline and output the reply content for the inquiry content. For the convenience of distinction, the reply content first generated by the language model can be referred to as the first reply content.

[0081] In the above reply generation method, the inquiry content to be replied can be obtained, and multiple reply guidelines for restricting the reply manner pre-configured can be determined. The multiple reply guidelines include reply guidelines for inquiry content in various professional fields and / or Q&A scenarios. Then, among the multiple reply guidelines, a target reply guideline related to the professional field and / or Q&A scenario of the inquiry content is matched, and the target reply guideline and the inquiry content are input into a trained language model. The language model generates a first reply content for the inquiry content according to the target reply guideline. In the embodiments of the present disclosure, on the one hand, by providing multiple configurable reply guidelines and inputting the target reply guideline and the inquiry content therein into the language model together, various reply constraint rules can be flexibly maintained, and the language model can flexibly output reply content related to the current Q&A scenario according to the currently dynamically input target reply guideline, improving the flexibility of output control. On the other hand, through multiple reply guidelines for inquiry content in various professional fields and Q&A scenarios and matching a target reply guideline therefrom, rich relevant constraint information can be provided for various Q&As, effectively making up for the deficiency of manually input constraint information and improving the matching degree of the output result with the Q&A scenario. Thus, the control efficiency of the output result of the language model can be effectively improved.

[0082] In an exemplary embodiment, some Q&A information may not correspond to a specific professional field or field rules, but reply to inquiry content in a specific scenario. In this regard, in the embodiments of the present disclosure, the reply manner of sample reply content in one or more Q&A scenarios can also be learned to determine the obtained reply guideline. Correspondingly, the multiple reply guidelines include a second reply guideline determined based on the reply manner of sample reply content in at least one Q&A scenario. For the convenience of distinction, this type of reply guideline is referred to as the second reply guideline. By learning the reply manner of the sample reply content, appropriate reply manners for various inquiry contents can be flexibly learned based on the existing sample reply contents. Thus, according to past sample experience, the language model is guided by the corresponding second reply guideline to output reply content that is more flexibly matched with the corresponding Q&A scenario, effectively solving problems that occur in various business scenarios in practice.

[0083] Correspondingly, as Figure 2 shown, before step S110, the following steps may further be included:

[0084] S210, collecting sample Q&A pairs in at least one Q&A scenario; the sample Q&A pairs include sample inquiry content, as well as positive sample reply content and negative sample reply content for the sample inquiry content.

[0085] In specific implementation, for at least one question-and-answer scenario, sample question-and-answer pairs in this question-and-answer scenario can be obtained. The sample question-and-answer pairs include sample inquiry content, as well as positive-sample reply content and negative-sample reply content for the sample inquiry content. Among them, the reply method of the positive-sample reply content to the sample inquiry content is better than that of the negative-sample reply content to the sample inquiry content. The negative-sample reply content is also called a bad case. In some examples, for the same sample inquiry content, at least one positive-sample reply content and at least one negative-sample reply content can be provided.

[0086] In some embodiments, the positive-sample reply content and / or the negative-sample reply content can be obtained by one or more methods such as manual selection and construction by staff. In some other embodiments, multiple sample reply contents for the same sample inquiry content can be determined, and then the positive-sample reply content and the negative-sample reply content among the multiple sample reply contents can be determined according to the respective popularity (such as interaction numbers like the number of likes or comments) of the multiple sample reply contents. For example, the sample reply content with the highest number of likes can be used as the positive-sample reply content, and the sample reply content with the number of likes lower than a preset threshold can be used as the negative-sample reply content.

[0087] S220, input the sample question-and-answer pairs into the trained language model, and the language model determines the implicit reply rules for the question-and-answer scenario according to the difference in the reply methods of the positive-sample reply content and the negative-sample reply content.

[0088] S230, configure the second reply guidance based on the implicit reply rules.

[0089] After obtaining the sample question-and-answer pairs, the sample question-and-answer pairs can be input into the language model. By comparing the difference in the reply methods of the positive-sample reply content and the negative-sample reply content, the language model can inductively learn the implicit reply rules for the sample inquiry content. Among them, the implicit reply rules can be implicit rules and norms when answering inquiry content.

[0090] Specifically, the language model can learn the reasonable or correct aspects of the reply method of the positive-sample reply content by comparing the difference in the reply methods of the positive-sample reply content and the negative-sample reply content, so as to determine the elements that should be possessed when replying to the inquiry content of this question-and-answer scenario. It can also learn some negative information that does not conform to the actual situation in the reply method of the negative-sample reply content, so as to determine the elements that should not be possessed (or need to be excluded or avoided) when replying to the inquiry content of this question-and-answer scenario, thereby obtaining the implicit reply rules for the question-and-answer scenario. Furthermore, based on the implicit reply rules, the second reply guidance can be configured. For example, the implicit reply rules can be configured as the second reply guidance, or after manually adjusting and confirming the implicit reply rules, the finally determined implicit reply rules can be used as the second reply guidance.

[0091] In this embodiment, by inputting sample question-and-answer pairs into a trained language model, the language model determines implicit reply rules for the question-and-answer scenario based on the differences in the reply methods of the positive-sample reply content and the negative-sample reply content. It can make full use of the inductive ability of the language model to quickly analyze the implicit reply rules in various question-and-answer scenarios. Facing different question-and-answer scenarios or complex contexts, it helps to supplement rule information that is not easily noticed by humans, improves the control fineness of the output results of the language model, and makes up for the deficiencies of manually written rules.

[0092] In an exemplary embodiment, in step S210, collecting sample question-and-answer pairs in at least one question-and-answer scenario may include the following steps:

[0093] Obtain sample inquiry content based on historical inquiry content in at least one question-and-answer scenario; determine multiple historical reply contents in response to the historical inquiry content, determine positive-sample reply content based on the target historical reply content selected by the account from the multiple historical reply contents, and determine negative-sample reply content based on the historical reply contents other than the target historical reply.

[0094] In this embodiment, sample question-and-answer pairs can be collected based on the historical inquiry content that has been proposed. Specifically, the historical inquiry content proposed by the account in at least one question-and-answer scenario can be obtained, and this historical inquiry content can be used as the sample inquiry content in the sample question-and-answer pair. On the other hand, multiple historical reply contents in response to this historical inquiry content can be obtained, where the historical reply content can be output by the language model or replied by real users. For example, the historical inquiry content and historical reply content can be obtained based on the question-and-answer information stored in the online system.

[0095] It can be understood that when providing historical reply content, one or more accounts can select the target historical reply content that is a suitable answer to the historical inquiry content from the multiple historical reply contents. For example, the historical reply content selected by the account that proposed the historical inquiry content can be used as the target historical reply content, or the historical reply content voted by multiple accounts can be used as the target historical reply content. Then, the negative-sample reply content can be determined based on the historical reply contents other than the target historical reply content.

[0096] In this embodiment, by relying on past real historical reply content, as well as the target historical reply content actually adopted by the user and other historical reply contents not adopted, it is possible to identify implicit reply rules in real scenarios in combination with real cases. While enabling post hoc induction and post hoc prevention, it helps to promptly solve problems when the inquiry content reappears in the same or similar question-and-answer scenarios.

[0097] Of course, in some other embodiments, the sample Q&A pairs can also be manually pre-constructed Q&A pairs. For example, positive sample response content and / or negative sample response content that are considered relatively important can be manually added, so as to achieve early prevention.

[0098] In an exemplary embodiment, the Q&A may involve information related to a specific professional field. For example, the inquiry content may include information related to the telecommunications business field or the legal knowledge field. In this regard, in this embodiment, a response guideline can be determined in advance according to the field rules corresponding to the professional field. For the convenience of distinction, this type of response guideline can be referred to as the first response guideline, and the first response guideline can be determined based on the field rules corresponding to the professional field involved in the Q&A.

[0099] Among them, the professional field can be an information field divided according to a specified dimension. For example, the professional field can be a knowledge field or a business field, such as the chemical field, the legal field, the e-commerce field, the logistics field, etc.; the first response guideline can be a guideline for guiding the language model to generate response content that conforms to the field rules. It can be understood that by obtaining the first response guideline, based on the existing, normative or professional field rules, more explicit constraint information can be provided for the information related to the professional field during the Q&A process, effectively making up for the deficiency of manually input constraint information.

[0100] The professional fields involved in the Q&A can include at least one business field, such as the e-commerce field, the financial business field, the telecommunications field, the legal field, etc.; before step S110, the method may further include:

[0101] Obtain the business processing rules corresponding to the business field; configure the first response guideline according to the business processing rules; the first response guideline is used to constrain the response content to be output to conform to the business processing rules when the response content to be output is related to the business field.

[0102] In specific implementation, the business field can have corresponding business processing rules, such as business norms, business systems, scenario constraints, etc. In some examples, the business processing rules of the business field can be information publicly disclosed by relevant institutions, and the business processing rules corresponding to the business field can be crawled from public websites through web crawler automation technology.

[0103] Furthermore, the first response guideline can be configured according to the business processing rules. Among them, the first response guideline can constrain the response content to be output to conform to the business processing rules of the business field when the response content to be output is related to the business field. In other words, by introducing the first response guideline, when replying to information related to the business field, the reply method can be restricted by the relevant rules in the business.

[0104] In this embodiment, by configuring the first reply guidance according to the business processing rules, on the one hand, when the relevant business fields are involved in the Q&A process, various existing business processing rules can be fully utilized to restrict the reply manner. Since the business processing rules are detailed, comprehensive, and accurate, it can effectively avoid the situation that the manually input constraint information is too simple to handle complex contexts. On the other hand, by restricting and constraining the content reply manner with the business processing rules, the matching degree between the reply content and the actual business processing logic can be effectively improved, making the reply content more in line with the reality, effectively avoiding the output reply of the language model containing abnormal information deviating from the reality, and improving the reliability, accuracy, and standardization of the reply content.

[0105] In one exemplary embodiment, in step S110, obtaining the inquiry content to be replied may include the following steps:

[0106] Obtain the business behavior information of the account to be identified for whether there is an abnormality; generate the inquiry content to be replied according to the business behavior information; the inquiry content is used to inquire whether there is an abnormality in the business behavior information.

[0107] In specific implementation, the business behavior information of the account to be identified can be obtained. Among them, the business behavior information is the behavior information in the specified business field, such as one or more of interactive behaviors, resource exchange behaviors, etc. The specified business field is the business field configured with the relevant first reply guidance, and the business behavior information can be determined by recording or summarizing the account behavior of the account in the business scenario.

[0108] Then, the inquiry content can be generated according to the business behavior information. The inquiry content can be used to inquire whether there is an abnormality in the business behavior information. For example, for the resource exchange record 1 of the account, the inquiry content "Does the resource exchange record 1 meet the interaction rules?" can be obtained.

[0109] It can be understood that since the corresponding first reply guidance is included in the pre-configured multiple reply guidances, by converting the business behavior information into the relevant inquiry content, the abnormality of the business behavior information can be accurately identified in combination with the flexibly configured multiple first reply guidances. While improving the accuracy of identifying abnormal business behavior information, the language model can be applied to abnormality identification, expanding the usage scenarios of the language model while improving the utilization rate of the language model.

[0110] In one exemplary embodiment, after step S130, the following steps may further be included:

[0111] In the case where the number of multiple target reply guidelines exceeds a preset threshold, determine the relevance of the multiple target reply guidelines to the professional field and / or Q&A scenario of the inquiry content; screen the multiple target reply guidelines according to the relevance to obtain at least one target reply guideline whose number does not exceed the preset threshold and whose relevance meets the relevance condition.

[0112] Among them, the preset threshold is determined according to the reply guideline processing ability of the trained language model.

[0113] In specific implementation, among the multiple reply guidelines provided in advance, multiple target reply guidelines relevant to the business field involved in the current inquiry content or the Q&A scenario to which the inquiry content belongs can be matched. For example, multiple target reply guidelines with a vector similarity exceeding the vector similarity threshold can be obtained by means of vector similarity.

[0114] In some embodiments, although it is also possible to input all the multiple target reply guidelines that have been matched into the language model, due to the limitation of the natural language processing ability of the language model, in the case of a large number of target reply guidelines, the language model may not be able to pay attention to all the target reply guidelines. Especially when the language model focuses on some main information, other information is often ignored by the language model, which is not conducive to the full play of the language model's capabilities.

[0115] In this regard, in this embodiment, the preset threshold of the guidelines can be determined in advance according to the reply guideline processing ability of the language model. In some examples, the preset threshold may be positively correlated with the reply guideline processing ability of the language model, that is, the stronger the reply guideline processing ability, the larger the preset threshold.

[0116] Furthermore, after obtaining multiple target reply guidelines, it can be determined whether the number of multiple target reply guidelines exceeds the preset threshold. If not, all the target reply guidelines obtained by matching can be input into the language model. If so, the multiple target reply guidelines obtained by matching can be further screened. Specifically, the relevance of the multiple target reply guidelines to the professional field and / or Q&A scenario of the inquiry content can be determined, and then, according to the relevance of each target reply guideline, the multiple target reply guidelines can be screened to obtain at least one target reply guideline whose number does not exceed the preset threshold and whose relevance meets the relevance condition. For example, all the target reply guidelines obtained by matching can be rearranged according to the relevance, and the first K (K is a positive integer) target reply guidelines after descending order can be used as at least one target reply guideline after screening.

[0117] Correspondingly, in step S140, inputting the target reply guideline and the inquiry content into the trained language model includes: inputting at least one target reply guideline and the inquiry content into the trained language model.

[0118] After screening out at least one target reply guidance whose quantity does not exceed a preset threshold, the at least one target reply guidance and the inquiry content can be input into the trained language model together.

[0119] In this embodiment, when the quantity of multiple target reply guidances exceeds the preset threshold, the multiple target reply guidances are screened again according to the relevance to obtain at least one target reply guidance whose quantity does not exceed the preset threshold and whose relevance meets the relevance condition, and then input into the language model. It is possible to dynamically streamline the reply guidance according to the actual processing ability of the language model, give priority to inputting the reply guidance more relevant to the current inquiry content into the language model, reduce the input of redundant guidance information to the language model, facilitate the language model to follow the instructions, generate replies that meet the rule requirements, and achieve better output control.

[0120] In an exemplary embodiment, after step S140, the method may further include the following steps:

[0121] Input the first reply content and the target reply guidance into the language model. When the first reply content does not conform to the target reply guidance, the language model adjusts the first reply content to obtain a second reply content that conforms to the target reply guidance; use the second reply content as the reply result of the inquiry content.

[0122] In specific implementation, the first reply content may contain some information that still does not meet the target reply guidance. For example, the first reply content may be a reply content generated by the language model in a generative manner (i.e., from scratch), that is, the language model generates the first reply content according to the target reply guidance. Due to being directly generated, the first reply content may partially deviate from the constraints and limitations of the target reply guidance.

[0123] In this regard, the language model in this embodiment may also have a rewriting function. After obtaining the first reply content, the first reply content and the target reply guidance can be input into the language model. The language model determines whether the first reply content conforms to the target reply guidance. If so, the first reply content can be determined as the reply result for the inquiry content. If not, the language model can adjust the first reply content to obtain a second reply content that conforms to the target reply guidance, such as performing output rewriting or making a fallback reply based on preset information, and then determine the second reply content as the reply result for the inquiry content.

[0124] Since the difficulty for the model to determine whether the response content meets the target response guidelines is often lower than that for the model to directly generate content that meets the target response guidelines, by inputting the first response content and the target response guidelines into the language model and instructing the model to reflect, it is possible to quickly identify situations where the target response guidelines are not met, determine the problems existing in the first response content, and make targeted optimization adjustments to achieve an overall better response result.

[0125] To enable those skilled in the art to better understand the above steps, the following provides an exemplary illustration of the embodiments of the present disclosure through an example, but it should be understood that the embodiments of the present disclosure are not limited thereto.

[0126] As Figure 3 shown, in this embodiment, a response system is provided, which may include a system configuration module, a generation control module, and a rewriting control module.

[0127] Among them, the system configuration module can build a control rule pool. On the one hand, existing industry norms or scenario constraints can be collected from public information as control rules. On the other hand, the existing badcase set of the online system can be obtained. Each badcase set stores two answers generated by the language model for the same query (such as Figure 3 "Answer 1" and "Answer 2" shown), and which of the two answers is more in line with the output requirements. Since real online queries are diverse and some queries lack corresponding industry norms, which may affect the language model to generate abnormal responses, by collecting badcases and summarizing them by the language model, implicit rules and norms can be extracted from the badcase set, so as to be able to solve more real online problems. In some examples, the vector features corresponding to the rules in the control rule pool can be obtained and stored in the vector database for downstream retrieval and application.

[0128] Upon receiving an inquiry content, such as the query input by the user, the information in the control rule pool can be matched with the inquiry content to obtain relevant control rules, and then the most relevant K control rules and the inquiry content are input into the language model to generate the first response content based on the control rules. Then, the K control rules and the generated first response content are input into the language model, and the language model is required to reflect on whether the generated first response content meets the rule requirements. If not, it is modified according to the rules to generate the second response content, and then the rewritten second response content is used as the response result for the inquiry content.

[0129] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turns with at least a part of other steps or steps or stages in other steps.

[0130] It can be understood that the same / similar parts among the various embodiments of the above method in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments. For the relevant parts, refer to the descriptions of other method embodiments.

[0131] Based on the same inventive concept, the embodiments of the present disclosure also provide a reply generation device for implementing the above-mentioned reply generation method.

[0132] Figure 4 It is a block diagram of a reply generation device shown according to an exemplary embodiment. Referring to Figure 4 , the device includes an inquiry content receiving unit 410, a candidate guidance determination unit 420, a matching unit 430, and a first reply obtaining unit 440.

[0133] The inquiry content receiving unit 410 is configured to obtain the inquiry content to be replied.

[0134] The candidate guidance determination unit 420 is configured to determine multiple reply guidelines pre-configured for restricting the reply manner of the inquiry content; the multiple reply guidelines include reply guidelines for inquiry content in various professional fields and / or question-and-answer scenarios.

[0135] The matching unit 430 is configured to match, among the multiple reply guidelines, a target reply guideline related to the professional field and / or the question-and-answer scenario of the inquiry content.

[0136] The first reply obtaining unit 440 is configured to input the target reply guideline and the inquiry content into a trained language model, and the language model outputs a first reply content for the inquiry content according to the target reply guideline.

[0137] In one embodiment, among the multiple reply guidelines, there is a second reply guideline determined based on the reply manner of sample reply content in at least one of the question-and-answer scenarios.

[0138] The device further includes a configuration acquisition unit configured to perform:

[0139] Collect sample question-and-answer pairs in at least one of the question-and-answer scenarios; the sample question-and-answer pairs include sample inquiry content, as well as positive sample reply content and negative sample reply content for the sample inquiry content;

[0140] Input the sample question-and-answer pairs into the trained language model, and the language model determines an implicit reply rule for the question-and-answer scenario according to the difference in the reply methods of the positive sample reply content and the negative sample reply content;

[0141] Configure the second reply guidance based on the implicit reply rule.

[0142] In one embodiment, the configuration acquisition unit is configured to perform:

[0143] Obtain the sample inquiry content based on the historical inquiry content in at least one of the question-and-answer scenarios;

[0144] Determine a plurality of historical reply contents feedback for the historical inquiry content, determine the positive sample reply content based on the target historical reply content selected by the account among the plurality of historical reply contents, and determine the negative sample reply content based on the historical reply contents other than the target historical reply content.

[0145] In one embodiment, among the multiple reply guidelines, there is a first reply guideline determined according to the domain rules corresponding to the professional field involved in the question-and-answer, and the professional field includes at least one business field;

[0146] The device further includes a configuration acquisition unit configured to perform:

[0147] Obtain the business processing rules corresponding to the business field;

[0148] Configure the first reply guideline according to the business processing rules; the first reply guideline is used to constrain the reply content to be output to conform to the business processing rules when the reply content to be output is related to the business field.

[0149] In one embodiment, the inquiry content receiving unit 410 is configured to perform:

[0150] Obtain business behavior information for which an account is to be identified as abnormal; the business behavior information is behavior information in the business field;

[0151] Generate the inquiry content to be replied according to the business behavior information; the inquiry content is used to inquire whether there is an abnormality in the business behavior information.

[0152] In one embodiment, the matching unit 430 is further configured to perform:

[0153] In the case where the number of multiple target reply guidelines exceeds a preset threshold, determine the relevance of the multiple target reply guidelines to the professional field and / or the Q&A scenario of the inquiry content;

[0154] Screen the multiple target reply guidelines according to the relevance to obtain at least one target reply guideline whose quantity does not exceed the preset threshold and whose relevance meets the relevance condition; the preset threshold is determined according to the reply guideline processing ability of the trained language model;

[0155] The inputting the target reply guideline and the inquiry content into the trained language model includes:

[0156] Input the at least one target reply guideline and the inquiry content into the trained language model.

[0157] In one embodiment, the device further includes a rewriting unit, and the rewriting unit is configured to perform:

[0158] Input the first reply content and the target reply guideline into the language model, and when the first reply content does not conform to the target reply guideline, the language model adjusts the first reply content to obtain a second reply content that conforms to the target reply guideline;

[0159] Use the second reply content as the reply result of the inquiry content.

[0160] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0161] Each module in the above reply generation device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0162] Figure 5FIG. 0 is a block diagram of an electronic device 500 for implementing a response generation method according to an exemplary embodiment. For example, the electronic device 500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0163] Referring to Figure 5 , the electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0164] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above-described method. In addition, the processing component 502 may include one or more modules to facilitate interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate interaction between the multimedia component 508 and the processing component 502.

[0165] The memory 504 is configured to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, videos, etc. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, optical disks, or graphene memory.

[0166] The power component 506 provides power to various components of the electronic device 500. The power component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.

[0167] The multimedia component 508 includes a screen that provides an output interface between the electronic device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of a touch or swipe action but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front camera and / or a rear camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0168] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.

[0169] The I / O interface 512 provides an interface between the processing component 502 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0170] The sensor component 514 includes one or more sensors for providing status assessments of various aspects of the electronic device 500. For example, the sensor component 514 can detect the on / off state of the electronic device 500, the relative positioning of components, such as the display and the keypad of the electronic device 500. The sensor component 514 can also detect a change in the position of the electronic device 500 or an electronic device 500 component, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the device 500, and a change in the temperature of the electronic device 500. The sensor component 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 514 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 514 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0171] The communication component 516 is configured to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0172] In an exemplary embodiment, the electronic device 500 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0173] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, and the above instructions can be executed by a processor 520 of the electronic device 500 to complete the above method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0174] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes instructions, and the above instructions can be executed by a processor 520 of the electronic device 500 to complete the above method.

[0175] Figure 6 is a block diagram of an electronic device 600 for implementing a reply generation method shown according to an exemplary embodiment. For example, the electronic device 600 can be a server. Referring to Figure 6 , the electronic device 600 includes a processing component 620, which further includes one or more processors, and memory resources represented by a memory 622 for storing instructions executable by the processing component 620, such as application programs. The application programs stored in the memory 622 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 620 is configured to execute instructions to perform the above method.

[0176] The electronic device 600 may further include: a power supply component 624 configured to perform power management of the electronic device 600, a wired or wireless network interface 626 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 628. The electronic device 600 may operate based on an operating system stored in the memory 622, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, or the like.

[0177] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as the memory 622 including instructions, and the above instructions can be executed by a processor of the electronic device 600 to complete the above method. The storage medium may be a computer-readable storage medium. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0178] In an exemplary embodiment, there is also provided a computer program product, and the computer program product includes instructions, and the above instructions can be executed by a processor of the electronic device 600 to complete the above method.

[0179] It should be noted that the above-mentioned apparatus, electronic device, computer-readable storage medium, computer program product, etc. may also include other implementation manners according to the description of the method embodiments. The specific implementation manners may refer to the description of the relevant method embodiments, and will not be elaborated here one by one.

[0180] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0181] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A reply generation method, characterized in that Including: Obtain the inquiry content to be replied; Determine multiple reply guidelines pre-configured for restricting the reply manner of the inquiry content; the multiple reply guidelines include reply guidelines for inquiry content in various professional fields and / or question-and-answer scenarios; Among the multiple reply guidelines, match the target reply guideline related to the professional field and / or the question-and-answer scenario of the inquiry content; Input the target reply guideline and the inquiry content into a trained language model, and the language model outputs a first reply content for the inquiry content according to the target reply guideline.

2. The method according to claim 1, characterized in that, Among the multiple reply guidelines, there is a first reply guideline determined based on the domain rules corresponding to the professional field involved in the question and answer, and the professional field includes at least one business field; Before the step of obtaining the inquiry content to be replied, it further includes: Obtain the business processing rules corresponding to the business field; Configure the first reply guideline according to the business processing rules; the first reply guideline is used to restrict the reply content to be output to conform to the business processing rules when the reply content to be output is related to the business field.

3. The method according to claim 2, wherein The obtaining the inquiry content to be replied includes: Obtain business behavior information of the account to be identified for whether there is an abnormality; the business behavior information is behavior information in the business field; Generate the inquiry content to be replied according to the business behavior information; the inquiry content is used to inquire whether there is an abnormality in the business behavior information.

4. The method according to claim 1, characterized in that, Among the multiple reply guidelines, there is a second reply guideline determined based on the reply manner of the sample reply content in at least one of the question-and-answer scenarios; Before the step of obtaining the inquiry content to be replied, it further includes: Collect sample question-and-answer pairs in at least one of the question-and-answer scenarios; the sample question-and-answer pairs include sample inquiry content, positive sample reply content, and negative sample reply content for the sample inquiry content; Input the sample question-and-answer pairs into the trained language model, and the language model determines the implicit reply rule for the question-and-answer scenario according to the reply manner difference between the positive sample reply content and the negative sample reply content; Configure the second reply guideline based on the implicit reply rule.

5. The method according to claim 4, wherein The collecting sample question-and-answer pairs in at least one of the question-and-answer scenarios includes: Obtain the sample inquiry content based on the historical inquiry content in at least one of the question-and-answer scenarios; Determine multiple historical reply contents feedback for the historical inquiry content, determine the positive sample reply content based on the target historical reply content selected by the account among the multiple historical reply contents, and determine the negative sample reply content based on the historical reply contents other than the target historical reply content.

6. The method according to claim 1, characterized in that, After matching the target reply guideline related to the professional field and / or the question-and-answer scenario of the inquiry content among the multiple reply guidelines, it further includes: In the case that the number of the multiple target reply guidelines exceeds a preset threshold, determine the relevance of the multiple target reply guidelines to the professional field and / or the question-and-answer scenario of the inquiry content; Filter multiple target response guidelines according to the relevance to obtain at least one target response guideline whose quantity does not exceed the preset threshold and whose relevance meets the relevance condition; the preset threshold is determined according to the response guideline processing ability of the trained language model; The inputting the target response guideline and the inquiry content into the trained language model includes: Input the at least one target response guideline and the inquiry content into the trained language model.

7. The method according to any one of claims 1 to 6, characterized in that, After the step of the language model generating a first response content for the inquiry content according to the target response guideline, it further includes: Input the first response content and the target response guideline into the language model, and when the first response content does not conform to the target response guideline, the language model adjusts the first response content to obtain a second response content that conforms to the target response guideline; Use the second response content as the response result of the inquiry content.

8. A reply generation device, characterized in that, It includes: An inquiry content receiving unit configured to execute obtaining the inquiry content to be replied; A candidate guideline determining unit configured to execute determining multiple response guidelines pre-configured for restricting the response manner of the inquiry content; the multiple response guidelines include response guidelines for inquiry contents in various professional fields and / or question-and-answer scenarios; A matching unit configured to execute matching out target response guidelines relevant to the professional field and / or the question-and-answer scenario of the inquiry content from the multiple response guidelines; A first response obtaining unit configured to execute inputting the target response guideline and the inquiry content into the trained language model, and the language model outputs a first response content for the inquiry content according to the target response guideline.

9. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the response generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the response generation method according to any one of claims 1 to 7.

11. A computer program product, comprising instructions, characterized in that, When the instructions are executed by the processor of the electronic device, the electronic device can execute the response generation method according to any one of claims 1 to 7.