Method and device for obtaining text based on interactive question and answer mode model and medium

By leveraging the semantic analysis ability of the interactive question-answer model, more accurate target feedback text is output from the preset strategy text, which solves the problem that the interactive question-answer model cannot query the information of the dedicated field, and realizes efficient and reliable query of text information.

CN120123486AActive Publication Date: 2025-06-10TIANJIN YITIAN DIGITAL SERVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510597072.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-10
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The interactive question-and-answer model relies on existing data during training and cannot query dedicated fields or undisclosed internal technical information, making it difficult for internal staff of the enterprise to query the required technical text.

Method used

By leveraging the semantic analysis capabilities of the interactive question-and-answer model, more accurate and consistent target feedback texts are output from the preset strategy text, achieving accurate expansion of keywords and efficient text query.

Benefits of technology

It improves the screening reliability of target feedback text and can accurately query more reliable text information from massive texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123486A_ABST
    Figure CN120123486A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text processing, in particular to a method and device for obtaining a text based on an interactive question and answer mode model.The method comprises the steps that a first semantic analysis text output by a given interactive question and answer mode model based on an input retrieval text is obtained, and a plurality of initial keywords are extracted, screening out a plurality of target keywords, obtaining a word weight corresponding to each target keyword, calculating the matching similarity between each preset strategy text and the plurality of target keywords based on the word weights, determining a plurality of first strategy texts according to the matching similarity, and re-executing the steps on the replacement text of the retrieval text, so as to obtain a plurality of second strategy texts. Obtaining a plurality of second strategy texts, and determining a target feedback text according to the repetition condition of the first strategy texts and the second strategy texts; according to the method, the powerful semantic analysis capability of the interactive question and answer model can be utilized, and the target feedback text which is more accurate and better meets the requirement can be output from the preset strategy texts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and in particular, to a method, device and medium for obtaining text based on an interactive Q&A model. Background Art

[0002] With the increasing development of AI big data, more and more interactive Q&A models are widely used by users, which brings great convenience to users' question queries and usage. However, there are still some defects that have not been overcome at present. For example, when the interactive Q&A model is trained, existing data is used as training data. When a user inputs a query question, the output content is all prior art or common knowledge. For some specific fields or internal technologies that have not been made public, they cannot be queried, which is not conducive to the internal staff of some enterprises to query technical texts. Therefore, it is necessary to provide an external library to establish a link with the interactive Q&A model, and how to accurately query the required text in a vast library is the problem to be solved at present. Summary of the Invention

[0003] In view of the above technical problems, the present invention provides a method, device and medium for obtaining text based on an interactive Q&A model, which can utilize the powerful semantic analysis ability of the interactive Q&A model to output more accurate and compliant target feedback text from a plurality of preset policy texts.

[0004] According to a first aspect of the present invention, there is provided a method for obtaining text based on an interactive Q&A model, including the following steps: S100, based on the retrieval text input by a user to a given interactive Q&A model, obtain the first semantic analysis text output by the given interactive Q&A model, and extract a plurality of initial keywords from the first semantic analysis text.

[0005] S200, screen a plurality of target keywords from the plurality of initial keywords according to the relevance between each initial keyword and the retrieval text, and obtain the word weight corresponding to each target keyword.

[0006] S300, based on the word weight corresponding to each target keyword, calculate the matching similarity between each preset policy text in a preset text library and the plurality of target keywords, and use the preset policy text with the corresponding matching similarity greater than the similarity threshold as the first policy text.

[0007] S400, use the replacement text corresponding to the retrieval text as the new retrieval text, and return to execute steps S100 - S300 to obtain a plurality of second policy texts; the replacement text corresponding to the retrieval text refers to the text obtained by performing a synonymous replacement on a given word object in the retrieval text through a preset large language model.

[0008] S500. When there is a duplication between the first policy text and the second policy text, the policy text with the highest corresponding matching similarity among all the duplicated first policy texts and all the duplicated second policy texts is used as the target feedback text; otherwise, the policy text with the highest corresponding matching similarity among all the first policy texts and all the second policy texts is used as the target feedback text.

[0009] According to a second aspect of the present invention, there is provided an apparatus for obtaining text based on an interactive Q&A model, the apparatus comprising: A first acquisition module, configured to obtain a first semantic analysis text output by a given interactive Q&A model based on a retrieval text input by a user to the given interactive Q&A model, and extract a plurality of initial keywords from the first semantic analysis text.

[0010] A second acquisition module, configured to screen out a plurality of target keywords from the plurality of initial keywords according to the relevance between each initial keyword and the retrieval text, and obtain the word weight corresponding to each target keyword.

[0011] A first calculation module, configured to calculate the matching similarity between each preset policy text in a preset text library and the plurality of target keywords based on the word weight corresponding to each target keyword, and use the preset policy text with the corresponding matching similarity greater than the similarity threshold as the first policy text.

[0012] A processing module, configured to use the replacement text corresponding to the retrieval text as a new retrieval text, and return to execute steps S100 - S300 to obtain a plurality of second policy texts; the replacement text corresponding to the retrieval text refers to the text obtained by performing a synonymous replacement on a given word object in the retrieval text through a preset large language model.

[0013] A first determination module, configured to, when there is a duplication between the first policy text and the second policy text, use the policy text with the highest corresponding matching similarity among all the duplicated first policy texts and all the duplicated second policy texts as the target feedback text; otherwise, use the policy text with the highest corresponding matching similarity among all the first policy texts and all the second policy texts as the target feedback text.

[0014] According to a third aspect of the present invention, there is provided a non - transitory computer - readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the above - mentioned method for obtaining text based on an interactive Q&A model.

[0015] The present invention has at least the following beneficial effects: The present invention provides a method for obtaining text based on an interactive Q&A model. First, a first semantic analysis text output by a given interactive Q&A model based on input retrieval text is obtained, and several initial keywords are extracted therefrom. In this process, the output result of the interactive Q&A model is not directly adopted, but the powerful semantic analysis ability of the model is utilized to obtain keywords from the semantic analysis text. In this process, the accurate expansion of keywords is realized, which is beneficial to querying more reliable text from the massive text in the preset text. Then, several target keywords are screened out according to the relevance between the initial keywords and the retrieval text, and the word weight corresponding to each target keyword is obtained. The matching similarity between each preset policy text and the several target keywords is calculated based on the word weight, so as to determine several first policy texts that more meet the retrieval requirements of the retrieval text according to the matching similarity. The replacement text of the retrieval text is re-executed the above steps to obtain several second policy texts, and the target feedback text is determined according to the repetition situation of the first policy text and the second policy text. In this way, the screening reliability of the target feedback text is improved as a whole. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the method for obtaining text based on an interactive Q&A model provided in Embodiment 1 of the present invention; Figure 2 It is a flowchart of step S200 provided in Embodiment 1 of the present invention; Figure 3 It is another flowchart of step S200 provided in Embodiment 1 of the present invention; Figure 4 It is a flowchart of step S300 provided in Embodiment 1 of the present invention; Figure 5 It is a flowchart of obtaining the retrieval text provided in Embodiment 1 of the present invention; Figure 6 It is a schematic structural diagram of the device for obtaining text based on an interactive Q&A model provided in Embodiment 2 of the present invention; Figure 7 It is a schematic structural diagram of the second acquisition module 200 provided in Embodiment 2 of the present invention; Figure 8 It is another schematic structural diagram of the second acquisition module 200 provided in Embodiment 2 of the present invention; Figure 9Schematic diagram of the first calculation module 300 provided in the second embodiment of the present invention. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present invention.

[0019] Embodiment 1 As Figure 1 shown, Embodiment 1 of the present invention provides a method for obtaining text based on an interactive Q&A model, including the following steps: S100, based on the retrieval text input by the user to a given interactive Q&A model, obtain the first semantic analysis text output by the given interactive Q&A model, and extract several initial keywords from the first semantic analysis text. In a specific implementation, the given interactive Q&A model outputs a first semantic analysis text and a result feedback text for the retrieval text, and this application does not directly use the result feedback text, but only obtains the first semantic analysis text. For example, the given interactive Q&A model can be the deepseek large language model, and the content it outputs includes the analysis logic and the result.

[0020] As mentioned above, since the reply results output by the Q&A models in the prior art are all prior art or common knowledge, and in order to obtain the non-public query text in the internal text library in this application, the output results of the interactive Q&A model are not directly used, but the powerful semantic analysis ability of the model is utilized to obtain keywords from the semantic analysis text. In this process, the accurate expansion of keywords is realized, so as to query more reliable text from the massive text in the preset text.

[0021] S200, screen several target keywords from several initial keywords according to the relevance of each initial keyword to the retrieval text, and obtain the word weight corresponding to each target keyword; it can be understood that: the initial keywords with a high degree of relevance to the retrieval text are screened out as target keywords.

[0022] Specifically, as Figure 2 shown, the step of screening several target keywords from several initial keywords according to the relevance of each initial keyword to the retrieval text includes the following steps: S201, screen out the initial keywords that are the same as any preset conventional word from several initial keywords to obtain the remaining several first keywords; it can be understood that: the preset conventional word is a habitual word with a low relevance to the technical field where the retrieval text is located.

[0023] S202. According to the position information of each initial keyword in the first semantic analysis text, obtain the reverse character ranking corresponding to each first keyword, and use a preset activation function to normalize the order values corresponding to the reverse character rankings of all first keywords to obtain the first keyword score corresponding to each first keyword. Among them, the independent variable coefficient of the preset activation function is related to the total number of characters in the retrieval text.

[0024] Specifically, the preset activation function uses the Sigmoid function.

[0025] Furthermore, in this embodiment, the preset activation function meets the following conditions: S = 2 / (1 + e -σx ), where x refers to the x-th reverse character ranking, S is the normalized value corresponding to x, σ is the preset independent variable coefficient, the value range of σ is between 0 and 1, and σ is negatively correlated with the total number of characters in the retrieval text.

[0026] In a preferred embodiment, σ = 2 / K, where K is the total number of characters in the retrieval text. That is, when K is larger, the first keyword scores corresponding to the more forward first keywords are closer to 1.

[0027] S203. Calculate the word frequency of each first keyword based on the first semantic analysis text, and determine the normalized value corresponding to the word frequency of each first keyword as the second keyword score corresponding to the first keyword itself.

[0028] S204. Determine the product of the first keyword score and the second keyword score as the final score of the first keyword, and screen out the first keywords with the final scores greater than the preset score threshold as target keywords. In a specific implementation, corresponding weights can also be set for the first keyword score and the second keyword score. Those skilled in the art set the preset score threshold according to actual needs, which will not be elaborated here.

[0029] As described above, when considering the relevance between the first keyword and the retrieval text, the position information and word frequency corresponding to the first keyword are introduced. Since in the semantic analysis text, usually the content in the front part of the text is the logical analysis keywords with high relevance to the retrieval text, such as the decomposition of problems, and the content in the back is the logical reasoning or examples of various possibilities. Therefore, the more forward the keyword, the greater the relevance to the retrieval text, and the higher the corresponding first keyword score. And the larger the word frequency, the higher the importance degree of the keyword. Therefore, these two factors are introduced and combined in this application, so that the selected target keywords have higher relevance to the retrieval text, in order to screen out a more compliant strategy text.

[0030] In a specific implementation manner, such as Figure 3As shown below, the word weights corresponding to each target keyword are obtained through the following steps: S210. Determine the normalization result of the final score corresponding to each target keyword as the first word confidence corresponding to the target keyword itself.

[0031] S220. Input the retrieval text sample into the given interactive Q&A model several times to obtain several second semantic analysis texts, calculate the TF-IDF value corresponding to each target keyword based on the several second semantic analysis texts, and determine the normalization result of the TF-IDF value corresponding to each target keyword as the second word confidence corresponding to the target keyword itself; those skilled in the art are aware of the specific process of calculating the TF-IDF value corresponding to a keyword, which will not be elaborated here.

[0032] S230. For any target keyword, calculate the weighted sum according to the first word confidence corresponding to the target keyword, the second word confidence, and the preset weights corresponding to the first word confidence and the second word confidence respectively to obtain the word weight corresponding to the target keyword.

[0033] As described above, when calculating the word weight corresponding to each target keyword, the first semantic analysis text where the target keyword is located and multiple second semantic analysis texts where the target keyword is located are considered, that is, the combination of its own factors and external factors makes the calculated word weight more reasonable and reliable.

[0034] S300. Based on the word weights corresponding to each target keyword, calculate the matching similarity between each preset strategy text in the preset text library and several target keywords, and use the preset strategy text with the corresponding matching similarity greater than the similarity threshold as the first strategy text; those skilled in the art set the similarity threshold according to actual needs, which will not be elaborated here.

[0035] Further, as Figure 4 shown, step S300 specifically includes the following steps: S301. Obtain the word vector corresponding to each target keyword, and calculate the matching similarity between each word vector and each preset strategy text; the vector dimensions of the word vector corresponding to the target keyword and the word vector corresponding to the preset strategy text are the same.

[0036] S302. For any preset strategy text, calculate the weighted sum based on the word weights corresponding to each target keyword and the matching similarity between each word vector and each preset strategy text to obtain the matching similarity between each preset strategy text and several target keywords.

[0037] As described above, when calculating the similarity between the target keyword and the predicted policy text, the importance of different keywords is considered. Since the word weights of the keywords with a high relevance to the retrieval text are also set relatively high, introducing word weights to calculate the matching similarity can obtain a policy text that better meets the retrieval requirements of the retrieval text.

[0038] S400. Use the replacement text corresponding to the retrieval text as the new retrieval text, and return to execute steps S100 - S300 to obtain several second policy texts; the replacement text corresponding to the retrieval text refers to the text obtained by performing synonym replacement on the given word object in the retrieval text through a preset large language model; it can be understood that: respectively execute steps S100 - S300 on several replacement texts corresponding to the retrieval text, and each replacement text corresponds to a second policy text.

[0039] As described above, considering the problem that due to the subjectivity of users, the words used in the retrieval text may be inaccurate and affect the query results, multiple synonym replacements are performed on the word objects in the retrieval text, making the obtained second policy texts more comprehensive and conducive to obtaining policy texts that meet the reply requirements.

[0040] S500. When there is a duplicate situation between the first policy text and the second policy text, use the policy text with the highest corresponding matching similarity among all the duplicate first policy texts and all the duplicate second policy texts as the target feedback text; conversely, use the policy text with the highest corresponding matching similarity among all the first policy texts and all the second policy texts as the target feedback text. For example, when there are two groups of identical first policy texts and second policy texts, select the policy text with the highest corresponding matching similarity from the two first policy texts and the two second policy texts as the target feedback text.

[0041] As described above, since both the first policy text and the second policy text are policy texts that meet the retrieval requirements queried from the retrieval text or the retrieval text after synonym transformation, when there is a situation where the second policy text is the same as the first policy text, it indicates that the possibility of this text meeting the requirements is greater. Therefore, screen from them and use the policy text with a higher matching similarity as the target feedback text. When they are all different, only screen out the target feedback text according to the matching similarity of each policy text, which overall improves the screening reliability of the target feedback text.

[0042] In an extended embodiment, as Figure 5 shown, before step S100, the following steps are further included: S10. Sequentially perform framing, windowing, and discrete Fourier transform processing on the collected user voice data to convert it into a spectrum signal, and construct a target Gaussian filter based on the frequency distribution in the spectrum signal.

[0043] S20. Perform a convolution operation on the spectral signal using the target Gaussian filter to obtain the denoised speech spectrum.

[0044] S30. Convert the denoised speech spectrum into a time-domain signal, and convert the time-domain signal after denoising processing into a retrieval text through speech recognition technology.

[0045] As described above, when considering that users may use voice input to retrieve questions, due to user accent problems or surrounding environment problems, voice noise is likely to occur. By constructing a Gaussian filter and denoising the voice, a more accurate and clear retrieval text can be obtained.

[0046] Embodiment 2 As Figure 6 shown, Embodiment 2 of the present application provides a device for obtaining text based on an interactive Q&A model, including: The first acquisition module 100 is used to obtain the first semantic analysis text output by the given interactive Q&A model based on the retrieval text input by the user to the given interactive Q&A model, and extract several initial keywords from the first semantic analysis text. In a specific implementation, the given interactive Q&A model outputs a first semantic analysis text and a result feedback text for the retrieval text, and this application does not directly use the result feedback text, but only obtains the first semantic analysis text. For example, the given interactive Q&A model can be the deepseek large language model, and the output content includes the analysis logic and results.

[0047] As described above, since the reply results output by the Q&A models in the prior art are all prior art or common knowledge, and in order to obtain the query text that is not publicly known in the internal text library in this application, the output results of the interactive Q&A model are not directly used, but the powerful semantic analysis ability of the model is utilized to obtain keywords from the semantic analysis text, and the accurate expansion of keywords is realized in this process, so as to query more reliable text from the massive text in the preset text.

[0048] The second acquisition module 200 is used to screen out several target keywords from several initial keywords according to the relevance between each initial keyword and the retrieval text, and obtain the word weight corresponding to each target keyword; it can be understood that: the initial keywords with a high degree of relevance to the retrieval text are screened out as target keywords.

[0049] Specifically, as Figure 7 shown, the second acquisition module 200 includes: The screening module 201 is used to screen out the initial keywords that are the same as any preset conventional word from several initial keywords to obtain the remaining several first keywords; it can be understood that: the preset conventional word is a habitual word, and the relevance to the technical field where the retrieval text is located is low.

[0050] The second calculation module 202 is configured to obtain the reverse character ranking corresponding to each first keyword according to the position information of each initial keyword in the first semantic analysis text, and perform normalization processing on the order values corresponding to the reverse character rankings of all first keywords by using a preset activation function to obtain the first keyword score corresponding to each first keyword; wherein, the independent variable coefficient of the preset activation function is related to the total number of characters in the retrieval text.

[0051] Specifically, the preset activation function adopts the Sigmoid function.

[0052] Furthermore, in this embodiment, the preset activation function meets the following conditions: S = 2 / (1 + e -σx ), where x refers to the x-th reverse character ranking, S is the normalized value corresponding to x, σ is the preset independent variable coefficient, the value range of σ is between 0 and 1, and σ is negatively correlated with the total number of characters in the retrieval text.

[0053] In a preferred embodiment, σ = 2 / K, where K is the total number of characters in the retrieval text; that is, when K is larger, the first keyword scores corresponding to the earlier first keywords are closer to 1.

[0054] The third calculation module 203 is configured to calculate the word frequency of each first keyword based on the first semantic analysis text, and determine the normalized value corresponding to the word frequency of each first keyword as the second keyword score corresponding to the first keyword itself.

[0055] The screening module 204 is configured to determine the product of the first keyword score and the second keyword score as the final score of the first keyword, and screen the first keywords with the corresponding final scores greater than the preset score threshold as target keywords; in a specific implementation, corresponding weights can also be set for the first keyword score and the second keyword score, and those skilled in the art set the preset score threshold according to actual needs, which will not be elaborated here.

[0056] As described above, when considering the relevance between the first keyword and the retrieval text, the position information and word frequency corresponding to the first keyword are introduced. Since in the semantic analysis text, usually the content in the front part of the text is the logical analysis keywords with high relevance to the retrieval text, such as the decomposition of the problem, and the content in the back is the logical reasoning or examples of multiple possibilities. Therefore, the earlier the keyword, the greater the relevance to the retrieval text, and the higher the corresponding first keyword score. Moreover, the larger the word frequency, the higher the importance of the keyword. Therefore, these two factors are introduced and combined in this application, so that the screened target keywords have higher relevance to the retrieval text, in order to screen out more compliant strategy texts.

[0057] In a specific implementation manner, such asFigure 8 As shown, the second acquisition module 200 further includes: A second determination module 210, configured to determine the normalization result of the final score corresponding to each target keyword as the first word confidence corresponding to the target keyword itself.

[0058] A third determination module 220, configured to input the retrieval text sample into a given interactive Q&A model multiple times to obtain a plurality of second semantic analysis texts, calculate the TF-IDF value corresponding to each target keyword based on the plurality of second semantic analysis texts, and determine the normalization result of the TF-IDF value corresponding to each target keyword as the second word confidence corresponding to the target keyword itself; those skilled in the art know the specific process of calculating the TF-IDF value corresponding to a keyword, which will not be elaborated here.

[0059] A fourth calculation module 230, configured to, for any target keyword, calculate the weighted sum according to the first word confidence corresponding to the target keyword, the second word confidence, and the preset weights corresponding to the first word confidence and the second word confidence respectively, to obtain the word weight corresponding to the target keyword.

[0060] As described above, when calculating the word weight corresponding to each target keyword, the first semantic analysis text where the target keyword is located and the multiple second semantic analysis texts where it is located are considered, that is, the combination of its own factors and external factors makes the calculated word weight more reasonable and reliable.

[0061] A first calculation module 300, configured to calculate the matching similarity between each preset policy text in the preset text library and a plurality of target keywords based on the word weight corresponding to each target keyword, and use the preset policy text with the corresponding matching similarity greater than the similarity threshold as the first policy text; those skilled in the art set the similarity threshold according to actual needs, which will not be elaborated here.

[0062] Further, as Figure 9 shown, the first calculation module 300 includes: A fifth calculation module 301, configured to obtain the word vector corresponding to each target keyword and calculate the matching similarity between each word vector and each preset policy text; the word vector corresponding to the target keyword and the word vector corresponding to the preset policy text have the same vector dimension.

[0063] A sixth calculation module 302, configured to, for any preset policy text, calculate the weighted sum based on the word weight corresponding to each target keyword and the matching similarity between each word vector and each preset policy text, to obtain the matching similarity between each preset policy text and a plurality of target keywords.

[0064] As described above, when calculating the similarity between the target keyword and the predicted policy text, the importance of different keywords is considered. Since the word weights of keywords with a high relevance to the retrieval text are also set relatively high, introducing word weights to calculate the matching similarity can obtain a policy text that better meets the retrieval requirements of the retrieval text.

[0065] The processing module 400 is configured to use the replacement text corresponding to the retrieval text as the new retrieval text, and return to execute steps S100 - S300 to obtain a number of second policy texts; the replacement text corresponding to the retrieval text refers to the text obtained by performing synonym replacement on the given word object in the retrieval text through a preset large language model; it can be understood that: the steps S100 - S300 are respectively executed for a number of replacement texts corresponding to the retrieval text, and each replacement text corresponds to a second policy text.

[0066] As described above, considering the problem that due to the subjectivity of the user, the words used in the retrieval text may be inaccurate, which affects the query result, multiple synonym replacements are performed on the word objects in the retrieval text, so that the obtained second policy text is more comprehensive, which is conducive to obtaining a policy text that meets the reply requirements.

[0067] The first determination module 500 is configured to, when there is a duplicate situation between the first policy text and the second policy text, use the policy text with the highest corresponding matching similarity among all the duplicate first policy texts and all the duplicate second policy texts as the target feedback text; otherwise, use the policy text with the highest corresponding matching similarity among all the first policy texts and all the second policy texts as the target feedback text. For example, when there are two groups of identical first policy texts and second policy texts, screen out the policy text with the highest corresponding matching similarity from the two first policy texts and the two second policy texts as the target feedback text.

[0068] As described above, since both the first policy text and the second policy text are policy texts that meet the retrieval requirements queried from the retrieval text or the retrieval text after synonymous transformation, when there is a situation where the second policy text is the same as the first policy text, it indicates that the possibility of this text meeting the requirements is greater. Therefore, screen from them and use the policy text with a higher matching similarity as the target feedback text. When they are all different, only screen out the target feedback text according to the matching similarity of each policy text. Overall, the screening reliability of the target feedback text is improved.

[0069] Embodiment III An embodiment of the present invention further provides a non-transitory computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to a method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiment.

[0070] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A method for acquiring text based on an interactive question-answering model, characterized in that: The method comprises the following steps: S100, based on a search text input by a user to a given interactive question-answering model, obtaining a first semantic analysis text output by the given interactive question-answering model, and extracting a number of initial keywords from the first semantic analysis text; S200, selecting a plurality of target keywords from a plurality of initial keywords according to the relevance of each initial keyword to the search text, and obtaining a word weight corresponding to each target keyword; S300, based on the word weight corresponding to each target keyword, calculating the matching similarity between each preset strategy text in the preset text library and a number of target keywords, and taking the preset strategy text whose corresponding matching similarity is greater than the similarity threshold as the first strategy text; S400, taking the replacement text corresponding to the search text as a new search text, returning to execute steps S100-S300, and obtaining a plurality of second strategy texts; the replacement text corresponding to the search text refers to the text obtained by performing synonymous replacement on a given word object in the search text through a preset large language model; S500, when there is a repetition between the first strategy text and the second strategy text, the strategy text with the highest corresponding matching similarity among all the repeated first strategy texts and all the repeated second strategy texts is used as the target feedback text; otherwise, the strategy text with the highest corresponding matching similarity among all the first strategy texts and all the second strategy texts is used as the target feedback text.

2. The method for acquiring text based on an interactive question-answering model according to claim 1, characterized in that: The step of selecting a plurality of target keywords from a plurality of initial keywords according to the relevance of each initial keyword to the search text comprises the following steps: S201, filtering out the initial keywords that are the same as any preset regular words from the initial keywords to obtain the remaining first keywords; S202, according to the position information of each initial keyword in the first semantic analysis text, obtain the reverse order ranking of characters corresponding to each first keyword, and use a preset activation function to normalize the order values ​​corresponding to the reverse order ranking of characters of all first keywords to obtain a first keyword score corresponding to each first keyword; wherein the independent variable coefficient of the preset activation function is related to the total number of characters in the search text; S203, calculating the word frequency of each first keyword based on the first semantic analysis text, and determining a normalized value corresponding to the word frequency of each first keyword as a second keyword score corresponding to the first keyword itself; S204: Determine the product of the first keyword score and the second keyword score as the final score of the first keyword, and select the first keyword whose corresponding final score is greater than a preset score threshold as the target keyword.

3. The method for acquiring text based on an interactive question-answering model according to claim 2, characterized in that: Obtain the word weight corresponding to each target keyword through the following steps: S210, determining the normalized result of the final score corresponding to each target keyword as the first word confidence corresponding to the target keyword itself; S220, inputting the search text sample into the given interactive question-answering model several times to obtain several second semantic analysis texts, and calculating the TF-IDF value corresponding to each target keyword based on the several second semantic analysis texts, and determining the normalized result of the TF-IDF value corresponding to each target keyword as the second word confidence corresponding to the target keyword itself; S230, for any target keyword, a weighted sum is calculated based on the first word confidence, the second word confidence corresponding to the target keyword, and the preset weights corresponding to the first word confidence and the second word confidence, respectively, to obtain the word weight corresponding to the target keyword.

4. The method for acquiring text based on an interactive question-answering model according to claim 1, characterized in that: Step S300 includes the following steps: S301, obtaining a word vector corresponding to each target keyword, and calculating a matching similarity between each word vector and each preset strategy text; the word vector corresponding to the target keyword and the word vector corresponding to the preset strategy text have the same vector dimension; S302, for any preset strategy text, based on the word weight corresponding to each target keyword and the matching similarity between each word vector and each preset strategy text, a weighted sum is calculated to obtain the matching similarity between each preset strategy text and several target keywords.

5. The method for acquiring text based on an interactive question-answering model according to claim 1, characterized in that: Before step S100, the following steps are also included: S10, converting the collected user voice data into a spectrum signal after performing frame division, windowing and discrete Fourier transform processing in sequence, and constructing a target Gaussian filter based on the frequency distribution in the spectrum signal; S20, performing a convolution operation on the spectrum signal using a target Gaussian filter to obtain a denoised speech spectrum; S30, converting the denoised speech spectrum into a time domain signal, and converting the denoised time domain signal into a retrieval text by using speech recognition technology.

6. A device for acquiring text based on an interactive question-answering model, characterized in that: The device comprises: A first acquisition module is used to acquire a first semantic analysis text output by a given interactive question-answering model based on a search text input by a user to the given interactive question-answering model, and extract a number of initial keywords from the first semantic analysis text; The second acquisition module is used to select a number of target keywords from the initial keywords according to the relevance of each initial keyword to the search text, and obtain the word weight corresponding to each target keyword; A first calculation module is used to calculate the matching similarity between each preset strategy text in the preset text library and a plurality of target keywords based on the word weight corresponding to each target keyword, and to use the preset strategy text whose corresponding matching similarity is greater than the similarity threshold as the first strategy text; A processing module, used to use the replacement text corresponding to the search text as a new search text, return to execute steps S100-S300, and obtain a plurality of second strategy texts; the replacement text corresponding to the search text refers to the text obtained by performing synonymous replacement on a given word object in the search text through a preset large language model; The first determination module is used to, when there is a repetition between the first strategy text and the second strategy text, use the strategy text with the highest corresponding matching similarity among all the repeated first strategy texts and all the repeated second strategy texts as the target feedback text; conversely, use the strategy text with the highest corresponding matching similarity among all the first strategy texts and all the second strategy texts as the target feedback text.

7. The device for acquiring text based on an interactive question-answer model according to claim 6, characterized in that: The second acquisition module includes: A screening module, used for screening out the initial keywords that are the same as any preset regular words from the initial keywords to obtain the remaining first keywords; The second calculation module is used to obtain the reverse order ranking of characters corresponding to each first keyword according to the position information of each initial keyword in the first semantic analysis text, and to normalize the order values ​​corresponding to the reverse order ranking of characters of all first keywords using a preset activation function to obtain a first keyword score corresponding to each first keyword; wherein the independent variable coefficient of the preset activation function is related to the total number of characters in the search text; A third calculation module is used to calculate the word frequency of each first keyword based on the first semantic analysis text, and determine the normalized value corresponding to the word frequency of each first keyword as the second keyword score corresponding to the first keyword itself; The screening module is used to determine the product of the first keyword score and the second keyword score as the final score of the first keyword, and screen the first keyword whose corresponding final score is greater than a preset score threshold as the target keyword.

8. The device for acquiring text based on an interactive question-answer model according to claim 7, characterized in that: The second acquisition module also includes: A second determination module is used to determine the normalized result of the final score corresponding to each target keyword as the first word confidence corresponding to the target keyword itself; The third determination module is used to input the search text sample into the given interactive question-answering model several times to obtain several second semantic analysis texts, and calculate the TF-IDF value corresponding to each target keyword based on the several second semantic analysis texts, and determine the normalized result of the TF-IDF value corresponding to each target keyword as the second word confidence corresponding to the target keyword itself; The fourth calculation module is used to calculate the word weight corresponding to any target keyword by weighted summing up the first word confidence, the second word confidence and the preset weights corresponding to the first word confidence and the second word confidence.

9. The device for acquiring text based on an interactive question-answer model according to claim 6, characterized in that: The first calculation module includes: A fifth calculation module is used to obtain a word vector corresponding to each target keyword, and calculate the matching similarity between each word vector and each preset strategy text; the word vector corresponding to the target keyword and the word vector corresponding to the preset strategy text have the same vector dimension; The sixth calculation module is used to calculate the weighted sum of any preset strategy text based on the word weight corresponding to each target keyword and the matching similarity between each word vector and each preset strategy text to obtain the matching similarity between each preset strategy text and several target keywords.

10. A non-transitory computer-readable storage medium, wherein at least one instruction or at least one program is stored in the storage medium, characterized in that: The at least one instruction or the at least one program is loaded and executed by the processor to implement the method for acquiring text based on an interactive question-and-answer model as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Semantic-based approximate text search method and device, computer equipment and medium

    CN113434636A

  • Text keyword extraction method and device, equipment and medium

    CN115983261A

  • Semantic retrieval method and system for automatically extracting keywords

    CN118468881A

  • Text matching method and device

    CN119128054A

  • Data recommendation method based on large model

    CN119719353A