Content generation method and device based on keyword semantics and electronic equipment
By constructing target query information and utilizing keyword extraction models and content generation patterns, the problem of information retrieval quality being affected in retrieval enhancement generation was solved, thereby improving the accuracy and effectiveness of the generated results.
Patent Information
- Application Number
- CN202511021768.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-24
AI Technical Summary
The enhancement of retrieval depends on the quality of information retrieval, which affects the accuracy and effectiveness of the results generated by the large language model.
By constructing target query information, using a keyword extraction model to generate an initial keyword information sequence, and combining content generation mode and content knowledge base, two generation modes are designed to improve the accuracy and effectiveness of the generated results.
It improves the accuracy and effectiveness of the generated results by enriching the input of the large language model, reducing redundant information, and ensuring the accuracy and robustness of the generated content.
Smart Images

Figure CN120910231A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer technology, and in particular, to a content generation method and device based on keyword semantics and an electronic device. BACKGROUND
[0002] Retrieval augmented generation (RAG) refers to a technology of enriching the input of a large language model (LLM) through information retrieval to improve the accuracy of the generation result of the large language model. However, retrieval augmented generation depends on the quality of information retrieval, and if no content is retrieved or the retrieved content is irrelevant, it will affect the accuracy and effectiveness of the generation result of the large language model.
[0003] The above information disclosed in this BACKGROUND section is only for the purpose of enhancing the understanding of the background of the present inventive concepts, and therefore, it can contain information that does not form the prior art that is already known in this field to those ordinary skilled ones. SUMMARY
[0004] The summary of the present disclosure is intended to introduce the concepts of the present inventive concepts in a simplified form, which will be described in detail in the detailed description section later. The summary of the present disclosure is not intended to identify key or essential features of the claimed technology nor is it intended to be used to limit the scope of the claimed technology.
[0005] Some embodiments of the present disclosure propose a content generation method and device based on keyword semantics and an electronic device to solve the technical problems mentioned in the background section.
[0006] In a first aspect, some embodiments of the present disclosure provide a keyword semantic-based content generation method, which comprises: constructing target query information according to user query information, basic guide information and scene guide information, wherein the user query information represents query content initiated by a target user, and the scene guide information matches the user query information; generating an initial keyword information sequence according to the target query information and a keyword extraction model, wherein the initial keyword information comprises a main keyword, an extended keyword group and keyword description information, and the keyword extraction model is a model for keyword extraction based on keyword semantics; determining a content generation mode according to the initial keyword information sequence, wherein the content generation mode comprises a first generation mode and a second generation mode; in response to the content generation mode being the first generation mode, generating feedback content according to a content generation model and the initial keyword information sequence; in response to the content generation mode being the second generation mode, generating a search vector matched with each initial keyword information in the initial keyword information sequence to obtain a search vector sequence; performing keyword search according to the search vector sequence and a pre-constructed content knowledge base to determine a target keyword sequence; and generating feedback content according to the target keyword sequence and the content generation model.
[0007] In a second aspect, some embodiments of the present disclosure provide a keyword semantic-based content generation device, which comprises: a construction unit configured to construct target query information according to user query information, basic guide information and scene guide information, wherein the user query information represents query content initiated by a target user, and the scene guide information matches the user query information; a first generation unit configured to generate an initial keyword information sequence according to the target query information and a keyword extraction model, wherein the initial keyword information comprises a main keyword, an extended keyword group and keyword description information, and the keyword extraction model is a model for keyword extraction based on keyword semantics; a determination unit configured to determine a content generation mode according to the initial keyword information sequence, wherein the content generation mode comprises a first generation mode and a second generation mode; a second generation unit configured to, in response to the content generation mode being the first generation mode, generate feedback content according to a content generation model and the initial keyword information sequence; a third generation unit configured to, in response to the content generation mode being the second generation mode, generate a search vector matched with each initial keyword information in the initial keyword information sequence to obtain a search vector sequence; a keyword search unit configured to perform keyword search according to the search vector sequence and a pre-constructed content knowledge base to determine a target keyword sequence; and a fourth generation unit configured to generate feedback content according to the target keyword sequence and the content generation model.
[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having stored thereon one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.
[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0010] The above various embodiments of the present disclosure have the following beneficial effects: through the keyword semantic-based content generation method of some embodiments of the present disclosure, the accuracy and effectiveness of the generation result are improved. Specifically, the reason that cannot guarantee the accuracy and effectiveness of the generation result is that the retrieval enhancement generation depends on the quality of information retrieval. Based on this, the keyword semantic-based content generation method of some embodiments of the present disclosure first constructs target query information according to user query information, basic guide information and scene guide information, wherein the user query information represents the query content initiated by the target user, and the scene guide information matches the user query information. In practice, the input of the large language model enriched by information retrieval should contain as much valid information as possible. However, the original input of the user (user query information) may have too much redundant information and poor semantic coherence. Based on this, the present disclosure improves the query information content from the perspective of general guidance and scene guidance by combining basic guide information and scene guide information, and guides the keyword extraction of the subsequent keyword extraction model. Secondly, according to the target query information and the keyword extraction model, an initial keyword information sequence is generated, wherein the initial keyword information includes: main keywords, extended keyword groups, keyword description information, and the keyword extraction model is a model for keyword extraction based on keyword semantics. Through keyword extraction, the core content of the user query information can be effectively refined. In addition, considering that the keywords contained in the user query information may have inaccurate expression, the present disclosure expands the keyword abundance by expanding the keywords during keyword extraction. Then, according to the initial keyword information sequence, a content generation mode is determined, wherein the content generation mode includes: a first generation mode and a second generation mode. Further, in response to the content generation mode being the first generation mode, a feedback content is generated according to a content generation model and the initial keyword information sequence. In addition, in response to the content generation mode being the second generation mode, a retrieval vector matching each initial keyword information in the initial keyword information sequence is generated to obtain a retrieval vector sequence. Then, according to the retrieval vector sequence and a pre-constructed content knowledge base, keyword retrieval is performed to determine a target keyword sequence. Finally, according to the target keyword sequence and the content generation model, a feedback content is generated. In practice, retrieval enhancement generation aims to improve the output accuracy and effectiveness of the large language model, but the retrieval process may increase the response time, therefore, the present disclosure designs two generation modes (i.e. the first generation mode and the second generation mode), that is, directly generating feedback content through extracted keywords (main keywords and extended keywords) and a content generation model, or generating feedback content through retrieval enhancement generation to obtain a keyword sequence and a content generation model. By designing two different processing methods of feedback content production accuracy, the robustness of content generation is enriched.In summary, the content generation method according to the present disclosure improves the accuracy and effectiveness of the generation result. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, aspects and advantages of the present disclosure will become more apparent by reference to the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements throughout. It should be noted that the drawings are schematic and elements and features do not necessarily appear to scale.
[0012] Figure 1 is a flowchart of some embodiments of the keyword semantic based content generation method according to the present disclosure;
[0013] Figure 2 is a diagrammatic view of the graph structure of a keyword graph;
[0014] Figure 3 is a diagrammatic view of some embodiments of the keyword semantic based content generation apparatus according to the present disclosure;
[0015] Figure 4 is a diagrammatic view of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and should not be construed as limiting the scope of protection of the present disclosure.
[0017] It should also be noted that, for the sake of brevity, only the parts of the drawings that are relevant to the present disclosure are shown. The embodiments and features in the present disclosure and the embodiments can be combined with each other where there is no conflict.
[0018] It should be noted that the terms “first”, “second”, and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0019] It should be noted that the terms “one”, “multiple” in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as “one or more”.
[0020] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0021] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0022] Reference Figure 1 Fig. 1 shows a flow 100 of some embodiments of the keyword semantic based content generation method according to the present disclosure. The keyword semantic based content generation method includes the following steps:
[0023] Step 101, constructing target query information according to user query information, basic guidance information and scene guidance information.
[0024] In some embodiments, the subject (e.g., a computing device) of the keyword semantic based content generation method can construct target query information according to user query information, basic guidance information and scene guidance information.
[0025] The user query information represents the query content initiated by the target user. The basic guidance information is guidance information for guiding the large language model (e.g., keyword extraction model) to process the user query information. The scene guidance information is matched with the user query information. The scene guidance information is guidance information for guiding the large language model (e.g., keyword extraction model) to process the user query information, which is matched with the query scene corresponding to the user query information. The target query information can represent the query information obtained by combining the user query information, the basic guidance information and the scene guidance information. The target user can be the user who initiates the query content.
[0026] As an example, the user query information, the basic guidance information and the scene guidance information can be combined in JSON (JavaScript Object Notation) format, and the specific format of the target query information can be as follows:
[0027] {“information type”: guidance information;
[0028] “information content”: basic guidance information;
[0029] “information type”: guidance information;
[0030] “information content”: scene guidance information;
[0031] “information type”: query information;
[0032] “information content”: user query information;}.
[0033] As a further example, the user query information can be "query the current hottest-selling electronic products". The basic guidance information can be "you are a text extraction expert, please extract the content, summarize and expand the content of the user query information, and extract the key query words and output, the output format is separated by ';', and correct the wrong words in the user query information". The scenario guidance information can be "you are an electronic product sales expert, please expand the user query information to make the subsequent content query more accurate". Therefore, the constructed target query information can be:
[0034] { "information type": guidance information;
[0035] "information content": "query the current hottest-selling electronic products". The basic guidance information can be "you are a text extraction expert, please extract the content, summarize and expand the content of the user query information, and extract the key query words and output, the output format is separated by ';', and correct the wrong words in the user query information";
[0036] "information type": guidance information;
[0037] "information content": "you are an electronic product sales expert, please expand the user query information to make the subsequent content query more accurate";
[0038] "information type": query information;
[0039] "information content": "query the current hottest-selling electronic products";}.
[0040] It should be noted that the above computing device can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the above-mentioned hardware devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made herein.
[0041] In some optional implementations of some embodiments, before the above constructing the target query information according to the user query information, the basic guidance information and the scenario guidance information, the above method further comprises:
[0042] Step S1: scanning the user query information queue.
[0043] Wherein, the above user query information queue is a message queue corresponding to the target user, used for temporarily storing the user query information initiated by the target user.
[0044] In practice, by setting a corresponding user query information queue for each user, the information isolation of query information between different users can be ensured, and the query blocking of the remaining users caused by the high-frequency query of a certain user can be avoided. In addition, the length of the information queue can be flexibly scaled according to the number of user query information initiated by the user. Specifically, when the target user initiates information query for the first time, a user query information queue corresponding to the target user is created. Whether there is unprocessed user query information in the user query information queue corresponding to the target user can be determined by periodic scanning.
[0045] Step S2: In response to the above-mentioned non-empty user query information queue and the invalid corresponding information token, the identity of the target user corresponding to the client is re-verified.
[0046] Among them, the information token is distributed and bound when the user query information queue is created, and is used to control the life cycle of the user query information queue.
[0047] In practice, although a corresponding user query information queue is set for each user, it can have the beneficial effects as described above in step S1, but as the number of users increases, the user query information queues that need to be created and managed also increase, especially when there are many low-frequency queries, which can result in a large number of idle message queues, thereby occupying computing resources. Therefore, by setting an information token to control the life cycle of the user query information queue. Specifically, the information token corresponding to each user query information queue is set with a queue life length, when new user query information is added to the user query information queue, the corresponding queue life length is updated, otherwise when the queue life length is 0, the user query information queue is marked as a state to be logged off, and when the information token is not updated after a predetermined time, the user query information queue is released. In this way, the efficient use of computing resources is ensured.
[0048] Step S3: In response to the identity re-verification, the token validity period of the information token is updated.
[0049] In practice, the above-mentioned execution subject can verify the online state of the target user through identity verification. When the target user is not online, it means that the identity verification fails, and when the target user is online, it means that the identity verification passes. In particular, when the identity verification passes, the token validity period of the information token is updated to extend the queue life length corresponding to the user query information queue.
[0050] Step S4: In response to the above-mentioned non-empty user query information queue and the valid corresponding information token, the user query information at the head position of the above-mentioned user query information queue is taken out.
[0051] Step S5: Determine the query scenario corresponding to the above-mentioned user query information.
[0052] In practice, considering the response speed limit of content generation (ensuring high response speed as much as possible), therefore, for example, the present disclosure can adopt a Word2Vec model as the backbone network to extract the scene semantic features of the user query information, and through a scene label classifier connected after the Word2Vec model, the query scene corresponding to the user query information is classified. Specifically, the Word2Vec model + scene label classifier is trained as a whole model in a supervised training manner. The training samples and sample labels are composed of the collected historical user query information and the corresponding scene labels. The query scene can be represented by the corresponding query scene identifier.
[0053] As an example, the query scene can be a "knowledge answering scene", or a "product search scene".
[0054] Step S6: recalling candidate scene guidance information corresponding to the above query scene from the pre-constructed scene guidance information library to obtain a candidate scene guidance information set.
[0055] Among them, the scene guidance information library can be a pre-constructed database that stores pre-set scene guidance information corresponding to different query scenes.
[0056] In practice, a recall algorithm can be used to recall the top K candidate scene guidance information with the highest recommendation value from the scene guidance information library as the candidate scene guidance information set.
[0057] Step S7: screening the candidate scene guidance information that meets the screening condition from the above candidate scene guidance information set as the scene guidance information.
[0058] Among them, the screening condition can be that the candidate scene guidance information corresponds to the highest recommendation value.
[0059] In some optional implementations of some embodiments, the above execution subject constructs the target query information according to the user query information, the basic guidance information and the scene guidance information, including:
[0060] Step S1: decomposing the above guidance information and the above scene guidance information into guidance tasks respectively to obtain a guidance task set.
[0061] In practice, the basic guidance information and the scene guidance information can have multiple task steps (guidance tasks). The guidance tasks can be decomposed in a guidance task decomposition manner to step-by-step guide and improve the rationality of the subsequent guidance process in sequence and the guidance efficiency. Specifically, for example, since the guidance information (basic guidance information and scene guidance information) is pre-configured, the guidance task decomposition can be performed in a preset decomposition rule manner. For another example, the guidance task decomposition can also be performed in combination with a large language model to take the guidance information as input.
[0062] As an example, the basic guidance information can be "You are a text extraction expert, please extract the content, summarize, and expand the content of the user query information, and extract the key query words and output, the output format is separated by ';', and correct the misspelled words in the user query information", which corresponds to 3 guidance tasks, namely guidance task A, guidance task B, and guidance task C. Among them, guidance task A is "You are a text extraction expert, please extract the content and summarize the user query information, and extract the key query words and output, the output format is separated by ';'. Guidance task B is "You are a text extraction expert, please expand the content of the user query information, and extract the key query words and output, the output format is separated by ';'. Guidance task C is "Correct the misspelled words in the user query information". The scene guidance information can be "You are an electronic product sales expert, please expand the user query information to make the subsequent content query more accurate", which corresponds to one guidance task, namely guidance task C. Among them, guidance task C is "You are an electronic product sales expert, please expand the user query information".
[0063] Step S2: determining guidance task description information corresponding to each guidance task in the guidance task set.
[0064] The guidance task description information includes guidance task features and guidance task types. The guidance task type can represent the task type of the guidance task.
[0065] As an example, when the large language model is not used for guidance task decomposition, first, the guidance task type corresponding to the guidance task can be determined in a template matching manner. Then, the semantic features of the guidance task are extracted to obtain the guidance task features. For another example, when the large language model is used for guidance task decomposition, the features of the guidance task extracted by the large language model in the task decomposition process can be taken as the guidance task features, and the corresponding guidance task type is classified and determined.
[0066] Step S3: performing approximate guidance task fusion on the guidance tasks in the guidance task set according to the guidance task description information corresponding to the guidance tasks to obtain a fused guidance task set.
[0067] In practice, the similarity between the guide tasks can be determined by feature similarity calculation (e.g., calculating cosine similarity), and the guide tasks with a similarity greater than a threshold value can be merged. In this way, the number of guide tasks can be compressed, and similar guide tasks can be merged to avoid repeated execution of similar guide tasks, thereby improving the execution efficiency of the guide tasks.
[0068] As an example, the guide task set can include guide task A, guide task B, guide task C, and guide task D. Among them, the feature similarity between guide task B and guide task D is greater than a threshold value, so the resulting merged guide task set can include merged guide task A (guide task A), merged guide task B (guide task C), and merged guide task C (a guide task obtained by merging guide task B and guide task D). In particular, merged guide task C can be "You are a text extraction expert and an electronic product sales expert. Please expand the content of the user query information and extract the key query words and output. The output format is separated by ';'to make the subsequent content query more accurate."
[0069] Step S4: Determine the guide task priority corresponding to each merged guide task in the merged guide task set.
[0070] The guide task priority represents the execution order of the merged guide task.
[0071] In practice, for different guide task types, the corresponding guide task priority can be set in advance. Therefore, the corresponding guide task priority can be determined according to the guide task type corresponding to the merged guide task.
[0072] Step S5: According to the guide task priority corresponding to the merged guide task, the merged guide tasks in the merged guide task set are rearranged to obtain a rearranged guide task sequence.
[0073] As an example, since merged guide task B (guide task C) is to correct the misspelled words in the user query information, the rearranged guide task sequence can be [merged guide task B, merged guide task A, merged guide task C].
[0074] Step S6: According to the rearranged guide task sequence, the basic guide information and the scene guide information are fused to obtain the fused guide information.
[0075] As an example, the fused guidance information can be "correct the misspelling in the user query information, and as a text extraction expert and an electronic product sales expert, perform content extraction, summarization, and content expansion on the user query information, extract the key query words therein, and output, with the output format being separated by';'to make subsequent content query more accurate".
[0076] Step S7: combining the user query information and the fused guidance information according to a preset combination format to obtain the target query information.
[0077] In practice, the user query information and the fused guidance information can be combined according to a preset combination format in JSON format to obtain the target query information.
[0078] Step 102: generating an initial keyword information sequence according to the target query information and a keyword extraction model.
[0079] In some embodiments, the execution subject can generate an initial keyword information sequence according to the target query information and a keyword extraction model.
[0080] The initial keyword information includes a main keyword, an extended keyword group, and keyword description information. The keyword extraction model is a model for keyword extraction based on keyword semantics. The main keyword is a keyword extracted from the user query information. The extended keywords in the extended keyword group are keywords similar in meaning to the main keyword. The keyword description information is used to describe the main keyword. In practice, the keyword description information can include part of speech and word position. The part of speech represents the part of speech of the main keyword. The word position represents the word position of the main keyword in the user query information. Specifically, the keyword extraction model can use a large language model. In particular, the large language model processes the user query information with the fused guidance information included in the target query information to obtain the initial keyword information sequence.
[0081] In some optional implementations of some embodiments, the execution subject generates an initial keyword information sequence according to the target query information and a keyword extraction model, including:
[0082] Step S1: inputting the target query information into the keyword extraction model to generate a main keyword sequence.
[0083] In practice, the keyword extraction model can use a large language model.
[0084] As an example, the main keyword sequence can be: ["query"; "hot-selling"; "electronic products"].
[0085] Step S2: for each stem keyword in the stem keyword sequence, the following processing steps are performed:
[0086] Step S21: extract the stem keyword feature corresponding to the stem keyword.
[0087] In practice, since the keyword extraction model performs semantic understanding in the process of keyword extraction of user query information, the semantic features corresponding to the stem keyword in the keyword extraction process can be used as the stem keyword feature corresponding to the stem keyword.
[0088] Step S22: randomly mask the word granularity of the content in the user query information except the stem keyword to obtain the masking position.
[0089] The masking position can represent the word position of the masked word in the user query information. The word position represents the position of the word in the user query information.
[0090] As an example, taking the stem keyword "electronic product" as an example, for the user query information "query the current hottest-selling electronic product", the word position corresponding to "query" can be [0], the word position corresponding to "current" can be [1], the word position corresponding to "hottest-selling" can be [2], and the word position corresponding to "electronic product" can be [3]. The randomly determined masking position can correspond to "current", and therefore the masking position can be [1].
[0091] Step S23: project and mask the position corresponding to the masking position in the scene semantic feature corresponding to the query scene to obtain the masked scene semantic feature.
[0092] In practice, since the extraction of the scene semantic feature is a semantic extraction based on the user query information, and the extraction process is a linear extraction process, it can be understood that the extracted scene semantic feature has a mapping relationship with the word position, and therefore the position corresponding to the masking position in the scene semantic feature can be updated to 0.
[0093] Step S24: perform extension keyword matching according to the stem keyword feature to obtain the candidate extension keyword group corresponding to the stem keyword.
[0094] In practice, the keyword feature matching method can be used to filter out keywords similar in meaning to the stem keyword as the candidate extension keyword group corresponding to the stem keyword.
[0095] Step S25: extract the extension keyword feature corresponding to each candidate extension keyword in the candidate extension keyword group.
[0096] In practice, the candidate extended keyword can be converted into a word vector as an extended keyword feature corresponding to the candidate extended keyword. In the process of converting the word vector, feature dimensionality needs to be increased to ensure that the extended keyword feature is consistent with the feature dimensionality of the masked scene semantic feature, facilitating subsequent feature similarity calculation.
[0097] Step S26: Determine the feature similarity between the extended keyword feature corresponding to each candidate extended keyword in the candidate extended keyword group and the masked scene semantic feature.
[0098] In practice, the feature similarity between the extended keyword feature corresponding to the candidate extended keyword and the masked scene semantic feature can be determined by calculating the cosine similarity. Specifically, the words in the user query information other than the main keyword can limit the main keyword, and too many limitations can lead to the generation of extended keywords with similar meanings to the main keyword. Therefore, by determining the masking position and projecting and masking the feature, the constraint of the word in the scene semantic feature dimension is reduced, thereby increasing the number of extended keywords obtained by extension.
[0099] Step S27: From the candidate extended keyword group, filter out the candidate extended keyword that meets the extended keyword filtering condition as the extended keyword group corresponding to the main keyword.
[0100] The extended keyword filtering condition can be that the feature similarity corresponding to the candidate extended keyword is greater than a preset threshold.
[0101] Step S28: Generate keyword description information corresponding to the main keyword based on the main keyword feature.
[0102] In practice, based on the main keyword feature, the part of speech of the main keyword can be determined through part of speech analysis. At the same time, the word position of the main keyword in the target query information is determined to construct the keyword description information corresponding to the main keyword.
[0103] Step S209: Generate the initial keyword information corresponding to the main keyword in the initial keyword information sequence based on the main keyword, the extended keyword group corresponding to the main keyword, and the keyword description information corresponding to the main keyword.
[0104] In practice, the main keyword, the extended keyword group corresponding to the main keyword, and the keyword description information corresponding to the main keyword can be combined according to a preset format to obtain the initial keyword information corresponding to the main keyword.
[0105] Step 103: Determine the content generation mode based on the initial keyword information sequence.
[0106] In some embodiments, the execution subject can determine the content generation mode according to the initial keyword information sequence.
[0107] The content generation mode includes a first generation mode and a second generation mode.
[0108] The first generation mode represents that the feedback content is generated by the content generation model only in combination with the initial keyword information sequence. The second generation mode represents that the feedback content is generated by the content generation model in combination with the content knowledge base for keyword retrieval on the basis of the initial keyword information sequence.
[0109] In practice, whether the initial keyword information sequence can meet the generation requirement of the feedback content in the query scene corresponding to the user query information can be determined by a preset mode matching rule. For example, the mode matching rule can include that the ratio of the first word quantity to the second word quantity is less than a ratio threshold. The second word quantity can be the total number of stem keywords or extended keywords included in the initial keyword information in the initial keyword information sequence. The second word quantity can be the number of hits of the stem keywords or the extended keywords included in the initial keyword information in the initial keyword information sequence and the professional words in a professional word library. The professional word library is a word library containing special domain-specific proper nouns constructed in advance. As the ratio of the first word quantity to the second word quantity is larger, the more professional words are represented, the more professional the query field and the query content corresponding to the user query information are, and at this time, it can be difficult to generate accurate feedback content in the first generation mode. In addition, other mode matching rules can also be set to select the content generation mode.
[0110] Specifically, keyword retrieval in combination with the content knowledge base requires a certain retrieval time. For relatively simple user query information, the initial keyword information sequence extracted on the basis thereof can generate relatively accurate feedback content in combination with the content generation model, and therefore, in this case, it is not necessary to further combine the content knowledge base for keyword retrieval, thereby improving the speed of content generation. For relatively professional user query information or relatively complex user query information, the initial keyword information sequence extracted on the basis thereof is difficult to generate accurate feedback content in combination with the content generation model, and at this time, the keywords need to be improved in combination with the content knowledge base to improve the accuracy of the feedback content generated by the subsequent content generation model. In particular, the content generation model can adopt a common large language model.
[0111] In some optional implementations of some embodiments, the execution subject determines the content generation mode according to the initial keyword information sequence, including:
[0112] Step S1: generating a keyword graph according to the initial keyword information sequence.
[0113] The keyword graph includes the same number of keyword layers as the number of initial keyword information in the initial keyword information sequence, and each keyword layer includes the number of keyword nodes +1 corresponding to the number of extended keywords in the extended keyword group included in the initial keyword information.
[0114] As an example, the initial keyword information sequence can include: initial keyword information A, initial keyword information B, and initial keyword information C. The initial keyword information A includes: main keyword AA. The extended keyword group included in the initial keyword information A includes: extended keyword AB, extended keyword AC, and extended keyword AD. The initial keyword information B includes: main keyword BA. The extended keyword group included in the initial keyword information B includes: extended keyword BB and extended keyword BC. The initial keyword information C includes: main keyword CA. The extended keyword group included in the initial keyword information C includes: extended keyword CB, extended keyword CC, extended keyword CD, and extended keyword CE. Referring to Figure 2 The graph structure diagram of the keyword graph is shown. The keyword graph includes: keyword layer A, keyword layer B, and keyword layer C. The distance between two adjacent keyword layers is determined by the word distance of the corresponding main keyword in the user query information (the edge length of the edge between the keyword layers in the keyword graph is represented by the number of intermediate interval words). For a keyword layer, the word distance between the main keyword and the extended keyword (the edge length of the edge between the main keyword and the extended keyword in the keyword graph) is determined by the word similarity. Thus, a graph structure with multiple main keywords as the radiation center is obtained.
[0115] Step S2: performing graph feature extraction on the keyword graph to generate keyword graph features.
[0116] In practice, a GNN (Graph Neural Networks) model can be used to perform graph feature extraction on the keyword graph to obtain keyword graph features.
[0117] Step S3: determining the content generation mode according to the keyword graph features and the content generation mode classifier.
[0118] The content generation mode classifier is a binary classifier. In particular, in the model training phase, the GNN model and the content generation mode classifier are supervised as a whole. The training sample is a keyword graph constructed for historical retrieval information, and the sample label is the corresponding content generation mode (first content generation mode or second content generation mode).
[0119] In response to the content generation mode being the first generation mode, the feedback content is generated according to the content generation model and the initial keyword information sequence.
[0120] In some embodiments, in response to the content generation mode being the first generation mode, the feedback content is generated according to the content generation model and the initial keyword information sequence.
[0121] In practice, the above execution subject can input the initial keyword information sequence into the above content generation model to obtain the feedback content. Specifically, the content generation model can be a large language model.
[0122] In response to the content generation mode being the second generation mode, a search vector matched with each initial keyword information in the initial keyword information sequence is generated to obtain a search vector sequence.
[0123] In some embodiments, in response to the content generation mode being the second generation mode, a search vector matched with each initial keyword information in the initial keyword information sequence is generated to obtain a search vector sequence. In practice, for example, the initial key information can be overall feature encoded to obtain the corresponding search vector. The main keywords and extended keyword groups included in the initial keyword information can also be feature encoded to obtain the corresponding search vector. In particular, in the feature encoding process, the output feature dimension needs to be constrained, such as 128 dimensions or 256 dimensions, to avoid excessively long feature dimensions, which would lead to high computational complexity in subsequent keyword retrieval combined with the search vector.
[0124] In some optional implementations of some embodiments, the above execution subject generates a search vector matched with each initial keyword information in the above initial keyword information sequence, including:
[0125] Step S1: Determine the scene matching degree of the main keywords included in the initial keyword information and the query scene as the initial voting probability corresponding to the main keywords.
[0126] In practice, the scene matching degree (similarity) of the main keywords and the query scene can be determined by an inter-word similarity calculation method (for example, calculating the cosine similarity between the feature vectors corresponding to the words) as the corresponding initial voting probability.
[0127] As an example, see the following pseudo code:
[0128] from sentence_transformers import SentenceTransformer
[0129] from sklearn.metrics.pairwise import cosine_similarity
[0130] model =
[0131] SentenceTransformer('paraphrase-multilingual-MiniLM-L12-v2')
[0132] vec_1 = model.encode("keyword")
[0133] vec_2 = model.encode("query scenario")
[0134] similarity=cosine_similarity([vec_1],[vec_2])[0][0]
[0135] Among them, "keywords" refers to the main keywords or extended keywords. "similarity" represents the degree of scene matching.
[0136] Step S2: Determine the matching degree between each extended keyword in the extended keyword group included in the initial keyword information and the scenario corresponding to the above query scenario, and use it as the initial voting probability for the extended keyword.
[0137] In practice, the initial voting probability corresponding to the main keyword can be calculated using step S2 above, and the initial voting probability corresponding to the extended keyword can be determined. This will not be elaborated further here.
[0138] Step S3: Create the first target number of voting objects.
[0139] The number of the first objective must be an odd number. For example, the default number of the first objective is 3.
[0140] Step S4: Using the first target number of voting objects, select the second target number of candidate keywords from the main keywords and extended keyword groups included in the initial keyword information based on the initial voting probability corresponding to the main keywords and the initial voting probability corresponding to the extended keywords.
[0141] In practice, each voting object votes for the stem keywords and the extended keywords according to the initial voting probability. In particular, the higher the initial voting probability, the higher the probability of the voting object voting for it (stem keyword or extended keyword), so as to obtain the top K (second target number) keywords (stem keywords or extended keywords) corresponding to the highest number of votes as the second target number of candidate keywords. For example, the second target number can be 2-4. In this way, the purpose of word expansion can be guaranteed, and the problem of word vector dimension explosion caused by too many words can be avoided.
[0142] Step S5: Construct a word vector corresponding to each candidate keyword in the second target number of candidate keywords to obtain a word vector sequence.
[0143] In practice, during the process of calculating the scene matching degree in steps S1 and S2, the keyword (stem keyword or extended keyword) has been constructed into a word vector. At this time, the word vector (vec_1) in the process of calculating the scene matching degree in steps S1 and S2 needs to be compressed (to reduce the calculation complexity in the subsequent retrieval process) as the corresponding word vector.
[0144] Step S6: Splice the word vectors in the above word vector sequence to obtain a spliced word vector.
[0145] As an example, the word vector sequence can include: word vector A, word vector B, word vector C. The spliced word vector = word vector A + word vector B + word vector C.
[0146] Step S7: Perform word vector compression on the above spliced word vector to obtain a retrieval vector corresponding to the initial keyword information.
[0147] In practice, the entire word vector compression is performed on the spliced word vector to ensure that the retrieval vector corresponding to the initial keyword information is consistent with the feature dimension of the word vector corresponding to the keyword in the content knowledge base, facilitating subsequent keyword retrieval.
[0148] Step 106: Perform keyword retrieval according to the retrieval vector sequence and the pre-constructed content knowledge base to determine the target keyword sequence.
[0149] In some embodiments, the above execution subject can perform keyword retrieval according to the retrieval vector sequence and the pre-constructed content knowledge base to determine the target keyword sequence. The content knowledge base is used to store pre-constructed keywords (guide words) and their corresponding word vectors.
[0150] In practice, for each search vector in the search vector sequence, the search vector is compared with the word vector corresponding to the keyword (prompt word) in the pre-constructed content knowledge base to retrieve the keyword (prompt word) as the target keyword sequence.
[0151] In some optional implementations of some embodiments, the above execution subject performs keyword retrieval according to the above search vector sequence and the pre-constructed content knowledge base to determine the target keyword sequence, including:
[0152] For each search vector in the above search vector sequence, keyword retrieval is performed in the above content knowledge base according to the above search vector by means of jump similarity calculation to retrieve the target keyword matching the above search vector.
[0153] In practice, conventional similarity calculation (for example, cosine similarity) solves the corresponding similarity by means of pair-wise vector value calculation. When the feature dimension of the vector is long or the keyword retrieval frequency is high, the calculation amount is large, thereby causing an increase in load. Therefore, the present disclosure does not perform similarity calculation on the preset proportion of vector values in the calculation process of the search vector and the word vector by means of jump similarity calculation, although some accuracy may be lost, but the corresponding calculation speed can be improved. In particular, the specific value of the preset proportion can be dynamically adjusted according to the keyword retrieval accuracy. It can also be dynamically adjusted according to the load pressure.
[0154] Step 107, generating feedback content according to the target keyword sequence and the content generation model.
[0155] In some embodiments, the above execution subject can generate feedback content according to the target keyword sequence and the content generation model. In practice, the above execution subject can take the target keyword sequence as the input of the content generation model to obtain the corresponding feedback content.
[0156] The above various embodiments of the present disclosure have the following beneficial effects: through the keyword semantic-based content generation method of some embodiments of the present disclosure, the accuracy and effectiveness of the generation result are improved. Specifically, the reason that cannot guarantee the accuracy and effectiveness of the generation result is that the retrieval enhancement generation depends on the quality of information retrieval. Based on this, the keyword semantic-based content generation method of some embodiments of the present disclosure first constructs target query information according to user query information, basic guide information and scene guide information, wherein the user query information represents the query content initiated by the target user, and the scene guide information matches the user query information. In practice, the input of the large language model enriched by information retrieval should contain as much valid information as possible. However, the original input of the user (user query information) may have too much redundant information and poor semantic coherence. Based on this, the present disclosure improves the query information content from the perspective of general guidance and scene guidance by combining basic guide information and scene guide information, and guides the keyword extraction of the subsequent keyword extraction model. Secondly, according to the target query information and the keyword extraction model, an initial keyword information sequence is generated, wherein the initial keyword information includes: main keywords, extended keyword groups, keyword description information, and the keyword extraction model is a model for keyword extraction based on keyword semantics. Through keyword extraction, the core content of the user query information can be effectively refined. In addition, considering that the keywords contained in the user query information may have inaccurate expression, the present disclosure expands the keyword abundance by expanding the keywords during keyword extraction. Then, according to the initial keyword information sequence, a content generation mode is determined, wherein the content generation mode includes: a first generation mode and a second generation mode. Further, in response to the content generation mode being the first generation mode, a feedback content is generated according to a content generation model and the initial keyword information sequence. In addition, in response to the content generation mode being the second generation mode, a retrieval vector matching each initial keyword information in the initial keyword information sequence is generated to obtain a retrieval vector sequence. Then, according to the retrieval vector sequence and a pre-constructed content knowledge base, keyword retrieval is performed to determine a target keyword sequence. Finally, according to the target keyword sequence and the content generation model, a feedback content is generated. In practice, retrieval enhancement generation aims to improve the output accuracy and effectiveness of the large language model, but the retrieval process may increase the response time, therefore, the present disclosure designs two generation modes (i.e. the first generation mode and the second generation mode), that is, directly generating feedback content through extracted keywords (main keywords and extended keywords) and a content generation model, or generating feedback content through retrieval enhancement generation to obtain a keyword sequence and a content generation model. By designing two different processing methods of feedback content production accuracy, the robustness of content generation is enriched.In summary, the content generation method of the present disclosure improves the accuracy and effectiveness of the generation result.
[0157] Further referring to Figure 3 , as an implementation of the method shown in the above figures, the present disclosure provides some embodiments of a content generation device based on keyword semantics, which correspond to the method embodiments shown in Figure 1 , and the content generation device based on keyword semantics can be applied in various electronic devices.
[0158] As shown in Figure 3 , the content generation device based on keyword semantics 300 of some embodiments includes a construction unit 301, a first generation unit 302, a determination unit 303, a second generation unit 304, a third generation unit 305, a keyword retrieval unit 305, and a fourth generation unit 306. The construction unit 301 is configured to construct target query information according to user query information, basic guidance information, and scene guidance information, wherein the user query information represents the query content initiated by a target user, and the scene guidance information matches the user query information. The first generation unit 302 is configured to generate an initial keyword information sequence according to the target query information and a keyword extraction model, wherein the initial keyword information includes a main keyword, an extended keyword group, and keyword description information, and the keyword extraction model is a model for keyword extraction based on keyword semantics. The determination unit 303 is configured to determine a content generation mode according to the initial keyword information sequence, wherein the content generation mode includes a first generation mode and a second generation mode. The second generation unit 304 is configured to generate feedback content according to a content generation model and the initial keyword information sequence in response to the content generation mode being the first generation mode. The third generation unit 305 is configured to generate a search vector matching each initial keyword information in the initial keyword information sequence to obtain a search vector sequence in response to the content generation mode being the second generation mode. The keyword retrieval unit 306 is configured to perform keyword retrieval according to the search vector sequence and a pre-constructed content knowledge base to determine a target keyword sequence. The fourth generation unit 307 is configured to generate feedback content according to the target keyword sequence and the content generation model. It can be understood that the units described in the content generation device based on keyword semantics 300 correspond to the respective steps in the method described with reference to Figure 1 . Therefore, the operations, features, and beneficial effects described above for the method also apply to the content generation device based on keyword semantics 300 and the units contained therein, which will not be described here again.
[0159] The following will be described with reference to Figure 4It illustrates a schematic diagram of the structure of an electronic device (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of this disclosure. Figure 4 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The memory may include a non-volatile storage medium and internal memory. The non-volatile storage medium may store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to perform any of the methods described above. The processor provides computational and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the execution of the computer program in the non-volatile storage medium; when executed by the processor, the computer program causes the processor to perform any of the methods described above. The network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the computer device to which the present disclosure is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0160] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0161] In one embodiment, the processor is configured to run a computer program stored in the memory to implement the following steps: constructing target query information according to user query information, basic guidance information and scene guidance information, wherein the user query information represents the query content initiated by a target user, and the scene guidance information is matched with the user query information; generating an initial keyword information sequence according to the target query information and a keyword extraction model, wherein the initial keyword information includes a main keyword, an extended keyword group and keyword description information, and the keyword extraction model is a model for keyword extraction based on keyword semantics; determining a content generation mode according to the initial keyword information sequence, wherein the content generation mode includes a first generation mode and a second generation mode; in response to the content generation mode being the first generation mode, generating feedback content according to a content generation model and the initial keyword information sequence; in response to the content generation mode being the second generation mode, generating a search vector matched with each initial keyword information in the initial keyword information sequence to obtain a search vector sequence; performing keyword search according to the search vector sequence and a pre-constructed content knowledge base to determine a target keyword sequence; and generating feedback content according to the target keyword sequence and the content generation model.
[0162] The embodiments of the present disclosure further provide a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed, the method can refer to each embodiment of the method of the present disclosure.
[0163] The computer readable storage medium can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card and the like.
[0164] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such a process, method, article or system. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or system including the element.
[0165] The above description is merely some of the preferred embodiments of the present disclosure, and an explanation of the principles of the technology employed. It will be appreciated by those skilled in the art that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalents
[0166] without departing from the above inventive concept. For example, the above-described features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) with similar functions are replaced with each other to form technical solutions. For example, the above-described features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) with similar functions are replaced with each other to form technical solutions.
[0167]
[0168] without departing from the above inventive concept. For example, the above-described features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) with similar functions are replaced with each other to form technical solutions. For example, the above-described features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) with similar functions are replaced with each other to form technical solutions.
Claims
1. A keyword semantic-based content generation method, characterized by, The method comprises the steps of: According to the user query information, the basic guide information and the scene guide information, the target query information is constructed, wherein the user query information represents the query content initiated by the target user, and the scene guide information matches the user query information; According to the target query information and the keyword extraction model, an initial keyword information sequence is generated, wherein the initial keyword information includes: main keywords, extended keyword groups, and keyword description information, and the keyword extraction model is a model based on keyword semantics for keyword extraction; According to the initial keyword information sequence, the content generation mode is determined, wherein the content generation mode includes: a first generation mode and a second generation mode; In response to the content generation mode being the first generation mode, feedback content is generated according to the content generation model and the initial keyword information sequence; In response to the content generation mode being the second generation mode, a search vector matching each initial keyword information in the initial keyword information sequence is generated to obtain a search vector sequence; According to the search vector sequence and the pre-constructed content knowledge base, keyword retrieval is performed to determine a target keyword sequence; According to the target keyword sequence and the content generation model, feedback content is generated.
2. The method of claim 1, wherein, Before the target query information is constructed according to the user query information, the basic guide information and the scene guide information, the method further comprises: Scanning the user query information queue, wherein the user query information queue is a message queue corresponding to the target user and used for temporarily storing the user query information initiated by the target user; In response to the user query information queue being non-empty and the corresponding information token being invalid, the identity of the client corresponding to the target user is re-verified, wherein the information token is distributed and bound when the user query information queue is created, and is used to control the life cycle of the user query information queue; In response to the identity re-verification being passed, the token validity period of the information token is updated; In response to the user query information queue being non-empty and the corresponding information token being valid, the user query information at the head position of the queue is taken out from the user query information queue; The query scene corresponding to the user query information is determined; The candidate scene guide information corresponding to the query scene is recalled from the pre-constructed scene guide information library to obtain a candidate scene guide information set; The candidate scene guide information that meets the screening condition is selected from the candidate scene guide information set as the scene guide information.
3. The method of claim 2, wherein, The keyword retrieval is performed according to the search vector sequence and the pre-constructed content knowledge base to determine the target keyword sequence, which comprises: For each search vector in the search vector sequence, the keyword retrieval is performed in the content knowledge base through the way of jump similarity calculation according to the search vector, so as to retrieve the target keyword matching the search vector.
4. The method of claim 3, wherein, The target query information is constructed according to the user query information, the basic guide information and the scene guide information, which comprises: The guide task decomposition is respectively performed on the basic guide information and the scene guide information to obtain a guide task set; determine guiding task description information corresponding to each guiding task in the set of guiding tasks, wherein the guiding task description information comprises guiding task characteristics and a guiding task type; perform approximate guiding task fusion on the guiding tasks in the set of guiding tasks according to the guiding task description information corresponding to the guiding tasks, to obtain a set of fused guiding tasks; determine a guiding task priority corresponding to each fused guiding task in the set of fused guiding tasks, wherein the guiding task priority represents an execution order of the fused guiding task; perform task rearrangement on the fused guiding tasks in the set of fused guiding tasks according to the guiding task priority corresponding to the fused guiding tasks, to obtain a sequence of rearranged guiding tasks; perform guiding information fusion on the basic guiding information and the scene guiding information according to the sequence of rearranged guiding tasks, to obtain fused guiding information; combine the user query information and the fused guiding information according to a preset combination format, to obtain the target query information.
5. The method of claim 4, wherein, The generating, according to the target query information and a keyword extraction model, of an initial keyword information sequence comprises: inputting the target query information into the keyword extraction model to generate a main keyword sequence; for each main keyword in the main keyword sequence, performing the following processing steps: extracting main keyword characteristics corresponding to the main keyword; performing word granularity random masking on content in the user query information other than the main keyword, to obtain a masking position; performing projection masking on a position corresponding to the masking position in scene semantic characteristics corresponding to the query scene, to obtain masked scene semantic characteristics; performing expansion keyword matching according to the main keyword characteristics, to obtain a candidate expansion keyword group corresponding to the main keyword; extracting expansion keyword characteristics corresponding to each candidate expansion keyword in the candidate expansion keyword group; determining a feature similarity between the expansion keyword characteristics corresponding to each candidate expansion keyword in the candidate expansion keyword group and the masked scene semantic characteristics; selecting, from the candidate expansion keyword group, a candidate expansion keyword that satisfies an expansion keyword selection condition as an expansion keyword group corresponding to the main keyword; generating keyword description information corresponding to the main keyword according to the main keyword characteristics; generating, in the initial keyword information sequence, initial keyword information corresponding to the main keyword according to the main keyword, the expansion keyword group corresponding to the main keyword, and the keyword description information corresponding to the main keyword.
6. The method of claim 5, wherein, The determining, according to the initial keyword information sequence, of a content generation mode comprises: generating a keyword graph according to the initial keyword information sequence, wherein the keyword graph comprises a same number of keyword layers as a number of initial keyword information in the initial keyword information sequence, and each keyword layer comprises a number of keyword nodes equal to a number of expansion keywords in an expansion keyword group included in the corresponding initial keyword information plus 1; performing graph feature extraction on the keyword graph to generate keyword graph characteristics; and According to the keyword graph feature and the content generation mode classifier, the content generation mode is determined.
7. The method of claim 6, wherein, The generating the retrieval vector matched with each initial keyword information in the initial keyword information sequence comprises: A scene matching degree of a stem keyword included in the initial keyword information and corresponding to the query scene is determined as an initial voting probability corresponding to the stem keyword; A scene matching degree of each extended keyword in an extended keyword group included in the initial keyword information and corresponding to the query scene is determined as an initial voting probability corresponding to the extended keyword; A first target number of voting objects are created, wherein the first target number is an odd number; A second target number of candidate keywords are selected by voting from the stem keyword and the extended keyword group included in the initial keyword information according to the initial voting probability corresponding to the stem keyword and the initial voting probability corresponding to the extended keyword through the first target number of voting objects; A word vector corresponding to each candidate keyword in the second target number of candidate keywords is constructed to obtain a word vector sequence; The word vectors in the word vector sequence are spliced to obtain a spliced word vector; The spliced word vector is compressed to obtain a retrieval vector corresponding to the initial keyword information.
8. A keyword semantic-based content generation apparatus, characterized by comprising: Comprise: A construction unit configured to construct target query information according to user query information, basic guidance information and scene guidance information, wherein the user query information represents query content initiated by a target user, and the scene guidance information matches the user query information; A first generation unit configured to generate an initial keyword information sequence according to the target query information and a keyword extraction model, wherein the initial keyword information comprises a stem keyword, an extended keyword group and keyword description information, and the keyword extraction model is a model for keyword extraction based on keyword semantics; A determination unit configured to determine a content generation mode according to the initial keyword information sequence, wherein the content generation mode comprises a first generation mode and a second generation mode; A second generation unit configured to generate feedback content according to a content generation model and the initial keyword information sequence in response to the content generation mode being the first generation mode; A third generation unit configured to generate a retrieval vector matched with each initial keyword information in the initial keyword information sequence to obtain a retrieval vector sequence in response to the content generation mode being the second generation mode; A keyword retrieval unit configured to perform keyword retrieval according to the retrieval vector sequence and a pre-constructed content knowledge base to determine a target keyword sequence; A fourth generation unit configured to generate feedback content according to the target keyword sequence and the content generation model.
9. An electronic device, comprising: One or more processors; Storage having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 7. A computer program is stored thereon, wherein the computer program is executed by a processor to implement the method of any one of claims 1 to 7.
10. A computer readable medium characterized by
Citation Information
Patent Citations
Method and system for automatic retrieval and classification of HS code
CN112765308A
Global intelligent retrieval system applied to smart enterprise platform
CN118839022A
Retrieval enhancement generation-based retrieval method, product, equipment and medium
CN119003795A
Search optimization method and device, electronic equipment, storage medium and program product
CN119025619A
Configuration method and system of network configuration text and electronic equipment
CN120067291A