Information searching method and device, equipment and medium
By matching the query information input by the user with the attribute label of the information point database, and combining the correlation model and the information summary model, the problem of low accuracy in information search is solved, accurate recall and sorting of information points is achieved, and the accuracy of information search is improved.
Patent Information
- Application Number
- CN202510397945.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, it is difficult to accurately establish a correlation between the unstructured query information input by the user and the matching of structured information points in the database, resulting in poor accuracy of information search.
By obtaining the query information input by the user, searching for information points, determining multiple first information points, and matching the query information with the attribute labels in the information point database, determining the target attribute labels, and accurately sorting and displaying using the correlation model and information summary model.
It improves the accuracy of the search results of information points, ensures that the recalled information points meet user demands, and improves the overall effect of information search.
Smart Images

Figure CN120336366A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to an information search method, apparatus, device, and medium. Background Art
[0002] With the development of large models, users can search for Points of Information (POIs) through large models. In related technologies, information points that match the query terms are recalled from a database based on the query statement input by the user and summarized and returned through a large model. Since the query statement is unstructured and the data in the database is structured, it is difficult to accurately establish an association during recall, and the rough ranking of information points during recall may have problems that do not meet the true demands of users, resulting in poor search accuracy and the need for improvement. Summary of the Invention
[0003] To solve the above technical problems, the present disclosure provides an information search method, apparatus, device, and medium.
[0004] An embodiment of the present disclosure provides an information search method, the method comprising:
[0005] Obtaining query information input by a user;
[0006] Performing an information point search based on the query information to determine a plurality of first information points;
[0007] Matching the query information with a plurality of attribute tags in an information point database to determine at least one target attribute tag corresponding to the query information;
[0008] Performing a relevance match between the at least one target attribute tag and the plurality of first information points to determine a plurality of second information points;
[0009] Using an information summarization model to determine a plurality of information point summarization results of the plurality of second information points, and displaying the plurality of information point summarization results.
[0010] An embodiment of the present disclosure further provides an information search apparatus, the apparatus comprising:
[0011] An obtaining module, configured to obtain query information input by a user;
[0012] A search module, configured to perform an information point search based on the query information to determine a plurality of first information points;
[0013] A tag module, configured to match the query information with a plurality of attribute tags in an information point database to determine at least one target attribute tag corresponding to the query information;
[0014] A matching module, configured to perform a relevance match between the at least one target attribute tag and the multiple first information points to determine multiple second information points;
[0015] A result module, configured to use an information summarization model to determine summarization results of the multiple second information points and display the summarization results of the multiple information points.
[0016] An embodiment of the present disclosure further provides an electronic device, which includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the information search method provided by the embodiment of the present disclosure.
[0017] An embodiment of the present disclosure further provides a computer-readable storage medium, which stores a computer program for executing the information search method provided by the embodiment of the present disclosure.
[0018] The technical solution provided by the embodiment of the present disclosure has the following advantages compared with the prior art: The information search solution provided by the embodiment of the present disclosure obtains query information input by a user; performs an information point search based on the query information to determine multiple first information points; matches the query information with multiple attribute tags in an information point database to determine at least one target attribute tag corresponding to the query information; performs a relevance match between the at least one target attribute tag and the multiple first information points to determine multiple second information points; uses an information summarization model to determine summarization results of the multiple second information points and display the summarization results of the multiple information points. By adopting the above technical solution, for the query information input by the user, multiple first information points that may be hit can be first searched, and the attribute tags corresponding to the query information can be determined. The attribute tags are used to perform a relevance match with the multiple first information points to determine the final multiple second information points. Furthermore, an information summarization model is used to determine the information point summarization results of the multiple second information points for display. For the unstructured query information input by the user, it can be converted into structured attribute tags defined in the database, which helps to better understand the query information and quickly establish an association with information points during recall. And based on the information points recalled for the first time, the attribute tags are used as the basis for secondary recall and accurate sorting, which can effectively filter out information points that do not meet the user's requirements, thereby improving the accuracy of the information point search results. Description of the Drawings
[0019] In combination with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original parts and elements are not necessarily drawn to scale.
[0020] Figure 1 Flow schematic diagram of an information search method provided by some embodiments of the present disclosure;
[0021] Figure 2 Schematic diagram of the process of determining a second information point provided by some embodiments of the present disclosure;
[0022] Figure 3 Flow schematic diagram of another information search method provided by some embodiments of the present disclosure;
[0023] Figure 4 Schematic diagram of the process of determining the label information of an information point provided by some embodiments of the present disclosure;
[0024] Figure 5 Schematic diagram of the process of constructing an information point database provided by some embodiments of the present disclosure;
[0025] Figure 6 Schematic diagram of the information search process provided by some embodiments of the present disclosure;
[0026] Figure 7 Schematic diagram of the structure of an information search device provided by some embodiments of the present disclosure;
[0027] Figure 8 Schematic diagram of the structure of an electronic device provided by some embodiments of the present disclosure. Detailed implementation manners
[0028] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the accompanying drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0029] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0030] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0031] It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules or units.
[0032] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0033] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes, and are not used to limit the scope of these messages or information.
[0034] To solve the problem of low accuracy in information search in related technologies, the embodiments of this disclosure provide an information search method, which will be introduced below in combination with specific embodiments.
[0035] Figure 1 It is a schematic flowchart of the information search method provided by some embodiments of this disclosure. This method can be executed by an information search device, which can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, this method includes:
[0036] Step 101, obtain the query information input by the user.
[0037] The information search method of the embodiments of this disclosure can be applied to an information search system. The information search system can search for multiple information points that meet the user's needs based on a large model. Here, the large model can be a large language model. The large language model is a processing model based on natural language. Through the large language model, the rules and structures of the language can be automatically learned, the meaning of the language can be understood, and coherent text with correct grammar and semantics can be generated according to the understood meaning.
[0038] Among them, the query information (query) can be the basic information for which the user needs to search for information points. When the user enters keywords, phrases or sentences on the input interface to search for information points, this input content is the query information. For example, the query information can be "Please query restaurants with good taste in City A". An information point can be any meaningful point on the map that is not of geographical significance, such as a restaurant, a store, a gas station, a station, etc.
[0039] Specifically, when the user needs to search for information points, the user can enter the query information on the input interface, and the information search device can obtain this query information.
[0040] Step 102, perform an information point search based on the query information to determine multiple first information points.
[0041] Among them, the first information points can be the information points that may match the query information screened from a large number of information points, which can be the information points obtained through rough screening. The number of the first information points can be relatively large. For example, the first information points can be 1,000.
[0042] Specifically, the information search device can search among multiple information points based on the query information. The specific search method is not limited. For example, it can be searched according to relevance to obtain multiple first information points.
[0043] In some embodiments, searching for information points based on the query information and determining multiple first information points may include: inputting the query information into an information rewriting model to output rewritten information; and searching for information points based on the rewritten information to obtain multiple first information points.
[0044] Among them, the rewritten information can be the information obtained by rewriting the query information. The information rewriting model can be a model for specific information rewriting processing, or it can be rewritten in other ways. The rewritten information can eliminate ambiguity, correct spelling mistakes, and more accurately reflect the semantic intention of the user's query. The information search device can input the query information into the information rewriting model for rewriting to obtain the rewritten information, and search among multiple information points based on the rewritten information to obtain multiple first information points. For example, when the user inputs several incomplete words as the query information, a complete sentence can be obtained as the rewritten information through rewriting. By rewriting the query information input by the user, the query intention can be made more clear, which helps to obtain more accurate search results.
[0045] Step 103: Match the query information with multiple attribute labels in the information point database to determine at least one target attribute label corresponding to the query information.
[0046] Among them, the information point database can be a database constructed after structuring the basic information such as comment data, introduction data, and other data of information points of local or third parties, and using a large model to confirm the label information. The information point database can be a vector database, which includes multiple label information and multiple attribute labels corresponding to multiple information points. The attribute label can be identification information used to describe the subjective and objective characteristics of the information point in multiple dimensions. In the embodiments of the present disclosure, multiple attribute labels can be preset to determine the corresponding attribute label for each information point. The formats of the multiple attribute labels are unified and stored in a unified set (schema), and the set is stored in a unified data set or data table. The target attribute label can be an attribute label that matches the query information input by the user, and the number of the target attribute labels is at least one.
[0047] Specifically, the information search device can obtain multiple attribute tags in the information point database, and perform similarity matching between the query information and each attribute tag, and determine at least one attribute tag that satisfies the similarity condition as at least one target attribute tag corresponding to the query information. Optionally, each attribute tag is provided with at least one synonym interpretation to participate in the matching when matching with the query information, effectively increasing the matching hit probability and improving the matching accuracy.
[0048] In some embodiments, matching query information with multiple attribute tags in an information point database to determine at least one target attribute tag corresponding to the query information may include: using a vector matching algorithm to determine multiple similarities between the query information and the multiple attribute tags in the information point database; extracting at least one attribute tag whose similarity is greater than or equal to a similarity threshold as the at least one target attribute tag corresponding to the query information based on the multiple similarities.
[0049] The vector matching algorithm can be a process of finding another vector that is most similar to a vector by calculating the similarity or distance between a vector and other vectors. This algorithm can convert the data to be matched into a vector representation, and then use the distance or similarity measurement method in the vector space to perform matching. The specific algorithm used is determined according to the actual situation. For example, the vector matching algorithm can be a cosine distance algorithm, a Euclidean distance algorithm, etc. The similarity threshold can be a pre-set minimum similarity value, which can be set according to the actual situation. For example, the similarity threshold can be 0.6.
[0050] Specifically, when determining at least one target attribute tag corresponding to the query information, the information search device can use a vector matching algorithm to calculate the similarity between each attribute tag, obtain multiple similarities, and compare each similarity with a similarity threshold. If a similarity is greater than or equal to the similarity threshold, the attribute tag corresponding to the similarity is the target attribute tag, and at least one target attribute tag is obtained. The vector matching technology can be used to convert the unstructured query information input by the user into a structured attribute tag defined in the database, which helps to better understand the query information and quickly establish associations with information points when recalling, recalling the information points that the user really demands, and improving the accuracy of information search.
[0051] Step 104: perform correlation matching on at least one target attribute tag and multiple first information points to determine multiple second information points.
[0052] Among them, the second information point can be the information point of the final output determined for the query information input by the user. The second information point can be the information point obtained by performing refined screening and refined sorting among the above-mentioned multiple first information points. The number of first information points is greater than the number of second information points. Or the second information point can also be the information point obtained by performing mixed screening and mixed sorting among the multiple third information points obtained by performing video search locally based on the query information and the multiple fourth information points obtained by performing text search on a third party after performing refined screening and refined sorting on the above-mentioned multiple first information points.
[0053] Specifically, after determining at least one target attribute tag, the information search device can perform a relevance match between the at least one target attribute tag and each first information point to determine the match result. The specific matching method is not limited. For example, it can be implemented by using statistical methods, rules or models, and extract multiple first information points with the top-ranked match results as the multiple second information points at this time.
[0054] In some embodiments, performing a relevance match between at least one target attribute tag and multiple first information points to determine multiple second information points may include: using a relevance model to determine the relevance score between at least one target attribute tag and each first information point; extracting a preset number of first information points with the top-ranked relevance scores among the multiple first information points as the multiple second information points.
[0055] The relevance model can be a model for analyzing whether there is a relevance between two data. The embodiments of the present disclosure are used to perform a relevance analysis on each first information point and at least one target attribute tag. Here, the relevance refers to the degree of closeness or tightness of the relationship, and outputs a relevance score. The relevance score can be a score for quantifying the relevance between two data. The higher the relevance score, the greater the relevance. The preset number can be the number determined in advance for the second information point and can be determined according to the actual situation. For example, the preset number can be 8.
[0056] Specifically, when the information search device determines multiple second information points, it can input at least one target attribute tag and each first information point into the relevance model, and output the relevance score between the first information point and the at least one target attribute tag. Specifically, the average value can be calculated after determining the relevance score between the first information point and each target attribute tag to obtain the final relevance score; then the multiple first information points can be sorted according to the relevance score, and a preset number of first information points with the top-ranked sorting results are extracted as the multiple second information points. For example, when the preset number is 8, the first 8 first information points with the highest relevance scores are extracted as the second information points.
[0057] In the above scheme, the correlation between at least one attribute label corresponding to the query information and the information point is analyzed through the correlation model, which more accurately reflects the real relationship between the information point and the label, ensuring higher efficiency.
[0058] For example, Figure 2 A schematic diagram of a second information point determination process provided in some embodiments of the present disclosure, such as Figure 2 As shown in the figure, the process of determining the second information point may include: recalling multiple first information points; rewriting the query information to obtain rewritten information; refining the logic, performing label matching based on the rewritten information to determine the corresponding target attribute label, and during the matching process, the interface service provided by the information database may be used to obtain multiple pre-defined attribute labels, and then the most matching preset number of second information points may be recalled based on the target attribute label, and the preset number of second information points may be output.
[0059] Optionally, at least one target attribute tag can be filtered in the process of determining the second information point. The specific filtering is implemented according to the sorting result of the correlation score. If the correlation scores of the preset number of first information points ranked at the top in the sorting result and a certain target attribute tag are all lower than the filtering threshold, that is, they are all irrelevant to a certain target attribute tag, then the target attribute tag can be deleted. The filtering threshold can be the minimum score for judging whether the information point is related to the attribute tag. If it is lower than the filtering threshold, it means that it is irrelevant and the attribute tag needs to be filtered. By filtering the tags of the query information, the accuracy of tag matching can be effectively guaranteed, avoiding the adverse effects on the information point search caused by matching errors.
[0060] Step 105: Determine multiple information point summary results of multiple second information points using the information summary model, and display the multiple information point summary results.
[0061] The information summary model may be a model that extracts important information from the relevant information of each information point searched based on the query information and logically integrates it to generate a summary text in the form of natural language. The information summary model may be a large language model. The information point summary result may be a summary text obtained for an information point using the information summary model, which is used to be displayed to the user on the information search interface.
[0062] Specifically, after the information search device determines multiple second information points, it can obtain the relevant information of each second information point, input the relevant information of the multiple second information points into the information summary model, and output multiple information point summary results. Then, the following processing can be performed on the multiple information point summary results: information splicing, calling the model to screen again, parsing reference classes, link processing, assembling display cards, and streaming return, etc., so as to output the multiple information point summary results to the large language model for adjustment and then display them to the user. The user can view the multiple information point summary results of the multiple second information points corresponding to the input query information.
[0063] In some embodiments, using the information summary model to determine the multiple information point summary results of multiple second information points may include: obtaining the supplementary information of each second information point; generating a summary prompt word based on the basic information, supplementary information of the multiple second information points, and the prompt word template of the information summary model; inputting the summary prompt word into the information summary model to output multiple information point summary results. Optionally, the supplementary information includes introduction information and comment information. Obtaining the supplementary information of each second information point includes: obtaining the label summary reason of the corresponding attribute label from the information point database based on the information point identifier of each second information point; generating introduction information and comment information based on the label summary reason of each second information point.
[0064] The supplementary information can be information such as images, introduction information, and evaluation information outside the basic information of the information point, used to supplement and explain the information point. The basic information of the information point can include information such as identification, name, category, location, and score for basic description. The introduction information can include the summary information and characteristic label information of the information point. The summary information can be obtained by summarizing the information point using a large model. The characteristic label information can be the characteristic labels hit by the information point and the label summary reason. The characteristic labels can be several labels with a relatively high degree of importance or specialness determined from multiple attribute labels, and the specific quantity can be set according to the actual situation. The label summary reason can be the reason why an information point has a certain attribute label, which can be determined according to multiple comment data corresponding to the attribute label. For example, if an information point is a store and has an attribute label of "tasty", the corresponding label summary reason can be obtained by summarizing reasons from multiple comments such as "The taste of this store is amazing". The comment information can include the label summary reasons of the preset number of attribute labels hit by the information point. When the number of attribute labels is less than the preset number, it can be supplemented using comment data. When supplementing, several with the highest reading volume can be extracted for supplementation, only as an example.
[0065] The prompt template can be a template for configuring the content and position included in the prompt, and is used to generate a prompt. The prompt template in the embodiments of the present disclosure is used to generate a summary prompt, and the prompt template includes multiple placeholders for filling different information. A prompt can be a text or statement fragment used to trigger and guide a model to execute a specific task to generate specific output content. The summary prompt can be used to guide an information summarization model to generate an information point summary result for a second information point.
[0066] Specifically, when the information search device determines the information point summary results of multiple second information points, for each second information point, corresponding supplementary information can be obtained. Specifically, the tag summary reason can be extracted from the information point database according to the information point identifier of the second information point, and the introduction information and comment information of the second information point can be generated according to the tag summary reason; then, the basic information and supplementary information of multiple second information points can be filled into the prompt template of the information summarization model to obtain a summary prompt, and the summary prompt is input into the information summarization model for summarization to obtain multiple information point summary results.
[0067] Exemplarily, for each second information point in the prompt template, its data structure can be set as basic information, introduction information, image, recommended content, on-site experience content, and comment information in the supplementary information. The on-site experience content can include key experience content and other experience content, etc. The comment information can include the tag summary reason corresponding to the attribute tag and other comment data, etc. The arrangement order of each data in the data structure is only an example.
[0068] In the above solution, by generating a summary prompt for the information summarization model through the prompt template, prompts with consistent formats can be generated efficiently, which is beneficial to improving the summarization efficiency of subsequent information point summary results; and when summarizing information points, introduction information and comment information generated based on the tag summary reason can be added to each information point, and relevant content of the attribute tag can be added to the summary result, effectively improving the diversity and integrity of the information obtained by users.
[0069] The information search solution provided by the embodiments of the present disclosure obtains the query information input by the user; searches for information points based on the query information to determine multiple first information points; matches the query information with multiple attribute tags in the information point database to determine at least one target attribute tag corresponding to the query information; performs a relevance match between the at least one target attribute tag and the multiple first information points to determine multiple second information points; uses an information summary model to determine multiple information point summary results of the multiple second information points, and displays the multiple information point summary results. By adopting the above technical solution, for the query information input by the user, multiple first information points that may be hit can be first searched, and the attribute tags corresponding to the query information can be determined. The attribute tags are used to perform a relevance match with the multiple first information points to determine the final multiple second information points. Furthermore, the information summary model is used to determine the information point summary results of the multiple second information points for display. The unstructured query information input by the user can be converted into structured attribute tags defined in the database, which helps to better understand the query information and quickly establish an association with the information points during recall. Moreover, based on the attribute tags as the basis for secondary recall and precise ranking on the basis of the first recall of information points, information points that do not meet the user's requirements can be effectively screened out, thereby improving the accuracy of the information point search results.
[0070] Exemplarily, Figure 3 FIG. 15 is a schematic flowchart of another information search method provided by some embodiments of the present disclosure. In a feasible implementation manner, before the above step 101, the information search method may further include constructing an information point database, which may be specifically as follows:
[0071] Step 301: Obtain the basic information of multiple information points and perform structured processing to obtain structured data of the multiple information points.
[0072] The basic information of the information points may include information such as identification, name, category, location, and score for basic description. The structured processing may be a process of converting unordered and unstructured data into ordered and structured formats. The embodiments of the present disclosure may convert the basic data of each information point into structured data in a unified format.
[0073] Specifically, the information search device may obtain the basic information of multiple information points from local and third parties, and perform structured processing on the basic information of each information point to obtain the corresponding structured data.
[0074] Step 302: Based on the comment data and a preset multiple attribute tags in the structured data of each information point, determine the tag information corresponding to each information point, where the tag information includes an attribute tag and a tag summary reason.
[0075] Among them, the label information may include at least one attribute label hit by an information point and the reason for summarizing the label of each attribute label. The reasons for summarizing the labels corresponding to the same attribute label for different information points may be different.
[0076] In some embodiments, based on the comment data and a plurality of preset attribute labels in the structured data of each information point, determining the label information corresponding to each information point includes: sequentially determining each information point as a to-be-processed information point; for each attribute label, extracting a plurality of comments corresponding to the attribute label from the comment data of the to-be-processed information point, and in response to determining that the relevance score of the plurality of comments and the attribute label is greater than a relevance threshold, determining the attribute label as the attribute label corresponding to the to-be-processed information point, and determining the reason for summarizing the label of the attribute label based on the plurality of comments; combining the attribute label corresponding to the to-be-processed information point and the reason for summarizing the label of the attribute label to obtain the corresponding label information.
[0077] The to-be-processed information point may be the information point currently being processed. Each information point is determined as a to-be-processed information point to determine the corresponding label information. The relevance threshold may be the lowest threshold for determining whether an information point hits an attribute label based on the relevance analysis of the comment of the information point and the attribute label. If it is greater than the relevance threshold, it means hitting; otherwise, it means not hitting.
[0078] Specifically, the information search device may extract comment data from the structured data for each information point. The comment data may include all comments of the information point, and extract a plurality of comments corresponding to the attribute label from the comment data for each attribute label, perform a relevance analysis on the plurality of comments and the attribute label to determine the corresponding relevance score. The specific determination method may adopt the above-mentioned relevance model to determine whether the relevance score is greater than the relevance threshold. If it is greater, it is determined that the information point hits the attribute label, and the attribute label is determined as the attribute label corresponding to the to-be-processed information point. For the attribute label, a corresponding reason for summarizing the label is determined based on the plurality of comments using a summary model. The summary model may be a model for determining a summary text based on a plurality of data. The embodiments of the present disclosure are used to determine the reason for summarizing the label based on a plurality of comments. If the relevance score is not greater than the relevance threshold, it is determined that the information point does not hit the attribute label, and the attribute label is not the attribute label corresponding to the to-be-processed information point; for all attribute labels of each information point, the above-mentioned determination of whether to hit is performed once. If it hits, the reason for summarizing the label is determined. Finally, at least one attribute label hit by each information point and the reason for summarizing the label of each attribute label can be determined, and combined to obtain the label information. By judging whether the attribute label of the information point hits or is valid, the accuracy of determining the attribute label of the information point can be improved, and further the accuracy of the reason for summarizing the label can be ensured.
[0079] Exemplarily, Figure 4Schematic diagram of the label information determination process for information points provided in some embodiments of the present disclosure, as Figure 4 shown. The figure shows the label information determination process for an attribute label of an information point, which may specifically include: for an attribute label of an information point, multiple corresponding comments can be recalled; low-quality comments can be filtered. Low-quality comments can be comments with a score less than the score threshold and / or a word count less than the word count threshold, which can improve the accuracy of subsequent analysis of the relevance between comments and the attribute label; determine the relevance scores of multiple ratings with the attribute label; if the relevance score is greater than the relevance threshold, determine the label summary reason corresponding to the attribute label; store the data in the database. The attribute labels are cycled in sequence until all multiple attribute labels in the information point database are judged, and the label information corresponding to the information point is stored in the information point database. The above process is performed once for all information points.
[0080] Step 303, construct an information point database based on the multiple label information corresponding to multiple information points and multiple attribute labels.
[0081] The information search device stores the multiple attribute labels and the multiple label information corresponding to multiple information points in the database, and encapsulates the data interface to construct an information point database.
[0082] Exemplarily, Figure 5 Schematic diagram of the information point database construction process provided in some embodiments of the present disclosure, as Figure 5 shown. The figure shows the information point database construction process from two dimensions: the data layer and the application layer. It may specifically include: at the data layer, the basic information of multiple local and third-party information points can be obtained, and the basic information of each information point can be structured to obtain the corresponding structured data. Then, data offline processing can be performed. Data offline processing may include determining the label information corresponding to each information point according to the comment data in the structured data of each information point and a preset multiple attribute labels; the multiple attribute labels come from a unified format set stored in a unified data table, and the upper layer of this data table can encapsulate data interfaces, such as a query interface and a relevance recall interface. The query interface is used to input the information point identifier of the information point to obtain the corresponding attribute label and label summary reason, and the relevance recall interface is used to input the information point identifier and attribute label of the information point to recall the corresponding multiple comments; the data interface is used to be called by the application layer and used in three stages: information summary, information point search ranking, and information point query in the above information search process.
[0083] In the above solution, by constructing an information point database, information points with a large amount of data can be analyzed in advance using a large model, and the attribute tags hit by each information point and the reasons for tag summary can be determined based on preset attribute tags. Subsequently, when processing the user's information search request, such as a life service request, the label information of the information points can be quickly obtained using engineering capabilities, avoiding the time-consuming process of analyzing information points from all angles in real time and the dependence on large model resources, saving a lot of time and model costs, providing support for rapid business iteration while improving the user experience. Moreover, the high-quality data in this information point database can be used in other business scenarios, enhancing scalability and applicability.
[0084] Next, a specific example is used to further illustrate the information search method of the embodiments of the present disclosure. Exemplarily, Figure 6 is a schematic diagram of the information search process provided by some embodiments of the present disclosure. As Figure 6 shown, the figure shows a complete information search process, which can specifically include the search preprocessing stage, the search link stage, the data service stage, and the search postprocessing stage in the figure. In the search preprocessing stage, the user inputs query information to an intelligent agent in the application, performs intent recognition, uses the local intelligent agent to rewrite the query information to obtain rewritten information, and inputs the rewritten information into the search link stage. In the search link stage, intent understanding can be performed based on the rewritten information, and three search links in the figure can be carried out, including vertical search for information points to recall multiple first information points, internal video vertical search to obtain multiple third information points, and external text search to obtain multiple fourth information points. For the multiple first information points, they can be combined with the data service stage to recall attribute tags for fine ranking to obtain multiple fifth information points. Then, the multiple fifth information points, multiple third information points, and multiple fourth information points can be mixed and sorted to obtain multiple second information points, such as 8 second information points. The mixing and sorting method is implemented according to text matching. Then, the multiple second information points can be output to the search postprocessing stage. In the search postprocessing and removal stage, data supplementation can be performed on the multiple second information points first to obtain information point basic data, images, introduction information, and comment information. The introduction information and comment information can be obtained by calling the query interface from the information point database through an online database. Then, information summary is performed to obtain multiple information point summary results, which are output to the large language model and then displayed to the user after adjustment.
[0085] In the data service stage, perform secondary recall and secondary sorting on multiple first information points. First, perform label matching based on the rewritten information and multiple attribute labels in the unified format set in the information point database to determine at least one corresponding target attribute label. Then, the at least one target attribute label can be correlated with the multiple first information points, and a preset number of first information points with the top-ranked correlation scores are extracted as multiple second information points. The multiple attribute labels can come from the unified format set after the labels are stored in the information point database, and the information point data can be synchronized with the online database.
[0086] This solution can perform label matching on the query information input by the user, convert the unstructured query information into structured input parameters, achieve a more accurate understanding of the query information, and the attribute labels can be used as the basis for accurate sorting of information points, improving the accuracy of accurate sorting, and further enhancing the accuracy of information search and the information search experience effect of users.
[0087] Figure 7 It is a schematic structural diagram of an information search device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 7 shown, the device includes:
[0088] An acquisition module 701, configured to acquire query information input by a user;
[0089] A search module 702, configured to perform information point search based on the query information to determine multiple first information points;
[0090] A label module 703, configured to match the query information with multiple attribute labels in the information point database to determine at least one target attribute label corresponding to the query information;
[0091] A matching module 704, configured to perform correlation matching between the at least one target attribute label and the multiple first information points to determine multiple second information points;
[0092] A result module 705, configured to use an information summarization model to determine multiple information point summarization results of the multiple second information points and display the multiple information point summarization results.
[0093] Optionally, the search module 702 is used to:
[0094] Input the query information into an information rewriting model, and output the rewritten information;
[0095] Perform information point search based on the rewritten information to obtain the multiple first information points.
[0096] Optionally, the label module 703 is used to:
[0097] Use a vector matching algorithm to determine multiple similarities between the query information and multiple attribute tags in the information point database;
[0098] Extract at least one attribute tag with a similarity greater than or equal to the similarity threshold from the multiple similarities as at least one target attribute tag corresponding to the query information.
[0099] Optionally, the matching module 704 is used for;
[0100] Use a correlation model to determine the correlation scores between the at least one target attribute tag and each of the first information points;
[0101] Extract a preset number of the first information points with the highest correlation scores from the multiple first information points as the multiple second information points, where the number of the first information points is greater than the number of the second information points.
[0102] Optionally, the result module 705 includes:
[0103] A supplement unit for obtaining supplementary information for each of the second information points;
[0104] A prompt word unit for generating summary prompt words based on the basic information, supplementary information of the multiple second information points, and the prompt word template of the information summary model;
[0105] An output unit for inputting the summary prompt words into the information summary model and outputting multiple information point summary results.
[0106] Optionally, the supplementary information includes introduction information and comment information, and the supplement unit is used to include:
[0107] Obtain the corresponding label summary reasons from the information point database based on the information point identifiers of each of the second information points;
[0108] Generate introduction information and comment information based on the label summary reasons of each of the second information points.
[0109] Optionally, the device further includes a database module for:
[0110] Obtain the basic information of multiple information points and perform structured processing to obtain structured data of multiple information points;
[0111] Based on the comment data in the structured data of each information point and multiple preset attribute tags, determine the label information corresponding to each information point, where the label information includes an attribute tag and a label summary reason;
[0112] Construct the information point database based on the multiple tag information corresponding to the multiple information points and the multiple attribute tags.
[0113] Optionally, the database module is specifically configured to:
[0114] Successively determine each of the information points as an information point to be processed;
[0115] For each of the attribute tags, extract multiple comments corresponding to the attribute tag from the comment data of the information point to be processed, and in response to determining that the relevance score of the multiple comments to the attribute tag is greater than the relevance threshold, determine the attribute tag as the attribute tag corresponding to the information point to be processed, and determine the tag summary reason of the attribute tag based on the multiple comments;
[0116] Combine the attribute tag corresponding to the information point to be processed and the tag summary reason of the attribute tag to obtain the corresponding tag information.
[0117] The information search device provided by the embodiments of the present disclosure can execute the information search method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0118] The embodiments of the present disclosure also provide a computer program product, including computer programs / instructions, which implement the information search method provided by any embodiment of the present disclosure when executed by a processor.
[0119] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure.
[0120] Specifically refer to Figure 8 , which shows a schematic structural diagram of an electronic device 800 suitable for implementing the embodiments of the present disclosure. The electronic device 800 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.
[0121] As Figure 8As shown, the electronic device 800 may include a processing device 801 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0122] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0123] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the information search method of the embodiment of the present disclosure are executed.
[0124] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an electrically erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0125] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as the HyperText Transfer Protocol (HTTP), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0126] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately without being assembled into the electronic device.
[0127] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain query information input by a user; perform an information point search based on the query information to determine a plurality of first information points; match the query information with a plurality of attribute tags in an information point database to determine at least one target attribute tag corresponding to the query information; perform a relevance match between the at least one target attribute tag and the plurality of first information points to determine a plurality of second information points; use an information summarization model to determine a plurality of information point summarization results of the plurality of second information points, and display the plurality of information point summarization results.
[0128] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network or a wide area network, or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0130] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0131] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field-Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0132] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, flash memories, optical fibers, portable compact disk read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0133] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to users and the authorization of users should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0134] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.
[0135] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0136] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An information search method, characterized in that, It includes: Obtain the query information input by the user; Based on the query information, conduct an information point search to determine multiple first information points; Match the query information with multiple attribute tags in the information point database to determine at least one target attribute tag corresponding to the query information; Perform a correlation match between the at least one target attribute tag and the multiple first information points to determine multiple second information points; Use an information summarization model to determine multiple information point summarization results of the multiple second information points and display the multiple information point summarization results.
2. The method according to claim 1, wherein Based on the query information, conduct an information point search to determine multiple first information points, including: Input the query information into an information rewriting model and output the rewritten information; Based on the rewritten information, conduct an information point search to obtain the multiple first information points.
3. The method according to claim 1, wherein Match the query information with multiple attribute tags in the information point database to determine at least one target attribute tag corresponding to the query information, including: Use a vector matching algorithm to determine multiple similarities between the query information and multiple attribute tags in the information point database; Extract at least one attribute tag with a similarity greater than or equal to the similarity threshold as at least one target attribute tag corresponding to the query information according to the multiple similarities.
4. The method according to claim 1, wherein Perform a correlation match between the at least one target attribute tag and the multiple first information points to determine multiple second information points, including; Use a correlation model to determine the correlation scores between the at least one target attribute tag and each of the first information points; Extract a preset number of first information points with the top-ranked correlation scores among the multiple first information points as the multiple second information points, where the number of the first information points is greater than the number of the second information points.
5. The method according to claim 1, wherein Use an information summarization model to determine multiple information point summarization results of the multiple second information points, including: Obtain the supplementary information of each of the second information points; Based on the basic information, supplementary information of the multiple second information points, and the prompt word template of the information summarization model, generate a summary prompt word; Input the summary prompt word into the information summarization model and output multiple information point summarization results.
6. The method according to claim 5, wherein The supplementary information includes introduction information and comment information. Obtaining the supplementary information of each of the second information points includes: Based on the information point identifier of each of the second information points, obtain the corresponding tag summary reason from the information point database; Generate introduction information and comment information based on the tag summary reason of each of the second information points.
7. The method according to claim 1, wherein The method further includes: Obtain the basic information of multiple information points and perform structured processing to obtain structured data of the multiple information points; Based on the comment data and multiple preset attribute tags in the structured data of each information point, determine the tag information corresponding to each information point, where the tag information includes an attribute tag and a tag summary reason; Construct the information point database based on the multiple tag information corresponding to the multiple information points and the multiple attribute tags.
8. The method according to claim 7, wherein Based on the comment data and multiple preset attribute tags in the structured data of each information point, determine the tag information corresponding to each information point, including: Determine each of the above information points as an information point to be processed in sequence; For each of the above attribute tags, extract multiple comments corresponding to the attribute tag from the comment data of the information point to be processed, and in response to determining that the relevance score of the multiple comments to the attribute tag is greater than the relevance threshold, determine the attribute tag as the attribute tag corresponding to the information point to be processed, and determine the tag summary reason of the attribute tag based on the multiple comments; Combine the attribute tag corresponding to the information point to be processed with the tag summary reason of the attribute tag to obtain the corresponding tag information.
9. An information search device, characterized in that It includes: An acquisition module for acquiring query information input by a user; A search module for performing information point search based on the query information to determine multiple first information points; A tag module for matching the query information with multiple attribute tags in an information point database to determine at least one target attribute tag corresponding to the query information; A matching module for performing relevance matching between the at least one target attribute tag and the multiple first information points to determine multiple second information points; A result module for using an information summary model to determine multiple information point summary results of the multiple second information points and display the multiple information point summary results.
10. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the information search method according to any one of claims 1-8 above.
11. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the information search method according to any one of claims 1-8 above.