Search Result Output Method, Device, Computer Equipment and Readable Storage Medium

By analyzing the content to be searched as the address and content part input by the user, combining word-cut strings, content labels and vectors, matching search results are extracted, which solves the problem of inaccurate search results based on keyword output, and improves the accuracy and efficiency of search results.

CN114417113BActive Publication Date: 2025-06-13RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210073949.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-06-13
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

The search results based on keyword output are relatively general and may not meet the real search needs of users, resulting in inaccurate search results, resulting in wasted search resources and low search efficiency.

Method used

By parsing the content to be searched input by the user into an address part and a content part, the first material to be output is obtained according to the address part, the second material to be output is obtained according to the word-cut string, content label and content vector of the content part, and the target material to be output matching the first material to be output is extracted from the second material to be output, and it is output as a search result.

Benefits of technology

It improves the accuracy and search efficiency of search results, ensures that the output search results are more in line with the user's search needs, and avoids the waste of search resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114417113B_ABST
    Figure CN114417113B_ABST
Patent Text Reader

Abstract

The present application discloses a search result output method, apparatus, computer device, and readable storage medium, which relate to the field of Internet technologies. The method parses the content to be searched into an address part and a content part, obtains the first material to be output according to the address part, obtains the second material to be output according to the word segmentation string, content tags, and content vector of the content part, determines the search result and outputs it, and sets different recall strategies for different parts to improve the accuracy and efficiency of the search. The method includes: obtaining the content to be searched and parsing the content to be searched into an address part and a content part; querying the address identifier of the address part and using the material mounted on the address identifier as the first material to be output; identifying the content parameters of the content part and obtaining the second material to be output according to the content parameters; extracting the target material to be output that matches the first material to be output from the second material to be output, and outputting the first material to be output and the target material to be output as the search result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technologies, and in particular, to a search result output method, apparatus, computer device, and readable storage medium. Background Art

[0002] In recent years, with the rapid development of Internet technologies, various Internet applications have been widely and deeply involved in various fields, and big data has shown an explosive growth. Massive amounts of data and information are scattered in the network space. When users need to obtain information and data, they can search for information through a search platform, so that the search platform can output relevant search content. As an important link between users and information, the search platform generally outputs a large number of relevant search results for the content searched by users, and provides as much information as possible for users to refer to.

[0003] In related technologies, the search platform obtains the search content input by the user, extracts keywords from the search content or directly uses the search content as a keyword, searches for materials including the keyword in the material pool, and outputs the searched materials as the search results of this time for the user to view.

[0004] In the process of implementing the present application, the applicant found that the related technologies have at least the following problems:

[0005] The search results output based on keywords are relatively general, and the search results may not meet the real search needs of users, resulting in inaccurate search results output, wasting search resources, and low search efficiency. Summary of the Invention

[0006] In view of this, the present application provides a search result output method, apparatus, computer device, and readable storage medium, mainly aiming to solve the problem that the search results output based on keywords are relatively general at present, the search results may not meet the real search needs of users, resulting in inaccurate search results output, wasting search resources, and low search efficiency.

[0007] According to a first aspect of the present application, a search result output method is provided, and the method includes:

[0008] Obtain the content to be searched, and parse the content to be searched into an address part and a content part;

[0009] Query the address identifier of the address part, and use the material mounted on the address identifier as the first material to be output;

[0010] Identify the content parameters of the content part, and obtain the second material to be output according to the content parameters, where the content parameters include the word segmentation string, content label, and content vector of the content part;

[0011] Extract the target material to be output that matches the first material to be output from the second material to be output, and output the first material to be output and the target material to be output as search results.

[0012] Optionally, the obtaining the content to be searched and parsing the content to be searched into an address part and a content part includes:

[0013] Receive the content to be searched input by the user, and identify the address string included in the content to be searched;

[0014] When the address string is successfully identified, use the address string as the address part, and use the other strings in the content to be searched except the address string as the content part. Wherein, if the content to be searched does not include the other strings, obtain a full content identifier and use the full content identifier as the content part;

[0015] When the address string is not successfully identified, obtain the current geographical location of the user, use the current geographical location as the address part, and use the content to be searched as the content part.

[0016] Optionally, the identifying the content parameters of the content part includes:

[0017] Perform word segmentation processing on the content part to obtain at least one string that makes up the content part, and use the at least one string as the segmented word string;

[0018] Query the product category, service category, and applicable scenarios to which the content part belongs, and use the product category label corresponding to the product category, the service category label corresponding to the service category, and the scenario label corresponding to the applicable scenario as the content label;

[0019] Map the content part to a semantic vector space, and generate a vector expression of the content part as the content vector based on a similarity calculation model constructed in the semantic vector space. The similarity calculation model is used to convert character content into a vector.

[0020] Optionally, the obtaining the second material to be output according to the content parameters includes:

[0021] Query the material pool for materials whose material names include the segmented word string as the first material set;

[0022] Obtain the materials labeled with the content label in the material pool simultaneously or respectively as the second material set;

[0023] Calculate the vector similarity between all the materials in the material pool and the content vector simultaneously or separately, and extract the materials whose vector similarity meets the similarity threshold from all the materials as the third material set;

[0024] Use the first material set, the second material set, and the third material set as the second material to be output.

[0025] Optionally, calculating the vector similarity between all the materials in the material pool and the content vector, and extracting the materials whose vector similarity meets the similarity threshold from all the materials as the third material set includes:

[0026] For each material in all the materials, obtain the material information of the material, map the material information to the semantic vector space, and generate a vector expression of the material information as the material vector of the material based on the similarity calculation model constructed in the semantic vector space;

[0027] Calculate the vector cosine value between the material vector and the content vector in the semantic vector space, and use the vector cosine value as the vector similarity between the material vector and the content vector;

[0028] Calculate the vector similarity for each material respectively to obtain a plurality of vector similarities;

[0029] Obtain the similarity threshold, extract the target vector similarities greater than or equal to the similarity threshold from the plurality of vector similarities, and use the materials indicated by the target vector similarities as the third material set;

[0030] Wherein, if the content vector is generated based on the content part set as the full content identifier, then use all the materials as the third material set.

[0031] Optionally, extracting the target material to be output that matches the first material to be output from the second material to be output includes:

[0032] Query a plurality of candidate stores that provide the second material to be output, obtain the store name of each candidate store in the plurality of candidate stores, and obtain a plurality of store names;

[0033] Compare the plurality of store names with the first material to be output, and extract the target store names that appear in the first material to be output from the plurality of store names;

[0034] Extract the materials provided by the target store names from the second material to be output as the target material to be output.

[0035] Optionally, outputting the first material to be output and the target material to be output as search results includes:

[0036] Associating the first material to be output and the target material to be output according to the supply relationship between the first material to be output and the target material to be output;

[0037] Outputting the associated first material to be output and target material to be output as the search results to a result page for display.

[0038] Optionally, the method further includes:

[0039] Obtaining a spatial map, parsing the spatial map to obtain a plurality of spatial addresses, where the plurality of spatial addresses include buildings, roads, business districts, and offline activity areas in the spatial map;

[0040] For each spatial address in the plurality of spatial addresses, setting an address identifier for the spatial address, where the address identifier is one or both of the address name and address parameters of the spatial address;

[0041] Identifying a target point of interest corresponding to the spatial address on the spatial map, obtaining a target store associated with the target point of interest, and mounting the target store as a material on the address identifier;

[0042] Obtaining a target store for each spatial address respectively, and completing the mounting of materials on the address identifiers of the plurality of spatial addresses.

[0043] According to a second aspect of the present application, a search result output device is provided, and the device includes:

[0044] An acquisition module, configured to acquire content to be searched, and parse the content to be searched into an address part and a content part;

[0045] A query module, configured to query the address identifier of the address part, and use the material mounted on the address identifier as the first material to be output;

[0046] An identification module, configured to identify content parameters of the content part, and obtain a second material to be output according to the content parameters, where the content parameters include a word segmentation string, a content label, and a content vector of the content part;

[0047] An output module, configured to extract a target material to be output that matches the first material to be output from the second material to be output, and output the first material to be output and the target material to be output as search results.

[0048] Optionally, the obtaining module is configured to receive the content to be searched input by the user and identify the address string included in the content to be searched; when the address string is successfully identified, use the address string as the address part and use the other strings in the content to be searched except the address string as the content part. Wherein, if the content to be searched does not include the other strings, obtain a full-content identifier and use the full-content identifier as the content part; when the address string is not successfully identified, obtain the current geographical location of the user and use the current geographical location as the address part and use the content to be searched as the content part.

[0049] Optionally, the identifying module is configured to perform word segmentation processing on the content part to obtain at least one string that makes up the content part, and use the at least one string as the segmented word string; query the commodity category, service category, and applicable scenario to which the content part belongs, and use the commodity category label corresponding to the commodity category, the service category label corresponding to the service category, and the scenario label corresponding to the applicable scenario as the content label; map the content part to a semantic vector space, and generate a vector expression of the content part as the content vector based on a similarity calculation model constructed in the semantic vector space, where the similarity calculation model is used to convert character content into a vector.

[0050] Optionally, the identifying module is configured to query in the material pool for materials whose material names include the segmented word string as a first material set; simultaneously or separately obtain in the material pool materials labeled with the content label as a second material set; simultaneously or separately calculate the vector similarity between all materials in the material pool and the content vector, and extract from all the materials the materials whose vector similarity meets the similarity threshold as a third material set; use the first material set, the second material set, and the third material set as the second material to be output.

[0051] Optionally, the recognition module is configured to, for each material in the all materials, obtain the material information of the material, map the material information to a semantic vector space, generate a vector expression of the material information as the material vector of the material based on a similarity calculation model constructed in the semantic vector space; calculate the vector cosine value between the material vector and the content vector in the semantic vector space, and use the vector cosine value as the vector similarity between the material vector and the content vector; calculate the vector similarity for each of the materials respectively to obtain a plurality of vector similarities; obtain the similarity threshold, extract the target vector similarities greater than or equal to the similarity threshold from the plurality of vector similarities, and use the materials indicated by the target vector similarities as the third material set; wherein, if the content vector is generated based on the content part set as the full content identifier, then use the all materials as the third material set.

[0052] Optionally, the output module is configured to query a plurality of candidate stores providing the second material to be output, obtain the store names of each candidate store in the plurality of candidate stores to obtain a plurality of store names; compare the plurality of store names with the first material to be output, and extract the target store names that appear in the first material to be output from the plurality of store names; extract the materials provided by the target store names from the second material to be output as the target material to be output.

[0053] Optionally, the output module is configured to associate the first material to be output and the target material to be output according to the supply relationship between the first material to be output and the target material to be output; output the associated first material to be output and target material to be output to the result page for display as the search result.

[0054] Optionally, the device further includes:

[0055] The parsing module is configured to obtain a space map, parse the space map to obtain a plurality of space addresses, and the plurality of space addresses include buildings, roads, business districts, and offline activity areas in the space map;

[0056] The setting module is configured to, for each space address in the plurality of space addresses, set an address identifier for the space address, and the address identifier is one or two of the address name and address parameters of the space address;

[0057] The mounting module is configured to identify a target point of interest corresponding to the space address on the space map, obtain a target store associated with the target point of interest, and mount the target store as a material on the address identifier;

[0058] The setting module is further configured to obtain a target store for each of the spatial addresses, respectively, and complete the material mounting of the address identifiers of the multiple spatial addresses.

[0059] According to a third aspect of the present application, there is provided a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the first aspects above are implemented.

[0060] According to a fourth aspect of the present application, there is provided a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the first aspects above are implemented.

[0061] By means of the above technical solutions, a search result output method, device, computer device, and readable storage medium provided by the present application parse the content to be searched input by the user into an address part and a content part, obtain the first material to be output according to the address part, obtain the second material to be output according to the word segmentation string, content label, and content vector of the content part, extract the target material to be output that matches the first material to be output from the second material to be output, and use the first material to be output and the target material to be output as the search result for output. Different material recall strategies are set for different parts, fully identifying the user's search intention, making the output search result fit the user's search requirements, improving the accuracy and search efficiency of the output search result, and avoiding waste of search resources.

[0062] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0064] Figure 1 It shows a schematic flowchart of a search result output method provided by an embodiment of the present application;

[0065] Figure 2A It shows a schematic flowchart of a search result output method provided by an embodiment of the present application;

[0066] Figure 2B It shows a schematic diagram of a search result output method provided by an embodiment of the present application;

[0067] Figure 3A shows a schematic structural diagram of a search result output device provided by an embodiment of the present application;

[0068] Figure 3B shows a schematic structural diagram of a search result output device provided by an embodiment of the present application;

[0069] Figure 4 shows a schematic structural diagram of a device of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0070] Hereinafter, exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0071] An embodiment of the present application provides a search result output method, as Figure 1 shown, the method includes:

[0072] 101. Obtain the content to be searched, and parse the content to be searched into an address part and a content part.

[0073] 102. Query the address identifier of the address part, and use the material mounted on the address identifier as the first material to be output.

[0074] 103. Identify the content parameters of the content part, and obtain the second material to be output according to the content parameters. The content parameters include the word segmentation string, content label, and content vector of the content part.

[0075] 104. Extract the target material to be output that matches the first material to be output from the second material to be output, and output the first material to be output and the target material to be output as search results.

[0076] For the method provided by the embodiment of the present application, the content to be searched input by the user is parsed into an address part and a content part. The first material to be output is obtained according to the address part. The second material to be output is obtained according to the word segmentation string, content label, and content vector of the content part. The target material to be output that matches the first material to be output is extracted from the second material to be output, and the first material to be output and the target material to be output are output as search results. Different material recall strategies are set for different parts, the search intention of the user is fully identified, so that the output search results fit the search needs of the user, the accuracy and search efficiency of the output search results are improved, and the waste of search resources is avoided.

[0077] An embodiment of the present application provides a search result output method, as follows Figure 2A shown, the method includes:

[0078] 201. Obtain the content to be searched, and parse the content to be searched into an address part and a content part.

[0079] The search platform can provide the largest entry for users to search for information, and is an important link connecting users and information. As the most commonly used search tool, the search platform has become an essential part of people's lives. When the search platform provides search services for users, it will obtain the search content input by the user, extract keywords from the search content or directly use the search content as a keyword, search for materials including the keyword in the material pool, and output the searched materials as the search results of this time for the user to view. However, the applicant realizes that in fact, the search content input by the user is complex. The search content may include store brand words such as "KFC" and "McDonald's", may also include address words such as "Xixi Road" and "Zhongguancun", and there are also category content words such as "rice noodles" and "milk tea". The search results output based on keywords are relatively general, and the search results may not match the user's true search needs, resulting in inaccurate search results output, wasting search resources, and low search efficiency.

[0080] Therefore, the present application proposes a search result output method, which parses the content to be searched input by the user into an address part and a content part, obtains the first material to be output according to the address part, obtains the second material to be output according to the word segmentation string, content label and content vector of the content part, extracts the target material to be output that matches the first material to be output from the second material to be output, and outputs the first material to be output and the target material to be output as search results, sets different material recall strategies for different parts, fully identifies the user's search intention, makes the output search results match the user's search needs, improves the accuracy and search efficiency of the output search results, and avoids wasting search resources.

[0081] In the embodiment of the present application, in order to realize the recall of materials in different parts, after the platform obtains the content to be searched input by the user, it is necessary to parse the content to be searched into an address part and a content part. The specific process is as follows:

[0082] First, receive the content to be searched input by the user, perform intention recognition on the content to be searched, and identify the address string included in the content to be searched, such as address strings such as "Zhongguancun" and "Xixi Road". Considering that sometimes the user's search may not include content related to the address, such as only searching for "KFC", therefore, in practical applications, there will be two situations: successful or failed identification of the address string

[0083] When the address string is successfully recognized, the platform will use the address string as the address part and the other strings in the content to be searched except the address string as the content part. For example, assume the content to be searched by the user is "Xixi Road KFC", then the address part is "Xixi Road" and the content part is "KFC". It should be noted that in actual applications, there is also a situation where the user only searches for an address. For example, the content to be searched entered by the user only includes "Zhongguancun". In this case, no other strings can be obtained after the address string is extracted. Therefore, in the embodiments of the present application, if the content to be searched does not include other strings, a full-content identifier is obtained and used as the content part. For example, if the full-content identifier is set to "all", then the content part is set to "all".

[0084] When the address string recognition fails, it means that the content to be searched entered by the user does not carry content related to the address. By default, the vicinity of the user is searched. Therefore, the platform will obtain the current geographical location of the user, use the current geographical location as the address part, and use the content to be searched as the content part. For example, if the content to be searched entered by the user is "KFC" and does not carry location-related content, the generated address part is "vicinity" and the content part is "KFC".

[0085] In summary, referring to Table 1 below, the parsing process of the address part and the content part for the content to be searched is exemplified as follows:

[0086] Content to be searched Address part Content part Fried chicken Nearby (default) Fried chicken Zhongguancun Zhongguancun All (default) Xixi Road milk tea Xixi Road Milk tea

[0087] Table 1

[0088] If the content to be searched is "fried chicken", then there is no address string related to the location involved. Therefore, the parsed address part is "vicinity" and the content part is "fried chicken"; if the content to be searched is "Zhongguancun", then there are no other strings related to the specific content to be searched. Therefore, the parsed address part is "Zhongguancun" and the content part is "all"; if the content to be searched is "Xixi Road milk tea", then the address part is "Xixi Road" and the content part is "milk tea".

[0089] 202. Query the address identifier of the address part and use the materials mounted on the address identifier as the first materials to be output.

[0090] In the embodiments of the present application, after the address part and the content part are determined, different strategies for material recall operations are started based on the two parts respectively. Among them, since the address part needs to obtain the materials that the user may want to query according to the actual position corresponding to the address part on the spatial map, in fact, the platform will mount the materials to some actual spatial addresses on the spatial map in advance, so that when recalling materials based on the address part, the actual spatial address corresponding to the address part can be directly queried, and the materials mounted on the spatial address are used as the materials for reference when outputting the search results later. The process of the platform mounting the address and the materials in advance is described below:

[0091] First, the platform obtains the spatial map and parses the spatial map to obtain multiple spatial addresses. Among them, the multiple spatial addresses include buildings, roads, business districts, and offline activity areas in the spatial map. In fact, the process of parsing the spatial map is the process of counting address words in the spatial map according to four types: points, lines, planes, and solids. Specifically, a building can be regarded as a "point", such as "Zhongguancun Subway Station"; a road can be regarded as a "line", such as "Xixi Road"; a business district can be a commercial entity with a three-dimensional structure concept and can be regarded as a "solid", such as "Xidan Joy City"; an offline activity area can be a large-area area such as a square on a plane and can be regarded as a "plane", such as "Wulin Square".

[0092] After determining the multiple spatial addresses, for each spatial address among the multiple spatial addresses, the platform sets an address identifier for the spatial address, and the address identifier is one or two of the address name and address parameters of the spatial address. Then, the platform identifies the target point of interest corresponding to the spatial address on the spatial map, obtains the target store associated with the target point of interest, mounts the target store as the material on the address identifier, obtains the target store for each spatial address respectively, and completes the material mounting of the address identifiers of the multiple spatial addresses. That is, each spatial address is mapped to a different address ID (Identity document, unique identification code), the POI (Point of Interest) of the address ID is determined, and the store in the POI is obtained as the target store and mounted on the address ID.

[0093] In this way, through the above process, a certain number of target stores will be mounted on the address identifier corresponding to each spatial address in the space. When the platform obtains the address part according to the content to be searched input by the user, it queries the address identifier of the address part and uses the materials mounted on the address identifier as the first materials to be output.

[0094] 203. Identify the content parameters of the content part.

[0095] In the embodiments of the present application, after the recall of materials is completed based on the address part, the platform begins to recall materials based on the content part. Among them, in the embodiments of the present application, the recall of materials based on the content part follows three recall strategies, namely text recall, label recall, and semantic vector recall. Therefore, it is necessary to implement the recognition of the content part to obtain the word segmentation string, content labels, and content vectors of the content part. Subsequently, text recall is performed based on the word segmentation string, label recall is performed based on the content labels, and semantic recall is performed based on the content vectors. The following describes the three recall strategies and the generation processes of the word segmentation string, content labels, and content vectors respectively:

[0096] Text recall can be considered as the basic recall logic and is also the earliest recall method of search engines. Its process is to segment the materials used for retrieval. For example, "fried chicken burger" will be segmented into two words, "fried chicken" and "burger", and stored in the inverted index. When the content to be searched entered by the user arrives, it will also be segmented. For example, when the user searches for "Xixi KFC", it will be segmented into two words, "Xixi KFC". Through the matching after segmentation, the result of "KFC Xixi Road Store" can be recalled. Text recall can better recall materials that match literally and ensure the basic recall logic. Therefore, in the embodiments of the present application, the content part will be segmented to obtain at least one string that composes the content part, and the at least one string will be used as the word segmentation string for subsequent text recall based on the word segmentation string. For example, assuming the content part is "fried chicken burger", the obtained word segmentation strings are "fried chicken" and "burger".

[0097] Furthermore, the process of label recall is to convert the user's search term into corresponding labels for recall. The label system of the platform is divided into three categories, namely food and beverage labels, life service labels, and scene labels. Among them, food and beverage labels include food and beverage categories, dishes, etc., such as Western food, hot pot, spicy fragrant pot, etc. Life service labels include categories related to life services, service contents, etc., such as massage, hair styling, ear picking, etc. In addition, there are also scene-related labels, such as birthday party, family dinner, afternoon tea, etc. The labels will be mounted on each material through offline mining. For example, the store "Haidilao" as a material will be mounted with multiple labels such as "hot pot" and "family dinner". After mounting, the store can be recalled by retrieving relevant label words. During the label recall process, the literal limit can be broken through to a certain extent, and more materials can be recalled by tagging. Therefore, in the embodiments of the present application, the platform will query the product category, service category, and applicable scene to which the content part belongs, and use the product category label corresponding to the product category, the service category label corresponding to the service category, and the scene label corresponding to the applicable scene as the content labels for subsequent label recall based on the content labels.

[0098] Furthermore, the method of semantic vector recall is to map both the content part of the user's search and the material information of the material to a semantic vector space. In semantic vector recall, a similarity calculation model for converting character content into vectors is constructed in the semantic vector space. This similarity calculation model will express the content part of the user's search and the material information as vector expressions in multiple dimensions, such as [0.2, 0.15, 0.34]. The goal of the similarity calculation model is to make the cosine value between the content part and the material information close to it larger (i.e., the similarity is larger), and the cosine value between the content part and the material information not close to it smaller (i.e., the similarity is smaller). Therefore, during the semantic vector recall process, materials with a relatively high similarity to the content part can be recalled by calculating the cosine value. It should be noted that semantic vector recall is different from text recall and label recall, and constructs a vector-based semantic expression method. This recall method is more flexible and can completely break through the literal and text for material recall. Therefore, in the embodiments of this application, the platform maps the content part to the semantic vector space, and based on the similarity calculation model constructed in the semantic vector space, generates a vector expression of the content part as the content vector for subsequent semantic vector recall based on the content vector.

[0099] 204. Obtain the second material to be output according to the content parameter.

[0100] In the embodiments of this application, after determining the content parameter, the platform starts text recall, label recall, and semantic vector recall according to the content parameter to obtain the second material to be output. Among them, when obtaining the second material to be output, the platform will query the materials in the material pool whose material names include the segmented string as the first material set. For example, based on the segmented string "fried chicken", "Korean fried chicken", "honey fried chicken", "soy sauce fried chicken", etc. can be recalled as the first material set.

[0101] Furthermore, the platform obtains the materials labeled with content tags in the material pool simultaneously or separately as the second material set. For example, assuming the content tag is "fast food", the recalled materials can be "KFC", "fried chicken", "hamburger", etc.

[0102] Further, the platform calculates the vector similarity between all the materials in the material pool and the content vector simultaneously or separately, and extracts the materials in all the materials whose vector similarity meets the similarity threshold as the third material set. Specifically, for each material in all the materials, the platform obtains the material information of the material, maps the material information to the semantic vector space, and based on the similarity calculation model constructed in the semantic vector space, generates a vector expression of the material information as the material vector of the material. Subsequently, the platform calculates the vector cosine value between the material vector and the content vector in the semantic vector space, takes the vector cosine value as the vector similarity between the material vector and the content vector, and calculates the vector similarity for each material separately to obtain multiple vector similarities. Finally, the platform obtains the similarity threshold, extracts the target vector similarities greater than or equal to the similarity threshold from the multiple vector similarities, and takes the materials indicated by the target vector similarities as the third material set. For example, assuming that the similarity threshold can be set to 90%, the materials with a vector similarity above 90% with the content vector are taken as the third material set. It should be noted that as described in step 201 above, some content parts are set as full-content identifiers. Therefore, if the content vector is generated based on the content parts set as full-content identifiers, all the materials are taken as the third material set.

[0103] After obtaining the first material set, the second material set, and the third material set through the above process, the platform takes the first material set, the second material set, and the third material set as the second materials to be output.

[0104] 205. Extract the target materials to be output that match the first materials to be output from the second materials to be output.

[0105] In the embodiment of the present application, after determining the second materials to be output, the platform extracts the target materials to be output that match the first materials to be output from the second materials to be output, and takes the target materials to be output as the search content finally presented to the user. Among them, when determining the target materials to be output, the platform queries multiple candidate stores that provide the second materials to be output, obtains the store names of each candidate store among the multiple candidate stores, and gets multiple store names. Subsequently, the multiple store names are compared with the first materials to be output, the target store names that appear in the first materials to be output are extracted from the multiple store names, and the materials provided by the target store names are extracted from the second materials to be output as the target materials to be output.

[0106] In fact, the process of determining the target material to be output is also the process of merging the materials recalled based on the address part and the content part. The merging is performed in an "AND" manner; for the first material set, the second material set, and the third material set recalled based on the three recall strategies in the content part, they are merged in an "OR" manner, that is, "text recall" or "tag recall" or "semantic vector recall". Specifically, it can be expressed in the formal formula shown in Formula 1 below:

[0107] Formula 1: The first material to be output AND (the first material set OR the second material set OR the third material set)

[0108] For example, assume that the first material to be output includes private dining restaurants, dessert shops, Manbao Wonton, Yoshinoya, KFC, McDonald's, and Burger King, and the first material set, the second material set, and the third material set include chicken leg burgers, original flavor chicken, French fries, and spicy and sour noodles. Then, the target store names obtained from the first material set, the second material set, and the third material set in the first material to be output are KFC, McDonald's, and Burger King. The materials such as chicken leg burgers, original flavor chicken, French fries provided by KFC, McDonald's, and Burger King and the target store names are used as the target materials to be output at the same time.

[0109] 206. Output the first material to be output and the target material to be output as search results.

[0110] In the embodiment of this application, after determining the first material to be output and the target material to be output, the platform outputs the first material to be output and the target material to be output as search results. Among them, since the first material to be output is a store in space and the target material to be output includes commodities, services, etc., when outputting the search results, the store and the commodities need to be associated and output so that the user can know which commodities are provided by which store. Therefore, the platform will associate the first material to be output and the target material to be output according to the supply relationship between them, and output the associated first material to be output and target material to be output as search results to the result page for display.

[0111] To sum up, see Figure 2B, in the embodiments of the present application, the content to be searched by the user is divided into two major categories. One is "where", that is, the address part, which refers to a location and can be the user's current location, the target location searched by the user, etc. It is divided into four types: point, line, surface, and body. For example, "Xixi Road", "Zhongguancun", etc.; the other is "what", that is, the content part, which refers to the content and can be a brand, category, such as "KFC", "massage", "western food", etc. Different recall methods are applied to the "where" and "what" parts in the embodiments of the present application. At the same time, three different recall strategies are applied to the "what" part, and a hierarchical recall system is formed based on the three different recall strategies, integrating three dimensions: text recall, label recall, and semantic vector recall, to solve the recall problem in search.

[0112] The method provided by the embodiments of the present application parses the content to be searched input by the user into an address part and a content part, obtains the first material to be output according to the address part, obtains the second material to be output according to the word segmentation string, content label, and content vector of the content part, extracts the target material to be output that matches the first material to be output from the second material to be output, and outputs the first material to be output and the target material to be output as search results. Different material recall strategies are set for different parts, fully identifying the user's search intention, making the output search results fit the user's search needs, improving the accuracy and search efficiency of the output search results, and avoiding waste of search resources.

[0113] Further, as Figure 1 a specific implementation of the method, the embodiments of the present application provide a search result output device, as Figure 3A shown. The device includes: an acquisition module 301, a query module 302, an identification module 303, and an output module 304.

[0114] The acquisition module 301 is used to acquire the content to be searched and parse the content to be searched into an address part and a content part;

[0115] The query module 302 is used to query the address identifier of the address part and use the material mounted on the address identifier as the first material to be output;

[0116] The identification module 303 is used to identify the content parameters of the content part and obtain the second material to be output according to the content parameters. The content parameters include the word segmentation string, content label, and content vector of the content part;

[0117] The output module 304 is used to extract the target material to be output that matches the first material to be output from the second material to be output and output the first material to be output and the target material to be output as search results.

[0118] In a specific application scenario, the obtaining module 301 is configured to receive the content to be searched input by a user, and identify an address string included in the content to be searched; when the address string is successfully identified, use the address string as the address part, and use other strings in the content to be searched except the address string as the content part, where if the content to be searched does not include the other strings, obtain a full-content identifier and use the full-content identifier as the content part; when the address string is not successfully identified, obtain the current geographical location of the user, use the current geographical location as the address part, and use the content to be searched as the content part.

[0119] In a specific application scenario, the identification module 303 is configured to perform word segmentation processing on the content part to obtain at least one string constituting the content part, and use the at least one string as the segmented word string; query the commodity category, service category, and applicable scenario to which the content part belongs, and use the commodity category label corresponding to the commodity category, the service category label corresponding to the service category, and the scenario label corresponding to the applicable scenario as the content label; map the content part to a semantic vector space, and generate a vector expression of the content part as the content vector based on a similarity calculation model constructed in the semantic vector space, where the similarity calculation model is used to convert character content into a vector.

[0120] In a specific application scenario, the identification module 303 is configured to query in a material pool for materials whose material names include the segmented word string as a first material set; simultaneously or separately obtain in the material pool materials labeled with the content label as a second material set; simultaneously or separately calculate the vector similarity between all materials in the material pool and the content vector, and extract from all the materials materials whose vector similarity meets a similarity threshold as a third material set; use the first material set, the second material set, and the third material set as the second material to be output.

[0121] In a specific application scenario, the recognition module 303 is configured to, for each material in all the materials, obtain the material information of the material, map the material information to a semantic vector space, and based on a similarity calculation model constructed in the semantic vector space, generate a vector expression of the material information as the material vector of the material; calculate the vector cosine value between the material vector and the content vector in the semantic vector space, and use the vector cosine value as the vector similarity between the material vector and the content vector; calculate the vector similarity for each of the materials respectively to obtain a plurality of vector similarities; obtain the similarity threshold, extract the target vector similarities greater than or equal to the similarity threshold from the plurality of vector similarities, and use the materials indicated by the target vector similarities as the third material set; wherein, if the content vector is generated based on the content part set as the full-content identifier, then all the materials are used as the third material set.

[0122] In a specific application scenario, the output module 304 is configured to query a plurality of candidate stores providing the second material to be output, obtain the store names of each candidate store in the plurality of candidate stores to obtain a plurality of store names; compare the plurality of store names with the first material to be output, and extract the target store names that appear in the first material to be output from the plurality of store names; extract the materials provided by the target store names from the second material to be output as the target material to be output.

[0123] In a specific application scenario, the output module 304 is configured to associate the first material to be output and the target material to be output according to the supply relationship between the first material to be output and the target material to be output; output the associated first material to be output and target material to be output to a result page for display as the search result.

[0124] In a specific application scenario, as Figure 3B shown, the device further includes: a parsing module 305, a setting module 306, and a mounting module 307.

[0125] The parsing module 305 is configured to obtain a space map, parse the space map to obtain a plurality of space addresses, and the plurality of space addresses include buildings, roads, business districts, and offline activity areas in the space map;

[0126] The setting module 306 is configured to, for each space address in the plurality of space addresses, set an address identifier for the space address, and the address identifier is one or two of the address name and address parameters of the space address;

[0127] The mounting module 307 is used to identify the target point of interest corresponding to the spatial address on the spatial map, obtain the target store associated with the target point of interest, and mount the target store as materials on the address identifier.

[0128] The setting module 306 is further used to obtain a target store for each spatial address respectively, and complete the material mounting of the address identifiers of the multiple spatial addresses.

[0129] The device provided by the embodiment of the present application parses the content to be searched input by the user into an address part and a content part, obtains the first material to be output according to the address part, obtains the second material to be output according to the word segmentation string, content label, and content vector of the content part, extracts the target material to be output that matches the first material to be output from the second material to be output, and outputs the first material to be output and the target material to be output as search results. Different material recall strategies are set for different parts, fully identifying the search intention of the user, making the output search results fit the search needs of the user, improving the accuracy and search efficiency of the output search results, and avoiding waste of search resources.

[0130] It should be noted that for other corresponding descriptions of each functional unit involved in the search result output device provided by the embodiment of the present application, reference can be made to Figure 1 and Figures 2A to 2B the corresponding descriptions therein, which will not be elaborated here.

[0131] In an exemplary embodiment, referring to Figure 4 , a computer device is further provided. The computer device includes a bus, a processor, a memory, and a communication interface, and may further include an input / output interface and a display device. Among them, each functional unit can complete mutual communication through the bus. The memory stores a computer program, and the processor is used to execute the program stored on the memory to execute the search result output method in the above embodiment.

[0132] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the search result output method are implemented.

[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented through hardware, or can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0134] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of a preferred embodiment scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present application.

[0135] Those skilled in the art can understand that the modules in the devices in the embodiment scenario can be distributed in the devices of the embodiment scenario according to the description of the embodiment scenario, or can be correspondingly changed and located in one or more devices different from the present embodiment scenario. The modules of the above embodiment scenario can be combined into one module, or can be further split into multiple sub-modules.

[0136] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the embodiment scenario.

[0137] The above disclosure only shows several specific embodiment scenarios of the present application. However, the present application is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present application.

Claims

1. A search result output method, characterized in that, it includes: Obtain the content to be searched, and parse the content to be searched into an address part and a content part; Query the address identifier of the address part, and use the material mounted on the address identifier as the first material to be output; Identify the content parameters of the content part, and according to the content parameters, obtain the second material to be output. The content parameters include the word segmentation string, content label and content vector of the content part. The second material to be output includes a first material set, a second material set and a third material set. The material names of the materials in the first material set include the word segmentation string. The materials in the second material set are marked with the content label. The vector similarity between the materials in the third material set and the content vector meets the similarity threshold; Extract the target material to be output that matches the first material to be output from the second material to be output, and output the first material to be output and the target material to be output as search results. The target material is provided by the store indicated by the target store name in the second material to be output, and the target store name is the store name that appears in the first material to be output.

2. The method according to claim 1, characterized in that, the obtaining the content to be searched and parsing the content to be searched into an address part and a content part includes: Receive the content to be searched input by the user, and identify the address string included in the content to be searched; When the address string is successfully identified, use the address string as the address part, and use the other strings in the content to be searched except the address string as the content part. Wherein, if the content to be searched does not include the other strings, obtain the full content identifier and use the full content identifier as the content part; When the address string fails to be identified, obtain the current geographical location of the user, use the current geographical location as the address part, and use the content to be searched as the content part.

3. The method according to claim 1, characterized in that, the identifying the content parameters of the content part includes: Perform word segmentation processing on the content part to obtain at least one string that makes up the content part, and use the at least one string as the word segmentation string; Query the commodity category, service category and applicable scenario to which the content part belongs, and use the commodity category label corresponding to the commodity category, the service category label corresponding to the service category and the scenario label corresponding to the applicable scenario as the content label; Map the content part to the semantic vector space, and based on the similarity calculation model constructed in the semantic vector space, generate the vector expression of the content part as the content vector. The similarity calculation model is used to convert character content into a vector.

4. The method according to claim 1, characterized in that, the obtaining the second material to be output according to the content parameters includes: Query the materials whose material names include the word segmentation string in the material pool as the first material set; Obtain the materials labeled with the content tags in the material pool simultaneously or separately as the second material set; Calculate the vector similarity between all the materials in the material pool and the content vector simultaneously or separately, and extract the materials with vector similarity meeting the similarity threshold from all the materials as the third material set; Use the first material set, the second material set, and the third material set as the second material to be output.

5. The method according to claim 4, wherein, the calculating the vector similarity between all the materials in the material pool and the content vector, and extracting the materials with vector similarity meeting the similarity threshold from all the materials as the third material set includes: For each material in all the materials, obtain the material information of the material, map the material information to the semantic vector space, and generate a vector expression of the material information as the material vector of the material based on the similarity calculation model constructed in the semantic vector space; Calculate the vector cosine value between the material vector and the content vector in the semantic vector space, and use the vector cosine value as the vector similarity between the material vector and the content vector; Calculate the vector similarity for each of the materials respectively to obtain a plurality of vector similarities; Obtain the similarity threshold, extract the target vector similarities greater than or equal to the similarity threshold from the plurality of vector similarities, and use the materials indicated by the target vector similarities as the third material set; wherein, if the content vector is generated based on the content part set as the full content identifier, then use all the materials as the third material set.

6. The method according to claim 1, wherein, the extracting the target material to be output that matches the first material to be output from the second material to be output includes: Query a plurality of candidate stores providing the second material to be output, obtain the store names of each candidate store in the plurality of candidate stores to get a plurality of store names; Compare the plurality of store names with the first material to be output, and extract the target store names that appear in the first material to be output from the plurality of store names; Extract the materials provided by the target store names from the second material to be output as the target material to be output.

7. The method according to claim 1, wherein, the outputting the first material to be output and the target material to be output as search results includes: Associate the first material to be output and the target material to be output according to the supply relationship between the first material to be output and the target material to be output; Output the associated first material to be output and target material to be output as the search results to a result page for display.

8. The method according to claim 1, wherein, the method further includes: Obtain a spatial map, parse the spatial map to obtain a plurality of spatial addresses, and the plurality of spatial addresses include buildings, roads, business districts, and offline activity areas in the spatial map; For each of the multiple spatial addresses, set an address identifier for the spatial address, where the address identifier is one or both of the address name and address parameters of the spatial address; Identify the target point of interest corresponding to the spatial address on the spatial map, obtain the target store associated with the target point of interest, and mount the target store as a material on the address identifier; Obtain a target store for each of the spatial addresses respectively, and complete the material mounting of the address identifiers of the multiple spatial addresses.

9. A search result output device, characterized in that, it includes: An acquisition module, configured to acquire the content to be searched, and parse the content to be searched into an address part and a content part; A query module, configured to query the address identifier of the address part, and use the material mounted on the address identifier as the first material to be output; An identification module, configured to identify the content parameters of the content part, and obtain a second material to be output according to the content parameters. The content parameters include the segmented string, content label, and content vector of the content part. The second material to be output includes a first material set, a second material set, and a third material set. The material names of the materials in the first material set include the segmented string, the materials in the second material set are labeled with the content label, and the vector similarity between the materials in the third material set and the content vector meets a similarity threshold; An output module, configured to extract the target material to be output that matches the first material to be output from the second material to be output, and output the first material to be output and the target material to be output as search results. The target material is provided by the store indicated by the target store name in the second material to be output, and the target store name is the store name that appears in the first material to be output.

10. A computer device, including a memory and a processor, where the memory stores a computer program, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Data searching method and system based on semantic analysis

    CN103353894A

  • Search based service information management system and method

    CN105117418A