Method, device and equipment for determining search result quality and storage medium

By acquiring search pairs from video search systems, extracting content information and generating vector representations, and calculating matching and novelty, the problem of diversity assessment in video search systems is solved, improving computational efficiency and accuracy.

CN114579410BActive Publication Date: 2025-12-09BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011378620.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-30
Publication Date
2025-12-09
Estimated Expiration
2040-11-30

AI Technical Summary

Technical Problem

Existing technologies face significant challenges in evaluating the diversity of video search systems, and their computational efficiency and accuracy need improvement.

Method used

By obtaining the search string and search results from the search pair, content information is extracted, word embedding is used to generate vector representations, and the relationship between matching and novelty is calculated to determine the quality information of the search results.

Benefits of technology

It effectively represents the diversity of search results, reduces the difficulty of diversity characterization, and improves computational efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114579410B_ABST
    Figure CN114579410B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device and equipment for determining the quality of search results, and a storage medium. The method comprises: for each search pair, extracting the content of the search string and the corresponding search result, respectively, obtaining the search content information corresponding to the search string and the resource content information corresponding to the search result; based on the search content information and the resource content information, determining the first relationship data between the search string and the search resource, and the second relationship data between the first search resource and each second search resource ranked in front; and based on the first relationship data and the second relationship data, determining the quality information of the search result. The present disclosure can effectively solve the problems of computational efficiency and computational accuracy in describing the diversity of search results.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to a search result quality determination method and device, equipment and a storage medium. BACKGROUND

[0002] An information retrieval system refers to a programmed system for searching and collecting, processing, storing and retrieving information according to specific information needs of a user. The main purpose of the information retrieval system is to provide a series of information services for the user, so that the user can find the content of interest. Information retrieval evaluation is an activity of evaluating the performance of the information retrieval system (mainly the ability to meet the information needs of the user). Through evaluation, the advantages and disadvantages of different technologies and the influence of different factors on the system can be evaluated, so as to continuously improve the research level in the field of information retrieval.

[0003] In the related art, the evaluation of the information retrieval system mainly focuses on improving the retrieval accuracy and importance, and the needs of different users may be different, so the diversity of the search results is also an important indicator for measuring the advantages and disadvantages of the search results. At present, there are few reports on the description of the diversity of the search results, especially for the video search system of the online application. Since online annotation is not possible, the description of the diversity of the search results brings certain difficulty, and the calculation efficiency and calculation accuracy of the description of the diversity of the search results of the system also need to be improved. SUMMARY

[0004] The present disclosure provides a search result quality determination method, device, equipment and storage medium to at least solve the problems of great difficulty in diversity evaluation for the video search system of the online application and at least one of the calculation efficiency and calculation accuracy of the diversity description in the related art. The technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a search result quality determination method is provided, comprising:

[0006] Obtaining at least one search pair, each search pair comprising a search string and a search result corresponding to the search string, the search result comprising a plurality of search resources sorted by search;

[0007] For each search pair, content extraction is performed on the search string and the corresponding search result in the search pair respectively, and search content information corresponding to the search string and resource content information corresponding to the search result are obtained respectively;

[0008] determine first relationship data between the search string and each search resource and second relationship data between each first search resource and each second search resource in the search result based on the search content information and the resource content information; the first relationship data is used to represent a matching degree between the search string and each search resource, and the second relationship data is used to represent a novelty degree of content of each first search resource in the search result;

[0009] determine quality information of the search result based on the first relationship data and the second relationship data; the quality information at least represents a diversity degree of the search result.

[0010] As an optional implementation, the step of extracting content of the search string and the corresponding search result in each search pair respectively to obtain search content information corresponding to the search string and resource content information corresponding to the search result respectively includes:

[0011] obtain a search sub-label corresponding to the search string in the search pair; the search sub-label is used to represent user search intention information covered by the search string;

[0012] extract content of the search sub-label corresponding to the search pair and each search resource in the search result respectively to obtain search content information corresponding to the search string and resource content information corresponding to the search result respectively.

[0013] As an optional implementation, the step of extracting content of the search sub-label corresponding to the search pair and each search resource in the search result respectively to obtain search content information corresponding to the search string and resource content information corresponding to the search result respectively includes:

[0014] perform word segmentation processing on the search sub-label to obtain search content information corresponding to the search string;

[0015] extract label dimension information and sentence dimension information of each search resource in the search result corresponding to the search string respectively;

[0016] merge the label dimension information and the sentence dimension information to obtain resource content information corresponding to the search result.

[0017] As an optional implementation, the step of determining first relationship data between the search string and each search resource and second relationship data between each first search resource and each second search resource in the search result based on the search content information and the resource content information includes:

[0018] perform word embedding processing on the search content information and the resource content information to obtain search vector representation of the search string and resource vector representation of each search resource respectively;

[0019] determine first relationship data between the search string and each search resource based on the resource vector representation of each search resource and the search vector representation of the search string;

[0020] determine relationship data between the Nth search resource and the top N-1 search resources based on the resource vector representation of the Nth search resource and the resource vector representation of the top N-1 search resources in the search result, where N is a positive integer;

[0021] determine the relationship data between the Nth search resource and the top N-1 search resources as second relationship data between each first search resource and each second search resource.

[0022] As an optional implementation, the step of determining the first relationship data between the search string and each search resource based on the resource vector representation of each search resource and the search vector representation of the search string comprises:

[0023] calculate a first similarity between each word vector in the resource vector representation of each search resource and each word vector in the search vector representation to determine a minimum similarity corresponding to each search resource;

[0024] determine the first relationship data between the search string and each search resource based on the size of the determined minimum similarity and a first preset similarity threshold.

[0025] As an optional implementation, the step of determining the relationship data between the Nth search resource and the top N-1 search resources based on the resource vector representation of the Nth search resource and the resource vector representation of the top N-1 search resources in the search result comprises:

[0026] determine a set of arrangement serial numbers of the top N-1 search resources from the search result based on arrangement orders of the search resources in the search result;

[0027] extract each word vector in the resource vector representation of each search resource belonging to the set of arrangement serial numbers to obtain a target word vector set;

[0028] calculate a similarity between each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource to determine a second similarity corresponding to each word vector in the target word vector set;

[0029] determine a first number of maximum similarities greater than or equal to a second preset similarity threshold;

[0030] Based on the determined first quantity, relationship data of the Nth search resource and the top N-1 search resources is determined.

[0031] As an optional implementation, the step of determining the quality information of the search result based on the first relationship data and the second relationship data comprises:

[0032] Based on the first relationship data, the second relationship data, and the ranking order of each search resource in the search result, sub-quality information of each search resource is determined.

[0033] The sub-quality information of each search resource in the search result is summed up to determine the quality information of the search result.

[0034] According to a second aspect of the embodiments of the present disclosure, a device for determining the quality of a search result is provided, comprising:

[0035] An acquisition module is configured to acquire at least one search pair, each search pair comprising a search string and a search result corresponding to the search string, the search result comprising a plurality of search resources sorted by search;

[0036] A content extraction module is configured to perform content extraction on the search string and the corresponding search result in each search pair respectively, and acquire search content information corresponding to the search string and resource content information corresponding to the search result respectively.

[0037] A relationship determination module is configured to determine first relationship data between the search string and each search resource, and second relationship data between each first search resource and each second search resource ranked in the front based on the search content information and the resource content information; the first relationship data is used to represent the matching degree between the search string and each search resource, and the second relationship data is used to represent the novelty degree of the content of each first search resource in the search result.

[0038] A quality determination module is configured to determine the quality information of the search result based on the first relationship data and the second relationship data, the quality information at least representing the diversity degree of the search result.

[0039] As an optional implementation, the content extraction module comprises:

[0040] An acquisition sub-module is configured to acquire a search sub-label corresponding to the search string in the search pair; the search sub-label is used to represent the user search intention information covered by the search string.

[0041] The content extraction submodule is configured to perform, for each search pair, content extraction on the search sub-label corresponding to the search pair and each search resource in the search result, respectively, to obtain search content information corresponding to the search string and resource content information corresponding to the search result.

[0042] As an optional implementation, the content extraction submodule is configured to perform:

[0043] segmenting the search sub-label to obtain the search content information corresponding to the search string;

[0044] extracting label dimension information and sentence dimension information of each search resource in the search result corresponding to the search string, respectively;

[0045] merging the label dimension information and the sentence dimension information to obtain the resource content information corresponding to the search result.

[0046] As an optional implementation, the relationship determination module includes:

[0047] a first processing submodule configured to perform word embedding processing on the search content information and the resource content information to obtain search vector representation of the search string and resource vector representation of each search resource, respectively;

[0048] a first relationship determination submodule configured to determine first relationship data between the search string and each search resource based on the resource vector representation of each search resource and the search vector representation of the search string;

[0049] a second relationship determination submodule configured to determine relationship data between an Nth search resource and the top N-1 search resources based on the resource vector representation of the Nth search resource and resource vector representations of the top N-1 search resources in the search result, where N is a positive integer;

[0050] a second processing submodule configured to use the relationship data between the Nth search resource and the top N-1 search resources as second relationship data between each first search resource and each second search resource in the top.

[0051] As an optional implementation, the first relationship determination submodule is configured to perform:

[0052] calculate a first similarity between each word vector in the resource vector representation of each search resource and each word vector in the search vector representation to determine minimum similarity corresponding to each search resource;

[0053] determine the first relationship data of the Nth search resource and the top N-1 search resources based on the determined first quantity.

[0054] As an optional implementation, the second relationship determining submodule is configured to perform:

[0055] determine a set of arrangement serial numbers of the top N-1 search resources from the search results based on the arrangement order of the search resources in the search results;

[0056] extract each word vector in the resource vector representation of each search resource belonging to the set of arrangement serial numbers to obtain a target word vector set;

[0057] calculate the similarity between each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource to determine the second similarity corresponding to each word vector in the target word vector set;

[0058] determine the first quantity of the maximum similarity greater than or equal to the second preset similarity threshold;

[0059] determine the relationship data of the Nth search resource and the top N-1 search resources based on the determined first quantity.

[0060] As an optional implementation, the quality determining module includes:

[0061] a first determining submodule configured to perform determining the sub-quality information of each search resource based on the first relationship data, the second relationship data, and the arrangement serial number of each search resource in the search results;

[0062] a second determining submodule configured to perform summing the sub-quality information of each search resource in the search results to determine the quality information of the search results.

[0063] According to a third aspect of the embodiments of the present disclosure, a storage medium is provided, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method for determining the quality of the search results as described in any of the above embodiments.

[0064] According to a fourth aspect of the embodiments of the present disclosure, an electronic device is provided, including:

[0065] a processor;

[0066] a memory for storing instructions executable by the processor;

[0067] wherein the processor is configured to execute the instructions to implement the method for determining the quality of the search results as described in any of the above embodiments.

[0068] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, which includes computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the method for determining the quality of search results provided in any of the embodiments.

[0069] The technical solutions provided by the embodiments of the present disclosure at least have the following beneficial effects:

[0070] The embodiments of the present disclosure acquire at least one search pair, each of the search pairs including a search string and a search result corresponding to the search string, and the search result including a plurality of search resources sorted by search; for each search pair, content extraction is performed on the search string and the corresponding search result in the search pair respectively, and search content information corresponding to the search string and resource content information corresponding to the search result are acquired respectively; based on the search content information and the resource content information, first relationship data between the search string and each search resource and second relationship data between each first search resource and each second search resource sorted in front are determined; the first relationship data is used to represent the matching degree between the search string and each search resource, and the second relationship data is used to represent the content novelty degree of each first search resource in the search result; and based on the first relationship data and the second relationship data, quality information of the search result is determined. Therefore, the quality information of the search result can effectively represent the diversity degree of the search result, which not only reduces the difficulty of depicting the diversity of the search result, but also improves the calculation efficiency and calculation accuracy of depicting the diversity of the search result.

[0071] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0072] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate an embodiment consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure, and do not constitute an improper limitation to the present disclosure.

[0073] Figure 1 is a schematic diagram of an implementation environment of a method for determining the quality of search results according to an exemplary embodiment.

[0074] Figure 2 is a flowchart of a method for determining the quality of search results according to an exemplary embodiment.

[0075] Figure 3is a flow chart of step S203 in a search result quality determination method according to an exemplary embodiment.

[0076] Figure 4 is a flow chart of step S205 in a search result quality determination method according to an exemplary embodiment.

[0077] Figure 5 is a partial flow chart of a search result quality determination method according to an exemplary embodiment.

[0078] Figure 6 is a partial flow chart of a search result quality determination method according to an exemplary embodiment.

[0079] Figure 7 is a block diagram of a search result quality determination apparatus according to an exemplary embodiment.

[0080] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0081] In order to make the ordinary person skilled in the art better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings.

[0082] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0083] Figure 1 is a schematic diagram of an implementation environment of a search result quality determination method according to an exemplary embodiment, see Figure 1 The implementation environment can include a first end 01 and a second end 02.

[0084] The first end 01 can be used to provide a background information retrieval service and generate matching search results based on search terms. The second end 02 can be used to provide a search result quality evaluation service for the first server 01 to determine the information retrieval performance of the first server 01.

[0085] The first end 01 and the second end 02 can be independent physical servers, can be a server cluster or a distributed system composed of multiple physical servers, can be a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform. Of course, the first end 01 and / or the second end 02 can also be a terminal, which can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart wearable device, a digital assistant, an augmented reality device, a virtual reality device, or an application running in an entity device, but is not limited thereto. The first end 01 and the second end 02 can be directly or indirectly connected through wired or wireless communication, and the present disclosure is not limited in this regard.

[0086] In an application scenario, the number of the first end 01 can be multiple, corresponding to different information retrieval algorithms, information retrieval models or information retrieval systems, etc. The quality of the search results of the multiple first ends 01 can be determined through the second end 02 to evaluate the retrieval performance of each information retrieval algorithm, information retrieval model or information retrieval system.

[0087] In another application scenario, the first end 01 can include an information retrieval model to be trained, and the model training effect or whether the training end condition is reached can be determined based on the search result quality determined by the second server, or the information retrieval model can be trained based on the target loss function constructed based on the determination method of the search result quality.

[0088] The search result quality determination method provided by the embodiments of the present disclosure can be executed by a search result quality determination device, which can be integrated in an electronic device such as a terminal or a server in the form of hardware or software, or can be executed by a server or a terminal alone, or can be executed by a terminal and a server cooperatively.

[0089] It should be noted that the application scenarios of the embodiments of the present disclosure include but are not limited to the above-mentioned application scenarios, and can also be applicable to other scenarios requiring search.

[0090] Figure 2 is a flowchart of a search result quality determination method according to an exemplary embodiment, as shown in Figure 2 The search result quality determination method is applied to an electronic device, and the second end in the above-mentioned implementation environment schematic diagram is taken as an example to illustrate the electronic device, which includes the following steps.

[0091] ​​​​​​In step S201, at least one search pair is acquired, each of the search pairs including a search string and a search result corresponding to the search string, the search result including a plurality of search resources sorted by search.

[0092] In an embodiment, the first end can acquire a search string to be searched, then retrieve a plurality of search resources meeting the conditions according to the search string, and sort the plurality of search resources by search to obtain a search result corresponding to the search string.

[0093] The search string is content to be searched, and can be used to reflect the search demand of a user. The search string includes at least one of a character, a letter, a symbol, a number, etc. The search string can be determined based on the search content input in an input box, or can be converted through a non-character input mode, for example, the content to be searched can be input through a voice input mode, then the input voice is converted into a search string according to voice recognition; for another example, an image or a video can be uploaded, then the image or the video is converted into a search string according to image recognition.

[0094] The search result includes a plurality of search resources, each of the search resources having a rank, for example, 1st, 2nd, etc. The priority of each search resource can be determined by ranking the search resources, and the search resources with higher ranks are preferentially displayed to the user to facilitate the user to quickly view. The search resource is an object retrieved by an information retrieval system, which can be a character written in natural language, or a media resource file such as voice, image, video, etc. According to different search objects, the search mode can be video search, text search, voice search, image search, or mixed search, etc.

[0095] In step S203, for each search pair, the search string and the corresponding search result in the search pair are respectively subjected to content extraction, and search content information corresponding to the search string and resource content information corresponding to the search result are respectively acquired.

[0096] In an embodiment, for the search string in each search pair, the search string can be subjected to word segmentation processing through a word segmentation tool to extract search content information corresponding to each search string. The search content information can be used to represent the actual search demand of a user corresponding to each search string. The content extraction of the search string can be extraction of the explicit content of the search string, or extraction of the implicit content or extended content of the search string. The word segmentation tool includes but is not limited to jieba, Ansj, mmseg4j, IKAnalyzer, paoding, etc.

[0097] For each search result corresponding to a search string, differential content extraction can be performed according to the object types (e.g. text, image, video, etc.) of the search resources in the search result, to extract resource content information corresponding to each search resource. The resource content information can be used to represent the content features of each search resource. For example, for a search resource of text type, content extraction can be performed by using the above-described word segmentation tool; for a search resource of image type, text content in the image can be recognized by using image recognition technology, and the text content contained in the image can be recognized by using OCR, and then content extraction is performed on the text content; for a search resource of video type, text information (e.g. video title, video abstract, tags, and text contained in pictures, etc.) and video frame information of the video can be extracted, and then content extraction is performed on the extracted text information and video frame information respectively.

[0098] In step S205, based on the search content information and the resource content information, first relationship data between the search string and each search resource, and second relationship data between each first search resource and each second search resource ranked in front are determined.

[0099] The first relationship data is used to represent the matching degree between the search string and each search resource, and the second relationship data is used to represent the content novelty degree of each first search resource in the search result.

[0100] In an embodiment, after obtaining the search content information of each search string and the resource content information of each search resource in the corresponding search result, the similarity between the resource content information of each search resource and the search content information of the search string is calculated, and then the first relationship data corresponding to each search resource is determined based on the similarity. The similarity between the Nth first search resource and the top N-1 second search resources is calculated, to determine the content repetition of the Nth first search resource with the top N-1 search resources, and then the second relationship data corresponding to the Nth first search resource is determined. Wherein, N is a positive integer.

[0101] For example, the first relationship data is used to represent the matching degree between the search string and each search resource. For example, the first relationship data can be a binary value (e.g. 0 or 1), wherein 1 represents that the matching degree of the search resource with the search string is high, and 0 represents that the matching degree of the search resource with the search string is low. Of course, the first relationship data can also be represented by other data, such as matching score, matching probability, etc.

[0102] The second relationship data is used to represent the novelty degree of the content of each first search resource in the search result. For example, the first relationship data can be a natural number, where 0 represents that there is no repeated content between the first N-1 search resources and the Nth search resource; 1 represents that there is one repeated content between the first N-1 search resources and the Nth search resource, 2 represents that there are two repeated contents between the first N-1 search resources and the Nth search resource, and so on. Of course, the second relationship data can also be represented by other data, such as a repetition score, a repetition probability, and the like.

[0103] In step S207, the quality information of the search result is determined based on the first relationship data and the second relationship data.

[0104] In an embodiment, the quality information of each search resource in the search result is determined based on the first relationship data and the second relationship data, then the quality information of each search resource in the search result is counted, and further the quality information of the search result is determined. The quality information at least represents the diversity degree of the search result. The quality information can be represented by a quality score or a quality level, which is not specifically limited in the present disclosure.

[0105] Since the first relationship data is used to reflect the matching degree between the search string and each search resource, and the second relationship data is used to reflect the novelty degree of the content of each first search resource in the search result, the quality information of the search result determined by the first relationship data and the second relationship data can reflect the diversity degree of the effective search resources in the search result.

[0106] In the embodiments of the present disclosure, at least one search pair is obtained, each search pair including a search string and a search result corresponding to the search string, the search result including a plurality of search resources sorted by search; for each search pair, content extraction is performed on the search string and the corresponding search result in the search pair respectively, and search content information corresponding to the search string and resource content information corresponding to the search result are obtained respectively; based on the search content information and the resource content information, first relationship data between the search string and each search resource, and second relationship data between each first search resource and each second search resource sorted in front are determined; the first relationship data is used to represent the matching degree between the search string and each search resource, and the second relationship data is used to represent the novelty degree of the content of each first search resource in the search result; and the quality information of the search result is determined based on the first relationship data and the second relationship data. Therefore, the diversity degree of the search result can be effectively represented by the quality information of the search result, which not only reduces the difficulty of depicting the diversity of the search result, but also improves the calculation efficiency and calculation accuracy of depicting the diversity of the search result.

[0107] In some embodiments, as shown in Figure 3 The step of extracting content of the search string and the search result in each search pair respectively to obtain search content information corresponding to the search string and resource content information corresponding to the search result in step S203 can include the following steps:

[0108] In step S301, a search sub-label corresponding to the search string in the search pair is obtained; the search sub-label is used to represent the user search intent information covered by the search string.

[0109] Optionally, in the search process, for each search string, a search sub-label corresponding to the search string can be obtained. The search sub-label is used to represent the user search intent information covered by the search string, for example, the search sub-label can be a lower-level search target of the content corresponding to the search string, or the synonymous content of the search string, the derivative content of the search string, the associated content of the search string, etc. The search sub-label can also be content determined based on the historical resource click situation corresponding to the search string, or content determined according to the previous search string and the subsequent search string of the search string.

[0110] For example, if the search string is the search word "Wang Zherong", the search sub-label corresponding to the search word "Wang Zherong" can include, for example, "national service", "mirror teaching", "Luban No. 7", "Han Xin teaching", "Agudo", etc.

[0111] It should be noted that in order to ensure the evaluation effect of the search result quality, the search sub-label here is not limited to the visible sub-label displayed on the terminal interface or application program, and includes more sub-labels. The number of search sub-labels can exceed 10, for example, 10-50, including but not limited to.

[0112] In step S303, for each search pair, content extraction is performed on the search sub-label corresponding to the search pair and each search resource in the search result, respectively, to obtain search content information corresponding to the search string and resource content information corresponding to the search result.

[0113] In an embodiment, for the search sub-label corresponding to the search string in each search pair, the search sub-label can be processed by a word segmentation tool to extract search content information corresponding to each search sub-label. The search content information can be used to represent the actual search demand of the user corresponding to each search sub-label. The word segmentation tool includes but is not limited to jieba, Ansj, mmseg4j, IKAnalyzer, paoding, etc.

[0114] For each search result corresponding to a search string, differential content extraction can be performed according to the object types (e.g., text, image, video, etc.) of the search resources in the search result, to extract resource content information corresponding to each search resource. The resource content information can be used to represent the content features of each search resource. For example, for a search resource of the text type, content extraction can be performed by using the above-described word segmentation tool; for a search resource of the image type, text content in the image can be identified by using image recognition technology, and the text content contained in the image can be identified by using OCR, and then the text content is subjected to content extraction; for a search resource of the video type, text information (e.g., video title, video abstract, label, and text contained in a picture) and video frame information of the video can be extracted, and then the extracted text information and video frame information are subjected to content extraction.

[0115] In the above embodiment, the search sub-labels corresponding to the search strings in the search pair are obtained, and then the search sub-labels corresponding to each search pair are subjected to content extraction, and the content information extracted from the search sub-labels is used as the search content information corresponding to the search string. Since the search sub-labels are used to reflect the user search intention information covered by the search string, more important information related to the user search intention can be obtained by using the search sub-labels, and thus the search content information contains more comprehensive user search intention information, which is beneficial to improving the accuracy and reliability of subsequent search result quality evaluation based on the search content information.

[0116] In some embodiments, the step S303 of performing content extraction on the search sub-labels corresponding to the search pair and each search resource in the search result, respectively, to obtain the search content information corresponding to the search string and the resource content information corresponding to the search result can include:

[0117] In step S3011, the search sub-labels are subjected to word segmentation processing, and the search content information corresponding to the search string is obtained.

[0118] Optionally, for any search string q, the TOP K search sub-labels corresponding thereto can be obtained, denoted as The K search sub-labels are subjected to word segmentation processing by using a word segmentation tool, and a search set is obtained after de-duplication processing. Taking the jieba tool as an example, the search set can be represented as:

[0119] In step S3013, the label dimension information and the sentence dimension information of each search resource in the search result corresponding to the search string are extracted, respectively.

[0120] Optionally, taking a video as an example of the search resource, the topic involved in the video information can be divided into three aspects: 1) a topic determined by recognizing a frame picture of the video; 2) a topic determined by converting a voice in the video into a text, a topic determined by converting into a text by using OCR recognition; and 3) a topic corresponding to a user input text such as a summary of the video and a title of the video. For any video p, the three aspects of the topic can be classified into two types of topic data:

[0121] a) label dimension content, for example, image classification labels, semantic classification labels, and the like.

[0122] Exemplarily, the label dimension information corresponding to the label dimension content can be obtained by extracting the label dimension content. The label dimension information can be a label set, denoted as term p .

[0123] b) sentence dimension content, for example, a text recognized by voice recognition, a text recognized by OCR recognition, and a user input text.

[0124] Exemplarily, the sentence dimension information can be obtained by performing a word segmentation processing on the sentence dimension content by using a word segmentation tool. The sentence dimension information can be a label set, denoted as sentence p .

[0125] In step S3015, the label dimension information and the sentence dimension information are merged to obtain resource content information corresponding to the search result.

[0126] Optionally, a set operation, for example, a union processing, is performed on the label dimension information and the sentence dimension information to merge the label dimension information and the sentence dimension information to obtain the resource content information corresponding to the search result. The resource content information topic p may be represented as: topic p = term p ∪(∪jieba(sentence p )).

[0127] The above embodiment, by performing word segmentation processing on the search sub-label, obtains search content information corresponding to the search string, which facilitates better reflecting the real user search intention and further improves the reliability and accuracy of the search result quality evaluation. Meanwhile, by performing information extraction on each search resource in the search result corresponding to the search string in the label dimension and the sentence dimension, and merging the extracted information in each dimension to obtain resource content information corresponding to the search result, such a resource content information not only reduces the overlap of each information in the resource content information, but also reduces the subsequent calculation amount, and makes the resource content information more fully reflect the difference between each search resource in the search result, thereby facilitating to reduce the difficulty of depicting the diversity of the search result, and further improving the reliability and accuracy of the search result quality evaluation.

[0128] In some embodiments, as shown in Figure 4 In step S205, the step of determining the first relationship data between the search string and each search resource, and the second relationship data between each first search resource and each second search resource ranked in front based on the search content information and the resource content information includes:

[0129] In step S401, the search content information and the resource content information are subjected to word embedding processing to obtain search vector representation of the search string and resource vector representation of each search resource, respectively.

[0130] The purpose of word embedding is to represent words as dense vectors with relative meanings, so as to facilitate computer reading and recognition. Specifically, ChineseWord2Vector, Glove, etc. word vector model can be used to perform word embedding processing on the search content information and the resource content information to obtain search vector representation of the search string and resource vector representation of each search resource, i.e. word vector of the search string and word vector of each search resource. The word vector can reflect the relative meaning between each word, and the words with similar meanings have similar word vectors.

[0131] By performing word embedding processing on the search content information and the resource content information, the words with similar meanings but different expressions can be uniformly represented by vectors, so as to facilitate subsequent judgment of the inclusion relationship between the search content information and the resource content information.

[0132] The search vector representation of the search string can be represented as:

[0133]

[0134] The resource vector representation of the search resource can be represented as:

[0135]

[0136] In step S403, based on the resource vector representation of each search resource and the search vector representation of the search string, the first relationship data between the search string and each search resource is determined.

[0137] Optionally, after obtaining the resource vector representation of each search resource and the search vector representation of the search string, the similarity between each search resource and the search vector representation corresponding to the search string can be calculated, and then the first relationship data between the search string and each search resource can be determined based on the calculated similarity.

[0138] In an optional embodiment, step S403 above, the step of determining the first relationship data between the search string and each search resource based on the resource vector representation of each search resource and the search vector representation of the search string, may include:

[0139] In step S4031, the first similarity between each word vector in the resource vector representation of each search resource and each word vector in the search vector representation is calculated, and the minimum similarity corresponding to each search resource is determined.

[0140] In step S4033, based on the determined minimum similarity and the first preset similarity threshold, the first relationship data between the search string and each search resource is determined.

[0141] In practical applications, for a search string q, the word vectors corresponding to its included topics are {c1, c2, c3, c4, c5}, and the corresponding sorted search results are {A, B, C}. For search resource A, the word vectors corresponding to its included topics are {c1, c2, c3, c4, c5}. a1 ,c a2}, then calculate the word vectors {c} respectively. a1 ,c a2 Word vector c in} a1 The first similarity scores with the word vectors in {c1,c2,c3,c4,c5} are obtained to obtain A11, A12, A13, and A14; the word vectors {c a1 ,c a2 Word vector c in} a2 The first similarity scores with the word vectors in {c1,c2,c3,c4,c5} are obtained as A21, A22, A23, and A24. If the numerical order of the first similarity scores is A11 > A12 > A13 > A14, and A22 > A21 > A24 > A23, then the minimum similarity score is determined from the calculated first similarity scores as the representative word vector c. a1 Similarity (e.g., A14) and representative word vector c a2 Similarity (e.g., A23).

[0142] Then, the two first similarity values (A14 and A23) are respectively compared with a first preset similarity threshold, and the comparison results are taken as the first relationship data between the search string and each search resource. If the comparison result is greater than or equal to the first preset similarity threshold, it is considered that the word vector pair used for similarity calculation is matched, that is, one word vector in the resource vector representation of the search resource matches the corresponding word vector in the search vector. If the comparison result is greater than or equal to the first preset similarity threshold, it is considered that the word vector pair used for similarity calculation is not matched, that is, one word vector in the resource vector representation of the search resource does not match all the corresponding word vectors in the search vector.

[0143] The first preset similarity threshold can be, but is not limited to, any value in the range of 0.2-0.5, which is not limited in the present application. The similarity can be, but is not limited to, cosine distance, Euclidean distance, etc. Taking the cosine distance as an example, the similarity calculation formula can be:

[0144]

[0145] In the above embodiment, the minimum similarity between each word vector in the resource vector representation of each search resource and each word vector in the search vector representation is calculated, and the first relationship data between the search string and each search resource is determined based on the size of the minimum similarity and the first preset similarity threshold, so that the matching degree between the word vector in the resource vector representation and each word vector in the search vector representation is intuitively reflected by the minimum similarity, and the first preset similarity threshold can be flexibly adjusted to adjust the first relationship data between the search string and each search resource, which is beneficial to improve the sensitivity and reliability of the first relationship data, and further improve the reliability and accuracy of the quality evaluation of the search result.

[0146] In step S405, the relationship data between the Nth search resource and the top N-1 search resources in the search result is determined based on the resource vector representation of the Nth search resource and the resource vector representations of the top N-1 search resources in the search result, where N is a positive integer.

[0147] Optionally, as shown in Figure 5 The step of determining the relationship data between the Nth search resource and the top N-1 search resources in the search result based on the resource vector representation of the Nth search resource and the resource vector representations of the top N-1 search resources in the search result can include:

[0148] In step S4051, the arrangement sequence set of the top N-1 search resources is determined from the search result based on the arrangement order of each search resource in the search result.

[0149] In step S4052, each word vector in the target word vector set is extracted from the resource vector representation of each search resource in the ranking sequence set, and a target word vector set is obtained.

[0150] In step S4053, the similarity between each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource is calculated, and the second similarity corresponding to each word vector in the target word vector set is determined.

[0151] In step S4054, the first number of the second similarity greater than or equal to the second preset similarity threshold is determined.

[0152] In step S4055, the relationship data between the Nth search resource and the top N-1 search resources is determined based on the determined first number.

[0153] The second similarity can reflect the maximum similarity between each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource.

[0154] In the above embodiment, the target word vector set corresponding to each word vector in the resource vector representation of the top N-1 search resources is constructed, and then the similarity between each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource is calculated to determine the second similarity corresponding to each word vector in the top N-1 search resources and each word vector in the Nth search resource. Then, the second similarity corresponding to each word vector in the top N-1 search resources is compared with the second preset similarity threshold, and the first number of the second similarity greater than or equal to the second preset similarity threshold is counted, and then the relationship data between the Nth search resource and the top N-1 search resources is determined based on the first number. Thus, the similarity between each word vector in the search vector representation is intuitively reflected by the second similarity, and the relationship data between the Nth search resource and the top N-1 search resources can be flexibly adjusted by the second preset similarity threshold, which is beneficial to improve the sensitivity and reliability of the second relationship data, and further improve the reliability and accuracy of the quality evaluation of the search result.

[0155] In step S407, the relationship data between the Nth search resource and the top N-1 search resources is used as the second relationship data between each first search resource and the top second search resource.

[0156] Optionally, for a search string q, the corresponding sorted search result is {A, B, C}, wherein the search resource A includes topic corresponding word vectors {c a1 ,c a2}; the search resource B includes topic corresponding word vectors {cb3 a6 b7}, search resource C, which contains the topic corresponding to each word vector {c c3 c4 c5}. For search resource C, the similarity between each word vector in search resource A and each word vector in search resource B and c c3 is calculated, that is, the similarity between each word vector in {c a1 a2 b3 a6 b7} and c c3 is calculated to determine the second similarity corresponding to each word vector. If it is determined that only the second similarity between c b3 and c c3 is greater than or equal to the second preset similarity threshold, it is determined that the first number is 1. If it is determined that there are two second similarities greater than or equal to the second preset similarity threshold, it is determined that the first number is 2. The greater the first number, the more similar topics exist in the first N-1 search resources to the Nth search resource, and therefore, the greater the first number, the lower the novelty of the corresponding word vector in the Nth search resource. Conversely, the smaller the first number or the first number is 0, the higher the novelty of the corresponding word vector in the Nth search resource.

[0157] In the above embodiments, by performing word embedding processing on search content information and resource content information, corresponding vector representations are obtained, so that content information with similar meanings or similar meanings can have the same or similar vector representations, facilitating accurate determination of the inclusion relationship between each word vector in the search content information and the resource content information. Then, based on the search vector representation of the search string and the resource vector representation of each search resource obtained by word embedding processing, the first relationship data between the search string and each search resource is determined, and the second relationship data between each first search resource and the second search resource with a high ranking is determined, thereby realizing more accurate expression of the first relationship data and the second relationship data, reducing the amount of calculation, and further improving the reliability, accuracy and calculation efficiency of the quality evaluation of the search result.

[0158] In some embodiments, as shown in Figure 6 , the step of determining the quality information of the search result based on the first relationship data and the second relationship data in the above step S207 can include:

[0159] In step S601, the sub-quality information of each search resource is determined based on the first relationship data, the second relationship data, and the ranking number of each search resource in the search result. ​​​​​​​​

[0160] In step S603, the quality information of the search result is determined by summing the sub-quality information of each search resource in the search result.

[0161] For example, the quality information M of the search result can be represented by the following formula:

[0162]

[0163] where i represents the ranking of the search resource, and j represents the jth word vector in the ith search resource. ij represents the first relationship data between the jth word vector and all word vectors in the search string. If there is a word vector in the search string that matches the jth word vector, then J ij is 1, otherwise J ij is 0. (1-α) rj represents the second relationship data between the jth word vector and all word vectors in the first i-1 search resources. α is a hyperparameter, which can take a value less than 1, for example 0.5; and rj represents the number of search resources in the first i-1 search resources that have a word vector matching the jth word vector.

[0164] To facilitate understanding, the aforementioned search string q and the word vectors in each search resource are simplified by numbers, for example, the word vectors corresponding to the topic contained in q are {1, 2, 3, 4, 5}, and the sub-quality information of each search resource in the search results {A, B, C} is:

[0165] The search resource A contains the word vectors {1, 2} corresponding to the topic,

[0166] The search resource B contains the word vectors {3, 6, 7} corresponding to the topic,

[0167] The search resource C contains the word vectors {3, 4, 5} corresponding to the topic,

[0168] In this embodiment, the quality information M of the search result is obtained by calculating the sum of Ma, Mb, and Mc, which represents the diversity degree of the search result.

[0169] The above embodiment determines the sub-quality information of each search resource based on the first relationship data, the second relationship data and the ranking number of each search resource in the search result, and the sub-quality information is used to reflect the quality of each search resource in the current search result. Then, the sub-quality information of each search resource in the search result is summed to determine the quality information of the search result, so that the diversity of the search result quality is more comprehensively and accurately expressed, the difficulty of depicting the diversity of the search result is reduced, and the calculation efficiency and accuracy of depicting the diversity of the search result are further improved.

[0170] Of course, in other embodiments, other mathematical operations can be performed on the sub-quality information of each search resource in the search result to obtain the quality information of the search result. For example, at least one mathematical operation such as but not limited to weighted summation operation, mean operation, variance operation, etc. can be performed on Ma, Mb and Mc to obtain the quality information M of the search result.

[0171] Figure 7 FIG. 7 is a block diagram of a device for determining the quality of a search result according to an exemplary embodiment. Referring to FIG. 7, Figure 7 The device includes:

[0172] The acquisition module 710 is configured to perform acquisition of at least one search pair, each search pair including a search string and a search result corresponding to the search string, the search result including a plurality of search resources sorted by search;

[0173] The content extraction module 720 is configured to perform, for each search pair, content extraction on the search string and the corresponding search result in the search pair respectively, and to obtain search content information corresponding to the search string and resource content information corresponding to the search result respectively;

[0174] The relationship determination module 730 is configured to perform determination of first relationship data between the search string and each search resource and second relationship data between each first search resource and each second search resource sorted in front based on the search content information and the resource content information; the first relationship data is used to represent the matching degree between the search string and each search resource, and the second relationship data is used to represent the content novelty degree of each first search resource in the search result;

[0175] The quality determination module 740 is configured to perform determination of quality information of the search result based on the first relationship data and the second relationship data, the quality information at least representing the diversity degree of the search result.

[0176] As an optional implementation, the content extraction module 720 includes:

[0177] an obtaining sub-module configured to perform obtaining a search sub-label corresponding to a search string in the search pair; the search sub-label is used to represent user search intention information covered by the search string;

[0178] a content extraction sub-module configured to perform, for each search pair, content extraction on the search sub-label corresponding to the search pair and each search resource in the search result, respectively, to obtain search content information corresponding to the search string and resource content information corresponding to the search result.

[0179] As an optional implementation, the content extraction sub-module is configured to perform:

[0180] performing word segmentation processing on the search sub-label to obtain search content information corresponding to the search string;

[0181] extracting label dimension information and sentence dimension information of each search resource in the search result corresponding to the search string, respectively;

[0182] merging the label dimension information and the sentence dimension information to obtain resource content information corresponding to the search result.

[0183] As an optional implementation, the relationship determination module 730 includes:

[0184] a first processing sub-module configured to perform word embedding processing on the search content information and the resource content information to obtain search vector representation of the search string and resource vector representation of each search resource, respectively;

[0185] a first relationship determination sub-module configured to perform, based on the resource vector representation of each search resource and the search vector representation of the search string, determination of first relationship data between the search string and each search resource;

[0186] a second relationship determination sub-module configured to perform, based on the resource vector representation of an Nth search resource and resource vector representations of the first N-1 search resources in the search result, determination of relationship data between the Nth search resource and the first N-1 search resources in the order, where N is a positive integer;

[0187] a second processing sub-module configured to perform, as second relationship data between each first search resource and each second search resource in the order, the relationship data between the Nth search resource and the first N-1 search resources in the order.

[0188] As an optional implementation, the first relationship determination sub-module is configured to perform:

[0189] calculate a first similarity between each word vector in the resource vector representation of each search resource and each word vector in the search vector representation, and determine a minimum similarity corresponding to each search resource;

[0190] determine first relationship data between the search string and each search resource based on the determined minimum similarity and a first preset similarity threshold

[0191] As an optional implementation, the second relationship determining submodule is configured to perform:

[0192] determine a set of ranking numbers of the first N-1 search resources from the search results based on the ranking order of the search resources in the search results;

[0193] extract each word vector in the resource vector representation of each search resource in the set of ranking numbers to obtain a target word vector set;

[0194] calculate a second similarity corresponding to each word vector in the target word vector set based on the similarity between each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource;

[0195] determine a first number of maximum similarities greater than or equal to a second preset similarity threshold;

[0196] determine relationship data between the Nth search resource and the first N-1 search resources based on the determined first number.

[0197] As an optional implementation, the quality determining module 740 includes:

[0198] a first determining submodule configured to determine sub-quality information of each search resource based on the first relationship data, the second relationship data, and the ranking number of each search resource in the search results;

[0199] a second determining submodule configured to sum the sub-quality information of each search resource in the search results to determine quality information of the search results.

[0200] As for the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.

[0201] Figure 8 is a block diagram of an electronic device according to an example embodiment. Referring to Figure 8The electronic device includes a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the method for determining the quality of search results in any of the above embodiments when executing the instructions stored on the memory.

[0202] The electronic device can be a device, a terminal, a server or a similar computing device with Bluetooth function. In this embodiment, the electronic device is taken as an example of a server, Figure 8 FIG. 1 is a block diagram of an electronic device for determining the quality of search results according to an example embodiment. The electronic device 1000 can have a large difference in configuration or performance, and can include one or more central processing units (CPUs) 1010 (the processor 1010 can include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA processing device), a memory 1030 for storing data, one or more storage media 1020 (such as one or more mass storage devices) for storing application programs 1023 or data 1022. The memory 1030 and the storage medium 1020 can be temporary storage or persistent storage. The programs stored in the storage medium 1020 can include one or more modules, each of which can include a series of instruction operations in the electronic device. Further, the central processing unit 1010 can be configured to communicate with the storage medium 1020 and execute a series of instruction operations in the storage medium 1020 on the electronic device 1000.

[0203] The electronic device 1000 can also include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1040, and / or one or more operating systems 1021, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0204] The input / output interface 1040 can be used to receive or send data via a network. The above-mentioned network can include a wireless network provided by a communication provider of the electronic device 1000. In one example, the input / output interface 1040 includes a network adapter (NIC) that can be connected to other network devices through a base station so as to communicate with the Internet. In an example embodiment, the input / output interface 1040 can be a radio frequency (RF) module for communicating with the Internet in a wireless manner.

[0205] Those of ordinary skill in the art can understand that Figure 8The illustrated structure is merely an example and does not limit the structure of the electronic device described above. For example, the electronic device 1000 can further include more or less components than those shown, or have a different configuration of components than those shown. Figure 8 The illustrated structure is merely an example and does not limit the structure of the electronic device described above. For example, the electronic device 1000 can further include more or less components than those shown, or have a different configuration of components than those shown. Figure 8 The illustrated structure is merely an example and does not limit the structure of the electronic device described above. For example, the electronic device 1000 can further include more or less components than those shown, or have a different configuration of components than those shown.

[0206] In an exemplary embodiment, a storage medium including instructions, for example, a memory including instructions, is also provided, and the instructions are executable by a processor of the electronic device 1000 to perform the above-described method. Alternatively, the storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.

[0207] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the method for determining the quality of search results provided in any one of the embodiments described above.

[0208] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the present disclosure cover any and all variations of the present disclosure including those variations that are now deemed to fall within the general principles of the present disclosure and including those variations that are deemed to fall within the patentably distinct and novel aspects of the present disclosure. The specification and examples given are intended as illustrative only and do not limit the true scope and spirit of the present disclosure, which is indicated by the following claims.

[0209] It should be understood that the present disclosure is not limited to the precise structures described and shown in the drawings, and that various modifications and changes can be made to the present disclosure without departing from its scope. The scope of the present disclosure is limited only by the claims that follow.

Claims

1. A method of determining the quality of search results, characterized by, The method comprises the following steps: obtaining at least one search pair, each search pair comprising a search string and a search result corresponding to the search string, the search result comprising a plurality of search resources sorted by search; for each search pair, respectively extracting the search string and the corresponding search result in the search pair to obtain search content information corresponding to the search string and resource content information corresponding to the search result; based on the search content information and the resource content information, determining first relationship data between the search string and each search resource, and second relationship data between each first search resource and each second search resource sorted in front; the first relationship data is used to represent the matching degree between the search string and each search resource, and the second relationship data is used to represent the content novelty degree of each first search resource in the search result; the first relationship data is determined based on resource vector representation of the search resource and search vector representation of the search string, and the second relationship data is determined based on resource vector representation of the Nth search resource and resource vector representation of the first N-1 search resources in the search result; based on the first relationship data and the second relationship data, determining quality information of the search result, the quality information at least representing the diversity degree of the search result.

2. The method of claim 1, wherein, The step of respectively extracting the search string and the corresponding search result in the search pair to obtain search content information corresponding to the search string and resource content information corresponding to the search result comprises: obtaining search sub-labels corresponding to the search string in the search pair; the search sub-labels are used to represent user search intent information covered by the search string; for each search pair, respectively extracting the search sub-labels corresponding to the search pair and each search resource in the search result to obtain search content information corresponding to the search string and resource content information corresponding to the search result.

3. The method of claim 2, wherein, The step of respectively extracting the search sub-labels corresponding to the search pair and each search resource in the search result to obtain search content information corresponding to the search string and resource content information corresponding to the search result comprises: performing word segmentation processing on the search sub-labels to obtain search content information corresponding to the search string; respectively extracting label dimension information and sentence dimension information of each search resource in the search result corresponding to the search string; merging the label dimension information and the sentence dimension information to obtain resource content information corresponding to the search result.

4. The method of claim 1-3, wherein, The step of determining the first relationship data between the search string and each search resource, and the second relationship data between each first search resource and each second search resource sorted in front based on the search content information and the resource content information comprises: performing word embedding processing on the search content information and the resource content information to obtain search vector representation of the search string and resource vector representation of each search resource, respectively; determine first relationship data of the search string and each search resource based on the resource vector representation of each search resource and the search vector representation of the search string; determine relationship data of the Nth search resource and the top N-1 search resources based on the resource vector representation of the Nth search resource and the resource vector representations of the top N-1 search resources in the search result, where N is a positive integer; use the relationship data of the Nth search resource and the top N-1 search resources as second relationship data of each first search resource and each second search resource.

5. The method of claim 4, wherein the step of determining first relationship data of the search string and each search resource based on the resource vector representation of each search resource and the search vector representation of the search string comprises: determining minimum similarity of each search resource based on first similarity of each word vector in the resource vector representation of each search resource and each word vector in the search vector representation; determining first relationship data of the search string and each search resource based on the determined minimum similarity and a first preset similarity threshold.

6. The method of claim 4, wherein the step of determining relationship data of the Nth search resource and the top N-1 search resources based on the resource vector representation of the Nth search resource and the resource vector representations of the top N-1 search resources in the search result comprises: determining a set of ranking numbers of the top N-1 search resources from the search result based on ranking order of each search resource in the search result; extracting word vectors in the resource vector representation of each search resource in the set of ranking numbers to obtain a target word vector set; determining second similarity of each word vector in the target word vector set based on similarity of each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource; determining a first number of maximum similarity greater than or equal to a second preset similarity threshold; determining relationship data of the Nth search resource and the top N-1 search resources based on the determined first number.

7. The method of claim 1-3, wherein, The step of determining quality information of the search result based on the first relationship data and the second relationship data comprises: determining sub-quality information of each search resource based on the first relationship data, the second relationship data, and ranking number of each search resource in the search result; summing sub-quality information of each search resource in the search result to determine quality information of the search result.

8. An apparatus for determining the quality of search results, characterized in that The method comprises: an obtaining module configured to obtain at least one search pair, each search pair comprising a search string and a search result corresponding to the search string, the search result comprising a plurality of search resources sorted by search. The content extraction module is configured to perform content extraction on the search string and the corresponding search result in each search pair respectively to obtain search content information corresponding to the search string and resource content information corresponding to the search result respectively; The relationship determination module is configured to determine first relationship data between the search string and each search resource and second relationship data between each first search resource and each second search resource in the search result based on the search content information and the resource content information; The first relationship data is used to represent the matching degree between the search string and each search resource, and the second relationship data is used to represent the novelty degree of the content of each first search resource in the search result; The first relationship data is determined based on resource vector representation of the search resource and search vector representation of the search string, and the second relationship data is determined based on resource vector representation of the Nth search resource and resource vector representation of the first N-1 search resources in the search result; The quality determination module is configured to determine quality information of the search result based on the first relationship data and the second relationship data, and the quality information at least represents the diversity degree of the search result.

9. The search result quality determination apparatus according to claim 8, characterized by, The content extraction module comprises: The obtaining submodule is configured to obtain search sub-labels corresponding to the search string in the search pair, and the search sub-labels are used to represent user search intention information covered by the search string; The content extraction submodule is configured to perform content extraction on the search sub-labels and each search resource in the search result in the search pair respectively to obtain search content information corresponding to the search string and resource content information corresponding to the search result.

10. The search result quality determination apparatus according to claim 9, wherein The content extraction submodule is configured to perform: performing word segmentation processing on the search sub-labels to obtain the search content information corresponding to the search string; extracting label dimension information and sentence dimension information of each search resource in the search result corresponding to the search string respectively; merging the label dimension information and the sentence dimension information to obtain the resource content information corresponding to the search result.

11. The apparatus of any of claims 8-10, wherein, The relationship determination module comprises: The first processing submodule is configured to perform word embedding processing on the search content information and the resource content information to obtain search vector representation of the search string and resource vector representation of each search resource respectively; The first relationship determination submodule is configured to determine the first relationship data between the search string and each search resource based on the resource vector representation of each search resource and the search vector representation of the search string; The second relationship determination submodule is configured to determine relationship data between the Nth search resource and the first N-1 search resources in the search result based on the resource vector representation of the Nth search resource and the resource vector representation of the first N-1 search resources, wherein N is a positive integer. The second processing submodule is configured to execute the relationship data between the Nth search resource and the N-1 search resources with high ranking as the second relationship data between each first search resource and each second search resource with high ranking.

12. The search result quality determination apparatus according to claim 11, wherein, The first relationship determining submodule is configured to execute: calculating the first similarity between each word vector in the resource vector representation of each search resource and each word vector in the search vector representation, and determining the minimum similarity corresponding to each search resource; determining the first relationship data between the search string and each search resource based on the size between the determined minimum similarity and the first preset similarity threshold.

13. The search result quality determination apparatus according to claim 11, wherein, The second relationship determining submodule is configured to execute: determining the ranking index set of the N-1 search resources from the search results based on the ranking order of each search resource in the search results; extracting each word vector in the resource vector representation of each search resource belonging to the ranking index set to obtain a target word vector set; calculating the similarity between each word vector in the target word vector set and each word vector in the resource vector representation of the Nth search resource, and determining the second similarity corresponding to each word vector in the target word vector set; determining the first number of the maximum similarity greater than or equal to the second preset similarity threshold; determining the relationship data between the Nth search resource and the N-1 search resources with high ranking based on the determined first number.

14. The apparatus for determining the quality of search results according to any one of claims 8-10, wherein, The quality determining module includes: a first determining submodule configured to determine the sub-quality information of each search resource based on the first relationship data, the second relationship data, and the ranking order of each search resource in the search results; a second determining submodule configured to sum the sub-quality information of each search resource in the search results to determine the quality information of the search results.

15. An electronic device, comprising: It includes: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the search result quality determination method according to any one of claims 1 to 7.

16. A storage medium, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the search result quality determination method according to any one of claims 1 to 7.

17. A computer program product comprising computer instructions, characterized in that, The computer instructions are executed by the processor to implement the search result quality determination method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Search result providing method and apparatus

    CN105117383A

  • Natural language querying with cascaded conditional random fields

    US20120254143A1