An information processing method, apparatus, electronic device, and storage medium

By constructing a click sequence in a search engine, identifying search terms that deviate from the semantic center, and determining bad examples of search engines, the problem of degradation in the quality of search results in the prior art is solved, and the user experience and the accuracy of search results are improved.

CN111814471BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010704980.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-21
Publication Date
2025-06-24
Estimated Expiration
2040-07-21

AI Technical Summary

Technical Problem

When existing search engines mine and deal with bad examples, it is difficult for them to accurately identify the degree of matching between user intentions and search results, resulting in a decline in the quality of search results and affecting user experience.

Method used

By obtaining the log information of the search engine, building a click sequence, determining the search term collection, and mapping it to the semantic space for encoding processing, identifying search terms that deviate from the semantic center, thereby determining bad examples of search engines.

Benefits of technology

It effectively improves the quality of search results, improves the user experience, and improves the accuracy of search engine recommendations through more accurate and comprehensive mining of bad examples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111814471B_ABST
    Figure CN111814471B_ABST
Patent Text Reader

Abstract

The present invention provides an information processing method, apparatus, electronic device and storage medium. The method includes: obtaining log information of a search engine, and determining a click sequence formed by a search term and a corresponding search result based on the log information of the search engine; determining a corresponding set of search terms according to different search results in different click sequences; mapping different search terms in the set of search terms to a matching semantic space to form corresponding search term vectors; determining search terms in the set of search terms that deviate from the semantic center based on the search term vectors; and determining a set formed by the corresponding search terms and search results as a bad example of the search engine according to the search terms that deviate from the semantic center. Thereby, it is possible to fully mine bad examples of the search engine, obtain more accurate and comprehensive bad examples, effectively improve the quality of search results, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information processing technology, and in particular to an information processing method, apparatus, electronic device, and storage medium. Background Art

[0002] In related technologies, the internal of a search engine is composed of many complex and coupled relevance strategies. The quantity, complexity, and mutual restriction relationships are relatively complex. Users can obtain relevant search results through the search engine for browsing. The search engine can continuously improve the accuracy of its own search results by learning and overcoming bad cases. Inaccurate bad cases may lead to a decline in the quality of search results and affect user experience. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide an information processing method, apparatus, electronic device, and storage medium, which can fully mine the bad cases of a search engine, obtain more accurate and comprehensive bad cases, effectively improve the quality of search results, and enhance user experience.

[0004] The technical solution of the embodiments of the present invention is implemented as follows:

[0005] Embodiments of the present invention provide an information processing method, including:

[0006] Obtain the log information of a search engine, and determine a click sequence formed by a search term and a corresponding search result based on the log information of the search engine;

[0007] Determine a corresponding set of search terms according to different search results in different click sequences;

[0008] Map different search terms in the set of search terms to a matching semantic space, and perform encoding processing on different search terms in the set of search terms to form corresponding search term vectors;

[0009] Based on the search term vectors, determine the search terms in the set of search terms that deviate from the semantic center;

[0010] According to the search terms that deviate from the semantic center, determine a set formed by the corresponding search terms and search results as the bad cases of the search engine.

[0011] Embodiments of the present invention further provide an information processing apparatus, including:

[0012] An information transmission module, configured to obtain the log information of a search engine, and determine a click sequence formed by a search term and a corresponding search result based on the log information of the search engine;

[0013] An information processing module, configured to determine a corresponding set of search terms according to different search results in different click sequences;

[0014] The information processing module is configured to map different search terms in the set of search terms to a matching semantic space, and perform encoding processing on different search terms in the set of search terms to form corresponding search term vectors;

[0015] The information processing module is configured to determine, based on the search term vectors, search terms in the set of search terms that deviate from the semantic center;

[0016] The information processing module is configured to determine, according to the search terms that deviate from the semantic center, a set formed by the corresponding search terms and search results as a bad example of the search engine.

[0017] In the above solution,

[0018] The information processing module is configured to trigger a corresponding word segmentation library according to the search term parameter information carried in the log information of the search engine;

[0019] The information processing module is configured to perform word segmentation processing on the log information of the search engine through the word dictionary of the triggered word segmentation library to form different word-level search terms and sentence-level search terms;

[0020] The information processing module is configured to determine search results corresponding to the different word-level search terms and sentence-level search terms respectively;

[0021] The information processing module is configured to perform noise reduction processing on the different word-level search terms and sentence-level search terms to form a click sequence corresponding to the log information of the search engine, where the click sequence includes a word-level search term and a corresponding search result, or a sentence-level search term and a corresponding search result.

[0022] In the above solution,

[0023] The information processing module is configured to determine the name of the word segmentation library used when performing word segmentation processing on the search term text;

[0024] The information processing module is configured to determine, according to the name of the word segmentation library, parameters of the word segmentation library that match the word-level feature vectors corresponding to the search term text, where the parameters of the word segmentation library include:

[0025] The type of the word segmentation library, the name of the word segmentation library, and the version of the word segmentation library.

[0026] In the above solution,

[0027] The information processing module is used to determine the search results in the click sequence based on the sorting of the click sequence;

[0028] The information processing module is used to traverse the log information of the search engine according to the search results, and determine the search terms that match the search results based on the behavior records of different target users;

[0029] The information processing module is used to combine the search terms corresponding to the same search result respectively to determine the corresponding set of search terms.

[0030] In the above solution,

[0031] The information processing module is used to determine the semantic space that matches the search term through any search term in the set of search terms;

[0032] The information processing module is used to determine at least one word-level vector corresponding to the target search term through the encoder of the information processing model;

[0033] The information processing module is used to map at least one word-level vector corresponding to the target search term to the semantic space, and perform iterative processing on the set of search terms until all search terms in the set of search terms are mapped to the matching semantic space, forming a search term vector that matches the set of search terms.

[0034] In the above solution,

[0035] The information processing module is used to obtain a first training sample set, where the first training sample set includes at least one set of statement samples with noise;

[0036] The information processing module is used to perform denoising processing on the first training sample set to form a corresponding second training sample set;

[0037] The information processing module is used to process the first training sample set through the information processing model to determine the initial parameters of the information processing model;

[0038] The information processing module is used to respond to the initial parameters of the information processing model, and process the second training sample set through the information processing model to determine the updated parameters of the information processing model;

[0039] The information processing module is used to iteratively update the encoder parameters and decoder parameters of the information processing model according to the updated parameters of the information processing model through the first training sample set and the second training sample set, so that the information processing model can encode the corresponding search terms.

[0040] In the above solution,

[0041] the information processing module is used to determine a dynamic noise threshold that matches the usage environment of the information processing model;

[0042] the information processing module is used to denoise the first training sample set according to the dynamic noise threshold to form a second training sample set that matches the dynamic noise threshold;

[0043] the information processing module is used to determine a fixed noise threshold corresponding to the information processing model, and denoise the first training sample set according to the fixed noise threshold to form a second training sample set that matches the fixed noise threshold.

[0044] In the above solution,

[0045] the information processing module is used to perform negative example processing on the first training sample set to form a negative example sample set corresponding to the first training sample set, where the negative example sample set includes bad example sample data in user behavior data and is used to adjust the encoder parameters and decoder parameters of the information processing model.

[0046] In the above solution,

[0047] the information processing module is used to substitute different statement samples in the second training sample set into the loss function corresponding to the autoencoder network composed of the encoder and decoder of the information processing model;

[0048] the information processing module is used to determine the parameters of the encoder and the corresponding decoder parameters in the information processing model as the update parameters of the information processing model when the loss function satisfies the convergence condition.

[0049] In the above solution,

[0050] the information processing module is used to determine the semantic center corresponding to the search term set based on the search term vector;

[0051] the information processing module is used to determine the distance from each search term in the search term set to the semantic center;

[0052] the information processing module is used to determine that the current search term deviates from the semantic center when the distance from the search term to the semantic center is greater than the corresponding distance threshold;

[0053] the information processing module is used to perform iterative processing on different search terms in the search term set until all search terms that deviate from the semantic center in the search term set are determined.

[0054] In the above solution,

[0055] the information processing module is used to determine that the bad case set of the search engine is an empty set when different search terms in the search term set do not exceed the distance threshold.

[0056] In the above solution, the device further includes:

[0057] a display module for displaying a user interface, where the user interface includes a perspective view of using the search engine in a corresponding software process from a first-person perspective, and the user interface further includes a display control component;

[0058] the display module is used to control the display of search results matching the search terms input by the user through the display control component.

[0059] An embodiment of the present invention further provides an electronic device, which includes:

[0060] a memory for storing executable instructions;

[0061] a processor for implementing the foregoing information processing method when running the executable instructions stored in the memory.

[0062] An embodiment of the present invention further provides a computer-readable storage medium storing executable instructions, and the executable instructions implement the foregoing information processing method when executed by a processor.

[0063] The embodiment of the present invention has the following beneficial effects:

[0064] The present invention obtains the log information of the search engine, and determines the click sequence formed by the search terms and the corresponding search results based on the log information of the search engine; determines the corresponding search term set according to different search results in different click sequences; maps different search terms in the search term set to a matching semantic space, and encodes different search terms in the search term set to form corresponding search term vectors; determines the search terms deviating from the semantic center in the search term set based on the search term vectors; determines the set formed by the corresponding search terms and search results as the bad cases of the search engine according to the search terms deviating from the semantic center. This application can fully mine the bad cases of the search engine, obtain more accurate and comprehensive bad cases, effectively improve the quality of search results, and improve the user experience. Description of the Drawings

[0065] Figure 1 It is a schematic diagram of the usage scenario of the information processing method provided by the embodiment of the present invention;

[0066] Figure 2 Schematic diagram of the composition structure of the server provided by an embodiment of the present invention;

[0067] Figure 3 Schematic diagram of bad case mining in related technologies in an embodiment of the present invention;

[0068] Figure 4 Schematic diagram of the influence of bad cases on search results in an embodiment of the present invention;

[0069] Figure 5 An optional flowchart of the information processing method provided by an embodiment of the present invention;

[0070] Figure 6A Schematic diagram of the click sequence formed by the search term and the corresponding search results in an embodiment of the present invention;

[0071] Figure 6B Schematic diagram of the composition of the click sequence in an embodiment of the present invention;

[0072] Figure 7 An optional flowchart of the information processing method provided by an embodiment of the present invention;

[0073] Figure 8 An optional information processing process diagram of the information processing model in an embodiment of the present invention;

[0074] Figure 9 An optional structural diagram of the encoder in the information processing model in an embodiment of the present invention;

[0075] Figure 10 Schematic diagram of vector splicing of the encoder in the information processing model in an embodiment of the present invention;

[0076] Figure 11 Schematic diagram of the encoding process of the encoder in the information processing model in an embodiment of the present invention;

[0077] Figure 12 Schematic diagram of the information processing effect in an embodiment of the present invention;

[0078] Figure 13 An optional flowchart of the information processing method provided by an embodiment of the present invention;

[0079] Figure 14 Schematic diagram of the information processing effect in an embodiment of the present invention. Detailed implementation manners

[0080] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0081] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0082] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.

[0083] 1) Responsive to, which is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations may be real-time or may have a set delay; without special instructions, there is no limitation on the execution order of the multiple executed operations.

[0084] 2) Word segmentation: Also known as tokenization, its function is to split the text information of a complete sentence into multiple words. For example: Liu Dehua is a Chinese singer. The result after word segmentation is: Liu Dehua, Chinese, singer.

[0085] 3) Word segmentation library: Also known as tokenization library, which refers to a specific word segmentation method. Different word segmentation libraries have their respective word dictionaries and can perform word segmentation processing on the corresponding text information according to their respective word dictionaries.

[0086] 4) Article: The text information included in different documents in Internet resources, such as word, html web pages, etc., or the text information in pictures.

[0087] 5) Word: It is a string that is determined by splitting the content of an article or a search term input by a user and logically forms a complete expression.

[0088] 6) Word dictionary: Stores all words, and each record consists of a word and a pointer pointing to an inverted list.

[0089] 7) Neural Network (NN): Artificial Neural Network (ANN), also simply referred to as neural network or neural-like network, is a mathematical model or computational model that mimics the structure and function of biological neural networks (the central nervous system of animals, especially the brain) in the fields of machine learning and cognitive science, and is used to estimate or approximate functions.

[0090] 8) BERT: Short for Bidirectional Encoder Representations from Transformers, it is a method for training language models using a large amount of text. This method is widely used in various natural language processing tasks, such as text classification, text matching, machine reading comprehension, etc.

[0091] 9) Artificial neural network: Also simply referred to as neural network (Neural Network, NN), it is a mathematical model or computational model that mimics the structure and function of biological neural networks in the fields of machine learning and cognitive science, and is used to estimate or approximate functions.

[0092] 10) Model parameters: They are a quantity that uses general variables to establish the relationship between a function and variables. In an artificial neural network, model parameters are usually real number matrices.

[0093] 11) Model training involves multi-class learning on an image dataset. This model can be constructed using deep learning frameworks such as TensorFlow and torch, and a multi-class model is formed by combining multiple layers of neural network layers such as CNN. The input of the model is a three-channel or original-channel matrix formed by reading an image through tools such as openCV, and the output of the model is multi-class probabilities, and finally the web page category is output through algorithms such as softmax. During training, the model approaches the correct trend through objective functions such as cross-entropy.

[0094] 12) Bidirectional Attention Neural Network Model (BERT Bidirectional Encoder Representations from Transformers): A bidirectional attention neural network model proposed by Google. Transformers: A new network structure that uses the attention mechanism to replace the traditional encoder-decoder mode that must rely on other neural networks.

[0095] Figure 1 It is a schematic diagram of the usage scenario of the information processing method provided by the embodiments of the present invention. Refer to Figure 1, different function-executable corresponding clients are set on the terminals (including terminal 10-1 and terminal 10-2). Among them, the corresponding clients are that the terminals (including terminal 10-1 and terminal 10-2) obtain different articles from the corresponding servers 200 through the network 300 for browsing, or obtain the applets or official accounts saved in the servers. When the terminals run the WeChat process, different contents such as moments, applets, articles, official accounts, novels, music, and emojis can be searched according to keywords through the provided search function. The terminals are connected to the servers 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and uses a wireless link to implement data transmission. Among them, the types of articles obtained by the terminals (including terminal 10-1 and terminal 10-2) from the corresponding servers 200 through the network 300 are different. For example: the terminals (including terminal 10-1 and terminal 10-2) can obtain the applets or official accounts matching the retrieval instruction A from the corresponding servers 200 through the network 300, or can also obtain the articles only matching the retrieval instruction A from the corresponding servers 200 through the network 300 for browsing.

[0096] In some embodiments of the present invention, the different types of applets saved in the server 200 can be written in software code environments of different programming languages, and the code objects can be different types of code entities. For example, in the software code of the C language, a code object can be a function. In the software code of the JAVA language, a code object can be a class, and in the OC language of the IOS side, it can be a piece of object code. In the software code of the C++ language, a code object can be a class or a function to execute the search terms from different terminals. Among them, the source of the retrieval instruction is not distinguished in this application. Among them, the applet in the WeChat process can trigger the search engine. The applet (MiniProgram) is a program that is developed based on a front-end-oriented language (such as JavaScript) and realizes services in a hypertext markup language (HTML, Hyper Text Markup Language) page. It is software that is downloaded by a client (such as a browser or any client embedding a browser core) via a network (such as the Internet) and interpreted and executed in the browser environment of the client, saving the steps of installation in the client. For example, the applet in the terminal can be awakened by a voice instruction to realize that various services such as air ticket purchase, task processing and production, and data display can be downloaded and run in the social network client.

[0097] The server 200 sends corresponding search results to the terminals (terminal 10-1 and / or terminal 10-2) via the network 300 according to the search terms input by the terminals. Therefore, as an example, the server 200 is used to obtain the log information of the search engine, and determine the click sequence composed of the search terms and the corresponding search results based on the log information of the search engine; determine the corresponding set of search terms according to the different search results in different click sequences; map the different search terms in the set of search terms to the matching semantic space, and perform encoding processing on the different search terms in the set of search terms to form the corresponding search term vectors; determine the search terms deviating from the semantic center in the set of search terms based on the search term vectors; and determine the set composed of the corresponding search terms and search results according to the search terms deviating from the semantic center as the bad examples of the search engine.

[0098] The structure of the server in the embodiments of the present invention will be described in detail below. The server can be implemented in various forms, such as a dedicated terminal with information processing functions, or a server with information processing functions, such as the Figure 1 server 200 mentioned above. Figure 2 FIG. is a schematic diagram of the composition structure of the server provided by the embodiments of the present invention. It can be understood that Figure 2 only shows the exemplary structure of the server rather than all structures, and can implement Figure 2 the partial structure or all structures shown according to needs.

[0099] The server provided by the embodiments of the present invention includes: at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. Each component in the server 20 is coupled together through a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 205.

[0100] Among them, the user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, a button, a touchpad, or a touch screen, etc.

[0101] It can be understood that the memory 202 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The memory 202 in the embodiments of the present invention is capable of storing data to support the operation of a terminal (such as 10-1). Examples of such data include: any computer programs for operating on the terminal (such as 10-1), such as an operating system and application programs. Among them, the operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs.

[0102] In some embodiments, the information processing device provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the information processing device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the information processing method provided by the embodiments of the present invention. For example, a processor in the form of a hardware decoding processor can employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0103] As an example of the information processing device provided by the embodiments of the present invention being implemented in a combination of software and hardware, the information processing device provided by the embodiments of the present invention can be directly embodied as a combination of software modules executed by the processor 201. The software modules can be located in a storage medium, and the storage medium is located in the memory 202. The processor 201 reads the executable instructions included in the software modules in the memory 202 and, in combination with necessary hardware (for example, including the processor 201 and other components connected to the bus 205), completes the information processing method provided by the embodiments of the present invention.

[0104] As an example, the processor 201 can be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0105] As an example of the information processing device provided by the embodiments of the present invention implemented in hardware, the device provided by the embodiments of the present invention can be directly implemented by a processor 201 in the form of a hardware decoding processor. For example, it can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components to execute and implement the information processing method provided by the embodiments of the present invention.

[0106] The memory 202 in the embodiments of the present invention is used to store various types of data to support the operation of the server 20. Examples of such data include: any executable instructions for operating on the server 20, such as executable instructions. The program for implementing the information processing method of the embodiments of the present invention can be included in the executable instructions.

[0107] In some other embodiments, the information processing device provided by the embodiments of the present invention can be implemented in software. Figure 2 Shown is the information processing device 2020 stored in the memory 202, which can be software in the form of a program and a plug-in, etc., and includes a series of modules. As an example of the program stored in the memory 202, it can include the information processing device 2020. The information processing device 2020 includes the following software modules: an information transmission module 2081 and an information processing module 2082. When the software modules in the information processing device 2020 are read into the RAM by the processor 201 and executed, the information processing method provided by the embodiments of the present invention will be implemented. The functions of each software module in the information processing device 2020 are introduced below:

[0108] The information transmission module 2081 is used to obtain the log information of the search engine and determine the click sequence composed of the search terms and the corresponding search results based on the log information of the search engine.

[0109] The information processing module 2082 is used to determine the corresponding set of search terms according to different search results in different click sequences.

[0110] The information processing module 2082 is used to map different search terms in the set of search terms to a matching semantic space and perform encoding processing on different search terms in the set of search terms to form corresponding search term vectors.

[0111] The information processing module 2082 is configured to determine, based on the search term vector, search terms in the search term set that deviate from the semantic center;

[0112] The information processing module 2082 is configured to determine, according to the search terms that deviate from the semantic center, a set composed of the corresponding search terms and search results as a bad example of the search engine.

[0113] Continue to combine Figure 2 Illustrate the information processing method provided by the embodiment of the present invention with the server 20 shown. During the execution of the search term, the prior art usually uses the badcase mining technology to achieve data correction, that is, to mine the cases where the doc returned by the query input by the user in the search does not match the query intention. Specifically, the query includes the search terms input by the user in the search system, usually short texts, and the doc refers to the results returned by the search system, which can be an account, such as a public account and a small program, or article information. In the prior art, references 3 and Figure 4 , Figure 3 is a schematic diagram of badcase mining in the related art in the embodiment of the present invention, Figure 4 is a schematic diagram of the influence of bad cases on search results in the embodiment of the present invention. Among them, the current mainstream method for mining search badcases is to analyze user retrieval log information, construct user behavior sequences, calculate features such as click-through rate, average click-through rate, page-turning rate, stay duration, and switching query, and combine statistical models to mine badcases. However, in this process, its defect is that when the user searches for a certain query, the target doc appears first, but because of curiosity, the user looks at other docs, but this is not the user's true viewing intention. Therefore, too high a confidence level is given to the user's click, without considering the situation where the user's click does not match the query intention. Second, some clicks of the user may be caused by account drainage. When the titles of two docs are highly similar, the doc ranked first may bring a lot of clicks due to drainage reasons, rather than the user's real selection experience; Third, only statistical information is carried out during the mining process, and the similarity of the semantic space is not considered, which affects the recommendation accuracy of the search system. Reference Figure 3 In the search home page, if the displayed results do not meet the user's expectations, it will greatly reduce the user experience and even lose users. For example: When a user of WeChat searches for the query "Bai Yansong", the first doc in the results is Zhengfan Reading. The user does not obtain a doc with a high matching degree at the most critical position, which reduces the user experience and even generates a negative impression on the corresponding search engine, thinking that the search engine technology is not good, affecting the user's use experience.

[0114] To overcome the above defects, see Figure 5 , Figure 5An optional flowchart of the information processing method provided by the embodiments of the present invention. Understandably, Figure 5 The steps shown can be executed by various servers running the information processing device. For example, it can be a dedicated terminal, server, or server cluster with a retrieval instruction processing function. The following will describe Figure 5 the steps shown.

[0115] Step 501: Obtain the log information of the search engine, and determine the click sequence formed by the search terms and the corresponding search results based on the log information of the search engine.

[0116] In some embodiments of the present invention, obtaining the log information of the search engine and determining the click sequence formed by the search terms and the corresponding search results based on the log information of the search engine can be achieved in the following manner:

[0117] According to the search term parameter information carried in the log information of the search engine, trigger the corresponding word segmentation library; perform word segmentation processing on the log information of the search engine through the word dictionary of the triggered word segmentation library to form different word-level search terms and sentence-level search terms; determine the search results corresponding to the different word-level search terms and sentence-level search terms respectively; perform noise reduction processing on the different word-level search terms and sentence-level search terms to form a click sequence corresponding to the log information of the search engine, where the click sequence includes a word-level search term and the corresponding search result, or a sentence-level search term and the corresponding search result. Among them, in combination with the description of the previous embodiments, different terminal devices (such as the terminal 10-1 and / or terminal 10-2 shown in the previous Figure 1 can provide a search bar for inputting keywords to be searched and a search button for searching data for the keywords to be searched on their respective corresponding search interfaces (such as web pages, information search APPs, and search mini-programs of WeChat). When the user enters a keyword in the search bar and the terminal device detects a click operation on the search button, it triggers the server to start the corresponding word segmentation instruction, and the word segmentation instruction carries the keyword in the search bar, and the server receives the word segmentation instruction. Or, the terminal device displays popular search keywords on the search interface. When a click operation on a popular search keyword is detected, the terminal device sends the word segmentation instruction to the server, and the word segmentation instruction carries the popular search keyword, and the server receives the word segmentation instruction. It should be noted that the present invention does not limit the triggering method of the word segmentation instruction. Among them, referring to Figure 6A , Figure 6ASchematic diagram of click sequence composed of search words and corresponding search results in an embodiment of the present invention, wherein the user behavior data is obtained, and the query-doc click sequence can be constructed by obtaining the user log stored in the server and associating the query-docpair by obtaining the user click behavior. Figure 6A As shown in the figure, the same doc may be clicked by users with different search queries, thereby associating query and doc. Query refers to the search term entered by the user in the search system, which is usually a relatively short text; doc refers to the result returned by the search system, which can be an account, such as a public account and mini program, or an article.

[0118] In some embodiments of the present invention, the search term text corresponding to the search term may be described in natural language, and there is a gap between its expression and the query requirements of the search system. The basis for the search system to retrieve the text content is to obtain documents including keywords through the inverted list, and the query requirements described in natural language cannot directly determine the keywords. Especially for Chinese, Chinese characters are the basic semantic units, and the smallest semantic unit with real meaning is the word; because there is no space between words like English words as a separator, it is not certain which characters constitute a word in a sentence, so it is an important task to segment Chinese text. In addition, for the search term text, it contains some things that are only valuable for natural language understanding, and for the search system, in order to query relevant content, it must be determined which are truly valuable retrieval bases. Therefore, by performing denoising on different word-level feature vectors in the previous embodiment, a word-level feature vector set corresponding to the search term text can be formed to avoid meaningless word-level feature vectors in the word-level feature vector set, such as "的", "地" and "得".

[0119] In some embodiments of the present invention, it is also possible to determine the name of the word segmentation library used when performing word segmentation on the search term text; according to the name of the word segmentation library, determine the parameters of the word segmentation library that match the word-level feature vector corresponding to the search term text, where the parameters of the word segmentation library include: the type of the word segmentation library, the name of the word segmentation library, and the version of the word segmentation library. Among them, the parameters of the word segmentation library include: the type of the word segmentation library, the name of the word segmentation library, and the version of the word segmentation library. Among them, since the word-level feature vectors formed when processing the same text information using different word segmentation libraries are not exactly the same, therefore, according to the name of the word segmentation library, determine the parameters of the word segmentation library that match the word-level feature vector corresponding to the search instruction text, so as to realize determining the parameters of the word segmentation library used for word segmentation of the search instruction text. For example: when the search instruction text is "The Story of Time by Luo Dayou mp3", after processing with word segmentation library A, a set of word-level feature vectors A (The Story of Time; Luo Dayou's mp3) corresponding to the search instruction text is formed; after processing with word segmentation library B, a set of word-level feature vectors B (The Story of Time; Luo Dayou; mp3) corresponding to the search instruction text is formed; after processing with word segmentation library A1, a set of word-level feature vectors A1 (Time; Story; Luo Dayou; mp3) corresponding to the search instruction text is formed.

[0120] Step 502: Determine the corresponding set of search terms according to different search results in different click sequences.

[0121] In some embodiments of the present invention, determining the corresponding set of search terms according to different search results in different click sequences can be achieved through the following methods:

[0122] Based on the sorting of the click sequences, determine the search results in the click sequences; according to the search results, traverse the log information of the search engine, and based on the behavior records of different target users, determine the search terms that match the search results; combine the search terms corresponding to the same search result respectively to determine the corresponding set of search terms. Refer to Figure 6B , Figure 6B is a schematic diagram of the composition of click sequences in the embodiments of the present invention. For doc1, it has been clicked by query1, query2... etc., then query1, query2... etc. queries are constructed into a query list. Thus, it can ensure the comprehensiveness of the search terms corresponding to the same search result, and avoid mining wrong bad cases that affect the accuracy of the search engine.

[0123] Step 503: Map different search terms in the set of search terms to a matching semantic space, and perform encoding processing on different search terms in the set of search terms to form corresponding search term vectors.

[0124] Continue to refer to Figure 7 , Figure 7 which is an optional flowchart of the information processing method provided by the embodiments of the present invention. It can be understood that Figure 7 the steps shown can be executed by various servers running the information processing device. For example, it can be a dedicated terminal, server, or server cluster with a retrieval instruction processing function. The following describes the steps Figure 7 shown.

[0125] Step 701: Determine a semantic space that matches the search term through any search term in the search term set.

[0126] Step 702: Determine at least one word-level vector corresponding to the target search term through the encoder of the information processing model.

[0127] Step 703: Map at least one word-level vector corresponding to the target search term to the semantic space, and perform iterative processing on the search term set until all search terms in the search term set are mapped to the matching semantic space, forming a search term vector that matches the search term set.

[0128] Among them, the information processing model can be a Bidirectional Encoder Representations from Transformers (BERT). Before using it, corresponding training is also required. Specifically, the training process includes:

[0129] Obtain a first training sample set, where the first training sample set includes at least one set of sentence samples with noise; perform denoising processing on the first training sample set to form a corresponding second training sample set; process the first training sample set through the information processing model to determine the initial parameters of the information processing model; in response to the initial parameters of the information processing model, process the second training sample set through the information processing model to determine the updated parameters of the information processing model; according to the updated parameters of the information processing model, iteratively update the encoder parameters and decoder parameters of the information processing model through the first training sample set and the second training sample set to enable the information processing model to encode corresponding search terms. When the information processing model is a bidirectional attention neural network, refer to Figure 8 and continue to refer to Figure 8 , Figure 8This is an optional schematic diagram of the information processing process of the information processing model in an embodiment of the present invention. Among them, both the encoder and decoder parts contain 6 encoders and decoders. The inputs entering the first encoder are combined with embedding and positional embedding. After passing through 6 encoders, the output is sent to each decoder in the decoder part; when the trained information processing model is deployed, only the encoder network needs to be used to determine at least one word-level vector corresponding to the target search term through the encoder of the information processing model.

[0130] Continue to refer to Figure 9 , Figure 9 This is an optional structural schematic diagram of the encoder in the information processing model in an embodiment of the present invention. Among them, its input consists of a query (Q) and a key (K) with a dimension of d and a value (V) with a dimension of d. All keys calculate the dot product of the query and apply the softmax function to obtain the weights of the values.

[0131] Continue to refer to Figure 9 , Figure 9 This is a vector schematic diagram of the encoder in the information processing model in an embodiment of the present invention. Among them, Q, K, and V are obtained by multiplying the vector x input to the encoder with W^Q, W^K, and W^V. The dimensions of W^Q, W^K, and W^V in the article are (512, 64), and then assume that the dimension of our inputs is (m, 512), where m represents the number of words. Therefore, the dimensions of Q, K, and V obtained after multiplying the input vector with W^Q, W^K, and W^V are (m, 64).

[0132] Continue to refer to Figure 10 , Figure 10 This is a schematic diagram of vector concatenation of the encoder in the information processing model in an embodiment of the present invention. Among them, Z0 to Z7 are the corresponding 8 parallel heads (with a dimension of (m, 64)), and then after concatenating these 8 heads, a matrix with a dimension of (m, 512) is obtained. Finally, after multiplying with W^O, an output matrix with a dimension of (m, 512) is obtained, and the dimension of this matrix is consistent with the dimension entering the next encoder.

[0133] Continue to refer to Figure 11 , Figure 11This is a schematic diagram of the encoding process of the encoder in the information processing model in the embodiment of the present invention, in which x1 is self-attentioned to the state of z1. The tensor that has passed the self-attention needs to be processed by the residual network and the Later Norm, and then enters the fully connected feedforward network. The feedforward network needs to perform the same operation, residual processing and normalization. Finally, the output tensor can enter the next encoder, and then this operation is iterated 6 times, and the result of the iterative processing enters the decoder.

[0134] In some embodiments of the present invention, denoising the first training sample set to form a corresponding second training sample set may be implemented in the following manner:

[0135] Determine a dynamic noise threshold that matches the use environment of the information processing model; perform denoising on the first training sample set according to the dynamic noise threshold to form a second training sample set that matches the dynamic noise threshold. The dynamic noise threshold that matches the use environment of the information processing model is different due to different use environments of the information processing model. For example, in the use environment of the search engine triggered by the WeChat applet, the dynamic noise threshold that matches the use environment of the information processing model needs to be smaller than the dynamic noise threshold of the search engine in the short video client.

[0136] In some embodiments of the present invention, a fixed noise threshold corresponding to the information processing model is further determined, and the first training sample set is denoised according to the fixed noise threshold to form a second training sample set matching the fixed noise threshold. Among them, when the information processing model is solidified in the corresponding hardware mechanism (such as a WeChat payment terminal), and the use environment is a WeChat applet triggering a search engine to realize a public account search, by fixing the fixed noise threshold corresponding to the information processing model, the training speed of the information processing model can be effectively improved, and the waiting time of the user can be reduced.

[0137] In some embodiments of the present invention, the first training sample set may be subjected to negative example processing to form a negative example sample set corresponding to the first training sample set, wherein the negative example sample set includes bad example sample data in the user behavior data, which is used to adjust the encoder parameters and decoder parameters of the information processing model. The negative example processing may recombine and adjust the corresponding order of the search terms and the search results, and the robustness of the information processing model may be effectively improved through the negative example sample set.

[0138] In some embodiments of the present invention, in response to the initial parameters of the information processing model, the second training sample set is processed by the information processing model to determine the update parameters of the information processing model, including:

[0139] Substitute different statement samples in the second training sample set into the loss function corresponding to the auto-encoding network composed of the encoder and decoder of the information processing model; when it is determined that the loss function satisfies the convergence condition, the parameters of the encoder and the corresponding decoder parameters in the information processing model are used as the update parameters of the information processing model. Thus, when the information processing model is trained, it can be deployed in the corresponding search engine server to realize the mining of bad cases in the search engine usage environment.

[0140] Step 504: Based on the search term vector, determine the search terms in the search term set that deviate from the semantic center.

[0141] In some embodiments of the present invention, based on the search term vector, determining the search terms in the search term set that deviate from the semantic center can be achieved by the following method:

[0142] Based on the search term vector, determine the semantic center corresponding to the search term set; determine the distance from each search term in the search term set to the semantic center; when the distance from the search term to the semantic center is greater than the corresponding distance threshold, determine that the current search term deviates from the semantic center; perform iterative processing on different search terms in the search term set until all search terms in the search term set that deviate from the semantic center are determined. For the obtained query embedding features, the semantic center of the query list can be calculated, and then the distance from each query in the query list to the semantic center can be calculated respectively. Set a threshold, and for queries exceeding the threshold, take query-doc as bad cases. Specifically, the cosine distance is used as the distance metric, and the calculation formula is:

[0143]

[0144] Thus, the information processing method provided by the present application backtracks and clusters from the doc side, and determines the queries that deviate from the query semantic space center under the same doc from the query-doc inverted index. By analyzing the user's click behavior data, it is determined that the queries used by users who click on the same doc are strongly correlated.

[0145] Step 505: According to the search terms that deviate from the semantic center, determine the set composed of the corresponding search terms and search results as the bad cases of the search engine.

[0146] In some embodiments of the present invention, when different search terms in the search term set do not exceed the distance threshold, it is determined that the bad case set of the search engine is an empty set.

[0147] Thus, by mining the bad cases that deviate from the center of the query semantic space in the query list that clicks on the same doc, the problem that the user's click does not match the query intention input by the user in the prior art is overcome, the bad cases of the search engine are fully mined. At the same time, the query is mapped to the semantic space, and the semantic information of the query is fully applied to ensure the user experience of using the search system and improve user stickiness.

[0148] Reference Figure 12 , Figure 12 is a schematic diagram of the information processing effect in the embodiment of the present invention. Among them, a user interface is displayed. The user interface includes a perspective screen of using the search engine in the corresponding software process from the first-person perspective. The user interface also includes a display control component; through the display control component, the search results matching the search terms input by the user are controlled to be displayed. For example: the search results provided by the user by inputting the search term "Bai Yansong" in the WeChat process are high-matching search result docs related to "Bai Yansong", and other search results as bad cases do not appear at the top of the search results.

[0149] Continue to refer to Figure 13 , Figure 13 is an optional flowchart of the information processing method provided by the embodiment of the present invention. Among them, Figure 13 the steps shown can be executed by a video server or a server cluster running the information processing device. The following will be described for Figure 13 the steps shown.

[0150] Step 1301: Obtain the user's behavior data and construct a query-doc click sequence.

[0151] Among them, refer to Figure 14 , among which, Figure 14 is a schematic diagram of the information processing effect in the embodiment of the present invention. As Figure 14 shown, the short video playing interface can be presented in the corresponding short video APP, or can be triggered through a WeChat mini-program (the information processing model can be encapsulated in the corresponding APP after training or saved in the form of a plug-in in the WeChat mini-program). As the short video application products continue to develop and increase, the carrying capacity of video information is much larger than that of text information. The short video can recommend relevant videos to the user in response to the search terms input by the user through the corresponding application program. Effective subsequent relevant video recommendations can effectively improve the user experience.

[0152] Step 1302: Based on the same doc as one category, construct a matching query set.

[0153] Step 1303: Through the information processing network model, map the query to the semantic space.

[0154] Step 1304: Calculate the queries in the same category that deviate from the semantic center.

[0155] Thus, it is possible to determine the corresponding bad cases by clustering the query-doc associations. Among them, when the user inputs the search term "Luo Dayou" during the short video process, the search results provided are the videos with high matching degrees related to "Luo Dayou", such as "Luogang Xiaozhen - Luo Dayou". Other search results as bad cases do not appear in the search result display interface, facilitating the user to select videos related to the search term "Luo Dayou" for viewing.

[0156] The present invention has the following beneficial technical effects:

[0157] The present invention obtains the log information of the search engine, and determines the click sequence composed of the search term and the corresponding search results based on the log information of the search engine; determines the corresponding search term set according to the different search results in different click sequences; maps the different search terms in the search term set to the matching semantic space, and encodes the different search terms in the search term set to form corresponding search term vectors; determines the search terms in the search term set that deviate from the semantic center based on the search term vectors; determines the set composed of the corresponding search terms and search results as the bad cases of the search engine according to the search terms that deviate from the semantic center. This application can fully mine the bad cases of the search engine, obtain more accurate and comprehensive bad cases, effectively improve the quality of search results, and improve the user experience.

[0158] The above is only an embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An information processing method, characterized in that, The method includes: Obtaining the log information of a search engine, and determining a click sequence composed of search terms and corresponding search results based on the log information of the search engine; Determining corresponding search term sets according to different search results in different click sequences; Mapping different search terms in the search term set to a matching semantic space, and performing encoding processing on different search terms in the search term set to form corresponding search term vectors; Determining the search terms deviating from the semantic center in the search term set based on the search term vectors; Determining a set composed of corresponding search terms and search results as bad examples of the search engine according to the search terms deviating from the semantic center; 2. The method according to claim 1, wherein The obtaining the log information of a search engine, and determining a click sequence composed of search terms and corresponding search results based on the log information of the search engine includes: Triggering a corresponding word segmentation library according to the search term parameter information carried in the log information of the search engine; Performing word segmentation processing on the log information of the search engine through the word dictionary of the triggered word segmentation library to form different word-level search terms and sentence-level search terms; Determining search results corresponding to the different word-level search terms and sentence-level search terms respectively; Performing denoising processing on the different word-level search terms and sentence-level search terms to form a click sequence corresponding to the log information of the search engine, where the click sequence includes a word-level search term and a corresponding search result, or a sentence-level search term and a corresponding search result.

3. The method according to claim 2, wherein The method further includes: Determining the name of the word segmentation library used when performing word segmentation processing on the search term text; Determining parameters of a word segmentation library that match the word-level feature vector corresponding to the search term text according to the name of the word segmentation library, where the parameters of the word segmentation library include: The type of the word segmentation library, the name of the word segmentation library, and the version of the word segmentation library.

4. The method according to claim 1, wherein The determining corresponding search term sets according to different search results in different click sequences includes: Determining search results in the click sequence based on the sorting of the click sequence; Traversing the log information of the search engine according to the search results, and determining search terms that match the search results based on the behavior records of different target users; Combining search terms corresponding to the same search result respectively to determine corresponding search term sets.

5. The method according to claim 1, wherein The mapping different search terms in the search term set to a matching semantic space includes: Determining a semantic space that matches the search term through any search term in the search term set; Determining at least one word-level vector corresponding to a target search term through an encoder of an information processing model; Mapping at least one word-level vector corresponding to the target search term to the semantic space, and performing iterative processing on the search term set until all search terms in the search term set are mapped to a matching semantic space to form search term vectors that match the search term set.

6. The method according to claim 5, characterized in that The method further includes: Obtain a first training sample set, where the first training sample set includes at least one set of statement samples with noise; Perform denoising processing on the first training sample set to form a corresponding second training sample set; Process the first training sample set through an information processing model to determine initial parameters of the information processing model; In response to the initial parameters of the information processing model, process the second training sample set through the information processing model to determine updated parameters of the information processing model; According to the updated parameters of the information processing model, iteratively update the encoder parameters and decoder parameters of the information processing model through the first training sample set and the second training sample set, so as to enable the information processing model to encode corresponding search terms.

7. The method according to claim 6, wherein The performing denoising processing on the first training sample set to form a corresponding second training sample set includes: Determine a dynamic noise threshold matching the usage environment of the information processing model; Perform denoising processing on the first training sample set according to the dynamic noise threshold to form a second training sample set matching the dynamic noise threshold; or, Determine a fixed noise threshold corresponding to the information processing model, and perform denoising processing on the first training sample set according to the fixed noise threshold to form a second training sample set matching the fixed noise threshold.

8. The method according to claim 6, wherein The method further includes: Perform negative example processing on the first training sample set to form a negative example sample set corresponding to the first training sample set, where the negative example sample set includes bad example sample data in user behavior data and is used to adjust the encoder parameters and decoder parameters of the information processing model.

9. The method according to claim 6, wherein The responding to the initial parameters of the information processing model, processing the second training sample set through the information processing model to determine the updated parameters of the information processing model includes: Substitute different statement samples in the second training sample set into a loss function corresponding to an autoencoder network composed of an encoder and a decoder of the information processing model; Determine the parameters of the encoder and the corresponding decoder parameters in the information processing model when the loss function satisfies the convergence condition as the updated parameters of the information processing model.

10. The method according to claim 1, wherein The determining the search terms deviating from the semantic center in the search term set based on the search term vector includes: Based on the search term vector, determine the semantic center corresponding to the search term set; Determine the distance from each search term in the search term set to the semantic center; When the distance from the search term to the semantic center is greater than the corresponding distance threshold, determine that the current search term deviates from the semantic center; Perform iterative processing on different search terms in the search term set until all search terms deviating from the semantic center in the search term set are determined.

11. The method according to claim 1, wherein The method further includes: When different search terms in the search term set do not exceed the distance threshold, determine that the bad example set of the search engine is an empty set.

12. An information processing apparatus, characterized in that, The device includes: An information transmission module, configured to obtain log information of a search engine, and determine a click sequence formed by a search term and a corresponding search result based on the log information of the search engine; An information processing module, configured to determine a corresponding set of search terms according to different search results in different click sequences; The information processing module is configured to map different search terms in the set of search terms to a matching semantic space, perform encoding processing on different search terms in the set of search terms, and form corresponding search term vectors; The information processing module is configured to determine, based on the search term vectors, search terms in the set of search terms that deviate from the semantic center; The information processing module is configured to determine, according to the search terms that deviate from the semantic center, a corresponding set formed by a search term and a search result as a bad example of the search engine.

13. The apparatus according to claim 12, wherein The information processing module is configured to trigger a corresponding word segmentation library according to search term parameter information carried in the log information of the search engine; The information processing module is configured to perform word segmentation processing on the log information of the search engine through a word dictionary of the triggered word segmentation library to form different word-level search terms and sentence-level search terms; The information processing module is configured to determine search results respectively corresponding to the different word-level search terms and sentence-level search terms; The information processing module is configured to perform denoising processing on the different word-level search terms and sentence-level search terms to form a click sequence corresponding to the log information of the search engine, where the click sequence includes a word-level search term and a corresponding search result, or a sentence-level search term and a corresponding search result.

14. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable instructions; A processor, configured to implement the information processing method according to any one of claims 1 to 11 when running the executable instructions stored in the memory.

15. A computer-readable storage medium storing executable instructions, characterized in that, The executable instructions, when executed by the processor, implement the information processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Recommendation method and device for keywords

    CN103136224A

  • Method and device for optimizing text search

    CN108932247A