Search word quality detection method and related device

By performing keyword analysis and semantic feature extraction for search terms, combining multi-dimensional quality evaluation, and using pre-trained models to determine search terms quality, the multi-dimensional coupling problem is solved, and the accuracy and reliability of search terms detection are improved.

CN120386906APending Publication Date: 2025-07-29BEIJING SOGOU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410122036.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, there is a coupling problem in the quality detection of multi-dimensional search terms, resulting in poor accuracy of detection results and affecting the user's search experience.

Method used

Through keyword analysis and semantic feature extraction based on search results, combined with multi-dimensional quality feature evaluation, and quality judgment is used to use pre-trained models to achieve multi-dimensional quality recognition of search terms.

Benefits of technology

It improves the accuracy and reliability of search term quality detection, enhances the richness of semantic features, and ensures the accuracy and reliability of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386906A_ABST
    Figure CN120386906A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a search word quality detection method and related device.The method comprises the steps that keyword analysis is conducted on all search results obtained on the basis of a to-be-detected search word, semantic feature extraction is conducted on the to-be-detected search word in combination with all obtained result keywords, and target semantic features are obtained; performing quality evaluation of at least one evaluation dimension on the to-be-detected search word, and performing quality feature extraction on the to-be-detected search word by using a quality evaluation value to obtain multi-dimensional quality features; and performing quality identification on the to-be-detected search word based on the multi-dimensional quality features and the target semantic features to obtain a quality inspection result. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. Quality identification is assisted through the search result and the quality evaluation value, and the accuracy and reliability of search word quality detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0002] With the development of computer technology and Internet technology, various platforms have gradually become indispensable information search channels. For example, various types of information can be searched using a general search system, product information can be searched using an e-commerce platform, videos can be searched using a video platform, etc. To identify illegal information and improve the search experience, it is usually necessary to perform quality detection on the search term (query).

[0003] In the related art, the quality of the search term is evaluated separately on multiple evaluation dimensions, and a corresponding quality evaluation value can be obtained for each evaluation dimension. Then, according to the score threshold set for each evaluation dimension respectively, it is determined whether there is a quality problem with the search term on that evaluation dimension.

[0004] However, due to the mutual coupling of the quality problems of multiple evaluation dimensions, there is a situation where the detection results indicate that there is no quality problem in each evaluation dimension, but when multiple evaluation dimensions are combined, the search term actually has a quality problem, resulting in poor accuracy of the search term, and further resulting in low-quality search terms on the platform, affecting the user search experience. SUMMARY OF THE INVENTION

[0005] Embodiments of the present application provide a search term quality detection method and related device to improve the accuracy and reliability of search term quality detection.

[0006] In a first aspect, an embodiment of the present application provides a search term quality detection method, including:

[0007] For each search result obtained based on the search term to be detected, based on the word frequencies of the words included in each search result, keyword parsing is performed on each search result to obtain each result keyword;

[0008] According to at least one evaluation dimension, the quality of the search term to be detected is evaluated separately to obtain the quality evaluation value of the corresponding evaluation dimension;

[0009] Based on each result keyword, semantic feature extraction is performed on the search term to be detected to obtain the target semantic feature of the search term to be detected;

[0010] Based on the at least one obtained quality evaluation value, quality feature extraction is performed on the search term to be detected to obtain the multi-dimensional quality feature corresponding to the search term to be detected;

[0011] Based on the target semantic feature and the multi-dimensional quality feature, quality identification is performed on the search term to be detected to obtain the quality inspection result of the search term to be detected.

[0012] In a second aspect, an embodiment of the present application provides a search term quality detection device, including:

[0013] A result keyword parsing unit, configured to perform keyword parsing on each search result obtained based on a search term to be detected, and obtain each result keyword based on the word frequencies of the words included in each search result;

[0014] A quality evaluation unit, configured to perform quality evaluation on the search term to be detected respectively according to at least one evaluation dimension, and obtain a quality evaluation value for the corresponding evaluation dimension;

[0015] A semantic feature extraction unit, configured to perform semantic feature extraction on the search term to be detected based on each result keyword, and obtain the target semantic feature of the search term to be detected;

[0016] A quality feature extraction unit, configured to perform quality feature extraction on the search term to be detected based on at least one obtained quality evaluation value, and obtain a multi-dimensional quality feature corresponding to the search term to be detected;

[0017] A quality recognition unit, configured to perform quality recognition on the search term to be detected based on the target semantic feature and the multi-dimensional quality feature, and obtain a quality inspection result of the search term to be detected.

[0018] As a possible implementation manner, when performing semantic feature extraction on the search term to be detected based on each result keyword to obtain the semantic feature of the search term to be detected, the semantic feature extraction unit is specifically configured to:

[0019] Concatenate each result keyword and the search term to be detected according to a set concatenation manner to obtain semantic input data;

[0020] Based on the semantic input data, use the semantic feature extraction network of the trained quality discrimination model to obtain the semantic feature of the search term to be detected.

[0021] As a possible implementation manner, the semantic feature extraction network includes a first feature extraction sub-network, and the first feature extraction sub-network is configured to extract the semantic feature of the search term to be detected;

[0022] When using the semantic feature extraction network of the trained quality discrimination model to obtain the target semantic feature of the search term to be detected based on the semantic input data, the semantic feature extraction unit is specifically configured to:

[0023] Perform word segmentation on the semantic input data according to a set word segmentation manner to obtain each segmented word;

[0024] Use the first feature extraction sub-network to perform feature extraction on each segmented word to obtain the segmented word feature of each segmented word;

[0025] Based on the obtained word segmentation features, obtain the target semantic features of the search term to be detected.

[0026] As a possible implementation, the semantic feature extraction network further includes a second feature extraction sub-network, and the second feature extraction sub-network is used to extract the target semantic features including the context information of each word segmentation;

[0027] When obtaining the target semantic features of the search term to be detected based on the obtained word segmentation features, the semantic feature extraction unit specifically is used for:

[0028] Based on the word segmentation features, use the second feature extraction sub-network to perform enhancement processing on the word segmentation features, and obtain the target semantic features including the context information.

[0029] As a possible implementation, when extracting the quality features of the search term to be detected based on the obtained at least one quality evaluation value to obtain the multi-dimensional quality features corresponding to the search term to be detected, the quality feature extraction unit specifically is used for:

[0030] Based on the obtained at least one quality evaluation value, construct quality input data;

[0031] Based on the quality input data, use the quality feature extraction network in the trained quality discrimination model to extract the quality features of the search term to be detected, and obtain the multi-dimensional quality features.

[0032] As a possible implementation, the quality feature extraction network includes multiple fully connected layers, and the number of nodes included in the last fully connected layer in the multiple fully connected layers exceeds the number of feature dimensions of the quality input data.

[0033] As a possible implementation, when constructing the quality input data based on the obtained at least one quality evaluation value, the quality feature extraction unit specifically is used for:

[0034] Directly splice the at least one quality evaluation value to construct the quality input data; or,

[0035] Based on the historical search records of the search term to be detected, obtain at least one search feature, and splice the at least one search feature and the at least one quality evaluation value to construct the quality input data.

[0036] As a possible implementation, when performing quality identification on the search term to be detected based on the target semantic features and the multi-dimensional quality features to obtain the quality inspection result of the search term to be detected, the quality identification unit specifically is used for:

[0037] Based on the target semantic features and the multi-dimensional quality features, use the classification network in the trained quality discrimination model to obtain the target detection value of the search term to be detected; the target detection value is used to characterize the probability that the search term to be detected belongs to an abnormal search term.

[0038] Based on the target detection value, in combination with a set detection threshold, obtain the quality inspection result of the search term to be detected.

[0039] As a possible implementation, the quality recognition unit is further configured to:

[0040] Obtain a training sample set, where the training sample set contains each training sample, and each training sample contains a sample search term.

[0041] Based on the training sample set, perform iterative training on the pre-trained quality discrimination model to obtain the trained quality discrimination model. Among them, in each iteration process, perform the following operations:

[0042] Input a batch of several training samples into the pre-trained quality discrimination model to obtain corresponding prediction results.

[0043] Based on the obtained several prediction results, in combination with the respective true results corresponding to the several training samples, obtain the model loss, and adjust the model parameters based on the model loss.

[0044] As a possible implementation, the quality recognition unit is further configured to:

[0045] If the detection result indicates that the search term to be detected is an abnormal search term, then filter the search term to be detected from the reference search term set.

[0046] When receiving a search term association indication carrying a target search term, based on the semantic similarity between each reference search term in the filtered reference search term set and the target search term, select at least one reference search term that meets the set similarity condition from the respective reference search terms.

[0047] Based on the at least one reference search term, perform search term recommendation for the target search term.

[0048] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above method.

[0049] Fourthly, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method in any of the above aspects.

[0050] Fifthly, an embodiment of the present application provides a computer program product. The program product includes a computer program, and the computer program is stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of the method in any of the above aspects.

[0051] In an embodiment of the present application, on the one hand, keyword parsing is performed on each search result obtained based on a search term to be detected, and semantic feature extraction is performed on the search term to be detected in combination with the obtained result keywords. Thus, the semantic information of the search term to be detected is enriched by the search results, thereby enhancing the semantic richness of the target semantic features, and further improving the accuracy and reliability of search term quality detection. On the other hand, quality evaluation is performed on the search term to be detected in at least one evaluation dimension, and quality feature extraction is performed on the search term to be detected by using the quality evaluation value. Thus, the search term information is expanded by the quality evaluation values of each evaluation dimension to assist in search term quality detection. And since the extracted multi-dimensional quality features can obtain information related to search term quality to the greatest extent, the accuracy and reliability of search term quality detection are improved.

[0052] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings

[0053] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0054] Figure 1 It is a schematic diagram of an application scenario provided in an embodiment of the present application;

[0055] Figure 2 It is a schematic diagram of the structure of a quality discrimination model provided in an embodiment of the present application;

[0056] Figure 3 It is a schematic diagram of the flowchart of a model training method provided in an embodiment of the present application;

[0057] Figure 4A schematic diagram of a search result provided in an embodiment of the present application;

[0058] Figure 5 A schematic diagram of a process for obtaining a quality evaluation value provided in an embodiment of the present application;

[0059] Figure 6 A flowchart of a process for detecting the quality of a search term provided in an embodiment of the present application;

[0060] Figure 7 A schematic diagram of the structure of a semantic feature extraction network provided in an embodiment of the present application;

[0061] Figure 8 A schematic diagram of the structure of another semantic feature extraction network provided in an embodiment of the present application;

[0062] Figure 9 A schematic diagram of the structure of a quality feature extraction network provided in an embodiment of the present application;

[0063] Fig.10 A flowchart of a method for detecting the quality of a search term provided in an embodiment of the present application;

[0064] Fig.11A A logical diagram of a search term filtering process provided in an embodiment of the present application;

[0065] Fig. 11B A schematic diagram of a search page provided in an embodiment of the present application;

[0066] Fig.12 A logical diagram of a process for detecting the quality of a search term provided in an embodiment of the present application;

[0067] Figure 13 A schematic diagram of the structure of a device for detecting the quality of a search term provided in an embodiment of the present application;

[0068] Fig.14 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without making creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0070] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0071] It can be understood that when the embodiments of the present application are applied to specific products or technologies, relevant permissions or consents need to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0072] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application are explained here first:

[0073] Bidirectional Encoder Representation from Transformers (BERT): A pre-trained natural language processing model.

[0074] Bidirectional Gated Recurrent Unit (Bi-GRU): Bi-GRU is used to process sequence data. It combines forward and backward information and processes the sequence through two independent GRU networks simultaneously. In this way, it can capture context information more comprehensively and improve the performance of the model. Bi-GRU has a gating mechanism that can control the flow of information, including update gates, reset gates, and candidate hidden states. It is widely used in natural language processing tasks. By processing the sequence bidirectionally, Bi-GRU can better understand the semantics and context of sentences and improve the accuracy of prediction and classification.

[0075] Wide&Deep structure: The application of the Wide&Deep structure in text classification tasks is to combine a linear model and a deep neural network. The linear model is used to process broad features, such as word frequency, part of speech, etc., while the deep neural network is used to learn deep features, such as word embeddings, semantic relationships, etc. By combining the two, the Wide&Deep structure can make full use of broad and deep features and improve the accuracy and performance of text classification.

[0076] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0077] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0078] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, evolved from the large language model in the field of NLP. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0079] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0080] A pre-training model (PTM), also known as a foundation model or a large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a vast amount of unlabeled data, and by leveraging the function approximation ability of the large-parameter DNN, the PTM extracts common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (such as ELMO, BERT, GPT), vision models (such as swin-transformer, ViT, V-MOE), speech models (such as VALL-E), and multi-modal models (such as ViBERT, CLIP, Flamingo, Gato) according to the data modalities they process. Among them, multi-modal models refer to models that establish feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence-generated content (AIGC) and can also serve as a general interface connecting multiple specific task models.

[0081] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content, conversational interactions, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0082] The solution provided in the embodiments of this application relates to the machine learning technology of artificial intelligence, mainly involving the application of pre-training models in the search term quality detection scenario. In the search term quality detection scenario, a deep learning model (i.e., a quality discrimination model) can be used to detect the quality of the search term to be detected. The quality discrimination model is obtained by fine-tuning a pre-training model. During the fine-tuning process, the search terms and related information in the embodiments of this application (such as search results, quality evaluation values of each evaluation dimension, search features, etc.) are used as training samples to iteratively train the pre-training model, and then the trained quality discrimination model is output.

[0083] Specifically, the quality discrimination model mainly includes three parts: a semantic feature extraction network, a quality feature extraction network, and a classification network. Among them, the semantic feature extraction network is used to extract semantic features of the search term to be detected based on each result keyword associated with the search term to be detected, so as to obtain the target semantic features of the search term to be detected; the quality feature extraction network is used to extract quality features of the search term to be detected based on the quality evaluation values of the search term to be detected in at least one evaluation dimension, so as to obtain the multi-dimensional quality features corresponding to the search term to be detected; the classification network is used to perform quality identification on the search term to be detected based on the target semantic features and the multi-dimensional quality features, so as to obtain the quality inspection result of the search term to be detected. The specific model training and model application processes are described below. It should be noted that the quality discrimination model in the embodiments of the present application can be trained online or offline, and no specific limitation is made here. In this article, an example of offline training is used for illustration.

[0084] Next, a brief introduction to the technical idea of the embodiments of the present application is given.

[0085] In the related art, the quality of the search term is evaluated separately in multiple evaluation dimensions, and corresponding quality evaluation values can be obtained for each evaluation dimension. Then, according to the score thresholds set for each evaluation dimension respectively, it is judged whether there are quality problems with the search term in this evaluation dimension.

[0086] However, due to the mutual coupling of the quality problems in multiple evaluation dimensions, there is a situation where the detection results indicate that there are no quality problems in each evaluation dimension, but considering multiple evaluation dimensions comprehensively, the search term actually has quality problems, resulting in poor accuracy of the search term, and further leading to low-quality search terms on the platform, affecting the user search experience.

[0087] In the embodiments of the present application, on the one hand, keyword parsing is performed on each search result obtained based on the search term to be detected, and combined with the obtained result keywords, semantic features of the search term to be detected are extracted, thereby enriching the semantic information of the search term to be detected through the search results, enhancing the semantic richness of the target semantic features, and further improving the accuracy and reliability of the search term quality detection; on the other hand, the quality of the search term to be detected is evaluated in at least one evaluation dimension, and the quality evaluation value is used to extract quality features of the search term to be detected, thereby expanding the search term information through the quality evaluation values of each evaluation dimension to assist in the search term quality detection, and since the extracted multi-dimensional quality features can obtain the information related to the search term quality to the greatest extent, the accuracy and reliability of the search term quality detection are improved.

[0088] The following briefly introduces the application scenarios applicable to the technical solutions of the embodiments of this application. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of this application and are not limiting. In the specific implementation process, the technical solutions provided by the embodiments of this application can be flexibly applied according to actual needs.

[0089] The solution provided by the embodiments of this application can be applicable to various search term quality detection scenarios. For example, the search term quality detection in browsers, social platforms, and video platforms. This solution can be used as a basic technology in various scenarios, including but not limited to scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0090] Refer to Figure 1 As shown, it is a schematic diagram of an application scenario provided in the embodiments of this application. This application scenario includes a terminal device 110 and a server 120. The number of terminal devices 110 can be one or more. The number of servers 120 can also be one or more. This application does not make specific limitations on the number of terminal devices 110 and servers 120.

[0091] In the embodiments of this application, the terminal device 110 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an Internet of Things device, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto. A client with a search function is installed in the terminal device 110. The client can be in the form of an application program, a mini program, a web page, etc., but is not limited thereto.

[0092] The server 120 is the background server corresponding to the client. The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0093] The terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make restrictions here.

[0094] The search term quality detection method mentioned in the embodiments of this application can be executed by the server or the terminal device, or jointly executed by the server and the terminal device, and no restrictions are imposed on this.

[0095] In some embodiments, the search term quality detection is based on a trained quality discrimination model. Therefore, before introducing the actual search term quality detection process, the model training process will be described first.

[0096] Next, the structural design of the quality discrimination model will be introduced first.

[0097] Refer to Figure 2 As shown, it is a schematic structural diagram of a quality discrimination model provided in an embodiment of the present application. The quality discrimination model includes three parts: a semantic feature extraction network, a quality feature extraction network, and a classification network. Among them, the semantic feature extraction network is used to extract semantic features of the search term to be detected, and obtain the target semantic features of the search term to be detected; the quality feature extraction network is used to extract quality features of the search term to be detected, and obtain the multi-dimensional quality features corresponding to the search term to be detected; the classification network is used to perform quality identification on the search term to be detected based on the target semantic features and multi-dimensional quality features, and obtain the quality inspection result of the search term to be detected.

[0098] As a possible implementation manner, the input data of the semantic feature extraction network (which can also be called semantic input data) includes the search term to be detected (i.e., the search term) and each result keyword obtained based on the search term to be detected (such as keyword 1, keyword 2, etc.), and the output data of the semantic feature extraction network is the target semantic features of the search term to be detected.

[0099] As a possible implementation manner, the input data of the quality feature extraction network (which can also be called semantic input data) includes the quality evaluation values of the search term to be retrieved in N evaluation dimensions, and the value of N is a positive integer. The N evaluation dimensions include, but are not limited to, any one or more of the evaluation dimensions such as law violation and expression ability. Exemplarily, corresponding single-dimensional evaluation models can be set in advance for the N evaluation dimensions respectively, and the quality evaluation values corresponding to the N evaluation dimensions are obtained respectively according to the single-dimensional evaluation models corresponding to the N evaluation dimensions. The semantic input data may also include posterior features, such as search feature 1, search feature 2, search feature M, etc.

[0100] As a possible implementation manner, the quality discrimination model is implemented using the Wide and Deep structure. Among them, the Deep part is the semantic feature extraction network, and the Wide part is the quality feature extraction network. The semantic feature extraction network is implemented using a deep neural network. The quality feature extraction network is implemented using a linear model or a shallow neural network.

[0101] As a possible implementation, the semantic feature extraction network includes: a first feature extraction sub-network. The first feature extraction sub-network is used to extract the semantic features of the search term to be detected. The first feature extraction sub-network can be implemented by a deep neural network such as BERT, but is not limited thereto. Exemplarily, based on the semantic input data, using the first feature extraction sub-network, each word segmentation feature is obtained, and directly based on each word segmentation feature, the target semantic feature is obtained. For example, any one of the word segmentation features is used as the target semantic feature.

[0102] The semantic feature extraction network may further include: a second feature extraction sub-network. The first feature extraction sub-network is used to extract the target semantic feature including context information. The second feature extraction sub-network can be implemented by, but not limited to, network structures such as Bi-GRU and LTSM that can process sequence data. In the case where the semantic feature extraction network includes the first feature extraction sub-network and the second feature extraction sub-network, based on the semantic input data, using the first feature extraction sub-network, the word segmentation features of each word are obtained, and then, based on the word segmentation features of each word, using the second feature extraction sub-network, the target semantic feature including context information is obtained. Through the second feature extraction sub-network, the model's representation of the context can be enhanced, thereby obtaining richer and more accurate semantic features.

[0103] As a possible implementation, the quality feature extraction network may include multiple fully connected layers. The number of nodes included in the last fully connected layer of the multi-layer fully connected layers exceeds the feature dimension number of the quality input data. Taking three fully connected layers as an example, the number of nodes included in each of the three fully connected layers is 20, 40, and 80 respectively. By gradually increasing the number of nodes, after inputting 20-dimensional features, 100-dimensional features can be obtained.

[0104] The data processing process executed by the above model will be introduced in detail in the subsequent process, so it will not be elaborated here too much.

[0105] In some embodiments, before training, hyperparameters such as the batch, number of epochs, and learning rate required for training the quality discrimination model can also be set. For example, set the batch of the quality discrimination model to 64, the number of epochs to 1000, and the learning rate to 0.0001, that is, perform iterative training 1000 times, and divide the training samples into 64 batches for learning each time. Of course, the parameter values here are only one possible example, and can be adjusted according to requirements in actual situations.

[0106] The training process of the quality discrimination model is a process of repeatedly training using training samples. During the iterative training process, all training samples are divided into specified batches, and training is performed based on the training samples of each sub-batch. Since the steps executed during the training for each batch in each iteration process are similar, only the training for one batch will be used as an example for illustration.

[0107] Refer to Figure 3 As shown, it is a schematic diagram of a model training method provided in an embodiment of the present application. This method is applied to an electronic device, which can be a terminal device or a server.

[0108] S301. Obtain a number of training samples of one batch, where each training sample contains a sample search term.

[0109] In the embodiments of the present application, both the sample search term and the search term to be detected are texts for information search. For example, they can be texts for searching videos, texts for searching commodity information, or texts for searching general information. The sample search term is the search term processed during the training process, and the search term to be detected is the search term processed during the actual application process. The search term can be composed of one or more strings, and there is no limitation on this.

[0110] Since the respective result keywords associated with the sample search term and the quality evaluation values of N evaluation dimensions associated with the sample search term are also required to be used in the subsequent data processing process, the training samples also need to include the respective result keywords associated with the sample search term and the quality evaluation values of N evaluation dimensions.

[0111] In some embodiments, as a possible situation, keyword parsing can be performed on each training sample in advance to obtain the corresponding respective result keywords, and then the obtained respective result keywords are added to the corresponding training sample. That is to say, each training sample contains the respective result keywords associated with the corresponding sample search term.

[0112] Among them, the respective result keywords associated with a sample search term are obtained by performing keyword parsing on the sample search term according to the word frequencies of the respective words included in each search result using each search result obtained based on the sample search term.

[0113] The search result is used to represent the relevant information of the result item returned after searching for the sample search term. For example, a search result can include any one or more of the title, abstract, and respective content keywords of a result item.

[0114] The recall methods of search results can be prefix recall, pinyin prefix recall, simple spelling recall, word - granularity recall, semantic recall based on deep learning, etc., but are not limited thereto. The embodiments of the present application do not limit the manner of obtaining multiple search results corresponding to a search term (including a sample search term and a search term to be detected).

[0115] Each search result obtained based on a sample search term can be the relevant information of all result items returned after searching for the sample search term, or can be the relevant information of some result items returned after searching for the sample search term, and this is not limited. In the actual application process, considering the existence of a large number of result items, to improve the detection efficiency, usually the relevant information of the first k items (such as the first 10 items) of the returned results is used as the search results respectively. Since generally each result item is sorted according to the relevance to the sample search term, therefore, performing subsequent processing on the first k result items can not only reduce the amount of data processing, thereby improving the data processing efficiency, but also avoid the influence of irrelevant or less relevant result items on the semantic extraction result, thereby improving the accuracy of the subsequently extracted semantic features, and further improving the detection accuracy.

[0116] As an example, if a search result contains: the title of a result item and each content keyword, then, the following method can be used to obtain each result keyword associated with the sample search term:

[0117] Perform word segmentation on the titles in each search result to obtain each title word segmentation, perform title word frequency statistics on each title word segmentation, and according to the title word frequency statistics result, select result keywords from each title word segmentation according to the set number of title selections;

[0118] Perform content word frequency statistics on each content keyword in each search result, and according to the content word frequency statistics result, select result keywords from each content keyword according to the set number of content selections.

[0119] In the embodiments of the present application, the word - segmentation method is not specifically limited, and can be implemented by methods such as a word - segmentation method based on deep learning, a statistical - based method, a dictionary - based method, etc. The content keyword can be composed of one or more characters. Each content keyword in the search result can be obtained by performing word segmentation on the XML file of the result item, but is not limited thereto.

[0120] For example, refer to Figure 4As shown in the figure, it is a schematic diagram of a process for selecting result keywords provided in an embodiment of the present application. Assume that the sample search term is "where to see the most beautiful sea". Search is performed for the sample search term to obtain the search results corresponding to each result item such as result item 1, result item 2, result item 3, etc. Among them, the search results of each result item are sorted according to the relevance to the sample search term. The search result of result item 1 has the highest relevance to the sample search term, and the search results of result items 2, 3, etc. have decreasing relevance to the sample search term in turn. Taking result item 1 as an example only, the search result of result item 1 includes the title "Recommendations for the Top Ten Sea-view Punch Card Spots" and each content keyword ["sea", "fish", "sand", "summer", "catch", "sunset", "good", "scenery", "ten", "bay"].

[0121] Next, assume that the set number of title selections is 3. Perform word segmentation on the titles of the first 10 result items to obtain each title word segmentation, and perform title word frequency statistics on each title word segmentation. Then, according to the title word frequency statistics results and in accordance with the set number of title selections, select 3 title word segmentations: see the sea # recommendation # punch card as result keywords, where # represents the separator between result keywords.

[0122] Then, assume that the set number of content selections is 6. Perform content word frequency statistics on each content keyword in the first 10 result items, and according to the content word frequency statistics results and in accordance with the set number of content selections, select 6 result keywords from each content keyword: beach # sunset # island # catch the sea # best # diving.

[0123] In the above implementation method, word frequency statistics are performed on the titles and content keywords of the result items, and then result keywords are selected according to the corresponding statistical results. Since the title is usually a summary of the main content and theme of the search term, and the content keyword can express the main content in the result item, therefore, the result keywords selected according to the title and content keyword can, to a certain extent, more accurately express the main content of the search result, thereby making semantic supplementation for the sample search term, and further avoiding improving the accuracy of subsequent semantic features in the case where the content of the sample search term itself is less.

[0124] It should be noted that in the embodiment of the present application, not only can word frequency statistics be performed on the titles and content keywords and result keywords be selected according to the corresponding statistical results, but also word frequency statistics can be performed on any one or more of the title, abstract, and content keywords, and result keywords can be selected according to the corresponding statistical results. There is no limitation on this. Since the process of selecting result keywords is similar, it will not be elaborated here.

[0125] As another possible scenario, if a training sample does not contain the result keywords associated with the corresponding sample search term, then after obtaining the sample search term, the result keywords associated with the sample search term can be generated according to the process of obtaining result keywords mentioned above. Then, the generated result keywords are added to the training sample, and subsequently, the training sample with the added result keywords is used for subsequent model training.

[0126] In some embodiments, as a possible case, quality assessment can be performed on each training sample in advance to obtain the quality assessment values of N evaluation dimensions associated with the corresponding sample search term. Then, the obtained quality assessment values of the N evaluation dimensions are added to the corresponding training sample. That is to say, each training sample contains the quality assessment values of N evaluation dimensions associated with the corresponding sample search term, and N is a positive integer.

[0127] Among them, the evaluation dimensions include but are not limited to any one or more of the following: lawbreaking, expression ability, etc. Expression ability includes but is not limited to any one or more of the following: typos, unclear expression, non-standard expression, awkward expression, etc.

[0128] As a possible implementation, the quality assessment values of N evaluation dimensions associated with a sample search term can be obtained in the following way: The sample search term is input into the trained single-dimensional evaluation models corresponding to the N evaluation dimensions respectively to obtain the quality assessment values corresponding to the N evaluation dimensions respectively.

[0129] Considering that some evaluation dimensions (such as expression ability) specifically include multiple sub-evaluation dimensions, as an example, a corresponding single-dimensional evaluation model can be trained for each sub-evaluation dimension, and then the quality assessment values of the corresponding sub-evaluation dimensions are obtained by using the single-dimensional evaluation models corresponding to each sub-evaluation dimension respectively. As another example, a single-dimensional evaluation model (i.e., a multi-classification model) that can output the quality assessment values corresponding to multiple sub-evaluation dimensions can also be trained for this evaluation dimension, so as to directly obtain the quality assessment values corresponding to multiple sub-evaluation dimensions by using the multi-classification model. Of course, corresponding multi-classification models can also be trained for the N evaluation dimensions, so as to directly output the quality assessment values corresponding to the N evaluation dimensions by using the multi-classification models.

[0130] Among them, the quality assessment value of an evaluation dimension is used to represent the probability that there are quality problems with the sample search term under the corresponding evaluation dimension. The quality assessment value can be represented by a numerical value or in the form of a grade, etc., and there is no limitation on this. In this article, only the case where the quality assessment value is represented by a numerical value is used as an example for illustration, and the larger the numerical value, the greater the probability that there are quality problems with the sample search term.

[0131] Refer to Figure 5 As shown in the figure, it is a logical schematic diagram of the process for obtaining a quality evaluation value provided in an embodiment of the present application. Assume that the evaluation dimensions include: violation of laws and disciplines, expression ability, etc. Among them, the expression ability includes: missing characters, misspelled characters, extra characters. The sample search term is "where to see the most beautiful sea". The sample search term is input into the single-dimensional evaluation model corresponding to violation of laws and disciplines, and a quality evaluation value 1 with a value of 0.096042 is obtained. The quality evaluation value 1 represents that the probability of the sample search term involving violation of laws and disciplines is 0.096042; the sample search term is input into the single-dimensional evaluation model corresponding to the expression ability, and a quality evaluation value 2 with a value of 0.797758, a quality evaluation value 3 with a value of 0, and a quality evaluation value 4 with a value of 0 are obtained. Among them, the quality evaluation value 2 represents that the probability of the sample search term having missing characters is 0.797758, the quality evaluation value 3 represents that the probability of the sample search term having misspelled characters is 0, and the quality evaluation value 4 represents that the probability of the sample search term having extra characters is 0.

[0132] As another possible situation, if a training sample does not contain the quality evaluation values of N evaluation dimensions associated with the corresponding sample search term, then, after obtaining the sample search term, according to the process for obtaining the quality evaluation value mentioned above, the quality evaluation values of N evaluation dimensions associated with the sample search term can be generated, and then, the generated quality evaluation values of the associated N evaluation dimensions are added to the training sample, and then the training sample with the added quality evaluation value is used for subsequent model training.

[0133] In some embodiments, to further improve the detection accuracy, in the embodiments of the present application, the posterior features of the sample search term can also be combined for modeling, and by extracting the implicit relationship between the posterior features and the search term quality, the quality detection of the search term can be assisted. Herein, the posterior features are used to describe the features of the data distribution related to the search of the search term (whether it is a sample search term or a search term to be detected), so the posterior features can also be called search features.

[0134] Specifically, as a possible implementation manner, for each training sample in a number of training samples, M search features can be obtained in advance based on the historical search records of the sample search term in the training sample, and the obtained M search features are added to the corresponding training sample, where M is a positive integer. That is to say, each training sample contains M search features associated with the corresponding sample search term.

[0135] Among them, the search features include but are not limited to: the number of search clicks (QV) within a preset time interval, the latest search time, the first search time, etc. Exemplarily, the preset time interval can be 1 day, 3 days, or 1 week, but is not limited thereto.

[0136] As another possible implementation, if a training sample does not contain the historical search records associated with the corresponding sample search term, then after obtaining the sample search term, the historical search records can be added to the training sample, and then the training sample with the added quality evaluation value can be used for subsequent model training.

[0137] In the above implementation, by adding posterior features that can dynamically reflect the search situation of the search term, it is possible to effectively refer to the feedback of various objects on the search term, and in the final quality detection, the quality of the search term can be detected more accurately.

[0138] S302. Input a number of training samples into the quality discrimination model to obtain corresponding prediction results.

[0139] Refer to Figure 6 As shown, it is a schematic flowchart of a quality discrimination method provided in an embodiment of the present application. Since the processing process of each training sample is similar, only the training sample x will be described below as an example. The training sample x can be any one of a number of training samples.

[0140] S601. Based on each result keyword in the training sample x, use the semantic feature extraction network to extract the semantic features of the sample search term x in the training sample x to obtain the target semantic features of the sample search term x.

[0141] In an embodiment of the present application, the training sample x includes: the sample search term x, each result keyword associated with the sample search term x, and N quality evaluation values associated with the sample search term x. The training sample x may further include: M search features associated with the sample search term x.

[0142] Specifically, according to the set splicing method, splice each result keyword in the training sample x and the sample search term x to obtain semantic input data, and then, based on the semantic input data, use the semantic feature extraction network to obtain the target semantic features of the sample search term x.

[0143] Among them, the set splicing method can be that the sample search term x is in the front and each result keyword is in the back, or each result keyword is in the front and the sample search term x is in the back, and there is no limitation on this. Considering that the BERT model may be used subsequently, when splicing, the start marker [CLS] and the end marker [SEP] can also be added. Among them, the start marker [CLS] is used to indicate the beginning of the semantic input data, and the end marker [SEP] is used to indicate the beginning and end of the semantic input data.

[0144] As a possible implementation, the semantic feature extraction network includes a first feature extraction subnetwork, which is used to extract semantic features. The target semantic features of the sample search term x can be obtained by, but not limited to, the following methods:

[0145] First, the semantic input data is segmented according to the set segmentation method to obtain each segmentation word;

[0146] Next, the first feature extraction sub-network is used to extract features from each word to obtain the word segmentation features of each word;

[0147] Finally, based on the obtained word segmentation features, the target semantic features of the sample search term x are obtained.

[0148] It should be noted that in the embodiments of the present application, the word segmentation method is not limited and will not be described in detail. When the first feature extraction sub-network adopts the BERT model, the tools provided by the BERT model are used for word segmentation. Since the sample search terms and the keywords of each result are more critical for the quality detection of the search terms, in the Chinese language scenario, the MacBERT model can be used as the first feature extraction sub-network. MacBERT is an improved pre-trained language representation model based on BERT, which can better adapt to the Chinese language scenario, thereby improving the effect of the pre-trained model in the search term quality detection task.

[0149] See Figure 7 As shown in FIG, it is a structural diagram of a semantic feature extraction network provided in an embodiment of the present application. Assuming that the target semantic feature is a 768-dimensional feature vector, first, the sample search term x and the M result keywords are spliced in a splicing manner in which the sample search term is placed in front and the result keywords are placed in the back to obtain semantic input data, and the semantic input data is segmented using BERT's own tools. The segmentation results include: q1, q2, q3, ..., q corresponding to the sample search term x. L There are L participles, k1, k2, k3, ..., k M There are M result keywords, and there is a start tag [CLS] before q1, k M After that there is an end marker [SEP].

[0150] Next, embedding is performed on the given vocabulary, mapping each word into a 768-dimensional word vector, where the start tag [CLS] is mapped to the word vector E [CLS] ,q1,q2,q3,…,q L Mapped to E [1] 、E [2] 、E [3] ,…,E [L] , k1, k2, k3, …, k MAre respectively mapped to word vectors E [L+1] 、E [L+2] 、E [L+3] 、…、E [L+M] ,and the end marker [SEP] is mapped to word vector E [SEP] 。

[0151] After that, the obtained word vectors are input into the multi-layer structure of BERT to obtain the tokenization features of each token. Among them, the tokenization feature of the start marker [CLS] is T [CLS] ,and the tokenization features of q1, q2, q3, …, q L are respectively T [1] 、T [2] 、T [3] 、…、T [L] ,and the tokenization features of k1, k2, k3, …, k M are respectively T [L+1] 、T [L+2] 、T [L+3] 、…、T [L+M] ,and the tokenization feature of the end marker [SEP] is T [SEP] 。It should be noted that only a two-layer structure is taken as an example for illustration in the figure. In the actual application process, the number of layers is not limited to two layers.

[0152] Finally, take S as the target semantic feature. Of course, T [1] 、T [2] 、…、T [L+M] in any item can also be taken as the target semantic feature, and C can also be taken as the target semantic feature. However, compared with the search terms and result keywords already existing in the semantic input data, symbols without obvious semantic information can relatively fairly fuse the semantic information of the search terms and result keywords in the semantic input data. Therefore, usually S or C is taken as the target semantic feature.

[0153] In the above implementation method, for each token obtained based on the semantic input data, using the first feature extraction sub-network such as BERT to extract features from the semantic input data, a target semantic feature that fuses the full-text semantic information of the semantic input data can be obtained, so as to make full use of the information of the search term itself, improve the semantic accuracy of the target semantic feature, and further improve the accuracy of subsequent quality inspection.

[0154] As another possible implementation method, in addition to the first feature extraction sub-network, the semantic feature extraction network also includes a second feature extraction sub-network, and the second feature extraction sub-network is used to extract the target semantic feature containing the context information of each token. The target semantic feature of the sample search term x can be obtained in the following ways but is not limited to them:

[0155] First, segment the semantic input data according to the set segmentation method to obtain each segment.

[0156] Next, use the first feature extraction sub-network to extract features from each segment to obtain the respective segment features of each segment.

[0157] After that, based on the segment features, use the second feature extraction sub-network to enhance the segment features of each segment to obtain the target semantic features containing context information.

[0158] In the embodiments of this application, the segmentation method is not limited and will not be elaborated here. When the first feature extraction sub-network adopts the BERT model, the tools provided by the BERT model are used for segmentation. Exemplarily, the MacBERT model can be specifically used as the first feature extraction sub-network.

[0159] The second feature extraction sub-network is used to enhance the model's representation of the context. Exemplarily, the Bi-GRU is adopted as the second feature extraction sub-network.

[0160] Refer to Figure 8 As shown, it is a schematic structural diagram of a semantic feature extraction network provided in the embodiments of this application. Assume that the target semantic feature is a 768-dimensional feature vector. First, splice the sample search term x and M result keywords in the splicing method where the sample search term is in the front and each result keyword is in the back to obtain the semantic input data, and use the tools provided by BERT to segment the semantic input data. The segmentation results include: q1, q2, q3,..., q L and other L segments, k1, k2, k3,..., k M and other M result keywords. There is a start marker [CLS] before q1, and there is an end marker [SEP] after k M .

[0161] Next, perform embedding through a given vocabulary to map each segment to a 768-dimensional word vector. Among them, the start marker [CLS] is mapped to the word vector E [CLS] , and q1, q2, q3,..., q L are respectively mapped to E [1] , E [2] , E [3] ,..., E [L] , and k1, k2, k3,..., k M are respectively mapped to the word vectors E [L+1] , E [L+2] , E [L+3] ,..., E [L+M] , and the end marker [SEP] is mapped to the word vector E [SEP] .

[0162] After that, the obtained word vectors are input into the multi-layer structure of BERT to obtain the token features of each token. Among them, the token feature of the start token [CLS] is C, and the token features of q1, q2, q3, …, q L are respectively T [1] , T [2] , T [3] , …, T [L] , k1, k2, k3, …, k M are respectively T [L+1] , T [L+2] , T [L+3] , …, T [L+M] , and the token feature of the end token [SEP] is S. It should be noted that only a two-layer structure is taken as an example for illustration in the figure, and in the actual application process, the number of layers is not limited to two layers.

[0163] Then, the token features of each token are input into the second feature extraction sub-network to obtain the target semantic features. The second feature extraction sub-network is composed of a 3-layer Bi-GRU, and each layer of Bi-GRU contains L + M + 2 nodes, and each node represents a Bi-GRU unit. The Bi-GRU unit can specifically include a forward GRU unit and a backward GRU unit (not shown in the figure), and the structure of the Bi-GRU unit will not be elaborated here.

[0164] Finally, the feature vector (with a length of 768) output by any node in the last hidden layer is used as the target semantic feature.

[0165] S602. Based on the N quality evaluation values in the training sample x, use the quality feature extraction network to extract the quality features of the sample search term x to obtain the multi-dimensional quality features of the sample search term x.

[0166] Specifically, when executing S602, the following methods can be adopted but are not limited to:

[0167] First, based on the N quality evaluation values in the training sample x, construct quality input data;

[0168] Next, based on the quality input data, use the quality feature extraction network to extract the quality features of the sample search term x to obtain the multi-dimensional quality features of the sample search term x.

[0169] Among them, constructing quality input data based on the N quality evaluation values in the training sample x includes but is not limited to the following methods:

[0170] Construction method 1: Directly splice the N quality evaluation values to construct quality input data.

[0171] Construction method 2: Based on the historical search records of the sample search term x, obtain M search features, and splice the N quality evaluation values and the M search features to construct quality input data.

[0172] It should be noted that in the embodiments of the present application, the splicing method is not limited and will not be elaborated here.

[0173] In construction method 1, by using the quality evaluation value as the model input, the quality of the search term is detected from each evaluation dimension, thereby improving the accuracy of quality detection. In construction method 2, by adding posterior features that can dynamically reflect the search situation of the search term, it is possible to effectively refer to the feedback of various objects on the search term, and thus in the final quality detection, the quality of the search term can be detected more accurately.

[0174] In some embodiments, since the quality evaluation value is usually a floating point number between 0 and 1, and the search feature is usually some integer greater than 1, for the convenience of subsequent model processing, the search feature can also be normalized, so as to convert the search feature into data with a value range of -1 to 1. Exemplarily, the normalization process can be to calculate the reference value (such as the average value) corresponding to each search feature through the historical search records of the sample search term x, and then, using the reference value, obtain the data after normalization by calculating (search feature - reference value) / reference value. If the value of the data after normalization is less than -1 or greater than 1, it is correspondingly converted to -1 and 1.

[0175] For example, the search features include: the number of search clicks within a preset time interval, the latest search time. Among them, the number of search clicks within the preset time interval is 7, and the latest search time is 14717838. After normalization, the number of search clicks within the preset time interval is 0.335219, and the latest search time is -0.674412.

[0176] In some embodiments, the quality feature extraction network includes multiple fully connected layers, and the number of nodes included in the last fully connected layer in the multiple fully connected layers exceeds the feature dimension of the quality input data.

[0177] Through multiple fully-connected layers, feature extraction and fusion are performed on multiple-dimensional quality features of the search term and the posterior features of the search term, so as to obtain the relationship between the quality features and the posterior features and the quality level of the search term. Since each node in the fully-connected layer is connected to all nodes in the previous layer, due to its fully-connected property, the features extracted previously in the input can be combined, so as to learn the interaction relationship between the features. In this article, the interaction relationship between the features can also be called cross features. Cross features refer to the combination of different features as new features input into the model to better capture the correlation between the features. The introduction of cross features can enhance the expression ability of the model.

[0178] See Figure 9 As shown, it is a schematic structural diagram of a quality feature extraction network provided in an embodiment of the present application. The quality feature extraction network is composed of three fully-connected layers. The three fully-connected layers include: fully-connected layer 1, fully-connected layer 2, and fully-connected layer 3. Fully-connected layer 1 contains 20 nodes, fully-connected layer 2 contains 40 nodes, and fully-connected layer 3 contains 100 nodes.

[0179] The input of each node in fully-connected layer 1 is quality input data, and the quality input data is composed of f [1] 、f [2] 、f [3] 、…、f

[20] . That is, the feature dimension of the quality input data is 20-dimensional, and f [1] 、f [2] 、f [3] 、…、f

[20] are respectively used to represent the quality evaluation value of law-breaking, the quality evaluation value of expression ability (specifically including the quality evaluation value of fewer characters, the quality evaluation value of misspelled characters, the quality evaluation value of more characters), the number of search clicks within a preset time interval, the latest search time, the earliest search time, etc.

[0180] The input of each node in fully-connected layer 2 is: the feature vectors output by each of the 20 nodes in fully-connected layer 1, and the input of each node in fully-connected layer 3 is: the feature vectors output by each of the 40 nodes in fully-connected layer 2.

[0181] Through fully-connected layer 3, multi-dimensional quality features composed of F [1] 、F [2] 、F [3] 、…、F

[100] can be obtained. The feature dimension of the multi-dimensional quality features is 100-dimensional, and F [1] 、F [2] 、F [3] 、…、F

[100] respectively represent the values of one dimension.

[0182] S603. Based on the target semantic features and multi-dimensional quality features, use a classification network to perform quality identification on the sample search term x, and obtain the quality inspection result of the sample search term x.

[0183] In the embodiments of the present application, the classification network can be implemented by, but not limited to, a linear regression model.

[0184] Specifically, as a possible implementation, based on the target semantic features and multi-dimensional quality features, use a classification network to obtain the target detection value of the sample search term x; based on the target detection value, combine the set detection threshold to obtain the quality inspection result of the sample search term x.

[0185] Among them, the target detection value is used to represent the probability that the sample search term x belongs to an abnormal search term. The target detection value can be represented by a numerical value or other forms, and there is no limitation on this.

[0186] It should be noted that in the embodiments of the present application, the search terms are divided into low-quality search terms and non-low-quality search terms. The quality inspection result of the sample search term x is used to represent whether the sample search term x is a low-quality search term, that is, the abnormal search term is a low-quality search term. However, in the actual application process, the search terms can also be divided into multiple categories, such as low-quality search terms, medium-high-quality search terms, high-quality search terms, etc. At this time, the abnormal search term can be a low-quality search term, or a low-quality search term + medium-high-quality search term. Obviously, by outputting the probability that the sample search term x belongs to an abnormal search term and combining the set detection threshold, a multi-classification problem can be transformed into a single-classification problem, so as to meet different classification needs, and because the calculation process is relatively simple, the quality detection efficiency can also be improved.

[0187] In some embodiments, it is also possible to obtain the detection values of each category of the sample search term x based on the target semantic features and multi-dimensional quality features, using a classification network, and then combine the corresponding thresholds to determine whether the sample search term x belongs to one or more categories, and there is no limitation on this.

[0188] As an example, a detection threshold can be set in advance. When the target detection value exceeds the set detection threshold, the quality inspection result of the sample search term x represents that the sample search term x is an abnormal search term.

[0189] For example, the input of the classification network is: 768-dimensional target semantic features and 100-dimensional multi-dimensional quality features. In the classification network, a linear regression model is used for calculation to obtain a calculation result, and an activation function (softmax function) is used to process the calculation result to obtain the target detection value of the sample search term. Exemplarily, the target detection value can be represented by the following formulas (1) and (2):

[0190] F = cat(F query, F wide ) Formula (1)

[0191] Score = softmax(WF + B) Formula (2)

[0192] Wherein, F query represents the target semantic feature, F wide represents the multi-dimensional quality feature, cat() is used to concatenate F query and F wide , F represents the feature vector composed of F query and F wide , Score represents the target detection value, W represents the feature matrix, WF represents W multiplied by F, and B is the bias vector.

[0193] S303. Based on the obtained several prediction results, combine with the true results corresponding to several training samples respectively to obtain the model loss.

[0194] In the embodiments of the present application, the true result corresponding to a training sample is used to characterize whether the sample search term included in the training sample belongs to an abnormal search term. For example, the true result of the sample search term "where to see the most beautiful sea" characterizes that the sample search term is an abnormal search term. Another example, the true result of the sample search term "recommend by the sea" characterizes that the sample search term is an abnormal search term.

[0195] The model loss can be calculated by loss functions such as the cross-entropy (Cross Entry) loss function, the quadratic loss function, and the absolute loss function, but is not limited thereto.

[0196] S304. Determine whether the quality discrimination model meets the convergence condition. If so, execute S305; otherwise, execute S306, and then perform the next batch of learning.

[0197] In the embodiments of the present application, the convergence condition may include at least one of the following conditions:

[0198] (1) The total loss value is not greater than the preset loss value threshold.

[0199] (2) The number of iterations reaches the preset upper limit value.

[0200] S305. Output the quality discrimination model.

[0201] S306. Adjust the model parameters based on the model loss.

[0202] That is to say, if the convergence condition is met, it is determined that the quality discrimination model has met the convergence condition, and the training ends. Otherwise, it is determined that the quality discrimination model has not met the convergence condition. Furthermore, it is necessary to continue adjusting the model parameters and use the adjusted identity recognition model to enter the next training process, that is, jump to S301 for execution.

[0203] To improve the training efficiency, before the process of Figure 3 training, the quality discrimination model can also be pre-trained respectively. After using the pre-trained quality discrimination model to participate in the Figure 3 training process, the training data used in the pre-training stage can be not limited to search terms, but also other types of texts. After pre-training, the quality discrimination model has converged, and then use the Figure 3 training process for fine-tuning learning to obtain the trained quality discrimination model.

[0204] Specifically, during the fine-tuning process, a training sample set is obtained, and each training sample contains a sample search term. Then, based on the training sample set, the pre-trained quality discrimination model is iteratively trained to obtain the trained quality discrimination model. Among them, in each iteration process, a batch of several training samples are input into the pre-trained quality discrimination model to obtain corresponding prediction results, and based on the obtained several prediction results, combined with the true results corresponding to each of the several training samples, the model loss is obtained, and the model parameters are adjusted based on the model loss.

[0205] Refer to Fig.10 shown in the figure, which is a schematic diagram of a search term quality detection method provided in an embodiment of the present application. This method is applied to an electronic device, and the electronic device can be a terminal device or a server. The specific process is as follows:

[0206] S1001: For each search result obtained based on the search term to be detected, keyword parsing is performed on each search result based on the word frequency of each word included in each search result to obtain each result keyword.

[0207] The process of obtaining each result keyword for the search term to be detected is similar to the process of obtaining each result keyword for the sample search term in the above text. For details, please refer to the above text and will not be elaborated here.

[0208] S1002: According to at least one evaluation dimension, the quality of the search term to be detected is evaluated respectively to obtain the quality evaluation value of the corresponding evaluation dimension.

[0209] The process of obtaining the quality evaluation value for the search term to be detected is similar to the process of obtaining the quality evaluation value for the sample search term in the above text. For details, please refer to the above text and will not be elaborated here.

[0210] S1003. Extract semantic features from the search term to be detected based on each result keyword, and obtain the target semantic features of the search term to be detected.

[0211] As a possible implementation, based on each result keyword, use the semantic feature extraction network in the trained quality discrimination model to obtain the target semantic features of the search term to be detected. For details, see S601, which will not be elaborated here.

[0212] S1004. Extract quality features from the search term to be detected based on the at least one obtained quality evaluation value, and obtain the multi-dimensional quality features corresponding to the search term to be detected.

[0213] As a possible implementation, based on the at least one obtained quality evaluation value, use the quality feature extraction network in the trained quality discrimination model to obtain the multi-dimensional quality features of the search term to be detected. For details, see S602, which will not be elaborated here.

[0214] S1005. Perform quality identification on the search term to be detected based on the target semantic features and the multi-dimensional quality features, and obtain the quality inspection result of the search term to be detected.

[0215] As a possible implementation, based on the target semantic features and the multi-dimensional quality features, use the classification network in the trained quality discrimination model to perform quality identification on the search term to be detected, and obtain the quality inspection result of the search term to be detected. For details, see S603, which will not be elaborated here.

[0216] It should be noted that in the embodiments of the present application, the execution order among S1001 to S1005 only needs to satisfy that S1001 is executed before S1003, S1002 is executed before S1004, and S1003 and S1004 are executed before S1005. For example, it can be executed in the order of S1001 - S1005, or in the order of S1001, S1003, S1002, S1004, S1005, or in the order of S1002, S1004, S1001, S1003, S1005, but it is not limited thereto.

[0217] In some embodiments, the quality inspection result can be used to recall the abnormal search terms in the reference search term set, so as to filter the abnormal search terms. When associating search terms, the abnormal search terms can be avoided from being presented to the user, thereby providing a better network environment for the user and improving the user experience.

[0218] Specifically, if the detection result of the search term to be detected indicates that the search term to be detected is an abnormal search term, then filter the search term to be detected from the reference search term set;

[0219] When a search term association indication carrying a target search term is received, at least one reference search term that meets the set similarity condition is selected from each reference search term in the filtered reference search term set based on the semantic similarity between each reference search term and the target search term;

[0220] Based on at least one reference search term, search term recommendation for the target search term is performed.

[0221] Among them, the search term to be detected can be a reference search term in the reference search term set that needs to be quality detected. The reference search term can be a historical search term or an automatically generated search term. One or more reference search terms are stored in the reference search term set.

[0222] Filtering the search term to be detected can mean deleting the search term to be detected from the reference search term set, or it can mean using a marking symbol in the reference search term to mark the search term to be detected. The marking symbol identifies the search term to be detected as an abnormal search term, and the reference search term with the marking symbol does not participate in the subsequent screening, so as to avoid recommending abnormal search terms.

[0223] The semantic similarity can be calculated using algorithms such as cosine similarity, or can be obtained using a neural network model based on natural language processing, and there is no limitation on this.

[0224] The set similarity condition includes but is not limited to at least one of the following: the corresponding semantic similarity belongs to the top n semantic similarities, that is, the top n reference search terms with the highest semantic similarity are selected; the corresponding semantic similarity exceeds the set semantic similarity threshold.

[0225] As an example, after obtaining the target search term, reference search terms with the search term to be detected as a prefix can also be retrieved from the historical search term library, the semantic similarity between these historical search terms and the search term to be detected is calculated, and then screening is performed from these historical search terms according to the set similarity condition.

[0226] In addition, the popularity of each reference search term is also obtained. The popularity can be statistically calculated based on historical search frequencies. For example, it can be statistically calculated based on the number of searches in the past period of time. Then, based on the two dimensions of semantic similarity and popularity, each reference historical search term is screened. When screening based on the dimensions of semantic similarity and popularity, the set similarity condition can include but is not limited to at least one of the following: the corresponding semantic similarity belongs to the top n semantic similarities; the corresponding semantic similarity exceeds the set semantic similarity threshold; the popularity is higher than the set popularity threshold.

[0227] As a possible implementation, the terminal device sends a search term association indication carrying the target search term to the server in response to the target search term entered by the target object on the search page. After receiving the search term association indication, the server filters out at least one reference search term that meets the set similarity condition from each reference search term in the filtered reference search term set based on the semantic similarity between each reference search term and the target search term.

[0228] Refer to Fig.11A As shown, it is a logical schematic diagram of a search term association process provided in an embodiment of the present application. Assume that the reference search term set includes reference search term 1 ("Why is the beach pink?"), reference search term 2 ("xx fish is highly poisonous"), reference search term 3 ("Popular tourist attractions"), reference search term 4 ("TV drama recommendations"), reference search term 5 ("TV drama complaints"), etc. Among them, reference search term 2 is the search term to be detected, and the detection result of reference search term 2 indicates that reference search term 2 is an abnormal search term. Therefore, reference search term 2 is filtered out from the reference search term set.

[0229] When receiving the search term association indication carrying the target search term "beach", calculate the semantic similarity between each reference search term such as reference search term 1, reference search term 3, reference search term 4, reference search term 5 and the target search term, and then filter out the reference search terms that meet the set similarity condition from each reference search term based on the calculated semantic similarity. Assume that the filtered reference search term is reference search term 1. Then, based on reference search term 1, perform a search term recommendation for the target search term, that is, reference search term 1 is the associated recommendation term for the target search term.

[0230] Refer to Fig. 11B As shown, it is a schematic diagram of a search interface provided in an embodiment of the present application. When the target object enters the target search term "beach" in the search bar, the associated recommendation terms of the target search term are presented below the search bar. The associated recommendation terms include "Why is the beach pink?", "Beach wallpaper", "Puppies playing happily on the beach", "Why isn't there a trash can on the beach", etc.

[0231] In the above implementation, if the search term to be detected belongs to an abnormal search term, then the search term to be detected is filtered out from the reference search term set. In this way, when performing search term association subsequently, it is possible to avoid recommending low-quality search terms as the associated terms of the target search term to the user, and at the same time, the recommended search terms meet the user's search intention, further improving the user's search experience.

[0232] Next, a specific embodiment is used to illustrate the search term quality detection process.

[0233] Refer to Fig.12As shown, assume that the search platform is a video platform, and the search term to be detected is "xx movie recommendations".

[0234] First, for each search result obtained based on the search term to be detected, based on the word frequency of each word included in each search result, keyword parsing is performed on each search result to obtain each result keyword. Fig.12 In the figure, circles are used to represent result keywords, pentagons are used to represent quality evaluation values, and triangles are used to represent posterior features. The number of circles, pentagons, and triangles is only taken as an example of 4, and the actual number is not limited.

[0235] Meanwhile, according to N evaluation dimensions, the quality of the search term to be detected is evaluated respectively to obtain the quality evaluation values of the corresponding evaluation dimensions. Also, the historical search records of the search term to be detected are obtained, and based on the historical search records of the search term to be detected, M search features are obtained.

[0236] Next, each result keyword, N quality evaluation values, and M search features are input into the quality discrimination model to obtain the quality inspection result of the search term to be detected. Among them, the quality discrimination model includes a semantic feature extraction network, a quality feature extraction network, and a classification network. The semantic feature extraction network specifically includes a first feature extraction sub-network and a second feature extraction sub-network.

[0237] Specifically, each result keyword and the search term to be detected are concatenated to obtain semantic input data. Based on the semantic input data, using the first feature extraction sub-network, the respective token features of each token in the semantic input data are obtained. And based on each token feature, using the second feature extraction sub-network, the token features are enhanced to obtain the target semantic features including context information; based on N quality evaluation values, using the quality feature extraction network, multi-dimensional quality features of the search term to be detected are obtained; based on the target semantic features and the multi-dimensional quality features, using the classification network, the quality of the search term to be detected is identified to obtain the quality inspection result indicating that the search term to be detected is a low-quality search term.

[0238] Based on the same inventive concept, an embodiment of the present application provides a search term quality detection device. As Figure 13 shown, it is a schematic structural diagram of the search term quality detection device 1300, which may include:

[0239] A result keyword parsing unit 1301, configured to perform keyword parsing on each search result obtained based on the search term to be detected, based on the word frequency of each word included in each search result, to obtain each result keyword;

[0240] A quality evaluation unit 1302, configured to evaluate the quality of the search term to be detected respectively according to at least one evaluation dimension to obtain the quality evaluation values of the corresponding evaluation dimensions;

[0241] A semantic feature extraction unit 1303, configured to perform semantic feature extraction on the to-be-detected search term based on the result keywords, so as to obtain the target semantic feature of the to-be-detected search term;

[0242] A quality feature extraction unit 1304, configured to perform quality feature extraction on the to-be-detected search term based on at least one obtained quality evaluation value, so as to obtain the multi-dimensional quality feature corresponding to the to-be-detected search term;

[0243] A quality recognition unit 1305, configured to perform quality recognition on the to-be-detected search term based on the target semantic feature and the multi-dimensional quality feature, so as to obtain the quality inspection result of the to-be-detected search term.

[0244] As a possible implementation manner, when performing semantic feature extraction on the to-be-detected search term based on the result keywords to obtain the semantic feature of the to-be-detected search term, the semantic feature extraction unit 1303 is specifically configured to:

[0245] Splice the result keywords and the to-be-detected search term according to a set splicing manner to obtain semantic input data;

[0246] Based on the semantic input data, use the semantic feature extraction network of the trained quality discrimination model to obtain the semantic feature of the to-be-detected search term.

[0247] As a possible implementation manner, the semantic feature extraction network includes a first feature extraction sub-network, and the first feature extraction sub-network is configured to extract the semantic feature of the to-be-detected search term;

[0248] When using the semantic feature extraction network of the trained quality discrimination model to obtain the target semantic feature of the to-be-detected search term based on the semantic input data, the semantic feature extraction unit 1303 is specifically configured to:

[0249] Segment the semantic input data according to a set word segmentation manner to obtain each segmented word;

[0250] Use the first feature extraction sub-network to perform feature extraction on each segmented word to obtain the segmented word feature of each segmented word;

[0251] Based on the obtained segmented word features, obtain the target semantic feature of the to-be-detected search term.

[0252] As a possible implementation manner, the semantic feature extraction network further includes a second feature extraction sub-network, and the second feature extraction sub-network is configured to extract the target semantic feature including the context information of each segmented word;

[0253] When obtaining the target semantic features of the search term to be detected based on the obtained word segmentation features, the semantic feature extraction unit 1303 specifically is used for:

[0254] Based on the word segmentation features, using the second feature extraction sub-network, perform enhancement processing on the word segmentation features to obtain the target semantic features including the context information.

[0255] As a possible implementation manner, when extracting the quality features of the search term to be detected based on the obtained at least one quality evaluation value to obtain the multi-dimensional quality features corresponding to the search term to be detected, the quality feature extraction unit 1304 specifically is used for:

[0256] Based on the obtained at least one quality evaluation value, construct quality input data;

[0257] Based on the quality input data, use the quality feature extraction network in the trained quality discrimination model to extract the quality features of the search term to be detected to obtain the multi-dimensional quality features.

[0258] As a possible implementation manner, the quality feature extraction network includes multiple fully connected layers, and the number of nodes included in the last fully connected layer in the multiple fully connected layers exceeds the feature dimension number of the quality input data.

[0259] As a possible implementation manner, when constructing the quality input data based on the obtained at least one quality evaluation value, the quality feature extraction unit 1304 specifically is used for:

[0260] Directly splice the at least one quality evaluation value to construct the quality input data; or,

[0261] Based on the historical search records of the search term to be detected, obtain at least one search feature, and splice the at least one search feature and the at least one quality evaluation value to construct the quality input data.

[0262] As a possible implementation manner, when performing quality identification on the search term to be detected based on the target semantic features and the multi-dimensional quality features to obtain the quality inspection result of the search term to be detected, the quality identification unit 1305 specifically is used for:

[0263] Based on the target semantic features and the multi-dimensional quality features, use the classification network in the trained quality discrimination model to obtain the target detection value of the search term to be detected; the target detection value is used to represent the probability that the search term to be detected belongs to an abnormal search term;

[0264] Based on the target detection value, in combination with a set detection threshold, obtain the quality inspection result of the search term to be detected.

[0265] As a possible implementation, the quality recognition unit 1305 is further configured to:

[0266] Obtain a training sample set, where the training sample set contains each training sample, and each training sample contains a sample search term;

[0267] Based on the training sample set, iteratively train a pre-trained quality discrimination model to obtain a trained quality discrimination model. Wherein, in each iteration process, the following operations are performed:

[0268] Input a number of training samples in a batch into the pre-trained quality discrimination model to obtain corresponding prediction results;

[0269] Based on the obtained number of prediction results, combine the true results corresponding to the number of training samples to obtain a model loss, and adjust the model parameters based on the model loss.

[0270] As a possible implementation, the quality recognition unit 1305 is further configured to:

[0271] If the detection result indicates that the search term to be detected is an abnormal search term, then filter the search term to be detected from the reference search term set;

[0272] When receiving a search term association indication carrying a target search term, based on the semantic similarity between each reference search term in the filtered reference search term set and the target search term, select at least one reference search term that meets the set similarity condition from the reference search terms;

[0273] Based on the at least one reference search term, perform a search term recommendation for the target search term.

[0274] For the convenience of description, the above parts are divided into each module (or unit) according to functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0275] Regarding the device in the above embodiments, the specific manner in which each unit executes the request has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0276] Those skilled in the art to which the present application pertains can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0277] Based on the same inventive concept, an embodiment of the present application further provides an electronic device. In one embodiment, the electronic device may be a server or a terminal device. Refer to Fig.14 As shown, it is a schematic structural diagram of a possible electronic device provided in an embodiment of the present application. Fig.14 In this, the electronic device 1400 includes: a processor 1410 and a memory 1420.

[0278] Among them, the memory 1420 stores a computer program executable by the processor 1410. By executing the instructions stored in the memory 1420, the processor 1410 can execute the above Figure 3 or Figure 8 shown steps.

[0279] The memory 1420 may be a volatile memory, such as a random-access memory (RAM); the memory 1420 may also be a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1420 is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1420 may also be a combination of the above memories.

[0280] The processor 1410 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1410 is used to implement the above search term quality detection method when executing the computer program stored in the memory 1420.

[0281] In some embodiments, the processor 1410 and the memory 1420 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.

[0282] In the embodiments of the present application, the specific connection medium between the above-mentioned processor 1410 and the memory 1420 is not limited. In the embodiments of the present application, taking the connection between the processor 1410 and the memory 1420 through a bus as an example, the bus is described in thick lines in Fig.14 The connection manners between other components are only for illustrative purposes and are not to be taken as a limitation. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of description, Fig.14 It is only described by a thick line, but it does not describe that there is only one bus or one type of bus.

[0283] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the above-mentioned search term quality detection method. In some possible implementation manners, various aspects of the search term quality detection method provided in the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the above-mentioned search term quality detection method. For example, the electronic device can execute as Figure 3 or Figure 8 the steps shown in.

[0284] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (Compact Disk Read Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0285] The program product of the embodiment of the present application can adopt a CD-ROM and include a computer program, and can run on an electronic device. However, the program product of the present application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a computer program, and the computer program can be used by or in combination with a command execution system, apparatus, or device.

[0286] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and the readable medium can send, propagate, or transmit a computer program for use by or in combination with a command execution system, apparatus, or device.

[0287] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0288] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A search term quality detection method, characterized in that, Including: For each search result obtained based on the search term to be detected, keyword parsing is performed on each search result based on the word frequencies of the words included in each search result to obtain result keywords for each result; The quality of the search term to be detected is evaluated respectively according to at least one evaluation dimension to obtain a quality evaluation value for the corresponding evaluation dimension; Based on the result keywords for each result, semantic feature extraction is performed on the search term to be detected to obtain the target semantic features of the search term to be detected; Based on at least one obtained quality evaluation value, quality feature extraction is performed on the search term to be detected to obtain the multi-dimensional quality features corresponding to the search term to be detected; Based on the target semantic features and the multi-dimensional quality features, quality identification is performed on the search term to be detected to obtain the quality inspection result of the search term to be detected.

2. The method according to claim 1, wherein The semantic feature extraction of the search term to be detected based on the result keywords for each result to obtain the semantic features of the search term to be detected includes: According to a set splicing method, the result keywords for each result and the search term to be detected are spliced to obtain semantic input data; Based on the semantic input data, the semantic feature extraction network of the trained quality discrimination model is used to obtain the semantic features of the search term to be detected.

3. The method according to claim 2, characterized in that, The semantic feature extraction network includes a first feature extraction sub-network, and the first feature extraction sub-network is used to extract the semantic features of the search term to be detected; The obtaining of the target semantic features of the search term to be detected based on the semantic input data by using the semantic feature extraction network of the trained quality discrimination model includes: According to a set word segmentation method, the semantic input data is segmented to obtain each segmented word; The first feature extraction sub-network is used to perform feature extraction on each segmented word to obtain the segmented word features of each segmented word; Based on the obtained segmented word features, the target semantic features of the search term to be detected are obtained.

4. The method according to claim 3, wherein The semantic feature extraction network further includes a second feature extraction sub-network, and the second feature extraction sub-network is used to extract the target semantic features including the context information of each segmented word; The obtaining of the target semantic features of the search term to be detected based on the obtained segmented word features includes: Based on the segmented word features, the second feature extraction sub-network is used to perform enhancement processing on the segmented word features to obtain the target semantic features including the context information.

5. The method according to any one of claims 1-4, characterized in that, The quality feature extraction of the search term to be detected based on at least one obtained quality evaluation value to obtain the multi-dimensional quality features corresponding to the search term to be detected includes: Based on at least one obtained quality evaluation value, quality input data is constructed; Based on the quality input data, the quality feature extraction network in the trained quality discrimination model is used to perform quality feature extraction on the search term to be detected to obtain the multi-dimensional quality features.

6. The method according to claim 5, characterized in that, The quality feature extraction network includes multiple fully-connected layers, and the number of nodes included in the last fully-connected layer in the multiple fully-connected layers exceeds the feature dimension number of the quality input data.

7. The method according to claim 5, characterized in that, The constructing of the quality input data based on at least one obtained quality evaluation value includes: Directly splice the at least one quality evaluation value to construct the quality input data; or, Based on the historical search records of the search term to be detected, obtain at least one search feature, and splice the at least one search feature and the at least one quality evaluation value to construct the quality input data.

8. The method according to any one of claims 1 to 4, characterized in that The quality identification of the search term to be detected based on the target semantic feature and the multi-dimensional quality feature to obtain the quality inspection result of the search term to be detected includes: Based on the target semantic feature and the multi-dimensional quality feature, use the classification network in the trained quality discrimination model to obtain the target detection value of the search term to be detected; the target detection value is used to characterize the probability that the search term to be detected belongs to an abnormal search term; Based on the target detection value and in combination with a set detection threshold, obtain the quality inspection result of the search term to be detected.

9. The method according to any one of claims 1-4, characterized in that, The quality discrimination model is trained in the following manner: Obtain a training sample set, where the training sample set contains each training sample, and each training sample contains a sample search term; Based on the training sample set, perform iterative training on the pre-trained quality discrimination model to obtain the trained quality discrimination model. Among them, in each iteration process, perform the following operations: Input a batch of several training samples into the pre-trained quality discrimination model to obtain corresponding prediction results; Based on the obtained several prediction results, in combination with the respective true results corresponding to the several training samples, obtain the model loss, and adjust the model parameters based on the model loss.

10. The method according to any one of claims 1-4, characterized in that, It further includes: If the detection result indicates that the search term to be detected is an abnormal search term, then filter the search term to be detected from the reference search term set; When receiving a search term association indication carrying a target search term, based on the semantic similarity between each reference search term in the filtered reference search term set and the target search term, select at least one reference search term that meets the set similarity condition from the each reference search term; Based on the at least one reference search term, perform a search term recommendation for the target search term.

11. The method according to any one of claims 1 to 4, characterized in that, Each search result includes: the title of a result item and each content keyword; The keyword parsing of each search result based on the word frequency of each vocabulary included in each search result to obtain each result keyword includes: Perform word segmentation on the titles in each search result to obtain each title word segment, perform title word frequency statistics on each title word segment, and according to the title word frequency statistics result, select result keywords from each title word segment according to the set number of title selections; Perform content word frequency statistics on each content keyword in each search result, and according to the content word frequency statistics result, select result keywords from each content keyword according to the set number of content selections.

12. A search term quality detection device, characterized in that, It includes: A result keyword parsing unit, configured to perform keyword parsing on each search result obtained based on the search term to be detected, based on the word frequency of each vocabulary included in each search result, to obtain each result keyword; A quality assessment unit, configured to perform quality assessment on each of the search terms to be detected according to at least one assessment dimension, and obtain a quality assessment value of the corresponding assessment dimension; A semantic feature extraction unit, configured to extract semantic features of the search term to be detected based on the result keywords, to obtain target semantic features of the search term to be detected; a quality feature extraction unit, configured to extract quality features of the search term to be detected based on the obtained at least one quality evaluation value, and obtain a multi-dimensional quality feature corresponding to the search term to be detected; The quality identification unit is used to perform quality identification on the search term to be detected based on the target semantic feature and the multi-dimensional quality feature to obtain a quality inspection result of the search term to be detected.

13. An electronic device, characterized in that, The method comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is enabled to perform the steps of any one of the methods of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The method comprises a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of any one of the methods of claims 1 to 11.

15. A computer program product, characterized in that, It includes a computer program, which is stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device performs the steps of any method described in claims 1 to 11.