Dynamic Abstract Determination Method and Apparatus, Computing Device, and Computer Storage Medium

By filtering keywords not included in the document title part and traversing the document body part, and determining part of the dynamic summary is solved, the problem of inaccurate dynamic summary in the prior art is solved, and the effect of presenting search content-related information is achieved more accurately.

CN113761125BActive Publication Date: 2025-06-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110577211.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-26
Publication Date
2025-06-03
Estimated Expiration
2041-05-26

AI Technical Summary

Technical Problem

When determining dynamic summary in the prior art, some keywords in the search content often appear repeatedly in the document title and dynamic summary, while other keywords do not appear in the document title and dynamic summary, resulting in the dynamic summary being inaccurate enough to effectively present information related to the overall search content.

Method used

By obtaining the current document searched based on the search content, extracting multiple keywords in the search content, filtering keywords not included in the document title part as the first keyword set, extracting keywords for each sentence in the document body part to form a second keyword set, traversing the sentences in the text part, determining the similarity between the first keyword set and the second keyword set of each sentence, and determining part of the dynamic summary in response to the similarity greater than the threshold.

Benefits of technology

Improve the accuracy of dynamic summary, so that document titles and dynamic summary can be sufficient to present information related to the overall search content, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113761125B_ABST
    Figure CN113761125B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for determining a dynamic abstract, a computing device, and a computer storage medium. The method includes: obtaining a current document searched based on search content, where the current document includes a title part and a body part; extracting a plurality of keywords from the search content; screening, from the plurality of keywords, the keywords not included in the title part of the current document as a first keyword set; extracting keywords for each sentence in the body part of the current document to correspondingly form a second keyword set for each sentence; traversing the sentences in the body part to determine the similarity between the first keyword set and the second keyword set of the traversed sentence; and in response to the similarity being greater than a similarity threshold, determining a part of the dynamic abstract for the current document based on the traversed sentence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of natural language processing, and in particular, to a method and apparatus for determining a dynamic summary, a computing device, and a computer storage medium. Background Art

[0002] With the development of computer technology, dynamic summaries are widely used, such as in fields like search result summary display, document key sentence marking, and display of content related to search content. As an example, for different search contents, the same document can have different dynamic summaries. Currently, in conventional methods for determining dynamic summaries, it is usually determined which sentences in a document should be used as the dynamic summary of the document based on the number of keywords of the search content contained in each sentence of the document.

[0003] However, in the dynamic summaries determined by conventional methods for determining dynamic summaries, some keywords in the search content often appear repeatedly in the document title and the dynamic summary, but some other keywords in the search content do not appear in either the document title or the dynamic summary. This makes the determined dynamic summary inaccurate and thus insufficient to present information related to the overall search content, and even seems to be far from the true query intention expressed by the search content. Summary of the Invention

[0004] In view of this, the present disclosure provides a method and apparatus for determining a dynamic summary, a computing device, and a computer storage medium, and it is expected to overcome some or all of the above-mentioned defects and other possible defects.

[0005] According to a first aspect of the present disclosure, there is provided a method for determining a dynamic summary, including: obtaining a current document searched based on a search content, the current document including a title part and a body part; extracting a plurality of keywords from the search content; screening, from the plurality of keywords, keywords not included in the title part of the current document as a first keyword set; extracting keywords for each sentence in the body part of the current document to correspondingly form a second keyword set for each sentence; traversing the sentences in the body part to determine a similarity between the first keyword set and the second keyword set of the traversed sentence; and in response to the similarity being greater than a similarity threshold, determining a part of the dynamic summary for the current document based on the traversed sentence.

[0006] In some embodiments, traversing the sentences in the body part and determining the similarity between the first keyword set and the second keyword set of the traversed sentence includes: determining the word vectors of the keywords in the first keyword set; determining the first feature vector of the first keyword set based on the word vectors of the keywords in the first keyword set; determining the word vectors of the keywords in the second keyword set of the traversed sentence; determining the second feature vector of the second keyword set based on the word vectors of the keywords in the second keyword set; and determining the similarity between the first keyword set and the second keyword set of the traversed sentence based on the first feature vector and the second feature vector.

[0007] In some embodiments, determining the first feature vector of the first keyword set based on the word vectors of the keywords in the first keyword set includes: performing bitwise accumulation on the word vectors of the keywords in the first keyword set to obtain the first feature vector of the first keyword set, and determining the second feature vector of the second keyword set based on the word vectors of the keywords in the second keyword set includes: performing bitwise accumulation on the word vectors of the keywords in the second keyword set to obtain the second feature vector of the second keyword set.

[0008] In some embodiments, determining the similarity between the first keyword set and the second keyword set of the traversed sentence based on the first feature vector and the second feature vector includes: determining the similarity between the first keyword set and the second keyword set of the traversed sentence based on the distance between the first feature vector and the second feature vector, where the distance includes one of cosine distance, Euclidean distance, and Manhattan distance.

[0009] In some embodiments, traversing the sentences in the body part and determining the similarity between the first keyword set and the second keyword set of the traversed sentence includes: traversing the sentences in the body part, and when the current number of words in the dynamic summary is less than the word count threshold, determining the similarity between the first keyword set and the second keyword set of the traversed sentence.

[0010] In some embodiments, in response to the similarity being greater than the similarity threshold, determining a part of the dynamic summary for the current document based on the traversed sentence includes: in response to the similarity being greater than the similarity threshold and the sum of the traversed sentence and the current number of words in the dynamic summary being greater than the word count threshold, determining a part of the traversed sentence as a part of the dynamic summary for the current document, such that the sum of the number of words in the part of the traversed sentence and the current number of words in the dynamic summary is equal to the word count threshold.

[0011] In some embodiments, extracting multiple keywords from the search content includes: segmenting the search content to obtain a first set of segmented words including multiple words; removing stop words from the multiple words in the first set of segmented words to obtain a second set of segmented words; determining the word weights of each word in the second set of segmented words; and removing words with word weights less than a word weight threshold from the second set of segmented words to obtain the multiple keywords in the search content.

[0012] In some embodiments, determining the word weight of each word in the second set of segmented words includes: determining the inverse document frequency value of each word in the second set of segmented words; and determining the inverse document frequency value of each word in the second set of segmented words as the word weight of that word.

[0013] In some embodiments, determining the word weight of each word in the second set of segmented words includes: determining the inverse document frequency value of each word in the second set of segmented words; and determining the word weight of that word based on at least one of the part of speech of each word in the second set of segmented words, the word position in the search content, the historical search times, and the historical click-through rate, and its inverse document frequency value.

[0014] In some embodiments, determining the inverse document frequency value of each word in the second set of segmented words includes: obtaining a query log, where the query log includes D search contents; for each corresponding word in the second set of segmented words, determining the number d of search contents in the query log that contain the corresponding word; determining the quotient of the total number D of search contents included in the query log and the number d of search contents in the query log that contain the corresponding word, and taking the logarithm of the quotient to obtain the inverse document frequency value of the corresponding segmented word.

[0015] In some embodiments, determining the word vectors of the keywords in the first keyword set includes: determining the word vectors of the keywords in the first keyword set based on a trained word embedding model, and wherein the trained word embedding model is trained in the following manner: obtaining a query log and segmenting the search contents in the query log to obtain multiple segmented words; training the word embedding model by using each corresponding segmented word in the multiple segmented words as the input of the word embedding model and using the context segmented words of the corresponding segmented word as the output of the word embedding model, or by using each corresponding segmented word in the multiple segmented words as the output of the word embedding model and using the context segmented words of the corresponding segmented word as the input of the word embedding model, to obtain the trained word embedding model.

[0016] According to a second aspect of the present disclosure, there is provided a dynamic abstract determination device, including: a current document acquisition module configured to acquire a current document searched based on search content, the current document including a title part and a body part; a first keyword extraction module configured to extract a plurality of keywords from the search content; a first keyword set determination module configured to screen out keywords not included in the title part of the current document from the plurality of keywords as a first keyword set; a second keyword extraction module configured to extract keywords from each sentence of the body part of the current document to correspondingly form a second keyword set for each sentence; a similarity determination module configured to traverse the sentences in the body part and determine the similarity between the first keyword set and the second keyword set of the traversed sentence; a dynamic abstract determination module configured to, in response to the similarity being greater than a similarity threshold, determine a part of the dynamic abstract for the current document based on the traversed sentence.

[0017] According to a third aspect of the present disclosure, there is provided a computing device, including: a memory configured to store computer-executable instructions; a processor configured to execute the method according to any one of claims 1-11 when the computer-executable instructions are executed by the processor.

[0018] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium storing computer-executable instructions, which when executed, execute any of the methods described above.

[0019] In the dynamic abstract determination method and device, computing device, and computer storage medium claimed by the present disclosure, by considering the keywords of the search content already included in the title of the current document, these keywords already included in the title of the current document are no longer considered when determining the dynamic abstract, avoiding the repeated occurrence of certain keywords in the document title and the dynamic abstract. Then, by traversing each sentence of the document body and comparing the similarity between the keywords of each sentence and the set of keywords of the search content not included in the title, it is efficiently determined which sentences are used as part of the dynamic abstract. Since the hit rate of the overall search content in the document title and the dynamic abstract is considered when determining the dynamic abstract, while avoiding the repeated occurrence of certain keywords in the document title and the dynamic abstract, the accuracy of the determined dynamic abstract is improved, making the document title and the dynamic abstract sufficient to present information related to the overall search content, thereby enhancing the user experience.

[0020] According to the embodiments described below, these and other advantages of the present disclosure will become clear, and these and other advantages of the present disclosure will be clarified with reference to the embodiments described below. Description of the Drawings

[0021] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings, in which:

[0022] Figure 1 An exemplary application scenario in which the technical solutions according to the embodiments of the present disclosure can be implemented is shown;

[0023] Figure 2 A schematic flowchart of a method for determining a dynamic abstract according to an embodiment of the present disclosure is illustrated;

[0024] Figure 3 A schematic flowchart of a method for determining the similarity between two keyword sets involved in the present disclosure according to an embodiment of the present disclosure is illustrated;

[0025] Figure 4 A schematic flowchart of a method for extracting multiple keywords from search content according to an embodiment of the present disclosure is illustrated;

[0026] Figure 5 An exemplary specific principle framework diagram of a word embedding model according to an embodiment of the present disclosure is illustrated;

[0027] Figure 6 A schematic effect diagram of a dynamic abstract determined using related technologies is illustrated;

[0028] Figure 7 A schematic effect diagram of a dynamic abstract determined using the method for determining a dynamic abstract according to an embodiment of the present disclosure is illustrated;

[0029] Figure 8 A schematic structural block diagram of a device for determining a dynamic abstract according to an embodiment of the present disclosure is illustrated;

[0030] Figure 9 An example system is illustrated, which includes an example computing device representing one or more systems and / or devices that can implement the various technologies described herein. Detailed implementation manners

[0031] The following description provides specific details of various embodiments of the present disclosure so that those skilled in the art can fully understand and implement the various embodiments of the present disclosure. It should be understood that the technical solutions of the present disclosure can be implemented without some of these details. In some cases, the present disclosure does not show or describe in detail some well-known structures or functions to avoid obscuring the description of the embodiments of the present disclosure with these unnecessary descriptions. The terms used in the present disclosure should be understood in the broadest reasonable manner, even if they are used in combination with specific embodiments of the present disclosure.

[0032] First, some terms involved in the embodiments of the present application are described to facilitate the understanding of those skilled in the art.

[0033] Dynamic abstract: A search engine term, which is a technology for dynamically displaying the main content of the retrieved document. For a search engine, in response to a user inputting search content, relevant text around the search content in the document is extracted according to the position where the search content appears in the document and returned as a dynamic abstract. Since a document may be recalled by different search contents, the dynamic abstract technology may form different dynamic abstracts for the same document according to different search contents.

[0034] Search content: That is, the meaning of query. It is a word, sentence or any appropriate content input by the user into the search engine in order to find a specific file, website, record or a series of records in the database for retrieving data from the database.

[0035] Query log: A log is a type of diary used to record the work done every day. In computer science, a log refers to the operation record of computer devices or software such as servers (Server log). When there are problems with computer devices and software, the log is an important basis for us to troubleshoot problems. The query log is used to record information related to the search content input by the user received from the client.

[0036] Stop Words: Refers to high-frequency words that do not carry any topic information, such as words like "de", "ye", "le". In information retrieval, to save storage space and improve search efficiency, it is best to filter out these words when processing natural language data (or text). Stop words are all manually input and not automatically generated. After generation, the stop words will form a stop word list. When filtering stop words later, the stop words in the document can be confirmed by querying the stop word list.

[0037] Tokenizer: It is a tool that analyzes a piece of text input by the user into a logical form. Common tokenizers include English tokenizers and Chinese tokenizers, etc. The tokenization process of an English tokenizer is generally: input text - keyword segmentation - stop word removal - morphological reduction - conversion to lowercase. A Chinese tokenizer splits a sequence of Chinese characters into individual words. In other words, tokenization is the process of recombining a continuous sequence of characters into a sequence of words according to certain specifications. In this process, stop words that do not affect the semantics can be identified. Commonly used tokenizers such as jieba tokenizer, Mmseg4j tokenizer, Ansj tokenizer, etc.

[0038] Natural Language Processing: Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies. Natural language processing is mainly applied to machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, Chinese OCR, and other aspects.

[0039] Word Embedding: Word embedding is a type of word representation where words with similar meanings have similar representations, and it is a general term for methods that map words to real-valued vectors. Conceptually, it means embedding a high-dimensional space with the dimension of the total number of words into a much lower-dimensional continuous vector space, and each word or phrase is mapped to a vector in the real number domain. Common word embeddings include Word2Vec, fastText, GloVe, etc. Word2Vec is a group of related models used to generate word vectors. These models are shallow and two-layer neural networks used to train and reconstruct linguistic word texts. After the Word2Vec model is trained, it can be used to map each word to a vector, which is the hidden layer of the neural network. The full name of GloVe is Global Vectors for Word Representation, which is a word representation tool based on global word frequency statistics (count-based & overall statistics). It can represent a word as a vector composed of real numbers, and these vectors capture some semantic features between words, such as similarity and analogy. fastText is a fast text classification algorithm that has two major advantages compared to neural network-based classification algorithms: it speeds up the training and testing speeds while maintaining high accuracy and does not require pre-trained word vectors.

[0040] Inverse document frequency: Inverse document frequency is a commonly used weighting technique in information retrieval and information mining. It is a statistical method used to evaluate the importance of a word for a document set or a single document in a corpus. The importance of a word increases proportionally with the number of times it appears in a document, but decreases inversely with the frequency of its appearance in the corpus. Various forms of inverse document frequency weighting are often applied by search engines as a measure or rating of the relevance between a document and a user query.

[0041] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0042] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, convolutional neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.

[0043] Deep learning (DL) is a new research direction in the field of machine learning (ML). It is introduced into machine learning to make it closer to the original goal - artificial intelligence (AI). Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analytical learning like humans, and be able to recognize data such as text, images, and sounds.

[0044] In the technical solution provided in this application, it involves natural language processing technology and mainly involves dynamic summary technology.

[0045] Figure 1 FIG. illustrates an exemplary application scenario 100 in which the technical solution according to an embodiment of the present disclosure can be implemented. As Figure 1 shown, the illustrated application scenario includes a terminal 110 and a server 120, and the terminal 110 is communicatively coupled to the server 120 via a network 130.

[0046] As an example, the terminal 110 can be used as an input device to input search content and send the search content to the server 120 via the network 130. The search content can be, for example, search content for searching relevant documents.

[0047] As an example, the server 120 can, for example, obtain the current document searched based on the search content, extract multiple keywords from the search content, and screen out the keywords that are not included in the title part of the current document as the first keyword set; then, the server 120 can extract keywords from each sentence in the body part of the current document to correspondingly form a second keyword set for each sentence; then, the server 120 can traverse the sentences in the body part to determine the similarity between the first keyword set and the second keyword set of the traversed sentence; finally, in response to the similarity being greater than the similarity threshold, the server 120 can determine a part of the dynamic summary for the current document based on the traversed sentence. The server 120 can send the determined dynamic summary to the terminal 110 for presentation.

[0048] The scenario described above is only one example in which the embodiments of the present disclosure can be implemented, and is not restrictive. For example, in some exemplary scenarios, the dynamic summary determination process may also be implemented on the terminal 110.

[0049] For example, the terminal 110 can be used as an input device to input search content, and save the current document obtained by searching based on the search content to the background of the terminal. Then, multiple keywords in the search content are extracted, and the keywords that are not included in the title part of the current document are screened out from the multiple keywords as the first keyword set; then, the terminal 110 can extract keywords for each sentence in the body part of the current document to correspondingly form a second keyword set for each sentence; then, the terminal 110 can traverse the sentences in the body part to determine the similarity between the first keyword set and the second keyword set of the traversed sentence; finally, in response to the similarity being greater than the similarity threshold, the terminal 110 can determine a part of the dynamic abstract for the current document based on the traversed sentence.

[0050] It should be noted that the terminal 110 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The server 120 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. The network 130 can be, for example, a wide area network (WAN), a local area network (LAN), a wireless network, a public telephone network, an intranet, and any other type of network well-known to those skilled in the art.

[0051] In some embodiments, the above application scenario 100 can be a distributed system composed of a cluster of terminals 110 and a server 120. The distributed system can, for example, form a blockchain system. For example, in the application scenario 100, the determination and storage of the dynamic abstract can both be carried out in the blockchain system to achieve the effect of decentralization. As an example, after determining the dynamic abstract, the dynamic abstract can be stored in the blockchain system for subsequent retrieval of the dynamic abstract from the blockchain system when performing the same search. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain is essentially a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0052] Figure 2The figure illustrates a schematic flowchart of a dynamic abstract determination method 200 according to an embodiment of the present disclosure. The dynamic abstract determination method can be implemented, for example, by a terminal 110 or a server 120 as shown in Figure 1 as follows. As shown in Figure 2 the method 200 includes the following steps.

[0053] In step 210, a current document searched based on search content is obtained. The current document includes a title part and a body part. The search content typically may include multiple words to more clearly characterize the user's query intention. The search engine can retrieve multiple relevant documents for the search content as search results of the search content. In an embodiment of the present disclosure, a relevant document can be obtained from the multiple relevant documents searched based on the search content as the current document.

[0054] In step 220, multiple keywords in the search content are extracted. In an embodiment of the present disclosure, various technical means can be used to extract the multiple keywords in the search content, which is not restrictive. For example, the various tokenizers described above can be used to tokenize the search content, and then multiple keywords in the search content are extracted from the tokens according to the importance of each token.

[0055] In step 230, keywords not included in the title part of the current document are screened out from the multiple keywords as a first keyword set. As an example, if the keywords "automobile", "maintenance", "standard", and "manual" are extracted from the search content, and the title of the current document as the search content result contains "automobile" and "maintenance", then the first keyword set will contain "standard" and "manual". At this time, the dynamic abstract is confirmed according to the first keyword set, and the dynamic abstract will contain one or both of "standard" and "manual", that is, the dynamic abstract and the title will contain 3 / 4 or 4 / 4 of the search content keywords, and the hit rate of the search content as a whole in the title and the dynamic abstract is 75%-100%. In contrast, in the conventional related art, the dynamic abstract is determined according to the keywords of the search content. The dynamic abstract is very likely to contain one or both of "automobile" and "maintenance", and does not contain "standard" and "manual". At this time, the dynamic abstract and the title will contain 1 / 4 or 2 / 4 of the search content keywords, and the hit rate of the search content as a whole in the title and the dynamic abstract is 25%-50%.

[0056] Since keywords already included in the title of the current document are no longer considered when determining the first keyword set, keywords already included in the title of the current document are no longer considered in subsequent steps of dynamic summary determination, avoiding the repeated appearance of certain keywords in the document title and dynamic summary while other keywords do not appear in either the document title or the dynamic summary, and improving the hit rate of the overall search content in the article title and dynamic summary. Therefore, step 230 takes into account the hit rate of the overall search content in the article title and dynamic summary, making the document title and the dynamic summary extracted by the method of the present disclosure sufficient to present information related to the overall search content, improving the accuracy of the determined dynamic summary, and enabling the document title and the dynamic summary to present information related to the overall search content, thus enhancing the user experience.

[0057] In step 240, keywords are extracted from each sentence in the body part of the current document to correspondingly form a second keyword set for each sentence. In an embodiment of the present disclosure, various technical means can be used to extract keywords from each sentence in the body part of the current document, which is not restrictive.

[0058] In some embodiments, the method of extracting keywords from each sentence in the body part of the current document can be the same as the method of extracting multiple keywords from the search content in step 220. As an example, the steps of extracting keywords from the sentence "The benefits of Internet technology are obvious" in the current document can be as follows: First, the sentence is segmented to obtain a first segmentation set including "Internet", "technology", "of", "benefits", "is", "obvious", "of"; then, stop words "of" and "is" are removed from the first segmentation set to obtain a second segmentation set; then, the word weight of each word in the second segmentation set is determined (which can be calculated according to historical occurrence times, for example), and here the word weights of "Internet", "technology", "benefits", and "obvious" are determined to be 0.7, 0.5, 0.6, and 0.2 respectively; finally, words with word weights less than a predetermined threshold are removed from the second segmentation set. For example, if the predetermined threshold is set to 0.4 here, then "obvious" is removed, and finally the keywords of this sentence are determined to be "Internet", "technology", and "benefits".

[0059] In step 250, the sentences in the body part are traversed to determine the similarity between the first keyword set and the second keyword set of the traversed sentence. Various different methods can be used to determine the similarity between the first keyword set and the second keyword set of the traversed sentence, which is not limited here.

[0060] As an example, the first keyword set includes "jet", "airplane", and "cost", and the second keyword set of the sentence traversed includes "airplane", "airport", "cost", and "hill". In some embodiments, the similarity is determined by comparing the ratio of the number of words simultaneously included in the first keyword set and the second keyword set to the number of words included in the first keyword set. For example, here, the first keyword set and the second keyword set simultaneously include 2 words: "airplane" and "cost", and the number of words included in the first keyword set is 4, so the similarity is 2÷4 = 0.5. In some other embodiments, based on the word vectors corresponding to the words in the first keyword set and the word vectors corresponding to the words in the second keyword set, a similarity matrix is determined, and the similarity is extracted from the similarity matrix.

[0061] In some embodiments, traversing the sentences in the body part and determining the similarity between the first keyword set and the second keyword set of the traversed sentence may include: traversing the sentences in the body part, and when the current number of words in the dynamic summary is less than the word count threshold, determining the similarity between the first keyword set and the second keyword set of the traversed sentence. As an example, the word count threshold of the dynamic summary is 100 words. Traversing the sentences in the body part, if the current number of words in the dynamic summary is less than the word count threshold, for example, the current number of words in the dynamic summary is 90 words, then determine the similarity between the first keyword set and the second keyword set of the traversed sentence; if the current number of words in the dynamic summary is not less than the word count threshold, for example, the current number of words in the dynamic summary is 100 words, then do not determine the similarity between the first keyword set and the second keyword set of the traversed sentence.

[0062] In step 260, in response to the similarity being greater than the similarity threshold, determine a part of the dynamic summary for the current document based on the traversed sentence. The similarity threshold can be preset as needed and is not restrictive.

[0063] In some embodiments, when determining a part of the dynamic summary for the current document based on the traversed sentence, a part or all of the traversed sentence can be determined as a part of the dynamic summary for the current document. As an example, some phrases, the main part of the sentence, or the traversed sentence as a whole can be extracted and determined as a part of the dynamic summary for the current document. As the sentences in the body part are traversed and a part of the dynamic summary for the current document is determined based on the traversed sentence, a final dynamic summary can be formed.

[0064] In some embodiments, in response to the similarity being greater than a similarity threshold, determining a part of the dynamic summary for the current document based on the traversed sentence may include: in response to the similarity being greater than the similarity threshold and the sum of the traversed sentence and the current number of words in the dynamic summary being greater than a word count threshold, determining a part of the traversed sentence as a part of the dynamic summary for the current document, such that the sum of the number of words in the part of the traversed sentence and the current number of words in the dynamic summary is equal to the word count threshold. As an example, assume that the word count threshold for the dynamic summary is 100 words, the current number of words in the dynamic summary is 95 words, and the traversed sentence is 15 words. Then, 5 words of the traversed sentence (which can be words extracted from the head, tail, or middle of the sentence, etc., and are not limited here) are determined as a part of the dynamic summary.

[0065] The method 200, by considering the keywords of the search content already included in the title of the current document, no longer considers these keywords already included in the title of the current document when determining the dynamic summary, avoiding the repeated occurrence of certain keywords in the document title and the dynamic summary. Then, each sentence in the document body is traversed, and the similarity between the keywords of the sentence and the set of keywords of the search content not included in the title is compared. Whether the sentence belongs to a part of the dynamic summary is judged based on whether the similarity is greater than the similarity threshold. Since, when determining the dynamic summary, the hit rate of the overall search content in the article title and the dynamic summary is considered, while avoiding the repeated occurrence of certain keywords in the document title and the dynamic summary, the accuracy of the determined dynamic summary is improved, such that the document title and the dynamic summary are sufficient to present information related to the overall search content, thereby enhancing the user experience.

[0066] Figure 3 FIG. illustrates a schematic flowchart of a method 300 for determining the similarity between two sets of keywords according to an embodiment of the present disclosure. The two sets of keywords include the first set of keywords and the second set of keywords of the traversed sentence. The method 300 can be used, for example, to implement the steps 250 described with reference to Figure 2 As shown in Figure 3 The method 300 includes the following steps.

[0067] In step 310, determine the word vectors of each keyword in the first keyword set. The word vectors of each keyword in the first keyword set can be determined using a trained word embedding model. The trained word embedding model can be obtained by training the word embedding model with an open-source corpus (such as the open corpus provided by Google) as the training set, or by training the word embedding model with a specific corpus (such as a training set established with words in a certain field) as the training set, which is not limited here. The word embedding model can be a common word embedding model, such as Word2Vec for embedding, fastText for word and character embedding, GloVe for global word embedding, etc., which is not limited here.

[0068] In some embodiments, determining the word vectors of each keyword in the first keyword set includes: determining the word vectors of each keyword in the first keyword set based on a trained word embedding model, and wherein the trained word embedding model is trained in the following manner: Obtain a query log (the query log records information related to the search content input by the user received from the client), and segment the search content in the query log to obtain multiple segmented words; Train the word embedding model by using each corresponding segmented word in the multiple segmented words as the input of the word embedding model and using the context segmented words of the corresponding segmented word as the output of the word embedding model, or by using each corresponding segmented word in the multiple segmented words as the output of the word embedding model and using the context segmented words of the corresponding segmented word as the input of the word embedding model, to obtain the trained word embedding model. Since the segmented words used to train the word embedding model come from the query log, and the query log records information related to the search content input by the user received from the client, the trained word embedding model can better extract the features in the search content, and the determined word vectors can better represent each word in the search content.

[0069] In step 320, based on the word vectors of each keyword in the first keyword set, determine the first feature vector of the first keyword set.

[0070] In some embodiments, based on the word vectors of each keyword in the first keyword set, determining the first feature vector of the first keyword set includes: performing bitwise accumulation on the word vectors of each keyword in the first keyword set to obtain the first feature vector of the first keyword set. As an example, the first keyword set contains the words "noodles", "easy", and "sticking to the pot". After determining the 200-dimensional word vectors corresponding to these three words respectively, perform bitwise accumulation on the three 200-dimensional word vectors to obtain a 200-dimensional vector, that is, obtain the first feature vector of the first keyword set.

[0071] In step 330, determine the word vectors of the keywords in the second keyword set of the traversed sentence. The word vectors of the keywords in the second keyword set can be determined using the trained word embedding model described above or other suitable word embedding models. The word embedding model can be, for example, a common word embedding model, such as Word2Vec, fastText for word and character embedding, GloVe for global word embedding, etc., which is not limited here.

[0072] In step 340, based on the word vectors of the keywords in the second keyword set, determine the second feature vector of the second keyword set. In some embodiments, based on the word vectors of the keywords in the second keyword set, determining the second feature vector of the second keyword set includes: performing bitwise accumulation on the word vectors of the keywords in the second keyword set to obtain the second feature vector of the second keyword set. As an example, the second keyword set includes the words "handmade noodles", "hung noodles", "sliced noodles", and "huguo". After determining the 200-dimensional word vectors corresponding to these four words respectively, perform bitwise accumulation on the four 200-dimensional word vectors to obtain a 200-dimensional vector, that is, obtain the second feature vector of the second keyword set.

[0073] In step 350, based on the first feature vector and the second feature vector, determine the similarity between the first keyword set and the second keyword set of the traversed sentence. As an example, the first feature vector and the second feature vector are 200-dimensional vectors respectively. Calculate the cosine similarity between the first feature vector and the second feature vector, and use the cosine similarity as the similarity between the first keyword set and the second keyword set of the traversed sentence. The cosine similarity between vectors is used to measure text similarity, which depends on the cosine distance between vectors. The cosine distance maps vectors into a vector space according to coordinate values. The formula for the cosine distance between vectors a and b is as follows:

[0074]

[0075] Let the coordinates of vectors a and b in a two-dimensional space be , then the expression of the cosine distance between vectors a and b in a two-dimensional space is as follows:

[0076]

[0077] Let the coordinates of vectors a and b in an n-dimensional space be a = (A 1 , A 2,……, A n ), b = (B 1 , B 2,……, B n), the cosine distance between vectors a and b in the n-dimensional space is expressed as follows:

[0078] .

[0079] In some embodiments, based on the first feature vector and the second feature vector, determining the similarity between the first keyword set and the second keyword set of the traversed sentence may include: determining the similarity between the first keyword set and the second keyword set of the traversed sentence based on the distance between the first feature vector and the second feature vector, where the distance includes one of cosine distance, Euclidean distance, and Manhattan distance. As an example, the distance may be selected as the Euclidean distance, that is, calculating the Euclidean distance between the first feature vector and the second feature vector, and using the Euclidean distance as the similarity between the first keyword set and the second keyword set of the traversed sentence.

[0080] The method 300 determines the similarity between the first keyword set and the second keyword set of the traversed sentence by comparing the first feature vector based on the first keyword set and the second feature vector based on the second keyword set. The similarity extracted in this way can more accurately characterize the similarity between the first keyword set and the second keyword set of the traversed sentence, so as to determine the dynamic summary according to the similarity in subsequent steps.

[0081] Figure 4 FIG. illustrates a schematic flowchart of a method 400 for extracting multiple keywords in a search content according to an embodiment of the present disclosure. The method 400 may be used, for example, to implement the steps 220 described with reference to Figure 2 As shown in Figure 4 , the method 400 includes the following steps.

[0082] In step 410, the search content is tokenized to obtain a first token set including multiple words. In the embodiments of the present disclosure, various technical means may be used to tokenize the search content, which is not restrictive. For example, various tokenizers described above (such as jieba tokenizer, Mmseg4j tokenizer, Ansj tokenizer, etc.) may be used to tokenize the search content.

[0083] In step 420, stop words are removed from the multiple words in the first token set to obtain a second token set. As described above, stop words refer to high-frequency words that do not carry any topic information, such as words like "de", "ye", "le". Removing stop words here can save storage space and improve search efficiency, and reduce interference for determining the dynamic summary. In some embodiments, the stop words in the first token set can be confirmed and removed by querying a pre-constructed stop word table.

[0084] In step 430, the word weights of each word in the second word segmentation set are determined. In some embodiments, determining the word weights of each word in the second word segmentation set may include: determining the inverse document frequency value of each word in the second word segmentation set; and determining the inverse document frequency value of each word in the second word segmentation set as the word weight of that word. The inverse document frequency is used to evaluate the importance of each word in the second word segmentation set.

[0085] In some embodiments, determining the word weights of each word in the second word segmentation set may include: determining the inverse document frequency value of each word in the second word segmentation set; and determining the word weight of that word based on at least one of the part of speech of each word in the second word segmentation set, the word position in the search content, the historical search times, and the historical click-through rate, and its inverse document frequency value. Since the word weight here takes into account not only the inverse document frequency value of each word in the second word segmentation set, but also comprehensively considers at least one of the part of speech of each word in the second word segmentation set, the word position in the search content, the historical search times, and the historical click-through rate, the finally determined word weight can more comprehensively and accurately reflect the importance of each word in the second word segmentation set.

[0086] In some embodiments, determining the inverse document frequency value of each word in the second word segmentation set may include: obtaining a query log, where the query log includes D search contents; for each corresponding word in the second word segmentation set, determining the number d of search contents in the query log that contain the corresponding word; determining the quotient of the total number D of search contents included in the query log and the number d of search contents in the query log that contain the corresponding word, and taking the logarithm of the quotient to obtain the inverse document frequency value of the corresponding word segmentation. As an example, the query log includes 1000 search contents, and the number of search contents in the query log that contain the corresponding word is 600, that is, D = 1000, d = 300. Therefore, the inverse document frequency value of the corresponding word segmentation is log e (D / d)=log e (1000 / 300)=1.204.

[0087] As an example, the inverse document frequency can be determined by querying an inverse document frequency dictionary. The inverse document frequency dictionary can be calculated based on the query log. The inverse document frequency (idf i ) of a word t i in the inverse document frequency dictionary is calculated according to the following formula:

[0088]

[0089] where the numerator |D| represents the total number of search contents in the query log, and the denominator represents the number of search contents that contain the word t i .

[0090] In step 440, words with a word weight less than the word weight threshold are removed from the second word segmentation set to obtain multiple keywords in the search content. The word weight threshold can be set as needed and is not restrictive. As an example, the second word segmentation set contains the word segments "Sima Qian", "Records of the Grand Historian", "Western Han Dynasty", and "prestige", and the corresponding word weights of these four words are 0.6, 0.5, 0.7, and 0.2 respectively. If the word weight threshold is set to 0.4, then "prestige" will be removed, and finally, the multiple keywords in the search content are determined to be "Sima Qian", "Records of the Grand Historian", and "Western Han Dynasty".

[0091] The method 400 extracts multiple keywords in the search content by performing word segmentation on the search content, removing stop words, determining weights, and removing words with weights less than the weight threshold. The keywords extracted in this way can more accurately and concisely represent the meaning of the search content compared to the search content, facilitating the determination of the dynamic summary based on the meaning of the search content in subsequent steps.

[0092] Figure 5 An exemplary specific principle framework diagram of a word embedding model according to an embodiment of the present disclosure is shown. As Figure 5 shown, the word embedding model is used to embed a high-dimensional space with a dimension equal to the number of all words into a much lower-dimensional continuous vector space. Each word or phrase is mapped to a vector in the real number domain, and it can be a three-layer neural network, including an input layer, a hidden layer, and an output layer.

[0093] The input layer is used to receive the input vector, and the input vector is usually a one-hot vector; the hidden layer processes the input vector by setting nodes. For example, if 300 features are used to represent a word (i.e., each word can be represented as a 300-dimensional vector), then 300 nodes will be set in the hidden layer, and its weights can be represented as a matrix with several rows (the number of rows depends on the dimension of the input vector) and 300 columns; the output layer is used to process the output of the hidden layer to output a probability distribution at this layer. The output layer can be a softmax regression classifier, and each of its nodes will output a value between 0 and 1 (probability), and the sum of the probabilities of all output layer neuron nodes is 1.

[0094] When training the word embedding model, it is usually trained based on pairs of words. The training samples are word pairs of (input word, output word), that is, using the input word to predict the output word, and both the input word and the output word are one-hot encoded vectors. Typically, the input word and the output word are words with a context relationship (such as adjacent words), that is, using the context relationship of each word in the corpus to obtain the trained word embedding model.

[0095] Figure 6Illustrated is a schematic effect diagram of a dynamic abstract determined using related technologies. As Figure 6 shown, in related technologies, the existing hit situation in the title is not considered in the dynamic abstract determination scheme, that is, in terms of presenting relevant information as a whole, the global hit experience is not concerned. This is more obvious in the overall abstract hit problem when the search engine faces some search contents with a relatively large length (such as Figure 6 shown). As an example, when a user searches for "Luofu Mountain hiking route", neither in the title nor in the selected fragments of the dynamic abstract does the word "route" appear, only "Luofu Mountain" and "hiking" are included. This is because when traversing the sentences in the text to determine the dynamic abstract in related technologies, only the similarity between the traversed sentences and the search content is considered, without considering that the title of the current document already contains some content of the search content, resulting in some content in the search content being repeated in the title and dynamic abstract of the current document, but some content not appearing in either the title or the dynamic abstract of the current document.

[0096] As Figure 6 shown, related technologies do not consider the hit rate of the overall search content in the article title and dynamic abstract, resulting in some keywords in the search content ("Luofu Mountain", "hiking") appearing repeatedly in the document title and dynamic abstract, but some other keywords in the search content ("route") not appearing in either the document title or the dynamic abstract. This makes the document title and the dynamic abstract extracted by the related technologies insufficient to present information related to the overall search content. As an example, in this article, the hit effect of the keywords in the search content in the current document is reflected in bold font, and it can also be achieved by highlighting relevant text, underlining the keyword font, italicizing or enlarging the keyword font, or marking the keyword font with a color different from the text or title color, etc., which is not limited here.

[0097] Figure 7 Illustrated is a schematic effect diagram of a dynamic abstract determined using the dynamic abstract determination method according to an embodiment of the present disclosure. This embodiment is also for the dynamic abstract of the current document searched for with "Luofu Mountain hiking route", and the document title of the current document already contains "Luofu Mountain" and "hiking". As Figure 7 shown, when the system selects the fragments included in the body dynamic abstract, it knows that "Luofu Mountain" and "hiking" actually already exist in the title, and is more inclined to extract those text fragments in the body that at least contain "route" as dynamic abstract fragments and bold them. This enables the user to determine that the content of this article tends to be about "Luofu Mountain hiking route" rather than "Luofu Mountain hiking" without clicking to read the full text.

[0098] Compare Figure 6 and Figure 7It can be seen that, compared with the conventional dynamic abstract determination method, the dynamic abstract determination method proposed in the present disclosure, by considering the keywords of the search content already included in the title of the current document ("Luofu Mountain", "hiking"), no longer considers these keywords already included in the title of the current document when determining the dynamic abstract, avoiding the repeated appearance of certain keywords in the document title and the dynamic abstract, improving the hit rate of the overall search content in the article title (hitting "Luofu Mountain", "hiking") and the dynamic abstract (hitting "route"), so that the document title and the dynamic abstract extracted by the method of the present disclosure can present information related to the overall search content. While avoiding the repeated appearance of certain keywords in the document title and the dynamic abstract, this improves the accuracy of the determined dynamic abstract, making the document title and the dynamic abstract sufficient to present information related to the overall search content, thereby enhancing the user experience.

[0099] At the same time, as can be seen from Figure 7 it, the words hit in the text are not necessarily exactly the same as the words in the search content not included in the title, but the semantics are basically the same. For example, Figure 7 in it, according to the keyword "route" in the search content, "line" in the text is hit. Although "route" and "line" are two words, their semantics are basically the same. This is because the present disclosure adopts a method of comparing the similarity between the keywords of the sentence and the set of keywords of the search content not included in the title, rather than requiring the keywords of the sentence to be exactly the same as the keywords of the search content not included in the title. This makes the determination of the dynamic abstract by the method shown in the present disclosure have better robustness and accuracy.

[0100] Figure 8 FIG. shows an exemplary structural block diagram of a dynamic abstract determination device 800 according to an embodiment of the present disclosure. As Figure 8 shown, the dynamic abstract determination device includes: a current document acquisition module 810, a first keyword extraction module 820, a first keyword set determination module 830, a second keyword extraction module 840, a similarity determination module 850, and a dynamic abstract determination module 860.

[0101] The current document acquisition module 810 is configured to acquire the current document searched based on the search content, and the current document includes a title part and a body part.

[0102] The first keyword extraction module 820 is configured to extract multiple keywords from the search content.

[0103] The first keyword set determination module 830 is configured to screen the keywords not included in the title part of the current document from the multiple keywords as the first keyword set.

[0104] The second keyword extraction module 840 is configured to extract keywords from each sentence in the body part of the current document, so as to correspondingly form a second keyword set for each sentence.

[0105] The similarity determination module 850 is configured to traverse the sentences in the body part and determine the similarity between the first keyword set and the second keyword set of the traversed sentence.

[0106] The dynamic summary determination module 860 is configured to, in response to the similarity being greater than the similarity threshold, determine a part of the dynamic summary for the current document based on the traversed sentence.

[0107] Figure 9 An example system 900 is illustrated, which includes an example computing device 910 representing one or more systems and / or devices that can implement the various technologies described herein. The computing device 910 can be, for example, a server of a service provider, a device associated with the server, a system on a chip, and / or any other suitable computing device or computing system. Referenced above Figure 8 The dynamic summary determination device 800 described can take the form of the computing device 910. Alternatively, the dynamic summary determination device 800 can be implemented as a computer program in the form of an application 916.

[0108] As illustrated, the example computing device 910 includes a processing system 911, one or more computer-readable media 912, and one or more I / O interfaces 913 that are communicatively coupled to each other. Although not shown, the computing device 910 may further include a system bus or other data and command transfer systems that couple the various components to each other. The system bus can include any one or combination of different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any one of various bus architectures. Various other examples are also contemplated, such as control and data lines.

[0109] The processing system 911 represents the functionality of performing one or more operations using hardware. Thus, the processing system 911 is illustrated as including hardware elements 914 that can be configured as processors, functional blocks, etc. This can include being implemented in hardware as an application-specific integrated circuit or other logic devices formed using one or more semiconductors. The hardware elements 914 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can be composed of (multiple) semiconductors and / or transistors (e.g., an electronic integrated circuit (IC)). In such a context, the instructions executable by the processor can be electronically executable instructions.

[0110] The computer-readable medium 912 is illustrated as including a memory / storage 915. The memory / storage 915 represents the memory / storage capacity associated with one or more computer-readable media. The memory / storage 915 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical disks, magnetic disks, etc.). The memory / storage 915 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disk, etc.). The computer-readable medium 912 may be configured in various other ways as further described below.

[0111] One or more I / O interfaces 913 represent the functionality that allows a user to input commands and information into the computing device 910 using various input devices and optionally also allows information to be presented to the user and / or other components or devices using various output devices. Examples of input devices include a keyboard, a cursor control device (e.g., a mouse), a microphone (e.g., for voice input), a scanner, a touch functionality (e.g., a capacitive or other sensor configured to detect physical touch), a camera (e.g., that can detect motion not involving touch as a gesture using visible or invisible wavelengths such as infrared frequencies), and so on. Examples of output devices include a display device (e.g., a monitor or a projector), a speaker, a printer, a network card, a haptic response device, etc. Thus, the computing device 910 may be configured in various ways as further described below to support user interaction.

[0112] The computing device 910 further includes an application 916. The application 916 may be, for example, a software instance of the dynamic digest determination device 800 and, in combination with other elements in the computing device 910, implements the techniques described herein.

[0113] Various techniques may be described herein in the general context of software hardware elements or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The terms “module,” “function,” and “component” as used herein generally represent software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, meaning that these techniques may be implemented on various computing platforms having a variety of processors.

[0114] Implementations of the described modules and techniques may be stored on or transmitted across some form of computer-readable medium. The computer-readable medium may include various media accessible by the computing device 910. By way of example and not limitation, the computer-readable medium may include “computer-readable storage media” and “computer-readable signal media.”

[0115] In contrast to mere signal transmission, carriers, or signals themselves, a "computer-readable storage medium" refers to a medium and / or device capable of persistently storing information, and / or a tangible storage device. Thus, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented in a method or technology suitable for storing information such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data. Examples of computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical storage devices, hard disks, cassette tapes, magnetic tapes, magnetic disk storage devices, or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture suitable for storing the desired information and accessible by a computer.

[0116] A "computer-readable signal medium" refers to a signal-bearing medium configured to send instructions to a computing device 910, such as via a network. Signal media typically may embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave, data signal, or other transmission mechanism. Signal media also include any information delivery medium. The term "modulated data signal" refers to a signal in which one or more of the characteristics are set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0117] As previously described, hardware elements 914 and computer-readable media 912 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments may be used to implement at least some aspects of the techniques described herein. Hardware elements may include integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and components of other hardware devices implemented in silicon or other hardware. In this context, a hardware element may serve as a processing device for executing program tasks defined by the instructions, modules, and / or logic embodied by the hardware element, as well as a hardware device for storing instructions for execution, e.g., the computer-readable storage medium previously described.

[0118] The foregoing combinations can also be used to implement the various techniques and modules described herein. Thus, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic on a computer-readable storage medium of some form and / or embodied by one or more hardware elements 914. The computing device 910 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium of the processing system and / or the hardware element 914, the implementation of the module as a module executable by the computing device 910 as software can be implemented at least in part in hardware. The instructions and / or functions can be executable / operable by one or more articles of manufacture (e.g., one or more computing devices 910 and / or processing systems 911) to implement the techniques, modules, and examples described herein.

[0119] In various embodiments, the computing device 910 can assume a variety of different configurations. For example, the computing device 910 can be implemented as a computer-like device including a personal computer, a desktop computer, a multi-screen computer, a laptop computer, a netbook, etc. The computing device 910 can also be implemented as a mobile device-like device including mobile devices such as mobile phones, portable music players, portable gaming devices, tablet computers, multi-screen computers, etc. The computing device 910 can also be implemented as a television-like device, which includes a device having or connected to a generally larger screen in a casual viewing environment. These devices include televisions, set-top boxes, game consoles, etc.

[0120] The techniques described herein can be supported by these various configurations of the computing device 910 and are not limited to the specific examples of the techniques described herein. The functionality can also be implemented in whole or in part using a distributed system, such as on a "cloud" 920 via a platform 922 as described below.

[0121] The cloud 920 includes and / or represents a platform 922 for resources 924. The platform 922 abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud 920. The resources 924 can include applications and / or data that can be used when performing computer processing on servers remote from the computing device 910. The resources 924 can also include services provided via the Internet and / or via a subscriber network such as a cellular or Wi-Fi network.

[0122] The platform 922 can abstract resources and functionality to connect the computing device 910 with other computing devices. The platform 922 can also be used to abstract the hierarchy of resources to provide a corresponding level of hierarchy for the demands encountered for the resources 924 implemented via the platform 922. Thus, in an interconnected device embodiment, the implementation of the functions described herein can be distributed throughout the system 900. For example, the functionality can be implemented in part on the computing device 910 and via the platform 922 that abstracts the functionality of the cloud 920.

[0123] The present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computing device executes the dynamic summary determination method provided in the above various alternative implementation manners.

[0124] It should be understood that, for clarity, embodiments of the present disclosure have been described with reference to different functional units. However, it will be apparent that, without departing from the present disclosure, the functionality of each functional unit can be implemented in a single unit, implemented in multiple units, or implemented as part of other functional units. For example, the functionality described as being performed by a single unit can be performed by multiple different units. Therefore, the reference to a specific functional unit is only considered as a reference to an appropriate unit for providing the described functionality, rather than indicating a strict logical or physical structure or organization. Thus, the present disclosure can be implemented in a single unit, or can be physically and functionally distributed among different units and circuits.

[0125] It will be understood that although the terms first, second, third, etc. may be used herein to describe various devices, elements, components or parts, these devices, elements, components or parts should not be limited by these terms. These terms are only used to distinguish one device, element, component or part from another device, element, component or part.

[0126] Although the present disclosure has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. On the contrary, the scope of the present disclosure is only limited by the appended claims. Additionally, although individual features may be included in different claims, these may possibly be advantageously combined, and including in different claims does not imply that a combination of features is not feasible and / or advantageous. The order of features in the claims does not imply that the features must be in any particular order in which they work. Further, in the claims, the word "comprising" does not exclude other elements, and the terms "a" or "an" do not exclude a plurality. The reference numerals in the claims are provided only as illustrative examples and should not be construed as limiting the scope of the claims in any way.

Claims

1. A method for determining a dynamic abstract, comprising: obtaining a current document searched based on search content, the current document including a title part and a body part; extracting a plurality of keywords from the search content; screening out the keywords not included in the title part of the current document from the plurality of keywords as a first keyword set; extracting keywords for each sentence in the body part of the current document to correspondingly form a second keyword set for each sentence; traversing the sentences in the body part to determine the similarity between the first keyword set and the second keyword set of the traversed sentence; in response to the similarity being greater than a similarity threshold, determining a part of the dynamic abstract for the current document based on the traversed sentence.

2. The method according to claim 1, wherein traversing the sentences in the body part to determine the similarity between the first keyword set and the second keyword set of the traversed sentence, comprises: determining word vectors of each keyword in the first keyword set; determining a first feature vector of the first keyword set based on the word vectors of each keyword in the first keyword set; determining word vectors of each keyword in the second keyword set of the traversed sentence; determining a second feature vector of the second keyword set based on the word vectors of each keyword in the second keyword set; determining the similarity between the first keyword set and the second keyword set of the traversed sentence based on the first feature vector and the second feature vector.

3. The method according to claim 2, wherein determining a first feature vector of the first keyword set based on the word vectors of each keyword in the first keyword set, comprises: performing bitwise accumulation on the word vectors of each keyword in the first keyword set to obtain a first feature vector of the first keyword set, and wherein, determining a second feature vector of the second keyword set based on the word vectors of each keyword in the second keyword set, comprises: performing bitwise accumulation on the word vectors of each keyword in the second keyword set to obtain a second feature vector of the second keyword set.

4. The method according to claim 2, wherein determining the similarity between the first keyword set and the second keyword set of the traversed sentence based on the first feature vector and the second feature vector, comprises: determining the similarity between the first keyword set and the second keyword set of the traversed sentence based on the distance between the first feature vector and the second feature vector, wherein the distance includes one of a cosine distance, an Euclidean distance, and a Manhattan distance.

5. The method according to claim 1, wherein traversing the sentences in the body part to determine the similarity between the first keyword set and the second keyword set of the traversed sentence, comprises: traversing the sentences in the body part, and when the current number of words in the dynamic abstract is less than a word number threshold, determining the similarity between the first keyword set and the second keyword set of the traversed sentence.

6. The method according to claim 1, wherein in response to the similarity being greater than a similarity threshold, a part of the dynamic summary for the current document is determined based on the traversed sentence, including: in response to the similarity being greater than the similarity threshold and the sum of the traversed sentence and the current number of words in the dynamic summary being greater than a word count threshold, a part of the traversed sentence is determined as a part of the dynamic summary for the current document, such that the sum of the number of words in the part of the traversed sentence and the current number of words in the dynamic summary is equal to the word count threshold.

7. The method according to claim 1, wherein a plurality of keywords in the search content are extracted, including: performing word segmentation on the search content to obtain a first word segmentation set including a plurality of words; removing stop words from the plurality of words in the first word segmentation set to obtain a second word segmentation set; determining the word weight of each word in the second word segmentation set; removing words with word weights less than a word weight threshold from the second word segmentation set to obtain the plurality of keywords in the search content.

8. The method according to claim 7, wherein determining the word weight of each word in the second word segmentation set, including: determining the inverse document frequency value of each word in the second word segmentation set; determining the inverse document frequency value of each word in the second word segmentation set as the word weight of the word.

9. The method according to claim 7, wherein determining the word weight of each word in the second word segmentation set, including: determining the inverse document frequency value of each word in the second word segmentation set; determining the word weight of the word based on at least one of the part of speech of each word in the second word segmentation set, the word position in the search content, the historical search times, and the historical click-through rate, and its inverse document frequency value.

10. The method according to claim 8 or 9, wherein determining the inverse document frequency value of each word in the second word segmentation set including: obtaining a query log, the query log including D search contents; for each corresponding word in the second word segmentation set, determining the number d of search contents in the query log that contain the corresponding word; determining the quotient of the total number D of search contents included in the query log and the number d of search contents in the query log that contain the corresponding word, and taking the logarithm of the quotient to obtain the inverse document frequency value of the corresponding word.

11. The method according to claim 2, wherein, determining the word vectors of the keywords in the first keyword set includes: determining the word vectors of the keywords in the first keyword set based on a trained word embedding model, and wherein the trained word embedding model is trained through the following method: obtaining a query log and performing word segmentation on the search contents in the query log to obtain a plurality of word segments; training the word embedding model by using each corresponding word segment in the plurality of word segments as the input of the word embedding model and using the context word segments of the corresponding word segment as the output of the word embedding model, or by using each corresponding word segment in the plurality of word segments as the output of the word embedding model and using the context word segments of the corresponding word segment as the input of the word embedding model, to obtain the trained word embedding model.

12. A dynamic summary determination device, including: A current document acquisition module, configured to acquire a current document searched based on search content, where the current document includes a title part and a body part; A first keyword extraction module, configured to extract a plurality of keywords from the search content; A first keyword set determination module, configured to screen out keywords not included in the title part of the current document from the plurality of keywords as a first keyword set; A second keyword extraction module, configured to extract keywords from each sentence in the body part of the current document to correspondingly form a second keyword set for each sentence; A similarity determination module, configured to traverse the sentences in the body part and determine the similarity between the first keyword set and the second keyword set of the traversed sentence; A dynamic abstract determination module, configured to, in response to the similarity being greater than a similarity threshold, determine a part of the dynamic abstract for the current document based on the traversed sentence.

13. A computing device, comprising: a memory configured to store computer-executable instructions; a processor configured to execute the method according to any one of claims 1-11 when the computer-executable instructions are executed by the processor.

14. A computer-readable storage medium storing computer-executable instructions that, when executed, execute the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Search engine, client thereof and method for searching page

    CN101661490A

  • Abstraction generation method and device, terminal equipment and storage medium

    CN110837556A