Method, apparatus, device and medium for processing user query content
By acquiring user query content and attribute data, and processing it using a knowledge-based scoring database, the system identifies users' knowledge needs, thus solving the problem of the narrow recognition range of knowledge graphs and enabling faster identification of knowledge needs and commercial profitability.
Patent Information
- Application Number
- CN202111164771.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In existing technologies, knowledge graphs have a narrow scope when identifying knowledge needs and cannot effectively identify knowledge points that do not appear in the knowledge graph, resulting in a slow speed of discovering new knowledge needs.
By acquiring users' query content and attribute data, and using a knowledge-based scoring database to process the query content and user attributes, we can identify whether the query content represents a knowledge requirement.
It improves the scope and speed of knowledge demand identification, enabling faster identification of users' knowledge needs and supporting service providers in providing targeted knowledge supplementation and commercial profitability.
Smart Images

Figure CN113901314B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and more particularly to natural language processing technology, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing user query content. Background Technology
[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0003] With the gradual upgrading of the information age, knowledge is being updated and iterated at an unprecedented speed. The definition and structure of knowledge, as well as the ways netizens express their knowledge needs, are all entering a rapid iteration phase. However, the emergence of new knowledge requires service providers to quickly identify advancements in knowledge and provide timely, targeted content supplements and accurate content matching to offer users high-quality services.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing user query content.
[0006] According to one aspect of this disclosure, a method for processing user query content is provided, comprising: acquiring user input query content and user attribute data; obtaining a first score corresponding to the input query content using a knowledge score database associated with the user query content; obtaining a second score corresponding to the user attribute data using a knowledge score database associated with the user attributes; and identifying whether the input query content is a knowledge requirement based on the first score and the second score.
[0007] According to another aspect of this disclosure, an apparatus for processing user query content is provided, comprising: a first unit configured to acquire user input query content and user attribute data; a second unit configured to obtain a first score corresponding to the input query content using a knowledge score database associated with the user query content; a third unit configured to obtain a second score corresponding to the user attribute data using a knowledge score database associated with user attributes; and a fourth unit configured to identify whether the input query content is a knowledge requirement based on the first score and the second score.
[0008] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the method described above for processing user query content.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described method for processing user query content.
[0010] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, implements the above-described method for processing user query content.
[0011] According to embodiments of this disclosure, the user's query content is first obtained. Then, through a database storing relevant data identified as historical query content, identification is performed based on the user's input and the user's own attributes. Finally, based on preset rules, it is determined whether the current user's query content constitutes a knowledge requirement. Thus, service providers can obtain feedback data using the above method, quickly identify the current market demand for specific knowledge, provide targeted knowledge supplementation, and achieve commercial profit through business activities such as paid knowledge services.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;
[0015] Figure 2 A flowchart is shown illustrating a method for processing user query content according to an exemplary embodiment of the present disclosure;
[0016] Figure 3 An exemplary embodiment of the present disclosure is shown in Figure 2 An example process for obtaining the first score corresponding to the input query content in the method;
[0017] Figure 4 An exemplary embodiment of the present disclosure is shown in Figure 2 An example process for obtaining the second score corresponding to the user's attribute data in the method;
[0018] Figure 5 A flowchart is shown for a method of processing user query content according to another exemplary embodiment of the present disclosure;
[0019] Figure 6 An exemplary embodiment of the present disclosure is shown in Figure 5 The flowchart illustrates an example process of statistically analyzing the input query content in the method.
[0020] Figure 7 A structural block diagram of an apparatus for processing user query content according to an exemplary embodiment of the present disclosure is shown;
[0021] Figure 8 A structural block diagram of another apparatus for processing user query content according to an exemplary embodiment of the present disclosure is shown; and
[0022] Figure 9 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0024] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same example of that element, while in other cases, based on the context, they may refer to different examples.
[0025] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0026] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0027] In related technologies, after a knowledge graph containing relevant knowledge points is manually constructed, the user's query is matched against each node in the knowledge graph using either exact or fuzzy matching. Based on the matching results, it is determined whether the user's query constitutes a knowledge requirement. Understandably, discovering new knowledge requirements relies heavily on the pre-built knowledge graph. If a relevant node in the knowledge graph matches the user's query, the knowledge requirement can be identified. However, this method also means that knowledge points not present in the knowledge graph will not be identified as knowledge requirements. Therefore, this technology suffers from a narrow scope for knowledge point identification and a slow speed in discovering new knowledge requirements.
[0028] To address the aforementioned problems, this disclosure provides a method for processing user query content. This method acquires the user's query content, processes the query content and the user's attribute data based on a pre-established knowledge scoring database, and finally identifies whether the user's query content constitutes a knowledge need. This method can alleviate, reduce, or even eliminate the aforementioned problems.
[0029] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0030] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0031] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of methods for processing user query content.
[0032] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.
[0033] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0034] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to display the user's query content page and retrieve the user's query content. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0035] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0036] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0037] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0038] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0039] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0040] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0041] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The data repository 130 may reside in various locations. For example, a data repository used by server 120 may be local to server 120, or it may be located remotely to server 120 and may communicate with server 120 via a network-based or dedicated connection. The data repository 130 may be of different types. In some embodiments, the data repository used by server 120 may be a database, such as a relational database. One or more of these databases may store, update, and retrieve data from and from the database in response to commands.
[0042] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0043] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0044] Figure 2 This is a flowchart illustrating a method 200 for processing user query content according to an exemplary embodiment of the present disclosure. Method 200 can be applied to... Figure 1 Server 120. Method 200 may include the following steps.
[0045] In step 201, the user's input query content and the user's attribute data are obtained.
[0046] According to some embodiments, user attribute data includes the user's unique identification information and location information.
[0047] In step 202, the first score corresponding to the input query content is obtained using a knowledge-based scoring database associated with the user's query content.
[0048] According to some embodiments, the sentence structure knowledge score database maintains multiple sentence structure templates and multiple sentence structure knowledge scores corresponding to the multiple sentence structure templates.
[0049] According to some embodiments, before executing method 200, establishing a sentence structure knowledge score database includes collecting user query inputs that are identified as knowledge needs in historical data, extracting sentence structure features of the inputs based on a natural language processing model, assigning a knowledge score to all obtained sentence structure features, and finally constructing a sentence structure knowledge score database by combining all sentence structures with their corresponding knowledge scores.
[0050] In one example, the sentence structure knowledge scoring database could include sentence structures such as patent number "CNXXXXXXXA", patent title "A method for XXX", and journal article title "Research on the filtering effect based on wavelet transform", etc. Furthermore, the database could also contain knowledge scores corresponding to all of the above sentence structures.
[0051] According to some embodiments, the vocabulary knowledge score database maintains multiple content fragment templates and multiple vocabulary knowledge scores corresponding to each of the multiple content fragment templates.
[0052] According to some embodiments, a vocabulary knowledge score database can be established before performing method 200. For example, the vocabulary knowledge score database can be established through the following process: collecting user query inputs identified as knowledge needs in historical data, extracting the input vocabulary based on a natural language processing model, assigning a knowledge score to each of the obtained vocabulary, and finally constructing a vocabulary knowledge score database by combining all the vocabulary and their corresponding knowledge scores.
[0053] In one example, the vocabulary knowledge score database could include words such as "way," "method," "practice," and "Fourier transform," etc. Similarly, the vocabulary knowledge score database would also include knowledge scores for all of the above words.
[0054] In step 203, a second score corresponding to the user's attribute data is obtained using a knowledge-based scoring database associated with user attributes.
[0055] According to some embodiments, the population knowledge score database maintains multiple user unique identifiers and multiple population knowledge scores corresponding to the multiple user unique identifiers.
[0056] According to some embodiments, a population knowledge score database can be established before executing method 200. For example, the population knowledge score database can be established through the following process: collecting attribute data of historical users, generating different user profiles for each historical user through a user profile model, assigning different knowledge scores to each user profile, and finally constructing a population knowledge score database by combining all user profiles with their corresponding knowledge scores.
[0057] According to some embodiments, in order to ensure that the collected historical users are unique, that is, one historical user can only correspond to one user profile data in the population knowledge score database, it is necessary to determine unique identification information for each historical user.
[0058] In one example, for historical user login operations on the web or mobile device, the account information is used as the unique user identifier. For historical user operations on the web without login, the web cookie information is used as the unique user identifier; for historical user operations on the mobile device without login, the mobile device's physical information is used as the unique user identifier.
[0059] In one example, if a user has both logged into their account and visited the web or mobile app as a guest (i.e., not logged into their account), their account information can be used as the unique user identifier.
[0060] In one example, historical users always access the web and mobile versions as guests, meaning they are not logged into an account. Therefore, the physical device information of the mobile device can be used as the unique user identification information.
[0061] According to some embodiments, the auxiliary knowledge score database maintains multiple location information and multiple auxiliary knowledge scores corresponding to the multiple location information.
[0062] According to some embodiments, an auxiliary knowledge score database may be established before executing method 200. For example, the auxiliary knowledge score database may be established by the following process: collecting location information of historical users when entering query content, assigning different knowledge scores to each location information, and finally constructing an auxiliary knowledge score database by combining all the location information with their corresponding knowledge scores.
[0063] In one example, the location information could be the IP address or point of interest of a historical user when they entered a query.
[0064] In step 204, based on the first score and the second score, it is determined whether the input query content is a knowledge requirement.
[0065] Figure 3 An exemplary embodiment of the present disclosure is shown in Figure 2 The example process of obtaining the first score corresponding to the input query content in method 200 (step 202). Step 202 may include the following steps.
[0066] In step 301, the input query content is analyzed to obtain at least one sentence structure of the input query content.
[0067] According to some embodiments, after obtaining input containing user query content, server 120 extracts the sentence structure features of user input through a natural language processing model.
[0068] In one example, the user's query could be a sentence like "a method for calculating surface integrals". The server 120 can extract the sentence features based on a natural language processing model, and finally obtain the sentence structure features of the user's query as "a method of XXX".
[0069] In step 302, the at least one sentence knowledge score corresponding to the at least one sentence template with the highest similarity to at least one sentence pattern is searched in the sentence pattern knowledge score database.
[0070] According to some embodiments, server 120 can use a natural language processing model to convert at least one sentence structure feature of the user's query content into at least one feature vector, and then convert sentence structure templates in the sentence structure knowledge score database into multiple feature vectors. Then, it calculates the Euclidean distance between each user's sentence structure feature vector and each sentence structure feature vector in the database, and determines multiple sentences in the database that have the highest similarity to the user's input, as well as the sentence structure knowledge scores corresponding to these multiple sentences.
[0071] In step 303, the input query content is sliced into words to obtain at least one content fragment of the input query content.
[0072] According to some embodiments, after obtaining the input containing the user's query content, the server 120 divides the user's query input into multiple content fragments through a natural language processing model.
[0073] In one example, the user's query could be a sentence like "What is the surface integral method?" The server 120 can use a natural language processing attention mechanism model to divide "surface integral method" into multiple content fragments such as "surface", "integral", "surface integral", "calculation", "method" and "what is".
[0074] In step 304, at least one vocabulary knowledge score is retrieved from the vocabulary knowledge score database for each content fragment template that has the highest similarity to at least one content fragment.
[0075] According to some embodiments, server 120 can use a natural language processing model to convert at least one content fragment feature of a user's query content into at least one feature vector, and then convert content fragment templates in a lexical knowledge score database into multiple content fragment feature vectors. Then, it calculates the Euclidean distance between the feature vector of at least one user's content fragment and each content fragment feature vector in the database, and determines multiple content fragments in the database that have the highest similarity to the user's input, along with the corresponding knowledge score for each content fragment.
[0076] In step 305, a weighted sum of at least one sentence structure knowledge score and a weighted sum of at least one vocabulary knowledge score are calculated.
[0077] In one example, a user's query might contain multiple sentence structures, such as both the patent title "A method for xxx" and the patent number "CNxxxxxxA". When building a sentence structure knowledge scoring database, it's necessary to assign and store a weight value for each sentence structure, indicating its importance within the query. Therefore, the sentence structure knowledge scoring database includes sentence templates and the corresponding weight and knowledge score for each sentence structure within those templates.
[0078] In one example, the user's input query contains at least one content fragment, such as multiple fragments like "surface," "integral," "surface integral," "calculate," "method," and "what." When building the vocabulary knowledge scoring database, it's also necessary to assign and store different weights for each content fragment, indicating its importance within the input query. In this case, the vocabulary knowledge scoring database includes content fragment templates and the corresponding weight and knowledge score for each content fragment within those templates.
[0079] According to some implementations, it is necessary to consider that multiple sentence structures in the user's input query content are not completely identical to the sentence structures in the sentence structure knowledge scoring database. That is, the feature vector of each user's query sentence structure has at least one similarity value with the sentence structure feature vector of the sentence structure knowledge scoring database.
[0080] In some implementations, it is also necessary to consider that multiple content fragments in the user's input query are not entirely identical to content fragments in the vocabulary knowledge scoring database. That is, the feature vector of each user's query content fragment must have at least one similarity value with the feature vector of the content fragments in the vocabulary knowledge scoring database.
[0081] In one example, the knowledge score, weight value, and similarity value of the sentence with the highest similarity to the user input are combined from multiple sentence knowledge score databases.
[0082] For example, the weighted sum of sentence structure knowledge scores = (N1 sentence structure score * N1 sentence structure knowledge requirement weight * similarity + N2 sentence structure score * N2 sentence structure knowledge requirement weight * similarity ... Nn sentence structure score * Nn sentence structure knowledge requirement weight * similarity), thus obtaining the weighted sum of sentence structure knowledge scores.
[0083] In one example, the knowledge score, weight value, and similarity value of the content segment with the highest similarity to the user input are combined from multiple vocabulary knowledge score databases.
[0084] According to the formula, the weighted sum of vocabulary knowledge scores = (N1 slice score * N1 slice knowledge requirement weight * similarity + N2 slice score * N2 slice knowledge requirement weight * similarity + ... Nn slice score * Nn slice knowledge requirement weight * similarity), thus obtaining the weighted sum of vocabulary knowledge scores.
[0085] In step 306, a first score is calculated based on the weighted sum of at least one sentence structure knowledge score and the weighted sum of at least one vocabulary knowledge score.
[0086] In one example, the first score = the weighted sum of sentence structure knowledge scores + the weighted sum of vocabulary knowledge scores.
[0087] Figure 4 An exemplary embodiment of the present disclosure is shown in Figure 2 The example process of obtaining the second score corresponding to the user's attribute data in method 200 (step 203). Step 203 may include the following steps.
[0088] In step 401, the user's unique identifier information is searched for in the population knowledge score database to obtain the corresponding population knowledge score.
[0089] According to some embodiments, a user's unique identifier can be account information, cookie information, or mobile device physical information. The system searches a population knowledge score database for identifiers that match the attribute data of the user containing the unique identifier. Then, using the identifiers in the population knowledge score database, a corresponding user profile is determined, and consequently, the corresponding population knowledge score is determined.
[0090] In step 402, the auxiliary knowledge score corresponding to the user's location information is searched in the auxiliary knowledge score database.
[0091] According to some embodiments, the IP address or POI associated with the user's attribute data, including location information, is searched in the auxiliary knowledge scoring database. The corresponding auxiliary knowledge score is then obtained through the location information in the auxiliary knowledge scoring database.
[0092] In step 402, a second score is calculated based on the population knowledge score and the auxiliary knowledge score.
[0093] In one example, the second score = population knowledge score + auxiliary knowledge score.
[0094] Figure 5 A flowchart of a method 500 for processing user query content according to another exemplary embodiment of the present disclosure is shown. Method 500 may include the following steps.
[0095] In step 505, it is determined whether the sum of the first score and the sum of the second score are greater than a threshold.
[0096] According to some implementation examples, the first score + the second score = the weighted sum of sentence structure knowledge scores + the weighted sum of vocabulary knowledge scores + population knowledge scores + auxiliary knowledge scores.
[0097] Steps 501 to 504 are related to the above. Figure 2 Steps 201 to 204 are described in the same way, and for the sake of brevity, they will not be repeated.
[0098] According to some embodiments, if the sum of the first score and the second score is greater than a threshold, then step 506 is executed. If the sum of the first score and the second score is less than the threshold, then step 507 is executed.
[0099] In step 506, in response to determining that the input query content is identified as a knowledge requirement, statistical analysis is performed on the input query content. Step 506 will be discussed in subsequent steps. Figure 6 Detailed description.
[0100] In step 507, the input query content is identified as a non-knowledge requirement.
[0101] Figure 6 An exemplary embodiment of the present disclosure is shown in Figure 5 The flowchart shows an example process (step 506) for performing statistical analysis on the input query content in method 500. Step 506 includes the following steps.
[0102] In step 601, the input query content is sliced into words to obtain at least one content fragment of the input query content.
[0103] In step 602, each content segment in at least one content segment is assigned a corresponding weight, which indicates the importance of the content slice in the input query content.
[0104] According to some embodiments, steps 601 to 602 can be the same as described above. Figure 2 Steps 303 to 304 are basically the same, the difference being that the input query content in step 601 has been determined as a knowledge requirement, but the input query content in steps 303 and 304 may not necessarily be a knowledge requirement.
[0105] In step 603, the number of historical searches for at least one content segment is counted.
[0106] In step 604, a weighted sum of the corresponding historical retrieval counts of at least one content segment and the corresponding weight of at least one content segment is calculated.
[0107] In one example, step 604 can be used to obtain the user value score of the user query content identified as a knowledge need. For example, the user value score = N1 slice retrieval volume * N1 slice knowledge need weight + N2 slice retrieval volume * N2 slice knowledge need weight + ... Nn slice retrieval volume * Nn slice knowledge need weight.
[0108] According to some embodiments, the historical access metrics of the knowledge carrier corresponding to each content segment in at least one content segment can also be counted to obtain the commercial value score of user query content identified as knowledge demand.
[0109] In one example, historical access metrics can be at least one of three metrics: download counts, citation counts, and purchase counts for books, patents, papers, standards, journals, and informally published documents. By statistically analyzing the historical access metrics for each content segment of user queries identified as knowledge needs, a commercial value score for the user's query content can be obtained.
[0110] Figure 7 A structural block diagram of an apparatus 700 for processing user query content according to an exemplary embodiment of the present disclosure is shown. Figure 7 As shown, the device 700 includes: a first unit 701 configured to acquire user input query content and user attribute data; a second unit 702 configured to obtain a first score corresponding to the input query content using a knowledge score database associated with the user query content; a third unit 703 configured to obtain a second score corresponding to the user attribute data using a knowledge score database associated with user attributes; and a fourth unit 704 configured to identify whether the input query content is a knowledge requirement based on the first score and the second score.
[0111] Figure 8 A structural block diagram of an apparatus 700 for processing user query content according to an exemplary embodiment of the present disclosure is shown. Figure 8 As shown, the device 800 includes: a first unit 801, a second unit 802, a third unit 803, a fourth unit 804, and a fifth unit 805. The fifth unit 805 is configured to perform statistical analysis on the input query content in response to determining that the input query content is identified as a knowledge requirement. The first unit 801, the second unit 802, the third unit 803, and the fourth unit 804 can be connected with... Figure 7 The first unit 701, the second unit 702, the third unit 703 and the fourth unit 704 are the same, and will not be described in detail here.
[0112] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0113] refer to Figure 9 The present invention describes a structural block diagram of an electronic device 900 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0114] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0115] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, output unit 907, storage unit 908, and communication unit 909. Input unit 906 can be any type of device capable of inputting information to device 900. Input unit 906 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 907 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 908 may include, but is not limited to, a hard disk and an optical disk. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, 1302.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0116] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as methods for processing user query content. For example, in some embodiments, the method for processing query content may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the methods for processing user query content described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured in any other suitable manner (e.g., by means of firmware) to perform the text recognition method and the training method of the text detection network model.
[0117] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0118] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0119] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0121] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0122] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0123] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0124] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A method for processing user query content, comprising: Obtain the user's input query content and the user's attribute data; The first score corresponding to the input query content is obtained by using a knowledge-based scoring database associated with the user's query content; A second score corresponding to the user's attribute data is obtained using a knowledge-based scoring database associated with user attributes; as well as Based on the first score and the second score, determine whether the input query content is a knowledge requirement; The knowledge-based scoring database associated with the user's query content includes: A sentence structure knowledge scoring database, which maintains multiple sentence structure templates and multiple sentence structure knowledge scores corresponding to each of the multiple sentence structure templates; and A vocabulary knowledge score database, which maintains multiple content fragment templates and multiple vocabulary knowledge scores corresponding to each of the multiple content fragment templates; Furthermore, the knowledge-based scoring database associated with user attributes includes: A population knowledge score database, which maintains multiple user unique identifiers and multiple population knowledge scores corresponding to each user unique identifier; and An auxiliary knowledge-based scoring database is maintained, which contains multiple location information and multiple auxiliary knowledge scores corresponding to each of the multiple location information.
2. The method as described in claim 1, wherein, The first score obtained corresponding to the input query content includes: Perform sentence structure analysis on the input query content to obtain at least one sentence structure of the input query content; Search the sentence knowledge score database for at least one sentence knowledge score corresponding to the at least one sentence template with the highest similarity to the at least one sentence pattern; The input query content is segmented into words to obtain at least one content fragment of the input query content; Search the vocabulary knowledge score database for at least one vocabulary knowledge score corresponding to the at least one content segment template that has the highest similarity to the at least one content segment; Calculate the weighted sum of the knowledge scores for at least one sentence structure and the weighted sum of the knowledge scores for at least one vocabulary word; and The first score is calculated based on the weighted sum of the at least one sentence structure knowledge score and the weighted sum of the at least one vocabulary knowledge score.
3. The method as described in claim 1, wherein, The user's attribute data includes the user's unique identifier and location information, and wherein obtaining the second score corresponding to the user's attribute data includes: Search the user's unique identifier for the user's knowledge score in the knowledge score database; Search the auxiliary knowledge score corresponding to the user's location information in the auxiliary knowledge score database; and The second score is calculated based on the population knowledge score and the auxiliary knowledge score.
4. The method of claim 1, wherein, The process of identifying whether the input query content is a knowledge requirement includes: Calculate the sum of the first score and the second score; and In response to determining that the sum of the first score and the second score is greater than a threshold, the input query content is identified as a knowledge requirement.
5. The method according to any one of claims 1-4, further comprising: In response to determining that the input query content is identified as a knowledge requirement, statistical analysis is performed on the input query content.
6. The method of claim 5, wherein, The statistical analysis of the input query content includes: The input query content is segmented into words to obtain at least one content fragment of the input query content; Each of the at least one content fragments is assigned a corresponding weight, the weight indicating the importance of the content fragment in the input query content; Count the number of historical searches for the at least one content segment; and Calculate the weighted sum of the corresponding historical retrieval counts of the at least one content segment and the corresponding weight of the at least one content segment.
7. The method of claim 6, wherein, The statistical analysis of the input query content also includes: Statistical analysis of historical access metrics for the knowledge carrier corresponding to each content segment in the at least one content segment.
8. The method of claim 7, wherein, The knowledge carriers include at least one of the following: books, patents, papers, standards, journals, and informally published documents, and the historical access metrics include at least one of the following: number of downloads, number of citations, and number of purchases.
9. An apparatus for processing user query content, comprising: The first unit is configured to obtain the user's input query content and the user's attribute data; The second unit is configured to obtain a first score corresponding to the input query content using a knowledge-based scoring database associated with the user's query content; The third unit is configured to obtain a second score corresponding to the user's attribute data using a knowledge-based scoring database associated with user attributes; as well as The fourth unit is configured to identify whether the input query content is a knowledge requirement based on the first score and the second score; The knowledge-based scoring database associated with the user's query content includes: A sentence structure knowledge scoring database, which maintains multiple sentence structure templates and multiple sentence structure knowledge scores corresponding to each of the multiple sentence structure templates; and A vocabulary knowledge score database, which maintains multiple content fragment templates and multiple vocabulary knowledge scores corresponding to each of the multiple content fragment templates; Furthermore, the knowledge-based scoring database associated with user attributes includes: A population knowledge score database, which maintains multiple user unique identifiers and multiple population knowledge scores corresponding to each of the multiple user unique identifiers; and An auxiliary knowledge-based scoring database is maintained, which contains multiple location information and multiple auxiliary knowledge scores corresponding to each of the multiple location information.
10. The apparatus of claim 9, wherein, The second unit includes: The first subunit is configured to perform sentence structure analysis on the input query content to obtain at least one sentence structure of the input query content; The second subunit is configured to search the sentence knowledge scoring database for at least one sentence knowledge score corresponding to the at least one sentence template that has the highest similarity to the at least one sentence pattern. The third subunit is configured to perform word segmentation on the input query content to obtain at least one content fragment of the input query content; The fourth subunit is configured to search the vocabulary knowledge score database for at least one vocabulary knowledge score corresponding to the corresponding content segment template that has the highest similarity to the at least one content segment. The fifth subunit is configured to calculate a weighted sum of the at least one sentence structure knowledge score and a weighted sum of the at least one vocabulary knowledge score; and The sixth subunit is configured to calculate the first score based on a weighted sum of the at least one sentence structure knowledge score and a weighted sum of the at least one vocabulary knowledge score.
11. The apparatus of claim 9, wherein, The user's attribute data includes the user's unique identifier and location information, and the third unit includes: The seventh subunit is configured to look up the user's unique identifier information in the population knowledge score database. The eighth subunit is configured to search for the auxiliary knowledge score corresponding to the user's location information in the auxiliary knowledge score database; And a ninth subunit, configured to calculate the second score based on the population knowledge score and the auxiliary knowledge score.
12. The apparatus of claim 9, wherein, The fourth unit includes: The tenth subunit is configured to calculate the sum of the first score and the second score; and The eleventh subunit is configured to identify the input query as a knowledge requirement in response to determining that the sum of the first score and the second score is greater than a threshold.
13. The apparatus of any one of claims 9-12, further comprising: The fifth unit is configured to perform statistical analysis on the input query content in response to determining that the input query content is identified as a knowledge requirement.
14. The apparatus of claim 13, wherein, The fifth unit includes: The twelfth subunit is configured to perform word slicing on the input query content to obtain at least one content fragment of the input query content; The thirteenth subunit is configured to assign a corresponding weight to each of the at least one content fragment, the weight indicating the importance of the content fragment in the input query content; The fourteenth subunit is configured to count the corresponding historical retrieval counts for the at least one content segment; and The fifteenth subunit is configured to calculate a weighted sum of the corresponding historical retrieval counts of the at least one content segment and the corresponding weight of the at least one content segment.
15. The apparatus of claim 14, wherein, The fifth unit also includes: The sixteenth subunit is configured to count the historical access metrics of the knowledge carrier corresponding to each of the at least one content segment.
16. The apparatus of claim 15, wherein, The knowledge carriers include at least one of the following: books, patents, papers, standards, journals, and informally published documents, and the historical access metrics include at least one of the following: number of downloads, number of citations, and number of purchases.
17. An electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
19. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-8.
Citation Information
Patent Citations
Intention recognition method and device, computer readable medium and electronic equipment
CN110069709A
Man-machine conversation processing method and device
CN111708869A