Knowledge vector library construction method, information processing method and electronic equipment

By building a knowledge vector library with multi-level category structure, the problem of quickly finding relevant information in information retrieval is solved, and efficient and accurate information retrieval and convenient user operations are achieved.

CN120106196APending Publication Date: 2025-06-06LENOVO (BEIJING) LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510205787.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the field of information retrieval, how to quickly find output information related to user input information from complicated information, improve the accuracy of information retrieval and the convenience of user operations.

Method used

By constructing a knowledge vector library, multiple knowledge information are obtained, and they are converted into knowledge vectors. The first and second elements of the knowledge vector of each knowledge information are determined, and clustered to form a knowledge vector library with a multi-level category structure.

Benefits of technology

It realizes rapid positioning of search results, improves the efficiency and accuracy of information retrieval, and reduces the complexity of user operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106196A_ABST
    Figure CN120106196A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge vector library construction method, an information processing method and electronic equipment, and relates to the technical field of information processing. The knowledge vector library construction method comprises the steps of obtaining multiple pieces of knowledge information, and converting each piece of knowledge information into a knowledge vector; a first element and a second element of the knowledge vector of each piece of knowledge information are determined, and the semantic feature level of the first element is lower than that of the second element; performing clustering processing on the plurality of first elements of the plurality of pieces of knowledge information to obtain a plurality of parent categories; for each parent category, performing clustering processing on a plurality of second elements of the knowledge vector corresponding to the parent category to obtain a plurality of subcategories; and forming a knowledge vector library according to the plurality of knowledge vectors of the plurality of pieces of knowledge information and the parent category and the subcategory to which each knowledge vector in the plurality of knowledge vectors belongs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of information processing technology, and in particular to a knowledge vector library construction method, an information processing method and an electronic device. Background Art

[0002] At present, in the field of information retrieval, information databases usually contain a lot of complicated information. However, how to quickly find output information related to user input information from these complicated information and improve the accuracy of information retrieval and the convenience of user operation has become a problem that still needs to be improved. Summary of the invention

[0003] In view of the above problems, the present disclosure provides a method for constructing a knowledge vector library, an information processing method and an electronic device.

[0004] According to the first aspect of the present disclosure, a method for constructing a knowledge vector library is provided, including: acquiring multiple knowledge information, converting each knowledge information into a knowledge vector; determining a first element and a second element of the knowledge vector of each knowledge information, wherein the semantic feature level of the first element is lower than the semantic feature level of the second element; clustering multiple first elements of the multiple knowledge information to obtain multiple parent categories; for each parent category, clustering multiple second elements of the knowledge vector corresponding to the parent category to obtain multiple subcategories; forming a knowledge vector library according to multiple knowledge vectors of the multiple knowledge information and the parent category and subcategory to which each knowledge vector in the multiple knowledge vectors belongs.

[0005] According to an embodiment of the present disclosure, the semantic feature level includes semantic condensation; determining the first element and the second element of the knowledge vector of each knowledge information includes: for the knowledge vector of each knowledge information, using a large model to summarize the knowledge vector according to the preset first semantic condensation and second semantic condensation, respectively, to obtain the first element and the second element of the knowledge vector, wherein the first semantic condensation is lower than the second semantic condensation.

[0006] According to an embodiment of the present disclosure, a knowledge vector library is formed based on multiple knowledge vectors of multiple knowledge information and the parent category and subcategory to which each knowledge vector in the multiple knowledge vectors belongs, including: for any target category in the multiple parent categories and the multiple subcategories, using a large model to perform semantic induction on the cluster center of the target category to obtain the category label of the target category; based on multiple knowledge vectors of multiple knowledge information and the category label of the parent category and the category label of the subcategory to which each knowledge vector in the multiple knowledge vectors belongs, a knowledge vector library is constructed.

[0007] According to an embodiment of the present disclosure, a knowledge vector library is constructed based on multiple knowledge vectors of multiple knowledge information and the category label of the parent category and the category label of the subcategory to which each knowledge vector in the multiple knowledge vectors belongs, including: sorting the multiple parent categories according to the number of first elements contained in each parent category, and selecting some parent categories with top sorting; screening some knowledge vectors belonging to some parent categories from the multiple knowledge vectors of the multiple knowledge information; and constructing the knowledge vector library based on the partial knowledge vectors and the category label of the parent category and the category label of the subcategory to which each knowledge vector in the partial knowledge vectors belongs.

[0008] The second aspect of the present disclosure provides an information processing method, including: obtaining input information and converting the input information into an input vector; determining a target parent category that matches the input vector and a plurality of first knowledge vectors corresponding to the target parent category based on a pre-constructed knowledge vector library; determining a target subcategory that matches the input vector and at least one second knowledge vector corresponding to the target subcategory from the plurality of first knowledge vectors, wherein the knowledge vector library stores a plurality of knowledge vectors and a parent category and a subcategory to which each of the plurality of knowledge vectors belongs, and a semantic feature level of the parent category is lower than a semantic feature level of the subcategory; and generating output information corresponding to the input information based on the at least one second knowledge vector.

[0009] According to an embodiment of the present disclosure, determining a target parent category that matches an input vector includes: determining the similarity between the input vector and the parent category to which each knowledge vector in the knowledge vector library belongs; and determining the parent category in the knowledge vector library whose similarity is higher than a preset threshold as the target parent category.

[0010] According to an embodiment of the present disclosure, output information corresponding to input information is generated based on at least one second knowledge vector, including: summarizing at least one second knowledge vector using a large model to generate corresponding knowledge information in a target domain; and determining the knowledge information in the target domain as output information.

[0011] According to an embodiment of the present disclosure, the information processing method also includes: determining a first input element and a second input element of an input vector, wherein a semantic feature level of the first input element is higher than a semantic feature level of the second input element; determining from a knowledge vector library a target parent category that matches the first input element, and a plurality of first knowledge vectors corresponding to the target parent category; determining from the plurality of first knowledge vectors a target subcategory that matches the second input element, and at least one second knowledge vector corresponding to the target subcategory.

[0012] The third aspect of the present disclosure provides a system for constructing a knowledge vector library, including: a knowledge information library, storing multiple knowledge information; a knowledge vector library; a construction module, which is communicated and connected with the knowledge information library and the knowledge vector library respectively, and is used to: obtain multiple knowledge information and convert each knowledge information into a knowledge vector; determine the first element and the second element of the knowledge vector of each knowledge information, wherein the semantic feature level of the first element is lower than the semantic feature level of the second element; cluster the multiple first elements of the multiple knowledge information to obtain multiple parent categories; for each parent category, cluster the multiple second elements of the knowledge vector corresponding to the parent category to obtain multiple subcategories; store the multiple knowledge vectors of the multiple knowledge information and the parent category and subcategory to which each knowledge vector in the multiple knowledge vectors belongs in the knowledge vector library.

[0013] The fourth aspect of the present disclosure provides an electronic device, including: an information input module, used to obtain input information; a memory, storing a knowledge vector library, the knowledge vector library storing multiple knowledge vectors and a parent category and a subcategory to which each of the multiple knowledge vectors belongs, the semantic feature level of the parent category being lower than the semantic feature level of the subcategory; a processor, communicatively connected to the information input module and the memory, respectively, and used to: convert the input information into an input vector; call the knowledge vector library, and determine from the knowledge vector library a target parent category that matches the input vector, and a plurality of first knowledge vectors corresponding to the target parent category; determine from the plurality of first knowledge vectors a target subcategory that matches the input vector, and at least one second knowledge vector corresponding to the target subcategory; and output output information corresponding to the input information based on the at least one second knowledge vector.

[0014] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0016] Figure 1 A diagram schematically illustrates an application scenario of a method for constructing a knowledge vector library and an information processing method according to an embodiment of the present disclosure;

[0017] Figure 2 A flowchart schematically shows a method for constructing a knowledge vector library according to an embodiment of the present disclosure;

[0018] Figure 3 Schematically shows a flow chart for determining a first element and a second element according to an embodiment of the present disclosure;

[0019] Figure 4A A flowchart of clustering the first element according to an embodiment of the present disclosure is schematically shown;

[0020] Figure 4B A flowchart of clustering multiple second elements according to an embodiment of the present disclosure is schematically shown;

[0021] Figure 5 Schematically shows a flow chart of forming a knowledge vector library according to an embodiment of the present disclosure;

[0022] Figure 6 The flowchart of the information processing method according to the embodiment of the present disclosure is schematically shown;

[0023] Figure 7 The schematic diagram shows a principle diagram of an information processing method according to an embodiment of the present disclosure;

[0024] Figure 8 A schematic diagram of determining a second knowledge vector according to an embodiment of the present disclosure is shown;

[0025] Fig. 9 A block diagram of a system for constructing a knowledge vector library according to an embodiment of the present disclosure is schematically shown;

[0026] Fig.10 A block diagram of an electronic device according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0029] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0031] Figure 1 The application scenario diagram of the knowledge vector library construction method and the information processing method according to the embodiment of the present disclosure is schematically shown.

[0032] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0033] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only examples).

[0034] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0035] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0036] It should be noted that the knowledge vector library construction method and information processing method provided in the embodiments of the present disclosure can generally be executed by the server 105. The knowledge vector library construction method and information processing method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to the implementation requirements.

[0038] The following will be based on Figure 1 The scene described by Figure 2~Figure 5 The method for constructing the knowledge vector library of the embodiment of the present disclosure is described in detail.

[0039] Figure 2 The flowchart of the method for constructing a knowledge vector library according to an embodiment of the present disclosure is schematically shown.

[0040] like Figure 2 As shown, the method for constructing a knowledge vector library in this embodiment includes operations S210 to S250. The method can be executed by a server.

[0041] In operation S210, a plurality of pieces of knowledge information are acquired, and each piece of knowledge information is converted into a knowledge vector.

[0042] In the embodiments of the present disclosure, knowledge information is a combination of systematic information that can be used to solve problems. For example, knowledge information can be multimodal information, which includes multiple types of data such as images, text, audio, and video. Knowledge information can be a text file, a word document, an excel spreadsheet, a PDF document, a picture file, an audio file, a video file, etc. For example, knowledge information can be fixed or updated in real time, and the present disclosure does not specifically limit this.

[0043] Each piece of knowledge information can be converted into a knowledge vector by embedding each word in each piece of knowledge information. The word embedding method may include WordEmbedding or Word2Vec. Each word in each piece of knowledge information can be converted into a vector representation of a fixed length, and the knowledge vector of each piece of knowledge information can be a collection of vector representations of each word. Therefore, the knowledge vector of each piece of knowledge information is essentially a vector combination of vector representations of multiple words constituting the knowledge information. To avoid redundancy, the following is uniformly described as a knowledge vector.

[0044] Taking text information as knowledge information as an example, the text information can be converted into a knowledge vector by first segmenting the text information through a tokenizer to obtain multiple words (tokens), and then converting each word into a fixed-dimensional word embedding vector through a text encoder (such as TokenEmbedding) as the knowledge vector corresponding to the text information.

[0045] Taking image information as an example, image information can be converted into knowledge vectors by extracting features such as color, texture, and shape of image information, and then converting these features into knowledge vectors. For complex and fuzzy image information, deep learning technology can also be used to learn image features by training a large amount of image data and converting image features into knowledge vectors.

[0046] Taking audio information as an example, the audio information can be converted into a knowledge vector by decomposing the audio information into multiple local features, such as Mel-frequency cepstral coefficients (MFCC) and spectrograms, and extracting vector representations of these features. The vector representation of the audio file can also be generated by performing a series of convolution operations on the spectrogram.

[0047] In operation S220, a first element and a second element of a knowledge vector of each piece of knowledge information are determined, wherein a semantic feature level of the first element is lower than a semantic feature level of the second element.

[0048] For each piece of knowledge information among the plurality of pieces of knowledge information, the knowledge information may be first converted into a knowledge vector, and then a first element and a second element of the knowledge vector may be determined.

[0049] In some embodiments, the first element and the second element may be partial vectors extracted from the knowledge vector of each knowledge information. In other embodiments, the first element and the second element may also be new vectors generated based on the knowledge vector of each knowledge information. It should be understood that the first element and the second element are both in vector form, so that the computer can understand and process them. To avoid redundancy, the first element and the second element are described below.

[0050] Semantic features are the value and meaning contained in the knowledge vector. The semantic feature level of the first element is lower than that of the second element. It can be understood that the semantic features contained in the first element are more complicated and redundant, while the semantic features contained in the second element are more concise, accurate and significant. Compared with the first element, the second element can accurately reflect the core viewpoint of knowledge information, remove redundant information, and express the most comprehensive content in the most concise and clear language.

[0051] For example, the first element is the complete content of the knowledge vector, and the second element is the summary and refinement of the knowledge vector. For another example, the first element is a relatively low degree of summary and refinement of the knowledge vector, and the second element may be a relatively high degree of summary and refinement of the knowledge vector.

[0052] Taking the knowledge information as text information as an example, after converting the text information into a knowledge vector, the summary content and title content of the knowledge vector can be extracted as the first element and the second element. The body content and summary content of the knowledge vector can also be extracted as the first element and the second element.

[0053] In operation S230, clustering is performed on the first elements of the plurality of pieces of knowledge information to obtain a plurality of parent categories.

[0054] Since multiple pieces of knowledge information are obtained and the first element of the knowledge vector of each piece of knowledge information is determined, multiple first elements of multiple pieces of knowledge information can be obtained, and all first elements are clustered to obtain multiple parent categories. Each first element can determine the corresponding parent category, and each parent category contains at least one first element.

[0055] In the embodiments of the present disclosure, both the parent category and the subcategory may be predefined. For example, before acquiring a plurality of knowledge information, a plurality of parent categories and a plurality of subcategories subordinate to each parent category may be defined.

[0056] The clustering process of this operation adopts an unsupervised clustering algorithm, and its purpose is to classify multiple first elements into different parent categories according to their similarities or distances, so that at least one first element in the same parent category has a higher similarity, while the first elements between different parent categories have a lower similarity.

[0057] Taking the summary content of the knowledge vector in which the knowledge information is text information and the first element is the text information as an example, the summary contents of the multiple knowledge vectors of the multiple text information obtained can be clustered, and the multiple summary contents can be divided into different technical fields. The divided technical fields are also used as parent categories, and different technical fields form multiple parent categories.

[0058] Exemplarily, the parent category may be a plurality of high-tech fields, including electronic information technology, biological and new medical technology, aerospace technology, new material technology, new energy and energy-saving technology, resource and environmental technology, etc. Thus, a plurality of first elements of a plurality of knowledge information may be divided into these different high-tech fields.

[0059] In operation S240, for each parent category, a clustering process is performed on a plurality of second elements of the knowledge vector corresponding to the parent category to obtain a plurality of subcategories.

[0060] For the multiple first elements in each parent category of the multiple parent categories, clustering processing may be performed on the multiple second elements of the knowledge vectors corresponding to the multiple first elements to obtain multiple subcategories.

[0061] The clustering process of this operation is different from the clustering process of the aforementioned operation S230 in that the clustering process targets different data. The clustering process of the aforementioned operation S230 targets multiple first elements of multiple knowledge information, while the clustering process of this operation targets the second elements of the knowledge vectors of multiple knowledge information corresponding to the same parent category.

[0062] In the embodiments of the present disclosure, a subcategory is a lower-level category subordinate to a parent category. A subcategory may be a specific technical branch under the same technical field. For example, when the parent category is electronic information technology, the subcategory may be a technical branch within electronic information technology, including software technology, computer and network technology, communication technology, new electronic components, microelectronics technology, information security technology, intelligent transportation technology, radio and television technology, etc. Thus, multiple second elements of multiple knowledge information may be divided into these different technical branches subordinate to the parent category (such as the high-tech field).

[0063] Taking the summary content and title content of the knowledge vector of the text information as an example, in which the knowledge information is text information and the first element and the second element are respectively the text information, after determining a certain technical field A to which the knowledge vector of the text information belongs, for all summary contents belonging to the technical field A, for example, M summary contents, where M is an integer greater than 1, the M title contents of the M knowledge vectors corresponding to the M summary contents can be clustered, and the M title contents can be divided into different technical branches belonging to the technical field A. The divided technical branches are also referred to as subcategories, and different technical branches form multiple subcategories.

[0064] It should be noted that the types and numbers of parent categories and subcategories can be set or adjusted according to actual conditions, and the present disclosure does not specifically limit the types and numbers of parent categories and subcategories.

[0065] In other embodiments, if there is only one first element in a parent category, it is not necessary to perform clustering on a second element of the knowledge vector corresponding to the first element, and the parent category and the subcategory can be set to be the same. In other embodiments, if there is only one subcategory subordinate to a parent category, it is not necessary to perform clustering on multiple second elements of the knowledge vector corresponding to the parent category, and only one subcategory subordinate to the parent category can be set.

[0066] In operation S250, a knowledge vector library is formed according to a plurality of knowledge vectors of a plurality of knowledge information and a parent category and a child category to which each of the plurality of knowledge vectors belongs.

[0067] Since the knowledge vector of each knowledge information can be attributed to the corresponding parent category and the subcategory belonging to the parent category, multiple knowledge vectors of multiple knowledge information and the parent category and subcategory to which each knowledge vector in the multiple knowledge vectors belongs can be associated and stored in the knowledge vector library, thereby completing the construction of the knowledge vector library.

[0068] Through the embodiments of the present disclosure, a knowledge vector library can be constructed based on the acquired multiple knowledge information, and can be graded and unsupervised self-classified to facilitate information query by users. Each knowledge vector in the knowledge vector library has a parent category to which it belongs and a subcategory subordinate to the parent category, thereby classifying multiple knowledge vectors into multiple levels. When a user needs to query relevant information, the query vector corresponding to the input information can be input into the constructed knowledge vector library, and several vectors with the closest distance or the highest similarity to the query vector can be searched in the vector library to obtain the retrieval results. In this way, the matching accuracy of the knowledge vector library can be improved, and since there is no need to rely on a large amount of manual material sorting, the convenience of user operation is also improved.

[0069] It should be understood that the embodiment of the present disclosure performs secondary classification of multiple knowledge vectors into parent categories and subcategories, thereby forming a knowledge vector library. By analogy, in other embodiments, multiple knowledge vectors can also be classified into multiple levels, such as three-level classification. For example, after each knowledge information is converted into a knowledge vector, the first element, the second element and the third element of the knowledge vector of each knowledge information are first determined, wherein the semantic feature levels of the first element, the second element and the third element are gradually increased. Then, the multiple first elements of the multiple knowledge information are clustered to obtain multiple first-level categories. For each first-level category, the multiple second elements of the knowledge vector corresponding to the first-level category are clustered to obtain multiple second-level categories. For each second-level category, the multiple third elements of the knowledge vector corresponding to the second-level category are clustered to obtain multiple third-level categories. Again, a knowledge vector library is formed based on the multiple knowledge vectors of the multiple knowledge information and the first-level category, the second-level category and the third-level category to which each knowledge vector in the multiple knowledge vectors belongs.

[0070] Exemplarily, the first element, the second element, and the third element may be the text content, abstract content, and title content of the knowledge vector of each knowledge information, respectively. When clustering, the multiple text contents are first clustered to obtain multiple first-level categories. For each first-level category, the multiple abstract contents of the knowledge vector corresponding to the first-level category are clustered to obtain multiple second-level categories. For each second-level category, the multiple title contents of the knowledge vector corresponding to the second-level category are clustered to obtain multiple third-level categories. Thus, a knowledge vector library with three-level classification can be constructed based on the multiple knowledge information obtained, which is convenient for users to query information.

[0071] In this way, the massive amount of knowledge information can be clustered in the order of text, abstract, and title, which can minimize the discrete data generated in each level during the clustering process. As the degree of condensation gradually increases, the level-by-level clustering itself will reduce a certain amount of discrete data or noise, making the accuracy of the clustering results at the next level higher.

[0072] Figure 3 The flowchart for determining the first element and the second element according to an embodiment of the present disclosure is schematically shown.

[0073] like Figure 3 As shown, in some embodiments, the above operation S220 determines the first element and the second element of the knowledge vector of each knowledge information, including: for the knowledge vector of each knowledge information, using the big model to summarize the knowledge vector according to the preset first semantic condensation and second semantic condensation, to obtain the first element and the second element of the knowledge vector, wherein the first semantic condensation is lower than the second semantic condensation.

[0074] In the embodiments of the present disclosure, the big model may be an artificial intelligence big model. An artificial intelligence big model is a machine learning model with ultra-large-scale parameters and complex computing structure. An artificial intelligence big model can process massive amounts of data and complete various complex tasks, such as natural language processing, image recognition, computer vision, etc. An artificial intelligence big model may include a language big model, a visual big model, and a multimodal big model, etc. The big model may run in a central processing unit (CPU) or a graphics processing unit (GPU).

[0075] Exemplarily, the large model may include a large language model (LLM), a visual large model, a multimodal large model, a decision large model, etc. Among them, the large language model (LLM) refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text.

[0076] For example, for any acquired knowledge information 301, the knowledge information 301 is first converted into a knowledge vector 302, and then the large model 303 is used to summarize and refine the knowledge vector 302 according to the preset first semantic condensation and second semantic condensation to obtain the first element 302a and the second element 302b of the knowledge vector 302.

[0077] It should be understood that the higher the semantic condensation of the knowledge vector, the more compact and concise the summarized elements are. The semantic condensation can be expressed in the form of a percentage. For example, the first semantic condensation and the second semantic condensation can be set to 40% and 80% respectively, and the large language model can be used to summarize and refine the knowledge vector of the knowledge information according to the semantic condensation of 40% and 80% respectively, to obtain the summary content and title content of the knowledge vector.

[0078] Through the embodiments of the present disclosure, a large model can be used to summarize and refine knowledge vectors with different semantic condensation levels to determine the first and second elements of the knowledge vectors. Through the large model, a large number of acquired knowledge vectors can be comprehensively analyzed, refined and summarized to ensure the degree of refinement of the content presented in the end, without relying on a large amount of manual material sorting, thereby improving the efficiency of hierarchical summarization and refinement of a large number of knowledge vectors.

[0079] Figure 4A The flowchart of clustering the first elements according to the embodiment of the present disclosure is schematically shown.

[0080] like Figure 4A As shown, for example, for the above operation S230, after obtaining knowledge vectors 1, 2, 3, ..., N of N knowledge information, where N is an integer greater than 1, and determining the first element 1a and the second element 1b of knowledge vector 1, the first element 2a and the second element 2b of knowledge vector 2, the first element 3a and the second element 3b of knowledge vector 3, ..., the first element Na and the second element Nb of knowledge vector N, the N first elements of the N knowledge information can be clustered to determine the corresponding parent categories, such as the first parent category F1 and the second parent category F2.

[0081] For example, according to actual needs, the first parent category F1 can be characterized as the field of electronic information technology, and the second parent category F2 can be characterized as the field of new energy and energy-saving technology.

[0082] In this way, the first element of each knowledge vector in knowledge vectors 1, 2, 3, ..., N can be attributed to one of the first parent category F1 and the second parent category F2, and each parent category in the first parent category F1 and the second parent category F2 contains at least one first element.

[0083] Figure 4B The flowchart of clustering multiple second elements according to an embodiment of the present disclosure is schematically shown.

[0084] like Figure 4BAs shown, for example, for the above operation S240, for any parent category among multiple parent categories, such as the first parent category F1, when the first elements 1a, 3a, 5a, 6a are included in the first parent category F1, the second elements 1b, 3b, 5b, 6b of the knowledge vectors 1, 3, 5, 6 corresponding to the first elements 1a, 3a, 5a, 6a can be clustered to determine multiple subcategories belonging to the first parent category F1, such as the first subcategory Z11 and the second subcategory Z12.

[0085] Taking the first parent category F1 representing the field of electronic information technology as an example, according to actual needs, the first subcategory Z11 can be represented as a branch of computer and network technology, and the second subcategory Z12 can be represented as a branch of communication technology.

[0086] In this way, each second element 1b, 3b, 5b, 6b of the knowledge vectors 1, 3, 5, 6 corresponding to the first parent category F1 can be assigned to one of the first subcategory Z11 and the second subcategory Z12, and each subcategory in the first subcategory Z11 and the second subcategory Z12 contains at least one second element.

[0087] The embodiments of the present disclosure use an unsupervised clustering algorithm to efficiently replace the tedious manual processing of massive data. After the massive data is vectorized, the cosine similarity function (including but not limited to other methods of measuring the similarity between vectors) can be used to perform unsupervised automatic aggregation between vectors, and the category is the vector feature point of the cluster center.

[0088] Figure 5 The flowchart of forming a knowledge vector library according to an embodiment of the present disclosure is schematically shown.

[0089] like Figure 5 As shown, in some embodiments, the above operation S250 forms a knowledge vector library according to multiple knowledge vectors of multiple knowledge information and the parent category and subcategory to which each knowledge vector in the multiple knowledge vectors belongs, and further includes operation S250a~operation S250b.

[0090] In operation S250a, for any target category among the plurality of parent categories and the plurality of child categories, a large model is used to perform semantic induction on the cluster center of the target category to obtain a category label of the target category.

[0091] For example, each time a parent category or a child category is obtained through clustering processing in the above operations S230 to S240, the semantic induction of the cluster center can be automatically performed on the parent category or the child category through the large language model to obtain a category label.

[0092] In operation S250b, a knowledge vector library is constructed according to the plurality of knowledge vectors of the plurality of knowledge information and the category label of the parent category and the category label of the child category to which each of the plurality of knowledge vectors belongs.

[0093] So far, the constructed knowledge vector library stores multiple knowledge vectors, and each knowledge vector has a category label of the parent category and a category label of the child category to which it belongs. The category label is generated after the cluster center is refined and polished by the large language model.

[0094] In some embodiments, the operation S250b constructs a knowledge vector library according to the multiple knowledge vectors of the multiple knowledge information and the category label of the parent category and the category label of the child category to which each knowledge vector in the multiple knowledge vectors belongs, further comprising:

[0095] Sort multiple parent categories according to the number of first elements contained in each parent category, and select some parent categories with higher rankings;

[0096] Filtering some knowledge vectors belonging to some parent categories from multiple knowledge vectors of multiple knowledge information;

[0097] A knowledge vector library is constructed according to the partial knowledge vectors and the category labels of the parent category and the subcategory to which each knowledge vector in the partial knowledge vectors belongs.

[0098] For example, the multiple parent categories obtained after clustering processing are the first parent category F1, the second parent category F2, and the third parent category F3. The number of first elements contained in the three parent categories is 5, 2, and 9, respectively. According to the predetermined sorting rules, the top two parent categories can be selected from the three parent categories.

[0099] The embodiments of the present disclosure do not specifically limit the predetermined sorting rules. For example, the predetermined sorting rule may be: the more first elements any parent category contains, the more typical and representative the parent category is among multiple parent categories. Thus, the first parent category F1 and the third parent category F3 with the highest ranking may be selected from the first parent category F1, the second parent category F2, and the third parent category F3 according to the number of first elements contained.

[0100] Through the embodiments of the present disclosure, some parent categories with top rankings can be selected from multiple parent categories obtained through clustering processing, thereby avoiding too much storage data in the knowledge vector library and improving retrieval efficiency.

[0101] The method for constructing a knowledge vector library provided by the embodiment of the present disclosure can adapt to various natural language level scenarios, and not only supports the category labels generated after the cluster center is refined and polished by a large language model, but also supports mixed scenarios where knowledge vector features are to be used downstream. In order to unify the use specifications, the storage data of the knowledge vector library can also be encapsulated in a specific format (such as json format) for easy expansion.

[0102] Based on the above-mentioned knowledge vector library construction method, the present disclosure also provides an information processing method. Figure 6~Figure 8 The information processing method is described in detail.

[0103] Figure 6 The flowchart of the information processing method according to the embodiment of the present disclosure is schematically shown.

[0104] like Figure 6 As shown, the information processing method 600 of this embodiment may include operations S610 to S630. The method may be executed by a server.

[0105] In operation S610 , input information is acquired and converted into an input vector.

[0106] In an embodiment of the present disclosure, corresponding to the knowledge information of the aforementioned operation S210, the input information may be multimodal information, which includes multiple types of data such as images, text, audio, and video. The input information may be a text file, a word document, an excel spreadsheet, a PDF document, a picture file, an audio file, a video file, and the like.

[0107] The input information may be converted into an input vector in the manner of converting each piece of knowledge information into a knowledge vector in the aforementioned operation S210, which will not be described in detail herein.

[0108] In operation S620, based on a pre-constructed knowledge vector library, a target parent category that matches the input vector and a plurality of first knowledge vectors corresponding to the target parent category are determined; from the plurality of first knowledge vectors, a target subcategory that matches the input vector and at least one second knowledge vector corresponding to the target subcategory are determined, wherein the knowledge vector library stores a plurality of knowledge vectors and a parent category and a subcategory to which each of the plurality of knowledge vectors belongs, and a semantic feature level of the parent category is lower than a semantic feature level of the subcategory.

[0109] The construction of the knowledge vector library can be completed by using the aforementioned operation S210 to operation S250, which will not be described in detail here.

[0110] In operation S630, output information corresponding to the input information is generated according to the at least one second knowledge vector.

[0111] Through the embodiments of the present disclosure, when a user needs to perform information retrieval based on input information, the input information can be first converted into an input vector, and then at least one second knowledge vector with a two-level category matching the input vector can be found from a pre-built knowledge vector library, and a retrieval result can be obtained based on the at least one second knowledge vector. Since each knowledge vector in the knowledge vector library has a parent category to which it belongs and a subcategory subordinate to the parent category, hierarchical retrieval can be performed in the order of the parent category and the subcategory, the retrieval scope can be quickly located, and the efficiency and accuracy of information retrieval can be improved.

[0112] In some embodiments, the above operation S620 determines the target parent category that matches the input vector, including: determining the similarity between the input vector and the parent category to which each knowledge vector in the knowledge vector library belongs; and determining the parent category in the knowledge vector library whose similarity is higher than a preset threshold as the target parent category.

[0113] For example, after obtaining the input vector, the cosine similarity function (including but not limited to other methods of measuring the similarity between vectors) can be used to calculate the similarity between the input vector and the parent category to which each knowledge vector in the knowledge vector library belongs, obtain multiple similarities, and then use the parent category with the highest similarity in the knowledge vector library as the target parent category.

[0114] In some embodiments, the above operation S630 generates output information corresponding to the input information based on at least one second knowledge vector, including: summarizing at least one second knowledge vector using a large model to generate corresponding knowledge information in the target domain; and determining the knowledge information in the target domain as output information.

[0115] Figure 7 The schematic diagram shows a principle diagram of an information processing method according to an embodiment of the present disclosure.

[0116] like Figure 7 As shown, for example, after obtaining input information 701, the input information 701 is converted into an input vector 702. Secondly, a pre-built knowledge vector library 703 is called, and the knowledge vector library 703 defines a first parent category F1, a second parent category F2, and a third parent category F3, wherein the subcategories subordinate to the first parent category F1 include the first subcategory Z11, the second subcategory Z12, and the third subcategory Z13; the subcategories subordinate to the second parent category F2 are the first subcategory Z21; and the subcategories subordinate to the third parent category F3 include the first subcategory Z31 and the second subcategory Z32.

[0117] Next, from the knowledge vector library 703, it can be determined that the target parent category that matches the input vector 702 is the third parent category F3, and multiple first knowledge vectors corresponding to the third parent category F3; from multiple first knowledge vectors, it can be determined that the target subcategory that matches the input vector 702 is the first subcategory Z31 and the second subcategory Z32, and at least one second knowledge vector corresponding to the target subcategory.

[0118] Then, the large model 705 may be used to assemble and polish at least one second knowledge vector to generate knowledge information in the target domain corresponding to the input information 701 as output information 706 .

[0119] In some embodiments, the operation S620 may further include:

[0120] Determining a first input element and a second input element of an input vector, wherein a semantic feature level of the first input element is higher than a semantic feature level of the second input element;

[0121] From the knowledge vector library, determine the target parent category that matches the first input element and multiple first knowledge vectors corresponding to the target parent category; from the multiple first knowledge vectors, determine the target subcategory that matches the second input element and at least one second knowledge vector corresponding to the target subcategory.

[0122] The first input element and the second input element of the input vector may be determined by adopting the aforementioned operation S220 to determine the first element and the second element of the knowledge vector of each piece of knowledge information, which will not be described in detail herein.

[0123] Figure 8 The schematic diagram shows a principle diagram of determining a second knowledge vector according to an embodiment of the present disclosure.

[0124] For example, in Figure 7 Based on the knowledge vector library constructed, Figure 8 As shown, after obtaining the input vector 901, the large model can be used to determine the first input element 801a and the second input element 801b of the input vector 901. Next, the pre-built knowledge vector library 802 is called, and the first input element 801a is input into the knowledge vector library 802, and it can be determined that the target parent category matched by the first input element 801a is the second parent category F2, and multiple first knowledge vectors corresponding to the second parent category F2. Next, for multiple first knowledge vectors, the second input element 801b is input into the knowledge vector library 802, and it can be determined that the target subcategory matched by the second input element 801b and subordinate to the second parent category F2 is the first subcategory Z21, and at least one second knowledge vector corresponding to the first subcategory Z21.

[0125] Then, the large model can be used to assemble and polish at least one second knowledge vector to generate knowledge information in the target domain corresponding to the input information as output information.

[0126] In some embodiments, when the pre-constructed knowledge vector library is a case where multiple knowledge vectors are classified into three levels, the above operation S620 may further include: determining the first input element, the second input element, and the third input element of the input vector, wherein the semantic feature levels of the first input element, the second input element, and the third input element decrease step by step. Then, from the knowledge vector library, determine the first-level target category that matches the first input element, and multiple first knowledge vectors corresponding to the first-level target category; from the multiple first knowledge vectors, determine the second-level target category that matches the second input element, and multiple second knowledge vectors corresponding to the second-level target category; from the multiple second knowledge vectors, determine the third-level target category that matches the third input element, and at least one second knowledge vector corresponding to the third-level target category.

[0127] Exemplarily, the first input element, the second input element and the third input element can be the title content, the abstract content and the text content of the input vector of the input information respectively. In the process of vector matching in the knowledge vector library, the matching can be carried out step by step in the order of title, abstract and text, and hierarchical retrieval can be carried out according to the principle of three-level step-by-step convergence. Since the results of the upper level will be used as the input range of the lower level, further convergence on this range, that is, the search range is gradually narrowed, thereby achieving a faster search convergence speed, improving the speed of vector matching, and thus improving the efficiency and accuracy of information retrieval.

[0128] Based on this, the information processing method of the embodiment of the present disclosure can be applied to fields such as information query, retrieval, and intelligent question and answer.

[0129] Based on the above-mentioned knowledge vector library construction method, the present disclosure also provides a knowledge vector library construction system. Fig. 9 The system is described in detail.

[0130] Fig. 9 The block diagram of a system for constructing a knowledge vector library according to an embodiment of the present disclosure is schematically shown.

[0131] like Fig. 9 As shown, the knowledge vector library construction system 900 of this embodiment includes a knowledge information library 910 , a construction module 920 and a knowledge vector library 930 .

[0132] The knowledge information base 910 stores a plurality of knowledge information;

[0133] The construction module 920 is communicated with the knowledge information library 910 and the knowledge vector library 930 respectively, and is used to: obtain multiple knowledge information and convert each knowledge information into a knowledge vector; determine the first element and the second element of the knowledge vector of each knowledge information, wherein the semantic feature level of the first element is lower than the semantic feature level of the second element; cluster the multiple first elements of the multiple knowledge information to obtain multiple parent categories; for each parent category, cluster the multiple second elements of the knowledge vector corresponding to the parent category to obtain multiple subcategories; store the multiple knowledge vectors of the multiple knowledge information and the parent category and subcategory to which each knowledge vector in the multiple knowledge vectors belongs in the knowledge vector library 930.

[0134] Exemplarily, the knowledge information base 910 may be a relational database or a search engine, which is not specifically limited in the present disclosure. For example, the knowledge information base 910 may be a MySQL database or an ES / Sor search engine.

[0135] The specific implementation logic of constructing the knowledge vector library of the construction module 920 has been described in detail in the embodiment of the construction method of the knowledge vector library, and will not be repeated here. The construction module 920 is applicable to the construction of various knowledge vector libraries, regardless of the knowledge type and application field.

[0136] The knowledge vector library 930 stores a plurality of knowledge vectors of a plurality of knowledge information and a parent category and a child category to which each of the plurality of knowledge vectors belongs.

[0137] It should be noted that the implementation method of the system part of the knowledge vector library construction is similar to the implementation method of the method part of the knowledge vector library construction mentioned above, and the technical effects achieved are also similar. For specific details, please refer to the implementation method part of the method of constructing the knowledge vector library mentioned above, which will not be repeated here.

[0138] Based on the above information processing method, the present disclosure also provides an electronic device. Fig.10 The electronic device is described in detail.

[0139] Fig.10 A block diagram of an electronic device according to an embodiment of the present disclosure is schematically shown.

[0140] like Fig.10 As shown, the electronic device 1000 of this embodiment includes an information input module 1010 , a processor 1020 and a memory 1030 .

[0141] Wherein, the information input module 1010 is used to obtain input information;

[0142] The memory 1030 stores a knowledge vector library, which stores a plurality of knowledge vectors and a parent category and a subcategory to which each of the plurality of knowledge vectors belongs, wherein the semantic feature level of the parent category is lower than the semantic feature level of the subcategory;

[0143] The processor 1020 is communicatively connected to the information input module 1010 and the memory 1030, respectively, and is used to: convert the input information into an input vector; call the knowledge vector library to determine, from the knowledge vector library, a target parent category that matches the input vector, and a plurality of first knowledge vectors corresponding to the target parent category; determine, from the plurality of first knowledge vectors, a target subcategory that matches the input vector, and at least one second knowledge vector corresponding to the target subcategory; and output output information corresponding to the input information based on the at least one second knowledge vector.

[0144] The information input module 1010 is used to receive input character information and generate key signal input related to user settings and function control of the electronic device.

[0145] The memory 1030 is a computer-readable storage medium that can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the information processing method in the embodiment of the present disclosure. The processor 1020 executes various functional applications and data processing of the electronic device by running the software programs, instructions and modules stored in the memory 1030, that is, implements the above-mentioned information processing method.

[0146] It should be noted that the implementation method of the electronic device part is similar to the implementation method of the above-mentioned information processing method part, and the technical effects achieved are also similar. For specific details, please refer to the implementation method part of the above-mentioned information processing method, which will not be repeated here.

[0147] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.

[0148] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A method for constructing a knowledge vector library, comprising: Acquire multiple pieces of knowledge information, and convert each piece of the knowledge information into a knowledge vector; Determining a first element and a second element of the knowledge vector of each piece of the knowledge information, wherein a semantic feature level of the first element is lower than a semantic feature level of the second element; Clustering the first elements of the plurality of pieces of knowledge information to obtain a plurality of parent categories; For each of the parent categories, clustering is performed on a plurality of the second elements of the knowledge vector corresponding to the parent category to obtain a plurality of subcategories; The knowledge vector library is formed according to the plurality of knowledge vectors of the plurality of knowledge information and the parent category and the child category to which each of the plurality of knowledge vectors belongs.

2. The method according to claim 1, wherein determining the first element and the second element of the knowledge vector of each piece of knowledge information comprises: For each knowledge vector of the knowledge information, a large model is used to summarize the knowledge vector according to a preset first semantic condensation degree and a preset second semantic condensation degree to obtain a first element and a second element of the knowledge vector, wherein the first semantic condensation degree is lower than the second semantic condensation degree.

3. The method according to claim 1, wherein the forming of the knowledge vector library according to the plurality of knowledge vectors of the plurality of knowledge information and the parent category and the subcategory to which each of the plurality of knowledge vectors belongs comprises: For any target category among the multiple parent categories and the multiple child categories, use the large model to perform semantic induction on the cluster center of the target category to obtain a category label of the target category; The knowledge vector library is constructed according to the plurality of knowledge vectors of the plurality of knowledge information and the category label of the parent category and the category label of the child category to which each of the plurality of knowledge vectors belongs.

4. The method according to claim 3, wherein the step of constructing the knowledge vector library according to the plurality of knowledge vectors of the plurality of knowledge information and the category label of the parent category and the category label of the child category to which each of the plurality of knowledge vectors belongs comprises: sorting the plurality of parent categories according to the number of first elements contained in each of the parent categories, and selecting some parent categories with higher sorting; Filtering some knowledge vectors belonging to some parent categories from among the plurality of knowledge vectors of the plurality of knowledge information; The knowledge vector library is constructed according to the partial knowledge vectors and the category label of the parent category and the category label of the child category to which each knowledge vector in the partial knowledge vectors belongs.

5. An information processing method, comprising: Obtaining input information, and converting the input information into an input vector; Based on a pre-constructed knowledge vector library, determine a target parent category that matches the input vector, and a plurality of first knowledge vectors corresponding to the target parent category; from the plurality of first knowledge vectors, determine a target subcategory that matches the input vector, and at least one second knowledge vector corresponding to the target subcategory, wherein the knowledge vector library stores a plurality of knowledge vectors and a parent category and a subcategory to which each of the plurality of knowledge vectors belongs, and a semantic feature level of the parent category is lower than a semantic feature level of the subcategory; Output information corresponding to the input information is generated according to the at least one second knowledge vector.

6. The method according to claim 5, wherein determining the target parent category that matches the input vector comprises: Determine the similarity between the input vector and the parent category to which each knowledge vector in the knowledge vector library belongs; The parent category in the knowledge vector library whose similarity is higher than a preset threshold is determined as the target parent category.

7. The method according to claim 5, wherein generating output information corresponding to the input information according to the at least one second knowledge vector comprises: Summarizing the at least one second knowledge vector using the large model to generate knowledge information in the corresponding target domain; The knowledge information in the target domain is determined as the output information.

8. The method according to claim 5, further comprising: Determining a first input element and a second input element of the input vector, wherein a semantic feature level of the first input element is higher than a semantic feature level of the second input element; From the knowledge vector library, Determine a target parent category that matches the first input element, and a plurality of first knowledge vectors corresponding to the target parent category; A target subcategory matched with the second input element and at least one second knowledge vector corresponding to the target subcategory are determined from the multiple first knowledge vectors.

9. A system for constructing a knowledge vector library, comprising: A knowledge information base stores a plurality of knowledge information; Knowledge vector library; The construction module is respectively connected to the knowledge information library and the knowledge vector library for: Acquire the plurality of pieces of knowledge information, and convert each piece of the knowledge information into a knowledge vector; Determining a first element and a second element of the knowledge vector of each piece of the knowledge information, wherein a semantic feature level of the first element is lower than a semantic feature level of the second element; Clustering the first elements of the plurality of pieces of knowledge information to obtain a plurality of parent categories; For each of the parent categories, clustering is performed on a plurality of the second elements of the knowledge vector corresponding to the parent category to obtain a plurality of subcategories; The plurality of knowledge vectors of the plurality of knowledge information and the parent category and the child category to which each of the plurality of knowledge vectors belongs are stored in the knowledge vector library.

10. An electronic device, comprising: An information input module, used for obtaining input information; A memory storing a knowledge vector library, wherein the knowledge vector library stores a plurality of knowledge vectors and a parent category and a subcategory to which each of the plurality of knowledge vectors belongs, wherein a semantic feature level of the parent category is lower than a semantic feature level of the subcategory; A processor is connected to the information input module and the memory for communication, and is used for: Converting the input information into an input vector; Calling the knowledge vector library, determining from the knowledge vector library a target parent category that matches the input vector, and a plurality of first knowledge vectors corresponding to the target parent category; determining from the plurality of first knowledge vectors a target subcategory that matches the input vector, and at least one second knowledge vector corresponding to the target subcategory; Output information corresponding to the input information is output according to the at least one second knowledge vector.

Citation Information

Cited By

  • Local knowledge base construction method and system applying artificial intelligence

    CN120562544A