AI (artificial intelligence) intelligent analysis method capable of enabling by applying large language model

Through the cooperation of users to build physical network structure and style supervision branches by themselves, the problem of poor user experience in the early stage of fine-tuning of large language models is solved, and the user experience is improved.

CN119938851AActive Publication Date: 2025-05-06NANJING HUIZHI INTERACTIVE ENTERTAINMENT NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510056208.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The problem of poor user experience in the early stages of fine-tuning of large language models.

Method used

Through specific interactions, users can build physical network structures that match their own tendencies or styles, and use style supervision branches to fine-tune the artificial intelligence model.

Benefits of technology

Effectively improve the user's initial experience, and optimize the performance of the artificial intelligence model through the cooperation of user-defined entity network structure and style supervision branches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938851A_ABST
    Figure CN119938851A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language models, and discloses an AI intelligent analysis method by applying a large language model, which comprises the following steps of: training an artificial intelligence model aiming at enabling the artificial intelligence model to have a universal function of generating text knowledge content; sorting according to the occurrence frequencies of the knowledge entities in the associated information, and extracting a plurality of knowledge entities as core entities according to the sequence of the occurrence frequencies from high to low; visualizing the core entities on the interactive interface, and constructing a second entity relationship between the core entities according to the operation of the user on the interactive interface; finely adjusting the artificial intelligence model; according to the method, before the user uses the large language model, the user constructs the entity network structure conforming to the tendency or style of the user by himself / herself through specific interaction, and then the artificial intelligence model is finely adjusted through the entity network structure and the style supervision branch, so that the initial experience of the user is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and more specifically, to a method for enabling AI intelligent analysis using a large language model. Background Art

[0002] With the development of the Internet and artificial intelligence technology, large language models can replace some simple mental work. For example, intelligent customer service assistants can use large language model technology to accurately understand user intentions, generate accurate answers, conduct multiple rounds of conversations, and continuously learn and optimize, thereby improving user experience. Large language models are generally trained to generate general text knowledge content first, and then fine-tune user needs and preferences through continuous interaction and learning with users. Training and fine-tuning rely on a large amount of interaction data, but this will inevitably cause poor user experience in the early stages of fine-tuning. Summary of the invention

[0003] The present invention provides an AI intelligent analysis method enabled by a large language model, which solves the technical problem of poor user experience in the early stage of fine-tuning of the large language model in the related art.

[0004] The present invention provides an AI intelligent analysis method using a large language model, comprising the following steps: training an artificial intelligence model, wherein the training goal is to enable it to have a universal function of generating text knowledge content;

[0005] Collecting relevant information of users and / or relevant information of application fields of artificial intelligence models;

[0006] Performing entity recognition from the collected association information to obtain knowledge entities and first entity relationships between the knowledge entities;

[0007] Sort the knowledge entities according to their frequencies of appearance in the associated information, and extract several knowledge entities as core entities in descending order of frequency of appearance;

[0008] Visualize the core entities on the interactive interface, and build second entity relationships between the core entities based on the user's operations on the interactive interface;

[0009] Extracting several knowledge entities that have a first entity relationship directly and / or indirectly with the core entity as extended entities, adding the extended entities to the interactive interface for visualization, and constructing a second entity relationship between the core entity and the extended entity or between the extended entities according to the user's operation on the interactive interface;

[0010] Fine-tune the artificial intelligence model. The artificial intelligence model includes a generation branch and a style supervision branch. The generation branch is used for text knowledge content generation; the style supervision branch includes a pattern learning layer, a feature fusion layer, and a second entity relationship restoration layer. The pattern learning layer inputs the structured representation of the core entity and the extended entity, and outputs the first hidden feature of the object. The feature fusion layer is used to fuse the first hidden feature of the object with the latent vector representation output by the encoder of the generation branch to obtain a fused feature matrix. The second entity relationship restoration layer inputs the fused feature matrix and outputs the element of the i-th row and j-th column of the second entity relationship matrix to indicate whether there is a correlation between the i-th object and the j-th object related to the generation of text knowledge content.

[0011] Furthermore, the entities are visualized as text boxes on the interactive interface, and the user's operation on the interactive interface includes connecting lines between the text boxes.

[0012] Further, the generation branch includes an embedding layer, a first encoder, a decoder, and a first output layer, wherein the embedding layer is used to map discrete text symbols into continuous vector representations;

[0013] The encoder is used to encode the embedded vector sequence, extract context information, and generate a latent vector representation of fixed length;

[0014] The decoder generates the target text sequence based on the latent vector representation generated by the encoder.

[0015] Furthermore, the goal of fine-tuning the artificial intelligence model is to minimize the difference between the second entity relationship matrix and the second entity relationships between the core entities and / or extended entities.

[0016] Furthermore, the structured representation includes a structural dimension, in which the basic structural elements included in the structural dimension are objects, and the structuring methods include:

[0017] Extract objects that can be associated with specific features;

[0018] Establishing a mapping between the extracted objects and their associated specific features;

[0019] For any two objects, if there is a correlation between the associated specific features related to the generation of textual knowledge content or there is a correlation between the objects related to the generation of textual knowledge content, an object relationship is established between the two objects.

[0020] Furthermore, the extracted objects that can be associated with specific features include: core entities and extended entities.

[0021] Furthermore, objects have specific characteristics associated with them as follows:

[0022] Specific features associated with core entities include: embedded representations of core entities;

[0023] Specific features of extending entity associations include: embedded representations of core entities.

[0024] Furthermore, the relevance related to textual knowledge content generation includes:

[0025] There is a second entity relationship between the core entity and the extended entity;

[0026] There is a second entity relationship between the extended entities;

[0027] There is a second entity relationship between the core entity and the core entity.

[0028] Furthermore, the contextual semantics of discrete text symbols are derived from the second entity relationships between core entities and / or extended entities.

[0029] The present invention provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, which, when executed by a processor, enable the processor to execute the aforementioned AI intelligent analysis method enabled by a large language model.

[0030] The beneficial effect of the present invention is that before the user uses the large language model, the present invention enables the user to build an entity network structure that conforms to his or her own inclinations or style through specific interactions, and then fine-tunes the artificial intelligence model through the entity network structure in conjunction with the style supervision branch, thereby effectively improving the user's initial experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a flow chart of the method for using a large language model to enable AI intelligent analysis in the present invention. DETAILED DESCRIPTION

[0032] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the present specification. Various examples may omit, replace, or add various processes or components as needed. In addition, the features described in some examples may also be combined in other examples.

[0033] At least one embodiment of the present invention discloses a method for using a large language model to enable AI intelligent analysis, such as Figure 1 As shown, the following steps are included:

[0034] Step 101, training the artificial intelligence model, the goal of the training is to enable it to have a universal function of generating text knowledge content;

[0035] The aforementioned universal function refers to: the text knowledge content generation function obtained by training under the general text knowledge content generation corpus, and the generated content does not have the characteristics of being biased towards the interests or styles of one or some specific users;

[0036] In one embodiment of the present invention, the aforementioned artificial intelligence model is used to assist users in generating new textual knowledge content, including but not limited to novels, scripts, essays, etc.

[0037] Step 102, collecting user-related information and / or related information of the application field of the artificial intelligence model;

[0038] Performing entity recognition from the collected association information to obtain knowledge entities and first entity relationships between the knowledge entities;

[0039] In one embodiment of the present invention, this step also requires pre-processing operations such as deduplication of the acquired knowledge entities.

[0040] In one embodiment of the present invention, the application fields of the artificial intelligence model include but are not limited to customer service question and answer, knowledge question and answer, novel generation, script generation, and essay generation.

[0041] Of course, these areas of the examples can be further subdivided. For example, script generation can be further subdivided into short video script generation or historical script generation. The examples provided do not limit them to being replaced by lower-level concepts.

[0042] In one embodiment of the present invention, the user's associated information is novels that the user has read, and the associated information of the application field of the artificial intelligence model is historical fiction novels in the novel library.

[0043] Step 103, sorting the knowledge entities according to their frequencies of appearance in the associated information, and extracting a number of knowledge entities as core entities in descending order of their frequencies of appearance;

[0044] Visualize the core entities on the interactive interface, and build second entity relationships between the core entities based on the user's operations on the interactive interface;

[0045] Step 104, extracting several knowledge entities that have a first entity relationship directly and / or indirectly with the core entity as extended entities, adding the extended entities to the interactive interface for visualization, and constructing a second entity relationship between the core entity and the extended entity or between the extended entities according to the user's operation on the interactive interface (consistent with the meaning of the entity relationship in step 103);

[0046] In one embodiment of the present invention, the entity is visualized as a text box on the interactive interface, and the user's operation on the interactive interface includes connecting lines between the text boxes and may also include annotating text on the connecting lines.

[0047] Step 105, fine-tuning the artificial intelligence model, where the artificial intelligence model includes a generation branch and a style supervision branch, where the generation branch includes an embedding layer, a first encoder, a decoder, and a first output layer, where the embedding layer is used to map discrete text symbols into continuous vector representations;

[0048] The encoder is used to encode the embedded vector sequence, extract context information, and generate a latent vector representation of fixed length;

[0049] The decoder generates the target text sequence based on the latent vector representation generated by the encoder.

[0050] Commonly used embedding methods include: word embedding (Word Embedding) such as Word2Vec, character embedding (Character Embedding), etc.

[0051] Common encoder structures include: recurrent neural networks (RNNs) such as LSTM (Long Short-Term Memory Network), GRU (Gated Recurrent Unit), etc., convolutional neural networks (CNN), and self-attention mechanisms (Self-Attention) such as the encoder part in Transformer.

[0052] The decoder usually uses an autoregressive approach to predict the next text unit based on the previously generated text.

[0053] Common decoder structures include: recurrent neural networks (RNNs) such as LSTM (Long Short-Term Memory Network), GRU (Gated Recurrent Unit), etc., and the decoder part in Transformer.

[0054] Although the specific structure of the generation branch of the artificial intelligence model is provided here, it should be understood that these structures are to assist in explaining the style supervision branch and do not exclude other models with text knowledge content generation capabilities, such as GPT-3, Bloom, and LLaMA (Large Language Model Meta AI).

[0055] The style supervision branch includes a pattern learning layer, a feature fusion layer, and a second entity relationship restoration layer. The pattern learning layer inputs the structured representation of the core entity and the extended entity. The structured representation includes a structural dimension. The basic structural elements contained in this structural dimension are objects. The structuring methods include:

[0056] Extract objects that can be associated with specific features;

[0057] Establishing a mapping between the extracted objects and their associated specific features;

[0058] For any two objects, if there is a correlation between the associated specific features related to the generation of textual knowledge content or there is a correlation between the objects related to the generation of textual knowledge content, then an object relationship is established between the two objects;

[0059] In some embodiments of the present invention, the extracted objects capable of associating specific features include: core entities, extended entities;

[0060] In some embodiments of the invention, objects have specific features associated with them as follows:

[0061] Specific features associated with core entities include: embedded representations of core entities;

[0062] Specific features of extending entity associations include: embedded representations of core entities.

[0063] In some embodiments of the present invention, the relevance related to the generation of textual knowledge content includes:

[0064] There is a second entity relationship between the core entity and the extended entity;

[0065] There is a second entity relationship between the extended entities;

[0066] There is a second entity relationship between the core entity and the core entity.

[0067] The pattern learning layer is a multi-layer structure, and the calculation formula of the lth layer is as follows:

[0068]

[0069]

[0070] in represents the aggregate representation of the vth object at the lth layer, and denote the first object recognition features of the vth object in the lth layer and the l-1th layer respectively, A collection of objects that represent the association with object v related to the generation of textual knowledge content; Represents the universal quantifier symbol; represents the first object recognition feature of the u-th object in the l-1th layer, , E represents the total number of layers, when l=1 , , and Represent the features associated with the u-th and v-th objects respectively, when l=E is equal to the first hidden feature of object v, Represents the aggregation function of layer l, such as average or maximum pooling function. represents the weight matrix of the lth layer, and σ represents the Sigmoid function.

[0071] The calculation formula of the feature fusion layer is as follows:

[0072]

[0073] in Represents the fusion feature matrix, whose t-th row vector is given by and Stitched together, represents the latent vector representation of the encoder output, represents the first hidden feature of object t, Represents the collection of all objects, Indicates splicing;

[0074] The calculation formula for the second entity relationship reduction layer is as follows:

[0075]

[0076] in Represents a binarization function. Any element will be binarized. The condition for assigning a value of 1 is that the element value is greater than 0.5. Represents the second entity relationship matrix. The element in the i-th row and j-th column of the second entity relationship matrix indicates whether there is a correlation between the i-th object and the j-th object related to the generation of text knowledge content. If the element value is 1, it means there is a correlation, otherwise it does not exist. represents transpose;

[0077] It should be noted that the discrete text symbols input into the fine-tuning artificial intelligence model are not input by the user through the interactive interface, but are discrete text symbols that express core entities and extended entities. The contextual semantics of these discrete text symbols come from the second entity relationship between the core entities and / or extended entities.

[0078] The goal of fine-tuning the artificial intelligence model is to minimize the difference between the second entity relationship matrix and the second entity relationships between the core entities and / or extended entities.

[0079] In some embodiments of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the aforementioned method of applying a large language model to enable AI intelligent analysis.

[0080] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation mode. The above-mentioned specific implementation mode is merely illustrative and not restrictive. Under the guidance of this embodiment, ordinary technicians in this field can also make more forms of equivalent embodiments, all of which are within the protection of this embodiment.

Claims

1. A method for AI intelligent analysis using a large language model, characterized in that: The following steps are involved: Train the AI ​​model to enable it to have the general functionality of generating textual knowledge content; Collecting relevant information of users and / or relevant information of application fields of artificial intelligence models; Performing entity recognition from the collected association information to obtain knowledge entities and first entity relationships between the knowledge entities; Sort the knowledge entities according to their frequencies of appearance in the associated information, and extract several knowledge entities as core entities in descending order of frequency of appearance; Visualize the core entities on the interactive interface, and build second entity relationships between the core entities based on the user's operations on the interactive interface; Extracting several knowledge entities that have a first entity relationship directly and / or indirectly with the core entity as extended entities, adding the extended entities to the interactive interface for visualization, and constructing a second entity relationship between the core entity and the extended entity or between the extended entities according to the user's operation on the interactive interface; Fine-tune the artificial intelligence model. The artificial intelligence model includes a generation branch and a style supervision branch. The generation branch is used for text knowledge content generation; the style supervision branch includes a pattern learning layer, a feature fusion layer, and a second entity relationship restoration layer. The pattern learning layer inputs the structured representation of the core entity and the extended entity, and outputs the first hidden feature of the object. The feature fusion layer is used to fuse the first hidden feature of the object with the latent vector representation output by the encoder of the generation branch to obtain a fused feature matrix. The second entity relationship restoration layer inputs the fused feature matrix and outputs the element of the i-th row and j-th column of the second entity relationship matrix to indicate whether there is a correlation between the i-th object and the j-th object related to the generation of text knowledge content.

2. According to claim 1, a method for AI intelligent analysis using a large language model is characterized in that: The entities are visualized as text boxes on the interactive interface, and the user's operations on the interactive interface include connecting lines between the text boxes.

3. According to claim 1, a method for AI intelligent analysis using a large language model is characterized in that: The generation branch includes an embedding layer, a first encoder, a decoder, and a first output layer, wherein the embedding layer is used to map discrete text symbols into continuous vector representations; The encoder is used to encode the embedded vector sequence, extract context information, and generate a latent vector representation of fixed length; The decoder generates the target text sequence based on the latent vector representation generated by the encoder.

4. According to claim 3, a method for AI intelligent analysis using a large language model is characterized in that: The contextual semantics of discrete text symbols comes from the secondary entity relationships between core entities and / or extended entities.

5. According to claim 1, a method for AI intelligent analysis using a large language model is characterized in that: The goal of fine-tuning the artificial intelligence model is to minimize the difference between the second entity relationship matrix and the second entity relationships between the core entities and / or extended entities.

6. According to claim 1, a method for AI intelligent analysis using a large language model is characterized in that: The structured representation includes a structural dimension, in which the basic structural elements are objects. The structuring methods include: Extract objects that can be associated with specific features; Establishing a mapping between the extracted objects and their associated specific features; For any two objects, if there is a correlation between the associated specific features related to the generation of textual knowledge content or there is a correlation between the objects related to the generation of textual knowledge content, an object relationship is established between the two objects.

7. According to claim 6, a method for AI intelligent analysis using a large language model is characterized in that: The extracted objects that can be associated with specific features include: core entities and extended entities.

8. According to claim 6, a method for AI intelligent analysis using a large language model is characterized in that: Objects have specific characteristics associated with them: Specific features associated with core entities include: embedded representations of core entities; Specific features of extending entity associations include: embedded representations of core entities.

9. According to claim 6, a method for AI intelligent analysis using a large language model is characterized in that: Relevance to textual knowledge content generation includes: There is a second entity relationship between the core entity and the extended entity; There is a second entity relationship between the extended entities; There is a second entity relationship between the core entity and the core entity.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor executes the method for enabling AI intelligent analysis using a large language model as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Knowledge-driven business operation graph construction method

    CN112507136A

  • Atlas extension method and device, electronic equipment and computer readable storage medium

    CN116401370A

  • Intelligent operation and maintenance management method and system based on knowledge graph

    CN116611813A

  • Electric power vector knowledge base enhanced retrieval method and system based on artificial intelligence

    CN118964648A

  • Method, apparatus and device for extracting information

    US20190122145A1