Data processing method and device based on large language model and computer program product

By analyzing and fusion data using large language models to generate high-accuracy knowledge graphs, the problems of low accuracy and insufficient generalization in the existing technology are solved, and the construction efficiency of graph data and information acquisition efficiency are improved.

CN120163227APending Publication Date: 2025-06-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320850.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When building a knowledge graph, the prior art has low accuracy and insufficient generalization, and relies on manual sorting of factors and relationships.

Method used

The data processing method based on the large language model is adopted to analyze the data to be parsed in the data set through the large language model, determine the relationship between entities, generate graph data, and fuse the graph data to generate the full amount of graph data.

Benefits of technology

It improves the efficiency and accuracy of the construction of graph data, ensures that the target object can quickly obtain the relationship between entities, and improves the efficiency of information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163227A_ABST
    Figure CN120163227A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device based on a large language model, electronic equipment, a storage medium and a computer program product, relates to the technical field of computers, in particular to the technical field of artificial intelligence large models, natural language understanding and knowledge maps, and can be applied to a knowledge map construction scene. According to the specific implementation scheme, the method comprises the steps of analyzing to-be-analyzed data in a data set through a large language model, and determining atlas data representing the relationship between entities in the to-be-analyzed data to obtain an atlas data set; fusing the atlas data in the atlas data set to obtain full-amount atlas data; and determining a relationship between the target type entities in the total atlas data, and generating relationship atlas data. According to the method, the construction efficiency and accuracy of the graph data are improved, the accuracy of the relation graph data is ensured, meanwhile, the target object can rapidly obtain the relation between the target type entities based on the relation graph data, and the information obtaining efficiency of the target object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, specifically to the fields of artificial intelligence large models, natural language understanding, and knowledge graph technologies. In particular, it relates to a data processing method, apparatus, electronic device, storage medium, and computer program product based on a large language model, which can be applied to the scenario of knowledge graph construction. Background Art

[0002] In certain specific application scenarios, it is often necessary to parse unstructured data such as natural language texts and images, and determine the relationships between various elements contained therein to construct a graph. Currently, element extraction is mainly carried out according to templates or keyword guidance, and the operation of establishing the relationships between elements depends on manual sorting of clear boundary ranges, resulting in low accuracy and insufficient generalization in the process of constructing graph data. Summary of the Invention

[0003] The present disclosure provides a data processing method, apparatus, electronic device, storage medium, and computer program product based on a large language model.

[0004] According to a first aspect, there is provided a data processing method based on a large language model, including: parsing, through the large language model, the data to be parsed in a data set to determine graph data representing the relationships between entities in the data to be parsed, obtaining a graph data set; fusing the graph data in the graph data set to obtain full-scale graph data; determining the relationships between target type entities in the full-scale graph data to generate relationship graph data.

[0005] According to a second aspect, there is provided a data processing apparatus based on a large language model, including: a parsing unit configured to parse, through the large language model, the data to be parsed in a data set to determine graph data representing the relationships between entities in the data to be parsed, obtaining a graph data set; a fusing unit configured to fuse the graph data in the graph data set to obtain full-scale graph data; a generating unit configured to determine the relationships between target type entities in the full-scale graph data to generate relationship graph data.

[0006] According to a third aspect, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.

[0007] According to a fourth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in any implementation manner of the first aspect.

[0008] According to a fifth aspect, there is provided a computer program product, including: a computer program which, when executed by a processor, implements the method described in any implementation manner of the first aspect.

[0009] According to the technology of the present disclosure, there is provided a data processing method and apparatus based on a large language model. By means of the large language model, the data to be parsed in the data set is parsed to determine the graph data representing the relationships between entities in the data to be parsed, obtaining a graph data set. Based on the powerful natural language understanding ability and logical analysis ability of the large language model, the construction efficiency and accuracy of various graph data are improved; the graph data with high accuracy in the graph data set is fused to obtain the full-scale graph data, and the relationships between the target type entities in the full-scale graph data are determined to generate relationship graph data. While ensuring the accuracy of the relationship graph data, the target object can quickly obtain the relationships between the target type entities based on the relationship graph data, improving the information acquisition efficiency of the target object.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is an exemplary system architecture diagram to which an embodiment according to the present disclosure can be applied;

[0013] Figure 2 is a flowchart of an embodiment of the data processing method based on a large language model according to the present disclosure;

[0014] Figure 3 a partial legend under the graph data modeling rules according to this embodiment;

[0015] Figure 4 is a schematic diagram of the process of a partial fusion operation according to this embodiment;

[0016] Figure 5 is a schematic diagram of an application scenario of the data processing method based on a large language model according to this embodiment;

[0017] Figure 6 is a schematic diagram of the 7-layer hierarchical structure relationship data in the field of criminal investigation according to this embodiment;

[0018] Figure 7 is a schematic diagram of the overall architecture of the data processing method based on a large language model according to the present disclosure;

[0019] Figure 8 is a flowchart of another embodiment of the data processing method based on a large language model according to the present disclosure;

[0020] Figure 9 is a structural diagram of an embodiment of the data processing apparatus based on a large language model according to the present disclosure;

[0021] Figure 10 is a schematic structural diagram of a computer system suitable for implementing the embodiments of the present disclosure. Detailed implementation manners

[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0023] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0024] Figure 1 Illustrates an exemplary architecture 100 to which the data processing method and apparatus based on a large language model according to the present disclosure can be applied.

[0025] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The terminal devices 101, 102, 103 are communicatively connected to form a topology network, and the network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0026] The terminal devices 101, 102, and 103 can be hardware devices or software that support network connections for data interaction and data processing. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connections and functions such as information acquisition, interaction, display, and processing, including but not limited to smartphones, tablets, e-book readers, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or can also be implemented as a single software or software module. No specific limitation is made here.

[0027] The server 105 can be a server that provides various services. For example, it can obtain the data to be parsed provided by the target object through the terminal devices 101, 102, and 103, determine the graph data representing the relationships between entities in the data to be parsed based on the large language model, extract the relationships between target type entities in each graph data, and generate the relationship graph data. Optionally, the server can feedback the relationship graph data to the terminal device. As an example, the server 105 can be a cloud server.

[0028] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can also be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or can also be implemented as a single software or software module. No specific limitation is made here.

[0029] It also needs to be noted that the data processing method based on the large language model provided by the embodiments of the present disclosure is generally executed by the server, but it does not exclude the possibility of being executed by the terminal device, or being executed by the server and the terminal device in cooperation with each other. Correspondingly, each part (such as each unit) included in the data processing device based on the large language model can be all set in the server, or can be all set in the terminal device, or can also be respectively set in the server and the terminal device.

[0030] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0031] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, the network, and the server. When the electronic device on which the data processing method based on the large language model runs does not need to perform data transmission with other electronic devices, the system architecture can only include the electronic device (such as the terminal device or the server) on which the data processing method based on the large language model runs.

[0031] Please refer to Figure 2 ,Figure 2 The flowchart of a data processing method based on a large language model provided by an embodiment of the present disclosure. Among them, in process 200, the following steps are included:

[0032] Step 201: Parse the data to be parsed in the data set through the large language model, determine the graph data representing the relationships between entities in the data to be parsed, and obtain a graph data set.

[0033] In this embodiment, the execution subject of the data processing method based on the large language model (for example, Figure 1 the server in) can obtain the graph data set from a remote or local location through a wired network connection method or a wireless network connection method, and parse the data to be parsed in the data set through the large language model, determine the graph data representing the relationships between entities in the data to be parsed, and obtain a graph data set.

[0034] The target object is, for example, an object such as a person, other intelligent devices, an artificial intelligence assistant, etc. The data to be parsed is generally unstructured data in the form of text, images, etc. It is necessary to parse the data to be parsed in the data set based on the powerful natural language understanding ability and logical analysis ability of the large language model to determine the graph data representing the relationships between entities in the data to be parsed.

[0035] The data to be parsed can be unstructured data in various application fields, including but not limited to the medical and health field, the financial field, the intelligent customer service field, the Internet of Things field, the social network field, the e-commerce field, the government management and security field.

[0036] In the medical and health field, the data to be parsed is, for example, medical records, medical images, medical literature, etc., and the entities are, for example, diseases, symptoms, drugs, treatment methods, medical devices, medical institutions, medical literature, etc. Build a medical knowledge graph, integrate the knowledge and entities in the medical field, mine medical literature and medical records through natural language processing technology, further extract various information such as text and pictures, promote the sharing and exchange of medical information, improve the accuracy and efficiency of disease diagnosis, and can also be applied to medical image diagnosis, combining different medical images with structured medical knowledge for automatic diagnosis and analysis.

[0037] In the financial field, the data to be parsed is, for example, financial news and social media data, transaction logs and customer service conversation records, contract documents and reports, etc., and the entities are, for example, customers, companies, securities, financial instruments, events, transaction records, financial institutions, etc. In business scenarios such as credit risk assessment and investment decision-making, it helps decision-makers quickly identify the relationships between entities such as customers, companies, and securities, mine industry information and asset value, and improve the accuracy and efficiency of decision-making.

[0038] In the field of intelligent customer service, the data to be parsed are, for example, user consultation records and feedback, product manuals, and frequently asked questions. Entities are, for example, user questions, product information, service content, solutions, and frequently asked questions. By establishing a natural language understanding and generation model, rapid matching and intelligent recommendation based on user needs and service content can be achieved. When a user calls the customer service for consultation, the knowledge graph technology can automatically match the existing data and inventory information through semantic analysis and achieve a certain degree of automatic response.

[0039] In the field of the Internet of Things, the data to be parsed are, for example, sensor data and device logs. Entities are, for example, devices, items, scenarios, sensor data, and user behaviors. By using devices such as sensors and cameras of intelligent hardware to model devices, items, scenarios, etc., and realizing the sharing of cloud information, an Internet of Things knowledge graph is constructed to achieve a new ecosystem that connects everything based on technology.

[0040] In the field of social networks, the data to be parsed are, for example, the content posted by users and user relationship data. Entities are, for example, users, social groups, interest tags, geographical locations, and interaction records. It is used to analyze the relationships in social networks, discover community structures, identify key nodes, etc., which is very important for fields such as social media, recommendation systems, and marketing.

[0041] In the e-commerce field, the data to be parsed are, for example, user evaluations and feedback, product pictures and descriptions. Entities are, for example, products, brands, merchants, user evaluations, purchase records, and promotional activities. By constructing a product knowledge graph to accurately match the user's purchase intention and the set of product candidates, it is widely applied to services such as search, shopping guide, platform governance, and intelligent question answering.

[0042] In the fields of government management and security, the data to be parsed are, for example, government documents and policies and regulations, security monitoring videos and intelligence reports. More specifically, for example, data such as case transcripts, case situations, on-site investigations, and police situations. Entities are, for example, personnel, organizations, events, geographical information, and policies and regulations. By analyzing the relationships between entities to obtain clues, etc., for example, in military and criminal intelligence analysis systems, multi-source heterogeneous information is integrated to conduct all-round real-time monitoring and analysis of personnel, equipment, and events, enabling dispatchers to grasp the battlefield situation in the first time and make predictions.

[0043] First, preprocess each piece of data to be parsed in the data set, including but not limited to:

[0044] OCR (Optical Character Recognition): For unstructured text data in image form (such as text in scanned documents, pictures, etc.), OCR technology needs to be used to convert it into an editable and parsable text format.

[0045] PDF (Portable Document Format) parsing: PDF documents usually contain various elements such as text, images, and tables, and specialized parsing techniques are required to extract the text information therein. This includes using PDF parsing libraries, OCR technology, and natural language processing technology to identify and extract the text content.

[0046] Then, based on the large language model, natural language understanding is performed on each data to be parsed in the data set, and the natural language understanding results are obtained. The specific operations include the following:

[0047] Word segmentation and part-of-speech tagging: This is the basic step of text parsing. The text is segmented into independent lexical units through word segmentation technology, and the part of speech of each word is determined through part-of-speech tagging, providing a basis for subsequent processing.

[0048] Named entity recognition: Identifying entities with specific meanings in the text, such as person names, place names, and organization names, helps to understand the text content.

[0049] Syntactic analysis: Analyzing the sentence structure, identifying the subject, predicate, object, etc. in the sentence, and their relationships, so as to understand the grammatical structure of the sentence.

[0050] Semantic analysis: Further understanding the meaning of the text, including identifying synonyms, antonyms, and understanding the context relationship, etc. This helps to more accurately grasp the information conveyed by the text.

[0051] Finally, for the operations of entity and entity attribute parsing, and relationship parsing between entities in the data to be parsed, multiple Prompts are generated respectively, and the Prompts are input into the large language model item by item for reasoning to determine the entities and entity attribute data, and the relationship data between entities in the data to be parsed. Taking the transcript data in the field of criminal investigation as an example, the Prompt is, for example, "According to the transcript content, find all [phone numbers], and the related attributes of the phone numbers (including related attributes such as phone numbers, holder names, holder identity numbers, etc.), and list the phone numbers item by item."

[0052] The large language model can be trained in the following ways:

[0053] First, key features in the text are extracted through machine learning algorithms. These features can be words, phrases, syntactic structures, etc., and are used for subsequent model training and prediction. Then, a large language model is trained using a large amount of labeled data in the required application domain, enabling it to identify and understand key information in unstructured text data. Finally, using specific labeled data corresponding to the target object, the model that has been trained in the relevant application domain is fine-tuned to adapt to more specific unstructured text data parsing tasks. This can greatly improve the development efficiency and performance of the model.

[0054] In some alternative implementation manners of this embodiment, the above-mentioned execution subject may execute the above step 201 in the following manner:

[0055] In the first step, through the large language model, the data to be parsed in the data set is parsed to determine the graph data representing the relationships between entities in the data to be parsed, and an initial graph data set is obtained.

[0056] In this implementation manner, the above-mentioned execution subject may generate the initial graph data set with reference to the construction process of the graph data in step 201 above, which will not be elaborated here.

[0057] In the second step, the graph data in the initial graph data set is combined with the graph data specified by the target object to obtain a graph data set.

[0058] The target object may specify at least one graph data according to its needs, and combine it with the graph data in the initial graph data set to obtain a graph data set. Taking the field of criminal investigation as an example, the graph data specified by the target object may be the original graph data obtained from previous cases, and it is fused with the graph data corresponding to the current cases.

[0059] In this implementation manner, the target object may specify existing graph data according to its own needs, so as to form a graph data set with the parsed graph data, thereby obtaining graph data with a more abundant amount of information through the subsequent graph data construction process, improving the flexibility of the graph construction process and the experience of the target object.

[0060] Step 202: Fuse the graph data in the graph data set to obtain full-scale graph data.

[0061] In this embodiment, the above-mentioned execution subject may fuse the graph data in the graph data set to obtain full-scale graph data.

[0062] As an example, first, the data of each graph is cleaned and formatted uniformly to remove duplicate, erroneous and redundant data to ensure the accuracy and consistency of the data. Then, natural language processing technology is used to identify and extract entities in each graph. Then, through the entity alignment algorithm, the entities referring to the same thing in different graphs are matched and associated to establish a unified entity identification. Then, the relationship between entities in each graph is analyzed, and the same or similar relationships are integrated, while retaining the unique relationships in different graphs. In the fusion process, factors such as the weight and credibility of the relationship can be considered to improve the accuracy and reliability of the full graph. Finally, the aligned entities and the fused relationships are integrated into a new graph to form the full graph data. At the same time, the structure of the graph is optimized to remove redundant edges and nodes to improve the readability and query efficiency of the graph.

[0063] As another example, first, a graph convolutional neural network (GCN) model suitable for graph data is designed, which can learn node features and edge features in graph structure data and capture complex dependencies between entities and relationships. Then, the entities and relationships in each graph are embedded in the same vector space, and through the learning process of the graph neural network, the entities and relationships in different graphs have similar representations in the embedding space. Then, the entities and relationships in different graphs are aligned using the similarity of the embedded vectors. Then, based on the embedded graph data, a graph matching algorithm is used to match different graphs to find the corresponding relationship between them. Then, the matched graph data is fused to generate a new full-volume graph. In the fusion process, technologies such as the graph attention mechanism can be used to further improve the accuracy and efficiency of the fusion. Finally, the fused graph data is constructed into a complete full-volume graph, and it is verified and evaluated to ensure that it can accurately reflect the overall picture and internal connection of each graph data. The quality and availability of the full-volume graph can be verified by query testing, visual analysis, etc.

[0064] In this embodiment, by analyzing the relationship analysis scenario requirements and domain data characteristics in specific business fields, graph data modeling rules that are more in line with the business fields are summarized. The graph data modeling rules support multiple import methods such as API (Application Programming Interface) and Kafka messages.

[0065] As an example, graph data modeling rules include the following principles:

[0066] 1. The entity nodes need to have unique identifiers and can be uniquely indexed in the graph data space. The rule for unique identifiers is "type + business identifier". Taking the field of criminal investigation as an example, in this embodiment, the naming rules for unique identifiers of more than 20 common entities in the field of criminal investigation are sorted out. For example, the unique identifier of a person node is "person + identity number".

[0067] 2. The node attributes only include the key information of the nodes, which is convenient for attribute retrieval and reduces the possibility of modifying the schema, because some attributes may need to be modeled as separate nodes as the business evolves. Modeling these nodes as node attributes in advance will cause schema changes and greatly increase the maintenance cost.

[0068] 3. The edges between nodes are generally divided into two types: attribute edges and behavior edges. Each edge fixedly includes attributes such as information source, creation time, and update time. This means that edges naturally have time attributes. This feature is mainly because attributes such as a person's residential address and work unit have the characteristic of changing over time, and the time attribute is very important in the process of analyzing relationships in a certain past time period.

[0069] Continue to refer to Figure 3 , which shows a partial legend 300 under the above graph data modeling rules.

[0070] In some alternative implementation manners of this embodiment, the above execution subject may execute the above step 202 in the following manner:

[0071] The first step is to determine the graph data to be fused from the graph data set.

[0072] The graph data set generally includes multiple graph data. For example, the above execution subject may generate a graph data sequence based on the graph data set to sequentially determine a graph data to be fused from the graph data sequence; or, the above execution subject may randomly determine a graph data from the graph data set as the graph data to be fused.

[0073] It can be understood that the determined graph data to be fused is generally graph data that has not undergone the fusion operation. The above execution subject may use a specific identifier to indicate the graph data that has undergone the fusion operation and the graph data that has not undergone the fusion operation in the graph data set, so as to determine the graph data to be fused from the graph data that has not undergone the fusion operation in the graph data set. Or, the above execution subject may delete the graph data that has undergone the fusion operation from the graph data set.

[0074] The second step is to determine the sub-graph data that has an associated relationship with the graph data to be fused from the graph data that has completed the fusion operation as of the current time.

[0075] The sub-graph data is at least part of the graph data in the graph spectrum data. As the fusion operation progresses, the number of graph data for which the fusion operation has been completed up to the current time gradually increases.

[0076] For each graph data for which the fusion operation has been completed up to the current time, the above-mentioned execution entity can determine whether there is an association relationship between the graph data and the graph data to be fused. In response to determining yes, sub-graph data associated with the graph data to be fused is determined from the graph data. Among them, existing associations include, for example, the same entities, the same relationships, etc.

[0077] The third step is to fuse the sub-graph data and the graph data to be fused to obtain fused sub-graph data.

[0078] In this implementation manner, the above-mentioned execution entity can fuse the sub-graph data and the graph data to be fused according to the associated entities and relationships between the sub-graph data and the graph data to be fused to obtain fused sub-graph data.

[0079] The fourth step is to update the current fused graph data according to the fused sub-graph data.

[0080] The fused graph data represents the relationships between entities in the graph data for which the fusion operation has been completed up to the current time. As an example, the above-mentioned execution entity can replace the part of the current fused graph data that has information overlap with the fused sub-graph data with the fused sub-graph data to obtain the updated fused graph data.

[0081] In this implementation manner, the above-mentioned execution entity can iteratively execute the fusion operations of the above-mentioned first step to the fourth step until the fusion operation of the graph data in the graph data set is completed, and determine the obtained fused graph data as the full-scale graph data.

[0082] Since at least two graph data are required to perform the fusion operation, therefore, in the first fusion operation, for the first graph data to be fused determined from the graph data set, there is no graph data for which the fusion operation has been completed up to the current time. At this time, the first graph data to be fused can be used as the current fused graph data in the first fusion operation.

[0083] In the second fusion operation, the second graph data is determined from the graph data set. The graph data for which the fusion operation has been completed up to the current time is the first graph data. Determine whether there is sub-graph data in the first graph data that is associated with the second graph data. If there is, fuse the sub-graph data and the second graph data to obtain fused sub-graph data, and update the current fused graph data (the first graph data) with the fused sub-graph data.

[0084] The above-mentioned execution entity can perform iterative execution with reference to the above process until the fusion operation of the graph data in the graph data set is completed, and the obtained fused graph data is determined as the full-scale graph data.

[0085] In this implementation manner, a specific implementation manner of graph fusion is provided. During the iterative process, the sub-graph data associated with the graph data to be fused is determined from the already fused graph data, and the graph is updated based on the fused sub-graph data obtained by fusing the two, improving the determination efficiency and accuracy of the full-scale graph data.

[0086] In some optional implementation manners of this embodiment, the above-mentioned execution entity can execute the above-mentioned second step (the determination process of the fused sub-graph data) in the following manner: According to the priorities and data update times of the data sources corresponding to the sub-graph data and the graph data to be fused respectively, the sub-graph data and the graph data to be fused are fused to obtain the fused sub-graph data.

[0087] During the fusion process of the graph data, there may be a situation of data inconsistency. At this time, it is necessary to determine the data with high confidence according to the priorities and data update times of the data sources corresponding to the sub-graph data and the graph data to be fused respectively.

[0088] For the graph data of different data sources, the party with a higher priority has a higher confidence; for the graph data with the same priority, the party with a smaller time difference between the data update time and the current time has a higher confidence.

[0089] In this implementation manner, a relationship table representing the corresponding relationship between different data sources and priorities is set in the above-mentioned execution entity or an electronic device communicatively connected to the above-mentioned execution entity, so that the above-mentioned execution entity can quickly determine the priority corresponding to the data source.

[0090] The data sources in different application scenarios are different, and it is necessary to determine their priorities specifically for different application scenarios. As an example, in the field of criminal investigation, the data sources include the case situation, transcripts, on-site inspections, and police situations, and their priorities generally decrease in the order of: on-site inspections, case situations, police situations, transcripts.

[0091] In this implementation manner, fully considering the priorities and data update times of the data sources corresponding to the sub-graph data and the graph data to be fused respectively during the graph data fusion process helps to improve the accuracy of the obtained fused sub-graph data.

[0092] In some optional implementation manners of this embodiment, the above-mentioned execution entity can execute the above-mentioned determination process of the fused sub-graph data in the following manner:

[0093] First, for target entities with consistent key attribute data in the sub-graph data and the graph data to be fused, in response to the sub-graph data and the graph data to be fused belonging to different data sources, and both the sub-graph data and the graph data to be fused including non-key attribute data of the target entity under the same non-key attribute, determine the fusion method of the non-key attribute data according to the type of the non-key attribute.

[0094] The fusion methods generally include overwriting and appending.

[0095] Target entities with consistent key attribute data in the sub-graph data and the graph data to be fused are generally the same entity. The non-key attribute data of the same entity may be different in the sub-graph data and the graph data to be fused. For example, for the non-key attribute data of the same entity, the attribute data under the same non-key attribute included in the sub-graph data is different from the attribute data under the same non-key attribute included in the graph data to be fused.

[0096] As an example, in response to the sub-graph data and the graph data to be fused belonging to different data sources, and both the sub-graph data and the graph data to be fused including non-key attribute data of the target entity under the same non-key attribute, perform intelligent analysis through a large language model according to the type of the non-key attribute to determine the fusion method of the non-key attribute data.

[0097] As another example, a relationship table representing the correspondence between the type of the non-key attribute and the fusion method is set in the above-mentioned execution subject or an electronic device communicatively connected to the above-mentioned execution subject, and the above-mentioned execution subject can determine the fusion method of the non-key attribute data through the relationship table.

[0098] Then, in response to the fusion method being overwriting, determine the target non-key attribute data from the non-key attribute data included in the sub-graph data and the graph data to be fused respectively according to the priority and the data update time, and use it as the attribute data of the target entity in the sub-graph data to be fused under the non-key attribute.

[0099] In this implementation, in response to the fusion method being overwriting, determine the target non-key attribute data with the highest execution degree from the non-key attribute data included in the sub-graph data and the graph data to be fused respectively according to the priority and the data update time, and use it as the attribute data of the target entity in the sub-graph data to be fused under the non-key attribute. For example, there is generally only one attribute data under non-key attributes such as a person's name and age, and the corresponding fusion method is overwriting.

[0100] In this implementation, a specific fusion method for entities in the graph data is provided, which further improves the accuracy of the sub-graph data to be fused by combining the fusion method, the priority, and the data update time.

[0101] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may also execute the above-mentioned determination process of the fused sub-graph data in the following manner: in response to the fusion manner being append, combine the non-critical attribute data included in the sub-graph data and the graph data to be fused, and obtain the attribute data of the target entity in the fused sub-graph data under the non-critical attribute.

[0102] In this implementation manner, in response to the fusion manner being append, combine the non-critical attribute data of the sub-graph data and the graph data to be fused under the same non-critical attribute, and obtain the attribute data of the target entity in the fused sub-graph data under the non-critical attribute. For example, there may be multiple pieces of attribute data under non-critical attributes such as a person's bank card number and mobile phone number, and the corresponding fusion manner is append.

[0103] In this implementation manner, another fusion manner of the entities in the graph data is provided, and by combining the fusion manner, priority, and data update time, the accuracy of the fused sub-graph data is further improved.

[0104] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may also execute the above-mentioned determination process of the fused sub-graph data in the following manner: for the target entity with consistent critical attribute data in the sub-graph data and the graph data to be fused, in response to the sub-graph data and the graph data to be fused belonging to the same data source, determine the target non-critical attribute data from the non-critical attribute data included in the sub-graph data and the graph data to be fused respectively according to the respective data update times of the sub-graph data and the graph data to be fused, and use it as the attribute data of the target entity in the fused sub-graph data under the non-critical attribute.

[0105] As an example, the above-mentioned execution entity determines the target non-critical attribute data with the highest confidence from the non-critical attribute data included in the sub-graph data and the graph data to be fused respectively according to the respective data update times of the sub-graph data and the graph data to be fused, and uses it as the attribute data of the target entity in the fused sub-graph data under the non-critical attribute.

[0106] In this implementation manner, another fusion manner of the entities in the graph data is provided, and by combining the fusion manner, priority, and data update time, the accuracy of the fused sub-graph data is further improved.

[0107] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may also execute the above-mentioned determination process of the fused sub-graph data in the following manner: for the relationship between the head entity and the tail entity, in response to the first relationship data between the head entity and the tail entity in the sub-graph data being consistent with the second relationship data between the head entity and the tail entity in the graph data to be fused, merge the first relationship data and the second relationship data, and use it as the relationship data between the head entity and the tail entity in the fused sub-graph data.

[0108] The merging of the first relational data and the second relational data can be performed only when they are the same. The determination conditions for the first relational data and the second relational data to be the same generally include: the first relational data and the corresponding head entity and tail entity of the first relational data are the same as the second relational data and the corresponding head entity and tail entity of the second relational data in sequence.

[0109] In this implementation, a specific fusion method for the relationships in the graph data is provided, which further improves the accuracy of the fused sub-graph data.

[0110] In some alternative implementation manners of this embodiment, the above execution subject may also execute the above process of determining the fused sub-graph data in the following manner: in response to the target relational data between the head entity and the tail entity being included in the sub-graph data or the graph data to be fused, the target relational data is used as the relational data between the head entity and the tail entity in the fused sub-graph data.

[0111] When only one of the sub-graph data and the graph data to be fused includes the target relational data between the head entity and the tail entity, the target relational data is used as the relational data between the head entity and the tail entity in the fused sub-graph data.

[0112] In this implementation, another fusion method for the relationships in the graph data is provided, which further improves the accuracy of the fused sub-graph data.

[0113] In some alternative implementation manners of this embodiment, the above execution subject may also execute the above fourth step (the update process of the fused graph data) in the following manner:

[0114] First, the entities included in the sub-graph data and the graph data to be fused respectively, and the relational data between the entities are deleted from the current fused graph data to obtain the graph data to be supplemented.

[0115] As an example, the above execution subject first determines the entities included in the sub-graph data and the graph data to be fused respectively, and the relational data between the entities; then, the determined entities and the relational data between the entities are deleted from the current fused graph data to obtain the graph data to be supplemented.

[0116] Then, the graph data to be supplemented and the fused sub-graph data are fused to obtain the updated fused graph data.

[0117] As an example, the above execution subject may merge the fused sub-graph data into the graph data to be supplemented to obtain the updated fused graph data.

[0118] In this implementation manner, by first deleting the entities included in the sub-graph data and the graph data to be fused respectively, as well as the relationship data between the entities, and then merging and fusing the sub-graph data, the convenience and update efficiency of the determination process of the fused graph data are improved.

[0119] Step 203: Determine the relationships between the target type entities in the full-scale graph data, and generate relationship graph data.

[0120] In this embodiment, the above execution subject can determine the relationships between the target type entities in the full-scale graph data and generate relationship graph data. The target type entity can be any type of entity among multiple types of entities in the full-scale graph data, and the number of target types can be one or multiple.

[0121] As an example, first, clarify the target type entities to be extracted, and use the named entity recognition algorithm in natural language processing technology, combined with domain knowledge and grammar rules, to accurately identify and extract these target type entities from the full-scale graph data. Then, analyze the semantic relationships between the target type entities. By constructing semantic models such as word vector models and semantic networks, understand the association methods and semantic meanings between the entities, and identify which entities have potential relationships. Then, according to the semantic analysis results, combined with predefined relationship patterns and rules, screen out the relationships that meet the requirements between the target type entities in the full-scale graph. For example, if the target type entities are "person" and "organization", then relationships such as the positions held by "person" in the "organization" can be extracted. Finally, construct the extracted target entities and their relationships into a new relationship graph data. During the construction process, optimize the graph, remove redundant relationships and nodes, ensure the simplicity and efficiency of the graph, and at the same time verify the accuracy and integrity of the graph.

[0122] As another example, first, for the graph structure corresponding to the full-scale graph data (where nodes represent entities and edges represent the relationships between entities), corresponding feature vectors are assigned to each node and edge. These features can include information such as the attributes of the entities and the types of relationships. Then, a graph neural network model suitable for relationship extraction is designed, such as a graph convolutional network (GCN), a graph attention network (GAT), the above-mentioned large language model, etc. By training the model, the embedding representations of the nodes and edges are learned to capture the complex dependencies between entities and relationships. Then, the target type entities and their related relationships are embedded into the same vector space. Using the trained graph neural network model, the relationships between the target entities are predicted and classified to obtain the possible relationship types and strengths between them. Finally, according to the prediction results, the relationship graph data between the target type entities is generated. The generated relationship graph can be verified and evaluated, for example, by comparing it with the known relationship graph, expert evaluation, etc., to ensure the accuracy and reliability of the extracted relationship graph data.

[0123] In some alternative implementation manners of this embodiment, the above-mentioned execution subject may execute the above step 203 in the following manner:

[0124] The first step is to determine the relationship data between the target type entities in the fused subgraph data and generate target relationship subgraph data.

[0125] In this implementation manner, the above-mentioned execution subject may refer to the manner of determining the relationships between the target type entities in the full-scale graph data in the above step 203 to determine the relationship data between the target type entities in the fused subgraph data, so as to generate target relationship subgraph data, which will not be elaborated here.

[0126] The second step is to update the current fused relationship graph data according to the fused subgraph data until the fusion operation on the graph data in the graph data set is completed, and the obtained fused relationship graph data is determined as the relationship graph data.

[0127] Among them, the fused relationship graph data represents the relationships between the target type entities in the graph data for which the fusion operation has been completed up to the current time.

[0128] In this implementation manner, during the iterative execution of the fusion operation, the above-mentioned execution subject concurrently executes the process of determining the relationship graph data.

[0129] As an example, in the first fusion operation, for the first data to be fused determined from the graph data set, the current fused relationship graph data represents the relationship data between the target type entities in the first data to be fused.

[0130] In the second fusion operation, a second graph data is determined from the graph data set. The graph data for which the fusion operation has been completed up to the current time is the first graph data. It is determined whether there is sub-graph data in the first graph data that is associated with the second graph data. If so, the sub-graph data and the second graph data are fused to obtain fused sub-graph data. The relationship data between the target type entities in the fused sub-graph data is determined to generate target relationship sub-graph data. The current fused relationship graph data, that is, the relationship data between the target type entities in the first data to be fused, is updated according to the target relationship sub-graph data.

[0131] Continue to refer to Figure 4 , which shows a schematic diagram of the process of partial fusion operation.

[0132] For the graph data 401 to be fused, the sub-graph data 402, 403, and 404 associated with the graph data 401 to be fused are determined; the graph data 401 - 404 are fused to obtain fused graph data 405, and it is stored in the full-scale graph data 406; and the relationships between the target type entities are determined from the fused graph data 405 to obtain fused relationship graph data 407.

[0133] In this implementation manner, a specific implementation manner of graph fusion is provided. During the iteration process, based on the relationship data between the target type entities in the fused sub-graph data, the fused relationship graph data is updated, improving the determination efficiency and accuracy of the relationship graph data.

[0134] In some alternative implementation manners of this embodiment, the above execution subject may execute the above second step (the update process of the fused relationship graph data) in the following manner:

[0135] First, the entities included in the sub-graph data and the graph data to be fused, as well as the relationship data between the entities, are deleted from the fused relationship graph data to obtain relationship graph data to be supplemented.

[0136] As an example, the above execution subject determines the entities included in the sub-graph data and the graph data to be fused, as well as the relationship data between the entities, and then deletes the determined entities and the relationship data between the entities from the fused relationship graph data to obtain relationship graph data to be supplemented.

[0137] Then, the relationship graph data to be supplemented and the target relationship sub-graph data are fused to obtain updated fused relationship graph data.

[0138] As an example, the above execution subject may merge the target relationship sub-graph data into the relationship graph data to be supplemented to obtain updated fused relationship graph data.

[0139] In this implementation manner, by first deleting the entities included in the sub-graph data and the graph data to be fused respectively, as well as the relationship data between the entities, and then merging the target relationship sub-graph data, the convenience and update efficiency of the determination process of the fused relationship graph data are improved.

[0140] In some optional implementation manners of this embodiment, the above-mentioned execution subject may execute the above-mentioned first step (the determination process of the target relationship sub-graph data) in the following manner:

[0141] First, determine the connected paths between the target type entities in the fused sub-graph data.

[0142] The target type entities may be connected by one or more edges to form a connected path.

[0143] Then, according to the preset connected path rule, screen the target connected paths from the connected paths.

[0144] As an example, the above-mentioned execution subject may determine the possible meta-paths between the target type entities according to the specific requirements in the application scenario, clarify the nodes and edges in the meta-path, so as to indicate that the relationship data between the determined target type entities is the relationship data represented by the meta-path.

[0145] Furthermore, based on the meta-path, determine the preset connected path rule to screen out the target connected paths that meet the preset connection rule from the connected paths. Taking the e-commerce user graph as an example, the preset connected path rule is used to represent that two users have purchased the same commodity, and its corresponding meta-path: user - purchase - commodity - purchase - user.

[0146] Finally, generate the target relationship sub-graph data according to the relationship data between the target type entities represented by the target connected paths.

[0147] In this implementation manner, extract the relationship data between the target type entities represented by the target connected paths from the fused sub-graph data to generate the target relationship sub-graph data.

[0148] In this implementation manner, a method for generating the target relationship sub-graph data based on the preset connected path rule is provided. The target object can flexibly limit the relationship data between the target type entities to be obtained by editing the preset connected path rule, which improves the adaptability of the target relationship sub-graph data to the user requirements.

[0149] Continue to refer to Figure 5 , Figure 5FIG. 500 is a schematic diagram of an application scenario of a data processing method based on a large language model according to this embodiment. First, the server 501 first obtains a data set from a database; then, through the large language model, each data to be parsed in the data set is parsed to determine graph data representing the relationships between entities in the data to be parsed, obtaining a graph data set; then, the graph data in the graph data set is fused to obtain full-scale graph data; finally, the relationships between target type entities in the full-scale graph data are determined to generate relationship graph data, and the relationship graph data is fed back to the terminal device 503 of the target user 502.

[0150] In this embodiment, a data processing method based on a large language model is provided. Through the large language model, the data to be parsed in the data set is parsed to determine graph data representing the relationships between entities in the data to be parsed, obtaining a graph data set. Based on the powerful natural language understanding ability and logical analysis ability of the large language model, the construction efficiency and accuracy of various graph data are improved; the highly accurate graph data in the graph data set is fused to obtain full-scale graph data, and the relationships between target type entities in the full-scale graph data are determined to generate relationship graph data. While ensuring the accuracy of the relationship graph data, it enables the target object to quickly obtain the relationships between target type entities based on the relationship graph data, improving the information acquisition efficiency of the target object.

[0151] In some optional implementation manners of this embodiment, the above execution entity may further perform the following operations: According to the relationship query request of the target object, determine the hierarchical structure relationship data including the target query entity from the relationship graph data.

[0152] The hierarchical structure relationship data represents the multi-level structure relationships between target type entities including the target query entity.

[0153] Continue to refer to Figure 6 , which shows a schematic diagram of the hierarchical structure relationship data at 7 levels in the field of criminal investigation. The 7 levels include, from bottom to top, the bottom layer, the zero-pack layer, the distribution layer, the transportation layer, the channel layer, the transition layer, and the top layer.

[0154] In this implementation manner, the hierarchical structure relationship data can be determined and displayed based on the relationship query request of the target object, enabling the target object to quickly understand the hierarchical structure relationship including the target query entity, which helps to further improve the information acquisition efficiency of the target object.

[0155] In some optional implementation manners of this embodiment, the above execution entity may further perform the following operations: First, according to the custom request of the target object for the graph analysis algorithm, determine the custom graph analysis algorithm; then, according to the custom graph analysis algorithm, analyze the full-scale graph data and / or the relationship graph data.

[0156] Continue to refer to Figure 7 , which shows a schematic diagram of the overall architecture of the data processing method based on the large language model.

[0157] The graph data analysis module for analyzing full-scale graph data and / or relationship graph data includes a real-time analysis module and an offline analysis module. The real-time analysis module provides millisecond-level real-time graph data query and supports Restful WEB API based on Gremlin query statements. The offline analysis module provides various graph analysis algorithm supports based on HugeGraph-Spark and HugeGraph-Computer, and with the support of HugeGraph-Spark, it supports flexible custom graph algorithms. The various graph analysis algorithms include but are not limited to the Louvain algorithm, connected component algorithm, strongly connected component algorithm, PageRank algorithm, breadth-first traversal algorithm, and single-source shortest path algorithm.

[0158] In this implementation, the custom function of the graph analysis algorithm is provided, which can meet the user's various graph analysis needs and further improve the adaptability of the graph analysis process to the user's needs.

[0159] The following is an explanation of the graph database part in the overall architecture:

[0160] The graph database uses HugeGraph. HugeGraph is an open-source graph database system. HugeGraph can store a large amount of vertices and edges, implements the Apache TinkerPop 3 framework, and supports the Gremlin query language. HugeGraph supports multi-user parallel operations. Users can input Gremlin query statements and obtain graph query results in a timely manner. They can also call the HugeGraph API in the user program for graph analysis or query. The main application scenario of this system is to solve the graph data storage and modeling analysis needs of the anti-fraud, threat intelligence, and black production crackdown operations faced by the security department. On this basis, it has been gradually expanded and supports more general graph applications.

[0161] HugeGraph has a complete toolchain component, supports data sharding, supports distributed deployment, and can support data scales of more than ten billion. It supports the rapid import of more than ten billion vertices and edges, and provides millisecond-level correlation query capabilities, that is, OLTP (Online Transaction Processing), and can be integrated with big data platforms such as Hadoop and Spark for offline analysis, that is, OLAP (Online Analytical Processing). It supports the Property Graph and Apache Gremlin query languages, has tool components such as import, export, backup, recovery, and visualization interfaces, provides a simple and easy-to-use RESTful API and Client, and can easily build various graph database-based applications and products.

[0162] HugeGraph mainly includes the following layers:

[0163] * 1. Application layer:

[0164] Hubble: A one-stop visualization analysis platform that covers the whole process from data modeling, to rapid data import, to online and offline data analysis, and to unified graph management, realizing a full-process wizard operation for graph applications.

[0165] Loader: A data import component that can convert data from multiple data sources into vertices and edges of a graph and batch import them into the graph database.

[0166] Tools: Command-line tools for deploying, managing, and backing up / restoring data in HugeGraph.

[0167] Computer: A distributed graph processing system (OLAP), which is an implementation of Pregel and can run on Kubernetes.

[0168] Client: A HugeGraph client written in Java. Users can use the Client to write Java code to operate on HugeGraph, and subsequent multi-language support such as Python, Go, and C++ will be provided according to needs.

[0169] 2. Graph engine layer:

[0170] REST Server: Provides RESTful APIs for querying Graph / Schema and other information, supports Gremlin and Cypher query languages, and provides APIs for service monitoring and operation and maintenance.

[0171] Graph Engine: Supports two types of graph computing, namely OLTP and OLAP. Among them, OLTP implements the Apache TinkerPop3 framework.

[0172] Backend Interface: Implements storing graph data into the backend.

[0173] 3. Storage Abstraction Layer:

[0174] Storage Backend: Supports multiple built-in storage backends (e.g., RocksDB / MySQL / HBase / …), and also allows users to extend custom backends without changing the existing source code.

[0175] In some optional implementation manners of this embodiment, the above execution subject may further perform the following operations: Edit the full-scale graph data and / or relationship graph data according to the editing operation of the target object.

[0176] After obtaining the full-scale graph data and / or relationship graph data, the full-scale graph data and / or relationship graph data can be displayed through the terminal device of the target object. Enter the key attribute data of the complete target type entity in the input box, click the search button to display the qualified list, and click the target type entity to display more detailed graph data.

[0177] The editing operations include but are not limited to adding and deleting operations of target type entities, and adding and deleting operations of the relationships between target type entities.

[0178] Taking the addition operation of the target type entity as an example, the display interface includes input boxes or selection boxes corresponding to entity elements, entity types, key attributes, etc. to instruct the target object to add a target type entity.

[0179] Taking the addition operation of the relationships between target type entities as an example, the target object selects the relationship graph in the display interface, enters the relationship graph page, and the relationship graph data is displayed to the target object. The target object can click the "Add Relationship" button, and a new relationship window will pop up on the page. The target object can further select the head entity, click the drop-down box to select the entity type, the corresponding entity list will be displayed, and the entity attributes will be checked. Then, the target object clicks the "Next" button in the display interface to select the tail entity. Click the drop-down box to select the entity type, the corresponding entity list will be displayed, and the entity attributes will be checked. Then, the target object clicks the "Next" button in the display interface to configure the relationship details. Click "Add Relationship" to configure information such as relationship type and relationship occurrence time.

[0180] In this implementation manner, according to the editing operations of the target object, the above-mentioned execution entity can edit the full-scale graph data and / or the relationship graph data, meeting the editing requirements of users.

[0181] To further illustrate the generation process of the full-scale graph data and the generation process of the relationship graph data of the present disclosure, the following processing logic is given:

[0182] 1. The acquisition logic of the fusion input parameters is as follows:

[0183] Suppose:

[0184] 1.1. There are a total of n graph data. The 0th graph data is the externally imported graph data specified by the target object, and the 1st to nth graph data are the graph data obtained by parsing n - 1 transcripts;

[0185] 1.2. G(v, e){x} = the xth original data graph;

[0186] 1.3. Gz(v, e){x} = the subgraph of the xth original data graph;

[0187] 1.4. Gr(v, e){Gzi~Gzj} is the graph that fuses Gzi~Gzj;

[0188] 1.5. Gt(v, e) is the full-scale graph data, and Gp(v, e) is the relationship graph data.

[0189] 2. The implementation logic of the data processing process is as follows:

[0190] 2.1. Parse the transcript i through the large language model to obtain the graph data G(v, e){i};

[0191] 2.2. Save the parsing result: Save G(v, e){i} to the original graph database (transcript_entity, transcript_relation, transcript_entity_detail, transcript_relation_detail tables);

[0192] 2.3. Construct the input parameters for the fusion result: From G(v, e)(0~n), find the subgraph data Gz(v, e){x, u,..., y} associated with G(v, e){i};

[0193] 2.4. Obtain the fusion result:

[0194] Fuse {G(v, e){i}, Gz(v, e){x}, Gz(v, e){u},...., Gz(v, e){y}} to obtain the fused subgraph Gr(v, e){G, Gzx,...Gzy};

[0195] 2.5. Save to the full-scale graph data:

[0196] a. Process Gt(v, e), and delete {G(v, e){i}, Gz(v, e){x}, Gz(v, e){u},...., Gz(v, e){y}}.V (entity nodes) in Gt(v, e), and the edges between {G(v, e){i}, Gz(v, e){x}, Gz(v, e){u},...., Gz(v, e){y}}.V;

[0197] b. Save Gr(v, e){G, Gzx,...Gzy} to Gt(v, e);

[0198] 2.6. Save to the relational graph data:

[0199] a. Obtain the target relational subgraph data Grp(v, e) according to the fused subgraph Gr(v, e){G, Gzx,...Gzy};

[0200] b. Process Gp(v, e), and delete {G(v, e){i}, Gz(v, e){x}, Gz(v, e){u},...., Gz(v, e){y}}.Vp (entity nodes) in Gp(v, e), and the edges between {G(v, e){i}, Gz(v, e){x}, Gz(v, e){u},...., Gz(v, e){y}}.Vp in Gp

[0201] c. Save Grp(v, e) to Gp(v, e).

[0202] Continue to refer to Figure 8 , which shows a schematic flow 800 of another embodiment of the data processing method based on a large language model according to the present disclosure. In the flow 800, the following steps are included:

[0203] Step 801, parse the data to be parsed in the data set through a large language model, determine the graph data representing the relationships between entities in the data to be parsed, and obtain an initial graph data set.

[0204] Step 802, combine the graph data in the initial graph data set and the graph data specified by the target object to obtain a graph data set.

[0205] Step 803, determine the graph data to be fused from the graph data set.

[0206] Step 804, determine the subgraph data associated with the graph data to be fused from the graph data that has completed the fusion operation up to the current time.

[0207] Step 805: According to the priorities and data update times of the data sources corresponding to the sub-graph data and the graph data to be fused respectively, fuse the sub-graph data and the graph data to be fused to obtain fused sub-graph data.

[0208] Step 806: Delete the entities included in the sub-graph data and the graph data to be fused respectively, as well as the relationship data between the entities, from the current fused graph data to obtain graph data to be supplemented.

[0209] Among them, the fused graph data represents the relationships between the entities in the graph data for which the fusion operation has been completed up to the current time.

[0210] Step 807: Fuse the graph data to be supplemented and the fused sub-graph data to obtain updated fused graph data.

[0211] Step 808: Determine the relationship data between the target type entities in the fused sub-graph data to generate target relationship sub-graph data.

[0212] Step 809: Delete the entities included in the sub-graph data and the graph data to be fused respectively, as well as the relationship data between the entities, from the fused relationship graph data to obtain relationship graph data to be supplemented.

[0213] Step 810: Fuse the relationship graph data to be supplemented and the target relationship sub-graph data to obtain updated fused relationship graph data.

[0214] Step 811: Until the fusion operation on the graph data in the graph data set is completed, determine the obtained fused graph data as the full-scale graph data, and determine the obtained fused relationship graph data as the relationship graph data.

[0215] Step 812: According to the relationship query request of the target object, determine the hierarchical structure relationship data including the target query entity from the relationship graph data.

[0216] The process 800 of the data processing method based on the large language model in this embodiment specifically illustrates the generation process of the full-scale graph data and the generation process of the relationship graph data, improving the construction efficiency and accuracy of the full-scale graph data and the relationship graph data.

[0217] Continue to refer to Figure 9 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a data processing device based on a large language model. This system embodiment corresponds to Figure 2 the method embodiment shown, and this system can be specifically applied to various electronic devices.

[0218] As shown in Figure 9As shown in the figure, the data processing device 900 based on the large language model includes: a parsing unit 901, configured to parse the data to be parsed in the data set through the large language model, determine the graph data representing the relationships between entities in the data to be parsed, and obtain a graph data set; a fusion unit 902, configured to fuse the graph data in the graph data set to obtain full-scale graph data; a generation unit 903, configured to determine the relationships between target type entities in the full-scale graph data and generate relationship graph data.

[0219] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: determine the graph data to be fused from the graph data set; determine the sub-graph data having an association relationship with the graph data to be fused from the graph data that has completed the fusion operation up to the current time; fuse the sub-graph data and the graph data to be fused to obtain fused sub-graph data; update the current fused graph data according to the fused sub-graph data until the fusion operation on the graph data in the graph data set is completed, and determine the obtained fused graph data as the full-scale graph data, where the fused graph data represents the relationships between entities in the graph data that has completed the fusion operation up to the current time.

[0220] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: fuse the sub-graph data and the graph data to be fused according to the priorities and data update times of the data sources corresponding to the sub-graph data and the graph data to be fused, respectively, to obtain fused sub-graph data.

[0221] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: for a target entity with consistent key attribute data in the sub-graph data and the graph data to be fused, in response to the sub-graph data and the graph data to be fused belonging to different data sources, and both the sub-graph data and the graph data to be fused include the non-key attribute data of the target entity under the same non-key attribute, determine the fusion method of the non-key attribute data according to the type of the non-key attribute; in response to the fusion method being overwrite, determine the target non-key attribute data from the non-key attribute data included in the sub-graph data and the graph data to be fused respectively according to the priority and data update time, and use it as the attribute data of the target entity in the non-key attribute in the fused sub-graph data.

[0222] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: in response to the fusion method being append, combine the non-key attribute data included in the sub-graph data and the graph data to be fused respectively to obtain the attribute data of the target entity in the non-key attribute in the fused sub-graph data.

[0223] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: for a target entity with consistent key attribute data in the sub-graph data and the graph data to be fused, in response to the sub-graph data and the graph data to be fused belonging to the same data source, determine target non-key attribute data from the non-key attribute data included in the sub-graph data and the graph data to be fused respectively according to the respective data update times of the sub-graph data and the graph data to be fused, and use the target non-key attribute data as the attribute data of the target entity in the fusion sub-graph data under the non-key attributes.

[0224] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: for the relationship between the head entity and the tail entity, in response to the first relationship data between the head entity and the tail entity in the sub-graph data being consistent with the second relationship data between the head entity and the tail entity in the graph data to be fused, merge the first relationship data and the second relationship data, and use the merged data as the relationship data between the head entity and the tail entity in the fusion sub-graph data.

[0225] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: in response to the target relationship data between the head entity and the tail entity being included in the sub-graph data or the graph data to be fused, use the target relationship data as the relationship data between the head entity and the tail entity in the fusion sub-graph data.

[0226] In some alternative implementation manners of this embodiment, the fusion unit 902 is further configured to: delete the entities included in the sub-graph data and the graph data to be fused respectively, and the relationship data between the entities, from the current fused graph data to obtain the graph data to be supplemented; fuse the graph data to be supplemented and the fusion sub-graph data to obtain the updated fused graph data.

[0227] In some alternative implementation manners of this embodiment, the generation unit 903 is further configured to: determine the relationship data between the target type entities in the fusion sub-graph data, and generate target relationship sub-graph data; update the current fused relationship graph data according to the fusion sub-graph data until the fusion operation on the graph data in the graph data set is completed, and determine the obtained fused relationship graph data as the relationship graph data, where the fused relationship graph data represents the relationship between the target type entities in the graph data for which the fusion operation has been completed up to the current time.

[0228] In some alternative implementation manners of this embodiment, the generation unit 903 is further configured to: delete the entities included in the sub-graph data and the graph data to be fused respectively, and the relationship data between the entities, from the fused relationship graph data to obtain the relationship graph data to be supplemented; fuse the relationship graph data to be supplemented and the target relationship sub-graph data to obtain the updated fused relationship graph data.

[0229] In some alternative implementation manners of this embodiment, the generating unit 903 is further configured to: determine a connection path between target type entities in the fused sub-graph data; screen a target connection path from the connection paths according to a preset connection path rule; and generate target relationship sub-graph data according to the relationship data between the target type entities represented by the target connection path.

[0230] In some alternative implementation manners of this embodiment, the parsing unit 901 is further configured to: parse the data to be parsed in the data set through a large language model, determine graph data representing the relationships between the entities in the data to be parsed, and obtain an initial graph data set; and combine the graph data in the initial graph data set and the graph data specified by the target object to obtain a graph data set.

[0231] In some alternative implementation manners of this embodiment, the above device further includes: a query unit (not shown in the figure), configured to determine hierarchical structure relationship data including a target query entity from the relationship graph data according to a relationship query request of the target object.

[0232] In some alternative implementation manners of this embodiment, the above device further includes: a customization unit (not shown in the figure), configured to determine a customized graph analysis algorithm according to a customization request of the target object for the graph analysis algorithm; and analyze the full-scale graph data and / or the relationship graph data according to the customized graph analysis algorithm.

[0233] In some alternative implementation manners of this embodiment, the above device further includes: an editing unit (not shown in the figure), configured to edit the full-scale graph data and / or the relationship graph data according to an editing operation of the target object.

[0234] In this embodiment, a data processing device based on a large language model is provided. By using the large language model to parse the data to be parsed in the data set, graph data representing the relationships between the entities in the data to be parsed is determined, and a graph data set is obtained. Based on the powerful natural language understanding ability and logical analysis ability of the large language model, the construction efficiency and accuracy of various graph data are improved; the highly accurate graph data in the graph data set is fused to obtain full-scale graph data, and the relationships between the target type entities in the full-scale graph data are determined to generate relationship graph data. While ensuring the accuracy of the relationship graph data, the target object can quickly obtain the relationships between the target type entities based on the relationship graph data, improving the information acquisition efficiency of the target object.

[0235] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the data processing method based on the large language model described in any of the above embodiments.

[0236] According to an embodiment of the present disclosure, the present disclosure further provides a readable storage medium, which stores computer instructions for enabling a computer to implement the data processing method based on the large language model described in any of the above embodiments when executed.

[0237] The embodiments of the present disclosure provide a computer program product, which can implement the data processing method based on the large language model described in any of the above embodiments when executed by a processor.

[0238] Figure 10 FIG. 1000 is a schematic block diagram of an example electronic device 1000 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0239] As Figure 10 shown, the device 1000 includes a computing unit 1001, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1002 or the computer program loaded from the storage unit 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0240] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as a keyboard, mouse, etc.; output unit 1007, such as various types of displays, speakers, etc.; storage unit 1008, such as a disk, optical disc, etc.; and communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0241] Computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1001 executes the various methods and processes described above, such as data processing methods based on large language models. For example, in some embodiments, the data processing method based on large language models can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by computing unit 1001, one or more steps of the data processing method based on large language models described above can be executed. Alternatively, in other embodiments, computing unit 1001 can be configured to execute the data processing method based on large language models in any other suitable manner (e.g., by means of firmware).

[0242] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0243] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus based on large language models, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0244] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0245] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0246] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0247] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to address the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services; it can also be a server of a distributed system, or a server combined with blockchain.

[0248] According to the technical solution of the embodiment of the present disclosure, there is provided a data processing method and apparatus based on a large language model. Through the large language model, according to the object background information of the target object, the query request of the target object is parsed to determine the true and complete object requirements of the target object; according to the object requirements, the query request is adjusted to generate an adjusted query request that can completely express the true needs of the target object; data query is performed according to the adjusted query request, and the data query result can be obtained conveniently, improving the acquisition efficiency, accuracy of the data query result, and the matching degree between the data query result and the target object.

[0249] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution provided by the present disclosure can be achieved, and no limitation is made herein.

[0250] The above specific implementation manners do not constitute a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A data processing method based on a large language model, comprising: Parsing the data to be parsed in the data set by using a large language model, determining graph data representing the relationship between entities in the data to be parsed, and obtaining a graph data set; Fusing the atlas data in the atlas data set to obtain full atlas data; Determine the relationship between target type entities in the full amount of graph data and generate relationship graph data.

2. The method according to claim 1, wherein: The fusing of the atlas data in the atlas data set to obtain the full atlas data includes: Determining the atlas data to be fused from the atlas data set; From the atlas data that have completed the fusion operation up to now, determine the sub-graph data that has an associated relationship with the atlas data to be fused; Fusing the sub-graph data and the to-be-fused atlas data to obtain fused sub-graph data; According to the fused sub-graph data, the current fused graph data is updated until the fusion operation of the graph data in the graph data set is completed, and the obtained fused graph data is determined as the full graph data, wherein the fused graph data represents the relationship between entities in the graph data that has completed the fusion operation up to the current time.

3. The method according to claim 2, wherein: The fusing the sub-graph data and the to-be-fused atlas data to obtain fused sub-graph data includes: According to the priorities and data update times of the data sources corresponding to the sub-graph data and the graph data to be fused, the sub-graph data and the graph data to be fused are fused to obtain the fused sub-graph data.

4. The method according to claim 3, wherein: The step of fusing the subgraph data and the graph data to be fused to obtain the fused subgraph data according to the priorities and data update times of the data sources respectively corresponding to the subgraph data and the graph data to be fused includes: For a target entity whose key attribute data is consistent in the subgraph data and the graph data to be fused, in response to the subgraph data and the graph data to be fused belonging to different data sources, and both the subgraph data and the graph data to be fused include non-key attribute data of the target entity under the same non-key attribute, determining a fusion method for the non-key attribute data according to the type of the non-key attribute; In response to the fusion mode being overlay, target non-critical attribute data are determined from the non-critical attribute data respectively included in the sub-graph data and the graph data to be fused according to the priority and the data update time, as the attribute data of the target entity in the fused sub-graph data under the non-critical attribute.

5. The method according to claim 4, wherein: The step of fusing the subgraph data and the graph data to be fused to obtain the fused subgraph data according to the priority and data update time of the data sources respectively corresponding to the subgraph data and the graph data to be fused, further comprising: In response to the fusion mode being appending, the non-key attribute data respectively included in the sub-graph data and the graph data to be fused are combined to obtain the attribute data of the target entity in the fused sub-graph data under the non-key attribute.

6. The method according to claim 4, wherein: The step of fusing the subgraph data and the graph data to be fused to obtain the fused subgraph data according to the priority and data update time of the data sources respectively corresponding to the subgraph data and the graph data to be fused, further comprising: For a target entity whose key attribute data is consistent with that in the sub-graph data and the graph data to be fused, in response to the sub-graph data and the graph data to be fused belonging to the same data source, according to the data update time corresponding to each of the sub-graph data and the graph data to be fused, target non-key attribute data are determined from the non-key attribute data respectively included in the sub-graph data and the graph data to be fused, as the attribute data of the target entity in the fused sub-graph data under the non-key attributes.

7. The method according to claim 3, wherein: The step of fusing the subgraph data and the graph data to be fused to obtain the fused subgraph data according to the priorities and data update times of the data sources respectively corresponding to the subgraph data and the graph data to be fused includes: Regarding the relationship between the head entity and the tail entity, in response to the first relationship data between the head entity and the tail entity in the sub-graph data, which is consistent with the second relationship data between the head entity and the tail entity in the graph data to be fused, the first relationship data and the second relationship data are merged as the relationship data between the head entity and the tail entity in the fused sub-graph data.

8. The method according to claim 7, wherein: The step of fusing the subgraph data and the graph data to be fused to obtain the fused subgraph data according to the priority and data update time of the data sources respectively corresponding to the subgraph data and the graph data to be fused, further comprising: In response to the subgraph data or the to-be-fused graph data including target relationship data between the head entity and the tail entity, the target relationship data is used as the relationship data between the head entity and the tail entity in the fused subgraph data.

9. The method according to claim 2, wherein: The updating of the current fused atlas data according to the fused sub-graph data comprises: Delete the entities respectively included in the sub-graph data and the graph data to be fused, as well as the relationship data between the entities from the current fused graph data, to obtain the graph data to be supplemented; The to-be-supplemented atlas data and the fused sub-atlas data are fused to obtain updated fused atlas data.

10. The method according to claim 2, wherein: Determining the relationship between target type entities in the full amount of graph data to generate relationship graph data includes: Determine the relationship data between the target type entities in the fused subgraph data, and generate target relationship subgraph data; According to the fused sub-graph data, the current fused relationship graph data is updated until the fusion operation of the graph data in the graph data set is completed, and the obtained fused relationship graph data is determined as the relationship graph data, wherein the fused relationship graph data represents the relationship between target type entities in the graph data that has completed the fusion operation up to the current time.

11. The method according to claim 10, wherein: The updating of the current fusion relationship graph data according to the fusion subgraph data includes: Delete the entities respectively included in the subgraph data and the graph data to be fused, as well as the relationship data between the entities, from the fused relationship graph data to obtain the relationship graph data to be supplemented; The to-be-supplemented relationship graph data and the target relationship subgraph data are fused to obtain updated fused relationship graph data.

12. The method according to claim 10, wherein: The determining of the relationship data between the target type entities in the fused subgraph data to generate the target relationship subgraph data includes: Determining connectivity paths between target type entities in the fused subgraph data; According to a preset connection path rule, selecting a target connection path from the connection paths; The target relationship subgraph data is generated according to the relationship data between the target type entities represented by the target connectivity path.

13. The method according to claim 1, wherein: The method of parsing the data to be parsed in the data set by using the large language model, determining the graph data representing the relationship between entities in the data to be parsed, and obtaining the graph data set includes: Parsing the data to be parsed in the data set by using a large language model, determining graph data representing the relationship between entities in the data to be parsed, and obtaining an initial graph data set; The atlas data set is obtained by combining the atlas data in the initial atlas data set and the atlas data specified by the target object.

14. The method according to any one of claims 1 to 13, wherein: Also includes: According to the relationship query request of the target object, the hierarchical structure relationship data including the target query entity is determined from the relationship graph data.

15. A data processing device based on a large language model, comprising: A parsing unit is configured to parse the data to be parsed in the data set through a large language model, determine graph data representing the relationship between entities in the data to be parsed, and obtain a graph data set; A fusion unit is configured to fuse the atlas data in the atlas data set to obtain full atlas data; A generating unit is configured to determine the relationship between target type entities in the full amount of graph data and generate relationship graph data.

16. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.

17. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 14.

18. A computer program product comprising: A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 14.