Structured data semantic representation method and device, equipment and medium
Patent Information
- Application Number
- CN202410687610.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-05-30
AI Technical Summary
[0004]然而,现有的多源结构化数据语义表征方法存在以下两方面问题:第一方面,同一种知识可以有不同的结构化数据表现形式,因此描述同一知识的不同结构化数据在语义表征空间内的距离应该更小,然而,现有方法均未考虑此类情况;以及第二方面,将结构化数据转换为文本的过程中,只能采用文本对数据结构进行简单描述,丧失了对于多跳路径、全局链接等深层结构信息的表征能力
[0020]从上述技术方案可以看出,本公开的实施例提供的一种结构化数据语义表征方法、装置、设备及介质至少具有以下有益效果其中之一:
Smart Images

Figure CN118585952B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of natural language processing, structured data semantic representation and large model technology, and in particular to a method, apparatus, device and medium for structured data semantic representation. Background Technology
[0002] Large models, with their massive parameters and deep network structures, can learn and understand more features and patterns, demonstrating astonishing natural language processing capabilities and are considered an important path towards general artificial intelligence. However, due to the different storage structure of text, large models have insufficient ability to understand multi-source structured data such as knowledge graphs, tables, and data records. With the rapid development of cue learning, researchers have proposed a series of effective algorithms that convert structured data into text format, add necessary cue information, and use large models for vector representation. However, such methods usually require the design of specialized cue templates for specific data formats, and cannot achieve unified semantic modeling for various different structured data.
[0003] To address the heterogeneity issue among multi-source structured data, some researchers propose converting different structured knowledge, user requests, and outputs into text-to-text format, and then using multi-task prefix parameter tuning to represent various data structures in a large model. Other researchers propose utilizing retrieval concepts, first constructing dedicated interfaces for accessing corresponding data types, then iteratively calling these interfaces to complete data collection, linearizing the data into corresponding text prompts, thereby enhancing the large model's ability to represent multi-source structured data.
[0004] However, existing semantic representation methods for multi-source structured data have the following two problems: First, the same knowledge can have different structured data representations, so the distance between different structured data describing the same knowledge in the semantic representation space should be smaller. However, existing methods do not consider this situation. Second, in the process of converting structured data into text, only text can be used to simply describe the data structure, losing the ability to represent deep structural information such as multi-hop paths and global links. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] To address at least one of the aforementioned technical problems in existing multi-source structured data semantic representation methods, embodiments of this disclosure provide a structured data semantic representation method, apparatus, device, and medium. Based on the correspondence between knowledge and multi-source structured data, the disclosed embodiments design independent structural information and text information encoding modules to achieve effective injection of deep structural information and adopt a joint fine-tuning strategy with a large model to complete the semantic representation of multi-source structured data.
[0007] (II) Technical Solution
[0008] In view of the above problems, embodiments of this disclosure provide a method, apparatus, device and medium for semantic representation of structured data.
[0009] According to a first aspect of this disclosure, a method for semantic representation of structured data is provided, the method comprising the following steps: inputting structured data into a target structured semantic representation model; using a graph neural network-based structural feature extractor to encode and extract structural information in the structured data to obtain a structural representation of the structured data; using a text feature extractor to encode and extract textual information in the structured data to obtain a textual representation of the structured data; and fusing the structural representation and the textual representation to obtain a semantic representation of the structured data, wherein the structured data includes at least one of knowledge graphs, tables, and data records.
[0010] In some exemplary embodiments, the structural representation and the textual representation are fused, including by weighting the structural representation and the textual representation according to preset weights.
[0011] In some exemplary embodiments, the method further includes the following steps: obtaining a target structural semantic representation model, wherein obtaining the target structural semantic representation model includes the following steps: obtaining structured data and text descriptions corresponding to the structured data, and constructing a semantic representation dataset; iteratively training the initial structural semantic representation model based on the semantic representation dataset to obtain an intermediate structural semantic representation model; and using a prefix fine-tuning method to perform end-to-end joint optimization of the intermediate structural semantic representation model and the large model to obtain the target structural semantic representation model.
[0012] In some exemplary embodiments, constructing a semantic representation dataset includes the following steps: constructing an original alignment dataset based on structured data and text descriptions; and dividing the original alignment dataset into a training set, a validation set, and a test set according to a preset ratio.
[0013] In some exemplary embodiments, the initial structural semantic representation model is iteratively trained based on the semantic representation dataset to obtain an intermediate structural semantic representation model. This includes iterative training using a first training method, wherein the first training method includes the following steps: training the initial structural semantic representation model with structured data in the training set as input data and text descriptions corresponding to the structured data as output targets; and validating the trained initial model using a validation set and saving the validation results.
[0014] In some exemplary embodiments, the process of iteratively training an initial structural semantic representation model based on a semantic representation dataset to obtain an intermediate structural semantic representation model further includes the following steps: selecting the initial training model corresponding to the verification result with the smallest error as the model to be tested; and testing the model to be tested using a test set, recording the test results, and obtaining the intermediate structural semantic representation model.
[0015] In some exemplary embodiments, a prefix fine-tuning method is used to jointly optimize the intermediate structural semantic representation model and the large model end-to-end, including the following steps: linearly projecting the text data to obtain a structurally enhanced prefix; adding a prefix to the attention layer of the large model to obtain a large model with an added prefix; and using the prefix fine-tuning method to jointly optimize the large model with the added prefix and the intermediate structural semantic representation model to obtain the target structural semantic representation model. The joint optimization process includes updating the parameters of the prefix of the attention layer of the large model and the intermediate structural semantic representation model, wherein the original parameters of the large model are kept frozen.
[0016] A second aspect of this disclosure provides a structured data semantic representation apparatus, comprising the following modules: an input module for inputting structured data into a target structured semantic representation model; a structure feature extraction module for encoding and extracting structural information from the structured data using a graph neural network-based structure feature extractor to obtain a structured representation of the structured data; a text feature extraction module for encoding and extracting text information from the structured data using a text feature extractor to obtain a textual representation of the structured data; and a fusion output module for fusing the structured representation and the textual representation to obtain a semantic representation of the structured data, wherein the structured data includes at least one of knowledge graphs, tables, and data records.
[0017] A third aspect of this disclosure provides an electronic device comprising: one or more processors and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the methods described above.
[0018] A fourth aspect of this disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the methods described above.
[0019] (III) Beneficial Effects
[0020] As can be seen from the above technical solutions, the structured data semantic representation method, apparatus, device, and medium provided by the embodiments of this disclosure have at least one of the following beneficial effects:
[0021] (1) Based on the correspondence between knowledge and multi-source structured data, an independent structural information and text information encoding module was designed to realize the effective injection of deep structural information.
[0022] (2) By linearly projecting the knowledge representation, a structure-enhanced prefix is generated and added to each attention layer of the large model to improve the accuracy of generating the corresponding knowledge description. Attached Figure Description
[0023] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0024] Figure 1 The illustration shows a flowchart of obtaining a target structural semantic representation model according to an embodiment of the present disclosure;
[0025] Figure 2 This illustration schematically depicts a process diagram for constructing a semantic representation dataset according to an embodiment of the present disclosure;
[0026] Figure 3 The illustration shows a schematic diagram of an iterative training process using a first training method according to an embodiment of the present disclosure;
[0027] Figure 4 The illustration shows a flowchart of iterative training of an initial structural semantic representation model based on a semantic representation dataset according to an embodiment of the present disclosure to obtain an intermediate structural semantic representation model.
[0028] Figure 5 A schematic diagram illustrating the structure of a structural semantic representation model according to an embodiment of the present disclosure is shown.
[0029] Figure 6 A schematic diagram of a structural feature extractor according to an embodiment of the present disclosure is shown.
[0030] Figure 7 A schematic diagram of a text feature extractor according to an embodiment of the present disclosure is shown.
[0031] Figure 8The illustration shows a flowchart of the process of using a prefix fine-tuning method to perform end-to-end joint optimization of the intermediate structure semantic representation model and the large model according to an embodiment of the present disclosure to obtain the target structure semantic representation model.
[0032] Figure 9 A schematic diagram illustrating the structure of a prefix fine-tuning method according to an embodiment of the present disclosure is shown.
[0033] Figure 10 The illustration shows a flowchart of a structured data semantic representation method according to an embodiment of the present disclosure;
[0034] Figure 11 This schematic diagram illustrates the structure of a device 700 for obtaining a target structural semantic representation model according to an embodiment of the present disclosure;
[0035] Figure 12 A schematic diagram illustrating the structure of a structured data semantic representation apparatus 800 according to an embodiment of the present disclosure is shown; and
[0036] Figure 13 A block diagram of an electronic device using a structured data semantic representation method according to an embodiment of the present disclosure is shown schematically.
[0037] Illustration:
[0038] 600 - Electronic device; 601 - Processor; 602 - Read-only memory (ROM); 603 - Random access memory (RAM); 604 - Bus; 605 - Input / output (I / O) interface; 606 - Input section; 607 - Output section; 608 - Storage section; 609 - Communication section; 610 - Driver; 611 - Removable media. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with specific embodiments and accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0040] Figure 1 The illustration shows a flowchart of obtaining a target structural semantic representation model according to an embodiment of the present disclosure.
[0041] like Figure 1 As shown, the method for obtaining a target structure semantic representation model according to an embodiment of this disclosure includes steps S110-S130.
[0042] In step S110, structured data and its corresponding text descriptions are obtained to construct a semantic representation dataset. Optionally, the structured data includes knowledge graphs, tables, and data records, with the same text description corresponding to these three external structured data representations: knowledge graphs, tables, and data records.
[0043] In some exemplary embodiments, step S110 includes steps S111-S112, see reference. Figure 2 .
[0044] In step S111, the original aligned dataset is constructed based on structured data and text descriptions.
[0045] For example, knowledge graph KG, table TB, and data record DI are three different types of structured data, all describing the same kind of knowledge. The textual description of this knowledge is represented by KT. Then the original aligned dataset D={KG,TB,DI,KT}.
[0046] In step S112, the original aligned dataset is divided into a training set, a validation set, and a test set according to a preset ratio.
[0047] For example, the original aligned dataset D={KG,TB,DI,KT} is split in an 8:1:1 ratio into a training set (Train), a validation set (Valid), and a test set (Test). The training set is used for model training; the validation set is used to verify the model's performance; and the test set is used to test the structural semantic representation model that achieves the final result on the validation set.
[0048] In step S120, the initial structural semantic representation model is iteratively trained based on the semantic representation dataset to obtain an intermediate structural semantic representation model. Optionally, the structural semantic representation model adopts a modular design that separates structure and content, encoding and fusing structural and textual information in the structured data separately. The structure of the structural semantic representation model is as follows: Figure 5 ,Depend on Figure 5 As can be seen, the structural semantic representation model according to the embodiments of this disclosure comprises three parts: a structural feature extractor (StructEmb), a text feature extractor (TextEmb), and feature fusion. The model input is multi-source structured data, and the output is a high-dimensional vector representation of knowledge. Preferably, the structural semantic representation model uses a matrix to initially represent the structural information of different structured data, and uses a graph neural network-based structural feature extractor for encoding to obtain the corresponding structural representation. Optionally, the structure of the structural feature extractor is as follows: Figure 6 As shown, for different structured data, the structural information S is represented in matrix form and encoded using a graph neural network-based structural feature extractor, StructEmb, to obtain the corresponding structural representation emb. s=StructEmb(S); This function encodes the text content in structured data using a Transformer-based text feature extractor to obtain the corresponding text representation. Optionally, the text feature extractor structure is as follows: Figure 7 As shown, for the textual description information C in structured data, the Transformer-based text feature extractor TextEmb is used for encoding to obtain the corresponding text representation emb. c =TextEmb(C); This function merges structural and textual representations to obtain a high-dimensional representation of knowledge.
[0049] In some exemplary embodiments, the iterative training in step S120 includes step S121 performing iterative training using a first training method. In some exemplary embodiments, step S121 includes steps S1211-S1212. See also... Figure 3 .
[0050] In step S1211, the initial structural semantic representation model is trained using structured data from the training set as input data and text descriptions corresponding to the structured data as output targets.
[0051] In step S1212, the initial trained model is validated using the validation set, and the validation results are saved.
[0052] For example, the model takes a knowledge graph (KG), a table (TB), and data records (DI) as common inputs, and a text description of the knowledge (KT) as the generation target for model training. Iterative training is performed on the training set (Train), and after each round of training, the model is validated on the validation set (Valid) to save the model's performance.
[0053] In some exemplary embodiments, after the iterative training is completed, step S120 further includes steps S122-S123, see reference. Figure 4 .
[0054] In step S122, the initial training model corresponding to the verification result with the smallest error is selected as the model to be tested.
[0055] In step S123, the test set is used to test the model to be tested, the test results are recorded, and the intermediate structure semantic representation model is obtained.
[0056] For example, after iterative training is completed, the model that performs best on the validation set is selected and tested on the test set, and the final test results are recorded.
[0057] In step S130, the prefix fine-tuning method is used to perform end-to-end joint tuning of the intermediate structural semantic representation model and the large model to obtain the target structural semantic representation model. Optionally, during the joint tuning process, the original parameters of the large model are kept frozen, and only the parameters of the prefix and structural semantic representation models are updated.
[0058] In some exemplary embodiments, step S130 includes steps S131-S133, see reference. Figure 8 .
[0059] In step S131, the text data is linearly projected to obtain a structurally enhanced prefix.
[0060] In step S132, a prefix is added to the attention layer of the large model to obtain a large model with an added prefix.
[0061] In step S133, a prefix fine-tuning method is used to jointly optimize the large model with added prefixes and the intermediate structure semantic representation model to obtain the target structure semantic representation model. The joint optimization process includes updating the parameters of the prefix and intermediate structure semantic representation model of the attention layer of the large model, while the original parameters of the large model remain frozen.
[0062] The structure of the prefix fine-tuning method is as follows: Figure 9 As shown, by Figure 9 It can be seen that by representing knowledge as emb K Linear projection is performed to generate a structurally enhanced prefix P. This prefix is added to each attention layer of the large model to generate the accuracy of the corresponding knowledge description. Joint optimization is performed, and the original parameters of the large model are kept frozen during the joint optimization process. Only the parameters of the prefix and the structural semantic representation model are updated.
[0063] Figure 10 The illustration shows a flowchart of a structured data semantic representation method according to an embodiment of the present disclosure.
[0064] like Figure 10 As shown, the structured data semantic representation method according to an embodiment of this disclosure includes steps S210-S240.
[0065] In step S210, structured data is input into the target structural semantic representation model. (See attached diagram for the structure of the structural semantic representation model.) Figure 5 Structured data includes at least one of knowledge graphs, tables, and data records.
[0066] In step S220, a graph neural network-based structural feature extractor is used to encode and extract structural information from the structured data to obtain a structural representation of the structured data.
[0067] In some exemplary embodiments, the structural feature extractor structure is as follows: Figure 6 As shown, for different structured data, the structural information S is represented in matrix form and encoded using a graph neural network-based structural feature extractor, StructEmb, to obtain the corresponding structural representation emb. s =StructEmb(S).
[0068] In step S230, a text feature extractor is used to encode and extract text information from the structured data to obtain a text representation of the structured data.
[0069] For example, a Transformer-based text feature extractor can be used to encode the text content in structured data to obtain the corresponding text representation.
[0070] In some exemplary embodiments, the text feature extractor structure is as follows: Figure 7 As shown, for the textual description information C in structured data, the Transformer-based text feature extractor TextEmb is used for encoding to obtain the corresponding text representation emb. c =TextEmb (C).
[0071] In step S240, the structural representation and textual representation are fused to obtain the semantic representation of the structured data.
[0072] In some exemplary embodiments, the structural representation and the textual representation are fused, including by weighting the structural representation and the textual representation according to preset weights.
[0073] For example, the text emb c and structural characterization of EMB s The fusion is performed using an averaging method to achieve a high-dimensional representation of knowledge, emb. K =0.5×emb c +0.5×emb s .
[0074] Figure 11 The diagram illustrates the structure of a device 700 for obtaining a target structural semantic representation model according to an embodiment of the present disclosure.
[0075] like Figure 11 As shown, the apparatus 700 for obtaining a target structural semantic representation model according to an embodiment of the present disclosure includes an acquisition module 710, a first training module 720, and a second training module 730.
[0076] The acquisition module 710 is used to acquire structured data and the corresponding text descriptions of the structured data, and to construct a semantic representation dataset, wherein the structured data includes at least one of knowledge graphs, tables and data records.
[0077] The first training module 720 is used to iteratively train the initial structural semantic representation model based on the semantic representation dataset to obtain the intermediate structural semantic representation model.
[0078] The second training module 730 is used to perform end-to-end joint tuning of the intermediate structure semantic representation model and the large model using the prefix fine-tuning method to obtain the target structure semantic representation model.
[0079] In some specific embodiments, any multiple modules among the acquisition module 710, the first training module 720, and the second training module 730 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functionality of one or more of these modules can be combined with at least some of the functionality of other modules and implemented in one module.
[0080] In some specific embodiments, at least one of the acquisition module 710, the first training module 720, and the second training module 730 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable method of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three methods. Alternatively, at least one of the acquisition module 710, the first training module 720, and the second training module 730 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0081] Figure 12 A schematic block diagram of a structured data semantic representation apparatus 800 according to an embodiment of the present disclosure is shown.
[0082] like Figure 12 As shown, the structured data semantic representation device 800 according to an embodiment of the present disclosure includes an input module 810, a structural feature extraction module 820, a text feature extraction module 830, and a fusion output module 840.
[0083] The input module 810 is used to input structured data into the target structural semantic representation model, wherein the structured data includes at least one of knowledge graphs, tables, and data records.
[0084] The structural feature extraction module 820 is used to encode and extract structural information from structured data using a graph neural network-based structural feature extractor to obtain a structural representation of the structured data.
[0085] The text feature extraction module 830 is used to encode and extract text information from structured data using a text feature extractor to obtain a text representation of the structured data.
[0086] The fusion output module 840 is used to fuse structural representations and textual representations to obtain semantic representations of structured data.
[0087] In some specific embodiments, any and multiple modules among the input module 810, structural feature extraction module 820, text feature extraction module 830, and fusion output module 840 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module.
[0088] In some specific embodiments, at least one of the input module 810, structural feature extraction module 820, text feature extraction module 830, and fusion output module 840 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable method of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three methods. Alternatively, at least one of the input module 810, structural feature extraction module 820, text feature extraction module 830, and fusion output module 840 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0089] Figure 13 A block diagram of an electronic device using a structured data semantic representation method according to an embodiment of the present disclosure is shown schematically.
[0090] like Figure 13As shown, an electronic device 600 according to an embodiment of this disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this disclosure.
[0091] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 executes various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 602 and / or RAM 603. It should be noted that the program may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also execute various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.
[0092] In some specific embodiments, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 may also include one or more of the following components connected to the I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.
[0093] One embodiment of this disclosure illustrates a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0094] In some specific embodiments, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.
[0095] One embodiment of this disclosure illustrates a computer program product including a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this disclosure.
[0096] When the computer program is executed by the processor 601, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0097] In some specific embodiments, the computer program may rely on tangible storage media such as optical storage devices or magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via communication section 609, and / or installed from removable medium 611. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0098] In some specific embodiments, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0099] In some specific embodiments, program code for executing the computer programs provided in the embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0101] Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0102] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
Claims
1. A method for semantic representation of structured data, characterized in that, The method includes the following steps: Obtain the semantic representation model of the target structure; Input structured data into the target structural semantic representation model; The structural information in the structured data is encoded and extracted using a graph neural network-based structural feature extractor to obtain the structural representation of the structured data. The text information in the structured data is encoded and extracted using a text feature extractor to obtain a text representation of the structured data; and The structural representation and the textual representation are fused to obtain the semantic representation of the structured data. The structured data includes at least one of knowledge graphs, tables, and data records; The process of obtaining the semantic representation model of the target structure includes the following steps: Obtain structured data and the corresponding text descriptions to construct a semantic representation dataset; Based on the aforementioned semantic representation dataset, the initial structural semantic representation model is iteratively trained to obtain an intermediate structural semantic representation model; and The intermediate structure semantic representation model and the large model are jointly tuned end-to-end using a prefix fine-tuning method to obtain the target structure semantic representation model. The method of using prefix fine-tuning to jointly optimize the intermediate structure semantic representation model and the large model end-to-end includes the following steps: The semantic representation of the structured data is linearly projected to obtain a structure-enhanced prefix; Adding the prefix to the attention layer of the large model yields a large model with an added prefix; A prefix fine-tuning method is used to jointly optimize the large model with added prefixes and the intermediate structural semantic representation model to obtain the target structural semantic representation model. The joint tuning process includes updating the prefix of the attention layer of the large model and the parameters of the intermediate structure semantic representation model, while the original parameters of the large model remain frozen.
2. The structured data semantic representation method according to claim 1, characterized in that, The process of fusing the structural representation and the text representation includes weighted summation of the structural representation and the text representation according to preset weights.
3. The structured data semantic representation method according to claim 1 or 2, characterized in that, The construction of the semantic representation dataset includes the following steps: Based on the structured data and the text description, construct the original aligned dataset; and The original aligned dataset is divided into a training set, a validation set, and a test set according to a preset ratio.
4. The structured data semantic representation method according to claim 3, characterized in that, The step of iteratively training the initial structural semantic representation model based on the semantic representation dataset to obtain an intermediate structural semantic representation model includes iterative training using a first training method, wherein the first training method includes the following steps: Using the structured data in the training set as input data and the text description corresponding to the structured data as the output target, the initial structural semantic representation model is trained; and The initial training model is validated using the validation set, and the validation results are saved.
5. The structured data semantic representation method according to claim 4, characterized in that, The step of iteratively training the initial structural semantic representation model based on the semantic representation dataset to obtain an intermediate structural semantic representation model further includes the following steps: The initial training model corresponding to the validation result with the smallest error is selected as the model to be tested; and The test set is used to test the model to be tested, the test results are recorded, and an intermediate structure semantic representation model is obtained.
6. A structured data semantic representation device, characterized in that, The device includes the following modules: The input module is used to obtain the target structural semantic representation model by inputting structured data into the target structural semantic representation model. The structural feature extraction module is used to encode and extract structural information from the structured data using a graph neural network-based structural feature extractor to obtain a structural representation of the structured data. The text feature extraction module is used to encode and extract text information in the structured data using a text feature extractor to obtain a text representation of the structured data. as well as The fusion output module is used to fuse the structural representation and the text representation to obtain the semantic representation of the structured data. The structured data includes at least one of knowledge graphs, tables, and data records; The process of obtaining the semantic representation model of the target structure includes the following steps: Obtain structured data and the corresponding text descriptions to construct a semantic representation dataset; Based on the aforementioned semantic representation dataset, the initial structural semantic representation model is iteratively trained to obtain an intermediate structural semantic representation model; and The intermediate structure semantic representation model and the large model are jointly tuned end-to-end using a prefix fine-tuning method to obtain the target structure semantic representation model. The method of using prefix fine-tuning to jointly optimize the intermediate structure semantic representation model and the large model end-to-end includes the following steps: The semantic representation of the structured data is linearly projected to obtain a structure-enhanced prefix; Adding the prefix to the attention layer of the large model yields a large model with an added prefix; A prefix fine-tuning method is used to jointly optimize the large model with added prefixes and the intermediate structural semantic representation model to obtain the target structural semantic representation model. The joint tuning process includes updating the prefix of the attention layer of the large model and the parameters of the intermediate structure semantic representation model, while the original parameters of the large model remain frozen.
7. An electronic device, wherein, include: One or more processors; as well as Storage device for storing one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Food safety knowledge graph construction and completion method based on graph neural network
CN115563297A
Text processing model training method and semantic information determining method and device
CN116090472A