Atlas acquisition method and object group extraction model training method
Patent Information
- Application Number
- CN202410324497.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-20
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-20
AI Technical Summary
[0003]相关技术中,可以基于预设的提示词范围从风险信息的文本内容中,对存在关联关系的抽象信息进行抽取,然而,由于抽象信息的抽取对于提示词范围存在一定程度的依赖性,因此,当风险信息的文本内容中存在预设的提示词范围之外的其他提示词时,该提示词对应的抽象信息存在可能无法被抽取,从而导致风险信息中的抽象信息抽取不完整
[0012]应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
Smart Images

Figure CN118227753B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, and in particular to artificial intelligence fields such as natural language processing and computer vision. Background Technology
[0002] With the development of society, people can extract key information from past risk information, thereby obtaining abstract information from that risk information and the relationships between abstract information.
[0003] In related technologies, abstract information with correlations can be extracted from the text content of risk information based on a preset range of prompt words. However, since the extraction of abstract information is dependent on the range of prompt words to a certain extent, when there are other prompt words outside the preset range of prompt words in the text content of risk information, the abstract information corresponding to the prompt words may not be extracted, resulting in incomplete extraction of abstract information in risk information. Summary of the Invention
[0004] This disclosure presents a method, apparatus, device, and medium for acquiring spectra.
[0005] According to a first aspect of this disclosure, a graph acquisition method is proposed, the method comprising: acquiring a preset reference relationship graph; acquiring update object text and extracting a first update object group from the update object text; identifying whether the first update object group satisfies the graph update conditions of the reference relationship graph; and, in response to identifying that the first update object group satisfies the graph update conditions, updating the reference relationship graph based on the first update object group to obtain an updated target relationship graph.
[0006] According to a second aspect of this disclosure, a training method for an object group extraction model is proposed. The method includes: acquiring a candidate object group extraction model to be trained and sample object text; acquiring a preset reference object set to extract sample object groups from the sample object text; and training the candidate object group extraction model based on the sample object text and the sample object groups until training is completed, thereby obtaining the trained target object group extraction model, wherein the target object group extraction model is used to implement the graph acquisition method proposed in the first aspect above.
[0007] According to a third aspect of this disclosure, a graph acquisition device is proposed, comprising: a first acquisition module for acquiring a preset reference relationship graph; a second acquisition module for acquiring update object text and extracting a first update object group from the update object text; an identification module for identifying whether the first update object group satisfies the graph update conditions of the reference relationship graph; and an update module for updating the reference relationship graph based on the first update object group in response to the identification that the first update object group satisfies the graph update conditions, thereby obtaining an updated target relationship graph.
[0008] According to the fourth aspect of this disclosure, a training apparatus for an object group extraction model is proposed. The apparatus includes: a third acquisition module for acquiring a candidate object group extraction model to be trained and sample object text; an extraction module for acquiring a preset set of reference objects to extract sample object groups from the sample object text; and a training module for training the candidate object group extraction model based on the sample object text and the sample object groups until training is completed, thereby obtaining the trained target object group extraction model. The target object group extraction model is used to implement the training apparatus for the object group extraction model proposed in the second aspect above.
[0009] According to a fifth aspect of this disclosure, an electronic device is proposed, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the map acquisition method proposed in the first aspect and / or the training method for the object group extraction model proposed in the second aspect.
[0010] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is proposed, wherein the computer instructions are used to cause the computer to execute the map acquisition method proposed in the first aspect and / or the training method of the object group extraction model proposed in the second aspect.
[0011] According to the seventh aspect of this disclosure, a computer program product is proposed, comprising a computer program that, when executed by a processor, implements the map acquisition method proposed in the first aspect and / or the training method for the object group extraction model proposed in the second aspect.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1 This is a schematic flowchart of a spectrum acquisition method according to an embodiment of the present disclosure;
[0015] Figure 2 This is a schematic flowchart of a spectrum acquisition method according to another embodiment of the present disclosure;
[0016] Figure 3 This is a schematic flowchart of a spectrum acquisition method according to another embodiment of the present disclosure;
[0017] Figure 4 This is a schematic flowchart of a spectrum acquisition method according to another embodiment of the present disclosure;
[0018] Figure 5 This is a flowchart illustrating a training method for an object group extraction model according to an embodiment of the present disclosure.
[0019] Figure 6 This is a schematic diagram of a candidate object group extraction model according to an embodiment of the present disclosure;
[0020] Figure 7 This is a flowchart illustrating a training method for an object group extraction model according to another embodiment of the present disclosure.
[0021] Figure 8 This is a schematic diagram of the structure of a spectrum acquisition device according to an embodiment of the present disclosure;
[0022] Figure 9 This is a schematic diagram of the structure of a training device for an object group extraction model according to an embodiment of the present disclosure;
[0023] Figure 10 This is a schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0025] Data processing is a fundamental aspect of systems engineering and automatic control. Data is a form of expression of facts, concepts, or instructions, which can be processed manually or by automated devices. After data is interpreted and given meaning, it becomes information. Data processing involves the acquisition, storage, retrieval, processing, transformation, and transmission of data. The basic purpose of data processing is to extract and derive valuable and meaningful data from large amounts of potentially chaotic and difficult-to-understand data.
[0026] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Its main applications include machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, and Chinese OCR.
[0027] Computer vision (CV) refers to machine vision that uses cameras and computers to replace human eyes for target recognition, tracking, and measurement, and further performs image processing. It uses various imaging systems instead of visual organs as input sensors, with computers replacing the brain to complete the processing and interpretation. This allows computer-processed images to be more suitable for human observation or transmission to instruments for detection. Computer vision research studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting 'information' from images or multidimensional data. Because perception can be seen as extracting information from sensory signals, computer vision can also be seen as the science of studying how to enable artificial systems to 'perceive' from images or multidimensional data.
[0028] Artificial Intelligence (AI) is a new branch of computer science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. It attempts to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Since its inception, AI has matured in both theory and technology, and its applications have expanded continuously. It is conceivable that future AI-driven technological products will serve as "containers" of human wisdom. AI can simulate the information processes of human consciousness and thought.
[0029] Figure 1This is a schematic flowchart of a spectrum acquisition method according to an embodiment of the present disclosure, as shown below. Figure 1 As shown, the method includes:
[0030] S101, Obtain the preset reference relationship map.
[0031] In this embodiment of the disclosure, people can extract abstract information from open source information, and through the analysis of the extracted abstract information, provide corresponding information support for subsequent operations in their work and life.
[0032] The open-source information can be publicly available news or other types of open-source information; no specific restrictions are imposed here.
[0033] As an example, let's define the open-source information as public information in the financial industry, and define the extraction of a first abstract information and a second abstract information from this public information, where the occurrence of the first abstract information leads to the occurrence of the second abstract information. In this example, the first abstract information can be marked as the cause object, and the second abstract information can be marked as the effect object of the cause object.
[0034] In this example, the first abstract information could be the restructuring of the major shareholder, and the second abstract information could be the change of the company's actual controller. Therefore, due to the existence of these cause and effect objects, it can be determined that a potential opportunity for merger and reorganization has emerged. In this scenario, the analysis results can provide certain information support for subsequent related operations.
[0035] In this embodiment of the disclosure, object groups consisting of cause objects and corresponding effect objects can be extracted from the disclosed object text. The object groups extracted from the disclosed object text may have a certain degree of correlation.
[0036] As an example, when the relationship is causal, multiple groups of objects may have one cause leading to multiple effects or multiple causes leading to one effect.
[0037] Based on this example, the target object group extraction model is set to extract object group 1 and object group 2 from the publicly available object text. Object group 1 includes cause object 1 and effect object 1, and object group 2 includes cause object 2 and effect object 2. In this case, effect object 1 is set to be the cause object of cause object 2.
[0038] In this example, the occurrence of cause object 1 leads to the occurrence of effect object 1, the occurrence of effect object 1 leads to the occurrence of cause object 2, and the occurrence of cause object 2 leads to the occurrence of effect object 2. In this scenario, object group 1 and object group 2 present a situation of multiple causes and one effect, that is, cause object 1 → effect object 1 → cause object 2 → effect object 2.
[0039] In this embodiment of the disclosure, a corresponding association graph can be constructed based on the association relationships between multiple object groups to achieve an intuitive presentation of the association relationships between multiple object groups.
[0040] Optionally, for any industry sector, the events that occur regularly in that sector can be marked as common abstract information, and the relationships between these common abstract information can be analyzed. The abstract information with relationships can be combined to obtain a group of common objects in that industry sector.
[0041] In this scenario, a graph construction algorithm based on relevant technologies can be used to process multiple common object groups within the industry domain based on the association relationships between multiple common object groups, thereby obtaining an association relationship graph between multiple common object groups, and marking this graph as a reference relationship graph within the industry domain.
[0042] It should be noted that, based on the reference relationship graph, information about the relationships between multiple events can be obtained intuitively. This can be understood as follows: in scenarios where the relationship is causal, for scenarios with one cause and multiple effects or multiple causes and one effect, the relationship between multiple abstract pieces of information can be intuitively obtained from the reference relationship graph.
[0043] S102, obtain the update object text, and extract the first update object group from the update object text.
[0044] In this embodiment of the disclosure, a preset object extraction method can be used to extract object groups from new publicly disclosed object text, and the extracted object groups can be merged into the reference relationship graph based on the association between the extracted object groups and the existing reference object groups in the reference relationship graph.
[0045] This allows the newly published object text to be marked as updated object text.
[0046] Optionally, the updated object text can be processed by a semantic recognition algorithm based on related technologies, and then the cause object and the corresponding effect object in the updated object text can be extracted according to the result of the algorithm processing, so as to obtain the object group in the updated object text.
[0047] Among them, the object group extracted from the updated object text can be marked as the first updated object group.
[0048] S103, Identify whether the first updated object group meets the map update conditions of the reference relationship map.
[0049] In this embodiment of the disclosure, the reference relationship graph has preset graph update conditions. For any object group, if the object group meets the graph update conditions, the object group can be updated to the reference relationship graph.
[0050] Optionally, the map update condition can be whether there is an association between the object group and the existing reference objects in the reference relationship map. In this scenario, when it is identified that there is an association between the first updated object group and the existing reference objects in the reference relationship map, it can be determined that the first updated object group meets the map update condition of the reference relationship map.
[0051] Optionally, the map update condition can also be a judgment condition corresponding to the association relationship between existing reference objects in the reference relationship map. When the association relationship between the first update object group is found to meet the judgment condition, it can be determined that the first update object group meets the map update condition.
[0052] S104, in response to the recognition that the first update object group meets the map update conditions, the reference relationship map is updated based on the first update object group to obtain the updated target relationship map.
[0053] In this embodiment of the disclosure, when it is identified that the first update object group meets the map update conditions of the reference relationship map, the first update object group can be merged into the reference relationship map to update the reference relationship map, and the updated map can be marked as the target relationship map.
[0054] Optionally, the first updated object group can be processed by a graph generation algorithm in related technologies, and then the association graph of the first updated object group can be obtained according to the result of the algorithm processing. The association graph can be merged into a reference relationship graph to update the reference relationship graph and obtain the updated target relationship graph.
[0055] Optionally, reference objects that have an association with the first updated object group in the reference relationship graph can be obtained, and the graph node corresponding to the first updated object group can be updated on the graph node to which the reference object belongs, thereby merging the first updated object group into the reference relationship graph to obtain the updated target relationship graph.
[0056] The graph acquisition method proposed in this disclosure acquires a preset reference relationship graph and an update object text, extracts a first update object group from the update object text, identifies whether the first update object group meets its corresponding graph update conditions, and when the first update object group is identified as meeting the graph update conditions, updates the reference relationship graph based on the first update object group to obtain the updated target relationship graph. In this disclosure, a first updated object group is extracted from the updated object text. When the first updated object group meets the graph update conditions of the reference relationship graph, the reference relationship graph is updated based on the first updated object group to obtain the updated target relationship graph. This improves the timeliness of updating the reference relationship graph and reduces the possibility that object groups in the disclosed object text cannot be completely extracted due to the limitation of the scope of reference objects corresponding to the reference relationship graph. By displaying the association relationship through the target relationship graph, the intuitiveness of the association relationship display is improved. In scenarios where the association relationship is a one-cause-multiple-effect and / or multiple-cause-one-effect relationship, the complexity of obtaining the association relationship is reduced, and the efficiency and accuracy of obtaining the association relationship are improved. In scenarios where the extracted object group is analyzed based on the association relationship, the efficiency and accuracy of the object group analysis are improved, the complexity of the pre-analysis preparation work is reduced, the efficiency of information acquisition is improved, and thus the accuracy and timeliness of the execution of related downstream tasks are improved.
[0057] In the above embodiments, the acquisition of the target relationship graph can be combined with... Figure 2 To understand further, Figure 2 Another embodiment of the spectrum acquisition method of this disclosure, such as Figure 2 As shown, the method includes:
[0058] S201, Obtain the preset reference relationship map.
[0059] Optionally, a preset set of reference objects and a reference object description for each reference object can be obtained.
[0060] In this embodiment of the disclosure, for any industry field, common events in that field can be marked as reference objects, thereby obtaining a reference object set composed of multiple reference objects.
[0061] Optionally, each reference object in the reference object set has descriptive information, and the descriptive information of the reference object can be marked as the reference object description of the reference object.
[0062] Optionally, the reference relationships between each reference object can be obtained.
[0063] In this embodiment of the disclosure, there may be a relationship between the reference objects in the reference object set. For example, regarding reference object 1 and reference object 2, the occurrence of reference object 1 will lead to the occurrence of reference object 2. In this example, it can be determined that there is a relationship between reference object 1 and reference object 2. The relationship can be a causal relationship. In this scenario, reference object 1 is the cause of reference object 2, and reference object 2 is the effect of reference object 1.
[0064] In this scenario, the relationships between reference objects can be marked as reference relationships.
[0065] Optionally, the reference object descriptions of each reference object and the reference relationships between each reference object are input into the graph database Bgraph. Through the graph construction capabilities of the graph database BGraph, a reference relationship graph of the reference object set is obtained.
[0066] In this embodiment of the disclosure, a reference relationship map of the reference object set, the reference object description of each reference object, and the reference association relationship between each reference object can be processed by a map construction algorithm in the related art, and then a reference relationship map of the reference object set can be obtained based on the result of the algorithm processing.
[0067] Optionally, the graph construction capabilities of the graph database (Bgraph) can be used to construct a reference relationship graph of the reference object set. In this process, the reference object descriptions of each reference object and the reference associations between each reference object are input into the graph database Bgraph. The graph database Bgraph performs graph construction processing on the reference object descriptions of each reference object and the reference associations between each reference object, thereby obtaining the reference relationship graph of the reference object set.
[0068] As an example, such as Figure 3 As shown, the set of reference objects includes Figure 3 The given objects are 1, 2, 3, 4, 5, 6, and 7. The occurrence of object 1 leads to the occurrence of objects 2 and 3; the occurrence of objects 3 and 4 leads to the occurrence of object 5; and the occurrence of object 5 leads to the occurrence of objects 6 and 7. Therefore, the... Figure 3 It can be seen that there is a causal relationship between object 1 and object 2, between object 1 and object 3, between object 3 and object 5, between object 4 and object 5, between object 5 and object 6, and between object 5 and object 7.
[0069] In this example, the relationships between objects 1, 2, 3, 4, 5, 6, and 7, as well as the object descriptions of each of these objects, can be input into the graph database Bgraph. Using Bgraph's graph construction capabilities, a graph can be constructed... Figure 3 The reference relationship diagram is shown.
[0070] It should be noted that, as an example, in a scenario where the set of reference objects is the set of reference objects corresponding to the financial industry, Figure 3 The object shown can be: 1, which can be related to policy regulation information; 2, which can be related to abnormal price fluctuations; 3, which can be related to increased orders; 4, which can be related to increased sales; 5, which can be related to year-on-year increases in operating revenue; 6, which can be related to year-on-year increases in profit; and 7, which can be related to increased performance. It can also be other objects with... Figure 3 The information objects showing the relationships are not specifically limited here.
[0071] S202, obtain the update object text, and extract the first update object group from the update object text.
[0072] Optionally, the update string sequence of the updated object text is extracted, and the update feature vector of each update string in the update string sequence is obtained.
[0073] In this embodiment of the disclosure, the updated object text can be information based on text. In this scenario, the updated object text can be processed by a text semantic string extraction algorithm in the related technology to extract the strings in the updated object text and mark them as updated strings, thereby obtaining an updated string sequence composed of multiple updated strings in the updated object text.
[0074] Furthermore, a string feature vector extraction algorithm from the relevant technology is obtained, and the feature vector extraction algorithm is applied to each updated string based on the algorithm to obtain the feature vector of each updated string, and then the feature vector is marked as the updated feature vector of each updated string.
[0075] Optionally, update factor objects can be extracted based on the update feature vectors of each update string.
[0076] In this embodiment of the disclosure, the cause object in the updated object text can be extracted from each updated string based on the update feature vector of each updated string, and marked as the update cause object in the updated object text.
[0077] Specifically, based on the update feature vector of each update string, the start string and end string of the update cause object are identified, and the update cause object string in the update string sequence is obtained based on the start string and end string of the update cause object to obtain the update cause object.
[0078] In this embodiment of the disclosure, the extraction position of the cause object to be extracted in the update object text can be determined by determining the start character and end character of the cause object in the update object text, and then the string of the cause object composed of the string at the extraction position can be obtained.
[0079] This can be understood as follows: for the cause object that needs to be extracted from the updated object text, the starting string and ending string of the cause object in the updated object text can be identified based on the update feature vector of each updated string, and marked as the starting string and ending string of the updated cause object. The string consisting of the starting string, the ending string, and the string in between is determined as the updated cause object string in the updated string sequence.
[0080] Furthermore, the event consisting of the characters corresponding to the update cause object string is identified as the update cause object in the update object text.
[0081] Optionally, update result objects can be extracted based on the update feature vectors of each update string.
[0082] In this embodiment of the disclosure, based on the update feature vector of each update string, the result object that has a relationship with the extracted cause object in the update object text can be extracted from each update string, and marked as the update result object of the update cause object in the update object text.
[0083] Specifically, for the update result object corresponding to the update cause object that needs to be extracted from the update object text, the start string and end string of the update result object can be identified based on the update feature vector of each update string. Then, based on the start string and end string of the update result object, the update result object string in the update string sequence can be obtained to obtain the update result object.
[0084] In this embodiment of the disclosure, the extraction position of the fruit object to be extracted in the updated object text can be determined by determining the starting character and the ending character of the updated fruit object in the updated object text, thereby obtaining the updated fruit object string composed of the strings at the extraction position.
[0085] This can be understood as follows: for the result objects that need to be extracted from the updated object text, the starting string and ending string of the result object in the updated object text can be identified based on the update feature vector of each updated string, and marked as the starting string and ending string of the updated result object. The string consisting of the starting string, the ending string, and the string in between is determined as the updated result object string in the updated string sequence.
[0086] Furthermore, the event consisting of the characters corresponding to the updated result object string is identified as the updated result object of the updated cause object in the updated object text.
[0087] Optionally, a first update object group is obtained based on the update cause object and the update effect object.
[0088] In this embodiment of the disclosure, the update cause object and the corresponding update result object extracted from the update object text can be combined, and the combined object is marked as the first update object group in the update object text.
[0089] As an example, if the update cause object extracted from the update object text is set to the order increase information object, and the update result object of the update cause object in the update object text is the year-on-year increase in intention achievement information object, then the order increase information object and the year-on-year increase in intention achievement information object can be combined to obtain the first update object group consisting of "order increase information object → year-on-year increase in intention achievement information object".
[0090] S203, Identify whether the first updated object group meets the map update conditions of the reference relationship map.
[0091] Optionally, the reference relationships between each reference object can be obtained from the reference relationship map.
[0092] In this embodiment of the disclosure, for any reference object, reference objects that have an association with the reference object can be obtained from the reference association relationships between the reference objects displayed in the reference relationship map.
[0093] As an example, such as Figure 3 As shown, through Figure 3 The displayed graph allows for intuitive understanding. Figure 3 The relationships between objects 1, 2, 3, 4, 5, 6, and 7 are shown in [the diagram / illustration]. Figure 3 In the scenario where the illustrated map is a reference relationship map, through Figure 3 This allows you to intuitively obtain the reference relationships between object 1, object 2, object 3, object 4, object 5, object 6, and object 7.
[0094] Optionally, obtain the update associations of the first update object group.
[0095] In this embodiment of the disclosure, there is an association between the cause object and the effect object of the update in the first update object group, and the association between the two can be marked as the update association relationship of the first update object group.
[0096] Optionally, for any reference association, the similarity between the reference association and the updated association can be obtained.
[0097] In this embodiment of the disclosure, the update association information of the first update object group can be compared with the reference association relationship between the existing reference objects in the reference relationship graph. The result of the similarity comparison can be used to identify whether the first update object group meets the graph update conditions of the reference relationship graph.
[0098] Among these features, the existing relationships between reference objects in the reference relationship graph can be marked as reference relationships.
[0099] In this scenario, for any reference association, the similarity between the updated association and the reference association can be calculated based on the similarity algorithm in the relevant technology. Then, the similarity between the updated association and the reference association can be obtained based on the result of the algorithm processing, and it can be marked as the relationship similarity.
[0100] Optionally, based on the similarity of relationships, it can be determined whether the first updated object group meets the map update conditions.
[0101] Specifically, in response to the fact that the similarity between any reference association and the updated association is greater than or equal to a preset similarity threshold, it is determined that the first updated object group meets the map update conditions.
[0102] In this embodiment of the disclosure, the similarity between each of the reference relationships in the reference relationship graph and the updated relationship can be obtained. When the similarity between any reference relationship and the updated relationship is greater than or equal to a preset similarity threshold, it can be determined that the first updated object group and the existing reference objects in the reference relationship graph may be of the same type of event.
[0103] In this scenario, it can be determined that the first updated object group satisfies the map update conditions of the reference relationship map.
[0104] Accordingly, in response to the fact that the similarity between each reference association and the updated association in the reference relationship graph is less than the relationship similarity threshold, it is indeed identified that the first updated object group does not meet the graph update conditions.
[0105] In this embodiment of the disclosure, the similarity between each of the reference relationships in the reference relationship graph and the updated relationship can be obtained. When the obtained similarity of each relationship is less than the preset relationship similarity threshold, it can be determined that the first updated object group and the existing reference objects in the reference relationship graph may be different types of events.
[0106] In this scenario, it can be determined that the first updated object group does not meet the map update conditions of the reference relationship map.
[0107] It should be noted that, in this embodiment of the disclosure, for the first update object group that does not meet the map update conditions, further processing can be performed based on a preset method to add it to the existing reference object range in the reference relationship map. This can be understood in conjunction with the following content:
[0108] Optionally, in response to the identification that the first updated object group does not meet the map update conditions, the first updated object group is added to the second updated object group set, and the second updated object group set is clustered to obtain the object clusters of the second updated object group set.
[0109] In this embodiment of the disclosure, when it is identified that the first update object group does not meet the map update conditions, it can be marked as the second update object group and added to the second update object group set.
[0110] This can be understood as the second update object group being the update object group that does not meet the map update conditions. The second update object group set includes update object groups that do not meet the map update conditions of the reference relationship map within a set time range.
[0111] In this scenario, based on a preset time interval, clustering algorithms in related technologies can be used to cluster each second updated object group, and the clusters obtained after clustering can be marked as object clusters of the second updated object group set.
[0112] Optionally, an object cluster tag can be obtained, wherein, in response to the object cluster tag satisfying the preset addition conditions of the reference object set, the object cluster tag is added as a new reference object to the reference object set.
[0113] In this embodiment of the disclosure, tag information can be obtained for the object clusters of the second updated object group set, and the obtained tag information can be determined as the object cluster tag of the object cluster.
[0114] As an example, the second updated object group can be clustered based on a time interval of ten natural days to obtain five object clusters. Furthermore, the label information of the five object clusters can be determined by a preset label acquisition algorithm, thereby obtaining the object cluster labels of the five object clusters.
[0115] Optionally, after obtaining the object cluster label of the object cluster, the event description information in the object cluster label can be compared with the preset addition conditions corresponding to the existing reference object set in the reference relationship graph. Then, based on the comparison result, it can be identified that the object cluster label meets the preset addition conditions, thereby determining whether the object cluster label can be added to the reference object set as a new reference object.
[0116] The preset conditions for adding the reference object set may include whether the object information in the object cluster tag is common information in the corresponding industry field, or the degree of influence of the object information in the object cluster tag in the corresponding industry field. No specific limitations are made here.
[0117] Optionally, the reference relationship graph can be updated with nodes based on the new reference object to obtain a new reference relationship graph.
[0118] In this embodiment of the disclosure, when an object cluster label is found to meet a preset condition, a new reference object can be generated based on the object cluster label, and the new reference object can be added to the existing set of reference objects in the reference relationship graph.
[0119] In this scenario, when the set of reference objects is updated, the corresponding reference relationship graph needs to be updated accordingly. This involves obtaining the reference object description of the new reference object, as well as the association between the new reference object and each reference object in the existing set of reference objects. This information is then input into the graph database Bgraph. Through Bgraph's graph construction capabilities, graph nodes corresponding to the new reference object are constructed based on the original reference relationship graph, thereby completing the node update of the reference relationship graph and obtaining the updated reference relationship graph.
[0120] S204, in response to the recognition that the first update object group meets the map update conditions, the reference relationship map is updated based on the first update object group to obtain the updated target relationship map.
[0121] In this embodiment of the disclosure, when the first update object group is identified as meeting the graph update conditions, the reference relationship graph can be updated based on the first update object group according to the preset graph update strategy, and the updated reference relationship graph can be marked as the target relationship graph.
[0122] Specifically, in response to the identification that the first updated object group meets the map update conditions, the associated object group of the first updated object group in the reference relationship map is obtained.
[0123] In this embodiment of the disclosure, a reference object group whose relationship with the update association of the first update object group is greater than or equal to a similarity threshold can be obtained from each reference association relationship in the reference relationship graph, and the reference object group is determined as the associated object group of the first update object group in the reference relationship graph.
[0124] Optionally, the associated object relationships of the associated object group are obtained, and the first updated object group is merged into the associated object group according to the updated associated relationships and associated object relationships, so as to update the reference relationship graph and obtain the target relationship graph.
[0125] In this embodiment of the disclosure, the association relationship between associated object groups can be marked as an associated object relationship. In this scenario, the relationship between the updated association relationship and the associated object relationship can be sorted out to obtain the association relationship between the objects included in the first updated object group and the associated object group.
[0126] As an example, the first update object group is set to include cause object 1 and effect object 1, and the associated object group includes cause object 2 and effect object 2, wherein effect object 1 is the cause object of cause object 2.
[0127] In this example, based on the update association relationship of the first updated object group and the association object relationship of the associated object group, it can be known that the association relationship between cause object 1, effect object 1, cause object 2 and effect object 2 is "cause object 1 → effect object 1 → cause object 2 → effect object 2".
[0128] Furthermore, based on this association, the objects in the first updated object group are associated and merged with the objects in the associated object group, thereby updating the reference relationship graph and obtaining the updated target relationship graph.
[0129] Based on the above example, a new graph node can be added upstream of cause object 2 in the associated object group as a graph node of effect object 1, and connected to the graph node of cause object 2. Similarly, a new graph node can be added upstream of the graph node of effect object 1 as a graph node of cause object 1, and connected to the graph node of effect object 1. This will enable the merging and updating of "cause object 1 → effect object 1" and "cause object 2 → effect object 2", thus obtaining the updated target relationship graph.
[0130] As another example, such as Figure 4 As shown, it is possible to obtain Figure 4 The text shows publicly available object text for a specific industry sector, and the object groups within the publicly available object text are extracted to obtain... Figure 4 The group of updated objects is shown.
[0131] Furthermore, to obtain Figure 4The similarity between the updated association of the shown updated object group and the existing reference association in the reference relationship map is compared with a preset similarity threshold.
[0132] like Figure 4 As shown, the similarity between each reference association in the reference relationship graph and the updated association of the updated object group is obtained. When any relationship has a similarity greater than or equal to a similarity threshold, it can be determined that... Figure 4 The shown update object group meets the preset map update conditions.
[0133] In this scenario, the associated object groups of the updated object group in the reference relationship graph can be obtained, and the relationship between each object in the updated object group and the associated object group can be constructed. Based on the relationship, the updated object group can be merged into the reference relationship graph to update the reference relationship graph and obtain the updated target relationship graph.
[0134] like Figure 4 As shown, the similarity between each reference relationship in the reference relationship graph and the updated relationship of the updated object group is obtained. When the similarity of each relationship is less than the preset similarity threshold, it can be determined that... Figure 4 The shown group of updated objects does not meet the map update conditions of the reference association map.
[0135] In this scenario, groups of updated objects that do not meet the graph update conditions can be clustered to obtain object clusters, and each object cluster can be labeled to obtain object cluster labels for each object cluster.
[0136] Optionally, it can identify whether the object cluster label meets the preset addition conditions of the reference object set in the reference relationship graph, and when the object cluster label meets the preset addition conditions, the object cluster label can be added as a new reference object to the reference object set.
[0137] The graph acquisition method proposed in this disclosure updates the reference relationship graph based on the first update object group that meets the graph update conditions, and obtains the target relationship graph. This improves the timeliness of updating the reference relationship graph and enhances the intuitiveness of displaying the association relationships between objects. The method also updates the set of reference objects in the reference relationship graph based on the second update object group, which improves the timeliness of updating the limited scope of the reference object set and reduces the possibility that object groups in the publicly disclosed object text cannot be completely extracted due to the limitation of the reference object scope.
[0138] This disclosure also proposes a training method for an object group extraction model, which can be combined with... Figure 5 , Figure 5This is a flowchart illustrating a training method for an object group extraction model according to an embodiment of the present disclosure, as shown below. Figure 5 As shown, the method includes:
[0139] S501, Obtain the candidate object group extraction model to be trained and the sample object text.
[0140] In this embodiment of the disclosure, the object group extraction model that needs to be trained can be marked as a candidate object group extraction model, such as... Figure 6 As shown, it can be Figure 6 The model shown is identified as a candidate object group extraction model.
[0141] In this scenario, a portion of the object text can be obtained from publicly available historical object text and used as training samples for the candidate object group extraction model, which are then labeled as sample object text.
[0142] S502, Obtain a preset set of reference objects to extract sample object groups from the sample object text.
[0143] In this embodiment of the disclosure, a reference object set composed of common information objects in the industry field can be obtained, and objects can be extracted from the sample object text based on each reference object in the reference object set as sample objects.
[0144] Furthermore, the relationships between the extracted sample objects are obtained, the cause and effect objects of the samples that have relationships among all sample objects are identified, and the combination of cause and effect objects is marked as the sample object group in the sample object text.
[0145] S503, based on the sample object text and sample object group, train the candidate object group extraction model until the training is completed, and obtain the trained target object group extraction model.
[0146] In this embodiment of the disclosure, the sample object group can be used as the label information of the sample object text, input into the candidate object group extraction model for model training, and the trained model can be marked as the target object group extraction model.
[0147] The target object group extraction model is used to achieve the above. Figures 1 to 4 The proposed method for obtaining the spectrum in the embodiments.
[0148] Optionally, a training termination condition can be set based on the training rounds. For the current round of model training, if the current training round of the candidate object group extraction model meets the preset training termination condition, the model training of the candidate object group extraction model can be terminated, and the model obtained after the last round of training can be determined as the trained target object group extraction model.
[0149] Optionally, a corresponding model training termination condition can be set based on the output results of the training rounds. For the current round of model training, if the output results of the candidate object group extraction model in this round meet the preset training termination condition, the model training of the candidate object group extraction model can be terminated, and the model obtained after the last round of training can be determined as the trained target object group extraction model.
[0150] The object group extraction model training method proposed in this disclosure involves acquiring sample object text, extracting sample object groups from the sample object text, and training a candidate object group extraction model based on the sample object groups until training is complete, resulting in a trained target object group extraction model. In this disclosure, the model training of the candidate object group extraction model based on sample object text and sample object groups enables the trained target object group extraction model to extract object groups from publicly available object text, improving the efficiency and accuracy of object group extraction from publicly available object text.
[0151] In the above embodiments, the model training for the object group extraction model can be combined with... Figure 7 To understand further, Figure 7 This is a flowchart illustrating a training method for an object group extraction model according to another embodiment of the present disclosure, as shown below. Figure 7 As shown, the method includes:
[0152] S701, Obtain a preset set of reference objects to extract sample object groups from sample object text.
[0153] Optionally, based on the reference object set, a sample object set is extracted from the sample object text, and based on the reference association relationship between each reference object in the reference object set, the sample object association relationship between each sample object in the sample object set is determined, and based on the sample object association relationship, a sample object group is obtained based on each sample object.
[0154] In this embodiment of the disclosure, information objects with similar descriptions can be extracted from the sample object text based on the reference object descriptions of each reference object in the reference object set as sample objects, thereby obtaining a sample object set composed of multiple extracted sample objects.
[0155] In this scenario, the reference relationships between each reference object in the reference object set can be obtained, and based on these reference relationships, the relationships between each sample object in the sample object set can be determined and marked as sample object relationships.
[0156] Furthermore, based on the relationships between the sample objects, the cause and effect objects that are related in the sample object set can be identified, and the cause and effect objects that are related can be grouped together and marked as a sample object group.
[0157] S702, using the candidate object group extraction model, extracts candidate object groups from the sample object text.
[0158] In this embodiment of the disclosure, sample object groups can be input as label information of sample object text into the candidate object group extraction model. The candidate object group extraction model extracts object groups from the sample object text and marks the extracted object groups as candidate object groups.
[0159] Among them, the candidate word segmentation layer, candidate encoder, candidate cause object recognition layer, and candidate effect object recognition layer of the candidate object group extraction model can be obtained.
[0160] As an example, such as Figure 6 As shown, the candidate object group extraction model can include Figure 6 The shown components are a word segmentation layer, an encoder, a cause object recognition layer, and an effect object recognition layer, which can be labeled as candidate word segmentation layer, candidate encoder, candidate cause object recognition layer, and candidate effect object recognition layer, respectively.
[0161] Optionally, the text of the input model as the sample object text is segmented into words through a candidate word segmentation layer to obtain a sequence of sample string texts. Then, the sample feature vector of each sample string in the sequence of sample string texts is extracted through a candidate encoder.
[0162] As an example, such as Figure 6 As shown, you can input the sample object text. Figure 6 The candidate tokenization layer shown performs word segmentation on the sample object text, and the resulting token strings are used as sample strings of the sample object text, thus obtaining a sequence of sample strings composed of each sample string in the sample object text.
[0163] Furthermore, such as Figure 6 As shown, you can input each sample string. Figure 6 The candidate encoder shown extracts features from each sample string and marks the extracted feature vectors as the sample feature vectors of each sample string.
[0164] The candidate encoder can be built based on the Language Representation Model (BERT) or other types of language models; no specific limitation is made here.
[0165] Optionally, candidate cause objects are extracted based on the sample feature vectors of each sample string through a candidate cause object identification layer.
[0166] In this embodiment of the disclosure, the candidate cause object identification layer can identify the start string and end string of the candidate cause object based on the sample feature vector of each sample string, and obtain the candidate cause object string in the sample string sequence based on the start string and end string of the candidate cause object, so as to obtain the candidate cause object.
[0167] As an example, such as Figure 6 As shown, it can be done through Figure 6 The candidate cause object identification layer shown determines the start and end positions of cause objects in the sample object text, thereby determining the extraction positions of cause objects that need to be extracted from the sample object text.
[0168] Specifically, the starting string of the cause object in the sample object text can be determined from each sample string as the starting string of the candidate cause object, and the ending string of the cause object can be determined from each sample string as the ending string of the candidate cause object, thus obtaining the candidate cause object string.
[0169] As an example, such as Figure 6 As shown, settings Figure 6 It can be shown that... Figure 6 The candidate cause object identification layer shown determines that sample string 11 corresponding to start 1 is the starting string of the candidate cause object, and sample string 14 corresponding to end 1 is the ending string of the candidate cause object, from sample string 11, sample string 12, sample string 13, sample string 14 and sample string 14. Then, sample string 11 can be used as the starting extraction position and sample string 14 can be used as the ending extraction position. Thus, the string composed of sample string 11, sample string 12, sample string 13 and sample string 14 is determined as the candidate cause object string extracted by the candidate cause object identification layer.
[0170] Furthermore, the object composed of candidate cause object strings is identified as the candidate cause object.
[0171] Optionally, candidate fruit objects are extracted based on the sample feature vectors of each sample string through a candidate fruit object recognition layer.
[0172] In this embodiment of the disclosure, the candidate fruit object recognition layer can identify the start string and end string of the candidate fruit object based on the sample feature vector of each sample string, and obtain the candidate fruit object string in the sample string sequence based on the start string and end string of the candidate fruit object, so as to obtain the candidate fruit object.
[0173] As an example, such as Figure 6 As shown, it can be done through Figure 6 The candidate fruit object recognition layer shown determines the start and end positions of the fruit objects in the sample object text, thereby determining the extraction positions of the fruit objects that need to be extracted from the sample object text.
[0174] Specifically, the starting string of the fruit object in the sample object text can be determined from each sample string as the starting string of the candidate fruit object, and the ending string of the fruit object can be determined from each sample string as the ending string of the candidate fruit object, thus obtaining the candidate fruit object string.
[0175] As an example, such as Figure 6 As shown, settings Figure 6 It can be shown that... Figure 6 The candidate fruit object recognition layer shown determines that sample string 21 corresponding to start 2 is the starting string of the candidate fruit object, and sample string 23 corresponding to end 2 is the ending string of the candidate fruit object, from sample string 21, sample string 22, sample string 23, sample string 24 and sample string 24. Then, sample string 21 can be used as the starting extraction position and sample string 23 can be used as the ending extraction position. Thus, the string composed of sample string 21, sample string 22 and sample string 23 is determined as the candidate fruit object string extracted by the candidate fruit object recognition layer.
[0176] Furthermore, the object composed of candidate result object strings is determined as the candidate result object.
[0177] Optionally, the candidate object group output by the candidate object group extraction model can be obtained based on the candidate cause object and the candidate effect object.
[0178] In this embodiment of the disclosure, after the candidate cause object is extracted by the candidate cause object identification layer and the candidate effect object that has a relationship with the candidate cause object is extracted by the candidate effect object identification layer, the extracted candidate cause object and candidate effect object can be combined, and the combined object group is determined as the candidate object group extracted by the candidate object group extraction model.
[0179] S703, based on the sample object group and the candidate object group, obtain the training loss of the candidate object group extraction model.
[0180] Optionally, the sample object group and the candidate object group can be processed by the loss value acquisition algorithm in related technologies, and then the corresponding loss value can be obtained as the training loss of the candidate object group extraction model based on the result of the algorithm processing.
[0181] S704, adjust the model parameters of the candidate object group extraction model according to the training loss, and return to obtain the next sample object text and the next sample object group to continue training the candidate object group extraction model with adjusted parameters until the training ends, and obtain the trained target object group extraction model.
[0182] In this embodiment of the disclosure, the model parameters in the candidate object group extraction model can be adjusted and optimized according to the training loss, thereby obtaining the adjusted candidate object group extraction model.
[0183] In this scenario, the next sample object text and the next sample object group in the next sample object text can be retrieved. The candidate object group extraction model with adjusted parameters can continue to be trained until the training termination condition is met, and the trained target object group extraction model is obtained.
[0184] For details regarding the selection of candidate object groups that meet the model training termination conditions, please refer to the relevant content in the above embodiments, which will not be repeated here.
[0185] The object group extraction model training method proposed in this disclosure determines the extraction position of candidate cause objects through a candidate cause object recognition layer and the extraction position of candidate effect objects through a candidate effect object recognition layer. Then, the model is trained based on the sample object text and the sample object group, so that the trained target object group extraction model has the ability to extract object groups from the published object text. In the scenario of object group extraction based on the target object group extraction model, the extraction efficiency and accuracy of object groups in the published object text are improved.
[0186] Corresponding to the spectrum acquisition methods proposed in the above embodiments, one embodiment of this disclosure also proposes a spectrum acquisition device. Since the spectrum acquisition device proposed in this disclosure corresponds to the spectrum acquisition methods proposed in the above embodiments, the implementation methods of the above spectrum acquisition methods are also applicable to the spectrum acquisition device proposed in this disclosure, and will not be described in detail in the following embodiments.
[0187] Figure 8 This is a schematic diagram of the structure of a spectrum acquisition device according to an embodiment of the present disclosure, as shown below. Figure 8 As shown, the map acquisition device 800 includes a first acquisition module 81, a second acquisition module 82, an identification module 83, and an update module 84, wherein:
[0188] The first acquisition module 81 is used to acquire a preset reference relationship map.
[0189] The second acquisition module 82 is used to acquire the updated object text and extract the first updated object group from the updated object text.
[0190] The identification module 83 is used to identify whether the first updated object group meets the map update conditions of the reference relationship map.
[0191] The update module 84 is used to update the reference relation graph based on the first update object group in response to the recognition that the first update object group meets the graph update conditions, so as to obtain the updated target relation graph.
[0192] In this embodiment of the disclosure, the second acquisition module 82 is further configured to: extract the update string sequence of the updated object text, and obtain the update feature vector of each update string in the update string sequence; extract the update cause object based on the update feature vector of each update string; extract the update effect object based on the update feature vector of each update string; and obtain the first update object group of the updated object text based on the update cause object and the update effect object.
[0193] In this embodiment of the disclosure, the second acquisition module 82 is further configured to: identify the start string and end string of the update cause object based on the update feature vector of each update string; and acquire the update cause object string in the update string sequence based on the start string and end string of the update cause object to obtain the update cause object.
[0194] In this embodiment of the disclosure, the second acquisition module 82 is further configured to: identify the start string and end string of the updated result object based on the update feature vector of each updated string; and acquire the updated result object string in the updated string sequence based on the start string and end string of the updated result object to obtain the updated result object.
[0195] In this embodiment of the disclosure, the identification module 83 is further configured to: obtain the reference association relationships between each reference object from the reference relationship graph; obtain the update association relationships of the first update object group; for any reference association relationship, obtain the relationship similarity between the reference association relationship and the update association relationship; and, based on the relationship similarity, identify whether the first update object group satisfies the graph update conditions.
[0196] In this embodiment of the disclosure, the identification module 83 is further configured to: determine that the first update object group meets the map update conditions in response to the similarity between any reference association and the updated association being greater than or equal to a preset relationship similarity threshold; and confirm that the first update object group does not meet the map update conditions in response to the similarity between each reference association and the updated association in the reference relationship map being less than the relationship similarity threshold.
[0197] In this embodiment of the disclosure, the update module 84 is further configured to: in response to recognizing that the first update object group meets the graph update conditions, obtain the associated object group of the first update object group in the reference relationship graph; obtain the associated object relationships of the associated object group; and, based on the update association relationships and associated object relationships, merge the first update object group into the associated object group to update the reference relationship graph and obtain the target relationship graph.
[0198] In this embodiment of the disclosure, the first acquisition module 81 is further configured to: acquire a preset set of reference objects and a reference object description for each reference object; acquire the reference relationships between each reference object; input the reference object descriptions of each reference object and the reference relationships between each reference object into the graph database BGraph, and obtain a reference relationship graph of the set of reference objects through the graph construction capabilities of the graph database BGraph.
[0199] In this embodiment of the disclosure, the update module 84 is further configured to: in response to identifying that the first update object group does not meet the map update conditions, add the first update object group to the second update object group set, and cluster the second update object group set to obtain object clusters of the second update object group set. Obtain object cluster tags for the object clusters. In response to the object cluster tags meeting the preset addition conditions of the reference object set, add the object cluster tags as new reference objects to the reference object set.
[0200] In this embodiment of the disclosure, the update module 84 is further configured to: update the nodes of the reference relationship graph according to the new reference object to obtain a new reference relationship graph.
[0201] The graph acquisition device disclosed herein acquires a preset reference relationship graph and an update object text, extracts a first update object group from the update object text, identifies whether the first update object group meets its corresponding graph update conditions, and when the first update object group is identified as meeting the graph update conditions, updates the reference relationship graph based on the first update object group to obtain the updated target relationship graph. In this disclosure, a first updated object group is extracted from the updated object text. When the first updated object group meets the graph update conditions of the reference relationship graph, the reference relationship graph is updated based on the first updated object group to obtain the updated target relationship graph. This improves the timeliness of updating the reference relationship graph and reduces the possibility that object groups in the disclosed object text cannot be completely extracted due to the limitation of the scope of reference objects corresponding to the reference relationship graph. By displaying the association relationship through the target relationship graph, the intuitiveness of the association relationship display is improved. In scenarios where the association relationship is a one-cause-multiple-effect and / or multiple-cause-one-effect relationship, the complexity of obtaining the association relationship is reduced, and the efficiency and accuracy of obtaining the association relationship are improved. In scenarios where the extracted object group is analyzed based on the association relationship, the efficiency and accuracy of the object group analysis are improved, the complexity of the pre-analysis preparation work is reduced, the efficiency of information acquisition is improved, and thus the accuracy and timeliness of the execution of related downstream tasks are improved.
[0202] Corresponding to the training methods of the object group extraction model proposed in the above embodiments, an embodiment of this disclosure also proposes a training device for the object group extraction model. Since the training device for the object group extraction model proposed in this disclosure corresponds to the training methods of the object group extraction model proposed in the above embodiments, the implementation methods of the above-mentioned object group extraction model training methods are also applicable to the training device for the object group extraction model proposed in this disclosure, and will not be described in detail in the following embodiments.
[0203] Figure 9 This is a schematic diagram of the structure of a spectrum acquisition device according to an embodiment of the present disclosure, as shown below. Figure 9 As shown, the training device 900 for the object group extraction model includes a third acquisition module 91, an extraction module 92, and a training module 93, wherein:
[0204] The third acquisition module 91 is used to acquire the candidate object group extraction model to be trained and the sample object text.
[0205] Extraction module 92 is used to obtain a preset set of reference objects in order to extract sample object groups from sample object text.
[0206] Training module 93 is used to train the candidate object group extraction model based on the sample object text and sample object groups until training is completed, resulting in a trained target object group extraction model. The target object group extraction model is used to implement the above... Figure 8 The spectrum acquisition device proposed in the embodiment.
[0207] In this embodiment of the disclosure, the extraction module 92 is further configured to: extract a set of sample objects from the sample object text based on a set of reference objects, and determine the sample object association relationships between each sample object in the sample object set based on the reference association relationships between each reference object in the set of reference objects. Based on the sample object association relationships, a sample object group is obtained based on each sample object.
[0208] In this embodiment of the disclosure, the training module 93 is further configured to: extract candidate object groups of sample object text using a candidate object group extraction model; obtain the training loss of the candidate object group extraction model based on the sample object groups and candidate object groups; adjust the model parameters of the candidate object group extraction model based on the training loss; and return to obtain the next sample object text and the next sample object group to continue training the parameter-adjusted candidate object group extraction model until training is completed, thereby obtaining a trained target object group extraction model.
[0209] In this embodiment of the disclosure, the training module 93 is further configured to: acquire the candidate word segmentation layer, candidate encoder, candidate cause object recognition layer, and candidate effect object recognition layer of the candidate object group extraction model. The candidate word segmentation layer segments the sample object text into words, obtaining a sequence of sample string sequences. The candidate encoder extracts features from each sample string in the sequence of sample string sequences, obtaining a sample feature vector for each sample string. The candidate cause object recognition layer extracts candidate cause objects based on the sample feature vectors of each sample string. The candidate effect object recognition layer extracts candidate effect objects based on the sample feature vectors of each sample string. Based on the candidate cause objects and candidate effect objects, the candidate object group output by the candidate object group extraction model is obtained.
[0210] In this embodiment of the disclosure, the training module 93 is further configured to: identify the start string and end string of the candidate cause object based on the sample feature vector of each sample string through the candidate cause object recognition layer; and obtain the candidate cause object strings in the sample string sequence based on the start string and end string of the candidate cause object to obtain the candidate cause object.
[0211] In this embodiment of the disclosure, the training module 93 is further configured to: identify the start string and end string of the candidate fruit object based on the sample feature vector of each sample string through the candidate fruit object recognition layer; and obtain the candidate fruit object string in the sample string sequence based on the start string and end string of the candidate fruit object to obtain the candidate fruit object.
[0212] The object group extraction model training device disclosed herein acquires sample object text, extracts sample object groups from the sample object text, and trains a candidate object group extraction model based on the sample object groups until training is complete, resulting in a trained target object group extraction model. In this disclosure, the model training of the candidate object group extraction model based on sample object text and sample object groups enables the trained target object group extraction model to extract object groups from publicly available object text, improving the efficiency and accuracy of object group extraction from publicly available object text.
[0213] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0214] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0215] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1009 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0216] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1006, such as various types of monitors, speakers, etc.; storage unit 1009, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0217] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the map acquisition method and / or the object group extraction model training method. For example, in some embodiments, the map acquisition method and / or the object group extraction model training method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1009. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the map acquisition method and / or object group extraction model training method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured by any other suitable means (e.g., by means of firmware) to perform a graph acquisition method and / or a training method for an object group extraction model.
[0218] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0219] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0220] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0221] To initiate interaction with a user account, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user account can submit input to the computer. Other types of devices can also be used to initiate interaction with the user account; for example, feedback submitted to the user account can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user account can be received in any form (including voice input, speech input, or tactile input).
[0222] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user account computer with a graphical user interface or web browser through which a user account can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0223] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0224] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0225] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for obtaining a map, wherein, The method includes: Obtain the preset reference relationship map; Obtain the updated object text, and extract the first updated object group from the updated object text; Identify whether the first updated object group satisfies the map update conditions of the reference relationship map; In response to the identification that the first update object group meets the map update condition, the reference relationship map is updated based on the first update object group to obtain the updated target relationship map; The step of obtaining the updated object text and extracting the first updated object group from the updated object text includes: Extract the update string sequence of the updated object text, and obtain the update feature vector of each update string in the update string sequence; Extract the update factor object based on the update feature vector of each update string; Extract the update result object based on the update feature vector of each updated string; Based on the update cause object and the update effect object, the first update object group of the update object text is obtained; The method further includes: In response to the identification that the first updated object group does not meet the map update conditions, the first updated object group is added to the second updated object group set, and the second updated object group set is clustered to obtain the object clusters of the second updated object group set; Obtain the object tag of the object cluster; In response to the object tag satisfying the preset addition conditions of the reference object set, the object tag is added as a new reference object to the reference object set.
2. The method according to claim 1, wherein, The step of extracting the update factor object based on the update feature vector of each update string includes: Based on the update feature vector of each update string, identify the start string and end string of the update cause object; Based on the start string and end string of the update cause object, obtain the update cause object string from the update string sequence to obtain the update cause object.
3. The method according to claim 1, wherein, The step of extracting the update result object based on the update feature vector of each updated string includes: Based on the update feature vectors of each updated string, identify the start string and end string of the updated result object; Based on the start string and end string of the updated result object, obtain the updated result object string from the updated string sequence to obtain the updated result object.
4. The method according to claim 1, wherein, The step of identifying whether the first updated object group satisfies the map update conditions of the reference relationship map includes: From the reference relationship map, obtain the reference association relationships between each reference object; Obtain the update associations of the first update object group; For any reference association, obtain the similarity between the reference association and the updated association; Based on the relationship similarity, determine whether the first updated object group satisfies the map update condition.
5. The method according to claim 4, wherein, The step of identifying whether the first updated object group satisfies the graph update condition based on the relationship similarity includes: In response to the fact that the similarity between any reference association and the updated association is greater than or equal to a preset similarity threshold, it is determined that the first updated object group satisfies the map update condition; In response to the fact that the similarity between each reference association in the reference relationship graph and the updated association is less than the relationship similarity threshold, it is confirmed that the first updated object group does not meet the graph update conditions.
6. The method according to claim 1, wherein, The step of updating the reference relation graph based on the first updated object group to obtain the updated target relation graph in response to recognizing that the first updated object group satisfies the graph update condition includes: In response to the identification that the first updated object group satisfies the map update condition, the associated object group of the first updated object group in the reference relationship map is obtained; Obtain the associated object relationships of the associated object group; Based on the updated association relationship and the associated object relationship, the first updated object group is merged into the associated object group to update the reference relationship graph and obtain the target relationship graph.
7. The method according to claim 1, wherein, The process of obtaining the preset reference relationship map includes: Obtain the preset set of reference objects, and the reference object description of each reference object; Obtain the reference relationships between various reference objects; The reference object descriptions of each reference object and the reference relationships between each reference object are input into the graph database Bgraph. Through the graph construction capabilities of the graph database Bgraph, the reference relationship graph of the reference object set is obtained.
8. The method according to claim 1, wherein, The step of adding the object tag as a new reference object to the reference object set in response to the object tag satisfying the preset addition conditions of the reference object set includes: Based on the new reference object, the nodes of the reference relationship graph are updated to obtain a new reference relationship graph.
9. A training method for an object group extraction model, wherein, The method includes: Obtain the candidate object group extraction model to be trained and the sample object text; Obtain a preset set of reference objects to extract sample object groups from the sample object text; Based on the sample object text and the sample object group, the candidate object group extraction model is trained until the training is completed, and a trained target object group extraction model is obtained, wherein the target object group extraction model is used to implement the map acquisition method according to any one of claims 1-8.
10. The method according to claim 9, wherein, The step of obtaining a preset set of reference objects to extract sample object groups from the sample object text includes: Based on the reference object set, extract the sample object set from the sample object text, and determine the sample object association relationship between each sample object in the sample object set based on the reference association relationship between each reference object in the reference object set; Based on the association relationship between the sample objects, the sample object group is obtained based on each sample object.
11. The method according to claim 9, wherein, The step of training the candidate object group extraction model based on the sample object text and the sample object group until training is completed, to obtain the trained target object group extraction model, includes: The candidate object group extraction model is used to extract candidate object groups from the sample object text. Based on the sample object group and the candidate object group, obtain the training loss of the candidate object group extraction model; The model parameters of the candidate object group extraction model are adjusted according to the training loss, and the next sample object text and the next sample object group are obtained to continue training the candidate object group extraction model with adjusted parameters until the training ends, and the trained target object group extraction model is obtained.
12. The method according to claim 11, wherein, The step of extracting candidate object groups from the sample object text using the candidate object group extraction model includes: Obtain the candidate word segmentation layer, candidate encoder, candidate cause object recognition layer, and candidate effect object recognition layer of the candidate object group extraction model; The candidate word segmentation layer is used to segment the sample object text to obtain a sample string sequence of the sample object text. The candidate encoder extracts features from each sample string in the sample string sequence to obtain the sample feature vector of each sample string. Candidate cause objects are extracted based on the sample feature vectors of each sample string through the candidate cause object identification layer. The candidate fruit object recognition layer extracts candidate fruit objects based on the sample feature vectors of each sample string. Based on the candidate cause objects and the candidate effect objects, the candidate object group output by the candidate object group extraction model is obtained.
13. The method according to claim 12, wherein, The step of extracting candidate cause objects through the candidate cause object identification layer based on the sample feature vectors of each sample string includes: The candidate cause object identification layer identifies the start string and end string of the candidate cause object based on the sample feature vector of each sample string. Based on the candidate cause object start string and the candidate cause object end string, obtain the candidate cause object string from the sample string sequence to obtain the candidate cause object.
14. The method according to claim 13, wherein, The step of extracting candidate fruit objects through the candidate fruit object recognition layer based on the sample feature vectors of each sample string includes: The candidate fruit object recognition layer identifies the start string and end string of the candidate fruit object based on the sample feature vector of each sample string. Based on the starting string and ending string of the candidate fruit object, obtain the candidate fruit object string from the sample string sequence to obtain the candidate fruit object.
15. A spectrum acquisition device, wherein, The device includes: The first acquisition module is used to acquire a preset reference relationship map; The second acquisition module is used to acquire the updated object text and extract the first updated object group from the updated object text; The identification module is used to identify whether the first updated object group satisfies the map update conditions of the reference relationship map; The update module is used to update the reference relation graph based on the first update object group in response to the recognition that the first update object group meets the graph update condition, so as to obtain the updated target relation graph. The second acquisition module is further configured to: Extract the update string sequence of the updated object text, and obtain the update feature vector of each update string in the update string sequence; Extract the update factor object based on the update feature vector of each update string; Extract the update result object based on the update feature vector of each updated string; Based on the update cause object and the update effect object, the first update object group of the update object text is obtained; The update module is also used for: In response to the identification that the first updated object group does not meet the map update conditions, the first updated object group is added to the second updated object group set, and the second updated object group set is clustered to obtain the object clusters of the second updated object group set; Obtain the object tag of the object cluster; In response to the object tag satisfying the preset addition conditions of the reference object set, the object tag is added as a new reference object to the reference object set.
16. A training apparatus for an object group extraction model, wherein, The device includes: The third acquisition module is used to acquire the candidate object group extraction model to be trained and the sample object text; The extraction module is used to obtain a preset set of reference objects in order to extract sample object groups from the sample object text. The training module is used to train the candidate object group extraction model based on the sample object text and the sample object group until the training is completed, so as to obtain the trained target object group extraction model, wherein the target object group extraction model is used to implement the map acquisition device of claim 15.
17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8 and / or 9-14.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8 and / or 9-14.
19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8 and / or 9-14.
Citation Information
Patent Citations
Entity updating method and system of knowledge graph
CN115809340A