Data verification method and device, electronic equipment and storage medium

By slicing the mind map and automatically judging the logical relationship categories using the classification model, the problem of difficult to guarantee the correctness of data verification caused by relying on manual verification in the prior art is solved, and efficient and accurate data verification and high-quality generation of mind maps is achieved.

CN119961260AActive Publication Date: 2025-05-09IFLYTEK CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510438933.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the prior art, manual experience is used to verify the logical relationship between data content, which makes it difficult to guarantee the correctness of data verification.

Method used

By slicing the mind map to be processed, slice features are extracted, and classification model trained based on the logical relationship category label of multiple sample mind maps and their slice data pairs is automatically judged for verification.

Benefits of technology

It realizes automatic and accurate judgment of the logical relationship categories of mind maps, significantly improves the efficiency and accuracy of data verification, reduces the impact of human factors on the correctness of data verification, and ensures high-quality generation of mind maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961260A_ABST
    Figure CN119961260A_ABST
Patent Text Reader

Abstract

The invention provides a data verification method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: obtaining a to-be-processed mind map; performing slicing processing on the mind map to be processed, and according to a segmentation result, obtaining slice features of each piece of slice data to be processed in the mind map to be processed; inputting the slice features of every two pieces of to-be-processed slice data into a classification model to obtain a logical relationship category between every two pieces of to-be-processed slice data; and according to the logical relationship category between every two pieces of to-be-processed slice data, verifying the to-be-processed mind map to obtain a target mind map corresponding to the to-be-processed mind map. According to the method, automatic and accurate judgment on the logical relationship category of the mind map is realized, the efficiency and accuracy of data verification are remarkably improved, the influence of human factors on the correctness of data verification is reduced, and high-quality generation of the mind map is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data verification method, device, electronic equipment and storage medium. Background Art

[0002] At present, the automatic generation of mind maps using large models requires a large amount of high-quality mind map data for training. The steps for screening and constructing high-quality mind map data are basically: retrieve data related to the problem through the retrieval system, and initially generate data that meets the format and content requirements through instruction tuning. In the process of data screening and annotation, the standards for high-quality data are specified, and data screening is completed through manual screening and modification.

[0003] However, when confirming whether the generated data is of high quality through manual annotation, the annotator needs to spend a lot of time to understand the content in order to determine whether the logical relationship between the data content is correct to generate high-quality mind map data. The inspection performance of this inspection method depends to a large extent on the level and status of the annotator. If the understanding is insufficient during annotation, it is easy to annotate logically flawed data as high-quality data, making it difficult to ensure the correctness of data verification. Summary of the invention

[0004] The present invention provides a data verification method, device, electronic device and storage medium, which are used to solve the defect that the logical relationship verification between data contents relies on manual experience in the prior art, resulting in difficulty in ensuring the correctness of data verification, and realize automatic and accurate data verification.

[0005] The present invention provides a data verification method, comprising: Get the mind map to be processed; Slice the mind map to be processed, and obtain slice features of each slice data to be processed in the mind map to be processed according to the slice result; Inputting the slice features of every two slice data to be processed into a classification model to obtain a logical relationship category between every two slice data to be processed; According to the logical relationship category between every two of the slice data to be processed, the mind map to be processed is verified to obtain a target mind map corresponding to the mind map to be processed; The classification model is obtained by training based on sample mind maps of multiple different data types and logical relationship category labels of sample slice data pairs corresponding to the sample mind maps.

[0006] According to a data verification method provided by the present invention, the classification model is trained based on the following steps: According to two sample slice data with normal logical relationship in the sample mind map, construct a positive sample slice data pair corresponding to the sample mind map, and mark the positive sample slice data pair as having a normal logical relationship; Preprocessing the sample mind map; According to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, construct a negative sample slice data pair corresponding to the sample mind map, and mark the negative sample slice data pair as an abnormal logical relationship; Iteratively training the initialization model according to the slice features of the sample slice data in the negative sample slice data pair, the slice features of the sample slice data in the positive sample slice data pair, the logical relationship category labels of the positive sample slice data pair, and the logical relationship category labels of the negative sample slice data pair to obtain the classification model; The preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols.

[0007] According to a data verification method provided by the present invention, the step of obtaining slice features of each slice data to be processed in the mind map to be processed according to the segmentation result includes: According to the segmentation result, obtaining content information and location information of each of the slice data to be processed; According to the content information, the location information, and the subject information of the mind map to be processed, the slice features of each slice data to be processed are obtained by combination.

[0008] According to a data verification method provided by the present invention, the slicing process of the mind map to be processed includes: The mind map to be processed is sliced ​​according to the parsing tool corresponding to the data format of the mind map to be processed.

[0009] According to a data verification method provided by the present invention, the step of obtaining a mind map to be processed includes: Inputting the preset mind map generated prompt text and each data to be processed into the large language model to obtain the initialized mind map corresponding to each data to be processed; According to the preset mind map function requirements, the mind map to be processed is obtained by screening out the multiple initialized mind maps.

[0010] According to a data verification method provided by the present invention, the method further includes: Obtaining target mind maps corresponding to the plurality of mind maps to be processed; Based on the target mind maps and data to be processed corresponding to the multiple mind maps to be processed, as well as the preset mind map generation prompt text, the large language model is iteratively trained to obtain a mind map generation model corresponding to the preset mind map function requirements, and the mind map generation model is used to generate a mind map that meets the preset mind map function requirements.

[0011] According to a data verification method provided by the present invention, the mind map to be processed is verified according to the logical relationship category between each two slice data to be processed to obtain a target mind map corresponding to the mind map to be processed, including: Acquire an abnormal slice data pair formed by two slice data to be processed whose logical relationship category is an abnormal logical relationship; Outputting annotation prompt information of the mind map to be processed according to the slice features of the slice data to be processed in each of the abnormal slice data pairs; Receive a verification operation corresponding to the mind map to be processed; the verification operation is generated by the verification end according to the annotation prompt information and is used to perform a mind map verification process; According to the verification operation, the mind map to be processed is verified to obtain the target mind map.

[0012] The present invention also provides a data verification device, comprising: An acquisition unit, used for acquiring the mind map to be processed; A slicing unit, used to perform slicing processing on the mind map to be processed, and obtain slicing features of each slicing data to be processed in the mind map to be processed according to the slicing result; A classification unit, used for inputting the slice features of every two slice data to be processed into a classification model to obtain a logical relationship category between every two slice data to be processed; A verification unit, used for verifying the mind map to be processed according to the logical relationship category between every two slice data to be processed, and obtaining a target mind map corresponding to the mind map to be processed; The classification model is obtained by training based on sample mind maps of multiple different data types and logical relationship category labels of sample slice data pairs corresponding to the sample mind maps.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned data verification methods when executing the computer program.

[0014] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the data verification method described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned data verification methods.

[0016] The data verification method, device, electronic device and storage medium provided by the present invention perform slicing processing on the mind map to be processed to extract slice features; then, a classification model pre-trained based on the logical relationship category labels of multiple sample mind maps and their slice data pairs is used to automatically determine the logical relationship category between every two slice data in the mind map to be processed, so as to verify the mind map to be processed according to the identified logical relationship category, thereby generating a high-quality target mind map, thereby achieving automatic and accurate judgment of the logical relationship category of the mind map, significantly improving the efficiency and accuracy of data verification, reducing the influence of human factors on the correctness of data verification, and ensuring the high-quality generation of mind maps. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a flowchart of the mind map generation model training provided by the prior art.

[0019] Figure 2 This is one of the flow charts of the data verification method provided by the present invention.

[0020] Figure 3 This is the second flow chart of the data verification method provided by the present invention.

[0021] Figure 4 It is a structural schematic diagram of the data verification device provided by the present invention.

[0022] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The large language model has achieved remarkable results in many natural language processing tasks due to its excellent general language understanding and generation capabilities. However, when the large language model realizes the understanding and generation functions of mind mapping tasks, it needs a large amount of high-quality mind mapping data for training. Figure 1 As shown in the figure, the steps of screening and constructing high-quality data for existing mind maps are basically: retrieve data related to the problem through the retrieval system, and initially construct and generate data that meets the format and content requirements through instruction tuning; then, in the process of data screening and annotation, specify the standards for high-quality data, and complete the data screening through manual screening and modification. After obtaining high-quality data for mind maps, the big model can be trained based on this data so that the big model can realize the functions of understanding and generating tasks for mind maps.

[0025] The main issues that need to be annotated in mind maps include: (1) the correctness of the content; (2) the rationality of the logical relationship of the content; (3) whether there are typos. Whether the generated data is of high quality through manual annotation depends largely on the status of the annotator. When annotating data, the annotator needs to read the entire article and verify it online. The correctness of the content and typos are judged through online retrieval, which can be searched and verified relatively easily. However, logical errors between the contents require the annotator to spend a lot of time to understand the content. If the understanding is not sufficient during annotation, it is easy to annotate the data with logical flaws as high-quality data. Therefore, the annotation of logical errors between the contents depends largely on the annotator's own understanding ability, level and status. However, the understanding ability, level and status of each annotator are different, and it is easy to annotate the data with logical flaws as high-quality data, resulting in the difficulty in ensuring the correctness of data verification. Therefore, how to improve the correctness of logical verification in mind map data is an important issue in improving the quality of mind map annotation data.

[0026] In this regard, the present application provides a data verification method, which verifies the logical relationship between the data of a mind map through a pre-trained classification model. Compared with completely relying on manual verification of mind map data through set mind map annotation representations, this method reduces manual dependence and effectively improves the intelligence and accuracy of data verification, thereby effectively reducing the time spent by annotators in verifying logical errors in mind maps, exposing possible logical errors in mind maps earlier, reducing the possibility of introducing low-quality data in the large model training process due to errors made by annotators, and improving the efficiency and accuracy of large model training.

[0027] It should be noted that the execution subject of the method may be a data verification device, which may be a mobile electronic device or a non-mobile electronic device. For example, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer, etc., and the non-mobile electronic device may be a server, a network attached storage, a personal computer, etc., which is not specifically limited in the present invention.

[0028] Figure 2 It is one of the flow charts of the data verification method provided by the present invention; Figure 2 As shown, the method includes step 210 , step 220 , step 230 and step 240 .

[0029] Step 210, obtaining the mind map to be processed.

[0030] The mind map to be processed here is the mind map that needs to be verified for data. For example, it can be obtained by automatically screening the mind map according to the mind map function requirements required for training the large model, or it can be formed by user input in real time according to data verification requirements, or it can be generated in real time according to the prompt text generated by the mind map. This embodiment does not make specific limitations on this.

[0031] The user input here can be information input through a command line interface, a graphical interface, touch input, a drop-down selection input, voice input, gesture input, visual input, brain-computer input, etc. This embodiment does not specifically limit this.

[0032] Step 220, slicing the mind map to be processed, and obtaining the slicing features of each slicing data to be processed in the mind map to be processed according to the slicing result.

[0033] Optionally, after obtaining the mind map to be processed, the mind map to be processed can be sliced ​​to divide the mind map to be processed into multiple slice data to be processed, and based on the segmentation results, slice features of each slice data to be processed are generated to facilitate subsequent better analysis of the logical relationship of the data in the mind map to be processed.

[0034] The slicing processing here can be implemented by slicing according to the data structure and / or data format of the mind map to be processed, so as to ensure the integrity of the slice data to be processed obtained by slicing.

[0035] The slice feature here can be formed by combining slice information formed by slices to form corresponding slice features; or, after extracting features from slice information formed by slices, corresponding slice features are formed, etc. This embodiment does not specifically limit this. The slice information includes but is not limited to content information, location information, subject information, etc. of slice data.

[0036] Step 230, inputting the slice features of each two slice data to be processed into the classification model to obtain the logical relationship category between each two slice data to be processed; wherein the classification model is trained based on sample mind maps of multiple different data types, and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0037] Optionally, after obtaining slice features of multiple slice data to be processed, the slice features of every two slice data to be processed can be combined in pairs and input into the classification model respectively, so that the classification model can apply the slice features of every two slice data to be processed to identify whether there is an unreasonable state in the logical relationship between every two slice data to be processed. If there is an unreasonable state in the logical relationship between any two slice data to be processed, the logical relationship category between the two slice data to be processed is determined as an abnormal logical relationship. If there is no unreasonable state in the logical relationship between any two slice data to be processed, the logical relationship category between the two slice data to be processed is determined as a normal logical relationship, so that errors in the mind map to be processed can be quickly located according to the logical relationship categories between different slice data to be processed, so as to assist the annotation personnel to quickly and accurately modify or re-annotate.

[0038] The classification model is a model used to classify the logical relationship between slice data, which can be pre-trained based on the following steps: First, collect sample mind maps in small batches with a data volume less than a set amount and covering multiple different data types; the sample mind maps here can be dynamically generated using a large language model, based on sample content data, under the guidance of map generation prompt text; or they can be pre-loaded and obtained from a database, which is not specifically limited in this embodiment.

[0039] After obtaining the sample mind map, the sample mind map may be segmented to obtain a plurality of sample slice data, and each two sample slice data are combined into a sample slice data pair. The logical relationship category between the two sample slice data in each sample slice data pair is labeled, thereby obtaining a logical relationship category label of each sample slice data pair.

[0040] Then, a sample data set constructed by using each sample slice data pair and the logical relationship category label of each sample slice data pair is used as a training data set to iteratively train the initialization model to obtain a classification model that can automatically and accurately verify the logical relationship of mind map data.

[0041] Step 240: verify the mind map to be processed according to the logical relationship category between every two slice data to be processed, and obtain the target mind map corresponding to the mind map to be processed.

[0042] Optionally, after obtaining the logical relationship category between each two slice data to be processed, data annotation reference recommendation information corresponding to the mind map to be processed can be given based on the logical relationship category between each two slice data to be processed, so as to verify the mind map to be processed, thereby obtaining a high-quality target mind map.

[0043] The verification here can be achieved by sending the data annotation reference recommendation information to the user end to receive the verification operation input by the user end based on the data annotation reference recommendation information, or it can be achieved by a data verification device or other controller automatically performing verification based on the data annotation reference recommendation information. This embodiment does not make any specific limitations on this.

[0044] The method provided in this embodiment performs slicing processing on the mind map to be processed to extract slice features; then, a classification model pre-trained based on logical relationship category labels of multiple sample mind maps and their slice data pairs is used to automatically determine the logical relationship category between every two slice data in the mind map to be processed, so as to verify the mind map to be processed according to the identified logical relationship category, thereby generating a high-quality target mind map, thereby achieving automatic and accurate judgment of the logical relationship category of the mind map, significantly improving the efficiency and accuracy of data verification, reducing the influence of human factors on the correctness of data verification, and ensuring high-quality generation of mind maps.

[0045] In some embodiments, the classification model is trained based on the following steps: According to two sample slice data with normal logical relationship in the sample mind map, construct a positive sample slice data pair corresponding to the sample mind map, and mark the positive sample slice data pair as having a normal logical relationship; Preprocessing the sample mind map; According to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, construct a negative sample slice data pair corresponding to the sample mind map, and mark the negative sample slice data pair as an abnormal logical relationship; Iteratively training the initialization model according to the slice features of the sample slice data in the negative sample slice data pair, the slice features of the sample slice data in the positive sample slice data pair, the logical relationship category labels of the positive sample slice data pair, and the logical relationship category labels of the negative sample slice data pair to obtain the classification model; The preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols.

[0046] Optionally, the classification model mainly implements the logical relationship verification of the slice data of the functional mind map. In the training data set for building the classification model, it is necessary to ensure that the data types of the sample mind map are widely covered to avoid the occurrence of undetermined logical relationships in the subsequent model use process. Therefore, when building the training data set, it is necessary to include a sufficient number of positive sample slice data pairs.

[0047] Among them, the normal logical relationship in the mind map includes that the contents under the same branch are in parallel or progressive relationship, and the contents under different branches are generally parallel or unrelated; while the abnormal logical relationship in the mind map includes that the contents under the same branch are neither parallel nor progressive, and / or the same content is introduced under different branches.

[0048] When using the sample mind map generated by the large model, the logical relationships of most sample slice data pairs are normal logical relationships. Therefore, two sample slice data with normal logical relationships can be directly selected from the sample mind map to construct a positive sample slice data pair.

[0049] It should be noted that in the process of labeling the logical relationship of the positive sample slice data pairs, there are generally parallel, progressive or irrelevant labels. In the process of labeling the positive sample slice data pairs, the accuracy of the labels of the positive sample slice data pairs should be guaranteed. Even if there are certain errors in the labeling, since the positive sample slice data pairs are marked with the labels of the positive sample slice data pairs, the final classification will also be classified into the classification of the positive sample, so it will not cause serious impact on the model training.

[0050] However, since negative sample slice data pairs account for a small proportion in the mind map, it is difficult to quickly obtain sufficient negative sample slice data pairs if negative sample slice data pairs are directly screened from the sample mind map. Therefore, in order to easily obtain sufficient negative sample slice data pairs, the sample mind map can be preprocessed to increase the data of abnormal logical relationships in the sample mind map. For example, some content in the sample mind map is repeated to increase the data of abnormal logical relationships with repeated content, or some sentences in the sample mind map are split based on preset symbols to increase the data of abnormal logical relationships of multiple split types.

[0051] Among them, the repetition processing can be to randomly select the content data of some slice data in the sample mind map, use the large model to fine-tune and modify the language expression to perform repetition processing on part of the content in the sample mind map, and then mark it as a negative sample slice data pair with repeated content on the sample mind map; the sentence splitting processing can be to select the content data of a sentence from the slice data of the sample mind map, split it into multiple sentences according to punctuation marks, and use this as a negative sample slice data pair with multiple splits of one sentence in the sample mind map.

[0052] It should be noted that the data annotation process of the negative sample slice data pairs needs to ensure the accuracy of the labels, because the logical relationship errors between the slices of the mind map need to be correctly pointed out during the use of the model.

[0053] For example, the representation of the positive sample slice data pair can be: [{"input1": "Core character: Paul Kochakin","input2": "Main plot","count1": "h1:1h2:2","count2": "h1:1h2:2h3:1","head": "How the Steel Was Tempered","target": "Progressive"},{"input1": "Author's personal background","input2": "Donya","count1": "h1:1h2:1","count2": "h1:1h2:3h3:1","head": "How the Steel Was Tempered","target": "Irrelevant"},{"input1": "Paul's growth experience","input2": "Joining the army and fighting","count1": "h1:1h2:1h3:1ul:1","count2": "h1:1h2:1h3:1ul:2","head": "How the Steel Was Tempered","target": "Parallel"}].

[0054] The representation of negative sample slice data pairs can be: [{"input1": "Technological development","input2": "Issuance of L3 autonomous driving test licenses","count1": "h1:1h2:1h3:2","count2": "h1:1h2:1h3:2ul:1","head": "Application of AI in automobiles","target": "Not relevant"},{"input1": "Verdun meat grinder","input2": "The battle was extremely fierce","count1": "h1:1h2:1h3:2lu:1","count2": "h1:1h2:1h3:2lu:2","head": "—Main battle, battle","target":"Forced split"},{"input1":"Dabai Electric Appliances' revenue in the first quarter increased by 10% year-on-year","input2":"Dabai Electric Appliances' revenue grew rapidly","count1": "h1:1h2:1h3:1ul:1","count2": "h1:1h2:3h3:2ul:1","head": "Financial Analysis of the Group","target": "Duplicated"}].

[0055] Among them, "input1" and "input2" are the content information of two slice data; "count1" and "count2" are the hierarchical structures of two slice data; "head" is the subject information; "target" is the logical relationship information; After obtaining the negative sample slice data pairs and the positive sample slice data pairs, the slice features of the sample slice data in the negative sample slice data pairs and the slice features of the sample slice data in the positive sample slice data pairs can be used as samples, and the logical relationship category labels of the positive sample slice data pairs and the logical relationship category labels of the negative sample slice data pairs can be used as labels to iteratively train the initialization model to obtain a classification model that can automatically and accurately identify normal logical relationship categories and abnormal logical relationship categories between different slice data pairs.

[0056] The method provided in this embodiment constructs positive and negative sample slice data pairs based on sample mind maps, and combines preprocessing technology to increase data of abnormal logical relationships, and trains a classification model that can accurately identify normal and abnormal logical relationships between slice data pairs in mind maps, thereby effectively improving the automation and accuracy of logical relationship verification and ensuring the coverage and robustness of the model on a wide range of data types.

[0057] In some embodiments, step 220 specifically includes: The mind map to be processed is sliced ​​according to the parsing tool corresponding to the data format of the mind map to be processed.

[0058] Optionally, when performing slicing processing, a parsing tool that matches the data format of the mind map to be processed can be used, and appropriate labels can be used to mark each branch data in the mind map to be processed to achieve slicing processing of the mind map to be processed.

[0059] For example, if the data format of the mind map to be processed is data in the lightweight markup language (abbreviated as markdown, or MD) format, the markdown parsing library can be used to slice the mind map to be processed.

[0060] The slice processing logic of the markdown parsing library can be implemented through the following code: “# Convert Markdown string to HTML html = markdown.markdown(markdown_string) # Parsing HTML with BeautifulSoup soup = BeautifulSoup(html, 'html.parser') def extract_structure(soup): structure = [] for element in soup.children: if element.name: structure.append({ 'tag': element.name, 'text': element.get_text(strip=True), 'children': extract_structure(element) if element.contents else [] }) Return structure ”.

[0061] For example, the mind map to be processed can be divided into the following formats through the markdown parsing library: [ {"text": "Introduction to Art Studies by Peng Jixiang, Fourth Edition, Mind Map", "count": "h1:1"}, {"text": "General Introduction to Art", "count": "h1:1h2:1"}, {"text": "The Essence and Characteristics of Art", "count": "h1:1h2:1h3:1"}, {"text": "The Essence of Art", "count": "h1:1h2:1h3:1ul:1"}, {"text": "Artistic Features", "count": "h1:1h2:1h3:1ul:2"}, {"text": "The Origin of Art", "count": "h1:1h2:1h3:2"}, {"text": "Five Views", "count": "h1:1h2:1h3:2ul:1"}, {"text": "Multi-determinism", "count": "h1:1h2:1h3:2ul:2"}, {"text": "The Function of Art and Art Education", "count": "h1:1h2:1h3:3"}, {"text": "Social Function", "count": "h1:1h2:1h3:3ul:1"}, {"text": "Art Education", "count": "h1:1h2:1h3:3ul:2"}, {"text": "Art in the Cultural System", "count": "h1:1h2:1h3:4"}, {"text": "Cultural phenomenon", "count": "h1:1h2:1h3:4ul:1"}, {"text": "Art and Philosophy", "count": "h1:1h2:1h3:4ul:2"}, {"text": "Art and Religion", "count": "h1:1h2:1h3:4ul:3"}, {"text": "Art and Morality", "count": "h1:1h2:1h3:4ul:4"}, {"text": "Art and Science", "count": "h1:1h2:1h3:4ul:5"}, {"text": "Art Types", "count": "h1:1h2:2"}, {"text": "Practical Art", "count": "h1:1h2:2h3:1"}, {"text": "Main categories", "count": "h1:1h2:2h3:1ul:1"}, {"text": "Aesthetic characteristics", "count": "h1:1h2:2h3:1ul:2"}, {"text": "Exquisite Appreciation", "count": "h1:1h2:2h3:1ul:3"}, {"text": "Plastic Arts", "count": "h1:1h2:2h3:2"}, {"text": "Main categories", "count": "h1:1h2:2h3:2ul:1"}, {"text": "Aesthetic characteristics", "count": "h1:1h2:2h3:2ul:2"}, {"text": "Exquisite Appreciation", "count": "h1:1h2:2h3:2ul:3"}, {"text": "Emoji Art", "count": "h1:1h2:2h3:3"}, {"text": "Main categories", "count": "h1:1h2:2h3:3ul:1"}, {"text": "Aesthetic characteristics", "count": "h1:1h2:2h3:3ul:2"}, {"text": "Exquisite Appreciation", "count": "h1:1h2:2h3:3ul:3"}, {"text": "Comprehensive Arts", "count": "h1:1h2:2h3:4"}, {"text": "Main categories", "count": "h1:1h2:2h3:4ul:1"}, {"text": "Aesthetic characteristics", "count": "h1:1h2:2h3:4ul:2"}, {"text": "Exquisite Appreciation", "count": "h1:1h2:2h3:4ul:3"}, {"text": "Language Arts", "count": "h1:1h2:2h3:5"}, {"text": "Main genre", "count": "h1:1h2:2h3:5ul:1"} ].

[0062] Among them, text is used to describe the content information related to each slice data to be processed. Count is used to describe the hierarchical relationship of each slice data to be processed in the mind map.

[0063] The method provided in this embodiment ensures the integrity and accuracy of information by using a parsing tool that matches the format of the mind map data to be processed for slicing and marking each branch data with appropriate labels, so that the logical relationship of the data in the mind map can be accurately parsed and verified based on the segmented data.

[0064] In some embodiments, step 220 further includes: According to the segmentation result, obtaining content information and location information of each of the slice data to be processed; According to the content information, the location information, and the subject information of the mind map to be processed, the slice features of each slice data to be processed are obtained by combination.

[0065] Alternatively, since the construction of mind map slice data will seriously affect the final effect of the classification model, such as how to determine whether the two slices belong to the same branch, how to determine whether the same keywords appear in the content of the two slices are a normal logical relationship or a content duplication. Therefore, in the process of constructing mind map slices, simply constructing a string of content information of the two slices to be processed and inputting it into the classification model for logical relationship classification will lead to large classification errors.

[0066] In this regard, the method provided in this embodiment constructs the slice features of each slice data to be processed by combining the content information, location information and subject information of the mind map to be processed of each slice data to be processed, and combines the slice features of two slice data to be processed as the input of the classification model, so as to help the classification model to more accurately determine whether the two slice data to be processed belong to the same branch, and whether the logical relationship between the two is normal or repeated in content, thereby reducing classification errors, improving the classification model's ability to understand and judge the logical relationship of mind map data, and significantly improving the accuracy and robustness of the classification model. The specific implementation steps are as follows: When acquiring the segmentation features, the segmentation results may be parsed to obtain the content information and location information of each slice data to be processed.

[0067] The step of obtaining the position information here includes: according to the set order corresponding to the slice label marked for each slice data to be processed during the slicing process, the position of each slice data to be processed in the mind map to be processed can be determined, thereby obtaining the corresponding position information.

[0068] The slice label here can be a character string used to mark the hierarchical relationship and position of content information of different slice data in the mind map, such as "h1, h2, h3, ul" format, etc. This embodiment does not make any specific limitation on this.

[0069] In the mind map verification, some repeated strings can be judged as content duplication, while some repeated strings cannot be judged as content duplication. For example, the words that appear in the theme information of the mind map are generally regarded as keywords. If the keywords appear repeatedly, they will not be judged as content duplication, while other identical contents will be judged as content duplication. Therefore, when determining the slice features to be input into the classification model, it is necessary not only to introduce content information and location information, but also to introduce the theme information of the mind map to be processed into the slice features, so as to avoid the misjudgment of the classification model caused by multiple appearances of the theme, thereby improving the accuracy of data verification.

[0070] The method provided in this embodiment constructs slice features by combining content information, location information and subject information, and uses the slice features as the input of the classification model to effectively distinguish normal logical relationships and content duplications within the same branch, reduce classification errors, and enhance the classification model's ability to understand and judge the structure of the mind map, thereby significantly improving the accuracy and robustness of the logical relationship verification of the mind map data.

[0071] In some embodiments, step 210 specifically includes: Inputting the preset mind map generated prompt text and each data to be processed into the large language model to obtain the initialized mind map corresponding to each data to be processed; According to the preset mind map function requirements, the mind map to be processed is obtained by screening out the multiple initialized mind maps.

[0072] Optionally, the mind map to be processed can be automatically screened through the following steps: Collecting the preset mind map to generate prompt text (prompt for short) and each data to be processed; inputting the preset mind map to generate prompt text and each data to be processed into the large language model, so that the large language model, under the guidance of the preset mind map to generate prompt text, uses the instruction call method to preliminarily generate the initialization mind map corresponding to each data to be processed; Furthermore, in order to generate high-quality sample mind maps to participate in the fine-tuning training of the large language model, so that the trained large language model has the preset mind map function requirements, the initialized mind map that matches the preset mind map function requirements can be screened from multiple initialized mind maps as the mind map to be processed.

[0073] The preset mind map generation prompt text here is used to guide the large language model to generate a mind map, which can be generated based on a large number of mind map questions collected from the Internet, etc.; the preset mind map generation prompt text includes but is not limited to the data type, data logic, architecture, format, etc. of the mind map to be generated, and this embodiment does not specifically limit this.

[0074] The data to be processed here may be information or knowledge points used to generate a mind map, including but not limited to text information, data tables, keywords and phrases, etc. This embodiment does not specifically limit this.

[0075] Among them, the large language model (LLM), referred to as the large model or large language model, refers to a natural language processing (NLP) model with a large parameter scale. The number of model parameters and / or the complexity of the model structure exceeds the preset threshold. The model is pre-trained on large-scale text data and has a high degree of semantic understanding and the ability to generate natural language.

[0076] The method provided in this embodiment generates prompt text and data to be processed by combining a large language model with a preset mind map, and automatically screens and generates a mind map that meets the functional requirements, so as to effectively improve the efficiency and quality of mind map generation, so that a large language model with the preset mind map functional requirements can be trained more quickly in the future.

[0077] In some embodiments, the method further comprises: Obtaining target mind maps corresponding to the plurality of mind maps to be processed; Based on the target mind maps and data to be processed corresponding to the multiple mind maps to be processed, as well as the preset mind map generation prompt text, the large language model is iteratively trained to obtain a mind map generation model corresponding to the preset mind map function requirements, and the mind map generation model is used to generate a mind map that meets the preset mind map function requirements.

[0078] Optionally, a large number of mind maps to be processed with a data volume exceeding a preset number can be collected and Figure 2 The data verification process shown performs corresponding data verification operations on a large number of collected mind maps to be processed, so as to obtain target mind maps corresponding to each mind map to be processed, thereby obtaining sample data with wide data type coverage, large data volume, and high quality that can be used to assist in the training of mind map generation models.

[0079] Subsequently, the processing data corresponding to each pending mind map and the preset mind map generation prompt text are taken as samples, and the target mind map corresponding to each pending mind map is used as a label. The large language model is iteratively trained to quickly obtain a mind map generation model, which can automatically and accurately generate a mind map that meets the preset mind map function requirements, thereby improving the degree of automation and accuracy of mind map generation and further optimizing the user experience.

[0080] In some embodiments, step 240 specifically includes: Acquire an abnormal slice data pair formed by two slice data to be processed whose logical relationship category is an abnormal logical relationship; Outputting annotation prompt information of the mind map to be processed according to the slice features of the slice data to be processed in each of the abnormal slice data pairs; Receive a verification operation corresponding to the mind map to be processed; the verification operation is generated by the verification end according to the annotation prompt information and is used to perform a mind map verification process; According to the verification operation, the mind map to be processed is verified to obtain the target mind map.

[0081] Figure 3 This is the second flow chart of the data verification method provided by the present invention.

[0082] like Figure 3 As shown, a classification model is obtained by training the logical relationship category labels of sample mind maps of multiple different data types and their corresponding sample slice data pairs, and the classification model is used to identify the logical relationship of the screened mind map to be processed corresponding to the functional requirements of the mind map generation model. After obtaining the logical relationship category between each two slice data to be processed in the mind map to be processed, the logical relationship category between each two slice data to be processed can be judged one by one to obtain an abnormal slice data pair formed by two slice data to be processed whose logical relationship category is an abnormal logical relationship, and the slice features of each abnormal slice data pair are recorded. And according to the slice features of the slice data to be processed in each abnormal slice data pair, the annotation prompt information of the mind map to be processed is generated to prompt the annotation personnel to locate the abnormal logical relationship data in the mind map to be processed according to the annotation prompt information, and retrieve the data with wrong data content in combination with the search tool, and input the verification operation of the mind map to be processed at the verification end according to the abnormal logical relationship data and the data with wrong data content.

[0083] After receiving the verification operation corresponding to the mind map to be processed, the mind map to be processed may be marked or modified according to the verification operation, thereby obtaining a high-quality target mind map.

[0084] The method provided in this embodiment realizes the data verification of the mind map by identifying abnormal logical relationship slice data pairs and generating annotation prompt information for human-computer interaction, which effectively improves the efficiency and accuracy of the data verification of the mind map, thereby being able to quickly and accurately generate high-quality target mind maps.

[0085] The data verification device provided by the present invention is described below. The data verification device described below and the data verification method described above can be referenced to each other.

[0086] Figure 4 Schematic diagram of the structure of the data verification device provided by the present invention. Figure 4 As shown, the device comprises: The acquisition unit 410 is used to acquire the mind map to be processed; The slicing unit 420 is used to perform slicing processing on the mind map to be processed, and obtain the slicing features of each slicing data to be processed in the mind map to be processed according to the slicing result; The classification unit 430 is used for inputting the slice features of every two slice data to be processed into the classification model to obtain the logical relationship category between every two slice data to be processed; The verification unit 440 is used to verify the mind map to be processed according to the logical relationship category between every two slice data to be processed, and obtain the target mind map corresponding to the mind map to be processed; The classification model is obtained by training based on sample mind maps of multiple different data types and logical relationship category labels of sample slice data pairs corresponding to the sample mind maps.

[0087] The device provided in this embodiment performs slicing processing on the mind map to be processed to extract slice features; then, a classification model pre-trained based on logical relationship category labels of multiple sample mind maps and their slice data pairs is used to automatically determine the logical relationship category between every two slice data in the mind map to be processed, so as to verify the mind map to be processed according to the identified logical relationship category, thereby generating a high-quality target mind map, thereby achieving automatic and accurate judgment of the logical relationship category of the mind map, significantly improving the efficiency and accuracy of data verification, reducing the influence of human factors on the correctness of data verification, and ensuring high-quality generation of mind maps.

[0088] In some embodiments, the apparatus further comprises a training unit for: According to two sample slice data with normal logical relationship in the sample mind map, construct a positive sample slice data pair corresponding to the sample mind map, and mark the positive sample slice data pair as having a normal logical relationship; Preprocessing the sample mind map; According to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, construct a negative sample slice data pair corresponding to the sample mind map, and mark the negative sample slice data pair as an abnormal logical relationship; Iteratively training the initialization model according to the slice features of the sample slice data in the negative sample slice data pair, the slice features of the sample slice data in the positive sample slice data pair, the logical relationship category labels of the positive sample slice data pair, and the logical relationship category labels of the negative sample slice data pair to obtain the classification model; The preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols.

[0089] In some embodiments, the slicing unit is specifically configured to: According to the segmentation result, obtaining content information and location information of each of the slice data to be processed; According to the content information, the location information, and the subject information of the mind map to be processed, the slice features of each slice data to be processed are obtained by combination.

[0090] In some embodiments, the slicing unit is further configured to: The mind map to be processed is sliced ​​according to the parsing tool corresponding to the data format of the mind map to be processed.

[0091] In some embodiments, the acquisition unit is specifically configured to: Inputting the preset mind map generated prompt text and each data to be processed into the large language model to obtain the initialized mind map corresponding to each data to be processed; According to the preset mind map function requirements, the mind map to be processed is obtained by screening out the multiple initialized mind maps.

[0092] In some embodiments, the training unit is further configured to: Obtaining target mind maps corresponding to the plurality of mind maps to be processed; Based on the target mind maps and data to be processed corresponding to the multiple mind maps to be processed, as well as the preset mind map generation prompt text, the large language model is iteratively trained to obtain a mind map generation model corresponding to the preset mind map function requirements, and the mind map generation model is used to generate a mind map that meets the preset mind map function requirements.

[0093] In some embodiments, the verification unit is specifically configured to: Acquire an abnormal slice data pair formed by two slice data to be processed whose logical relationship category is an abnormal logical relationship; Outputting annotation prompt information of the mind map to be processed according to the slice features of the slice data to be processed in each of the abnormal slice data pairs; Receive a verification operation corresponding to the mind map to be processed; the verification operation is generated by the verification end according to the annotation prompt information and is used to perform a mind map verification process; According to the verification operation, the mind map to be processed is verified to obtain the target mind map.

[0094] The device provided by the present invention is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific processes and detailed contents, which will not be repeated here.

[0095] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the data verification method, which includes: obtaining a mind map to be processed; slicing the mind map to be processed, and obtaining the slice features of each slice data to be processed in the mind map to be processed according to the slicing result; inputting the slice features of each two slice data to be processed into the classification model to obtain the logical relationship category between each two slice data to be processed; verifying the mind map to be processed according to the logical relationship category between each two slice data to be processed, and obtaining the target mind map corresponding to the mind map to be processed; wherein the classification model is obtained by training based on a plurality of sample mind maps of different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0096] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0097] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data verification method provided by the above-mentioned methods, which method includes: obtaining a mind map to be processed; slicing the mind map to be processed, and obtaining the slice features of each slice data to be processed in the mind map to be processed according to the slicing result; inputting the slice features of every two slice data to be processed into a classification model to obtain the logical relationship category between every two slice data to be processed; verifying the mind map to be processed according to the logical relationship category between every two slice data to be processed, and obtaining a target mind map corresponding to the mind map to be processed; wherein the classification model is obtained by training based on sample mind maps of multiple different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0098] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the data verification method provided by the above-mentioned methods, the method comprising: obtaining a mind map to be processed; slicing the mind map to be processed, and obtaining slice features of each slice data to be processed in the mind map to be processed according to the slicing result; inputting the slice features of every two slice data to be processed into a classification model to obtain the logical relationship category between every two slice data to be processed; verifying the mind map to be processed according to the logical relationship category between every two slice data to be processed to obtain a target mind map corresponding to the mind map to be processed; wherein the classification model is obtained by training based on sample mind maps of multiple different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0099] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0100] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data verification method, characterized in that: include: Get the mind map to be processed; Slice the mind map to be processed, and obtain slice features of each slice data to be processed in the mind map to be processed according to the slice result; Inputting the slice features of every two slice data to be processed into a classification model to obtain a logical relationship category between every two slice data to be processed; According to the logical relationship category between every two of the slice data to be processed, the mind map to be processed is verified to obtain a target mind map corresponding to the mind map to be processed; The classification model is obtained by training based on sample mind maps of multiple different data types and logical relationship category labels of sample slice data pairs corresponding to the sample mind maps.

2. The data verification method according to claim 1, characterized in that: The classification model is trained based on the following steps: According to two sample slice data with normal logical relationship in the sample mind map, construct a positive sample slice data pair corresponding to the sample mind map, and mark the positive sample slice data pair as having a normal logical relationship; Preprocessing the sample mind map; According to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, construct a negative sample slice data pair corresponding to the sample mind map, and mark the negative sample slice data pair as an abnormal logical relationship; Iteratively training the initialization model according to the slice features of the sample slice data in the negative sample slice data pair, the slice features of the sample slice data in the positive sample slice data pair, the logical relationship category labels of the positive sample slice data pair, and the logical relationship category labels of the negative sample slice data pair to obtain the classification model; The preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols.

3. The data verification method according to claim 1, characterized in that: The step of obtaining slice features of each slice data to be processed in the mind map to be processed according to the segmentation result includes: According to the segmentation result, obtaining content information and location information of each of the slice data to be processed; According to the content information, the location information, and the subject information of the mind map to be processed, the slice features of each slice data to be processed are obtained by combination.

4. The data verification method according to any one of claims 1 to 3, characterized in that: The slicing process of the mind map to be processed includes: The mind map to be processed is sliced ​​according to the parsing tool corresponding to the data format of the mind map to be processed.

5. The data verification method according to any one of claims 1 to 3, characterized in that: The step of obtaining the mind map to be processed includes: Inputting the preset mind map generated prompt text and each data to be processed into the large language model to obtain the initialized mind map corresponding to each data to be processed; According to the preset mind map function requirements, the mind map to be processed is obtained by screening out the multiple initialized mind maps.

6. The data verification method according to claim 5, characterized in that: The method further comprises: Obtaining target mind maps corresponding to the plurality of mind maps to be processed; Based on the target mind maps and data to be processed corresponding to the multiple mind maps to be processed, as well as the preset mind map generation prompt text, the large language model is iteratively trained to obtain a mind map generation model corresponding to the preset mind map function requirements, and the mind map generation model is used to generate a mind map that meets the preset mind map function requirements.

7. The data verification method according to any one of claims 1 to 3, characterized in that: The method of verifying the mind map to be processed according to the logical relationship category between each two slice data to be processed to obtain a target mind map corresponding to the mind map to be processed includes: Acquire an abnormal slice data pair formed by two slice data to be processed whose logical relationship category is an abnormal logical relationship; Outputting annotation prompt information of the mind map to be processed according to the slice features of the slice data to be processed in each of the abnormal slice data pairs; Receive a verification operation corresponding to the mind map to be processed; the verification operation is generated by the verification end according to the annotation prompt information and is used to perform a mind map verification process; According to the verification operation, the mind map to be processed is verified to obtain the target mind map.

8. A data verification device, characterized in that: include: An acquisition unit, used for acquiring the mind map to be processed; A slicing unit, used to perform slicing processing on the mind map to be processed, and obtain slicing features of each slicing data to be processed in the mind map to be processed according to the slicing result; A classification unit, used for inputting the slice features of every two slice data to be processed into a classification model to obtain a logical relationship category between every two slice data to be processed; A verification unit, used for verifying the mind map to be processed according to the logical relationship category between every two slice data to be processed, and obtaining a target mind map corresponding to the mind map to be processed; The classification model is obtained by training based on sample mind maps of multiple different data types and logical relationship category labels of sample slice data pairs corresponding to the sample mind maps.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the data verification method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data verification method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Document image recognition method and device, electronic equipment and storage medium

    CN117765556A

  • Mind mapping generation method and device based on large language model and electronic equipment

    CN117933377A

  • Intelligent verification method and system for electric power work ticket content

    CN119202680A

  • Causal atlas formation model construction method based on adaptive context learning

    CN119293212A

  • AI-driven multilingual intelligent document abstract and mind map generation system

    CN119494394A