Data verification method, device, electronic device and storage medium

By slicing the mind map and automatically judging the logical relationship using the classification model, the problem of inaccurate data verification caused by manual verification is solved, and efficient and accurate mind map generation is achieved.

CN119961260BActive Publication Date: 2025-07-22IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510438933.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the prior art, relying on manual experience to verify the logical relationship between data contents, making it difficult to guarantee the correctness of data verification. Especially in the process of screening and constructing high-quality data of mind maps, the level and status of the markers have a great impact, and it is easy to label data with defects in logic as high-quality data.

Method used

By slicing the mind map to be processed, slice features are extracted, and a classification model trained on the logical relationship category label based on multiple sample mind maps and their slice data, the logical relationship categories between each two slice data are automatically judged to generate a high-quality target mind map.

Benefits of technology

It realizes automatic and accurate judgment of the logical relationship of mind maps, significantly improves the efficiency and accuracy of data verification, reduces the impact of human factors on the correctness of data verification, and ensures high-quality generation of mind maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961260B_ABST
    Figure CN119961260B_ABST
Patent Text Reader

Abstract

The present invention provides a data verification method, apparatus, electronic device, and storage medium, which relate to the technical field of data processing. The method includes: obtaining a to-be-processed mind map; performing slicing processing on the to-be-processed mind map, and according to the slicing result, obtaining the slicing features of each to-be-processed slice data in the to-be-processed mind map; inputting the slicing features of every two to-be-processed slice data into a classification model to obtain the logical relationship category between every two to-be-processed slice data; verifying the to-be-processed mind map according to the logical relationship category between every two to-be-processed slice data to obtain a target mind map corresponding to the to-be-processed mind map. The present invention realizes the automatic and accurate judgment of the logical relationship category of the mind map, significantly improves the efficiency and accuracy of data verification, reduces the influence of human factors on the correctness of data verification, and ensures the high-quality generation of the mind map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data verification method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, to implement the automatic generation function of mind maps using large models, a large amount of high-quality mind map data is required for training. Generally, the steps for screening and constructing high-quality mind map data are basically as follows: retrieving data related to the problem through a retrieval system, and initially generating data that meets the format and content requirements through instruction tuning. The criteria for high-quality data are specified during the data screening and annotation process, and the data screening is completed through manual screening and modification.

[0003] However, when manually annotating to confirm whether the generated data is high-quality, the annotator needs to spend a lot of time understanding the content to judge whether the logical relationship between the data contents is correct for the verification and generation of high-quality mind map data. The verification performance of this verification method largely depends on the level and state of the annotator. If the understanding is insufficient during annotation, it is easy to label data with logical flaws as high-quality data, resulting in difficulty in ensuring the correctness of data verification. Summary of the Invention

[0004] The present invention provides a data verification method, apparatus, electronic device, and storage medium to solve the defect in the prior art that the verification of the logical relationship between data contents depends on manual experience, resulting in difficulty in ensuring the correctness of data verification, and to achieve automatic and accurate data verification.

[0005] The present invention provides a data verification method, including:

[0006] Obtaining a to-be-processed mind map;

[0007] Performing slicing processing on the to-be-processed mind map, and according to the slicing result, obtaining the slicing features of each to-be-processed slice data in the to-be-processed mind map;

[0008] Inputting the slicing features of every two to-be-processed slice data into a classification model to obtain the logical relationship category between every two to-be-processed slice data;

[0009] Verifying the to-be-processed mind map according to the logical relationship category between every two to-be-processed slice data to obtain a target mind map corresponding to the to-be-processed mind map;

[0010] Wherein, the classification model is trained based on multiple sample mind maps of different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0011] A data verification method provided by the present invention, the classification model is trained based on the following steps:

[0012] According to two sample slice data with normal logical relationships in the sample mind map, construct a pair of positive sample slice data corresponding to the sample mind map, and label the pair of positive sample slice data as having a normal logical relationship;

[0013] Preprocess the sample mind map;

[0014] According to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, construct a pair of negative sample slice data corresponding to the sample mind map, and label the pair of negative sample slice data as having an abnormal logical relationship;

[0015] According to the slice features of the sample slice data in the pair of negative sample slice data, the slice features of the sample slice data in the pair of positive sample slice data, and the logical relationship category labels of the pair of positive sample slice data and the logical relationship category labels of the pair of negative sample slice data, perform iterative training on the initialized model to obtain the classification model;

[0016] Wherein, the preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols.

[0017] A data verification method provided by the present invention, the obtaining the slice features of each to-be-processed slice data in the to-be-processed mind map according to the segmentation result includes:

[0018] According to the segmentation result, obtain the content information and position information of each to-be-processed slice data;

[0019] According to the content information, position information, and the theme information of the to-be-processed mind map, combine to obtain the slice features of each to-be-processed slice data.

[0020] A data verification method provided by the present invention, the slicing the to-be-processed mind map includes:

[0021] Slice the to-be-processed mind map according to the parsing tool corresponding to the data format of the to-be-processed mind map.

[0022] A data verification method provided by the present invention, the obtaining the to-be-processed mind map includes:

[0023] Input the preset mind map generation prompt text and each to-be-processed data into a large language model to obtain an initialized mind map corresponding to each to-be-processed data;

[0024] Filter the to-be-processed mind map from multiple initialized mind maps according to the preset mind map function requirements.

[0025] According to a data verification method provided by the present invention, the method further includes:

[0026] Obtain the target mind maps corresponding to multiple to-be-processed mind maps;

[0027] Based on the target mind maps corresponding to multiple to-be-processed mind maps, the to-be-processed data, and the preset mind map, generate prompt texts to iteratively train the large language model to obtain a mind map generation model corresponding to the preset mind map function requirements, where the mind map generation model is used to generate a mind map that meets the preset mind map function requirements.

[0028] According to a data verification method provided by the present invention, the verification of the to-be-processed mind map according to the logical relationship category between every two to-be-processed slice data to obtain the target mind map corresponding to the to-be-processed mind map includes:

[0029] Obtain an abnormal slice data pair formed by two to-be-processed slice data with an abnormal logical relationship category;

[0030] Output the annotation prompt information of the to-be-processed mind map according to the slice features of the to-be-processed slice data in each abnormal slice data pair;

[0031] Receive the verification operation corresponding to the to-be-processed mind map; the verification operation is an operation generated by the verification end according to the annotation prompt information for performing mind map verification processing;

[0032] Verify the to-be-processed mind map according to the verification operation to obtain the target mind map.

[0033] The present invention also provides a data verification device, including:

[0034] An acquisition unit for acquiring a to-be-processed mind map;

[0035] A slicing unit for slicing the to-be-processed mind map and obtaining the slice features of each to-be-processed slice data in the to-be-processed mind map according to the slicing result;

[0036] A classification unit for inputting the slice features of every two to-be-processed slice data into a classification model to obtain the logical relationship category between every two to-be-processed slice data;

[0037] A verification unit, configured to verify the to-be-processed mind map according to the logical relationship categories between every two pieces of the to-be-processed slice data, so as to obtain a target mind map corresponding to the to-be-processed mind map;

[0038] Wherein, the classification model is trained based on a plurality of sample mind maps of different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0039] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the data verification method as described in any one of the above is implemented.

[0040] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the data verification method as described in any one of the above is implemented.

[0041] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the data verification method as described in any one of the above is implemented.

[0042] The data verification method, device, electronic device, and storage medium provided by the present invention slice the to-be-processed mind map to extract slice features; subsequently, use a classification model pre-trained based on the logical relationship category labels of multiple sample mind maps and their slice data pairs to automatically determine the logical relationship category between every two slice data in the to-be-processed mind map, and verify the to-be-processed mind map according to the identified logical relationship category, so as to generate a high-quality target mind map, thereby realizing the automatic and accurate determination of the logical relationship category of the mind map, significantly improving the efficiency and accuracy of data verification, reducing the influence of human factors on the correctness of data verification, and ensuring the high-quality generation of the mind map. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 is a schematic flowchart of the training of a mind map generation model provided by the prior art.

[0045] Figure 2 is one of the schematic flowcharts of the data verification method provided by the present invention.

[0046] Figure 3 This is the second flow chart of the data verification method provided by the present invention.

[0047] Figure 4 It is a structural schematic diagram of the data verification device provided by the present invention.

[0048] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] The large language model has achieved remarkable results in many natural language processing tasks due to its excellent general language understanding and generation capabilities. However, when the large language model realizes the understanding and generation functions of mind mapping tasks, it needs a large amount of high-quality mind mapping data for training. Figure 1 As shown in the figure, the steps of screening and constructing high-quality data for existing mind maps are basically: retrieve data related to the problem through the retrieval system, and initially construct and generate data that meets the format and content requirements through instruction tuning; then, in the process of data screening and annotation, specify the standards for high-quality data, and complete the data screening through manual screening and modification. After obtaining high-quality data for mind maps, the big model can be trained based on this data so that the big model can realize the functions of understanding and generating tasks for mind maps.

[0051] The main problems that need to be marked in the mind map include: (1) the correctness of the content; (2) the rationality of the logical relationship of the content; (3) the existence of typos. And whether the generated data is of high quality is confirmed by manual marking, which largely depends on the state of the marking personnel. When marking data, the marking personnel need to read the whole text and conduct online verification. For the correctness of the content and the judgment of typos, it can be easily searched and verified through network retrieval. However, for the logical errors between the contents, it takes a lot of time for the marking personnel to understand the content. If the understanding is not sufficient during marking, it is easy to mark the data with logical flaws as high-quality data. Therefore, the marking of logical errors between the contents largely depends on the understanding ability, level and state of the marking personnel themselves. However, the understanding ability, level and state of each marking personnel are different, and it is easy to mark the data with logical flaws as high-quality data, resulting in difficulty in ensuring the correctness of data verification. Therefore, how to improve the logical verification correctness in the mind map data is an important issue for improving the quality of the marked data of the mind map.

[0052] In response to this, the present application provides a data verification method. This method conducts logical relationship verification between the data of the mind map through a pre-trained classification model. Compared with completely relying on manual data verification of the mind map through the set mind map marking representation, it reduces the dependence on manual work, effectively improves the intelligence and accuracy of data verification, and thus effectively reduces the time spent by the marking personnel in verifying the logical errors of the mind map, exposes the possible logical errors of the mind map earlier, reduces the possibility of introducing low-quality data during the training process of the large model due to the mistakes of the marking personnel, and improves the efficiency and accuracy of the large model training.

[0053] It should be noted that the execution subject of this method can be a data verification device, which can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a notebook computer, a handheld computer, an in-vehicle electronic device, a wearable device, a super mobile personal computer, etc., and the non-mobile electronic device can be a server, a network attached storage, a personal computer, etc. The present invention does not make specific limitations.

[0054] Figure 2 is one of the flow diagrams of the data verification method provided by the present invention; as Figure 2 shown, this method includes step 210, step 220, step 230 and step 240.

[0055] Step 210, obtain the mind map to be processed.

[0056] The mind map to be processed here is the mind map for which data verification is required. For example, it can be obtained by automatically screening the mind map according to the functional requirements of the mind map for training based on the large model, or it can be formed by the user's real-time input according to the data verification requirements, or it can be generated in real time according to the prompt text generated by the mind map. This embodiment does not make specific limitations on this.

[0057] The user input here can be information input through a command line interface, a graphical interface, touch input, drop-down selection input, voice input, gesture input, visual input, brain-computer input, etc. This embodiment does not make specific limitations on this.

[0058] Step 220: Perform slicing processing on the mind map to be processed, and according to the slicing result, obtain the slice features of each slice data to be processed in the mind map to be processed.

[0059] Optionally, after obtaining the mind map to be processed, slicing processing can be performed on the mind map to be processed to slice the mind map to be processed into multiple slice data to be processed, and based on the slicing result, generate the slice features of each slice data to be processed, so as to better analyze the data logical relationship in the mind map to be processed subsequently.

[0060] The slicing processing here can be implemented by slicing according to the data structure and / or data format of the mind map to be processed, etc., to ensure the integrity of the slice data to be processed obtained by slicing.

[0061] The slice features here can be formed by combining the slice information formed by slicing to form corresponding slice features; or, after feature extraction is performed on the slice information formed by slicing, corresponding slice features are formed, etc. This embodiment does not make specific limitations on this. The slice information includes but is not limited to the content information, position information, theme information, etc. of the slice data.

[0062] Step 230: Input the slice features of every two slice data to be processed into the classification model to obtain the logical relationship category between every two slice data to be processed; wherein, the classification model is trained based on multiple sample mind maps of different data types and the logical relationship category labels of the corresponding sample slice data pairs.

[0063] Optionally, after obtaining the slice features of multiple slices of data to be processed, the slice features of every two slices of data to be processed can be combined pairwise and input into the classification model respectively, so that the classification model can apply the slice features of every two slices of data to be processed to identify whether there is an unreasonable state in the logical relationship between every two slices of data to be processed. If there is an unreasonable state in the logical relationship between any two slices of data to be processed, the logical relationship category between these two slices of data to be processed is determined as an abnormal logical relationship. If there is no unreasonable state in the logical relationship between any two slices of data to be processed, the logical relationship category between these two slices of data to be processed is determined as a normal logical relationship, so as to quickly locate the errors in the mind map to be processed according to the logical relationship categories between different slices of data to be processed subsequently, and assist the annotator to modify or re-annotate quickly and accurately.

[0064] The classification model therein is a model for classifying the logical relationship between slice data, and it can be pre-trained based on the following steps:

[0065] First, collect small batches of sample mind maps with a data volume less than a set number and covering multiple different data types; the sample mind maps here can be dynamically generated by using a large language model according to the sample content data under the guidance of the mind map generation prompt text; or they can be obtained by pre-loading in the database, and this embodiment does not make specific limitations on this.

[0066] After obtaining the sample mind maps, the sample mind maps can be sliced to obtain multiple sample slice data, and every two sample slice data are combined into a pair of sample slice data. And label the logical relationship category between the two sample slice data in each pair of sample slice data, thereby obtaining the logical relationship category labels of each pair of sample slice data.

[0067] Immediately afterwards, use the sample data set constructed by each pair of sample slice data and the logical relationship category labels of each pair of sample slice data as the training data set to iteratively train the initialized model, so as to obtain a classification model that can automatically and accurately verify the logical relationship of mind map data.

[0068] Step 240, verify the mind map to be processed according to the logical relationship category between every two slices of data to be processed, and obtain the target mind map corresponding to the mind map to be processed.

[0069] Optionally, after obtaining the logical relationship categories between every two slices of data to be processed, the data annotation reference recommendation information corresponding to the mind map to be processed can be given according to the logical relationship categories between every two slices of data to be processed, so as to verify the mind map to be processed, and thus a high-quality target mind map can be obtained.

[0070] The verification here can be implemented by sending the data annotation reference recommendation information to the user side to receive the verification operation input by the user side according to the data annotation reference recommendation information, or can be automatically verified by a data verification device or other controllers according to the data annotation reference recommendation information. This embodiment does not make specific limitations on this.

[0071] The method provided in this embodiment extracts slice features by slicing the mind map to be processed; subsequently, using a classification model trained in advance based on the logical relationship category labels of multiple sample mind maps and their pairs of slice data, automatically determine the logical relationship category between every two slice data in the mind map to be processed, so as to verify the mind map to be processed according to the identified logical relationship category, thereby generating a high-quality target mind map. Thus, the automatic and accurate judgment of the logical relationship category of the mind map is realized, the efficiency and accuracy of data verification are significantly improved, the influence of human factors on the correctness of data verification is reduced, and the high-quality generation of the mind map is ensured.

[0072] In some embodiments, the classification model is trained based on the following steps:

[0073] According to two sample slice data with normal logical relationships in the sample mind map, construct the positive sample slice data pair corresponding to the sample mind map, and label the positive sample slice data pair as the normal logical relationship;

[0074] Preprocess the sample mind map;

[0075] According to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, construct the negative sample slice data pair corresponding to the sample mind map, and label the negative sample slice data pair as the abnormal logical relationship;

[0076] According to the slice features of the sample slice data in the negative sample slice data pair, the slice features of the sample slice data in the positive sample slice data pair, as well as the logical relationship category label of the positive sample slice data pair and the logical relationship category label of the negative sample slice data pair, perform iterative training on the initialized model to obtain the classification model;

[0077] Wherein, the preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols.

[0078] Optionally, for the logical relationship verification of the sliced data of the function mind map mainly implemented by the classification model, it is necessary to ensure a wide coverage of the data types of the sample mind maps in the training dataset for constructing the classification model, so as to avoid logical relationships that cannot be judged during the subsequent model usage. Therefore, when constructing the training dataset, it is necessary to include a sufficient number of positive sample sliced data pairs.

[0079] Among them, the normal logical relationships existing in the mind map include the relationship of juxtaposition or progression among the contents under the same branch, and generally the relationship of juxtaposition or irrelevance among the contents under different branches; while the abnormal logical relationships existing in the mind map include that the contents under the same branch are neither juxtaposed nor progressive, and / or the same content is introduced under different branches.

[0080] When using the sample mind maps generated by the large model, the logical relationships of most of the sample sliced data pairs are normal logical relationships. Therefore, two sample sliced data with normal logical relationships can be directly selected from the sample mind maps to construct positive sample sliced data pairs.

[0081] It should be noted that during the process of labeling the logical relationships of the positive sample sliced data pairs, there are generally labels of juxtaposition, progression or irrelevance. During the process of labeling the positive sample sliced data pairs, the accuracy of the labels of the positive sample sliced data pairs should be ensured. Even if there are certain errors in the labeling, since the positive sample sliced data pairs are marked with the labels of the positive sample sliced data pairs, the final classification will still be classified into the positive sample classification, so it will not cause serious impact on model training.

[0082] Since the proportion of negative sample sliced data pairs in the mind map is relatively small, it is difficult to quickly obtain a sufficient amount of negative sample sliced data pairs if directly screening negative sample sliced data pairs from the sample mind maps. Therefore, in order to conveniently obtain a sufficient amount of negative sample sliced data pairs, the sample mind maps can be preprocessed to add data with abnormal logical relationships to the sample mind maps. For example, some contents in the sample mind maps can be duplicated to increase the data with abnormal logical relationships of content duplication, or based on preset symbols, some sentences in the sample mind maps can be split to increase the data with abnormal logical relationships of one sentence being split into multiple sentences.

[0083] Among them, the duplication process can be randomly selecting the content data of some sliced data in the sample mind maps, using the large model to fine-tune and modify the language expression to duplicate some contents in the sample mind maps, and then calibrating them as negative sample sliced data pairs with content duplication on the sample mind maps; the sentence splitting process can be selecting the content data of a sentence from the sliced data of the sample mind maps and splitting it into multiple sentences according to punctuation marks, and using this as the negative sample sliced data pair of one sentence being split into multiple sentences in the sample mind maps.

[0084] It should be noted that the data annotation process for negative sample slice data pairs needs to ensure the accuracy of the labels, because during the use of the model, it is necessary to correctly point out the logical relationship errors between the slices of the mind map.

[0085] Exemplarily, the representation form of positive sample slice data pairs can be:

[0086] [{"input1": "Core figure: Pavel Korchagin", "input2": "Main plot", "count1": "h1: 1h2: 2", "count2": "h1: 1h2: 2h3: 1", "head": "How the Steel Was Tempered", "target": "Progressive"}, {"input1": "Author's personal background", "input2": "Tonia", "count1": "h1: 1h2: 1", "count2": "h1: 1h2: 3h3: 1", "head": "How the Steel Was Tempered", "target": "Irrelevant"}, {"input1": "Pavel's growth experience", "input2": "Joining the army and fighting", "count1": "h1: 1h2: 1h3: 1ul: 1", "count2": "h1: 1h2: 1h3: 1ul: 2", "head": "How the Steel Was Tempered", "target": "Parallel"}].

[0087] The representation form of negative sample slice data pairs can be:

[0088] [{"input1": "Technological development","input2": "Issuance of L3 autonomous driving test licenses","count1": "h1:1h2:1h3:2","count2": "h1:1h2:1h3:2ul:1","head": "Application of AI in automobiles","target": "Not relevant"},{"input1": "Verdun meat grinder","input2": "The battle was extremely fierce","count1": "h1:1h2:1h3:2lu:1","count2": "h1:1h2:1h3:2lu:2","head": "—Main battle, battle","target":"Forced split"},{"input1":"Dabai Electric Appliances' revenue in the first quarter increased by 10% year-on-year","input2":"Dabai Electric Appliances' revenue grew rapidly","count1": "h1:1h2:1h3:1ul:1","count2": "h1:1h2:3h3:2ul:1","head": "Financial Analysis of the Group","target": "Duplicated"}].

[0089] Among them, "input1" and "input2" are the content information of two slice data; "count1" and "count2" are the hierarchical structures of two slice data; "head" is the subject information; "target" is the logical relationship information;

[0090] After obtaining the negative sample slice data pairs and the positive sample slice data pairs, the slice features of the sample slice data in the negative sample slice data pairs and the slice features of the sample slice data in the positive sample slice data pairs can be used as samples, and the logical relationship category labels of the positive sample slice data pairs and the logical relationship category labels of the negative sample slice data pairs can be used as labels to iteratively train the initialization model to obtain a classification model that can automatically and accurately identify normal logical relationship categories and abnormal logical relationship categories between different slice data pairs.

[0091] The method provided in this embodiment constructs positive and negative sample slice data pairs based on sample mind maps, and combines preprocessing technology to increase data of abnormal logical relationships, and trains a classification model that can accurately identify normal and abnormal logical relationships between slice data pairs in mind maps, thereby effectively improving the automation and accuracy of logical relationship verification and ensuring the coverage and robustness of the model on a wide range of data types.

[0092] In some embodiments, step 220 specifically includes:

[0093] The mind map to be processed is sliced according to the parsing tool corresponding to the data format of the mind map to be processed.

[0094] Optionally, when performing slicing processing, a parsing tool that matches the data format of the mind map to be processed can be used, and appropriate tags are used to label each branch data in the mind map to be processed, so as to realize the slicing processing of the mind map to be processed.

[0095] For example, if the data format of the mind map to be processed is data in the format of lightweight markup language (abbreviated as markdown, or MD), the parsing library of markdown can be used to perform slicing processing on the mind map to be processed.

[0096] The slicing processing logic of the parsing library of markdown can be specifically implemented through the following code:

[0097] "# Convert Markdown string to HTML

[0098] html = markdown.markdown(markdown_string)

[0099] # Use BeautifulSoup to parse HTML

[0100] soup = BeautifulSoup(html, 'html.parser')

[0101] def extract_structure(soup):

[0102] structure = []

[0103] for element in soup.children:

[0104] if element.name:

[0105] structure.append({

[0106] 'tag': element.name,

[0107] 'text': element.get_text(strip=True),

[0108] 'children': extract_structure(element) if element.contents else []

[0109] })

[0110] return structure

[0111] ”.

[0112] Exemplarily, through a Markdown parsing library, the mind map to be processed can be split into the following format:

[0114] {"text": "Mind Map of Peng Jixiang's Introduction to Art, Fourth Edition", "count": "h1:1"},

[0115] {"text": "General Introduction to Art", "count": "h1:1h2:1"},

[0116] {"text": "The Essence and Characteristics of Art", "count": "h1:1h2:1h3:1"},

[0117] {"text": "The Essence of Art", "count": "h1:1h2:1h3:1ul:1"},

[0118] {"text": "The Characteristics of Art", "count": "h1:1h2:1h3:1ul:2"},

[0119] {"text": "The Origin of Art", "count": "h1:1h2:1h3:2"},

[0120] {"text": "Five Views", "count": "h1:1h2:1h3:2ul:1"},

[0121] {"text": "Plural Determinism", "count": "h1:1h2:1h3:2ul:2"},

[0122] {"text": "The Functions of Art and Art Education", "count": "h1:1h2:1h3:3"},

[0123] {"text": "Social Functions", "count": "h1:1h2:1h3:3ul:1"},

[0124] {"text": "Art Education", "count": "h1:1h2:1h3:3ul:2"},

[0125] {"text": "Art in the Cultural System", "count": "h1:1h2:1h3:4"},

[0126] ​{"text": "Cultural Phenomena", "count": "h1:1h2:1h3:4ul:1"},

[0127] {"text": "Art and Philosophy", "count": "h1:1h2:1h3:4ul:2"},

[0128] {"text": "Art and Religion", "count": "h1:1h2:1h3:4ul:3"},

[0129] {"text": "Art and Morality", "count": "h1:1h2:1h3:4ul:4"},

[0130] {"text": "Art and Science", "count": "h1:1h2:1h3:4ul:5"},

[0131] {"text": "Types of Art", "count": "h1:1h2:2"},

[0132] {"text": "Applied Art", "count": "h1:1h2:2h3:1"},

[0133] {"text": "Main Types", "count": "h1:1h2:2h3:1ul:1"},

[0134] {"text": "Aesthetic Features", "count": "h1:1h2:2h3:1ul:2"},

[0135] {"text": "Appreciation of Fine Works", "count": "h1:1h2:2h3:1ul:3"},

[0136] {"text": "Plastic Art", "count": "h1:1h2:2h3:2"},

[0137] {"text": "Main Types", "count": "h1:1h2:2h3:2ul:1"},

[0138] {"text": "Aesthetic Features", "count": "h1:1h2:2h3:2ul:2"},

[0139] {"text": "Appreciation of Fine Works", "count": "h1:1h2:2h3:2ul:3"},

[0140] {"text": "Expressive Arts", "count": "h1:1h2:2h3:3"},

[0141] {"text": "Main Types", "count": "h1:1h2:2h3:3ul:1"},

[0142] {"text": "Aesthetic Features", "count": "h1:1h2:2h3:3ul:2"},

[0143] {"text": "Appreciation of Fine Works", "count": "h1:1h2:2h3:3ul:3"},

[0144] {"text": "Comprehensive Arts", "count": "h1:1h2:2h3:4"},

[0145] {"text": "Main Types", "count": "h1:1h2:2h3:4ul:1"},

[0146] {"text": "Aesthetic Features", "count": "h1:1h2:2h3:4ul:2"},

[0147] {"text": "Appreciation of Fine Works", "count": "h1:1h2:2h3:4ul:3"},

[0148] {"text": "Linguistic Arts", "count": "h1:1h2:2h3:5"},

[0149] {"text": "Main Genres", "count": "h1:1h2:2h3:5ul:1"}

[0150] .

[0151] Among them, text is used to describe the content information related to each slice of data to be processed. count is used to describe the hierarchical relationship of each slice of data to be processed in the mind map.

[0152] The method provided in this embodiment can ensure the integrity and accuracy of information by using a parsing tool that matches the data format of the mind map to be processed for slicing and labeling each branch of data with appropriate tags, so that the data logic relationship in the mind map can be accurately parsed and verified based on the sliced data.

[0153] In some embodiments, step 220 further includes:

[0154] According to the segmentation result, obtain the content information and position information of each piece of data of the slice to be processed;

[0155] According to the content information, position information, and the theme information of the mind map to be processed, combine to obtain the slice features of each piece of data of the slice to be processed.

[0156] Optionally, since the construction of the mind map slice data will seriously affect the final effect of the classification model. For example, how to determine whether these two slices belong to the same branch, and how to determine whether the same keyword appears in the content of the two slices is a normal logical relationship or a duplication in content. Therefore, if only the content information of two pieces of data of the slice to be processed is simply constructed into a string and input into the classification model for the logical relationship category during the construction of the mind map slice, it will lead to a large classification error.

[0157] In response to this, the method provided in this embodiment constructs the slice features of each piece of data of the slice to be processed by combining the content information, position information of each piece of data of the slice to be processed, and the theme information of the mind map to be processed, and combines the slice features of two pieces of data of the slice to be processed as the input of the classification model, so as to help the classification model more accurately determine whether two pieces of data of the slice to be processed belong to the same branch, and whether the logical relationship between the two is normal or a duplication in content. Furthermore, it can reduce the classification error, improve the understanding and judgment ability of the classification model for the logical relationship of the mind map data, and significantly improve the accuracy and robustness of the classification of the classification model. The specific implementation steps are as follows:

[0158] When obtaining the segmentation features, it can be to parse the segmentation result to obtain the content information and position information of each piece of data of the slice to be processed.

[0159] The step of obtaining the position information here includes: according to the set order corresponding to the slice label marked for each piece of data of the slice to be processed during the slicing process, the position of each piece of data of the slice to be processed in the mind map to be processed can be determined, and thus the corresponding position information can be obtained.

[0160] The slice label here can be a string used to mark the hierarchical relationship and position of the content information of different slice data in the mind map, such as in the format of "h1, h2, h3, ul", etc. This embodiment does not make specific limitations on this.

[0161] Since there are cases in the mind map verification where some repeated strings can be judged as content duplicates, while some repeated strings cannot be judged as content duplicates. For example, the words appearing in the topic information of the mind map are generally regarded as keywords. If the keywords appear repeatedly, they will not be judged as content duplicates, while the repetition of other identical content will be judged as content duplicates. Therefore, when determining the slice features input to the classification model, not only content information and location information need to be introduced, but also the topic information of the mind map to be processed needs to be introduced into the slice features to avoid misjudgment of the classification model caused by the repeated appearance of the topic, thereby improving the accuracy of data verification.

[0162] The method provided in this embodiment constructs slice features by combining content information, location information, and topic information, and uses them as the input of the classification model to effectively distinguish the normal logical relationship and content duplication within the same branch, reduce classification errors, enhance the classification model's understanding and judgment ability of the mind map structure, and thus significantly improve the accuracy and robustness of the mind map data logical relationship verification.

[0163] In some embodiments, step 210 specifically includes:

[0164] Input the preset mind map generation prompt text and each data to be processed into the large language model to obtain the initial mind map corresponding to each data to be processed;

[0165] According to the preset mind map function requirements, screen out the mind map to be processed from multiple initial mind maps.

[0166] Optionally, the mind map to be processed can be automatically screened out through the following steps:

[0167] Collect the preset mind map generation prompt text (abbreviated as prompt) and each data to be processed; input the preset mind map generation prompt text and each data to be processed into the large language model, so that the large language model, under the guidance of the preset mind map generation prompt text, uses the method of instruction call to initially generate the initial mind map corresponding to each data to be processed;

[0168] Furthermore, in order to generate high-quality sample mind maps to participate in the fine-tuning training of the large language model, so that the trained large language model has the preset mind map function requirements, the initial mind maps that match the preset mind map function requirements can be screened out from multiple initial mind maps as the mind maps to be processed.

[0169] The preset mind map generation prompt text here is used to guide a large language model to generate a mind map, which can be generated based on questions of a large number of mind maps collected from the Internet, etc.; the preset mind map generation prompt text includes but is not limited to the data type, data logic, architecture, format, etc. of the mind map to be generated, and this embodiment does not make specific limitations on this.

[0170] The data to be processed here can be information or knowledge points for mind map generation, including but not limited to text information, data tables, keywords and phrases, etc. This embodiment does not make specific limitations on this.

[0171] Among them, the large language model (LLM) is abbreviated as the large model or large language model, which refers to a natural language processing (NLP) model with a huge number of parameters. The number of model parameters and / or the complexity of the model structure exceed a preset threshold. This model is pre-trained with a large amount of text data and has a high semantic understanding ability and the ability to generate natural language.

[0172] The method provided in this embodiment combines a large language model with a preset mind map generation prompt text and data to be processed, and automatically screens and generates a mind map that meets the functional requirements, so as to effectively improve the efficiency and quality of mind map generation, so that a large language model with the functional requirements of the preset mind map can be trained more quickly subsequently.

[0173] In some embodiments, the method further includes:

[0174] Obtain the target mind maps corresponding to the multiple data to be processed mind maps;

[0175] Based on the target mind maps corresponding to the multiple data to be processed mind maps, the data to be processed, and the preset mind map generation prompt text, perform iterative training on the large language model to obtain a mind map generation model corresponding to the preset mind map functional requirements. The mind map generation model is used to generate a mind map that meets the preset mind map functional requirements.

[0176] Optionally, a large number of data to be processed mind maps with a data volume exceeding a preset number can be collected, and according to the Figure 2 data verification process shown, perform corresponding data verification operations on the collected large number of data to be processed mind maps to obtain the target mind maps corresponding to each data to be processed mind map, thereby obtaining sample data with a wide coverage of data types, a large amount of data, and high quality for assisting the training of the mind map generation model.

[0177] Subsequently, using the data to be processed corresponding to each mind map to be processed and the preset mind map generation prompt text as samples, and using the target mind map corresponding to each mind map to be processed as labels, iteratively train a large language model to quickly obtain a mind map generation model. This mind map generation model can automatically and accurately generate mind maps that meet the functional requirements of the preset mind map, thereby improving the automation and accuracy of mind map generation and further optimizing the user experience.

[0178] In some embodiments, step 240 specifically includes:

[0179] Obtain a pair of abnormal slice data formed by two slices of data to be processed with an abnormal logical relationship category;

[0180] Output annotation prompt information for the mind map to be processed according to the slice features of the slices of data to be processed in each pair of abnormal slice data;

[0181] Receive the verification operation corresponding to the mind map to be processed; the verification operation is an operation generated by the verification end according to the annotation prompt information and is used for mind map verification processing;

[0182] Verify the mind map to be processed according to the verification operation to obtain the target mind map.

[0183] Figure 3 It is the second flowchart of the data verification method provided by the present invention.

[0184] As Figure 3 shown, after training a classification model through sample mind maps of multiple different data types and the logical relationship category labels of their corresponding sample slice data pairs, and using the classification model to identify the logical relationships of the slices of data to be processed corresponding to the functional requirements of the mind map generation model, after obtaining the logical relationship category between every two slices of data to be processed in the mind map to be processed, it is possible to judge one by one the logical relationship category between every two slices of data to be processed to obtain a pair of abnormal slice data formed by two slices of data to be processed with an abnormal logical relationship category, and record the slice features of each pair of abnormal slice data. And according to the slice features of the slices of data to be processed in each pair of abnormal slice data, generate annotation prompt information for the mind map to be processed to prompt the annotator to locate the abnormal logical relationship data in the mind map to be processed according to the annotation prompt information, combine the retrieval tool to retrieve the data with incorrect data content, and input the verification operation for verifying the mind map to be processed at the verification end based on the abnormal logical relationship data and the data with incorrect data content.

[0185] After receiving the verification operation corresponding to the mind map to be processed, it is possible to mark or modify the mind map to be processed according to the verification operation, thereby obtaining a high-quality target mind map.

[0186] The method provided in this embodiment realizes the data verification of the mind map through identifying abnormal logical relationship slice data pairs and generating annotation prompt information for human-computer interaction, effectively improving the efficiency and accuracy of the data verification of the mind map, so as to quickly and accurately generate a high-quality target mind map.

[0187] The data verification device provided by the present invention will be described below. The data verification device described below can be mutually corresponding and referred to the data verification method described above.

[0188] Figure 4 It is a schematic structural diagram of the data verification device provided by the present invention. As Figure 4 shown, the device includes:

[0189] An obtaining unit 410 is configured to obtain a mind map to be processed;

[0190] A slicing unit 420 is configured to perform slicing processing on the mind map to be processed, and according to the slicing result, obtain the slice features of each to-be-processed slice data in the mind map to be processed;

[0191] A classification unit 430 is configured to input the slice features of every two to-be-processed slice data into a classification model to obtain the logical relationship category between every two to-be-processed slice data;

[0192] A verification unit 440 is configured to verify the mind map to be processed according to the logical relationship category between every two to-be-processed slice data, and obtain a target mind map corresponding to the mind map to be processed;

[0193] Wherein, the classification model is trained based on a plurality of sample mind maps of different data types and the logical relationship category labels of the corresponding sample slice data pairs.

[0194] The device provided in this embodiment extracts slice features by performing slicing processing on the mind map to be processed; subsequently, using a classification model pre-trained based on the logical relationship category labels of a variety of sample mind maps and their slice data pairs, automatically judge the logical relationship category between every two slice data in the mind map to be processed, and verify the mind map to be processed according to the identified logical relationship category, so as to generate a high-quality target mind map, thereby realizing the automatic and accurate judgment of the logical relationship category of the mind map, significantly improving the efficiency and accuracy of data verification, reducing the influence of human factors on the correctness of data verification, and ensuring the high-quality generation of the mind map.

[0195] In some embodiments, the device further includes a training unit configured to:

[0196] Construct a positive sample slice data pair corresponding to the sample mind map according to two sample slice data with normal logical relationships in the sample mind map, and label the positive sample slice data pair as having a normal logical relationship;

[0197] Preprocess the sample mind map;

[0198] Construct a negative sample slice data pair corresponding to the sample mind map according to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, and label the negative sample slice data pair as having an abnormal logical relationship;

[0199] Iteratively train the initialized model according to the slice features of the sample slice data in the negative sample slice data pair, the slice features of the sample slice data in the positive sample slice data pair, the logical relationship class label of the positive sample slice data pair, and the logical relationship class label of the negative sample slice data pair to obtain the classification model;

[0200] Wherein, the preprocessing includes content duplication processing and / or sentence splitting processing based on a preset symbol.

[0201] In some embodiments, the slicing unit is specifically configured to:

[0202] Obtain the content information and position information of each of the to-be-processed slice data according to the segmentation result;

[0203] Combine to obtain the slice features of each of the to-be-processed slice data according to the content information, position information, and the theme information of the to-be-processed mind map.

[0204] In some embodiments, the slicing unit is further configured to:

[0205] Slice the to-be-processed mind map according to the parsing tool corresponding to the data format of the to-be-processed mind map.

[0206] In some embodiments, the obtaining unit is specifically configured to:

[0207] Input a preset mind map generation prompt text and each to-be-processed data into a large language model to obtain an initialized mind map corresponding to each to-be-processed data;

[0208] Screen and obtain the to-be-processed mind map from multiple initialized mind maps according to the preset mind map function requirements.

[0209] In some embodiments, the training unit is further configured to:

[0210] Obtain target mind maps corresponding to multiple of the to-be-processed mind maps;

[0211] Based on the target mind maps corresponding to multiple of the to-be-processed mind maps and the to-be-processed data, and the preset mind map, generate prompt text, and perform iterative training on the large language model to obtain a mind map generation model corresponding to the preset mind map functional requirements, where the mind map generation model is used to generate a mind map that meets the preset mind map functional requirements.

[0212] In some embodiments, the verification unit is specifically configured to:

[0213] Obtain an abnormal slice data pair formed by two to-be-processed slice data with a logical relationship category of abnormal logical relationship;

[0214] According to the slice features of the to-be-processed slice data in each abnormal slice data pair, output annotation prompt information for the to-be-processed mind map;

[0215] Receive a verification operation corresponding to the to-be-processed mind map; the verification operation is an operation generated by a verification end according to the annotation prompt information and used for mind map verification processing;

[0216] According to the verification operation, verify the to-be-processed mind map to obtain the target mind map.

[0217] The device provided by the present invention is used to execute the above method embodiments. For the specific process and detailed content, please refer to the above embodiments and will not be elaborated here.

[0218] Figure 5 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute a data verification method, which includes: obtaining a mind map to be processed; performing slicing processing on the mind map to be processed, and according to the slicing result, obtaining the slicing features of each slice data to be processed in the mind map to be processed; inputting the slicing features of every two slice data to be processed into a classification model to obtain the logical relationship category between every two slice data to be processed; verifying the mind map to be processed according to the logical relationship category between every two slice data to be processed to obtain the target mind map corresponding to the mind map to be processed; wherein, the classification model is trained based on a plurality of sample mind maps of different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0219] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0220] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data verification method provided by each of the above methods. The method includes: obtaining a mind map to be processed; performing slicing processing on the mind map to be processed, and according to the slicing result, obtaining the slicing features of each slice data to be processed in the mind map to be processed; inputting the slicing features of every two slice data to be processed into a classification model to obtain the logical relationship category between every two slice data to be processed; verifying the mind map to be processed according to the logical relationship category between every two slice data to be processed to obtain a target mind map corresponding to the mind map to be processed; wherein, the classification model is trained based on a plurality of sample mind maps of different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0221] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the data verification method provided by each of the above methods. The method includes: obtaining a mind map to be processed; performing slicing processing on the mind map to be processed, and according to the slicing result, obtaining the slicing features of each slice data to be processed in the mind map to be processed; inputting the slicing features of every two slice data to be processed into a classification model to obtain the logical relationship category between every two slice data to be processed; verifying the mind map to be processed according to the logical relationship category between every two slice data to be processed to obtain a target mind map corresponding to the mind map to be processed; wherein, the classification model is trained based on a plurality of sample mind maps of different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps.

[0222] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data verification method, characterized in that, Including: Obtain the mind map to be processed; Perform slicing processing on the mind map to be processed, and according to the slicing result, obtain the slicing features of each slice data to be processed in the mind map to be processed; Input the slicing features of every two slice data to be processed into a classification model to obtain the logical relationship category between every two slice data to be processed; According to the logical relationship category between every two slice data to be processed, verify the mind map to be processed to obtain the target mind map corresponding to the mind map to be processed; Wherein, the classification model is trained based on sample mind maps of multiple different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps; The classification model is trained based on the following steps: According to two sample slice data with normal logical relationships in the sample mind map, construct a positive sample slice data pair corresponding to the sample mind map, and label the positive sample slice data pair as a normal logical relationship; Perform preprocessing on the sample mind map; According to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, construct a negative sample slice data pair corresponding to the sample mind map, and label the negative sample slice data pair as an abnormal logical relationship; According to the slicing features of the sample slice data in the negative sample slice data pair, the slicing features of the sample slice data in the positive sample slice data pair, and the logical relationship category labels of the positive sample slice data pair and the logical relationship category labels of the negative sample slice data pair, perform iterative training on the initialized model to obtain the classification model; Wherein, the preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols; the duplication processing is to randomly select the content data of some slice data in the sample mind map, and use a large model to fine-tune and modify the language expression to perform duplication processing on some content in the sample mind map; the sentence splitting processing selects the content data of a sentence from the slice data of the sample mind map and splits it into multiple sentences according to punctuation marks.

2. The data verification method according to claim 1, wherein The obtaining the slicing features of each slice data to be processed in the mind map to be processed according to the slicing result includes: According to the slicing result, obtain the content information and position information of each slice data to be processed; According to the content information, position information, and the theme information of the mind map to be processed, combine to obtain the slicing features of each slice data to be processed.

3. The data verification method according to any one of claims 1-2, characterized in that, The performing slicing processing on the mind map to be processed includes: Perform slicing processing on the mind map to be processed according to the parsing tool corresponding to the data format of the mind map to be processed.

4. The data verification method according to any one of claims 1-2, characterized in that, The obtaining the mind map to be processed includes: Input the preset mind map generation prompt text and each data to be processed into a large language model to obtain the initialized mind maps corresponding to each data to be processed; According to the preset mind map function requirements, screen and obtain the mind map to be processed from multiple initialized mind maps.

5. The data verification method according to claim 4, wherein The method further includes: Obtain the target mind maps corresponding to multiple to-be-processed mind maps; Based on the target mind maps corresponding to multiple to-be-processed mind maps, the to-be-processed data, and the preset mind map, generate prompt texts to iteratively train the large language model to obtain a mind map generation model corresponding to the preset mind map functional requirements, where the mind map generation model is used to generate mind maps that meet the preset mind map functional requirements.

6. The data verification method according to any one of claims 1-2, characterized in that The step of obtaining the target mind map corresponding to the to-be-processed mind map by verifying the to-be-processed mind map according to the logical relationship category between every two pieces of to-be-processed slice data includes: Obtain abnormal slice data pairs formed by two pieces of to-be-processed slice data with an abnormal logical relationship category; Output the annotation prompt information of the to-be-processed mind map according to the slice features of the to-be-processed slice data in each abnormal slice data pair; Receive the verification operation corresponding to the to-be-processed mind map; the verification operation is generated by the verification end according to the annotation prompt information and is used for mind map verification processing; Verify the to-be-processed mind map according to the verification operation to obtain the target mind map.

7. A data verification device, characterized in that, It includes: An acquisition unit for acquiring to-be-processed mind maps; A slicing unit for slicing the to-be-processed mind map and obtaining the slice features of each piece of to-be-processed slice data in the to-be-processed mind map according to the slicing result; A classification unit for inputting the slice features of every two pieces of to-be-processed slice data into a classification model to obtain the logical relationship category between every two pieces of to-be-processed slice data; A verification unit for verifying the to-be-processed mind map according to the logical relationship category between every two pieces of to-be-processed slice data to obtain the target mind map corresponding to the to-be-processed mind map; Among them, the classification model is trained based on sample mind maps of multiple different data types and the logical relationship category labels of the sample slice data pairs corresponding to the sample mind maps; the classification model is trained according to the following steps: construct a positive sample slice data pair corresponding to the sample mind map according to two sample slice data with normal logical relationships in the sample mind map, and label the positive sample slice data pair as a normal logical relationship; preprocess the sample mind map; construct a negative sample slice data pair corresponding to the sample mind map according to two sample slice data with abnormal logical relationships in the preprocessed sample mind map, and label the negative sample slice data pair as an abnormal logical relationship; perform iterative training on the initialized model according to the slice features of the sample slice data in the negative sample slice data pair, the slice features of the sample slice data in the positive sample slice data pair, the logical relationship category label of the positive sample slice data pair, and the logical relationship category label of the negative sample slice data pair to obtain the classification model; among them, the preprocessing includes content duplication processing and / or sentence splitting processing based on preset symbols; the duplication processing is to randomly select the content data of some slice data in the sample mind map, and use the large model fine-tuning to modify the language expression to perform duplication processing on some content in the sample mind map; the sentence splitting processing selects the content data of a sentence from the slice data of the sample mind map and splits it into multiple sentences according to punctuation marks.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data verification method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data verification method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent verification method and system for electric power work ticket content

    CN119202680A

  • AI-driven multilingual intelligent document abstract and mind map generation system

    CN119494394A