Big model-based multi-modal digital publishing intelligent proofreading system and method
The intelligent proofreading system for multimodal digital publishing based on a large model solves the structural problems in traditional proofreading techniques, achieves accurate semantic recognition and structural judgment, and improves proofreading efficiency and content integration capabilities.
Patent Information
- Application Number
- CN202511449327.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Traditional multimodal digital publishing intelligent proofreading technology lacks a hierarchical subdivision mechanism, making it difficult to deeply analyze structural problems such as the separation of subject and object semantics, misalignment of paragraph themes, and fluctuations in terminology collocation. This results in chaotic structural tags in the proofreading output, reducing efficiency and content integration capabilities.
The multimodal digital publishing intelligent proofreading system based on a large model identifies subject-verb-object combinations through a sample construction module, identifies image action paths and text subject-verb combinations through a multimodal analysis module, analyzes paragraph semantic connection features through a structure recognition module, and identifies term collocation offsets through a term verification module, generating accurate annotation structure merging records.
It improves the accuracy of semantic recognition and structural judgment, reveals issues such as paragraph content misalignment and topic jumps, identifies changes in terminology collocation, generates tag merging suggestions, and improves the efficiency of the review results and the ability to integrate content.
Smart Images

Figure CN120930637B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent document review technology, and in particular to a multimodal digital publishing intelligent review system and method based on a large model. Background Technology
[0002] The field of intelligent document review technology encompasses the automated analysis, review, and correction of digital document content. It is particularly relevant for reviewing multimodal content such as text, images, and videos in publishing, office, and media settings. The core of this technology lies in using language understanding, visual recognition, and format parsing to automatically identify and structure errors in syntax, formatting inconsistencies, and text-image discrepancies. The results are then fed back to the original document through an integrated interface. Intelligent document review has gradually expanded from single-text processing to comprehensive multimodal content review, forming a system technology framework that includes content feature extraction, multi-dimensional error detection, standard rule matching, and semantic consistency analysis, emphasizing the precision of algorithm processing. The system integrates with office scenarios. Among them, the multimodal digital publishing intelligent review system based on a large model refers to a technical solution that unifies the input and preprocessing of publishing content in different formats such as text, images, and videos, calls on the semantic and visual recognition capabilities of a batch parameter pre-trained model to complete the identification and review of content errors, and outputs the results back to office software such as PPT or Word in the form of annotations or direct modifications. Specifically, it involves multiple processing steps such as text segmentation and structure parsing, image feature extraction, video keyframe extraction, model-based semantic logic consistency analysis, correlation detection between image, video and text content, error type labeling, review result format generation, automatic content injection, and user rule customization configuration.
[0003] Traditional multimodal digital publishing intelligent proofreading technology relies on language understanding and visual recognition to proofread documents. However, it lacks a hierarchical subdivision mechanism due to its unified processing flow. This makes it difficult to deeply analyze structural issues such as subject-object semantic separation, paragraph topic misalignment, and terminology collocation fluctuations. In complex sentence structures or cross-paragraph semantic jump scenarios, there are problems with overlapping error annotations and insufficient stripping of redundant tags. As a result, the structural tags in the proofreading output are difficult to form complete and mergeable annotation units. In the actual document annotation process, problems such as duplicate multiple tags, missing error labels, and chaotic annotation structures are prone to occur, reducing the efficiency of the proofreading results and the ability to integrate content. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a multimodal digital publishing intelligent review system and method based on a large model.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a multimodal digital publishing intelligent proofreading system based on a large model includes:
[0006] The sample construction module is based on the training sample document set. It identifies the subject-verb-object combination relationship of continuous sentence groups in the sample text, analyzes word order stability and sentence construction path, combines part-of-speech changes and semantic structure features, filters training samples, trains and obtains the intelligent review model, and generates the model configuration dataset.
[0007] The multimodal analysis module calls the model configuration dataset, analyzes the input document, identifies the behavioral objects of image action paths and text subject verb combinations, compares the consistency of action-carrying role expression in the text and image, identifies subject-object semantic separation paragraphs, and obtains multimodal semantic offset information;
[0008] Based on the multimodal semantic offset information, the structure recognition module analyzes the semantic connection features of paragraphs and the coverage of the preceding and following themes, compares the repetition of topic word chains, determines the structural continuity relationship, identifies semantic breaks, topic jumps and content offset positions, and obtains paragraph structure misalignment identifiers.
[0009] The terminology verification module analyzes the input document based on the paragraph structure misalignment identifier, identifies the occurrence position of each term in multiple paragraphs, analyzes the contextual collocation differences of terms in multiple paragraphs, judges the changes in part-of-speech continuation and collocation expression fluctuations, and obtains term collocation offset tags.
[0010] As a further aspect of the present invention, the model configuration dataset includes a semantic span label set, a perturbation structure group sequence, and a training sample syntax set. The multimodal semantic offset information includes a text-image role pointing reference table, a behavior expression conflict sequence, and an action semantic offset identifier. The paragraph structure misalignment identifier specifically includes a paragraph continuation anomaly group, a title chain affiliation offset set, and a hierarchical jump positioning index. The term collocation offset label specifically includes a context combination variation record, a collocation consistency distribution map, and a term relocation index sequence.
[0011] As a further aspect of the present invention, the sample construction module includes:
[0012] The structural perturbation identification submodule acquires the training sample document set, identifies the subject-verb-object combination of continuous sentence groups, analyzes the changes in word order and sentence structure mapping, determines the combination frequency of the main components under the perturbation, extracts the main component intersection area and combination conflict area, calculates the combination perturbation amplitude, filters the main structural perturbation significant area, and generates the sentence group main perturbation scalar set.
[0013] The semantic span annotation submodule calls the active interference dynamic scalar set of the sentence group, analyzes the dependency connection trajectory between keywords and the dynamic structural phrases within the sentence group, records the index interval of the keyword covered phrases, annotates the semantic span range in the sentence group, and obtains the semantic span interval data of keywords.
[0014] The training data generation submodule filters training samples based on the keyword semantic span range data, trains and obtains the intelligent review model, and generates the model configuration dataset.
[0015] As a further aspect of the present invention, the multi-mode analysis module includes:
[0016] The action path recognition submodule obtains the model configuration dataset, analyzes the action execution path of each image in the input document and the combination of text subject verbs, establishes the pairing relationship between image actions and text behavior objects, and generates an image-text action pairing matrix.
[0017] The role consistency judgment submodule compares the expression of the action-carrying role in the subject position in each group of image and text data according to the image and text action pairing matrix, analyzes the correspondence between image action entities and text behavior subjects, calculates subject-object semantic consistency offset value, and generates subject-object semantic offset coefficient sequence.
[0018] The semantic offset annotation submodule calls the subject-object semantic offset coefficient sequence to determine the consistency of subject-object semantics, filters semantically separated paragraphs, marks the degree of semantic difference between text and image, and obtains multimodal semantic offset information.
[0019] As a further aspect of the present invention, the structure recognition module includes:
[0020] The word grouping and connection recognition submodule obtains the multimodal semantic offset information, identifies the semantic word grouping and connection features of the subject-object semantic separation paragraph, analyzes the main theme coverage and logical connection expression of the preceding and following paragraphs, and generates paragraph word grouping and connection data;
[0021] The structural chain attribution judgment submodule compares the attribution repetition of the topic word chain formed by the title keywords in the structural hierarchy based on the paragraph word connection data, judges the continuity of the paragraph theme in the chapter structure, analyzes the topic chain structural jump and content offset, calculates the chapter structure offset intensity, filters the structural jump interval, and generates a topic chain offset coefficient sequence.
[0022] The hierarchical misalignment positioning submodule calls the topic chain offset coefficient sequence to determine the continuity of paragraph topics in the chapter structure and the coherence of paragraph relationships, identify semantic breaks, topic jumps and content offset positions, and obtain paragraph structure misalignment identifiers.
[0023] As a further aspect of the present invention, the terminology verification module includes:
[0024] The terminology distribution recognition submodule obtains the paragraph structure misalignment identifier, detects the distribution position of each term in multiple paragraphs in the input document, filters the occurrence frequency of each term at multiple levels, establishes terminology distribution statistics, and generates terminology paragraph distribution data.
[0025] The contextual collocation analysis submodule analyzes the frequency distribution of verb and noun combinations of terms in the context before and after each paragraph based on the term paragraph distribution, detects the change trajectory of part-of-speech continuation patterns, calculates the cumulative number of word substitutions in the collocation structure, and obtains the term collocation fluctuation coefficient.
[0026] The collocation consistency judgment submodule calls the term collocation fluctuation coefficient to compare the collocation stability of each term in multiple paragraphs, judges the collocation consistency of the term in multiple paragraph contexts, and obtains the term collocation offset label.
[0027] As a further aspect of the present invention, the system further includes:
[0028] The annotation output module calls the terminology pairing offset label and paragraph structure misalignment identifier, extracts the position information of each error label, filters multiple types of error labels that are adjacent in the same sentence group, analyzes the overlapping position of the word groups associated with the error labels, judges the repetition of multiple errors in the word coverage level, adjusts the annotation structure and merges the suggested content in the label set, and generates a record of annotation structure merging.
[0029] The annotation structure merging record includes an annotation merging path table, a tag association matrix, and an annotation position compression mapping set.
[0030] As a further aspect of the present invention, the annotation output module includes:
[0031] The error tag extraction submodule calls the term pairing offset tag and paragraph structure misalignment identifier to extract the position information of each error tag, filters multiple types of error tags that are adjacent in the same sentence group, and obtains the error tag position sequence.
[0032] The overlap relationship analysis submodule analyzes the overlapping positions of related word groups of error tags within the same sentence group based on the error tag position sequence, determines the repetition of error tags in the word coverage level, and obtains the error overlap relationship matrix.
[0033] The tag merging output submodule calls the error overlap matrix, adjusts the annotation structure, merges the suggested content in the tag set, and generates a record of the merged annotation structure.
[0034] A multimodal digital publishing intelligent review method based on a large model, wherein the method is executed based on the aforementioned multimodal digital publishing intelligent review system based on a large model, includes the following steps:
[0035] S1: Based on the training sample document set, identify the subject-verb-object combination relationship of continuous sentence groups in the sample text, analyze the word order stability and sentence construction path, combine part-of-speech changes and semantic structure features, filter training samples, train and obtain the intelligent review model, and generate the model configuration dataset.
[0036] S2: Call the model to configure the dataset, analyze the input document, identify the behavior object of the combination of image action path and text subject verb, compare the consistency of the expression of the role carried by the action in the image and text, identify the subject-object semantic separation paragraph, and obtain multimodal semantic offset information;
[0037] S3: Based on the multimodal semantic offset information, analyze the semantic connection features of paragraphs and the coverage of the main ideas before and after, compare the repetition of the topic word chain, determine the structural continuity relationship, identify the semantic break, topic jump and content offset position, and obtain the paragraph structure misalignment identifier.
[0038] S4: Based on the paragraph structure misalignment identifier, analyze the input document, identify the occurrence position of each term in multiple paragraphs, analyze the contextual collocation differences of terms in multiple paragraphs, determine the part-of-speech continuation changes and collocation expression fluctuations, and obtain term collocation offset tags;
[0039] S5: Call the terminology pairing offset label and paragraph structure misalignment identifier, extract the position information of each error label, filter multiple types of error labels that are adjacent in the same sentence group, analyze the overlapping position of the error label related word groups, determine the repetition of multiple errors in the word coverage level, adjust the annotation structure and merge the suggested content in the label set, and generate the annotation structure merge record.
[0040] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0041] In this invention, a model configuration dataset is generated by parsing the subject-verb-object combination relationship in the text and jointly filtering the parts of speech, sentence patterns, and semantic structures. This improves the accuracy of semantic recognition and structural judgment. By combining the comparison relationship between image action paths and text behavior objects, local features of semantic offset and separation of image and text expression are identified. Based on semantic offset information, the logical structure and topic connection of paragraphs are analyzed, effectively revealing the content misalignment and topic jump between paragraphs. The collocation change trajectory and part-of-speech continuation fluctuation of terms in cross-paragraph contexts are identified. Furthermore, the overlapping and redundancy of tags are analyzed at the sentence group level to form tag merging suggestions and construct annotation structure records. Attached Figure Description
[0042] Figure 1 This is a system flowchart of the present invention;
[0043] Figure 2 This is a flowchart of the sample construction module of the present invention;
[0044] Figure 3 This is a flowchart of the multi-mode analysis module of the present invention;
[0045] Figure 4 This is a flowchart of the structure recognition module of the present invention;
[0046] Figure 5 This is a flowchart of the terminology verification module of the present invention;
[0047] Figure 6 This is a flowchart of the annotation output module of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0049] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0050] Please see Figure 1 This invention provides a technical solution: a multimodal digital publishing intelligent proofreading system based on a large model, comprising:
[0051] The sample construction module is based on the training sample document set. It identifies the subject-verb-object combination relationship of continuous sentence groups in the sample text, analyzes word order stability and sentence construction path, combines part-of-speech changes and semantic structure features, filters training samples, trains and obtains the intelligent review model, and generates the model configuration dataset.
[0052] The multimodal analysis module calls the model configuration dataset, analyzes the input document, identifies the behavioral objects of image action paths and text subject verb combinations, compares the consistency of action-carrying role expression in the text and image, identifies subject-object semantic separation paragraphs, and obtains multimodal semantic offset information;
[0053] The structure recognition module analyzes the semantic connection features of paragraphs and the coverage of the main ideas before and after based on multimodal semantic offset information, compares the repetition of topic word chains, judges the structural connection relationship, identifies semantic breaks, topic jumps and content offset positions, and obtains paragraph structure misalignment indicators.
[0054] The terminology verification module analyzes the input document based on paragraph structure misalignment identifiers, identifies the occurrence position of each term in multiple paragraphs, analyzes the differences in contextual collocation of terms in multiple paragraphs, judges the changes in part-of-speech continuation and collocation expression fluctuations, and obtains term collocation offset tags.
[0055] The annotation output module calls term collocation offset tags and paragraph structure misalignment identifiers to extract the position information of each error tag, filters multiple types of error tags that are adjacent in the same sentence group, analyzes the overlapping position of error tag-related phrases, determines the repetition of multiple errors in the word coverage level, adjusts the annotation structure and merges the suggested content in the tag set, and generates a record of merged annotation structure.
[0056] The model configuration dataset includes a semantic span label set, a perturbation structure group sequence, and a training sample grammar set. Multimodal semantic offset information includes a text-image role reference table, a behavioral expression conflict sequence, and an action semantic offset identifier. Paragraph structure misalignment identifiers specifically include paragraph continuation anomaly groups, title chain affiliation offset sets, and hierarchical jump positioning indexes. Term collocation offset labels specifically include context combination variation records, collocation consistency distribution maps, and term relocation index sequences. Annotation structure merging records include an annotation merging path table, a label association matrix, and an annotation position compression mapping set.
[0057] Please see Figure 2 The sample construction module includes:
[0058] The structural perturbation identification submodule acquires the training sample document set, identifies subject-verb-object combinations in continuous sentence groups, analyzes changes in word order and sentence structure mapping, determines the combination frequency of core components under order perturbation, and extracts core intersection regions and combination conflict regions using the following formula:
[0059] ;
[0060] Calculate the combined perturbation amplitude, screen out the regions with significant perturbations in the main structure, and generate a set of active perturbation scalars for the sentence group;
[0061] in, Let be the combined perturbation amplitude of the i-th sentence group. The subject-verb-object combination weight in the j-th statement of the i-th group is obtained by marking the positions of the subject, verb, and object in the sentence and normalizing their relative frequency. The normalized value of the number of word order switching operations for the j-th statement in the i-th group is obtained by recording the number of word order adjustment operations and normalizing it with the total number of words in the statement. The normalized value of the sentence structure label for the j-th statement in the i-th group is obtained by extracting the sentence type category label of the statement and calculating its proportion in the same sentence type group. This is used to assign a semantic pattern number to the j-th statement in the i-th group. It is a dimensionless classification code, and its value is assigned according to the semantic motivation combination method. is the average depth of the semantic motivation nesting level of the j-th statement, which is obtained by analyzing the depth of the nested path tree structure between semantic motivations. i is the sentence group number, representing the i-th group of sentences in the training sample document set. j is the index number of the statement in the sentence group, representing the j-th statement in the i-th sentence group. n is the total number of sentence groups in the training sample document set, which is obtained by counting the number of all sentence groups divided according to semantic continuity in the input corpus.
[0062] The training sample document set obtained by the structural perturbation recognition submodule, taking a digital textbook review scenario as an example, firstly, multiple sentence groups formed by semantic continuity in the textbook content are sequentially labeled as group i. For example, if a sentence group includes five consecutive sentences, the actual positions of the subject, predicate, and object in each sentence are manually or automatically labeled. For example, in the first sentence "Students use software to learn", the labeling result is: the subject "students" is in position 1, the predicate "uses" is in position 2, and the object "software" is in position 3. Then, the relative frequency of the subject-verb-object order in all sentences is counted. For example, the combination of "subject-verb-object" accounts for 70% of the total, and the weights are normalized. =0.7, then record the number of word order adjustments. If the sentence is edited twice and the total number of words in the sentence is 10, then the normalized value of the word order switching count is... Then, sentence structure is marked. For example, if a sentence belongs to the "active declarative" type and this sentence type accounts for 80% of the total sentences in this group, then the normalized value of the sentence structure mark is... =0.8, and further based on the combination of semantic motivations, the semantic patterns of the statements are classified into pattern numbers. =3 (for example, pattern 1 is a simple statement, and pattern 3 is a complex statement), then calculate the nested path tree structure depth of semantic motivations. If the nested motivation depths in the statements are 2, 3, and 4 levels respectively, then the average depth is... Finally, based on these data, the combined disturbance amplitude is calculated using a formula. The formula parameters are explained below: The perturbation amplitude of the i-th group of sentences is used to measure the degree of perturbation of the word order of the sentence group; This represents the positional weight of the subject-verb-object combination of the j-th statement in the i-th group, such as 0.7 in the example above; This is the normalized value for the number of word order switching times of the j-th statement in the i-th group, such as 0.2 in the example above; Assign a normalized value to the sentence structure of the j-th statement in the i-th group, as in example 0.8 above; It is used to assign semantic pattern numbers, which are dimensionless and take integer classification codes, such as 1, 2, 3, depending on the semantic complexity of the sentence. represents the average depth of the semantic motivation nesting level of the j-th statement, obtained through actual analysis, such as 3 mentioned above; n is the total number of sentence groups in the training sample document, obtained by statistical analysis of actual textbook content, such as n=5 in this example. The actual calculation example is as follows (i=1), assuming the first sentence group contains 3 statements, and its actual data is shown in Table 1:
[0063] Table 1. Calculation parameters for sentence group combination perturbation.
[0064]
[0065] Using the above data, the specific calculation process of the formula is as follows:
[0066] ;
[0067] First part of the calculations:
[0068] ;
[0069] Part Two Calculations:
[0070] ;
[0071] The final result is:
[0072] ;
[0073] Among them, the combination perturbation amplitude refers to the stability deviation measure of the structural combination constructed around the subject-verb-object semantic core within a given sentence group after being affected by word order perturbation, semantic motivation drift, sentence nesting mismatch, etc. It is a numerical indicator that measures the degree to which the word order perturbation of the sentence group disrupts the consistency of the main expression. The parameter essentially comprehensively evaluates the combination concentration, semantic cohesion, and distribution shift trend of language elements in the sentence group along the main structural line. It is used to identify the local areas most prone to main ambiguity in the training samples and serves as a direct reference for screening samples with high risk of semantic misleading. Specifically, when the combination perturbation amplitude of a sentence group is much greater than that of other samples, its main information expression path shows abnormalities such as jumps, fragmentation, or multiple intention intersections. Its training value is significant, and it is given priority as an abnormal main perturbation sample for robust model training, determining which language segments contribute the most to the structural learning of the intelligent proofreading model. Combination perturbation amplitude The threshold was set at 0.5, determined after manual review of 30 sets of actual textbook sentence groups. When the main structure perturbation of the sentence group is considered a significant perturbation, the calculated value of 0.693 in this example indicates that the first group of sentences belongs to the region of significant perturbation, thereby generating the main perturbation scalar set of the sentence group.
[0074] The semantic span annotation submodule calls the active perturbation dynamic scalar set of the sentence group, analyzes the dependency connection trajectory between keywords and the dynamic structural phrases within the sentence group, records the index interval of the keyword covered phrases, annotates the semantic span range in the sentence group, and obtains the semantic span interval data of keywords.
[0075] The semantic span annotation submodule, based on the active perturbation scalar set of sentence groups, takes sentence groups marked as significantly perturbed during the review of a digital textbook (e.g., the first group of sentences with a perturbation amplitude of 0.693) as its execution target. First, it selects keywords within the sentence group, such as "intelligent interactive platform." Through sentence-by-sentence analysis, it specifically records the lexical position index of the keywords within each sentence. Assuming a keyword is located at the 3rd to 7th lexical position index in a sentence, it uses grammatical dependency analysis to sequentially determine the direct dependency relationship between each word within that index interval and the motivational structure (such as subject-verb-object or verb-object phrases) in the sentence, thereby confirming the specific semantic range covered by the keyword. If the 3rd word is "utilize" and the 7th word is "teaching interaction," then that interval is determined to be the keyword's semantic span. The semantic span of the keyword "intelligent interaction platform" is defined, and the start and end indices of this span are marked. Then, the above process is repeated for other sentences in the sentence group, recording the keyword coverage index range and span length (e.g., span length is 5). Subsequently, all keyword coverage ranges in the sentence group are statistically analyzed, and repeated or overlapping positions are merged (e.g., positions 3-7 and 5-9 are merged into position 3-9). The correspondence between the merged range and the semantic motivation in the sentence group is checked again to determine the semantic consistency of the merged span range. Redundant word indices within the span range that have no obvious semantic dependence on the keyword are deleted (e.g., the 8th index "and" is removed). Finally, accurate keyword semantic span range data is formed (e.g., the final range is position 3-7, and the span length is still 5).
[0076] The training data generation submodule filters training samples based on keyword semantic span range data, trains and obtains the intelligent review model, and generates the model configuration dataset.
[0077] The training data generation submodule, based on the specific content of the aforementioned keyword semantic span range data, such as in the digital textbook review scenario where the semantic span length of the keyword "intelligent interactive platform" is 5 words, first checks the specific numerical value of the keyword semantic span length in each sentence group in the textbook. For example, in a textbook review sample document with 200 sentence groups, the actual semantic span length of the keywords in each sentence group is counted sequentially. The semantic span value of the keywords obtained for each sentence group is compared one by one with the aforementioned span length of 5. Sentence groups with actual span values between 4 and 6 words (e.g., a sentence group with a keyword span length of 4, 5, or 6 words, meeting the set screening criteria) are selected as valid training samples. Then, the selected valid training samples undergo further content consistency checks, deleting sentence groups that, although meeting the keyword semantic span length criteria, have obviously discontinuous semantic content or logical inconsistencies (e.g., a sentence group with a keyword span of 5 words, but with semantic gaps or logical contradictions within the sentences). After the initial screening, for example, 120 training samples that meet the requirements are selected from 200 sentence groups. These 120 sentence groups are then used for training. Specifically, the index position and span length of the keyword semantic span interval are used as the core inputs. Through multiple rounds of data iteration calculation (e.g., 50 iterations, each iteration repeatedly executes the review model fitting process using all sentence group data within the sample, and the error rate of each round of iteration calculation gradually decreases from the initial 20%), the iteration process is terminated when the error accuracy drops to the set error threshold (e.g., the error rate is less than or equal to 10%, which is determined through historical data statistical analysis and actual textbook review experiments). This results in a fully trained intelligent review model. Finally, the relevant parameters of the trained model are configured, and the parameter values of the trained model are recorded (e.g., the semantic span length judgment threshold is 4 to 6 words, and the upper limit of the error rate is 10%). Finally, an intelligent review model configuration dataset that can be actually used in digital textbook review scenarios is generated.
[0078] Please see Figure 3 The multi-mode analysis module includes:
[0079] The action path recognition submodule acquires the model configuration dataset, analyzes the action execution path of each image in the input document and the combination of text subject verbs, establishes the pairing relationship between image actions and text behavior objects, and generates an image-text action pairing matrix.
[0080] The action path recognition submodule is based on the textbook review model included in the model configuration dataset. Using the text and images in the digital textbook as input, it first extracts action entities from each image in the textbook. Specifically, through action semantic segmentation, it marks the positions of the action-related entities in the image. For example, if an image contains two action roles, "the teacher points to the screen" and "the student looks up to observe," it clarifies the position of the roles in the image coordinates and the type of action. Furthermore, it records the spatial path trajectory of each role performing the action within the image, obtaining the corresponding position sequence data of the roles and actions. Simultaneously, it uses the corresponding text paragraph content below the image, such as "The teacher displays the projector, and the students pay attention to observe..." The text analysis process involves marking the semantic information of each subject and corresponding verb combination within the text, extracting the subjects "teacher" and "student" and the verbs "display" and "observe," and establishing a mapping relationship between the image action roles and the text action subjects. Entities in the two types of data are paired one by one; for example, the action entity "teacher" in the image is paired with the subject "teacher" in the text, and the action entity "student" in the image is paired with the subject "student" in the text. Finally, based on the pairing relationship, the specific positional relationship between the action roles and subjects of each image-text combination in the document is recorded, forming a matrix-formatted image-text action pairing matrix data according to the actual number of pairs.
[0081] The role consistency judgment submodule compares the expression of the role carrying the action in the subject position in each group of image and text data based on the image-text action pairing matrix, and analyzes the correspondence between the image action entity and the text behavior subject, using the formula:
[0082] ;
[0083] Calculate the subject-object semantic consistency offset value and generate a subject-object semantic offset coefficient sequence;
[0084] in, The subject-object semantic consistency offset value represents the degree of semantic consistency offset of the k-th image-text behavior pair. For the first in the image The normalized semantic encoding value of each action character in the k-th group is obtained by extracting semantic features and normalizing them after obtaining the character entity through image semantic segmentation and character recognition model. For the first in the text The normalized semantic encoding values of the subject roles in the k-th group are obtained by extracting semantic features from the subject-verb combination through syntactic analysis and then normalizing them. For the first in the image The semantic order displacement difference between each action role and its corresponding text behavior in group k is calculated by mapping the difference between the time axis or narrative order of the image and text content. M is the number of semantically aligned entities in each image-text behavior pair, which is determined by the minimum mapping number between the number of image roles and the number of text subjects. k is the index number of the image-text alignment object, and k is the index number of the image-text semantic combination group;
[0085] The role consistency judgment submodule uses the image-text action pairing matrix data as a basis to perform expression analysis of the role carried by the action group by group. For each image-text action pairing item, the semantic feature encoding value of the action role in the image is first extracted. By using actual character characteristics, such as action type, posture, and spatial position, feature vector encoding is used for transformation and normalization. For example, if the character action is "teacher pointing at the screen," the normalized encoded value would be... =0.75, the normalized coded value for the role "student looking up and observing". =0.65, and for the paired text subject role, such as the text subject "teacher displays the projector", word vector encoding values are extracted based on the text semantic lexical features and normalized to obtain the result. =0.7, normalized code value for "students' attention to observation" =0.6, next we compare the shift difference in semantic order between the image and the text. For example, in terms of time or narrative order, if the teacher's action is first in the image and the subject of the same character is also first in the text, then the difference in order displacement is... In the image, the student's action is in the second position, and in the text, the corresponding subject is in the third position. The difference is... The minimum number of entities mapped to roles in each image-text behavior pair is determined to be M=2. Substitute the parameter values into the formula: The meanings of the parameters in the formula are as follows: , is the subject-object semantic consistency offset value of the kth group of image-text behavior pairs, which measures the semantic consistency between images and text; For the k-th image Normalized semantic encoding value of each action role (normalized value of action feature); For the k-th text Normalized semantic encoding values of subject roles (normalized values of text lexical features); This represents the semantic order displacement difference (entity position difference) between the image character and the text subject; M is the minimum number of characters in each image-text pair, obtained through character counting. A specific implementation example is as follows: Substituting the aforementioned parameters into the formula, assuming the parameters for the first image-text pair are: =0.75, =0.7, =0, =0.65, =0.6, If =1, then the calculation is as follows:
[0086] Calculation formula numerator:
[0087] ;
[0088] Calculation formula denominator:
[0089] ;
[0090] Calculate the offset value :
[0091] ;
[0092] The subject-object semantic consistency offset value is a quantitative indicator that measures the degree of difference in semantic expression and structural sequence between the action performer in the image and the subject in the text. This parameter characterizes whether the subject in the image and the subject in the text description are consistent when describing the same behavioral event, and the degree of their positional offset in the semantic space. A larger value indicates a greater difference in semantic matching and structural alignment between the subject and object behavioral descriptions in the image and text, suggesting a cross-modal understanding conflict. It is used to filter semantically conflicting paragraphs between the image and text, providing target paragraphs for subsequent structural recognition and annotation output modules. Embedded as a scoring factor in the multimodal comparison engine, it supports behavioral alignment anomaly detection, effectively improving the system's ability to identify inconsistent content between images and text. It possesses numerical characteristics with unified dimensions and comparable indicators, serving as the core data output for multimodal behavioral consistency judgment in publication review. Experiments determined the consistency offset threshold to be 0.05, i.e. A value less than or equal to 0.05 is considered to indicate that the subject and object semantic expressions are consistent. The semantic offset value for the first group of text-image behavior pairs is 0.026, indicating that the semantic expressions of this group of text-image pairs are consistent. All calculation results are summarized to form a sequence of subject-object semantic offset coefficients. The formula uses parameters... , and displacement difference The participation of [the system / component] enables accurate measurement of the semantic consistency between roles in text and images.
[0093] The semantic offset annotation submodule calls the subject-object semantic offset coefficient sequence to determine the consistency of subject-object semantics, filters semantically separated paragraphs, marks the degree of semantic difference between text and image, and obtains multimodal semantic offset information;
[0094] The semantic offset annotation submodule calls the subject-object semantic offset coefficient sequence, such as the sequence of multiple sets of text and image behavior data in the review of digital textbooks. First, using a set consistency offset threshold of 0.05 as the criterion, each offset coefficient value in the sequence is compared with the threshold to determine the semantic consistency of each group of text and image behaviors. For example, if the first group's offset value of 0.026 is less than 0.05, it is judged as consistent and is not marked. If the second group's offset value of 0.062 is greater than 0.05, it is judged as semantically inconsistent and marked as a subject-object semantic separation paragraph. The difference in text and image behaviors in this paragraph is recorded as 0.062. The third group's value of 0.015 is less than the threshold, and the fourth group's value of 0.089 is greater than the threshold. The semantic difference in all groups of text and image behaviors in the textbook document is marked in this way. All marked difference data are then summarized to complete the multimodal semantic offset information annotation of the textbook review document, and finally, accurate multimodal semantic offset information is obtained.
[0095] Please see Figure 4 The structure recognition module includes:
[0096] The word grouping and connection recognition submodule acquires multimodal semantic offset information, identifies semantic word grouping and connection features of subject-object semantically separated paragraphs, analyzes the main theme coverage and logical connection expression of preceding and following paragraphs, and generates paragraph word grouping and connection data;
[0097] The word connection recognition submodule, based on the subject-object semantic separation paragraphs marked in multimodal semantic offset information, takes a chapter marked as semantically separated in a digital textbook review scenario as an example. It analyzes the semantic connection relationships between adjacent words in each separated paragraph. Specifically, it first selects adjacent word pairs from the separated paragraphs, such as the paragraph sentence "After students complete the experimental operation, the experimental equipment should be returned promptly." It judges the semantic connection between adjacent word pairs word by word, such as "complete" and "experimental operation," and "experimental equipment" and "return promptly," extracting semantic connection features and marking the validity of the word pair connections. Then, it records the connection features of each marked pair and stores them using vector encoding. Simultaneously, it continues to extract the main content of the preceding and following paragraphs, using frequently occurring core words in each paragraph, such as "experimental equipment." Based on terms such as "equipment", "return", and "experimental records", the intersection and difference of the main vocabulary sets of the preceding and following paragraphs are compared and calculated to determine the actual scope of the main theme coverage of each paragraph. For example, if the current paragraph covers the main vocabulary "experimental equipment" and "return", and the previous paragraph covers the main vocabulary "experimental equipment" and "experimental records", then the intersection of the main theme coverage is "experimental equipment", and the difference is "return" and "experimental records". Then, by analyzing the logical expression relationship between the end and beginning sentences of the preceding and following paragraphs word by word, the connection features are recorded to see if there are obvious logical jumps. For example, if the previous paragraph ends with "experimental records" but the next paragraph begins directly with "equipment return", it is marked as a logical jump connection. After analyzing the entire chapter paragraph by paragraph through the above steps, a complete paragraph word connection data is formed.
[0098] The structural chain attribution judgment submodule compares the repetition of the topic word chain formed by the title keywords in the structural hierarchy based on the paragraph word connection data, judges the continuity of the paragraph theme in the chapter structure, and analyzes the topic chain structural jumps and content offsets, using the following formula:
[0099] ;
[0100] Calculate the chapter structure offset strength, filter the structure jump intervals, and generate a topic chain offset coefficient sequence;
[0101] in, For chapter structure offset strength, Let q be the length of the keyword chain in the title of the qth and vth paragraphs. This is obtained by extracting keywords from the paragraph titles and counting the number of their chain connections. The length of the main idea word chain for paragraph v in chapter q is obtained by constructing and counting word chains after obtaining the main idea words of the paragraph through word segmentation and part-of-speech determination. Let be the normalized value of the main idea coverage area of paragraph v in chapter q. This value is obtained by calculating the ratio of the area covered by the main idea words in this paragraph to the total length of the paragraph. The semantic connection difference item in paragraph v of chapter q is obtained by the intersection difference of the word connection relationship between this paragraph and the previous paragraph. N is the total number of paragraphs in the chapter, which is obtained by counting the number of paragraphs in chapter q. q is the chapter index number, which represents the number of the chapter being processed. v is the paragraph index number, which represents the paragraph sequence number within the chapter.
[0102] The structure chain attribution judgment submodule calls the paragraph word grouping connection data, using the specific chapter and paragraph structure of the digital textbook as an example to demonstrate the detailed calculation process. For example, taking Chapter 3 of the textbook, which contains 4 paragraphs, as an example, the following process is performed for each paragraph: First, extract the keywords of each paragraph title. For example, if the title of the first paragraph is "Requisition and Recording of Experimental Equipment", the length of the title keyword chain is... Next, extract the main idea word chain from the main text of the paragraph, and obtain it through Chinese word segmentation and part-of-speech tagging. For example, the length of the main idea word chain of the main text of the paragraph is... Further normalize the coverage area of the paragraph's main idea. The calculation involves counting the number of sentences covered by the topic word (assuming a paragraph has 10 sentences and the topic word covers 7), calculating the coverage area as 7 / 10 = 0.7; and finally, calculating the semantic connectivity difference between this paragraph and the preceding paragraph. This refers to comparing the difference between the current paragraph and the previous paragraph's set of main ideas. If the previous paragraph has 2 main ideas and the current paragraph has 3, with 1 main idea being shared, then the difference value is... Substitute the above parameters into the formula: The specific explanations of each parameter in the formula are as follows: The structural offset intensity of Chapter q is used to measure the degree of structural offset between paragraphs. The length of the keyword chain in the title of paragraph v of chapter q is obtained by counting the number of title keywords; The length of the main idea word chain in the main text is obtained by counting the number of main idea words in each paragraph. The normalized value of the coverage area of the main idea words in paragraph v of chapter q is calculated from the ratio of the number of sentences covered by the main idea words to the total number of sentences. The semantic connection difference term is obtained by intersecting and subtracting the sets of thematic words of the preceding and following paragraphs; N is the total number of paragraphs in the current chapter, obtained by counting the number of paragraphs within the chapter. Taking Chapter 3 as an example, assume that the parameters are calculated for each of the four paragraphs:
[0103] Paragraph 1: ;
[0104] Paragraph 2: ;
[0105] Paragraph 3: ;
[0106] Paragraph 4: ;
[0107] The specific calculation process involves substituting the above data into the formula as follows:
[0108] Calculation of structural offset strength in Chapter 3 :
[0109] ;
[0110] Specifically, it can be elaborated as follows:
[0111] ;
[0112] calculate:
[0113] ;
[0114] Among them, the chapter structure offset strength is an important indicator for measuring the rationality of the internal structure of each chapter in a document. The higher the value, the more severe the break, misalignment, or logical jump in the topic word chain of paragraphs in that chapter. This parameter not only reflects the degree of matching between the title and the main idea, but also covers the semantic concentration of paragraphs in the chapter and the continuity of connections between adjacent paragraphs. In practical use, this parameter is often used to locate structural breaks, assess the risk of content jumps, and can serve as a structural early warning signal in intelligent proofreading systems for subsequent reorganization, review, or visual annotation. Experimental statistics show that a structural offset strength threshold of 0.4 is set, i.e. This is marked as a significant structural shift interval. A calculated value of 0.4355 indicates a significant structural shift within this chapter, thus determining the structural shift interval and obtaining the topic chain shift coefficient sequence. The formula uses parameters... , , , The combined operation effectively captures the degree of jump in the internal structure of chapters and paragraphs.
[0115] The hierarchical misalignment positioning submodule calls the topic chain offset coefficient sequence to determine the continuity of paragraph topics in the chapter structure and the coherence of paragraph relationships, identify semantic breaks, topic jumps and content offset positions, and obtain paragraph structure misalignment identifiers.
[0116] The hierarchical misalignment positioning submodule calls the topic chain offset coefficient sequence, using the structural offset coefficient values obtained from Chapter 3 of the digital textbook. Based on a value of 0.4355, the structural continuity of each chapter and paragraph is determined one by one. First, for the structural offset coefficient sequence determined above, the offset coefficient value obtained for each chapter is compared with the set structural offset threshold of 0.4 to determine the continuity between paragraphs. For example, the offset coefficient value obtained for Chapter 3 is 0.4355, which is higher than the threshold of 0.4. Therefore, specific positioning operations are performed, including analyzing the semantic breakpoints between paragraphs within the chapter. For example, the position from the end of the second paragraph to the beginning of the third paragraph is determined, where the main idea jumps abruptly from "experimental record" to "equipment return". This position is recorded as the topic jump position. Then, by analyzing the specific content offset within the paragraph, such as the difference between the main idea of the fourth paragraph and the previous paragraph exceeding 30% of the main idea coverage area, it is marked as a content offset position. Through detailed analysis of each paragraph, all marked semantic breaks, topic jumps, and content offset positions are recorded one by one, and finally, the accurate paragraph structural misalignment mark of the chapter is obtained.
[0117] Please see Figure 5 The terminology verification module includes:
[0118] The terminology distribution recognition submodule obtains paragraph structure misalignment identifiers, detects the distribution position of each term in multiple paragraphs in the input document, filters the frequency of each term at multiple levels, establishes terminology distribution statistics, and generates terminology paragraph distribution data.
[0119] The terminology distribution identification submodule, based on paragraph structure misalignment identifiers, starts from the actual content of digital textbook review. For example, a textbook may contain terms such as "smart terminal," "interactive platform," and "behavioral path." The specific steps are as follows: First, it searches all chapters of the textbook, recording the index of each term's appearance in each paragraph. For example, the term "smart terminal" appears in paragraph 2 of Chapter 1, paragraph 4 of Chapter 3, and paragraphs 1 and 2 of Chapter 4; the term "interactive platform" appears in paragraph 3 of Chapter 1, paragraph 2 of Chapter 2, and paragraph 1 of Chapter 3. The term "behavioral path" appears in paragraph 4 of Chapter 2, paragraph 3 of Chapter 3, and paragraphs 2 and 4 of Chapter 5. The frequency of occurrence of the term within each chapter level is then calculated. For example, "smart terminal" appears twice in Chapter 4, accounting for 40% of the total number of paragraphs in this chapter (5 paragraphs), while the term "interactive platform" appears only once in Chapter 3, accounting for 25% of the total number of paragraphs in this chapter (4 paragraphs). Then, terms with a frequency of more than 20% at each level are selected, and frequency statistics are performed by chapter. The distribution data of all terms in the textbook is summarized to form the term paragraph distribution.
[0120] The contextual collocation analysis submodule analyzes the frequency distribution of verb and noun combinations of terms in the context before and after each paragraph based on the term paragraph distribution, detects the change trajectory of part-of-speech continuation patterns, calculates the cumulative number of word substitutions in the collocation structure, and obtains the term collocation fluctuation coefficient.
[0121] The contextual collocation analysis submodule, based on the distribution of terms across paragraphs, uses the textbook term "smart terminal" as an example. It analyzes the frequency distribution of verb and noun combinations in the context before and after the term appears in each paragraph. For instance, in Chapter 4, Paragraph 1, the contextual combinations before and after the term "smart terminal" are "using a smart terminal for learning" and "smart terminals have interactive functions." The frequency of the verbs "use" and "have" combined with the nouns "learn" and "interactive functions" is extracted. Further analysis is performed when the term appears in Chapter 4, Paragraph 2, in the phrase "configuring a smart terminal and recording data." The frequency of the verb "configure" and the noun "data" combination is extracted. The frequency distribution of verb and noun combinations in each paragraph is statistically analyzed. Furthermore, grammatical part-of-speech tagging is used to detect changes in part-of-speech connections before and after the term, such as "verb + term + verb" or "verb + term + noun," to determine the frequency of each term combination. When a term appears, its specific part-of-speech combination pattern is analyzed. The differences in part-of-speech patterns in the context before and after the term are compared one by one. For example, the pattern in paragraph 1 of Chapter 4 is "verb + term + verb", and in paragraph 2 it is "verb + term + noun", which is recorded as one pattern change. Then, the cumulative number of word substitutions in the collocation structure before and after the term is calculated. Specifically, the differences in the collocation words in the context of the term in adjacent paragraphs are compared one by one. For example, from paragraph 1 to paragraph 2, the verb before and after the term changes from "use" to "configure", and the noun changes from "learn" to "data", so the cumulative number of substitutions is 2. The total number of substitutions for the term is calculated by counting all paragraphs in the same way. For example, the term "smart terminal" has a total of 6 word substitutions in the textbook. The number of pattern changes and word substitutions obtained above are normalized and used as the term collocation fluctuation coefficient. For example, the normalized collocation fluctuation coefficient is 0.35.
[0122] The collocation consistency judgment submodule calls the term collocation fluctuation coefficient to compare the collocation stability of each term in multiple paragraphs, judges the collocation consistency of the term in the context of multiple paragraphs, and obtains the term collocation offset label.
[0123] The collocation consistency judgment submodule calls the term collocation fluctuation coefficient. Taking the terms "smart terminal," "interactive platform," and "behavioral path" as examples, it compares the values of the collocation fluctuation coefficient of each term in different paragraphs. For example, the collocation fluctuation coefficient of "smart terminal" is 0.35, that of "interactive platform" is 0.25, and that of "behavioral path" is 0.42. The collocation fluctuation threshold is set to 0.3. Through statistical experiments on actual data from the textbook, it is determined that when the collocation fluctuation coefficient exceeds 0.3, the term collocation stability is considered poor. Based on this standard, the terms "smart terminal" and "behavioral path" are judged to have poor collocation stability, while "interactive platform" has good collocation stability. Then, "smart terminal" and "behavioral path" are marked as term collocation offset labels. The paragraphs where the terms marked with collocation offset labels are located are further recorded to obtain the term collocation offset labels.
[0124] Table 2. Statistical Table of Terminology Combinations and Volatility Coefficients
[0125]
[0126] Table 2 lists the statistical results of the specific collocation fluctuation coefficients of each term in the textbook review scenario. "Smart terminal" and "behavioral path" significantly exceed the threshold of 0.3, so they are marked as term collocation offset labels. "Interactive platform" is not marked within the threshold range.
[0127] Please see Figure 6 The annotation output module includes:
[0128] The error tag extraction submodule calls term collocation offset tags and paragraph structure misalignment indicators to extract the position information of each error tag, filters multiple types of error tags that are adjacent in the same sentence group, and obtains the error tag position sequence.
[0129] The error tag extraction submodule uses terminology offset tags (such as "smart terminal" and "behavioral path") marked during textbook review and paragraph structure misalignment markers (such as the topic jump position in paragraph 2 of chapter 3) as input. The specific steps are as follows: First, each term offset tag is searched and located paragraph by paragraph. For example, "smart terminal" is located in sentence 5 of paragraph 1 of chapter 4, and "behavioral path" is located in sentence 3 of paragraph 2 of chapter 5. Simultaneously, the paragraph structure misalignment marker position (such as sentence 1 of paragraph 2 of chapter 3) is located. The position of each tag is extracted using the sentence as the smallest positioning unit. The process begins by comparing all sentence groups in the textbook one by one to identify multiple error tags within the same paragraph or sentence group. For example, the fifth sentence of the first paragraph of Chapter 4 contains both a term offset tag for "smart terminal" and a paragraph structure misalignment marker. All sentence groups with multiple adjacent error tags within the same sentence group are then selected, and the position information of each error tag is recorded sequentially. For instance, the position record for the fifth sentence of the first paragraph of Chapter 4 is (4-1-5), indicating that there are multiple adjacent error tags in the fifth sentence of the first paragraph of Chapter 4. By summarizing all position records, a complete sequence of error tag position data is finally obtained.
[0130] The overlap relationship analysis submodule analyzes the overlapping positions of related word groups with error tags within the same sentence group based on the error tag position sequence, determines the repetition of error tags in the word coverage level, and obtains the error overlap relationship matrix.
[0131] The overlapping relationship analysis submodule is based on the error label position sequence. Taking the sentence group of sentence 5 in paragraph 1 of Chapter 4 (marked as 4-1-5) as an example, the specific execution steps include analyzing the position of the phrase associated with each error label in the sentence group one by one. For example, the phrase position corresponding to the error label of the term "smart terminal" in this sentence group is the 3rd to 5th words, and the phrase position corresponding to the paragraph structure misalignment mark is the 4th to 6th words. The overlapping position of the two error label phrases is determined to be the 4th to 5th words. After recording the specific overlapping position, the overlapping relationship of other error label phrases is analyzed, such as the sentence group of sentence 3 in paragraph 2 of Chapter 5 (marked as 5-2-3). If the term "behavioral path" appears as the 2nd to 4th word and another label appears as the 1st to 3rd word, then the overlapping position is the 2nd to 3rd word. After analyzing each sentence group, the repetition of the overlapping position at the word coverage level is further determined. For example, if the overlapping position (4th to 5th word) in the 4-1-5 sentence group is covered by multiple labels at the same time, it belongs to the area with a high repetition coverage level and is marked as a high repetition area. If the overlapping area is covered by only two labels, it is marked as a low repetition area. The overlapping position and repetition coverage level information of all sentence groups are recorded and then the error overlap relationship matrix data is generated after sorting, as shown in Table 3.
[0132] Table 3 Error Overlap Relationship Matrix
[0133]
[0134] Table 3 lists the overlapping positions and repetition levels of error tags in different sentence groups during the textbook review process. Among them, the overlapping positions in sentence group 4-1-5 have a higher coverage level and should be merged first in subsequent processing.
[0135] The tag merging output submodule calls the error overlap matrix, adjusts the annotation structure, merges the suggested content in the tag set, and generates a record of the merged annotation structure.
[0136] The tag merging output submodule calls the error overlap matrix data given in Table 3. Taking sentence group 4-1-5 as an example, the specific execution steps include first adjusting the annotation structure within the sentence group one by one according to the overlapping position (the 4th to 5th words) and repetition level (high) information recorded in the overlap matrix. The suggested content of the originally separate term offset tags and paragraph structure misalignment tags is merged. For example, the original suggested content of the term "smart terminal" tag is "check the semantic consistency of the term", and the suggested content of the structure misalignment tag is "check the logic of structural jumps". After merging, the new annotation suggested content is "check the semantic consistency of the term and the logic of structural jumps". The annotation structure of this merge is recorded, and the merging process of other sentence groups such as 5-2-3 is continued. The corresponding suggested content of the merge is adjusted and merged in turn. The tag merging processing of sentence groups within all overlap matrix data is completed, and finally a complete annotation structure merging record is formed.
[0137] The intelligent proofreading method for multimodal digital publishing based on a large model is executed based on the aforementioned intelligent proofreading system for multimodal digital publishing based on a large model, and includes the following steps:
[0138] S1: Based on the training sample document set, identify the subject-verb-object combination relationship of continuous sentence groups in the sample text, analyze the word order stability and sentence construction path, combine part-of-speech changes and semantic structure features, filter training samples, train and obtain the intelligent review model, and generate the model configuration dataset.
[0139] S2: Call the model configuration dataset, analyze the input document, identify the behavior object of the combination of image action path and text subject verb, compare the consistency of the expression of the role carried by the action in the text and image, identify the subject-object semantic separation paragraph, and obtain multimodal semantic offset information;
[0140] S3: Based on multimodal semantic offset information, analyze the semantic connection features of paragraphs and the coverage of the main ideas before and after, compare the repetition of topic word chains, determine the structural connection relationship, identify semantic breaks, topic jumps and content offset positions, and obtain paragraph structural misalignment indicators;
[0141] S4: Based on the paragraph structure misalignment identifier, analyze the input document, identify the occurrence position of each term in multiple paragraphs, analyze the contextual collocation differences of terms in multiple paragraphs, determine the changes in part of speech and collocation expression fluctuations, and obtain term collocation offset tags;
[0142] S5: Call the terminology collocation offset label and paragraph structure misalignment identifier, extract the position information of each error label, filter multiple types of error labels that are adjacent in the same sentence group, analyze the overlapping position of the word groups associated with the error labels, determine the repetition of multiple errors in the word coverage level, adjust the annotation structure and merge the suggested content in the label set, and generate the annotation structure merge record.
[0143] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A large model-based multi-modal digital publishing intelligent proofreading system, characterized in that, The system comprises: The sample construction module identifies the subject-predicate-object combination relationship of a continuous sentence group in a sample text based on a training sample document set, analyzes the word order stability and sentence structure path, filters the training samples, trains and obtains an intelligent proofreading model, and generates a model configuration dataset; The sample construction module comprises: The structure disturbance identification submodule obtains a training sample document set, identifies the subject-predicate-object combination of a continuous sentence group, analyzes the word group arrangement order change and sentence structure mapping, judges the combination frequency of the main components under the order disturbance, extracts the main stem interlaced area and combination conflict area, calculates the combination disturbance amplitude, filters the main stem structure disturbance significant area, and generates a sentence group main stem disturbance scalar set; The semantic span annotation submodule calls the sentence group main stem disturbance scalar set, analyzes the dependence connection track of the key words and the cause structure word group in the sentence group, records the key word coverage word group index interval, annotates the semantic span range in the sentence group, and obtains key word semantic span interval data; The training data generation submodule filters the training samples according to the key word semantic span interval data, trains and obtains an intelligent proofreading model, and generates a model configuration dataset; The multi-modal analysis module calls the model configuration dataset, analyzes the input document, identifies the behavior object of the image action path and the text subject-verb combination, compares the expression consistency of the action carrying roles in the image and the text, identifies the subject-object semantic separation paragraph, and obtains multi-modal semantic deviation information; The structure identification module analyzes the paragraph semantic connection characteristics and the front and back main idea coverage range based on the multi-modal semantic deviation information, compares the theme word chain attribution repeatability, judges the structure connection relationship, identifies the semantic break, theme jump and content deviation position, and obtains the paragraph structure dislocation identifier; The term verification module analyzes the input document according to the paragraph structure dislocation identifier, identifies the appearance position of each term in multiple paragraphs, analyzes the context collocation difference of the term in multiple paragraphs, judges the part-of-speech continuity change and collocation expression fluctuation, and obtains the term collocation deviation tag.
2. The large model-based multi-modal digital publishing intelligent proofreading system according to claim 1, characterized in that, The model configuration dataset comprises a semantic span tag set, a disturbance structure group sequence, and a training sample sentence set. The multi-modal semantic deviation information comprises an image-text role pointing comparison table, a behavior expression conflict sequence, and an action semantic deviation identifier. The paragraph structure dislocation identifier is specifically a paragraph connection abnormal group, a title chain attribution deviation set, and a hierarchical jump positioning index. The term collocation deviation tag is specifically a context combination variation record, a collocation consistency distribution diagram, and a term repositioning index sequence.
3. The large model-based multi-modal digital publishing intelligent proofreading system according to claim 1, characterized in that, The multi-modal analysis module comprises: The action path identification submodule obtains the model configuration dataset, analyzes the action execution path of each image and the text subject-verb combination in the input document, establishes the pairing relationship between the image action and the text behavior object, and generates an image-text action pairing matrix; The role consistency judgment submodule compares the expression of the action carrying roles in the subject position in each group of image-text data based on the image-text action pairing matrix, analyzes the corresponding relationship between the image action entity and the text behavior subject, calculates the subject-object semantic consistency deviation value, and generates a subject-object semantic deviation coefficient sequence; The semantic deviation marking sub-module calls the host-guest semantic deviation coefficient sequence, judges the consistency of the host-guest semantics, filters the semantic separation paragraphs, and marks the semantic deviation degree to obtain the multi-modal semantic deviation information.
4. The large model-based multi-modal digital publishing intelligent proofreading system according to claim 3, characterized in that, The structure identification module comprises: The word group connection identification sub-module obtains the multi-modal semantic deviation information, identifies the semantic word group connection features of the host-guest semantic separation paragraphs, analyzes the main idea coverage range and the logical connection expression mode of the previous and subsequent paragraphs, and generates the paragraph word group connection data; The structure chain attribution judgment sub-module compares the attribution repeatability of the theme word chain composed of the title keywords in the structure hierarchy according to the paragraph word group connection data, judges the connection performance of the paragraph theme in the chapter structure, analyzes the theme chain structure jump and content deviation, calculates the chapter structure deviation strength, filters the structure jump interval, and generates the theme chain deviation coefficient sequence; The hierarchical dislocation positioning sub-module calls the theme chain deviation coefficient sequence, judges the connection performance of the paragraph theme in the chapter structure and the paragraph relationship coherence degree, identifies the semantic break, theme jump and content deviation position, and obtains the paragraph structure dislocation identification.
5. The large model-based multi-modal digital publishing intelligent proofreading system according to claim 4, characterized in that, The term verification module comprises: The term distribution identification sub-module obtains the paragraph structure dislocation identification, detects the distribution position of each term in the input document in multiple paragraphs, filters the appearance frequency of each term in multiple hierarchies, establishes the term distribution statistics, and generates the term paragraph distribution amount; The context collocation analysis sub-module analyzes the frequency distribution of the verb and noun combination of the term in the context before and after each paragraph based on the term paragraph distribution amount, detects the change trajectory of the part-of-speech connection mode, calculates the cumulative number of vocabulary replacement in the collocation structure, and obtains the term collocation fluctuation coefficient; The collocation consistency judgment sub-module calls the term collocation fluctuation coefficient, compares the collocation stability of each term in multiple paragraphs, judges the collocation consistency of the term used in the context of multiple paragraphs, and obtains the term collocation deviation label.
6. The large model-based multi-modal digital publishing intelligent proofreading system according to claim 1, wherein, The system further comprises: The annotation output module calls the term collocation deviation label and the paragraph structure dislocation identification, extracts the position information of each error label, filters multiple types of error labels adjacent in position in the same sentence group, analyzes the overlapping position of the error label associated word group, judges the repetition of multiple errors in the word coverage hierarchy, adjusts the annotation structure and merges the recommended content in the label set, and generates a batch note structure merging record; The batch note structure merging record comprises an annotation merging path table, a label association matrix, and a batch note position compression mapping set.
7. The large model-based multi-modal digital publishing intelligent proofreading system according to claim 6, characterized in that, The annotation output module comprises: The error label extraction sub-module calls the term collocation deviation label and the paragraph structure dislocation identification, extracts the position information of each error label, filters multiple types of error labels adjacent in position in the same sentence group, and obtains an error label position sequence; The overlapping relationship analysis sub-module analyzes the overlapping position of the error label associated word group in the same sentence group based on the error label position sequence, judges the repetition of the error label in the word coverage hierarchy, and obtains an error overlapping relationship matrix; The label merging output sub-module calls the error overlapping relationship matrix, adjusts the annotation structure and merges the recommended content in the label set, and generates a batch note structure merging record.
8. A large model-based multi-modal digital publishing intelligent proofreading method, characterized in that, The method is used to realize the large model-based multi-modal digital publishing intelligent proofreading system according to any one of claims 1-7, comprising the following steps: S1: Based on the training sample document set, the subject-predicate-object combination relationship of the continuous sentence group in the sample text is identified, the word order stability and the sentence structure path are analyzed, the part-of-speech changes and semantic structure features are combined, the training samples are screened, the intelligent proofreading model is trained and obtained, and the model configuration data set is generated; S2: Call the model configuration data set, analyze the input document, identify the action object of the image action path and the text subject-verb combination, compare the consistency of the action carrying role expression in the image and the text, identify the subject-object semantic separation paragraph, and obtain the multi-modal semantic offset information; S3: Based on the multi-modal semantic offset information, analyze the paragraph semantic connection features and the front and back main theme coverage range, compare the theme word chain attribution repeatability, judge the structure connection relationship, identify the semantic fracture, theme jump and content offset position, and obtain the paragraph structure dislocation identifier; S4: According to the paragraph structure dislocation identifier, analyze the input document, identify the appearance position of each term in multiple paragraphs, analyze the context collocation difference of the term in multiple paragraphs, judge the part-of-speech connection change and collocation expression fluctuation, and obtain the term collocation offset label; S5: Call the term collocation offset label and paragraph structure dislocation identifier, extract the position information of each error label, screen multiple types of error labels with adjacent positions in the same sentence group, analyze the overlapping position of the error label associated word group, judge the repetition of multiple errors in the word coverage level, adjust the annotation structure and merge the recommended content in the label set, and generate the annotation structure merging record.
Citation Information
Patent Citations
Document intelligent writing and analysis system based on knowledge graph
CN120449833A
Method and system for processing data based on large language model
CN120471023A