AI-based ancient literature anthology analysis method and system

CN122596066APending Publication Date: 2026-08-18山东电子职业技术学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610827315.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]因此,本发明提供了基于AI的古代文学作品集分析方法及系统,解决现有技术无法准确反映作者长期稳定形成的意象传播路径、情绪递进规律及修辞耦合结构、无法从作者长期稳定认知传播规律层面对作品真实性进行深层判别的问题

Benefits of technology

[0051] The beneficial effects of this invention are as follows: By constructing a literary semantic Gaussian field, a unified representation of the continuous semantic propagation relationship across sentence elements and paragraphs in ancient literary works is achieved. This effectively depicts the cognitive diffusion pattern that the author has formed over a long period of time. By establishing a multi-directional cognitive propagation structure and a dynamic iterative propagation mechanism, the evolution path of imagery, the trajectory of emotional progression, and the rhetorical coupling relationship in the work can be accurately captured, significantly improving the ability to extract deep literary cognitive features. Secondly, by using the cognitive Gaussian distance analysis mechanism to comprehensively consider the covariance coupling relationship between different semantic directions, the accuracy of identifying imitation texts, partially imitated texts, and hybrid generated texts is improved. Therefore, this invention not only improves the accuracy and robustness of the authenticity analysis of ancient literature, but also has good structural interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122596066A_ABST
    Figure CN122596066A_ABST
Patent Text Reader

Abstract

The application discloses an ancient literature works set analysis method and system based on AI, relates to the technical field of natural language processing, and comprises the following steps: obtaining a real work set and a to-be-tested work, and performing standardization processing to generate a standardized text; after decoding based on the standardized text and extracting a character vector sequence, a literature semantic unit is constructed, and a literature semantic unit set is formed; after analyzing the propagation direction based on the literature semantic unit set and constructing a complete literature semantic Gaussian field, a double matrix is used for iteration, and a literature semantic Gaussian field is output; the literature semantic Gaussian field is hierarchically divided and randomly sampled, semantic sampling points are generated, semantic contribution degrees are calculated, the semantic sampling points are compressed or cloned, an enhanced semantic Gaussian field is generated, a multi-directional cognitive propagation structure is constructed by expanding the spherical harmonic propagation, and a cognitive Gaussian field of the to-be-tested work and a target work is generated to perform Mahalanobis operation, so that a Mahalanobis distance deviation value is obtained; and the application improves the accuracy and robustness of the ancient literature authenticity analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to an AI-based method and system for analyzing ancient literary works. Background Technology

[0002] With the continuous development of artificial intelligence, natural language processing, and digital humanities research methods, automated analysis of ancient literary works based on computational models has become an important research direction in the field of literary information processing. Early research on ancient literature mainly relied on manual collation, textual comparison, statistical analysis of language styles, and empirical textual research. Researchers manually summarized word habits, rhetorical structures, allusion patterns, and emotional expression methods to identify works. With the development of statistical language models and machine learning methods, researchers began to introduce techniques such as word frequency statistics, N-gram analysis, support vector machine classification, and topic models to model author style characteristics, achieving some progress in areas such as author identification, authenticity verification, and style transfer analysis of ancient poetry. In recent years, deep semantic modeling techniques, represented by Transformer, have leveraged their long-distance dependency capture capabilities to extract semantic association information across sentences, paragraphs, and even entire texts, providing a new technical path for the deep cognitive structure analysis of ancient literary works.

[0003] Existing technologies struggle to depict the long-distance semantic diffusion relationships across sentences and paragraphs in ancient literary works, failing to accurately reflect the author's long-term stable imagery dissemination paths, emotional progression patterns, and rhetorical coupling structures. Furthermore, the lack of a unified continuous cognitive space modeling mechanism prevents in-depth judgment of the work's authenticity from the perspective of the author's long-term stable cognitive dissemination patterns. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides an AI-based method and system for analyzing ancient literary works, which solves the problems that existing technologies cannot accurately reflect the long-term stable imagery dissemination path, emotional progression pattern, and rhetorical coupling structure formed by the author, and cannot make in-depth judgments on the authenticity of the work from the perspective of the author's long-term stable cognitive dissemination pattern.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] Firstly, this invention provides an AI-based method for analyzing collections of ancient literary works, which includes:

[0008] Obtain a collection of real works and the works to be tested, perform standardization processing, and generate standardized text;

[0009] After decoding the character vector sequence extracted from the standardized text, literary semantic units are constructed, forming a set of literary semantic units;

[0010] Based on the set of literary semantic units, a complete literary semantic Gaussian field is constructed for propagation direction analysis. Then, a dual matrix is ​​established to perform iteration and output the literary semantic Gaussian field.

[0011] The semantic Gaussian field of literature is hierarchically divided and randomly sampled to generate semantic sampling points and calculate semantic contribution. The semantic sampling points are compressed or cloned to generate an enhanced semantic Gaussian field. A multi-directional cognitive propagation structure is constructed by expanding spherical harmonic propagation, and the cognitive Gaussian fields of the work to be tested and the target work are generated. The Mahalanobis operation is performed to obtain the Mahalanobis distance deviation value.

[0012] By performing probability mapping on the Mahalanobis distance deviation value, the probability of forgery is generated to judge the authenticity of the work under test, and the final analysis result is obtained.

[0013] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of obtaining the real works set and the works to be tested, performing standardization processing, and generating standardized text includes:

[0014] By having the user input a collection of the target author's real works and the works to be tested;

[0015] All text works in both the real works collection and the works to be tested are uniformly converted to UTF-8 encoding format; the text works in UTF-8 encoding format are uniformly standardized to obtain standardized text works, including standardized real works and standardized works to be tested.

[0016] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of decoding the character vector sequence extracted from standardized text to construct literary semantic units and form a set of literary semantic units is as follows:

[0017] Read all the characters of the standardized text, construct a character sequence according to the original character order, obtain the character embedding vectors, and combine them to form a character vector sequence;

[0018] Extract the contextual semantic information of the character vector sequence and fuse it to generate a fused contextual semantic state vector sequence;

[0019] The sequence of fused context semantic state vectors is decoded to form a set of sentence elements;

[0020] Five types of literary semantic units are extracted from the sentence element set to form a set of literary semantic units.

[0021] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of constructing a complete literary semantic Gaussian field based on the set of literary semantic units, performing propagation direction analysis, establishing a dual matrix to perform iteration, and outputting the literary semantic Gaussian field is as follows:

[0022] After encoding the literary semantic units in the set of literary semantic units and constructing literary Gaussian units, a continuous semantic propagation connection is established between the literary Gaussian units to form a complete literary semantic Gaussian field.

[0023] Generate propagation direction for literary Gaussian units and establish a dual matrix, including a rotation matrix and a diffusion scale matrix;

[0024] Based on the rotation matrix and the diffusion scale matrix, a semantic covariance matrix is ​​generated for the literary Gaussian unit. The propagation path is updated iteratively according to the propagation overlap relationship between units, and a stable literary semantic Gaussian field of author cognition is output.

[0025] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of hierarchically dividing and randomly sampling the literary semantic Gaussian field includes:

[0026] Based on the stable author's cognition of the literary semantic Gaussian field, after obtaining the semantic density, rhetorical complexity, emotional fluctuation rate and allusion dissemination rate in the field, the entire literary semantic Gaussian field is hierarchically divided to form semantic layers, including the general semantic layer, the rhetorical semantic layer, the emotional semantic layer and the allusion semantic layer.

[0027] In each semantic layer, semantic propagation paths are generated based on the propagation connection relationships between literary Gaussian units, and semantic sampling regions are randomly selected in the propagation paths for random sampling to generate several semantic sampling points.

[0028] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of calculating semantic contribution by compressing or cloning semantic sampling points includes:

[0029] Based on the semantic sampling points, calculate the semantic contribution of each semantic sampling region;

[0030] By using preset contribution thresholds and preset judgment thresholds, it is determined whether the semantic region is a high-complexity semantic region or a low-complexity semantic region. In the case of a high-complexity semantic region, the semantic nodes are cloned, and in the case of a low-complexity semantic region, the semantic nodes are compressed.

[0031] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of constructing a multi-directional cognitive propagation structure using spherical harmonic propagation includes:

[0032] For the enhanced literary semantic Gaussian field, a fixed semantic direction set is established, which includes emotional direction, rhetorical direction, historical direction, rhythmic direction and imagery direction.

[0033] Based on each fixed semantic direction, the enhanced literary Gaussian unit is subjected to spherical harmonic propagation expansion processing to establish the semantic propagation function in the corresponding direction;

[0034] Perform directional propagation modeling based on each semantic propagation function to obtain directional semantic propagation results corresponding to different fixed semantic directions;

[0035] Calculate the directional coupling propagation strength between different directions based on the directional semantic propagation results corresponding to different fixed semantic directions;

[0036] Based on the directional coupling propagation strength and directional semantic propagation results, the multi-directional cognitive propagation structure is calculated.

[0037] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of generating cognitive Gaussian fields of the test work and the target work and performing Mahalanobis operations includes:

[0038] Based on the multi-directional cognitive propagation structure, Gaussian center, Gaussian covariance matrix and cognitive propagation density are generated;

[0039] Based on Gaussian center, Gaussian covariance and cognitive propagation density, generate cognitive Gaussian field of the work to be tested and cognitive Gaussian field of the target work.

[0040] Based on the cognitive Gaussian field of the work under test and the cognitive Gaussian field of the target work, the Mahalanobis distance deviation value between the cognitive Gaussian fields is obtained.

[0041] As a preferred embodiment of the AI-based ancient literary works analysis method of the present invention, the step of performing probability mapping on the Mahalanobis distance deviation value to generate a forgery probability for judging the authenticity of the work under test, and obtaining the final analysis result, includes:

[0042] The Sigmoid function is used to perform probability mapping on the comprehensive cognitive bias value to form the spurious probability.

[0043] The probability of forgery is compared with a preset forgery detection threshold;

[0044] If the probability of forgery is greater than or equal to the preset forgery judgment threshold, the work to be tested is judged to be suspected of being forgery; otherwise, the work to be tested is judged to be genuine.

[0045] Secondly, this invention provides an AI-based system for analyzing collections of ancient literary works, including:

[0046] The data acquisition and processing module is used to acquire a collection of real works and works to be tested, perform standardized processing, and generate standardized text.

[0047] The decoding construction module is used to construct literary semantic units after decoding the character vector sequence extracted from the standardized text, thus forming a set of literary semantic units;

[0048] The analysis and iteration module is used to construct a complete literary semantic Gaussian field based on the set of literary semantic units, perform propagation direction analysis, establish a dual matrix to perform iteration, and output the literary semantic Gaussian field.

[0049] The enhanced computation module is used to hierarchically divide and randomly sample the literary semantic Gaussian field, generate semantic sampling points and calculate semantic contribution, compress or clone the semantic sampling points, generate an enhanced semantic Gaussian field, construct a multi-directional cognitive propagation structure by expanding spherical harmonic propagation, and generate cognitive Gaussian fields of the work to be tested and the target work for Mahalanobis operation to obtain the Mahalanobis distance deviation value.

[0050] The mapping and judgment module is used to perform probability mapping on Mahalanobis distance deviation values, generate forgery probabilities to judge the authenticity of the work under test, and obtain the final analysis results.

[0051] The beneficial effects of this invention are as follows: By constructing a literary semantic Gaussian field, a unified representation of the continuous semantic propagation relationship across sentence elements and paragraphs in ancient literary works is achieved. This effectively depicts the cognitive diffusion pattern that the author has formed over a long period of time. By establishing a multi-directional cognitive propagation structure and a dynamic iterative propagation mechanism, the evolution path of imagery, the trajectory of emotional progression, and the rhetorical coupling relationship in the work can be accurately captured, significantly improving the ability to extract deep literary cognitive features. Secondly, by using the cognitive Gaussian distance analysis mechanism to comprehensively consider the covariance coupling relationship between different semantic directions, the accuracy of identifying imitation texts, partially imitated texts, and hybrid generated texts is improved. Therefore, this invention not only improves the accuracy and robustness of the authenticity analysis of ancient literature, but also has good structural interpretability. Attached Figure Description

[0052] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Fig. 1 This is a flowchart of the AI-based method for analyzing ancient literary works in Example 1.

[0054] Fig. 2 This is a structural diagram of the AI-based ancient literature collection analysis system in Example 1.

[0055] Fig. 3 This is a flowchart for determining the authenticity of the work to be tested in Example 1. Detailed Implementation

[0056] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0057] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0058] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0059] Example 1, referring to Figs. 1-3 This is the first embodiment of the present invention, which provides an AI-based method for analyzing ancient literary works, including the following steps:

[0060] S1. Obtain a collection of real works and the works to be tested, perform standardization processing, and generate standardized text.

[0061] Specifically, the user inputs a collection of the target author's real works and the works to be tested;

[0062] All text works in both the real works collection and the works to be tested are uniformly converted to UTF-8 encoding format; then, the text works in UTF-8 encoding format are uniformly standardized to obtain standardized text works, including standardized real works and standardized works to be tested.

[0063] It should be noted that when uniformly converting text to UTF-8 encoding, the system reads the original text character by character and maps each character to a UTF-8 standard byte sequence. When encountering a character that cannot be correctly mapped, the system does not directly delete the corresponding character, but first queries the ancient Chinese variant character database. If a corresponding ancient character encoding is found, extended character mapping is performed; if no corresponding mapping is found, the current character is marked as an abnormal character.

[0064] After completing the unification of UTF-8 encoding, the system further performs standardization processing on text works uniformly. The standardization processing in this embodiment includes simplification and traditionalization unification processing, variant character normalization processing, ancient and modern character mapping processing, punctuation unification processing, and whitespace character unification processing;

[0065] In the simplification and traditionalization unification processing, the system first establishes a simplification-traditionalization mapping dictionary. The simplification-traditionalization mapping dictionary includes modern simplification-traditionalization mapping relationships and ancient variant traditional character mapping relationships. For example, "國", "囯", and "国" are uniformly mapped to "国"; "學", "斈", and "学" are uniformly mapped to "学"; "說", "説" are uniformly mapped to "说";

[0066] In the variant character normalization processing, the system further establishes a corresponding relationship table for variant characters. For example, a unified mapping relationship is established between "餘" and "余", "於" and "于", and "爲" and "为". When performing mapping, the system scans the text content character by character. When a variant character is detected, it is automatically replaced with the standard glyph;

[0067] In the ancient and modern character mapping processing, the system further performs modern semantic mapping on ancient specific glyphs. For example, ancient Chinese function words such as "矣", "焉", "兮", etc. remain in their original forms unchanged, while "云" and "雲" are distinguished according to the semantic context. When "云" means "say", "云" remains unchanged; when "云" means a weather phenomenon, it is mapped to "雲";

[0068] In the punctuation unification processing, the system first deletes duplicate punctuation, redundant punctuation, and illegal punctuation introduced during modern editing. For example, for consecutive multiple full stops, multiple commas, and abnormal whitespace punctuation combinations, the system performs normalization processing. At the same time, the system retains the original ancient literary sentence-breaking structure and does not force the introduction of modern long sentence segmentation methods;

[0069] In the whitespace character unification processing, the system deletes redundant spaces, tab characters, repeated line breaks, and invisible control characters in the text. For example, when consecutive multiple spaces are detected, they are uniformly compressed into a single space; when consecutive multiple line break characters are detected, they are uniformly compressed into a single line break.

[0070] S2. After decoding by extracting the character vector sequence based on the standardized text, construct literary semantic units to form a set of literary semantic units.

[0071] S2.1. Read all the characters of the standardized text, construct a character sequence in the original character order, and then obtain the embedding vectors of the characters for combination to form a character vector sequence.

[0072] Specifically, based on the standardized text work, all characters in the standardized text work are read, and a character sequence is constructed in the original character order; the embedding vectors of each character in the character sequence are obtained through a pre-trained encoding model, and the embedding vectors are combined in the original character order to form a character vector sequence;

[0073] It should be noted that: the pre-trained encoding model preferably uses the pre-trained bert-ancient-chinese model, and the training corpus preferably uses the digitalized texts of ancient books of past dynasties.

[0074] S2.2. Extract and fuse the context semantic information of the character vector sequence to generate a fused context semantic state vector sequence.

[0075] Specifically, use the BiLSTM network to extract and fuse the context semantic information of the character vector sequence from the forward and backward directions respectively to obtain a fused context semantic state vector sequence;

[0076] It should be noted that: the BiLSTM network can be set as a bidirectional long short-term memory network. When specifically executed, the forward LSTM network propagates the semantic state character by character in the original character order to generate the corresponding forward hidden state; subsequently, the system uses the backward LSTM network to propagate the semantic state forward from the last character of the text, thereby gradually establishing a backward semantic dependence structure. The forward context semantic state and the backward context semantic state are concatenated to form a complete context semantic representation.

[0077] S2.3. Decode the fused context semantic state vector sequence to form a set of sentence elements.

[0078] Specifically, decode the fused context semantic state vector sequence through a conditional random field to obtain the optimal sentence breaking boundary sequence and form a set of sentence elements.

[0079] It should be noted that: when decoding the fused context semantic state vector sequence through a conditional random field, a corresponding boundary label set should be established for each character in the character sequence first. The boundary label set includes:

[0080] B label, M label, E label, S label.

[0081] Among them: the B label is used to indicate that the current character is the starting character of a sentence element, the M label is used to indicate that the current character is in the middle of a sentence element, the E label is used to indicate that the current character is the ending character of a sentence element, and the S label is used to indicate that the current character forms an independent sentence element alone;

[0082] For example, for the sentence element "Raise the cup to invite the bright moon", its label sequence is marked as: "Raise / B, cup / M, invite / M, bright / M, moon / E". And for the independent mood word "Xi", it is marked as: "Xi / S";

[0083] After establishing the boundary label set, the system further calculates the emission probability of each character corresponding to different boundary labels according to the fused context semantic state vector. Subsequently, the fused context semantic state vector corresponding to the current character is read, and the probability values of the current character belonging to each boundary label are generated through a linear mapping layer;

[0084] After calculating the character boundary emission probability, the system further establishes the label transition constraint relationship.

[0085] It should be noted that the label transition relationship in this embodiment is not established randomly, but is predefined based on the sentence-breaking rules of ancient Chinese. Specifically, it includes: after the B label, it is allowed to transition to the M label or the E label; after the M label, it is allowed to transition to the M label or the E label; after the E label, it is allowed to transition to the B label or the S label; after the S label, it is allowed to transition to the B label or the S label. At the same time: the system prohibits the following illegal transitions:

[0086] Direct transition from the E label to the M label, direct transition from the B label to the S label, and direct transition from the M label to the B label.

[0087] After establishing the label transition relationship, the system further performs joint path scoring on the entire character sequence.

[0088] During the scoring process, the system not only considers the emission probability of each character corresponding to the boundary label, but also considers the transition probability between adjacent labels.

[0089] For example: when the character "月" is marked as the E label: the system not only calculates the probability that "月" belongs to the E label itself; but also further calculates whether the label of the previous character allows transition to the E label.

[0090] If the previous character is the M label, the system increases the total score of the current path; if the previous character is the S label, the system decreases the total score of the current path.

[0091] Subsequently, the system further uses the Viterbi decoding algorithm to perform dynamic programming search on all label paths.

[0092] During the specific search process: the system starts from the starting position of the character sequence and calculates character by character: the optimal cumulative path score of the current character in all possible label states. Subsequently: the system further combines the optimal path result of the previous character and retains the label path with the highest cumulative score. After completing the traversal of all characters: the system backtracks the optimal path from the last character to obtain the optimal boundary label sequence of the entire character sequence.

[0093] The sentence-breaking boundary is to use the corresponding character as the end position of the current sentence element and start constructing a new sentence element from the next character;

[0094] Suppose that when the system detects the S tag: the current character is directly output as an independent sentence element. Finally: the system forms a set of sentence elements, for example, "Raising my cup, I invite the bright moon", "With my shadow we're three on this night".

[0095] S2.4. Extract five types of literary semantic units from the set of sentence elements to form a set of literary semantic units.

[0096] Specifically, based on the set of sentence elements, five types of literary semantic units are extracted, including image semantic units, rhetorical semantic units, emotional semantic units, allusion semantic units, and rhythm semantic units;

[0097] Combine all literary semantic units to form a set of literary semantic units.

[0098] It should be noted that: the author's cognitive communication laws in ancient literature are mainly reflected in five aspects: image diffusion, rhetorical structure, emotional communication, historical allusions, and rhythm. Therefore, the five types of literary semantic units can ensure the stability of authenticity analysis;

[0099] When extracting image semantic units, the system first identifies natural images, character images, and spatial images based on the ancient literature image dictionary.

[0100] When extracting rhetorical semantic units, first identify metaphor, intertextuality, hyperbole, metonymy, and antithesis structures. For example, when detecting "My white hair is three thousand zhang long", it is identified as a hyperbole rhetorical device; when the system detects "The moon in Qin time and the pass in Han time", it is identified as having an intertextuality structure;

[0101] When extracting emotional semantic units, the system first identifies emotional carrier words such as sadness, joy, worry, sorrow, loneliness, resentment, etc. based on the ancient literature emotion dictionary, and further judges the overall emotional tendency in combination with the context propagation direction. For example, words such as "alone", "cold", "lonely", "night" will form a sad propagation structure in most cases, so the system marks the corresponding sentence element as a negative emotional propagation unit;

[0102] When extracting allusion semantic units, the system performs allusion matching based on the ancient literature knowledge graph. For example, when the system detects words such as "Chibi", "Changting", "Yangguan", etc., it further determines whether they belong to the historical allusion propagation structure in combination with the context semantics. If the current sentence element forms a stable association with historical figures, historical events, or classic literary images, the system establishes the corresponding allusion propagation node;

[0103] When extracting rhythm semantic units, the system first performs level and oblique tone analysis and rhyme analysis. Specifically, the system can identify the corresponding level and oblique tone attributes of each character based on the Middle Ancient Rhythm Database and analyze the rhyme relationship between sentence elements. For example, in a five-character regulated poem, analyze whether the last characters of even-numbered sentences belong to the same rhyme group and record the corresponding rhythm propagation structure.

[0104] S3. Based on the set of literary semantic units, construct a complete literary semantic Gaussian field, perform propagation direction analysis, establish a dual matrix to perform iteration, and output the literary semantic Gaussian field.

[0105] S3.1 Encode the literary semantic units in the set of literary semantic units, construct literary Gaussian units, and establish continuous semantic propagation connections between literary Gaussian units to form a complete literary semantic Gaussian field.

[0106] Specifically, the Transformer encoding model is used to encode each literary semantic unit in the set of literary semantic units into a semantic feature vector; then, literary Gaussian units are constructed based on the semantic feature vectors.

[0107] Establish continuous semantic propagation connections between literary Gaussian units based on contextual relationships to form a complete literary semantic Gaussian field.

[0108] It should be noted that the Transformer encoding model preferably adopts a multi-head self-attention structure;

[0109] The semantic feature vectors are not directly used for subsequent authenticity analysis. Instead, they are first input into a Gaussian parameter generation network to generate corresponding literary Gaussian units.

[0110] In this embodiment, the literary Gaussian unit includes: a semantic center, a semantic diffusion direction, a semantic density, and a propagation direction parameter. When constructing the semantic center, the system first performs linear mapping on the semantic feature vector through a fully connected layer to generate the semantic center vector; subsequently, it further constructs the semantic diffusion direction.

[0111] It should be noted that the semantic diffusion direction in this embodiment is not generated by random initialization, but is dynamically generated based on the context semantic propagation relationship;

[0112] In the specific implementation process, the system first reads the contextual adjacent semantic units of the current literary semantic unit in the original text; then, the system calculates the co-occurrence propagation relationship between the current semantic unit and its adjacent semantic units. After completing the contextual propagation statistics, the system further constructs a semantic propagation direction vector.

[0113] Specifically, the system performs a weighted accumulation of the propagation difference vectors between the current semantic unit and all adjacent semantic units. Then, the system performs a weighted average of each propagation difference based on its propagation frequency. This ultimately forms the main propagation direction of the current literary semantic unit.

[0114] It should also be noted that the direction of propagation does not represent a fixed geometric direction; rather, it represents the main semantic diffusion trend of current literary semantics in a high-dimensional cognitive space. Furthermore, establishing semantic density parameters is not simply equivalent to word frequency.

[0115] In the specific implementation process: The system first counts the number of times the current literary semantic unit appears in the target author's collection of works. Then, the system further counts the degree of clustering of the current literary semantic unit across different sentence elements. Finally, the system calculates the centrality of the current literary semantic unit within the overall propagation network.

[0116] For example, if a large number of other semantic units propagate to "bright moon", then its semantic density is increased. At the same time, the system further reads the attention weights in the Transformer attention matrix.

[0117] For example, if multiple semantic units show high attention to "bright moon," it indicates that "bright moon" is a core semantic node in the current literary dissemination structure. Subsequently, the system integrates frequency of occurrence, clustering degree, network centrality, and attention intensity to generate the final semantic density parameters.

[0118] After completing the semantic density construction, the system further establishes the directional propagation parameters.

[0119] It should be noted that the directional propagation parameter describes the ability of current literary semantics to propagate and change across different cognitive directions. For example, the phrase "bright moon" in Li Bai's works may primarily propagate towards romanticism, loneliness, and drinking. However, in Du Fu's works, it propagates more towards homesickness and patriotism. Therefore, it is necessary to establish directional propagation parameters. In the specific implementation process, the system first establishes a fixed set of semantic directions. This set includes emotional direction, rhetorical direction, historical direction, imagery direction, and rhythmic direction.

[0120] Subsequently, the system statistically analyzes the propagation intensity of the current literary semantic unit in each direction.

[0121] For example, regarding "bright moon," the system calculates: the intensity of its propagation towards emotion; the intensity of its propagation towards imagery; and the intensity of its propagation towards history.

[0122] Subsequently, the system further analyzed the rate of change of the current literary semantics in different directions.

[0123] For example, if the image of "bright moon" continuously spreads in the direction of "homesickness" across different works, then the corresponding directional propagation parameter should be increased. Ultimately, the system generates a set of directional propagation parameters corresponding to the current literary semantic unit.

[0124] After constructing the semantic center, semantic diffusion direction, semantic density, and directional propagation parameters, the system combines them to generate complete literary Gaussian units.

[0125] Subsequently, continuous semantic propagation connections are established between literary Gaussian units based on contextual relationships. When establishing these connections, the system first reads the adjacent Gaussian units before and after the current literary Gaussian unit. Then, the system calculates the common propagation frequency between different Gaussian units and further calculates the propagation distance. The propagation weights are updated based on the consistency of the propagation direction, ultimately forming a complete literary semantic Gaussian field.

[0126] S3.2 Generate the propagation direction for the literary Gaussian unit and establish a dual matrix, including a rotation matrix and a diffusion scale matrix.

[0127] Specifically, based on the complete literary semantic Gaussian field, the propagation direction is analyzed for each literary Gaussian unit, the propagation direction is generated, and the rotation matrix and diffusion scale matrix are established;

[0128] It should be noted that when performing propagation direction analysis for each literary Gaussian unit, a local propagation neighborhood is first established for each literary Gaussian unit.

[0129] For example, in one implementation, the system uses the current literary Gaussian unit as the central node and searches for neighboring literary Gaussian units with contextual propagation relationships within the literary semantic Gaussian field. Subsequently, it counts the number of times the current literary Gaussian unit co-occurs with other literary Gaussian units in the target author's actual works collection, and further analyzes its propagation stability within sentences, paragraphs, and cross-paragraph structures.

[0130] Among them, the intra-sentence propagation stability is used to represent the stable diffusion of the current literary semantics within a single sentence; the intra-paragraph propagation stability is used to represent the diffusion pattern of the current literary semantics in a local context; and the cross-paragraph propagation stability is used to represent the long-distance diffusion ability of the current literary semantics in the entire literary work.

[0131] Local propagation weights are generated based on the three types of propagation stability, and a propagation neighborhood is established based on the propagation weights;

[0132] After the propagation neighborhood is established, the system begins to perform propagation direction analysis. Specifically, the system first performs directional clustering analysis on the current literary Gaussian unit and all literary Gaussian units in its propagation neighborhood. In this embodiment, the system consistently uses a "directional recursive clustering method" to perform propagation direction analysis, that is, starting from the current literary Gaussian unit, recursively searching for the main propagation paths in its propagation neighborhood, performing directional statistics on all propagation paths, and selecting the path with the most propagation as the main propagation direction.

[0133] When establishing the rotation matrix, direction axes are created based on the set of propagation directions, and different propagation directions are mapped to different propagation dimensions in the high-dimensional semantic space. For example, the emotion direction corresponds to the first propagation dimension, the rhetoric direction to the second propagation dimension, the historical direction to the third propagation dimension, the imagery direction to the fourth propagation dimension, and the rhythm direction to the fifth propagation dimension. Subsequently, the system generates the direction rotation relationship based on the diffusion intensity of the current literary Gaussian unit in each propagation direction.

[0134] In practice, the system first calculates the semantic change rate as the current literary Gaussian unit diffuses in each propagation direction. Then, it establishes directional rotation relationships based on the change rates in different directions.

[0135] In the specific implementation process, the propagation span of the current literary Gaussian unit in different propagation directions is first statistically analyzed. Based on the propagation span, a diffusion scale corresponding to the propagation direction is established, thus forming a diffusion scale matrix.

[0136] S3.3. Based on the rotation matrix and diffusion scale matrix, generate a semantic covariance matrix for the literary Gaussian unit, and update the propagation path according to the propagation overlap relationship between units to iterate and output a stable literary semantic Gaussian field of author cognition.

[0137] Specifically, based on the rotation matrix and the diffusion scale matrix, the semantic covariance matrix of each literary Gaussian unit is generated as an anisotropic semantic propagation structure. Subsequently, the propagation path is updated according to the propagation overlap relationship between different literary Gaussian units, and the propagation structure is iterated. When the maximum number of iterations is reached, a stable literary semantic Gaussian field of author cognition is output.

[0138] It should be noted that the semantic covariance matrix of each literary Gaussian unit is generated as an anisotropic semantic propagation structure by establishing the main propagation direction based on the rotation matrix, establishing the diffusion range in different directions based on the diffusion scale matrix, and then jointly fusing the two to generate the final semantic covariance structure.

[0139] In this embodiment, the semantic covariance matrix is ​​not a fixed matrix, but a dynamic propagation matrix. That is, if the propagation neighborhood of the current literary Gaussian unit changes, the semantic covariance structure will also be updated synchronously.

[0140] Secondly, the propagation overlap relationship is used to describe the degree of cross-propagation between different literary Gaussian units. In the specific implementation process, the number of common diffusions of different literary Gaussian units in each propagation direction is first counted, and their common propagation stability in different contexts is further counted. Subsequently, if the degree of propagation overlap between two literary Gaussian units is higher than a preset threshold, the system strengthens the propagation connection between them; if the degree of propagation overlap is lower than the preset threshold, the system reduces their propagation connection weight.

[0141] It should be noted that the preset threshold is set by performing statistical distribution analysis on the set of propagation overlap values ​​and calculating the mean and standard deviation of propagation overlap.

[0142] A preferred approach is to use the mean of propagation overlap minus 0.5 times the standard deviation of propagation overlap to generate the preset threshold.

[0143] S4. The literary semantic Gaussian field is hierarchically divided and randomly sampled to generate semantic sampling points and calculate semantic contribution. The semantic sampling points are compressed or cloned to generate an enhanced semantic Gaussian field. A multi-directional cognitive propagation structure is constructed by expanding the spherical harmonic propagation and generating cognitive Gaussian fields of the work to be tested and the target work. The Mahalanobis operation is performed to obtain the Mahalanobis distance deviation value.

[0144] Specifically, based on the stable author's cognition of the literary semantic Gaussian field, after obtaining the semantic density, rhetorical complexity, emotional fluctuation rate and allusion dissemination rate in the field, the entire literary semantic Gaussian field is hierarchically divided to form semantic layers, including the general semantic layer, the rhetorical semantic layer, the emotional semantic layer and the allusion semantic layer.

[0145] It should be noted that semantic density is the number of effective propagation connections in a statistical local region (such as the high-frequency coupling between "bright moon" and "homesickness").

[0146] Rhetorical complexity: the nesting depth and propagation span between statistical rhetorical structures (such as continuous nesting of metonymy and hyperbole).

[0147] Emotional volatility: This refers to the number of emotional changes, the magnitude of those changes, and the frequency of directional shifts along the emotional transmission path.

[0148] Allusion propagation rate: Statistical analysis of the propagation span, depth, and number of branches of allusion semantics across multiple sentence elements;

[0149] Based on the four types of statistical results, semantic clustering is performed on all literary Gaussian units in the entire literary semantic Gaussian field.

[0150] Ideally, K-Means clustering algorithm should be used for semantic clustering, with a fixed four-layer semantic structure, such as ordinary semantic layer, rhetorical semantic layer, emotional semantic layer and allusion semantic layer.

[0151] Furthermore, in each semantic layer, a semantic propagation path is generated based on the propagation connection relationship between literary Gaussian units, and a semantic sampling region is randomly selected in the propagation path for random sampling to generate several semantic sampling points.

[0152] It should be noted that: semantic propagation paths are generated based on the propagation connections between literary Gaussian units, and semantic sampling regions are randomly selected within these paths. Specifically, a semantic propagation network is established based on the propagation connections between literary Gaussian units. In this network, each literary Gaussian unit corresponds to a semantic propagation node, and the propagation edges between nodes represent stable contextual propagation relationships between two literary Gaussian units. After establishing the semantic propagation network, the propagation region is gradually expanded along the propagation edges, starting from the high-propagation-density nodes in the current semantic layer, thus forming a complete emotion propagation path.

[0153] It should be noted that random sampling is not completely random, but rather random region sampling constrained by the propagation path. Specifically, a starting node is first randomly selected in the propagation path, and then a fixed-length propagation region is extended forward or backward according to a preset propagation span, and this propagation region is used as the semantic sampling region;

[0154] It should be noted that semantic sampling points are not individual character positions, but rather the locations of the propagation centers within the entire continuous semantic propagation region. Each semantic sampling point corresponds to a complete local cognitive propagation structure.

[0155] S4.1 Calculate the semantic contribution of each semantic sampling region based on the semantic sampling points.

[0156] It should be noted that the calculation of semantic contribution requires first normalizing the semantic density, rhetorical complexity, emotional fluctuation rate, and allusion propagation rate in the current semantic sampling area to obtain normalized values ​​for each parameter; then, the normalized values ​​are weighted and fused using preset weights (i.e., the normalized values ​​are multiplied by the corresponding weights and the products are added together) to form the semantic contribution.

[0157] It should be noted that the preset weights include semantic density weight, rhetorical complexity weight, emotional fluctuation rate weight, and allusion dissemination rate weight.

[0158] Preferably, the semantic density weight can be set to 0.15, the rhetorical complexity weight to 0.35, the emotion fluctuation rate weight to 0.25, and the allusion propagation rate weight to 0.25. With these weight settings, the present invention achieves a authenticity analysis structure that is "centered on the rhetorical propagation structure, supplemented by emotion propagation and allusion propagation, and using semantic density for regional localization." Furthermore, this structure effectively enhances the author's ability to identify long-term cognitive propagation structures, detect deep propagation anomalies, and analyze long-distance semantic diffusion.

[0159] S4.2. Based on the preset contribution threshold and the preset judgment threshold, determine whether the semantic region is a high-complexity semantic region or a low-complexity semantic region. Under the result of the high-complexity semantic region, perform cloning processing on the semantic nodes, and under the result of the low-complexity semantic region, perform compression processing on the semantic nodes.

[0160] Specifically, the semantic contribution is compared with a preset contribution threshold;

[0161] If the semantic contribution is greater than the preset contribution threshold, the current semantic region is determined to be a high semantic contribution region; if the semantic contribution is less than or equal to the preset contribution threshold, the current semantic region is determined to be a low semantic contribution region and is frozen.

[0162] Based on the regions with high semantic contribution, the corresponding literary Gaussian units are obtained, and the semantic change rate of the literary Gaussian units is calculated (i.e., the semantic contribution is subtracted).

[0163] The magnitude of the semantic change rate is compared with a preset judgment threshold.

[0164] If the magnitude of the semantic change rate is greater than the preset judgment threshold, the current semantic region is determined to be a high-complexity semantic region, and semantic node cloning is performed. If the magnitude of the semantic change rate is less than or equal to the preset judgment threshold, the current semantic region is determined to be a low-complexity semantic region, and semantic node compression is performed.

[0165] An enhanced literary Gaussian field is generated based on semantic node cloning or semantic node compression.

[0166] It should be noted that the preset contribution threshold is dynamically generated based on the overall semantic propagation distribution in the target author's actual works collection;

[0167] Preferably, the contribution distribution of all semantic sampling regions in the target author's authentic works is statistically analyzed, and then contribution intervals are established based on the contribution distribution results. In this embodiment, a percentile interval method is preferably used to determine the contribution threshold, that is, all semantic contributions are sorted from high to low, and the contribution boundary corresponding to the top 40% is selected as the preset contribution threshold. This ensures the stability of the authenticity analysis among different authors;

[0168] Secondly, the preset judgment threshold is also dynamically generated. Specifically, the system first statistically analyzes the semantic change rate distribution of all high semantic contribution regions in the target author's actual works, and then generates a complexity judgment interval based on the change rate distribution results. In this embodiment, the judgment threshold is preferably generated using a median offset method, that is, a fixed offset is added to the median of all change rates to avoid ordinary semantic fluctuations being misidentified as high complexity regions.

[0169] In this embodiment, the fixed offset is preferably generated using a "standard deviation proportional offset method". For example, firstly, the standard deviation of all semantic change rate moduli is calculated to represent the overall fluctuation of semantic change rates in the target author's actual works. Subsequently, the system multiplies the standard deviation by a fixed proportional coefficient to generate the final offset.

[0170] Preferably, the fixed ratio coefficient is between 0.3 and 0.5, and more preferably 0.4. At this value, if the ratio coefficient is lower than 0.3, the offset is too small, and a large number of natural fluctuations in ordinary semantic regions will be misidentified as high-complexity regions, resulting in a significant increase in the number of node clones and excessive complexity of the propagation network.

[0171] If the proportional coefficient is higher than 0.5, the offset is too large, and some real high-complexity regions may not reach the complexity judgment threshold, thus making it impossible to effectively extract the stable features of the real author in the complex cognitive propagation structure.

[0172] Therefore, setting the scaling factor to 0.4 achieves a better balance between complexity recognition accuracy and propagation structure stability.

[0173] It should also be noted that the semantic node cloning in this embodiment is not a simple copying of literary Gaussian units, but rather a local propagation decomposition of the current complex semantic propagation structure. Specifically, the system first identifies the core propagation nodes in the current high-complexity region. For example, in the propagation chain of "bright moon—hometown—loneliness—Chang'an," the system may identify the "hometown" node as simultaneously undertaking the functions of emotional propagation and historical propagation, thus determining it as a core propagation node. Subsequently, the system performs propagation decomposition on the core propagation nodes based on different propagation directions.

[0174] After cloning nodes, the system further re-establishes the propagation connections between the cloned nodes. For example, a single propagation node that was originally connected to multiple propagation directions can now be connected to the corresponding propagation paths after cloning, thus achieving fine-grained separation of complex cognitive propagation structures.

[0175] Conversely, in this embodiment, semantic node compression processing is performed on low-complexity semantic regions. Specifically, the system first identifies continuous low-change propagation nodes in the current region. For example, in a general landscape description region, images such as "mountain," "water," and "wind" may propagate continuously and stably without significant rhetorical changes or emotional fluctuations. Therefore, the system classifies such continuous propagation nodes as compressible propagation structures. Subsequently, the system performs semantic fusion processing on continuous low-change propagation nodes. This fusion processing is not a simple deletion of nodes, but rather an aggregation of the propagation attributes (e.g., propagation direction, semantic center, and propagation density) of multiple low-complexity nodes to form new compressed propagation nodes.

[0176] Furthermore, for the enhanced literary semantic Gaussian field, a fixed semantic direction set is established, which includes emotional direction, rhetorical direction, historical direction, rhythmic direction and imagery direction;

[0177] Based on each fixed semantic direction, the enhanced literary Gaussian unit is subjected to spherical harmonic propagation expansion processing to establish the semantic propagation function in the corresponding direction;

[0178] Perform directional propagation modeling based on each semantic propagation function to obtain directional semantic propagation results corresponding to different fixed semantic directions;

[0179] Calculate the directional coupling propagation strength between different directions based on the directional semantic propagation results corresponding to different fixed semantic directions;

[0180] Based on the directional coupling propagation strength and directional semantic propagation results, the multi-directional cognitive propagation structure is calculated.

[0181] The semantic propagation function is expressed as follows:

[0182]

[0183] In the formula, Indicates the first An enhanced literary Gaussian unit in the current fixed semantic direction The direction of semantic propagation function value, This represents the order of the maximum spherical harmonic expansion. Represents the spherical harmonic order. Represents the spherical harmonic dimension index. This represents the direction propagation coefficient of the corresponding spherical harmonic basis function. Represents spherical harmonic basis functions. Indicates the current fixed semantic direction;

[0184] The expression for obtaining the directional semantic propagation results corresponding to different fixed semantic directions is:

[0185]

[0186] In the formula, Indicates the current fixed semantic direction directional semantic propagation results This represents the total number of Gaussian units in the enhanced literature. Indicates the first The propagation weight of an enhanced Gaussian unit for literature. Indicates the first An enhanced literary Gaussian unit;

[0187] The expression for calculating the directional coupling propagation strength between different directions is:

[0188]

[0189] In the formula, Indicates fixed semantic direction With fixed semantic direction The directional coupling propagation strength between them Represents the conditional propagation probability. Indicates fixed semantic direction directional semantic propagation results Indicates fixed semantic direction The result of directional semantic propagation;

[0190] The multidirectional cognitive transmission structure is expressed as follows:

[0191]

[0192] In the formula, This represents a multi-directional cognitive transmission structure. Indicates the number of fixed semantic directions. This indicates the transpose operation.

[0193] It should be noted that after establishing a fixed set of semantic directions, the system does not directly perform directional propagation modeling. Instead, it first extracts propagation features corresponding to different fixed semantic directions based on the enhanced literary semantic Gaussian field. For example, in the emotion direction, the system extracts the intensity of emotion progression, the rate of emotion decay, the span of emotion diffusion, and the continuity of emotion propagation. In the rhetoric direction, the system extracts the depth of rhetorical nesting, the level of rhetorical recursion, the strength of rhetorical coupling, and the frequency of rhetorical jumps. In the historical direction, the system extracts the stability of allusion associations, the span of historical event propagation, the degree of character propagation coupling, and the strength of cross-dynasty allusion associations. In the phonetics direction, the system extracts the continuity of tonal propagation, the stability of rhyme propagation, the progressive rules of meter, and the frequency of rhythmic changes. In the imagery direction, the system extracts the intensity of core image diffusion, the degree of image aggregation, the span of image propagation, and the stability of image progression.

[0194] The directional propagation coefficients are five, corresponding to fixed semantic directions;

[0195] Ideally, the propagation coefficients for the emotion direction are set to 0.30, rhetoric direction to 0.25, history direction to 0.15, rhythm direction to 0.10, and imagery direction to 0.20. This ensures that the propagation results in different directions maintain a stable proportional relationship. For example, when the propagation result in the emotion direction is abnormally enhanced, since the emotion direction propagation coefficient is fixed at 0.30, the system will automatically limit the proportion of the emotion direction propagation in the overall cognitive propagation structure, thereby preventing the emotion direction from overshadowing other propagation directions.

[0196] It should also be noted that the spherical harmonic basis functions are fixedly complex spherical harmonic functions. In specific implementation, the fixed semantic direction is mapped to a unit sphere. In the unit sphere, the polar angle is used to describe the intensity of the propagation direction, and the azimuth angle is used to describe the direction of propagation diffusion. This allows literary propagation to unfold through spherical harmonic functions. Subsequently, the directional propagation coefficient is combined with the spherical harmonic basis functions to generate the directional semantic propagation function. ;

[0197] The spherical harmonic basis functions are expressed as follows:

[0198]

[0199] In the formula, Represents the values ​​of spherical harmonic basis functions. Indicates the polar angle. Indicates the direction angle. This represents the associated Legendre polynomial. Represents the cosine function. Indicates the direction term of the complex exponent;

[0200] Secondly, this embodiment preferably employs a low-order stable spherical harmonic expansion mechanism. Specifically, a third-order directional propagation expansion structure is fixed. The reason for using a third-order expansion is that the long-term cognitive propagation patterns in ancient literature are usually concentrated in low-order stable directional propagation. If a higher-order propagation expansion is used, it is easy to introduce a large number of local semantic oscillations, thereby reducing the stability of the authenticity analysis.

[0201] It should be further explained that if the propagation weight is too small, it is considered that the current literary Gaussian unit's propagation contribution is insufficient; if the propagation weight is too small, it is considered that there is abnormal clustering in local propagation.

[0202] Preferably, the propagation weight can be set to a value range of 0.02 to 0.15, and the present invention can take 0.15 as the default value; thereby ensuring that the calculation process of propagation results in different fixed semantic directions is repeatable and interpretable.

[0203] It should be further explained that the conditional propagation probability is obtained statistically based on the co-propagation relationship between fixed semantic directions in the enhanced literary semantic Gaussian field; for example, the number of propagation activations occurring in fixed semantic directions is counted, and the fixed semantic directions are statistically analyzed. With fixed semantic direction The number of consecutive activations within the same semantic propagation path, the same semantic sampling region, or adjacent literary Gaussian units is counted; subsequently, the conditional propagation probability is generated by the ratio between the number of consecutive activations and the number of propagation activations.

[0204] Furthermore, based on the multi-directional cognitive propagation structure, Gaussian center, Gaussian covariance matrix, and cognitive propagation density are generated;

[0205] Based on Gaussian center, Gaussian covariance and cognitive propagation density, generate cognitive Gaussian field of the work to be tested and cognitive Gaussian field of the target work.

[0206] Based on the cognitive Gaussian field of the work under test and the cognitive Gaussian field of the target work, the Mahalanobis distance deviation value between the cognitive Gaussian fields is obtained.

[0207] The Gaussian center is expressed as:

[0208]

[0209] In the formula, This indicates the Gaussian center of the work to be tested. Represents the Gaussian central mapping matrix. Represents the Gaussian center bias term;

[0210] The Gaussian covariance matrix is ​​expressed as:

[0211]

[0212] In the formula, This represents the Gaussian covariance matrix of the work to be tested. Represents the covariance mapping matrix. This represents the transpose of the covariance mapping matrix. This represents the transpose matrix of the multidirectional cognitive propagation structure;

[0213] Cognitive propagation density, expressed as:

[0214]

[0215] In the formula, This indicates the cognitive dissemination density of the work being tested. Indicates the magnitude of the vector;

[0216] Generate the cognitive Gaussian field of the work to be tested and the cognitive Gaussian field of the target work, expressed as follows:

[0217]

[0218]

[0219] In the formula, This represents the cognitive Gaussian field of the work to be tested. This represents the cognitive Gaussian field of the target work. Indicates the Gaussian center of the target work. This represents the Gaussian covariance matrix of the target work. Indicates the cognitive dissemination density of the target work;

[0220] The Mahalanobis distance deviation value is expressed as follows:

[0221]

[0222] In the formula, This represents the Mahalanobis distance deviation value. Let represent the Gaussian covariance inverse matrix.

[0223] It should be noted that: Gaussian central mapping matrix Gaussian center bias term This is achieved through linear regression learning on all semantic Gaussian fields in the target work set, ensuring the comparability of cognitive centers between different works. For example, if a work is biased towards the emotional direction in a multi-directional propagation structure, its Gaussian center will shift towards the emotional propagation axis;

[0224] And the covariance mapping matrix It can be calculated that the target work , The mean matrix of the target work and the mean matrix of all target works are obtained, and then generalized eigenvalue decomposition is performed on the two mean matrices to generate the initial covariance mapping matrix. .

[0225] S5. Perform probability mapping on the Mahalanobis distance deviation value to generate the forgery probability and judge the authenticity of the work to be tested, and obtain the final analysis result.

[0226] Specifically, the Sigmoid function is used to perform probability mapping on the comprehensive cognitive bias value to form the spurious probability;

[0227] The probability of forgery is compared with a preset forgery detection threshold;

[0228] If the probability of forgery is greater than or equal to the preset forgery judgment threshold, the work to be tested is judged to be suspected of being forgery; otherwise, the work to be tested is judged to be genuine.

[0229] The probability mapping is expressed as:

[0230]

[0231] In the formula, This represents the probability of a fake. Represents the natural constant. This represents the probability scaling factor. This indicates the deviation from the truth value.

[0232] It should be noted that the authenticity deviation value is essentially the statistical upper bound of the Mahalanobis distance within the target work;

[0233] A better approach is to set the distance to the 95th percentile of the Mahalanobis distance between real works.

[0234] The preset threshold for judging forgery is dynamically set based on the probability distribution of forgery between pairs of works in the real collection.

[0235] For example, in one implementation, all target works are paired up to calculate the probability of forgery, and the 95th percentile is used as a threshold. If the probability of forgery between the test work and the target works exceeds this threshold, it is judged as a suspected forgery.

[0236] This embodiment also provides an AI-based system for analyzing collections of ancient literary works, including:

[0237] The data acquisition and processing module is used to acquire a collection of real works and works to be tested, perform standardized processing, and generate standardized text.

[0238] The decoding construction module is used to construct literary semantic units after decoding the character vector sequence extracted from the standardized text, thus forming a set of literary semantic units;

[0239] The analysis and iteration module is used to construct a complete literary semantic Gaussian field based on the set of literary semantic units, perform propagation direction analysis, establish a dual matrix to perform iteration, and output the literary semantic Gaussian field.

[0240] The enhanced computation module is used to hierarchically divide and randomly sample the literary semantic Gaussian field, generate semantic sampling points and calculate semantic contribution, compress or clone the semantic sampling points, generate an enhanced semantic Gaussian field, construct a multi-directional cognitive propagation structure by expanding spherical harmonic propagation, and generate cognitive Gaussian fields of the work to be tested and the target work for Mahalanobis operation to obtain the Mahalanobis distance deviation value.

[0241] The mapping and judgment module is used to perform probability mapping on Mahalanobis distance deviation values, generate forgery probabilities to judge the authenticity of the work under test, and obtain the final analysis results.

[0242] In summary, by constructing a literary semantic Gaussian field, a unified representation of the continuous semantic propagation relationship across sentences and paragraphs in ancient literary works is achieved. This effectively depicts the long-term stable cognitive diffusion pattern formed by the author. By establishing a multi-directional cognitive propagation structure and a dynamic iterative propagation mechanism, the evolution path of imagery, the trajectory of emotional progression, and the rhetorical coupling relationship in the work can be accurately captured, significantly improving the ability to extract deep literary cognitive features. Secondly, by using the cognitive Gaussian distance analysis mechanism to comprehensively consider the covariance coupling relationship between different semantic directions, the recognition accuracy of imitation texts, partially imitated texts, and hybrid generated texts is improved. Therefore, this invention not only improves the accuracy and robustness of ancient literary authenticity analysis but also has good structural interpretability.

[0243] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An AI-based method for analyzing ancient literary works, characterized by: include, Obtain a collection of real works and the works to be tested, perform standardization processing, and generate standardized text; After decoding the character vector sequence extracted from the standardized text, literary semantic units are constructed, forming a set of literary semantic units; Based on the set of literary semantic units, after constructing a complete literary semantic Gaussian field and analyzing the propagation direction, a dual matrix is ​​established to perform iteration and output the literary semantic Gaussian field. The semantic Gaussian field of literature is hierarchically divided and randomly sampled to generate semantic sampling points and calculate semantic contribution. The semantic sampling points are compressed or cloned to generate an enhanced semantic Gaussian field. A multi-directional cognitive propagation structure is constructed by expanding spherical harmonic propagation, and the cognitive Gaussian fields of the work to be tested and the target work are generated. The Mahalanobis operation is performed to obtain the Mahalanobis distance deviation value. By performing probability mapping on the Mahalanobis distance deviation value, the probability of forgery is generated to judge the authenticity of the work under test, and the final analysis result is obtained.

2. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: The process of obtaining a set of real works and the works to be tested, and then performing standardized processing to generate standardized text includes: By having the user input a collection of the target author's real works and the works to be tested; All text works in both the real works collection and the works to be tested are uniformly converted to UTF-8 encoding format; the text works in UTF-8 encoding format are uniformly standardized to obtain standardized text works, including standardized real works and standardized works to be tested.

3. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: After decoding the character vector sequence extracted from the standardized text, literary semantic units are constructed to form a set of literary semantic units, as detailed below: Read all the characters of the standardized text, construct a character sequence according to the original character order, obtain the character embedding vectors, and combine them to form a character vector sequence; Extract the contextual semantic information of the character vector sequence and fuse it to generate a fused contextual semantic state vector sequence; The sequence of fused context semantic state vectors is decoded to form a set of sentence elements; Five types of literary semantic units are extracted from the sentence element set to form a set of literary semantic units.

4. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: After constructing a complete literary semantic Gaussian field based on the set of literary semantic units and performing propagation direction analysis, a dual-matrix iterative process is established to output the literary semantic Gaussian field, as detailed below: After encoding the literary semantic units in the set of literary semantic units and constructing literary Gaussian units, a continuous semantic propagation connection is established between the literary Gaussian units to form a complete literary semantic Gaussian field. Generate propagation direction for literary Gaussian units and establish a dual matrix, including a rotation matrix and a diffusion scale matrix; Based on the rotation matrix and the diffusion scale matrix, a semantic covariance matrix is ​​generated for the literary Gaussian unit. The propagation path is updated iteratively according to the propagation overlap relationship between units, and a stable literary semantic Gaussian field of author cognition is output.

5. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: The hierarchical partitioning and random sampling of the Gaussian field of literary semantics includes: Based on the stable author's cognition of the literary semantic Gaussian field, after obtaining the semantic density, rhetorical complexity, emotional fluctuation rate and allusion dissemination rate in the field, the entire literary semantic Gaussian field is hierarchically divided to form semantic layers, including the general semantic layer, the rhetorical semantic layer, the emotional semantic layer and the allusion semantic layer. In each semantic layer, semantic propagation paths are generated based on the propagation connection relationships between literary Gaussian units, and semantic sampling regions are randomly selected in the propagation paths for random sampling to generate several semantic sampling points.

6. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: The calculation of semantic contribution involves compressing or cloning semantic sampling points, including: Based on the semantic sampling points, calculate the semantic contribution of each semantic sampling region; By using preset contribution thresholds and preset judgment thresholds, it is determined whether the semantic region is a high-complexity semantic region or a low-complexity semantic region. In the case of a high-complexity semantic region, the semantic nodes are cloned, and in the case of a low-complexity semantic region, the semantic nodes are compressed.

7. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: The multi-directional cognitive propagation structure constructed using spherical harmonic propagation includes: For the enhanced literary semantic Gaussian field, a fixed semantic direction set is established, which includes emotional direction, rhetorical direction, historical direction, rhythmic direction and imagery direction. Based on each fixed semantic direction, the enhanced literary Gaussian unit is subjected to spherical harmonic propagation expansion processing to establish the semantic propagation function in the corresponding direction; Perform directional propagation modeling based on each semantic propagation function to obtain directional semantic propagation results corresponding to different fixed semantic directions; Calculate the directional coupling propagation strength between different directions based on the directional semantic propagation results corresponding to different fixed semantic directions; Based on the directional coupling propagation strength and directional semantic propagation results, the multi-directional cognitive propagation structure is calculated.

8. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: The generation of cognitive Gaussian fields for the test work and the target work is performed using Markovian computation, including: Based on the multi-directional cognitive propagation structure, Gaussian center, Gaussian covariance matrix and cognitive propagation density are generated; Based on Gaussian center, Gaussian covariance and cognitive propagation density, generate cognitive Gaussian field of the work to be tested and cognitive Gaussian field of the target work. Based on the cognitive Gaussian field of the work under test and the cognitive Gaussian field of the target work, the Mahalanobis distance deviation value between the cognitive Gaussian fields is obtained.

9. The AI-based method for analyzing ancient literary works as described in claim 1, characterized in that: The process of performing probability mapping on the Mahalanobis distance deviation value to generate a forgery probability for judging the authenticity of the work under test, and obtaining the final analysis results, includes: The Sigmoid function is used to perform probability mapping on the comprehensive cognitive bias value to form the spurious probability. The probability of forgery is compared with a preset forgery detection threshold; If the probability of forgery is greater than or equal to the preset forgery judgment threshold, the work to be tested is judged to be suspected of being forgery; otherwise, the work to be tested is judged to be genuine.

10. An AI-based system for analyzing ancient literary works, based on the AI-based method for analyzing ancient literary works as described in any one of claims 1 to 9, characterized in that: include, The data acquisition and processing module is used to acquire a collection of real works and works to be tested, perform standardized processing, and generate standardized text. The decoding construction module is used to construct literary semantic units after decoding the character vector sequence extracted from the standardized text, thus forming a set of literary semantic units; The analysis and iteration module is used to construct a complete literary semantic Gaussian field based on the set of literary semantic units, perform propagation direction analysis, establish a dual matrix to perform iteration, and output the literary semantic Gaussian field. The enhanced computation module is used to hierarchically divide and randomly sample the literary semantic Gaussian field, generate semantic sampling points and calculate semantic contribution, compress or clone the semantic sampling points, generate an enhanced semantic Gaussian field, construct a multi-directional cognitive propagation structure by expanding spherical harmonic propagation, and generate cognitive Gaussian fields of the work to be tested and the target work for Mahalanobis operation to obtain the Mahalanobis distance deviation value. The mapping and judgment module is used to perform probability mapping on Mahalanobis distance deviation values, generate forgery probabilities to judge the authenticity of the work under test, and obtain the final analysis results.