Multi-mode text reading system based on multi-semantic personality mapping
The multi-modal text reading system, which uses multi-semantic personality mapping, solves the problem that existing tools cannot dynamically adjust text presentation and multi-language synchronization. It enables collaborative reading with multiple languages, multiple modes, and multiple personality expressions, improving the flexibility and immersion of the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MENGKEDA (HONG KONG) TECHNOLOGY CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-05
AI Technical Summary
Existing tools cannot dynamically adjust the way text is presented according to user habits, multilingual versions cannot achieve consistent synchronous jumps, lack a unified framework to support multiple reading strategies, and cannot present different expressive personalities according to user preferences, thus limiting flexibility and immersion.
This paper presents a multi-modal text reading system based on multi-semantic personality mapping, including a semantic parsing module, a style configuration module, a multi-language mapping module, a reading mode module, and a user preference management module. It realizes a comprehensive reading experience with multi-language, multi-modal, and personalized expression. Through semantic parsing, style configuration, structural alignment, and reading mode switching, it generates personalized text or audio output.
It enables the collaborative presentation of multilingual, multimodal, and multipersonal expressions within the same framework, enhancing the flexibility and immersive experience of cross-language reading and meeting users' personalized needs.
Smart Images

Figure CN121981125A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing and text presentation, and more specifically, to a multimodal text reading system based on multi-semantic personality mapping. Background Technology
[0002] In cross-language reading, text learning, and content consumption, users often need to read in different ways depending on the task type, reading habits, or semantic comprehension level, such as close reading, comparative reading, and exploratory skimming. However, existing tools generally suffer from the following problems: 1) Fixed reading methods make it difficult to dynamically adjust text presentation according to user habits; 2) Multilingual versions cannot achieve consistent synchronous navigation; 3) Lack of a unified framework to support multiple reading strategies (such as linear reading, comparative reading, and partial display); 4) Inability to present different expressive personalities according to user preferences (such as serious, relaxed, academic, and colloquial), thus limiting flexibility and immersion.
[0003] Therefore, a system with a unified structure, scalability, and the ability to present text in multiple dimensions is needed to support a comprehensive reading experience that supports multiple languages, multiple modes, and personalized expression. Summary of the Invention
[0004] This invention provides a multi-modal text reading system based on multi-semantic personality mapping to at least address the technical problem of poor reading experience.
[0005] According to one aspect of the present invention, a multi-mode text reading system based on multi-style semantic mapping is provided, comprising: a semantic parsing module for performing semantic structure analysis on input text to obtain sentence-level, phrase-level, or word-level semantic annotation information; a style configuration module for rewriting the semantic annotations in a personalized manner according to expression parameters selected by the user, wherein the expression parameters include at least one of the following: lexical formality, speech rate, and tone intensity; a multilingual mapping module for mapping the text to at least one target language and generating a structure-aligned multilingual version; a reading mode module for providing multiple reading presentation modes based on user selection; a presentation control module for generating the final text or audio output according to the semantic analysis results, target language version, personal style, and reading mode; and a user preference management module for recording user reading behavior and parameter selection and providing adaptive configuration.
[0006] According to another aspect of the present invention, a multi-modal text reading method based on multi-style semantic mapping is also provided, comprising: performing semantic structure analysis on input text to obtain a multi-layer semantic structure at the sentence level, phrase level, or word level; rewriting the multi-layer semantic structure in a personalized manner based on the expression style selected by the user to obtain a personalized rewritten text; mapping the personalized rewritten text to at least one target language and aligning the different language versions structurally; and presenting the structurally aligned different language versions in the form of text or audio / video in response to the reading mode selected by the user, wherein the reading mode includes at least one of the following: word presentation, phrase presentation, masked reading, segmented exposure reading, and left and right channel comparison.
[0007] In this embodiment of the invention, a semantic parsing module generates a multi-layered semantic structure of the input text; a style configuration module rewrites the semantic structure in a personalized way according to the user's selected expression style; a multilingual mapping module aligns the structures between different language versions; a reading mode module provides multiple reading modes such as word presentation, phrase presentation, masked reading, segmented exposure reading, and left / right channel contrast; a presentation control module generates corresponding text or audio output based on the semantic structure, target language, and personality style; and a user preference management module records and adapts to the user's reading behavior. This invention enables the collaborative presentation of multiple languages, modes, and personalities within a single framework, enhancing the flexibility and immersive experience of cross-language reading, thereby solving the technical problem of poor reading experience. Attached Figure Description
[0008] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0009] Figure 1 This is a schematic diagram of an optional multi-mode text reading system based on multi-style semantic mapping according to an embodiment of the present invention;
[0010] Figure 2 This is an optional style control flowchart according to an embodiment of the present invention;
[0011] Figure 3 This is an optional semantic model and style configuration relationship diagram according to an embodiment of the present invention;
[0012] Figure 4 This is an optional multi-mode reading presentation flowchart according to an embodiment of the present invention;
[0013] Figure 5 This is a schematic diagram of an optional user interaction and preference control according to an embodiment of the present invention;
[0014] Figure 6 This is a flowchart of an optional multi-mode text reading method based on multi-style semantic mapping according to an embodiment of the present invention;
[0015] Figure 7 A schematic diagram of the structure of a computer device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0016] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0018] This invention aims to provide a system that can simultaneously support: structural mapping of multilingual texts, dynamic switching of multiple reading modes, and multidimensional presentation of personality-styled expressions, thereby constructing a unified, multi-layered text reading system and realizing real-time mapping between the four layers of "semantics, language, personality, and pattern".
[0019] This invention constructs a semantic parsing module, a language mapping module, a style configuration module, a reading mode module, a presentation control module, and a user preference management module, enabling the system to automatically generate corresponding text presentations or audio expressions based on users' personalized choices in language, personality, and reading style, ultimately achieving a multi-dimensional and customizable reading experience.
[0020] The main advantages of this invention are: First, it supports multilingual switching based on a unified semantic model, achieving text structure alignment; second, it can flexibly adapt to the needs of different users through the combination of multiple reading modes, and provides users with the ability to choose their own expression style by combining style configuration; in addition, the modules adopt an arrangable and combinable design, making the system as a whole highly scalable.
[0021] like Figure 1 As shown, the multi-style semantic mapping-based multi-modal text reading system comprises three main parts: a user interface layer, a reading engine layer, and an output feedback layer. The reading engine layer includes a semantic parsing module, a style configuration module, a multilingual mapping module, and a reading mode module; the output feedback layer includes a presentation control module and a user preference management module. All modules interact efficiently and collaborate on data through a unified semantic annotation layer, jointly supporting a personalized reading experience. The multi-semantic personality can be a set of semantic presentation strategies preset by the system, selected by the user, or configured based on the reading scenario. The reading mode module may also include composite presentation modes dynamically generated based on semantic structure.
[0022] The semantic parsing module is responsible for performing deep semantic analysis and structured annotation on the input text, providing a unified semantic foundation for the system.
[0023] The style configuration module allows users to select a specific "expressive personality" (such as "British Gentleman," "American Relaxed," etc.). The system will then automatically adjust vocabulary preferences, tone style, rhythm, emotional curve, and syntax to reflect this individual expression in reading aloud or text presentation. Figure 2 As shown, the style control process includes a set of style vector parameters set by the user or generated by the system through learning; controlling the output rhythm, sentence structure, tone, vocabulary, and other expressive features; the specific process includes: parameter input → style space positioning → output style construction path. A multi-semantic personality can be a set of semantic presentation strategies preset by the system, selected by the user, or configured based on the reading scenario; in this application, personality is an engineered attribute of the "strategy set / configuration object" and does not involve personality learning or generation.
[0024] The multilingual mapping module is used to achieve structural synchronization between different language versions, including sentence-level and phrase-level mapping, punctuation and paragraph structure alignment, and supports parallel jumping between multiple languages to ensure the immediacy of language switching during reading.
[0025] Figure 3 The process of cross-language semantic mapping and the location synchronization mechanism between multilingual content are shown, including steps such as unified semantic graph, anchor structure extraction, mapping table construction, and jump pointer synchronization. It demonstrates how the system can achieve seamless jumps between different languages while maintaining semantic consistency.
[0026] The reading mode module is used to present text in different structural ways according to user needs. Specifically, it includes: 1) Word-Level Presentation Mode: Displaying text word by word, suitable for basic understanding or fine-grained learning scenarios. 2) Phrase-Level Presentation Mode: Displaying text by phrases or semantic units to improve reading fluency. 3) Masked Reading Mode: Obscuring parts of words or phrases to guide users to predict or reinforce memory. 4) Segmental Exposure Mode: Presenting text in segments for structured skimming and key point extraction. 5) BinauralContrast Mode: Playing audio of different languages or styles simultaneously or in contrast between the left and right channels to enhance cross-language comparison and comprehension.
[0027] Figure 4 The system demonstrates its switching mechanism between different reading modes (word, phrase, masking, segmented exposure, bilingual audio); this includes mode selection judgment, content segmentation logic, semantic node recognition, and rendering distribution. During runtime, the system automatically schedules different reading modes based on semantics or user behavior. Mode switching can be triggered based on user interaction, reading progress status, or preset reading strategies. This switching behavior has a clear engineering trigger source, rather than an abstract judgment.
[0028] The presentation control module is responsible for taking into account the user's selected reading mode, expression personality, and language version to generate the final text rendering or audio output.
[0029] The user preference management module records users' historical behavior to remember their personality preferences, automatically restore reading modes, and recommend the most suitable presentation methods, thereby improving the continuity of system use and the personalized experience.
[0030] Figure 5 It demonstrates the complete feedback loop of the user preference learning module, including: user behavior collection (swiping speed, dwell time, language switching, etc.) → preference model update → output strategy adjustment; showing how the system continuously learns to optimize the personalized output rhythm and style adaptation.
[0031] According to an embodiment of the present invention, a method embodiment of a multi-mode text reading method based on multi-style semantic mapping is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0032] like Figure 6 As shown, the method includes the following steps:
[0033] Step S602, semantic parsing.
[0034] The system receives raw text data as input. The text can be monolingual or contain multilingual content. The system first calls the semantic parsing module to perform hierarchical semantic annotation on the text, generating a multi-layered semantic structure including word, phrase, sentence, and paragraph levels. During parsing, the module utilizes a multi-layered attention network based on the Transformer architecture to extract semantic dependency edges and core semantic nodes from contextual relationships. The above model structure is only an example; other equivalent semantic modeling, style control, or parameter scheduling methods can be used for specific implementations.
[0035] To achieve accurate semantic segmentation, a multi-channel parsing strategy is adopted, which involves running a main parser and an auxiliary parser simultaneously in the semantic vector space. The main parser is responsible for generating the global semantic backbone, while the auxiliary parser performs local completion on low-frequency words or structurally ambiguous parts. By dynamically fusing the attention weight matrix, a unified semantic embedding result is obtained.
[0036] Specifically, for the input text sequence The system first generates the corresponding initial embedding matrix. ,in For word count, The dimension is vector. Then, semantic weights are calculated based on a multi-head attention mechanism:
[0037]
[0038] Where A represents the attention weight matrix, used to measure the semantic relevance between words in the input sequence; Q represents the query matrix, obtained by linear transformation of the embedding vectors of the input sequence; K represents the key matrix, used to calculate the matching relationship between words; K T This represents the transpose of the key matrix, used to multiply with the query matrix to form a match score; d k The dimension of each attention head is used to scale the computation results to prevent excessive gradients; softmax(·) represents the normalization function used to transform the matching scores into a probability distribution. The final semantic representation is as follows:
[0039]
[0040] Where S represents the final semantic representation matrix, which contains the feature information after context fusion; A represents the attention weight matrix; and V represents the value matrix, which is the original semantic vector of each word in the input sequence.
[0041] After layer normalization and context residual correction, the complete semantic structure tree is obtained. Each node contains semantic labels, dependencies, and edge weights. During the parsing phase, the system simultaneously constructs a timestamp index table and a semantic anchor index table to facilitate rapid invocation of subsequent cross-language and multi-modal modules.
[0042] This method introduces a contextual semantic relevance matrix and a dynamic anchor mechanism, which can accurately maintain semantic continuity in multilingual texts.
[0043] Step S604, style configuration.
[0044] Based on the user's selected target personality type, the semantic structure is stylized and rewritten. The personality type can be a preset template (such as "academic rationality", "gentle narration", "light and cheerful speech", "analytical expert", etc.) or it can be generated by the system through self-learning from historical preferences.
[0045] Traditional style configuration methods often rely on fixed style dictionaries or simple tone shifts, which can easily lead to style rigidity. This invention proposes a dynamic style configuration algorithm based on style control vectors. This algorithm achieves adaptive changes in multiple dimensions such as tone, vocabulary, and syntactic rhythm by adjusting the control vectors in real time.
[0046] First, a basic style vector is defined for each target personality type. ,in The system considers stylistic features, including tone weighting, sentiment intensity, sentence length preference, pause rhythm, and lexical abstraction. During text generation, the system monitors changes in the semantic context and calculates the sentiment density of the current semantic block. and logical complexity .
[0047] The new control vector is then calculated using the following adaptive adjustment formula. :
[0048]
[0049] in, This is the sentiment density vector for the current paragraph; Let this be a vector of logical complexity. and These are the global mean; , These are learnable parameters used to control personalized sensitivity.
[0050] in, Pi represents the dynamic style control vector generated on the i-th semantic block; P0 represents the base style vector of the target personality; α represents the affective sensitivity adjustment parameter, used to control the adjustment magnitude of style in the affective dimension; β represents the logical sensitivity adjustment parameter, used to control the adjustment magnitude of style in the logical complexity dimension; Ei i This represents the sentiment density vector of the current semantic block; Ci represents the average vector of global sentiment density; Ci represents the logical complexity vector of the current semantic block. A vector representing the average global logical complexity.
[0051] Using the methods described above, the system can dynamically fine-tune the style of different semantic blocks while maintaining overall personality consistency. For example, when the system detects that the emotional intensity of a sentence segment is higher than the global average, it can automatically increase the weight of the emotional coloring of words and slow down the speaking pace, thereby making the output personality more natural and vivid.
[0052] In addition, to prevent style drift, the system introduces a personality consistency loss function. :
[0053]
[0054] in Indicates the Kullback-Leibler divergence. and These are the distributions of the current and baseline styles, respectively; These are the consistency weight coefficients. Optimization is achieved through backpropagation. The model is able to stably maintain personality consistency. This represents the personality consistency loss value, used to constrain the deviation between the dynamically adjusted style control vector and the original personality template. λ represents the squared Euclidean distance between style vectors, used to reflect the degree of deviation; λ represents the consistency constraint weight coefficient, used to balance the contributions of the two parts of the loss. This represents the Kullback-Leibler divergence, used to measure the current style distribution. Compared with the baseline style distribution Differences
[0055] Finally, the system integrates style control vectors with semantic tree node embeddings to generate a semantic expression structure imbued with personality traits. This improved dynamic adjustment algorithm not only reflects the user's individual preferences but also adaptively adjusts the language rhythm according to changes in the text's theme.
[0056] Among them, the sentiment density Ei is calculated by matching the sentiment dictionary, and the logical complexity Ci is obtained based on the syntactic tree depth and the number of clauses.
[0057] Step S606: Structure alignment and position synchronization.
[0058] Based on the semantic parsing results and the personalized expression structure, multilingual mapping is performed to ensure structural alignment between different language versions, enabling users to seamlessly jump to the same semantic position when switching reading languages.
[0059] First, using a semantic anchor index table, an anchor alignment map is established between the source and target languages. Assume the source language text is... The target language text is Then each sentence-level node Corresponding to one or more Through semantic similarity function Calculate the matching score.
[0060] When the matching score exceeds the set threshold At that time, establish anchor point mapping relationships:
[0061]
[0062] Where M represents the set of anchor mappings between the source language and the target language; si represents the i-th semantic node in the source language text; tj represents the j-th semantic node in the target language text; f(s) i , t j ) represents the semantic similarity function, used to calculate the similarity between two nodes, usually in the form of cosine similarity; θm represents the semantic matching threshold, only when the similarity is higher than this threshold are the node pairs considered semantically equivalent.
[0063] To avoid confusion in many-to-one or one-to-many mappings, a maximum matching optimization algorithm is employed, based on constraints. , To maintain structural uniqueness.
[0064] After establishing the anchor point map, a synchronized jump pointer table is generated. When a user triggers a language switching operation during reading, the presentation control module can directly synchronize the reading position from the source language to the corresponding sentence segment in the target language based on the pointer table, achieving seamless cross-language jumps.
[0065] Compared to existing technologies, this method considers not only lexical alignment during mapping but also introduces a semantic energy matrix and context dependency coefficients, resulting in more stable structural matching. Especially in cases of asymmetric bilingual sentence structures, the system can maintain semantic consistency through minimum energy path search.
[0066] Step S608: Switch reading mode.
[0067] The system dynamically switches between different reading modes based on user selection or system recommendations. It includes multiple reading modes such as word presentation, phrase presentation, masked reading, segmented exposure, and dual-channel audio comparison.
[0068] First, the hierarchical relationship of nodes in the semantic structure tree is read, and the optimal mode is determined based on user preference parameters. If the user selects "masked reading mode," the system performs probabilistic masking on the node words, with the masking probability determined by the node's semantic weight.
[0069]
[0070] in This represents the semantic importance of a node; a higher value indicates that the semantics are more critical and less likely to be obscured. This is the mode adjustment coefficient. This represents the probability that the i-th semantic node is occluded. This ensures that key content remains visible while auxiliary information is partially hidden, thus strengthening the user's active memory.
[0071] If the user switches to "segmented exposure mode," the system calculates the exposure sequence based on semantic hierarchy. Specifically, the module generates the exposure function according to the priority of each segment. ,in Control the exposure speed, These are the rhythm modulation parameters.
[0072] In "Dual-channel contrast mode," audio signals of different languages or different personality styles are output synchronously through the left and right channels, with signal synchronization controlled by a time alignment function.
[0073]
[0074] in , These represent the left and right channel timestamps, respectively. This is the tolerance error threshold.
[0075] During operation, the module automatically schedules different reading modes based on the real-time status of semantic nodes. For example, when it detects a decrease in user attention or excessive dwell time, the system can switch from segmented mode to phrase mode to restore rhythm and fluency.
[0076] Step S610, present control.
[0077] The aforementioned multi-dimensional information is then integrated to generate the final output. The output format can be text rendering, voice reading, or visual overlay display.
[0078] First, the personalized rewritten semantic tree is fed into the presentation synthesizer to perform semantic-to-form mapping. For text output, the corresponding syntactic template and mood marker are selected based on the current state of the style control vector. If the output is speech, a neural speech synthesizer is used to map the style parameters into acoustic features, such as timbre fundamental frequency, formants, and speech rate curves.
[0079] In acoustic synthesis, characteristic modulation functions are used:
[0080]
[0081] in As the final signal, Based on the basic speech waveform, Indicates the emotion modulation factor. This is the modulation intensity coefficient.
[0082] If the presentation is a combination of text and images, paragraph indentation, font weighting, or color markings are automatically generated based on semantic hierarchy to highlight the semantic intensity at different levels. This module also supports cross-device synchronous output, such as displaying text on the home screen and playing audio through headphones, thereby achieving a collaborative reading experience of visual and auditory senses.
[0083] During this phase, the system continuously receives feedback data from the user preference module to adjust output details. For example, if the system detects that a user spends a long time on a certain personality style, it will increase the weight of that style in the real-time output.
[0084] Step S612, User Preference Management and Adaptive Learning.
[0085] By continuously recording user behavior data, feedback and learning are provided for the reading process. Behavioral data includes page scrolling speed, dwell time, language switching frequency, mode switching frequency, and volume adjustment habits.
[0086] The module first normalizes these raw behavioral signals to construct behavioral feature vectors. The user preference distribution was then calculated using a time-weighted average model.
[0087]
[0088] in For the time when the behavior occurs, This is a time decay factor used to emphasize the importance of recent behavior. Let represent the feature value of the i-th action.
[0089] The system dynamically updates user profiles based on this distribution and automatically restores the previous language version, personality style, and reading mode when the next reading session begins. If the system detects that a user exhibits differentiated tendencies at different times (e.g., a preference for fast mode during the day and an immersive mode at night), it will automatically create multi-time-period configuration sets to achieve environmental adaptation.
[0090] At the same time, the module feeds back the preference parameters to the style configuration module to adjust the initial state of the style control vector, so that the system can gradually form a unique personality expression curve for each user over a long period of use.
[0091] The method of this invention achieves the collaborative operation of semantic parsing, multilingual structural alignment, personalized rewriting, multi-modal reading presentation, and user preference self-learning within the same technical framework. Among these, the dynamic adjustment algorithm for style control vectors brings significant advantages to the system, enabling reading output to be adjusted in real time based on content semantics and user behavior, rather than being limited to static templates.
[0092] bottom of form
[0093] Figure 7 A schematic diagram of a computer device suitable for implementing embodiments of the present disclosure is shown. It should be noted that... Figure 7 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0094] like Figure 7 As shown, the computer device includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage section 1008 into a random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0095] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 1010 as needed so that computer programs read from it can be installed into storage section 1008 as needed.
[0096] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A multi-modal text reading system based on multi-style semantic mapping, characterized in that, include: The semantic parsing module is used to perform semantic structure analysis on the input text to obtain semantic annotation information at the sentence, phrase, or word level. The style configuration module is used to rewrite semantic annotations based on the expression parameters selected by the user, wherein the expression parameters include at least one of the following: lexical formality, speech rate, and tone intensity; The multilingual mapping module is used to map text to at least one target language and generate a structure-aligned multilingual version. The reading mode module is used to provide multiple reading presentation methods based on user selection; The presentation control module is used to generate the final text or audio output based on semantic analysis results, target language version, personality style, and reading pattern. The user preference management module records user reading behavior and parameter selections and provides adaptive configuration.
2. The system according to claim 1, characterized in that, The reading mode module includes: word presentation mode, phrase presentation mode, masked reading mode, segmented exposure reading mode, and left and right channel contrast mode.
3. The system according to claim 1, characterized in that, The style configuration module is used to adjust at least one expressive feature, including: lexical style, tone curve, sentence rhythm, syntactic structure, or timbre features.
4. The system according to claim 1, characterized in that, The multilingual mapping module performs cross-language structural alignment based on semantic annotation information, including sentence-level alignment, phrase alignment, or word-level alignment.
5. The system according to claim 1, characterized in that, The presentation control module can generate audio output based on the target language version and personality style, and supports synchronous or contrastive presentation of different languages or styles between the left and right channels.
6. The system according to claim 1, characterized in that, The user preference management module automatically recommends personality style, reading mode, or language version based on the user's historical operations.
7. The system according to claim 1, characterized in that, The semantic parsing module can identify the semantic boundaries of text, including paragraph boundaries, sentence boundaries, phrase boundaries, and word boundaries.
8. A multi-modal text reading method based on multi-style semantic mapping. Its characteristic is that... include Based on the semantic structure obtained by the semantic parsing module, cross-language position matching is performed in the multilingual mapping module, enabling users to jump synchronously between texts in different languages. Based on semantic annotation information, the text is stylized and rewritten so that the output content conforms to the target personality expression characteristics selected by the user. Depending on the selected reading mode module, the text will be output in the form of word presentation, phrase presentation, masking presentation, segmented exposure presentation, or left and right channel comparison. Based on the user's selected target language version, personality expression characteristics, and reading mode, the corresponding text or audio output is generated through the presentation control module; Recommendation parameters are automatically generated based on the user's historical reading path, language preferences, and personality choices, and the reading presentation modes are prioritized.
9. A multi-modal text reading method based on multi-style semantic mapping, characterized in that, include: Perform semantic structure analysis on the input text to obtain multi-level semantic structures at the sentence, phrase, or word level; Based on the user-selected expression style, the multi-layer semantic structure is rewritten in a personalized way to obtain the personalized rewritten text; The personalized rewritten text is mapped to at least one target language, and the different language versions are structurally aligned. In response to the user's selected reading mode, the target language is presented in the form of text or audio / video, wherein the reading mode includes at least one of the following: word presentation, phrase presentation, masked reading, segmented exposure reading, and left and right channel comparison.
10. The method according to claim 9, characterized in that, After presenting the target language in text or audio / video format, the method further includes recording and adapting the user's reading behavior.