A digital reading platform construction system supporting multi-modal content interaction

By constructing modules for multimodal content acquisition, interactive parsing, and platform construction, the problems of chaotic resource management and low interactive accuracy on digital reading platforms have been solved, enabling efficient management and precise interaction of heterogeneous resources, and improving system stability and user experience.

CN122451027APending Publication Date: 2026-07-24CHINA FOCUS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-27
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing digital reading platforms lack a complete XR multimodal resource and AI lightweight model processing mechanism, making it difficult to adapt to diverse reading scenarios. They suffer from chaotic resource management, insufficient semantic parsing accuracy, inability to achieve precise binding interaction, and issues such as loading delays and running lag.

Method used

The system constructs a multimodal content acquisition module, a multimodal interaction parsing module, and a digital reading platform construction module to achieve the classification and organization of heterogeneous resources, semantic parsing, and generation of interactive commands. It also establishes a standardized multimodal content feature database, adopts a progressive intent matching and scenario-based weighted scoring mechanism, and configures platform operation and management components.

Benefits of technology

It enables efficient management and precise allocation of heterogeneous resources, improves the recognition accuracy and system stability of multimodal interactions, optimizes resource management efficiency, and ensures flexible scheduling of user interaction needs and an immersive reading experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451027A_ABST
    Figure CN122451027A_ABST
Patent Text Reader

Abstract

The application discloses a kind of digital reading platform construction systems of supporting multimodal content interaction, it is related to digital reading technical field, the present application includes multimodal content acquisition module, multimodal interaction analysis module and digital reading platform construction module, the present application is classified and is regularized by from multiple channels acquisition heterogeneous reading resources, standardization multimodal content feature database is constructed by mode extraction feature, pre-processing and semantic analysis are carried out to user multimodal interaction input, user interaction intent is identified and interaction execution instruction is generated by progressive threshold matching and scene weighting score mechanism, based on feature database, build hierarchical platform framework, configure control component and establish multimodal content display and interaction response logic, the present application improves the resource management efficiency of digital reading platform, interaction accuracy, immersive experience and running reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital reading technology, specifically to a digital reading platform construction system that supports multimodal content interaction. Background Technology

[0002] With the rapid development of digital information technology, XR extended reality, and artificial intelligence, digital reading has been fully integrated into people's daily lives. Users' reading needs are no longer limited to traditional single text reading, but are gradually shifting towards XR immersive multimodal reading that integrates text, audio, video, and images. At the same time, the demand for AI-reliable retrieval, lightweight content model generation, and spatial object binding interaction is also increasing. This places higher demands on the content management, interaction accuracy, XR scene adaptation, and intelligent service capabilities of digital reading platforms.

[0003] Most existing digital reading platforms lack robust XR multimodal resources and lightweight AI model processing mechanisms, making them ill-suited for the diverse reading scenarios of the new era. They generally lack a comprehensive resource management system, resulting in disorganized and chaotic collection and categorization of XR space materials and AI-reliable content, lacking hierarchical organization and filtering processes. Furthermore, their user interaction modes are relatively simplistic, with insufficient semantic parsing accuracy for multimodal input information such as voice, natural language, and XR space gestures, making it difficult to accurately identify and match users' true reading intentions, and even more difficult to achieve precise binding and interaction between spatial objects and reading content. In addition, most platforms lack reliable AI content generation and XR immersive multimodal linkage capabilities, and their platform architecture design has not been specifically optimized for XR space computing and multimodal resource operation. This leads to problems such as loading delays, adaptation failures, and stuttering during resource retrieval, transmission, and display, severely impacting the smoothness of multimodal interaction. Although some platforms have attempted to introduce multimedia reading content, they have not yet formed a complete construction system from standardized processing of XR resources, intelligent analysis of spatial interaction, AI-based reliable content scheduling to platform architecture construction. This makes it impossible to meet users' needs for intelligent, immersive, and multi-modal digital reading. Therefore, developing a targeted platform construction system is of great practical significance. Summary of the Invention

[0004] The purpose of this invention is to provide a digital reading platform construction system that supports multimodal content interaction, thereby solving the problems existing in the background technology.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention provides a digital reading platform construction system that supports multimodal content interaction, including a multimodal content acquisition module, a multimodal interaction parsing module, and a digital reading platform construction module; The multimodal content acquisition module is used to collect multimodal heterogeneous reading resources from multiple channels, classify and organize the collected multimodal heterogeneous reading resources, extract data features of the multimodal heterogeneous reading resources, and build a multimodal content feature database based on the extracted data features; The multimodal interaction parsing module is used to collect multimodal interaction input information from users through smart terminals, perform semantic parsing on the collected interaction input information, obtain user interaction requirements based on the parsing results, and generate interaction execution instructions based on user interaction requirements. The digital reading platform construction module is used to build a digital reading platform framework based on a multimodal content feature database, configure platform operation and management components, and then combine interactive execution commands to establish multimodal content display and interactive response logic, forming a digital reading platform that supports multimodal content interaction.

[0006] The beneficial effects of this invention are as follows: (1) This invention achieves efficient management and accurate retrieval of heterogeneous reading resources by constructing a standardized multimodal content feature database. It adopts a three-level classification mode to regulate heterogeneous resources, and combines feature extraction and normalization processing to establish an association index between feature data and corresponding resources and a multi-dimensional auxiliary retrieval mechanism. This not only ensures the standardization and consistency of resource data, but also greatly improves the response speed of resource retrieval and retrieval, and avoids interference from invalid resources such as repeated collection and file corruption. At the same time, the hierarchical partitioned storage architecture provides data support for cross-modal resource linkage, enabling the platform to flexibly schedule different types of multimodal resources according to user interaction needs, thereby improving the stability and scalability of system operation and optimizing the resource management efficiency and service capabilities of the digital reading platform.

[0007] (2) This invention achieves accurate identification and scene adaptation of user multimodal interaction intent through progressive semantic intent matching and scene-based weighted scoring mechanism. It effectively breaks through the technical limitations of traditional digital reading platforms, such as single interaction mode, low intent recognition accuracy, and fragmented response logic. The system extracts three key information types for multimodal input: core words, operation behavior, and resource orientation. Combined with threshold judgment based on big data training and experimental optimization, it performs three progressive checks through core word matching, operation behavior adaptation, and resource type consistency. Then, it calculates the total matching score with scene-optimized weighted coefficients. This not only ensures the accuracy of intent recognition but also improves the adaptability to different reading scenarios and user interaction habits. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a schematic diagram of the system structure connection of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] Reference Figure 1 As shown, the present invention provides a digital reading platform construction system that supports multimodal content interaction, characterized in that it includes: a multimodal content acquisition module, a multimodal interaction parsing module, and a digital reading platform construction module; The multimodal content acquisition module is used to collect multimodal heterogeneous reading resources from multiple channels, classify and organize the collected multimodal heterogeneous reading resources, extract data features of the multimodal heterogeneous reading resources, and build a multimodal content feature database based on the extracted data features; In the above embodiments, the specific method for collecting multimodal heterogeneous reading resources from multiple channels and classifying and organizing the collected multimodal heterogeneous reading resources is as follows: Heterogeneous reading resources were collected from e-book libraries, multimedia reading resource libraries, picture book libraries, lightweight XR resource material libraries, and lightweight AI content model libraries. These heterogeneous reading resources included text, audio, video, images, XR spatial models, and independent AI content mini-model resources. Invalid resources that were repeatedly collected, had damaged files, incomplete content, or originated from illegal sources were removed to obtain initially screened reading resource data. Then, a multi-level classification model was adopted to classify and organize the reading resources. A three-level classification model was set up, where the first level of classification divided resource categories according to resource modality and XR spatial attributes, the second level of classification was processed according to content theme, reading scenario, and XR spatial scenario, and the third level of classification was classified according to file encoding, storage format, and model specifications to obtain multimodal reading resources after classification and organization.

[0012] It should be noted that the system employs a multi-channel approach to cover mainstream resource types across all digital reading scenarios. The e-book library focuses on in-depth text-based reading resources; the multimedia reading resource library emphasizes audio and video resources; the picture book library focuses on visual and educational resources combining text and images; the lightweight XR resource library focuses on immersive resources such as XR spatial models and 3D interactive scenes; and the lightweight AI content model library focuses on independent AI content models and AI-generated intelligent resources. These five complementary resource libraries achieve comprehensive coverage of all types of multimodal reading resources, avoiding the limitations of scenario adaptation caused by a single resource type, while simultaneously meeting the new reading demands of XR immersive reading and AI intelligent interaction. Initial screening of collected resources to remove invalid resources reduces redundant data in subsequent feature extraction and storage, lowering system computing and storage costs.

[0013] It should be noted that the three-level classification model follows a hierarchical logic based on modal attributes, application scenarios, and storage specifications: The first-level classification is based on resource modality and XR space attributes, which can quickly distinguish the processing logic of different types of resources and adapt to the specific processing needs of XR space resources and AI model resources; the second-level classification is based on content theme, reading scenario, and XR space scenario, which can realize the scenario-based aggregation of resources and facilitate the accurate call to resources based on user reading scenarios and XR interaction scenarios; the third-level classification is based on file encoding, storage format, and model specifications, which can adapt to the playback and display needs of different terminals, while meeting the standardized management requirements of AI models and XR space models, ensuring the compatibility and availability of resources in multi-terminal and multi-interaction scenarios.

[0014] In the above embodiments, the specific method for extracting data features from multimodal heterogeneous reading resources and constructing a multimodal content feature database based on the extracted data features is as follows: Based on the categorized and normalized multimodal reading resources, data features adapted to XR spatial computing and AI content model calls are extracted for each modality. Specifically, for text-based reading resources, keywords, content attributes, and length levels are extracted; for audio-based resources, duration, timbre tags, and content summaries are extracted; for video-based resources, image parameters, duration nodes, and associated text features are extracted; for image-text resources, layout, image-text ratio, and annotation information features are extracted; for XR-based resources, spatial coordinates, object IDs, and scene binding features are extracted; and for AI model-based resources, model identifiers and trusted source features are extracted. The extracted data features are normalized to obtain standardized reading resource data features. These standardized features are then hierarchically and partitioned according to the resource categories corresponding to the three-level classification model. An association index is established for each standardized feature data with the corresponding reading resource, XR spatial object, and AI model. The standardized feature data and association indexes are integrated to form a standardized multimodal content feature database adapted to XR spatial computing and trusted AI scheduling.

[0015] It should be noted that differentiated feature extraction strategies are adopted for different modalities of reading resources because of the fundamental differences in the information carrying dimensions and presentation formats of text, audio, video, images, XR spatial models, and AI content models. Specifically, text, audio / video, and image-based resources correspond to the feature extraction needs of traditional reading scenarios; XR resources extract spatial coordinates, object IDs, and scene binding features to adapt to the technical requirements of XR spatial computation and spatial object binding interaction, ensuring accurate resource retrieval in immersive XR reading scenarios; AI model resources extract model identifiers and trusted source features to adapt to the trusted scheduling and security management requirements of AI content models, preventing the misuse of unauthorized models. Normalizing the extracted feature data eliminates the differences in dimensions and numerical ranges between different feature dimensions, bringing all types of feature data to the same level, facilitating subsequent feature retrieval, matching, and calculation.

[0016] The multimodal interaction parsing module is used to collect multimodal interaction input information from users through smart terminals, perform semantic parsing on the collected interaction input information, obtain user interaction requirements based on the parsing results, and generate interaction execution instructions based on user interaction requirements. In the above embodiments, the specific method for performing semantic parsing on the collected interactive input information and then obtaining user interaction requirements based on the parsing results is as follows: The collected multimodal interactive input information is preprocessed to filter noise, invalid characters, and abnormal operation data. Voice information is translated into a standardized text format, and touch, XR gaze, gesture, and spatial location entry information are converted into instructional text descriptions adaptable to XR interaction forms, resulting in standardized user interaction information. This standardized user interaction information is then segmented to extract key information, including core vocabulary, operation behavior, resource reference, XR spatial objects, and coordinates. Subsequently, a digital reading scene semantic database adapted to XR spatial computing and AI reading scenarios is invoked. The extracted key information is matched with the corresponding reading interaction intent in the semantic database, and the multimodal resource type, reading scenario, and XR spatial binding relationship corresponding to the user interaction intent are labeled to form user interaction requirements.

[0017] It should be noted that the semantic database has a built-in standardized reading interaction intent set, intent-keyword mapping table, operation behavior matching rules, resource type association rules, reading scene tag library and XR space binding relationship rule library that integrate XR space computing and AI capabilities, forming a semantic matching system that covers the entire digital reading scenario and is adapted to XR immersive interaction. The reading interaction intent set includes full-scenario interaction intents such as resource retrieval, modal switching, content playback / pause, content annotation, content collection, page navigation, and XR spatial object triggering / location, covering the complete reading link from resource search to XR immersive interactive operation. The intent-keyword mapping table establishes a multi-dimensional mapping relationship between core words and interaction intents. For example, words such as "play" and "read aloud" are associated with the "content playback" intent, "annotate" and "underline" are associated with the "content annotation" intent, and "focus on object" and "location coordinates" are associated with the "XR spatial object interaction" intent, improving the matching efficiency of core words and intents. The operation behavior matching rules define the adaptation logic of different operation behaviors and intents. For example, the voice command "pause", the touch operation "click the pause button", and the XR gesture confirmation are all adapted to the "content pause" intent, ensuring the consistency of intent recognition for multimodal operation behaviors. The XR spatial binding relationship rule library solidifies the binding logic of spatial object ID, coordinates and interaction intents, providing data support for accurate matching of XR immersive interaction.

[0018] It should be noted that in the preprocessing of multimodal interactive input information, noise filtering for environmental noise in voice input can be achieved using spectral subtraction and wavelet denoising algorithms; abnormal data such as accidental touches and repetitive operations in touch input can be eliminated by removing touch data that is repeatedly clicked within a short period of time or exceeds the effective operation area; for XR gaze, gesture, and spatial location entry type inputs, abnormal data can be eliminated by removing operation data that deviates from the preset gaze trajectory, does not conform to the standard gesture trajectory, or exceeds the effective interaction range of XR space. Translating voice information into standardized text and converting touch / XR gaze / gesture / spatial location entry type information into instructional text descriptions can unify the input information of different modalities into a text format, eliminating the interference of modal differences on intent recognition; when performing word segmentation processing on standardized user interaction information, a dedicated word segmentation dictionary adapted to digital reading and XR interaction scenarios can be used to prioritize the identification of core words related to the reading scenario, XR spatial object identifiers, and coordinate information, such as picture book names, resource types, operation instructions, and spatial object IDs, thereby improving the accuracy of key information extraction.

[0019] It should be noted that generating interaction execution instructions based on user interaction needs transforms users' multimodal interaction needs into standardized operation instructions that the platform can recognize and execute. This achieves a precise mapping from user needs to the platform's business logic. The user interaction needs are marked with the core interaction intent, multimodal resource types, and target reading scenarios. When generating instructions, the needs are structurally decomposed, and the operation behavior, resource type, resource identifier, and reading scenario information are encapsulated into standardized instruction fields.

[0020] It's important to note that generating interaction execution commands based on user interaction needs involves transforming users' multimodal interaction requirements into standardized operation commands that the platform can recognize and execute. This achieves a precise mapping from user needs to the platform's business logic. User interaction requirements are already labeled with core interaction intents, multimodal resource types, target reading scenarios, and XR space binding relationships. When generating commands, the requirements are structurally decomposed, encapsulating operation behaviors, resource types, resource identifiers, reading scenario information, XR space object IDs, and binding coordinate information into standardized command fields. Simultaneously, it adapts to the content scheduling requirements of AI reading scenarios, providing complete command support for subsequent XR immersive multimodal content display and spatial interaction responses, ensuring the system can accurately respond to users' multimodal interaction needs.

[0021] In the above embodiments, the specific method for further invoking the digital reading scenario semantic database and matching the extracted key information with the corresponding reading interaction intent in the semantic database is as follows: It utilizes a dedicated semantic database for digital reading scenarios that integrates AI capabilities, extracting key words, operational behaviors, resource type matching, and spatial object matching thresholds from the semantic database. The key word matching threshold is... The operational behavior adaptation threshold is derived from the combination of digital reading interaction scenarios and VR / MR content production and interaction scenario big data training and iteration. Based on actual tests in reading interaction scenarios and VR / MR immersive interaction scenarios, the resource type matching threshold was determined. To adapt to the full-match hard threshold required for VR / MR content production resource matching, the extracted key information is matched and calculated with the reading interaction intent in the semantic database, and the spatial object matching threshold is calculated. Bind a specific threshold to the XR space using the formula. Calculate the core vocabulary matching degree, where For core vocabulary matching, To match the number of characters, This is the total number of characters in the standard keywords. If the core vocabulary match is valid, then the formula is used to determine that the match is valid. Calculate the operational behavior adaptation degree, where For operational behavior adaptation, To match the number of operations, To extract the total number of operations, if If the operation is deemed to be consistent with the basic reading interaction intent, then according to the formula... Computational resource consistency, where For resource consistency, For the number of overlapping resource types, This refers to the total number of resource types. If the resource points to the same content as the reading interaction intent, then according to the formula... Calculate the spatial object matching degree, where For spatial object matching degree, The Euclidean distance between the user interaction point and the reference coordinates of the spatial object. Bind a valid domain radius to a spatial object, if Then the spatial object is matched with the reading interaction intent, and then the formula is used to determine the relationship. Calculate the total matching score, where The total matching score. The weighting coefficients for scenario optimization are derived from big data analysis of digital reading interaction scenarios and actual measurement of user interaction behavior. The reading interaction intent with the highest total matching score is selected as the target reading interaction intent that matches the user's reading interaction intent.

[0022] It should be noted that the core vocabulary matching threshold is derived through iterative training using big data from digital reading interaction scenarios. Based on massive amounts of user interaction data, it statistically correlates the core vocabulary matching degree with the accuracy of intent recognition, selecting a critical value that maximizes the distinction between valid and invalid matches. This ensures the accuracy of core semantic recognition while avoiding the missed detection of valid intents due to excessively high thresholds. The operational behavior adaptation threshold is derived from actual testing in reading interaction scenarios. It is determined by testing the adaptation of different operational behaviors and intents in various reading scenarios, such as children before bedtime or commuting while studying. The threshold is selected to ensure that the operational behavior and intent are compatible. Figure 1 The optimal threshold for consistency avoids misjudgment of intent due to deviations in operational behavior. The resource type matching threshold is set to a hard threshold of full match because resource type is one of the core indicators of user interaction needs. Mismatched resource types will directly lead to interaction responses that do not meet user needs. Therefore, it is required that the resource type associated with the intent be completely consistent to ensure the accuracy of the interaction response. The spatial object matching threshold is an XR space binding-specific threshold, which is derived from actual measurements based on the spatial positioning characteristics of XR immersive interaction scenarios. It is used to determine whether the distance between the user interaction point and the reference coordinates of the spatial object meets the binding interaction requirements, ensuring accurate identification of XR space interaction intent and adapting to the spatial interaction scenario needs of R / MR content creation and interaction.

[0023] It should be noted that the scene optimization weighting coefficients were derived from big data analysis and actual user interaction behavior testing. The highest weighting, reflecting core keywords, is the foundation for identifying user interaction intent; The weighting is secondary; it balances the auxiliary verification role of operational behavior and resource allocation on intent. The weighted adaptation meets the specific scenario requirements of XR spatial interaction, reflecting the core position of spatial object matching in immersive interaction. By calculating the total matching score through weighted summation, the contribution of three types of key information to intent recognition can be comprehensively balanced, avoiding intent recognition errors caused by single-dimensional bias, and improving the comprehensiveness and scenario adaptability of intent recognition.

[0024] It should be noted that: the core vocabulary matching degree is the matching degree of VR / MR content creation and interaction scenario; the number of matched characters is the number of matched characters related to VR / MR content creation and interaction; the total number of standard keyword characters is the total number of standard keyword characters related to VR / MR content creation and interaction; the operation behavior adaptation degree is the adaptation degree to VR / MR interaction form; the number of matching operation items is the number of matching operation items related to VR / MR interaction; the total number of extracted operation items is the total number of extracted operation items related to VR / MR interaction; the resource pointing consistency degree is the consistency degree to the resource requirements of VR / MR content creation; the number of overlapping resource types is the number of overlapping resource types related to VR / MR content creation; the total number of pointing resource types is the total number of pointing resource types related to VR / MR content creation; the effective domain radius of spatial object binding is the effective range threshold of XR spatial interaction; the total matching score is the total matching score adapted to VR / MR content creation and interaction scenario; and the scene optimization weighting coefficient is obtained from big data analysis of digital reading interaction scenarios and actual measurement of user interaction behavior to ensure that all parameters and thresholds are in line with the actual application requirements of VR / MR immersive interaction and multimodal content interaction.

[0025] The digital reading platform construction module is used to build a digital reading platform framework based on a multimodal content feature database, configure platform operation and management components, and then combine interactive execution commands to establish multimodal content display and interactive response logic, forming a digital reading platform that supports multimodal content interaction.

[0026] In the above embodiments, the specific method for building a digital reading platform framework based on a multimodal content feature database and configuring platform operation and management components is as follows: A layered digital reading platform framework is established based on a standardized multimodal content feature database. This framework comprises a data layer, an XR space layer, an interaction layer, and a presentation layer. A data interaction interface is established between the data layer and the standardized multimodal content feature database to adapt to XR content creation resource retrieval and AI model invocation. The interaction layer is connected to the multimodal interaction parsing module, and a two-way channel for receiving and responding to interaction execution commands is configured. The XR space layer is used for virtual scene rendering and spatial object binding management. The presentation layer sets up display areas for multimodal content such as text, audio, video, graphics, XR 3D animation, and AR annotations, and establishes cross-modal and XR space linkage display entry points. Furthermore, platform operation and control components are configured, including a resource scheduling unit, a spatial content scheduling unit, a data verification unit, and an operation monitoring unit. All control components are linked and adapted with the platform framework's data layer, XR space layer, interaction layer, and presentation layer.

[0027] It should be noted that: the data layer focuses on the connection and resource storage interaction of the standardized multimodal content feature database, and is responsible for the reading, writing, indexing, and persistent management of resource data. Its configuration is adapted to the data interaction interface for XR content production resource retrieval and AI model calling, which can meet the needs of efficient retrieval and reliable scheduling of XR space materials and lightweight AI content models; the XR space layer focuses on virtual scene rendering and spatial object binding management, and is responsible for the construction of XR immersive scenes and the binding and association of spatial objects with reading content, providing underlying support for XR space interaction; the interaction layer focuses on user multimodal interaction processing and command generation, and connects to the multimodal interaction parsing module to complete the reception, parsing and feedback channel of interaction commands; the presentation layer focuses on multimodal content presentation and user operation feedback, sets up dedicated display areas for text, audio, video, graphics, XR 3D animation and AR annotation multimodal content, and builds cross-modal linkage and XR space linkage display entrances to meet the content display needs of different reading scenarios.

[0028] It should be noted that the linkage and adaptation between the control components and the platform's four-layer framework are achieved through bidirectional transmission of command flow, data flow, and state flow. For example, in an XR immersive picture book reading scenario, when a user triggers a spatial object through an XR gesture, the interaction layer receives the parsed execution command. The resource scheduling unit then listens in real time and parses the operation behavior, resource type, and scene priority, and subsequently initiates a resource retrieval request to the data layer. Based on feature data indexing, it locates the corresponding XR 3D animation, AR annotation resources, and bound picture book content. Simultaneously, the spatial content scheduling unit completes the binding verification between the spatial object and the reading content, and XR... Scene content loading ensures accurate response to XR space interactions; after the data verification unit confirms that the data has passed verification, it notifies the display layer and XR space layer to allocate resources to the corresponding controls and configures the linkage rules for synchronized AR annotations and linked picture book text paragraphs in XR3D animation; the operation monitoring unit collects the running status of each layer in real time. When XR scene rendering delay is detected, the rendering optimization mechanism is triggered to temporarily adjust the rendering precision to match the interaction rhythm. At the same time, the process is recorded to the operation log to ensure accurate resource scheduling, strict verification, and controllable status, so as to achieve efficient collaboration between multimodal resources and XR space interactions.

[0029] In the above embodiments, the control component includes a resource scheduling unit, a data verification unit, and an operation monitoring unit; The resource scheduling unit is used to retrieve and allocate multimodal resources of the data layer, interaction layer and presentation layer according to the platform's interactive execution instructions and multimodal resource call requirements; The data verification unit is used to verify the multimodal resources retrieved from the standardized multimodal content feature database by the data layer, and to confirm that the reading resources can be transmitted normally to the display layer and are adapted to the resolution of the corresponding display area and meet the linkage requirements of the cross-modal linkage display entry. The operation monitoring unit is used to collect the operation status of the data layer, interaction layer, and presentation layer in real time, monitor the transmission status of the data interaction interface, the reception and feedback status of interaction commands, and record them to form a platform operation log.

[0030] It should be noted that: the data verification unit is responsible for verifying the availability and adaptability of resources to avoid display failures caused by invalid resources being transmitted to the presentation layer; the operation monitoring unit is responsible for monitoring the system's operating status in real time, promptly detecting and handling anomalies, ensuring the continuous and stable operation of the system, and linking and adapting the management components with the platform's three-layer framework to achieve deep integration of management logic and business logic. For example, the resource scheduling unit allocates resources to the presentation layer according to the instructions of the interaction layer, the data verification unit feeds back the verification results to the data layer and the presentation layer, and the operation monitoring unit monitors the operating status of each layer and triggers the anomaly handling mechanism to ensure that all aspects of the platform operate collaboratively and efficiently.

[0031] In the above embodiments, the specific method for establishing multimodal content display and interactive response logic by combining interactive execution instructions is as follows: The system extracts operational behaviors, multimodal resource types, target reading scenarios, and XR space object information from interactive execution commands. It retrieves matching reading resources, XR space materials, and AI content models from a standardized multimodal content feature database. Combining the target reading scenario and XR space scenario, it establishes multimodal content display logic, setting adaptation scenarios for text, audio, video, graphics, XR 3D animation, and AR annotations. Simultaneously, it establishes cross-modal linkage and XR space linkage rules to form multimodal content display logic. Based on user operation behaviors and XR space interaction behaviors, it establishes interactive response logic, setting exclusive execution rules compatible with XR interaction forms for various operation behaviors. The display logic and response logic are synchronized to the platform data layer, XR space layer, interaction layer, and display layer, and linked with the operation control components to form a complete XR immersive multimodal content display and spatial interaction response logic.

[0032] It should be noted that when establishing multimodal content display logic, differentiated display rules are set for different modal resources: text resources adopt adaptive font size and line spacing to adapt to different terminal screen sizes and display resolutions; audio resources are equipped with progress bars, volume control, and playback mode switching controls to meet users' auditory interaction needs; video resources support full-screen playback, speed adjustment, and subtitle synchronization functions to adapt to different terminal playback resolutions and improve the audiovisual experience; graphic resources provide page turning, zooming, and annotation functions to adapt to visual reading needs; XR 3D animation resources adapt to XR spatial layer rendering logic, support spatial perspective switching, 3D model scaling and rotation, and spatial anchor point positioning display to adapt to XR immersive interactive scenarios; AR annotation resources support spatial overlay display and virtual-real fusion linkage, and can dynamically adjust the display perspective according to the user's spatial position to adapt to AR augmented reality reading needs; AI-generated content resources support real-time rendering and dynamic update display to adapt to AI intelligent interactive reading scenarios.

[0033] It should be noted that cross-modal linkage rules, such as synchronously highlighting corresponding text and image segments during audio playback and synchronously displaying accompanying text during video playback, can achieve collaborative presentation of multimodal content and enhance the user's immersive reading experience. XR spatial linkage rules, such as synchronously loading corresponding AR annotations when XR 3D animations are triggered and dynamically displaying content bound to spatial objects based on user interaction, can achieve virtual-real fusion linkage of multimodal content in XR spatial scenarios, enhancing the immersive reading effect. Interactive response logic is established based on user operation behavior and XR spatial interaction behavior, adapting to different user interaction habits. For example, corresponding response rules are set for pause commands from voice input and page-turning behaviors from touch operations. For XR gaze, gesture, and spatial location entry interactions, dedicated spatial trigger response rules are set, such as gaze focusing triggering XR 3D animation playback, gesture dragging adjusting the position of spatial objects, and spatial location entry triggering AR annotation display, ensuring the smoothness and intuitiveness of XR spatial interaction operations. Linking with the operation and management components means that the execution of the display and response logic depends on the resource scheduling, spatial content scheduling, data verification, and operation monitoring of the management components. For example, the display logic can only be executed after the resource scheduling unit allocates resources, the XR spatial content can only be displayed after the spatial content scheduling unit completes the spatial object binding verification, the resources can only be displayed after the data verification unit passes the verification, and the operation monitoring unit monitors the status of the display and response process and triggers alarms or rollback mechanisms when anomalies occur to ensure the stability and reliability of the interactive experience.

[0034] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A system for building a digital reading platform that supports multimodal content interaction, characterized in that, include: The multimodal content acquisition module is used to collect multimodal heterogeneous reading resources from multiple channels, classify and organize the collected multimodal heterogeneous reading resources, extract data features of the multimodal heterogeneous reading resources, and build a multimodal content feature database based on the extracted data features; The multimodal interaction parsing module is used to collect multimodal interaction input information from users through smart terminals, perform semantic parsing on the collected interaction input information, obtain user interaction requirements based on the parsing results, and generate interaction execution instructions based on user interaction requirements. The digital reading platform construction module is used to build a digital reading platform framework based on a multimodal content feature database, configure platform operation and management components, and then combine interactive execution commands to establish multimodal content display and interactive response logic, forming a digital reading platform that supports multimodal content interaction.

2. The digital reading platform construction system supporting multimodal content interaction according to claim 1, characterized in that, The specific method for collecting multimodal heterogeneous reading resources from multiple channels and classifying and organizing the collected multimodal heterogeneous reading resources is as follows: Heterogeneous reading resources were collected from e-book libraries, multimedia reading resource libraries, picture book libraries, lightweight XR resource material libraries, and lightweight AI content model libraries. These heterogeneous reading resources included text, audio, video, images, XR spatial models, and independent AI content mini-model resources. Invalid resources that were repeatedly collected, had damaged files, incomplete content, or originated from illegal sources were removed to obtain initially screened reading resource data. Then, a multi-level classification model was adopted to classify and organize the reading resources. A three-level classification model was set up, where the first level of classification divided resource categories according to resource modality and XR spatial attributes, the second level of classification was processed according to content theme, reading scenario, and XR spatial scenario, and the third level of classification was classified according to file encoding, storage format, and model specifications to obtain multimodal reading resources after classification and organization.

3. The digital reading platform construction system supporting multimodal content interaction according to claim 2, characterized in that, The specific method for extracting data features from multimodal heterogeneous reading resources and constructing a multimodal content feature database based on the extracted data features is as follows: Based on the categorized and normalized multimodal reading resources, data features adapted to XR spatial computing and AI content model calls are extracted for each modality. Specifically, for text-based reading resources, keywords, content attributes, and length levels are extracted; for audio-based resources, duration, timbre tags, and content summaries are extracted; for video-based resources, image parameters, duration nodes, and associated text features are extracted; for image-text resources, layout, image-text ratio, and annotation information features are extracted; for XR-based resources, spatial coordinates, object IDs, and scene binding features are extracted; and for AI model-based resources, model identifiers and trusted source features are extracted. The extracted data features are normalized to obtain standardized reading resource data features. These standardized features are then hierarchically and partitioned according to the resource categories corresponding to the three-level classification model. An association index is established for each standardized feature data with the corresponding reading resource, XR spatial object, and AI model. The standardized feature data and association indexes are integrated to form a standardized multimodal content feature database adapted to XR spatial computing and trusted AI scheduling.

4. The digital reading platform construction system supporting multimodal content interaction according to claim 1, characterized in that, The specific method for performing semantic parsing on the collected interactive input information and then obtaining user interaction requirements based on the parsing results is as follows: The collected multimodal interactive input information is preprocessed to filter noise, invalid characters, and abnormal operation data. Voice information is translated into a standardized text format, and touch, XR gaze, gesture, and spatial location entry information are converted into instructional text descriptions adaptable to XR interaction forms, resulting in standardized user interaction information. This standardized user interaction information is then segmented to extract key information, including core vocabulary, operation behavior, resource reference, XR spatial objects, and coordinates. Subsequently, a digital reading scene semantic database adapted to XR spatial computing and AI reading scenarios is invoked. The extracted key information is matched with the corresponding reading interaction intent in the semantic database, and the multimodal resource type, reading scenario, and XR spatial binding relationship corresponding to the user interaction intent are labeled to form user interaction requirements.

5. A digital reading platform construction system supporting multimodal content interaction according to claim 4, characterized in that, The next step involves calling a digital reading scenario semantic database and matching the extracted key information with the corresponding reading interaction intent in the semantic database. The specific method is as follows: It utilizes a dedicated semantic database for digital reading scenarios that integrates AI capabilities, extracting key words, operational behaviors, resource type matching, and spatial object matching thresholds from the semantic database. The key word matching threshold is... The operational behavior adaptation threshold is derived from the combination of digital reading interaction scenarios and VR / MR content production and interaction scenario big data training and iteration. Based on actual tests in reading interaction scenarios and VR / MR immersive interaction scenarios, the resource type matching threshold was determined. To adapt to the full-match hard threshold required for VR / MR content production resource matching, the extracted key information is matched and calculated with the reading interaction intent in the semantic database, and the spatial object matching threshold is calculated. Bind a specific threshold to the XR space using a formula. Calculate the core vocabulary matching degree, where For core vocabulary matching, To match the number of characters, This is the total number of characters in the standard keywords. If the core vocabulary match is valid, then the formula is used to determine that the match is effective. Calculate the operational behavior adaptation degree, where For operational behavior adaptation, To match the number of operations, To extract the total number of operations, if If the operation is deemed to be consistent with the basic reading interaction intent, then according to the formula... Computational resource consistency, where For resource consistency, The number of overlapping resource types, This refers to the total number of resource types. If the resource points to the same content as the reading interaction intent, then according to the formula... Calculate the spatial object matching degree, where For spatial object matching degree, The Euclidean distance between the user interaction point and the reference coordinates of the spatial object. Bind a valid domain radius to a spatial object, if Then the spatial object is matched with the reading interaction intent, and then the formula is used to determine the relationship. Calculate the total matching score, where The total matching score. The weighting coefficients for scenario optimization are derived from scenario big data analysis and actual user interaction behavior testing. The reading interaction intent with the highest total matching score is selected as the target reading interaction intent that matches the user's digital reading interaction intent.

6. A digital reading platform construction system supporting multimodal content interaction according to claim 1, characterized in that, The specific method for building a digital reading platform framework based on a multimodal content feature database and configuring platform operation and management components is as follows: A layered digital reading platform framework is established based on a standardized multimodal content feature database. This framework comprises a data layer, an XR space layer, an interaction layer, and a presentation layer. A data interaction interface is established between the data layer and the standardized multimodal content feature database to adapt to XR content creation resource retrieval and AI model invocation. The interaction layer is connected to the multimodal interaction parsing module, and a two-way channel for receiving and responding to interaction execution commands is configured. The XR space layer is used for virtual scene rendering and spatial object binding management. The presentation layer sets up display areas for multimodal content such as text, audio, video, graphics, XR 3D animation, and AR annotations, and establishes cross-modal and XR space linkage display entry points. Furthermore, platform operation and control components are configured, including a resource scheduling unit, a spatial content scheduling unit, a data verification unit, and an operation monitoring unit. All control components are linked and adapted with the platform framework's data layer, XR space layer, interaction layer, and presentation layer.

7. A digital reading platform construction system supporting multimodal content interaction according to claim 6, characterized in that, The control components include a resource scheduling unit, a data verification unit, and an operation monitoring unit; The resource scheduling unit is used to retrieve and allocate multimodal resources of the data layer, interaction layer and presentation layer according to the platform's interactive execution instructions and multimodal resource call requirements; The data verification unit is used to verify the multimodal resources retrieved from the standardized multimodal content feature database by the data layer, and to confirm that the reading resources can be transmitted normally to the display layer and are adapted to the resolution of the corresponding display area and meet the linkage requirements of the cross-modal linkage display entry. The operation monitoring unit is used to collect the operation status of the data layer, interaction layer, and presentation layer in real time, monitor the transmission status of the data interaction interface, the reception and feedback status of interaction commands, and record them to form a platform operation log.

8. A digital reading platform construction system supporting multimodal content interaction according to claim 6, characterized in that, The specific method for establishing multimodal content display and interactive response logic by combining interactive execution instructions is as follows: The system extracts operational behaviors, multimodal resource types, target reading scenarios, and XR space object information from interactive execution commands. It retrieves matching reading resources, XR space materials, and AI content models from a standardized multimodal content feature database. Combining the target reading scenario and XR space scenario, it establishes multimodal content display logic, setting adaptation scenarios for text, audio, video, graphics, XR 3D animation, and AR annotations. Simultaneously, it establishes cross-modal linkage and XR space linkage rules to form multimodal content display logic. Based on user operation behaviors and XR space interaction behaviors, it establishes interactive response logic, setting exclusive execution rules compatible with XR interaction forms for various operation behaviors. The display logic and response logic are synchronized to the platform data layer, XR space layer, interaction layer, and display layer, and linked with the operation control components to form a complete XR immersive multimodal content display and spatial interaction response logic.