Information processing method and device, electronic equipment and computer readable storage medium
By using long sequence modeling and knowledge network construction, the system automatically integrates note information, solving the problem of manual dependence in existing technologies. It enables in-depth mining and structured presentation of multi-hop causal relationships, improving the efficiency of information retrieval and management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JRD COMM (SHENZHEN) LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-31
AI Technical Summary
Existing note-taking applications rely on manual intervention when integrating discrete note information, which makes it impossible to deeply explore multi-hop causal relationships. As a result, the integrated content lacks causal reasoning ability and has low retrieval and management efficiency.
By acquiring multiple related information, long sequence modeling is performed to obtain the semantics of nodes. Algorithms such as Transformer-XL are used to capture contextual information, infer multi-hop logical relationships, and construct a knowledge network to automatically achieve logical integration and structured presentation.
It improves the efficiency of retrieving and reusing multiple related information, realizes the intuitive presentation of multi-dimensional logical context through automated processing, reduces manual operation, and improves the coherence and accuracy of information.
Smart Images

Figure CN122489601A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to an information processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] In digital office and learning scenarios, note-taking applications have become core tools for users to record information and integrate content. To improve the efficiency of retrieving and managing fragmented notes, it is necessary to integrate and process discrete note information. However, mainstream note-taking applications rely heavily on manual intervention for integrating discrete note information, requiring users to manually add tags or create content links for each note. Their intelligent processing capabilities for note information are still in a rudimentary stage, typically only able to simply aggregate multiple notes as static attachments or independent pages, unable to deeply mine and identify potential multi-hop causal relationships between note information, ultimately resulting in the integrated content lacking causal reasoning ability. Summary of the Invention
[0003] This application provides an information processing method, apparatus, electronic device, and computer-readable storage medium that can automatically and deeply mine multi-hop logical relationships between related information, and realize the logical integration and structured presentation of multiple related information.
[0004] In a first aspect, embodiments of this application provide an information processing method, including: Obtain multiple related information about the target content; Long-sequence modeling is performed on multiple pieces of associated information to obtain the semantics of multiple nodes; the semantics of the nodes carry contextual information. The semantics of multiple nodes are inferred to obtain multiple thought processes corresponding to the multiple nodes; Based on multiple thought processes, a knowledge network for the target content is constructed; the nodes in the knowledge network are labeled with logical relationships.
[0005] Secondly, embodiments of this application provide an information processing apparatus, including: The information acquisition module is used to acquire multiple related information about the target content; The semantic acquisition module is used to perform long sequence modeling on multiple related information to obtain the semantics of multiple nodes; the semantics of the nodes carry contextual information. The semantic reasoning module is used to reason about the semantics of multiple nodes to obtain multiple thought processes corresponding to the multiple nodes; The network construction module is used to construct a knowledge network of the target content based on multiple thought processes; the nodes in the knowledge network are labeled with logical relationships.
[0006] Thirdly, embodiments of this application also provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the information processing method described above.
[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the information processing method described above.
[0008] Fifthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.
[0009] The embodiments of this application have the following beneficial effects: It can acquire multiple related information about the target content, perform long-sequence modeling on these related information according to a set time window, and obtain the semantics of multiple nodes, thereby preserving the temporal correlation of the related information and giving the generated node semantics contextual attributes. Semantics carrying contextual information can improve the coherence and accuracy of subsequent reasoning. Reasoning on the semantics of multiple nodes yields multiple thought streams corresponding to multiple nodes, which can uncover potential multi-hop logical relationships between the semantics of nodes. These multiple thought streams can present the multi-dimensional logical context of the target content. Based on these multiple thought streams, a knowledge network of the target content is constructed, and the nodes in the knowledge network are labeled with logical relationships. This can transform discrete thought streams into a structured knowledge network, intuitively presenting the logical relationships between related information. Without manual operation, it can automatically complete the logical integration and structured presentation of related information, thereby integrating and improving the retrieval and reuse efficiency of multiple related information of the target content. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of the steps of an information processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the thought process provided in one embodiment of this application; Figure 3 This is a schematic diagram of the principle architecture of an information processing method provided in an embodiment of this application; Figure 4 This is a timing diagram of an information processing method provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an information processing device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0013] Note-taking apps have become core tools for users' daily recording, work, and study organization. Mainstream note-taking apps like OneNote, Evernote, Notion, and Obsidian generally support recording various materials such as structured text, images, and attachments. They also continuously optimize knowledge organization features such as tags, links, and hierarchical structures to improve the efficiency of information retrieval and management. Hierarchical knowledge organization refers to organizing note content into layers and categories, giving discrete note information clear categories and hierarchical relationships.
[0014] To meet the growing demand for information gathering, some note-taking apps support multimodal input, allowing users to input photos, handwriting, audio recordings, drawings, and other multimodal note information. They also extend this to include multi-document collaboration and search functions. However, most note-taking apps still have a relatively rudimentary approach to the integration and intelligent processing of multimodal note information. They typically treat different types of note information as static attachments or independent pages, failing to achieve efficient cross-modal content linking and unified understanding.
[0015] To gain a deeper understanding of the multi-hop logical relationships between multiple note information and to uniformly organize multiple note information, one embodiment of this application, such as Figure 1 As shown, an information processing method is provided. Although the logical order is illustrated in the step diagram, in some cases, the steps shown or described may be performed in a different order than that shown in the figures. These will be described in detail below. It should be noted that the order of description of the following embodiments is not intended to limit the priority of the embodiments.
[0016] according to Figure 1The information processing method shown includes at least steps S110 to S140, which are described in detail below: In step S110, multiple related information of the target content is obtained.
[0017] The target content can be information from various scenarios, such as conference content, academic papers, videos, courseware, or project plans. Multiple related information items can be supplementary information that has semantic, logical, or contextual connections to the target content. These related information items can be multimodal, including images, text, audio, and handwritten / drawn content.
[0018] For example, when the target content is conference content, related information may include, but is not limited to: multiple notes taken by one or more participants, audio recordings of each speaker's presentation, and / or background information related to the conference. When the target content is an academic paper, related information may include, but is not limited to: user-entered notes, excerpts of references, and / or supplementary literature in related fields. When the target content is a film or video, related information may include, but is not limited to: reviews by one or more film critics, video segments or keyframes, video scripts, and / or plot summaries. When the target content is courseware, related information may include, but is not limited to: photos of classroom blackboard writing, audio or video clips of lectures, multimodal notes taken by students regarding the courseware, and / or classroom Q&A records. When the target content is a project proposal, related information may include, but is not limited to: meeting minutes of the project proposal, market analysis reports, research data, and / or technical documents.
[0019] The associated information may carry time information, which may include, but is not limited to: the generation time of the associated information (such as the time of note recording, audio recording, hand-drawn creation, and publication time of the literature), the collection time of the associated information (such as the time of taking a photo of classroom blackboard writing, and the time of collecting survey data), the time of the associated information relative to the target content (such as the time of a video clip in a film or television video), and / or the time of the associated information annotated by the user.
[0020] In step S120, long sequence modeling is performed on multiple related information to obtain the semantics of multiple nodes.
[0021] Long sequence modeling refers to methods for capturing long-range correlations and contextual dependencies between data, thereby extracting data semantics. Long sequence modeling can overcome the computational and storage bottlenecks of traditional models, transforming scattered relational information into structured semantic nodes, enabling the effective long-distance transmission and integration of information across time periods. The semantics of a node carry contextual information, which refers to the set of historical inputs, subsequent content, and / or cross-segment relational information associated with that node. This contextual information can be used to assist the model in understanding the semantics and logical relationships of the current node. By performing long sequence modeling on relational information, multi-turn, multi-hop information between multiple relational pieces of information can be captured, which helps in subsequent reasoning of the logical relationships between nodes.
[0022] In one embodiment, a set time window can be obtained, with each time window considered as a node. The semantics of the node are determined based on the semantics of the associated information within each time window. Long sequence modeling can divide multiple pieces of associated information into multiple time windows based on the set time windows. For each time window, the hidden state of that time window can be calculated. The hidden state is a vector representation generated by the model after processing the input at a certain position in the sequence, containing current input information and historical context information. It is used to pass context to assist in the calculation of subsequent sequence positions. The hidden state of this time window is cached as historical memory information and passed to the next time window. In each subsequent time window, the semantics of the current time window are calculated by combining the historical memory passed from the previous time window with the local information of the current time window (the various associated information within the current time window). The semantics of the current time window is then used as new memory and passed to the next time window for calculation. In this way, the semantics of each node extracted through long sequence modeling can contain context information.
[0023] In this embodiment, long-sequence modeling of related information can be achieved based on the Transformer ExtraLong (Transformer-XL) deep learning algorithm for long-sequence dependency modeling. First, the various related information pieces are context-bounded. Then, Transformer-XL can effectively capture semantic information across time windows through a segmented loop mechanism and relative position encoding. Transformer-XL can input related information carrying time information into the model in temporal segments according to a set time window, utilizing the algorithm's memory mechanism to retain the semantics of preceding segments, achieving coherent transmission of information across segments. Simultaneously, Transformer-XL can employ an attention mechanism to calculate the contextual dependencies between related information, transforming scattered related information into structured nodes and obtaining the semantic information of each node. Transformer-XL demonstrates excellent performance in event evolution, inference chain reuse, and temporal inference capabilities.
[0024] In this embodiment, the contextual logic of associated information can also be captured based on long sequence modeling algorithms such as Longformer, Reformer, or BERT-Long to obtain the semantics of each node.
[0025] In another embodiment, each piece of related information can be defined as a node, and the semantics of that related information can be defined as the semantics of the corresponding node. Long sequence modeling can calculate the hidden state of each piece of related information, cache the hidden state as historical memory information, and pass it to the next piece of related information. For each subsequent piece of related information, the semantics of the current piece of related information is calculated by combining the historical memory passed from the previous piece of related information with the current piece of related information, and the semantics of the current piece of related information is used as new memory to continue to be passed to the next piece of related information for calculation. Thus, the semantics of each piece of related information extracted through long sequence modeling can contain contextual information.
[0026] In step S130, semantic reasoning is performed on multiple nodes to obtain multiple thought flows corresponding to multiple nodes.
[0027] A thought process flow is a semantic chain formed by reasoning about the semantics of nodes and connecting multiple nodes. It contains implicit logical relationships and can present the temporal evolution and logical connection between nodes.
[0028] Multi-hop reasoning can be performed on the semantic, contextual, and temporal information of nodes to uncover potential connections and evolutionary patterns between nodes, resulting in multiple thought streams. Each thought stream can contain multiple nodes, and these nodes possess potential logical relationships. Logical relationships can include, but are not limited to, at least one of causal relationships, temporal relationships, and contextual relationships. Contextual relationships refer to the associations between the semantics of different nodes based on temporal sequence, scene background, and / or semantic continuity, providing contextual support for understanding the complete meaning of information. During the reasoning process, multiple thought streams can be derived through multiple rounds of reasoning based on the semantic relationships between various related information within the same time window, as well as the semantic relationships between different time windows. Each thought stream corresponds to a complete semantic logical chain.
[0029] In step S140, a knowledge network of the target content is constructed based on multiple thought processes.
[0030] A knowledge network is a structured and visualized knowledge system that integrates multiple related pieces of information, and can fully present the logical relationships between these related pieces of information. A knowledge network can include nodes and edges. Nodes obtained from long sequence modeling can be identified as nodes in the knowledge network, and edges between nodes can be determined based on the logical relationships between them, thus obtaining the knowledge network.
[0031] Based on the identified multiple thought processes, a knowledge network can be formed by integrating the nodes according to semantic and temporal relationships. Each node in the knowledge network is labeled with its logical relationships and can also include the temporal information of the original associations between nodes, thus clearly presenting the inherent connections between them.
[0032] By employing the technical solution of this application embodiment, multiple related information of the target content can be obtained. Long-sequence modeling of these related information is performed according to a set time window to obtain the semantics of multiple nodes, thereby preserving the temporal correlation of the related information and enabling the generated node semantics to possess contextual attributes. Semantics carrying contextual information can improve the coherence and accuracy of subsequent reasoning. Reasoning on the semantics of multiple nodes yields multiple thought processes corresponding to those nodes, which can uncover potential multi-hop logical relationships between the semantics of the nodes. These multiple thought processes can present the multi-dimensional logical context of the target content. Based on these multiple thought processes, a knowledge network of the target content is constructed, and the nodes in the knowledge network are labeled with logical relationships. This transforms discrete thought processes into a structured knowledge network, intuitively presenting the logical relationships between related information. No manual operation is required; the logical integration and structured presentation of related information can be automatically completed, thereby improving the efficiency of retrieval and reuse of multiple related information of the target content.
[0033] Based on the above technical solution, as an embodiment, the associated information can be multimodal information. To facilitate machine understanding, the associated information can be machine-understandable information obtained by processing the original associated information. Multiple original associated information pieces of the target content can be obtained, and these multiple original associated information pieces can include, but are not limited to, at least one modality among images, image-based text, audio, and text. The multiple original associated information pieces are processed according to the processing methods corresponding to their respective modalities to obtain multiple machine-understandable associated information pieces.
[0034] When the original associated information is in image mode, image recognition can be performed on the original associated information to obtain image information; when the original associated information is in image-text mode, optical character recognition can be performed on the original associated information to obtain first text associated information; when the original associated information is in audio mode, automatic speech recognition can be performed on the original associated information to obtain second text associated information. Machine-understandable associated information includes the image information, the first text associated information, and / or the second text associated information.
[0035] The system can obtain raw association information from multiple modalities. The raw association information for the image modality can be a visual medium, such as a photograph of a whiteboard, a hand-drawn sketch, courseware illustrations, or a real-life scene photograph. This information can include visual information such as objects, structures, and / or graphics within the image. The information for the image-text modality can be text content presented as an image, such as scanned book pages, handwritten manuscript pages, or screenshots of text paragraphs, where the text information is attached to the image medium. The information for the audio modality can include, but is not limited to, recorded and produced audio. The information for the text modality can be structured information directly in the form of text symbols, such as manually entered notes, copied document fragments, and / or edited annotations. This text modality information can be directly parsed by the machine without additional processing.
[0036] The modalities of each original relational information can be automatically identified through format checking or large-scale modeling. Based on the identified modalities, a content understanding AI engine calls different models to process the original relational information of each modality, resulting in processed relational information for each modality. This processed relational information is machine-understandable. Machine-understandable information refers to information that, after structured processing or format standardization, can be parsed by computers, have its features extracted, and be transformed into processable data.
[0037] When the original association information is in the image modality, scene recognition or object detection algorithms (such as YOLOv7, ViT Vision Transformer) can be used to identify events and structural relationships within the image, extracting key information from the image into image nodes and their attribute features, thus obtaining machine-readable image information. When the original association information is in the image-text modality, OCR models (such as CRNN or Transformer-based OCR) can be used to extract and transcode the text content in the image, obtaining standardized first text association information. When the original association information in the image-text modality is a handwritten formula or chart, image segmentation and structured parsing can be used to format logical nodes, labels, arrows, etc., in the original association information into a usable structure for machine understanding. When the original association information is in the audio modality, SOTA (State-of-the-Art) ASR models (such as DeepSpeech, wav2vec2.0, or API services) can be used to transcribe the original association information into text fragments in real time with high accuracy, obtaining second text association information. Optionally, when processing the original association information of audio modalities, scene-customized hot words and / or technical terms can be added to expand the recognition rate of the original association information.
[0038] By processing the original association information of different modalities, machine-understandable association information can be formed, providing data support for subsequent long sequence modeling.
[0039] The technical solution adopted in this application can support the analysis and processing of multimodal original association information. By processing the original association information of different modalities, the original association information of multimodalities can be transformed into semantically consistent and structured association information, thereby improving the accuracy of information processing and facilitating the subsequent organic integration of multiple association information. The machine-understandable data format can ensure the accuracy of long sequence modeling and improve processing efficiency.
[0040] Based on the above technical solution, as an embodiment, when each time window is regarded as a node, long sequence modeling is performed on multiple related information to obtain the semantics of multiple nodes. This may include: obtaining a set time window; determining at least one related information corresponding to each time window according to the time information of the time window and each related information; performing multimodal fusion on the related information corresponding to each time window to obtain the fused information corresponding to each time window; for each time window, performing fusion reasoning on the fused information of the time window and the semantics of the previous time window to obtain the semantics of the time window, and determining each time window as a node.
[0041] The length of the time window can be set according to actual needs; optionally, the time window can be a fixed duration, or it can be dynamically adjusted according to the time distribution of associated information to divide multiple pieces of associated information into appropriate time windows. Each time window can include a start time and an end time. The time window to which each piece of associated information belongs can be determined based on the range of the time information carried by each piece of associated information, thereby determining the associated information corresponding to each time window. If a time window does not contain any associated information, that time window can be directly discarded.
[0042] For each time window, multimodal fusion can be performed on the various related information within that window. The unstructured raw related information has already been transformed into machine-understandable related information using various image recognition, OCR, ASR, and other technologies. By integrating the image features, textual semantics, and other information from multiple modalities through a multimodal fusion algorithm, the fused information for that time window can be obtained. This fused information can eliminate the semantic fragmentation problem caused by the heterogeneous formats of related information from different modalities.
[0043] When modeling long sequences of related information, each time window can be defined as a node. For each node, its hidden state is calculated, and this hidden state is cached as historical memory and passed to the next node. At each subsequent node, the semantics of the current node are calculated by combining the historical memory passed from the previous node with the fused information of the current node. This semantics is then used as new memory and passed to the next node to calculate the semantics of the next node. In this way, the semantics of each node extracted through long sequence modeling can contain contextual information. Furthermore, the start and end times of each time window can be used to determine the temporal information of the node corresponding to that time window.
[0044] By adopting the technical solution of this application embodiment, multimodal association information is divided and classified through time windows, which can avoid semantic interference between information across time periods, improve the temporal accuracy and logical coherence of information processing when modeling long sequences, and reduce the computational complexity of long sequence modeling; fusing multimodal data within the time window can eliminate modal differences and ensure the semantic integrity of the fused information; and performing fusion reasoning based on the semantics of the previous time window can improve the semantic coherence and obtain the contextual information of the nodes.
[0045] Based on the above technical solution, as an example, reasoning about the semantics of multiple nodes to obtain multiple thought flows corresponding to multiple nodes may include: reasoning about the contextual information of the semantics of multiple nodes to obtain cross-item semantic association features; determining the semantic arrangement logic between multiple nodes based on the cross-item semantic association features; and connecting multiple nodes according to the semantic arrangement logic to obtain thought flows.
[0046] An entry is an independent unit of information obtained by dividing related information according to time windows or scenarios. It is the basic unit for carrying semantic information and conducting cross-entry association analysis. For example, it can be divided according to time windows, such as integrating the meeting speech recordings and simultaneous handwritten notes from 9:00 to 9:10 into one entry, and the meeting speech recordings and simultaneous handwritten notes from 9:10 to 9:20 into another entry; it can also be divided according to scenarios, such as integrating the illustrations and lecturer audio related to Chapter 1 of the courseware into one entry, and the illustrations and lecturer audio related to Chapter 2 of the courseware into another entry.
[0047] The semantic contextual information carried by nodes can include relationships such as temporal sequence, scene background, and semantic continuity between items. Deep learning algorithms or large models can be used to infer the contextual information of each node to determine cross-item semantic association features between different items. Cross-item semantic association features reflect the deep semantic association attributes of different nodes and can be obtained through contextual reasoning of multimodal association information scattered across different items. During the reasoning process, deep learning algorithms (such as Longformer, Transformer-XL) or large models can be used to mine the deep associations between corresponding nodes of different items, breaking through the limitations of single-item information and extracting cross-item semantic association features. Cross-item semantic association features may include, but are not limited to, at least one of the following: thematic evolution trends, event development rhythm, and signal flow paths of multiple items.
[0048] The theme evolution trend refers to the trajectory of core issues as time progresses or scenarios change, reflecting shifts in information focus. It can be determined by analyzing the semantic connections between different items to extract patterns, thus identifying the evolution trend of the target content from the initial topic to subsequent topics. For example, if the target content is a project workshop, and the initial items revolve around the theme of "technical problems," this theme gradually transitions to "solutions," and eventually becomes "implementation plans." This demonstrates an evolution trend from problem analysis to solution implementation.
[0049] The pacing of an event refers to the frequency, intervals, and speed of events related to the target content as time progresses. Event development can be determined by the temporal information and semantic relationships carried by the entries, thus establishing the pacing. For example, if the target content is courseware, and the first chapter has fewer related information while the second chapter has more, then the pacing of the event can be determined to be gradually accelerating.
[0050] Signal flow path refers to the trajectory of information propagation between different items. This trajectory can be determined by mining the semantic connections between nodes, identifying the information's path from its source to subsequent carriers. For example, if the target content is a movie, and the associated information is its release date, and this associated information travels from the movie's announcement text to the film critic's analysis and finally to the audience's comments, then the corresponding signal flow path is from the movie's source to the professionals and finally to the user.
[0051] Based on cross-item semantic association features, the semantic arrangement logic between multiple nodes can be determined. Different types of cross-item semantic association features can correspond to different semantic arrangement logics. When the cross-item semantic association features focus on the theme evolution trend, the semantic arrangement logic of each node can be determined according to the order of theme progression, branching, or aggregation, such as arranging nodes in a paper scenario according to the theme structure of "research background - core argument - evidence - conclusion". When the cross-item semantic association features focus on the rhythm of event development, the semantic arrangement logic of each node can be determined according to the time sequence, such as determining the semantic arrangement logic of each node according to the order of speeches at a conference. When the cross-item semantic association features focus on the signal flow path, the semantic arrangement logic of each node can be determined according to the direction and hierarchical relationship of information transmission, such as determining the semantic arrangement logic of each node from video frames to bullet comments to film reviews in a video scenario.
[0052] The nodes can be arranged and connected according to a defined semantic arrangement logic to obtain a thought process. During the connection process, scattered nodes can be used as key nodes in the thought process, and the association attributes and cross-item semantic association features between the node and other nodes in the thought process are preserved, thus presenting a complete path except for knowledge transfer and supplementation. The final generated thought process can include the core semantics of multimodal association information, as well as highlight the logical hierarchy through orderly arrangement.
[0053] The technical solution adopted in this application can visualize the semantic logic between nodes through thought process flow, connect scattered nodes, intuitively show cross-item semantic relationship features, transform fragmented node information into an ordered semantic chain, thereby presenting the inherent relationship of the target content; and can reduce the complexity of subsequent knowledge network integration, and improve the efficiency and accuracy of knowledge system construction.
[0054] Based on the above technical solution, as an example, constructing a knowledge network of target content based on multiple thought processes can include: reasoning about each thought process to determine the category and logical relationship of each node in the thought process; establishing edges between each node according to the category and logical relationship of each node in each thought process, and marking the logical relationship on each node to obtain the knowledge network.
[0055] A thought process flow is a semantic chain that connects multiple nodes, formed by reasoning about their semantics, and contains implicit logical relationships. Therefore, thought processes flow can be further reasoned and analyzed to determine the categories and logical relationships of the nodes. Figure 2 This is a schematic diagram of the thought process provided in an embodiment of this application; as shown... Figure 2As shown, nodes in a thought process flow can have different categories. The categories of nodes in a thought process flow can include, but are not limited to, at least one of: inspiration, fact, material, conclusion, hypothesis, and evidence. The logical relationships between nodes can include, but are not limited to, at least one of: causal relationship, temporal relationship, and contextual relationship. The category of nodes in a thought process flow can be determined by their semantic attributes. For example, in a meeting thought process flow, a solution concept can be categorized as an inspiration node, industry data as a fact node, case references as material nodes, and the final solution as a conclusion node. The logical relationships between nodes can be determined based on semantic connection and contextual reasoning. For example, an inspiration node and its corresponding final solution node form a reasoning relationship, and sequential event nodes form a temporal relationship.
[0056] A knowledge network can be constructed by processing thought processes using incremental graph neural networks (GNNs). Each node in the thought process can be identified as a node in the knowledge network, and edges between nodes can be determined based on their logical relationships, thus achieving structured connections between nodes. Edges between nodes can be automatically established according to logic such as "fact—inference—conclusion—evidence." When constructing a knowledge network based on thought processes, the relationships between nodes within each thought process and between nodes in different thought processes can be considered. For example, conclusion nodes on the same topic in different thought processes can be connected through contextual relationships to form the final knowledge network. Logical relationships between nodes that are directly or indirectly connected to each node can be labeled, as well as the category information of the node, thus visually presenting the logical relationships between nodes.
[0057] By adopting the technical solution of this application embodiment, dividing nodes into different categories can help clarify the logical relationship between nodes, avoid confusion in information association, make the implicit logical relationship contained in the thought process clearer, and thus form a knowledge network for the entire thought process, realizing the fusion and interoperability of information from multiple thought processes, and providing a structured data foundation for subsequent tasks such as information retrieval, intelligent reasoning and content generation.
[0058] Based on the above technical solution, as an embodiment, the information processing method may further include: when new related information is obtained, obtaining the semantics, content summary and tags of the new related information; generating a new node based on the semantics, content summary and tags of the new related information, and determining the category and attributes of the new node; evaluating the logical relationship between the new node and other nodes based on the category and attributes of the new node, and establishing an edge between the new node and other nodes based on the logical relationship.
[0059] Incremental GNNs can perform local computation and parameter updates on the affected subgraphs or nodes when new nodes, edges, node attributes, and / or edge attributes are added, resulting in an updated knowledge network. Incremental GNNs do not require full graph retraining when updating the knowledge network, enabling efficient and real-time model iteration and graph representation learning.
[0060] When new original association information is acquired, its modality can be determined, and the new association information can be processed according to the corresponding modality to obtain new association information that can be understood by the machine. Incremental GNN can extract the semantics of the new association information and generate a content summary of the new association information. According to the topic, scenario, or modality of the new association information, the tags of the new association information can be determined. For example, if the new association information is supplementary audio of a meeting speech, after ASR transcription and semantic extraction, the speech semantics and summary, as well as tags such as "project progress," "meeting supplement," and "audio," can be obtained, providing basic data for the generation of subsequent nodes.
[0061] Incremental GNNs can generate new nodes based on the semantics, summaries, and labels of new associated information, or update existing nodes to obtain new nodes. Combining the node category system, the category to which the new node belongs is determined, and the node's attributes are identified. Node attributes can include at least one of the following: content text, source, time, and modality type; the content text can be the semantic content of the node. Based on the category and attributes of the new node, the logical relationships between the new node and other nodes are inferred, thereby establishing or updating the edges between the new node and other nodes, and labeling these logical relationships on the new node.
[0062] In one embodiment, AI can automatically recommend and complete causal chains, enrich node labels, and add keywords to a knowledge network.
[0063] The technical solution adopted in this application embodiment, which updates the knowledge network based on incremental GNN, can be adapted to incremental information access scenarios, realize the dynamic update of the knowledge network in an efficient and real-time manner, improve the update efficiency of the knowledge network, and provide the latest structured data foundation for subsequent tasks such as information retrieval, intelligent reasoning and content generation.
[0064] Figure 3 This is a schematic diagram illustrating the principle architecture of an information processing method provided in an embodiment of this application; as follows: Figure 3As shown, the multimodal input acquisition layer can collect raw multimodal correlation information. The content understanding AI engine layer processes this raw correlation information according to its corresponding modality to obtain machine-understandable correlation information. The logical reasoning and knowledge network layer then performs semantic extraction and logical reasoning on the correlation information to construct a knowledge network. On one hand, the visualization and interaction layer can visually display the knowledge network, allowing users to intuitively understand the deep logical relationships within the correlation information. It also supports user interaction with the knowledge network, such as retrieval, manipulation, and feedback. On the other hand, the content generation and output layer supports direct content generation and output based on the knowledge network, such as generating audio, text, or video content from the knowledge network.
[0065] The multimodal acquisition layer can include text input, image input, handwriting / drawing input, and audio input modules. The text input module supports direct editing and copying of standard text data. The image input module supports taking photos with a mobile phone camera / importing from albums, including photos of notes, book pages, and whiteboards. The handwriting / drawing input module supports integration with tools such as touchpads or styluses to capture unstructured expressions such as user-generated drawings, arrows, and mind maps. The audio input module supports real-time voice recording, meeting / discussion recording, and simultaneous voice note-taking. The multimodal acquisition layer can automatically sort the raw, related information according to the data modality and perform basic formatting and noise reduction locally, ensuring the consistency of the data structure fed into the content understanding AI engine layer.
[0066] Based on the above technical solution, as an example, from a visualization perspective, on the one hand, it can visualize the thought process, knowledge network and / or logical relationship, and on the other hand, it can also support visualization operations on the knowledge network.
[0067] It can visualize thought processes in the form of linear semantic chains, which can be displayed as thought flow views, causal chain views, or timeline views. It can also visualize knowledge networks, supporting zooming and hierarchical expansion. Nodes can be categorized by different colors or shapes, and the logical relationships corresponding to edges can be distinguished. For example, edges corresponding to causal relationships can be set to red, and edges corresponding to temporal relationships can be set to green. The logical relationships corresponding to nodes can be directly visualized, for example, through pop-ups or sidebar annotations on the visualized knowledge network nodes, helping users quickly identify the logical relationships between nodes.
[0068] The visualization interface supports user operations on the knowledge network, including but not limited to at least one of drag-and-drop, editing, and merging. Drag-and-drop allows adjusting node positions and reconstructing the network layout to suit individual user viewing habits; editing allows modifying node categories, attributes, and logical relationships between edges, as well as correcting semantic annotation deviations; merging merges semantically redundant or closely related nodes into a single node, removes redundant edges, and optimizes the network structure. All editing operations are synchronized back to the knowledge network in real time, ensuring consistency between the visualization and the underlying data, thereby optimizing the knowledge network through user actions and meeting user needs.
[0069] The technical solution adopted in this application can reduce the user's understanding threshold through visualization. The intuitive presentation of the visualization can help users quickly understand relevant information. Visual operation can optimize the knowledge network according to user needs. The real-time linkage between the visualization operation and the underlying data of the knowledge network can improve the network optimization efficiency.
[0070] Building upon the aforementioned technical solutions, as an example, to achieve efficient retrieval and utilization of structured information within the knowledge network, a search interface can be provided to users. This search interface supports natural language question-and-answer interaction. Users do not need to master specific search syntax; they can directly input search information using colloquial and everyday natural language. The system uses natural language processing technology to parse the semantics of the user's input, thereby extracting the user's search needs and transforming them into search instructions recognizable by the knowledge network. For example, users can directly input search information such as "What are the reasons that led to this conclusion?" or "What are the most closely related inspirations to event X?" to conduct searches, thus significantly reducing the threshold for search operations and adapting to various user scenarios such as meeting debriefing, courseware organization, and project research, thereby improving the convenience of interaction.
[0071] After obtaining the search information input by the user through the search interface, a multi-dimensional search can be performed on the node attributes and logical relationships of the knowledge network based on this information. This multi-dimensional search yields search results. The multi-dimensional search includes at least one of keyword search, semantic similarity search, and logical relevance search. Keyword search can quickly locate nodes and edges containing target keywords by matching node tags, content text, and other text information. Semantic similarity search can determine semantic associations based on deep learning algorithms, returning nodes that are semantically similar to the search information (similarity greater than a similarity threshold). Logical relevance search can filter nodes that have a direct or indirect logical connection with the search target based on the various logical relationships included in the knowledge network. Multi-dimensional search methods can be used individually or in combination, displaying search results sorted by relevance, and can also include node categories, logical relationships, and sources, facilitating users to quickly filter effective information. In this way, multi-dimensional matching can overcome the limitations of single-dimensional search, providing users with an efficient and accurate channel for information acquisition.
[0072] Based on the above technical solution, as an example, relevant outputs for the target content can be automatically generated according to the knowledge network. The knowledge network can be used as input to an Artificial Intelligence Generated Content (AIGC) model to automatically generate relevant outputs for the target content. These relevant outputs may include, but are not limited to, audio, graphic, text, or video content; specifically, they may include at least one of dynamic web page content, scrolling subtitle animations, presentation documents, and short video scripts.
[0073] Specifically, AIGC models can ensure the accuracy, logic, and relevance of output content by relying on the semantics, logical relationships, and attribute information of nodes in the knowledge network. AIGC models can deeply analyze the topic evolution trends, node categories, and interconnections within the knowledge network, thereby generating relevant output. For example, AIGC models can automatically invoke text-to-speech (TTS) models based on the links in the knowledge network, automatically add images, automatically generate subtitles, and generate scripts according to the rhythm, transforming the knowledge network into audio-visual-text hybrid content or explanatory videos. Users can customize the style, speaking speed, and structural organization to meet diverse teaching, knowledge sharing, and review scenarios.
[0074] By adopting the technical solution of this application embodiment, relevant outputs are generated based on the knowledge network, which can enrich the content output format, adapt to different scenario needs, improve content production efficiency, reduce manual creation costs, and meet users' content output needs.
[0075] Figure 4 This is a timing diagram of an information processing method provided in an embodiment of this application; see reference. Figure 4 The acquisition end can obtain the user's input multimodal raw association information and transmit it to the AI understanding engine. The AI understanding engine can call the models corresponding to different modalities to process the raw association information of different modalities, obtaining processed machine-understandable association information. Long sequence modeling algorithms can be used to model the association information in long sequences, obtaining the semantics of each node, which carries contextual information. Through semantic reasoning and analysis of the nodes, multiple thought processes can be constructed, thereby building a knowledge network. Nodes in this knowledge network can be labeled with causal relationships, temporal relationships, and / or topic information. The generated knowledge network is saved to a knowledge network graph library. The knowledge network can be visualized at the output end, and users can perform visual operations on the visualized knowledge network, including searching, questioning, editing, and dragging. Audio, text, or video content can be generated from the knowledge network with one click, and the generated content can be visualized at the output end.
[0076] The technical solution of this application supports a unified underlying framework for multimodal heterogeneous inputs such as images, image-based text, audio, and text. It integrates processing modules for different modalities, including OCR, ASR, and image AI, enabling automatic parsing, semantic understanding, and cross-modal information aggregation of original related information from different modalities. Based on the Transformer-XL model, it can complete long text context modeling across time sequences and topics, accurately tracking the evolution of inspiration and knowledge nodes to obtain a thought process flow. Incremental GNN technology is introduced to dynamically transform fragmented knowledge into a knowledge network, automatically annotating logical relationship links and achieving full traceability of the reasoning process. By integrating scattered information into a logically coherent thought process flow or knowledge network, it supports the visualization of logical relationships such as temporal evolution and causal reasoning. It also provides auxiliary functions such as natural language semantic retrieval, AI intelligent question answering, and traceable recall, and supports one-click conversion of the knowledge network into a visually appealing video or dynamic content with mixed audio, image, and text, facilitating knowledge reuse and output.
[0077] In one embodiment, each model parameter possesses high adaptability and self-optimization capabilities; each AI model and data structure can be optimized and customized according to the actual domain and user preferences. The OCR, ASR, and image understanding models support customized training based on user data, enabling targeted optimization of vocabulary recognition accuracy in specialized domains; the Transformer-XL model can adapt to the continuous expansion needs of the user's knowledge base while automatically reclaiming and cleaning up expired content; the GNN model supports efficient management of high-dimensional attribute nodes and can achieve self-organized dynamic adjustment of complex link relationships.
[0078] In terms of scalability, it can provide flexible multi-terminal extension interfaces, which not only support multi-platform input and data synchronization of API, web, mobile and desktop terminals, but also seamlessly connect with external knowledge bases, calendars, project management and collaboration platforms and other third-party systems, fully meeting the needs of functional expansion and data interoperability in multiple scenarios.
[0079] The technical solution adopted in this application realizes a closed-loop process of collecting fragmented related information—automatic AI parsing—dynamic generation of knowledge graphs—causal link annotation—human-computer interaction—multimedia content review. It employs mainstream AI technologies and a platform architecture, ensuring both efficient multimodal integration and a high degree of intelligence and customization for knowledge tracing and reuse. The technical solution adopted in this application transforms various materials into semantically connected and structured "knowledge nodes" through a unified multimodal acquisition architecture and a deep AI analysis chain, achieving the organic integration of fragmented information flow. Relying on the synergistic effect of Transformer-XL and incremental GNN, it can automatically sort out the long-term context and reasoning links of fragmented content, automatically associating facts, materials, reasoning, inspiration, conclusions, etc. with logical and causal relationships, allowing users to obtain the source context of the "knowledge flow" without manual step-by-step sorting. Incremental GNN can continuously adjust the network structure, and new related information can be quickly attached, realizing automatic determination of node attributes and spontaneous growth of strong and weak correlation links, greatly enhancing the "long-term vitality" and upgrade capability of the note knowledge system. At the same time, it supports graph chain tracing, causal questioning, and hierarchical knowledge visualization, helping users to recall key logics such as "why did I think of this" and "which evidence leads to this conclusion" in one sentence, significantly improving the efficiency of knowledge application and innovative creation. It can also generate explanatory videos through custom scripts, audio-visual text mixing, and AI dubbing, with AI connecting materials, conclusions, and reasoning paths, directly transforming the knowledge network into disseminable multimedia content.
[0080] The technical solution adopted in this application can greatly reduce the threshold for knowledge organization and review, realize the automatic organization of fragmented resources and the automatic formation of causal chains, and allow users to quickly build, recall and reuse knowledge systems without tedious editing. This solution can efficiently support the innovation and dissemination of complex knowledge, and provide great convenience for scenarios such as scientific research, project promotion, innovative design, and subject teaching that require tracing the source of reasoning and the flow of evidence. At the same time, it has multimodal and full-scene compatibility, and can integrate inspiration and evidence materials in multiple forms such as photography, recording, handwriting, and text, covering unstructured input scenarios such as meeting minutes, classroom notes, and inspiration stenography. It can also enhance the intelligence level of the entire chain of "finding information → organizing ideas → outputting content" by combining natural language interaction and graph link AI reasoning, and improve content retrieval and decision support capabilities. In addition, this solution makes it easy for creators, lecturers and teams to quickly transform personal or team knowledge into outputs such as training, lectures and online sharing, broaden the scope of knowledge reuse and content dissemination, and help improve personal brand and organizational effectiveness.
[0081] Alternatively, lightweight deep learning models (such as LSTM and GRU alternative variants) can be used to process and label the original association information, but their functionality is limited and they are difficult to support complex knowledge networks and multimodal connections.
[0082] Alternatively, the semantics of nodes can be extracted using RNN convolutional structures or information retrieval-weighted tag systems, and the nodes can be linked together, but it is difficult to automatically reason about knowledge evolution and causal chains.
[0083] Alternatively, a simple knowledge network can be built using ordinary relational databases and weakly manually structured tags / hyperlinks, but this can easily lead to hierarchical confusion, relationship distortion, and a heavy burden of manual maintenance over time.
[0084] Alternatively, supervised causal discovery algorithms (such as Bayesian networks and Granger causal analysis) can be used to track the relationships between knowledge nodes, but these are mostly limited to engineering applications with simple structures and large amounts of data, and are difficult to adapt to the needs of flexible and dynamic expansion of personal knowledge systems.
[0085] Alternatively, simple methods such as exporting web pages, batch compiling presentation documents, or using third-party tools to split fragments and then synthesize them can be used to generate audio, image, text, or video content. However, it is difficult to achieve a high level of customization that allows for dynamic tracking of audio, image, and text links and one-click automatic generation.
[0086] Alternatively, the deep learning model can be replaced with a locally run lightweight inference engine, or the weights of the incremental GNN network can be simplified to adapt to the edge computing environment, but this may reduce the overall level of intelligence and the effectiveness of knowledge organization.
[0087] Alternatively, the system integration between modules can also adopt a microservice or plug-in deployment architecture to adapt to enterprise knowledge bases and team collaboration platforms. However, if the core of AI-driven automatic causal reasoning is removed, the overall knowledge flow restructuring and review value will be greatly reduced.
[0088] To facilitate better implementation of the information processing method of this application, this application also provides an information processing apparatus based on the above-described information processing method. The meanings of the terms used are the same as in the above-described information processing method, and specific implementation details can be found in the descriptions of the method embodiments.
[0089] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of the information processing device provided in the embodiments of this application, wherein the information processing device includes: The information acquisition module 501 is used to acquire multiple related information of the target content; The semantic acquisition module 502 is used to perform long sequence modeling on multiple related information to obtain the semantics of multiple nodes; the semantics of the nodes carry context information. The semantic reasoning module 503 is used to reason about the semantics of multiple nodes to obtain multiple thought flows corresponding to the multiple nodes; The network construction module 504 is used to construct a knowledge network of the target content based on multiple thought processes; the nodes in the knowledge network are labeled with logical relationships.
[0090] In one embodiment, the semantic reasoning module 503 is specifically used to perform: The semantic context information of multiple nodes is inferred to obtain cross-entry semantic association features; the cross-entry semantic association features include at least one of the topic evolution trend, event development rhythm and signal flow path of multiple entries; Based on the cross-entry semantic association features, determine the semantic arrangement logic among the multiple nodes; The nodes are connected according to the semantic arrangement logic to obtain the thought process flow.
[0091] In one embodiment, the network construction module 504 is specifically configured to perform: Reasoning is performed on each of the aforementioned thought processes to determine the category and logical relationship of each node on the thought process; the category of the node includes at least one of inspiration, fact, material, conclusion, hypothesis, and evidence; the logical relationship includes at least one of causal relationship, temporal relationship, and contextual relationship; Based on the categories and logical relationships of the nodes in each of the aforementioned thought flows, edges are established between the nodes, and the logical relationships are labeled on each node to obtain the knowledge network.
[0092] In one embodiment, the device further includes: The new information acquisition module is used to acquire the semantics, content summary, and tags of the new related information when new related information is acquired. The new node generation module is used to generate new nodes based on the semantics, content summary, and tags of the new association information, and to determine the category and attributes of the new nodes; the attributes of the nodes include at least one of the following: content text, source, time, and modality type; The new edge establishment module is used to evaluate the logical relationship between the new node and other nodes based on the category and attributes of the new node, and to establish an edge between the new node and other nodes based on the logical relationship.
[0093] In one embodiment, the information acquisition module 501 is specifically used to perform: Obtain multiple original association information of the target content; the multiple original association information includes at least one modality among images, image-based text, audio, and text; The original association information is processed according to the corresponding modality to obtain machine-understandable association information.
[0094] In one embodiment, processing the multiple original association information pieces according to the corresponding modality processing method to obtain multiple association information pieces that can be understood by a machine includes one or more of the following steps: Image information is obtained by performing image recognition on the original association information of the image modality; Optical character recognition is performed on the original association information of the image-text modality to obtain the first text association information; Automatic speech recognition is performed on the original association information of the audio modality to obtain the second text association information.
[0095] In one embodiment, the semantic acquisition module 502 is specifically used to perform: Get the set time window; Based on the time window and the time information of each of the associated information, at least one of the associated information corresponding to each of the time windows is determined; The associated information corresponding to each time window is fused using multimodal methods to obtain the fused information corresponding to each time window. For each of the time windows, the fusion information of the time window and the semantics of the previous time window are fused and reasoned to obtain the semantics of the time window, and each of the time windows is determined as the node.
[0096] In one embodiment, the device further includes: The display module is used to visually represent the thought process, the knowledge network, and / or the logical relationships. An operation support module is used to support visualization operations on the knowledge network; the visualization operations include at least one of drag and drop, editing and merging.
[0097] In one embodiment, the device further includes: An interface providing module is used to provide a retrieval interface; the retrieval interface supports natural language question-and-answer interaction. The retrieval module is used to obtain retrieval information input based on the retrieval interface, and to perform multidimensional retrieval in the knowledge network based on the retrieval information to obtain retrieval results; the multidimensional retrieval includes at least one of keyword retrieval, semantic similarity retrieval, and logical relevance retrieval.
[0098] In one embodiment, the device further includes: The output generation module is used to automatically generate relevant outputs for the target content based on the knowledge network. The relevant outputs include at least one of dynamic web page content, scrolling subtitle animations, presentation documents, and short video scripts.
[0099] By employing the technical solution of this application embodiment, multiple related information of the target content can be obtained. Long-sequence modeling of these related information is performed according to a set time window to obtain the semantics of multiple nodes, thereby preserving the temporal correlation of the related information and enabling the generated node semantics to possess contextual attributes. Semantics carrying contextual information can improve the coherence and accuracy of subsequent reasoning. Reasoning on the semantics of multiple nodes yields multiple thought processes corresponding to those nodes, which can uncover potential multi-hop logical relationships between the semantics of the nodes. These multiple thought processes can present the multi-dimensional logical context of the target content. Based on these multiple thought processes, a knowledge network of the target content is constructed, and the nodes in the knowledge network are labeled with logical relationships. This transforms discrete thought processes into a structured knowledge network, intuitively presenting the logical relationships between related information. No manual operation is required; the logical integration and structured presentation of related information can be automatically completed, thereby improving the efficiency of retrieval and reuse of multiple related information of the target content.
[0100] For specific limitations regarding the information processing device, please refer to the limitations on the information processing method above, which will not be repeated here. Each module in the aforementioned information processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.
[0101] In addition, this application also provides an electronic device, such as Figure 6 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically: The electronic device may include components such as a processor 601 with one or more processing cores and a memory 602 with one or more computer-readable storage media. Those skilled in the art will understand that... Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 601 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, thereby providing overall monitoring of the electronic device. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 601.
[0102] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0103] In one embodiment, the electronic device further includes a power supply 603 that supplies power to the various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 603 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0104] In one embodiment, the electronic device may further include an input unit 604, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0105] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 602 according to the following instructions, and the processor 601 runs the applications stored in the memory 602, thereby implementing the steps in any of the information processing methods provided in the embodiments of this application.
[0106] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0107] In one embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods described in any embodiment of this application.
[0108] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of this application.
[0109] In some embodiments, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the methods described in any embodiment of this application.
[0110] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0111] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0112] Therefore, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the information processing methods provided in this application.
[0113] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0114] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0115] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the information processing methods provided in this application, the beneficial effects that any of the information processing methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0116] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0117] The above provides a detailed description of an information processing method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An information processing method characterized by comprising: include: Obtain multiple related information about the target content; Long-sequence modeling is performed on multiple pieces of related information to obtain the semantics of multiple nodes; The semantics of the nodes carry contextual information; The semantics of multiple nodes are inferred to obtain multiple thought processes corresponding to the multiple nodes; Based on multiple thought processes, a knowledge network for the target content is constructed; the nodes in the knowledge network are labeled with logical relationships.
2. The method according to claim 1, characterized in that, The semantic reasoning of the multiple nodes to obtain multiple thought processes corresponding to the multiple nodes includes: The semantic context information of multiple nodes is inferred to obtain cross-entry semantic association features; the cross-entry semantic association features include at least one of the topic evolution trend, event development rhythm and signal flow path of multiple entries; Based on the cross-entry semantic association features, determine the semantic arrangement logic among the multiple nodes; The nodes are connected according to the semantic arrangement logic to obtain the thought process flow.
3. The method according to claim 1, characterized in that, The construction of the knowledge network for the target content based on multiple thought processes includes: Reasoning is performed on each of the aforementioned thought processes to determine the category and logical relationship of each node in the thought process; the category of the node includes at least one of inspiration, fact, material, conclusion, hypothesis, and evidence; the logical relationship includes at least one of causal relationship, temporal relationship, and contextual relationship; Based on the categories and logical relationships of the nodes in each of the aforementioned thought flows, edges are established between the nodes, and the logical relationships are labeled on each node to obtain the knowledge network.
4. The method according to claim 1, characterized in that, The method further includes: When new related information is obtained, its semantics, content summary, and tags are acquired. Based on the semantics, content summary, and tags of the new associated information, a new node is generated, and the category and attributes of the new node are determined; the attributes of the node include at least one of the following: content text, source, time, and modality type. The logical relationship between the new node and other nodes is evaluated based on the category and attributes of the new node, and an edge is established between the new node and other nodes based on the logical relationship.
5. The method according to claim 1, characterized in that, The acquisition of multiple related information of the target content includes: Obtain multiple original association information of the target content; the multiple original association information includes at least one modality among images, image-based text, audio, and text; The original association information is processed according to the corresponding modality to obtain machine-understandable association information.
6. The method according to claim 5, characterized in that, The step of processing multiple original association information pieces according to the corresponding modality processing method to obtain multiple association information pieces that can be understood by a machine includes one or more of the following steps: Image information is obtained by performing image recognition on the original association information of the image modality; Optical character recognition is performed on the original association information of the image-text modality to obtain the first text association information; Automatic speech recognition is performed on the original association information of the audio modality to obtain the second text association information.
7. The method according to claim 1, characterized in that, The step of performing long-sequence modeling on multiple pieces of associated information to obtain the semantics of multiple nodes includes: Get the set time window; Based on the time window and the time information of each of the associated information, at least one of the associated information corresponding to each of the time windows is determined; The associated information corresponding to each time window is fused using multimodal methods to obtain the fused information corresponding to each time window. For each of the time windows, the fusion information of the time window and the semantics of the previous time window are fused and reasoned to obtain the semantics of the time window, and each of the time windows is determined as the node.
8. The method according to claim 1, characterized in that, The method further includes: Visualize the thought process, the knowledge network, and / or the logical relationships. The system supports visualization operations on the knowledge network, including at least one of drag-and-drop, editing, and merging.
9. The method according to claim 1, characterized in that, The method further includes: A search interface is provided; the search interface supports natural language question-and-answer interaction. Obtain the retrieval information input based on the retrieval interface, perform a multi-dimensional retrieval in the knowledge network based on the retrieval information, and obtain the retrieval results; The multidimensional search includes at least one of keyword search, semantic similarity search, and logical relevance search.
10. The method according to claim 1, characterized in that, The method further includes: Based on the knowledge network, the relevant output of the target content is automatically generated; The relevant outputs include at least one of dynamic web page content, scrolling subtitle animations, presentation documents, and short video scripts.
11. An information processing device, characterized in that, include: The information acquisition module is used to acquire multiple related information about the target content; The semantic acquisition module is used to perform long sequence modeling on multiple related information to obtain the semantics of multiple nodes; the semantics of the nodes carry contextual information. The semantic reasoning module is used to reason about the semantics of multiple nodes to obtain multiple thought processes corresponding to the multiple nodes; The network construction module is used to construct a knowledge network of the target content based on multiple thought processes; the nodes in the knowledge network are labeled with logical relationships.
12. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the information processing method as described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the information processing method as described in any one of claims 1 to 10.