Multi-point conference large screen interactive management system

The multi-point collaborative meeting large-screen interactive management system enables precise capture of user intent and structured parsing of content objects, solving the difficulties in managing and tracing interaction processes and decision-making processes in existing remote meeting systems, and improving meeting efficiency and the credibility of conclusions.

CN120994150BActive Publication Date: 2026-07-31NANJING AOZHUO HI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING AOZHUO HI TECH CO LTD
Filing Date
2025-06-20
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing remote conferencing systems cannot understand the inherent structure and semantics of screen content, resulting in inefficient interaction processes, discussion branches, and decision-making processes that cannot be systematically managed and traced in a structured manner.

Method used

Through modules for establishing identity-based interaction coordinates, parsing shared content objects, predicting user interaction intent, managing parallel topic conversations, capturing multimodal decision events, and generating decision process associations, the system achieves accurate capture of user intent and structured parsing of content objects, automatically establishing logical relationships between discussion branches and decision processes.

Benefits of technology

It enables precise capture of user intent and structured management of interactive events, improving meeting efficiency, ensuring orderly management of discussion branches and visual traceability of the decision-making process, and enhancing the usability and credibility of meeting conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994150B_ABST
    Figure CN120994150B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-point collaborative meeting large-screen interactive management system, relating to the field of network communication technology. The system includes: a shared content object parsing module, which parses shared content into an interactive content object tree; a user interaction intent prediction module, used to monitor user intent and generate initial decision nodes as the starting point for discussion; a parallel topic session management module, used to create parallel discussion sessions with automatic context inheritance; a decision process association generation module, which constructs a decision graph to display the logical relationships between topics; and a meeting state machine control module, used to manage the focused, divergent, and convergent meeting process states. This invention significantly improves the collaborative efficiency and decision-making quality of complex topics by providing comprehensive, structured, traceable, and visualized management of the core interactions, discussion branches, and decision-making processes of the meeting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network communication, and in particular to a multi-point collaborative conferencing large-screen interactive management system. Background Technology

[0002] Existing remote conferencing systems, such as various video conferencing software and online collaboration platforms, have played a significant role in improving cross-regional communication efficiency. However, their core technologies largely rely on screen pixel stream sharing. In this mode, all participants see a uniform image, and the system itself cannot understand the inherent structure and semantics of the content displayed on the screen. User interactions, such as mouse pointer gestures, are unstructured, transient information that cannot be automatically captured and utilized by the system.

[0003] When multiple topics need to be discussed in parallel during a meeting, traditional systems typically use a "group discussion" function, which divides participants into different virtual rooms. However, this division is coarse-grained, lacks automatic context inheritance, and requires members to verbally reconfirm the topics after grouping, leading to inefficiency. Furthermore, the main meeting room cannot intuitively and systematically monitor the discussion status and progress of each group.

[0004] More importantly, key viewpoints, provisional conclusions, and decision-making processes generated during meetings are often scattered across chat logs, oral recordings, or personal notes, lacking a unified, structured mechanism for preservation. Post-meeting minutes compilation is time-consuming and labor-intensive, and due to the fragmented information, it's difficult to reconstruct a complete chain of decision-making logic, leading to the loss of important information and difficulties in subsequent accountability. Therefore, existing technologies exhibit significant efficiency, management, and traceability bottlenecks when handling complex, multi-threaded collaborative tasks. Summary of the Invention

[0005] This invention provides a multi-point collaborative meeting large-screen interactive management system, which aims to solve the technical problem that existing meeting interactions rely on unstructured audio and video streams, resulting in the inability to systematically manage and structurally trace the interaction process, discussion branches, and decision-making process.

[0006] In view of the above problems, the present invention provides a multi-point collaborative meeting large-screen interactive management system, including:

[0007] The identity-based interactive coordinate establishment module establishes a mapping relationship between each client terminal accessing the system and the main coordinate system of the display screen, and renders an identity-based cursor on the display screen that can uniquely identify the user's identity based on the user identity information of each client terminal.

[0008] The shared content object parsing module receives content shared by any of the client terminals and parses the shared content into a content object tree composed of content objects carrying bounding box coordinate information;

[0009] The user interaction intent prediction module monitors the position of the identity cursor in real time. When any of the identity cursors and the bounding box of a certain content object meet the preset association rules, it determines that the user has an interaction intent with the content object, generates an intent trigger event containing user identity information and the identity information of the content object, and generates an initial decision node as the starting point of the discussion based on the intent trigger event.

[0010] The parallel topic conversation management module responds to the instruction generated based on the intent triggering event, instantiates a branch conversation object for parallel discussion, and automatically copies the content object pointed to by the intent triggering event as the initial discussion material for the branch conversation object;

[0011] The multimodal decision event capture module captures specified information during the meeting based on user operations and encapsulates the specified information into standardized decision node objects;

[0012] The decision process association generation module organizes all generated decision node objects into a decision graph data structure. Based on the branch session object ID and the triggered content object identity information, it automatically establishes parent-child relationships between decision nodes and generates structured meeting minutes based on the decision graph.

[0013] Preferably, the system further includes: a conference state machine control module, communicatively connected to the parallel topic session management module and the decision process association generation module, used for:

[0014] Responding to the creation event of the branch session object, the main meeting interface is switched from focused state to divergent state;

[0015] Responding to the event that marks the final conclusion within the branch session object, the main conference interface is switched from the divergent state to the convergent state.

[0016] The technical solution provided in this application has at least the following technical effects or advantages:

[0017] The shared content object parsing module transforms unstructured screen-shared content into a content object tree that can be understood and interacted with by the system. Combined with the user interaction intent prediction module, it achieves accurate capture of user intent and instantly solidifies it by generating initial decision nodes. This elevates simple instructions into structured, traceable interactive event starting points, improving the depth and effectiveness of the interaction.

[0018] Through the parallel topic conversation management module and the meeting state machine control module, the system achieves automatic inheritance of context information and orderly management of parallel discussions. The system can intuitively present the creation and association of discussion branches, and guide the meeting to smoothly switch between focused, divergent, and convergent states through the state machine, thus solving the problems of chaotic process and low efficiency in traditional group discussions.

[0019] By introducing initial decision nodes and constructing decision graphs, the entire process of a meeting, from topic initiation and discussion to final conclusion, is structured and visualized. The decision process association generation module can automatically establish logical relationships between events based on chronology and semantics, making meeting outcomes no longer scattered information points, but rather complete, traceable, and logically clear decision assets, greatly improving the usability and credibility of meeting conclusions. Attached Figure Description

[0020] Figure 1 This is an architecture diagram of the multi-point collaborative meeting large-screen interactive management system of the present invention. Detailed Implementation

[0021] This invention relates to a multi-point collaborative meeting large-screen interactive management system, aiming to solve the technical problem that existing meeting interactions rely on unstructured audio and video streams, resulting in the inability to systematically manage and structurally trace the interaction process, discussion branches, and decision-making process.

[0022] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0023] See Figure 1 The following is an architecture diagram of a multi-point collaborative meeting large-screen interactive management system. This embodiment of the invention provides a multi-point collaborative meeting large-screen interactive management system deployed on a central processing server. The system includes:

[0024] The identity-based interactive coordinate establishment module establishes a mapping relationship between each client terminal connected to the system and the main coordinate system of the display screen, and renders an identity-based cursor on the display screen that can uniquely identify the user's identity based on the user identity information of each client terminal.

[0025] Specifically, during central processing server initialization, the visible area of ​​the display screen is defined as a two-dimensional Cartesian principal coordinate system M with a resolution of WxH. When a client terminal (ClientID: C1) connects, the identity-based interactive coordinate establishment module performs the following steps:

[0026] 1. Receive the resolution wxh of the shared window coordinate system C of C1.

[0027] 2. Calculate and store a 3x3 affine transformation matrix T1 for the transformation from coordinate system C to principal coordinate system M.

[0028] 3. Receive data packets from C1 through a persistent TCP connection. The data packet format is {clientID:C1,op:'move',pos:[x,y]}.

[0029] 4. On the server side, find the transformation matrix T1 associated with C1, perform an affine transformation operation on the coordinate point [x,y], and obtain the coordinates [x',y'] in the principal coordinate system.

[0030] 5. Generate a rendering instruction {render_type:'cursor',user_id:U1,pos:[x',y']} and push the instruction into the rendering engine's command queue. The rendering engine will then draw the identity cursor associated with user U1 at the [x',y'] position on the display screen.

[0031] The shared content object parsing module receives content shared by any of the client terminals and parses the shared content into a content object tree composed of content objects carrying bounding box coordinate information.

[0032] Specifically, when a client terminal shares a PowerPoint file, it sends the file handle or file stream to the server, not screen pixels. The shared content object parsing module is then triggered, executing the following steps:

[0033] 1. Determine the file type and call a document parser based on the Apache POI or Aspose.Slides library.

[0034] 2. The parser traverses the DOM (Document Object Model) of the PPT file, and for each individual Shape (such as a text box or chart), it extracts its type, Z-order, position, and size information.

[0035] 3. Instantiate a ContentObject data structure in memory for each Shape, which contains the following fields: {objectID:UUID,type:'chart',boundingBox:[x,y,w,h],rawData:[data_pointer]}.

[0036] 4. All pointers to instantiated ContentObjects are stored in a graph data structure to form a content object tree. The structure of this tree reflects the hierarchical relationship of the file content and is cached for high-speed querying by other modules.

[0037] The user interaction intent prediction module monitors the position of the identity cursor in real time. When any of the identity cursors and the bounding box of a content object meet the preset association rules, it determines that the user has an interaction intent with the content object, generates an intent trigger event containing user identity information and the identity information of the content object, and generates an initial decision node as the starting point of the discussion based on the intent trigger event.

[0038] Specifically, the user interaction intent prediction module runs a separate polling thread that performs the following operations at a frequency of 60Hz:

[0039] 1. Get the cursor positions [x', y'] of all currently active users.

[0040] 2. Traverse each ContentObject in the content object tree and perform bounding box collision detection.

[0041] 3. If cursor C1 is detected to be within the bounding box of object O1, then in a hash table with {C1_O1} as the key, increment or update the timestamp of a hover timer field.

[0042] 4. If the cumulative hover time recorded by the timer exceeds a preset time threshold (e.g., 1500 milliseconds), the user interaction intent prediction module determines that an intent has been generated.

[0043] 5. Immediately invoke the constructor of DecisionNode to create an initial decision node with its contentType field set to 'TOPIC' and its contentData field storing the objectID of O1. This node is written to the DecisionGraph table in the database, and a unique nodeID is returned.

[0044] 6. Finally, generate and broadcast an intent-triggered event, with the event message body being {userID:

[0045] U1,objectID:O1,initialNodeID:returned_nodeID}.

[0046] The parallel topic conversation management module, in response to the instruction generated based on the intent-triggered event, instantiates a branch conversation object for parallel discussion, and automatically copies the content object pointed to by the intent-triggered event as the initial discussion material for the branch conversation object.

[0047] Specifically, an event listener subscribes to the intent-triggered event. When the user clicks "Create Branch," the parallel topic session management module is activated and the CreateBranchSession function is executed:

[0048] 1. Generate a globally unique branchID.

[0049] 2. Insert a new record into the BranchSessions table in the database, containing the branchID and the parentObjectID obtained from the intent-triggered event.

[0050] 3. Send an API request to the media server (such as Mediasoup) to create a separate media room and obtain a unique roomID.

[0051] 4. Instantiate a new whiteboard data model in memory, retrieve the corresponding ContentObject from the content object tree using parentObjectID, perform a deep copy on it, and inject the copied data into the new whiteboard data model.

[0052] 5. Push the branchID and roomID to the client that initiated the creation via WebSocket.

[0053] The multimodal decision event capture module captures specified information during the meeting process based on user operations and encapsulates the specified information into standardized decision node objects.

[0054] Specifically, when a user selects a whiteboard area on the client interface and clicks "Pin," the client sends a POST / api / capture request to the server. The request body is as follows:

[0055] {type:'snapshot',sessionID:current_branchID,rect:[x1,y1,x2,y2]}.

[0056] After the multimodal decision event capture module receives the request:

[0057] 1. Locate the corresponding session framebuffer based on the sessionID.

[0058] 2. Instructs the rendering engine to extract the rectangular region specified by rect from the framebuffer and encode it as a PNG format byte stream.

[0059] 3. Upload the byte stream to object storage (such as S3) to obtain a URL.

[0060] 4. Call the DecisionNode constructor to create a new decision node object with the contentType field set to 'IMAGE_SNAPSHOT' and the contentData field set to the URL.

[0061] 5. Store the object in the database.

[0062] The decision process association generation module organizes all generated decision node objects into a decision graph data structure. Based on the branch session object ID and the identity information of the triggered content object, it automatically establishes parent-child relationships between decision nodes and generates structured meeting minutes based on the decision graph.

[0063] Specifically, the underlying data structure of the decision process association generation module is an adjacency table in a graph database (such as Neo4j) or a relational database that supports directed edges. When the multimodal decision event capture module generates a new decision node N_new, the decision process association generation module is triggered. It extracts the creatorID and sessionID from the metadata of N_new, then performs a database query to find the previous node N_prev, created by the same creatorID within the same sessionID, with the most recent timestamp. Once found, the decision process association generation module creates an edge in the graph pointing from N_prev to N_new, with the edge type marked as 'SEQUENTIAL'. The establishment of other associations is detailed below.

[0064] Furthermore, when establishing the mapping relationship, the identity-based interactive coordinate establishment module is specifically used to: obtain the coordinate system of the shared window on the client terminal, calculate the transformation matrix from the coordinate system to the main coordinate system, and convert the operation coordinates on the client terminal into the corresponding coordinates in the main coordinate system in real time through the transformation matrix.

[0065] Specifically, the transformation matrix is ​​a 3x3 affine transformation matrix capable of handling translation, scaling, and rotation simultaneously (although rotation is less commonly used in screen-sharing scenarios). The matrix calculation is performed once when the client first connects or when its shared window size changes. The calculation result is cached in a hash table on the server using the clientID as the key to support efficient real-time coordinate transformations, avoiding recalculation with each coordinate data packet arrival, thus ensuring low latency for multi-user cursor movement.

[0066] Furthermore, the identity cursor has exclusive visual attributes bound to the user's identity information, including the user's exclusive color, shape, or attached name tag.

[0067] Specifically, in the user database table, each user ID has a `cursor_style` field, which is a JSON object, for example, `{color:'#FF0000',shape:'pointer_v2',show_name:true}`. When the identity-based interactive coordinate creation module generates cursor rendering instructions, it pushes this JSON object to the rendering engine. Based on these attributes, the rendering engine selects the corresponding SVG or bitmap texture from the preset resource library for drawing, thereby achieving personalization of the user's cursor.

[0068] Furthermore, the shared content object parsing module is specifically used for:

[0069] When the shared content is a document type, the document object model parser is invoked to parse the text boxes, images, charts, or tables in the document into the content object.

[0070] When the shared content is a bitmap image that cannot be structured and parsed, a computer vision-based parser is invoked to generate the content object for the identified text blocks or visual elements through optical character recognition technology or object detection algorithms.

[0071] Specifically, the shared content object parsing module internally implements a StrategyPattern. It maintains a mapping from file extensions to specific parser implementations. For example, .pptx maps to the PptxParser class, and .pdf maps to the PdfParser class. When the extension cannot be determined or the content is a raw bitmap stream, the default CvParser class is called. The CvParser class sends the image data to a separate Python microservice running pre-trained models (such as YOLOv5 for object detection and Tesseract for OCR) via an internal RPC call, asynchronously receiving the returned JSON data containing the recognition results and coordinates, and then converting it into a ContentObject structure.

[0072] Furthermore, the preset association rules include:

[0073] The time that the coordinates of the identified cursor remain within the bounding box of a certain content object exceeds a preset hovering time threshold.

[0074] The identified cursor completes a predefined gesture trajectory within the bounding box of a certain content object.

[0075] Specifically, for gesture trajectory determination, the user interaction intent prediction module stores a series of consecutive coordinate points [x', y'] within a short period (e.g., within 500ms) in a temporary buffer. When the cursor moves out of the object's bounding box or the mouse button is released, the sequence of coordinate points in this buffer is passed to a gesture recognizer. This recognizer can use a simple template matching algorithm to compare the input trajectory with preset gesture templates such as "drawing a circle" or "double-clicking," calculating the distance or similarity between the trajectories. If the similarity exceeds a preset threshold (e.g., 0.85), the corresponding gesture type is returned, thereby triggering intent determination.

[0076] Furthermore, the parallel topic conversation management module is also used to: dynamically draw a connecting line from the edge of the bounding box of the content object that triggers the branch conversation object on the main interface of the display screen, and generate an interactive node icon representing the branch conversation object at the end of the connecting line.

[0077] Specifically, when the parallel topic session management module creates a branch session, it pushes an instruction {render_type:'branch_link',start_obj:O1,end_node_id:B1} to the rendering engine. Upon receiving this instruction, the rendering engine obtains the coordinates P_start of the bounding box center point of the content object O1 and calculates a dynamic docking position P_end for the node icon representing branch session B1. Subsequently, the rendering engine uses a Bézier curve algorithm to calculate and draw a smooth connecting line with P_start and P_end as endpoints.

[0078] Furthermore, the branch session object is a data structure that includes: a branch session ID for uniquely identifying the parallel discussion, a parent object ID pointing to the content object that triggered the parallel discussion, a dynamic member list recording the current participants, and an independent communication channel isolated from the main venue's audio, video, and interaction.

[0079] Specifically, the branch session object is an instantiated class on the server side. Its dynamic member list is an array or linked list storing userIDs. The independent communication channel field stores a unique roomID string obtained from the media server. When a new user requests to join the branch, the system adds their userID to the member list and uses the roomID to obtain signaling from the media server, enabling them to subscribe to and publish audio and video streams for that channel, thereby achieving communication isolation at the network layer.

[0080] Furthermore, when capturing the specified information, the multimodal decision event capture module specifically performs at least one of the following operations:

[0081] Copy the text content of the chat message;

[0082] Allows users to select a rectangular area on the whiteboard and take a screenshot of that area;

[0083] Call the speech-to-text engine to convert a specified audio dialogue into text.

[0084] Specifically, for speech-to-text conversion, when a user presses the "Record Voice" button, the client's WebRTC stack sends the user's audio stream to the server via a separate MediaStreamTrack. A speech processing service on the server (which can integrate a third-party ASRSDK) receives the audio stream, performs real-time speech-to-text processing, and returns the converted text segment to the client via WebSocket for display. After user confirmation, the multimodal decision event capture module encapsulates it into a decision node object with contentType 'ASR_TEXT'.

[0085] Furthermore, the decision process association generation module is specifically used to:

[0086] Within the same branch session object, at least one of the following operations is performed according to preset association rules:

[0087] Establish a temporal successor association between a newly created decision node object and its direct upstream decision node object;

[0088] A new decision node, used to supplement or refute an existing decision node, is semantically branched with the existing decision node.

[0089] After a branch session ends, the decision node object marked as the final conclusion within that session is associated with the initial decision node generated by the intent-triggered event that triggered that branch session.

[0090] Specifically, when establishing a connection, the decision process association generation module analyzes the creation context of the new decision node N_new:

[0091] 1. Temporal association: By default, the decision process association generation module finds the node N_prev created by N_new.creatorID in N_new.sessionID with the closest timestamp that is less than N_new.timestamp, and establishes an edge of type 'SEQUENTIAL' from N_prev->N_new.

[0092] 2. Semantic Association: If the client includes the reply_to_nodeID field in the request to create N_new, the decision process association generation module will not establish a temporal association, but will instead establish an edge of type 'SEMANTIC_BRANCH' from reply_to_nodeID to N_new.

[0093] 3. Parent-child association: When a branch session ends and the user marks a node N_conclusion as the conclusion, the decision process association generation module will look up the initial decision node N_initial of the session (the ID of this node is stored in its metadata when the branch session is created) and establish an edge of type 'CONCLUSION' from N_initial to N_conclusion to complete the logical loop.

[0094] Furthermore, the system also includes: a conference state machine control module, communicatively connected to the parallel topic session management module and the decision process association generation module, used for:

[0095] Responding to the creation event of the branch session object, the main meeting interface is switched from focused state to divergent state;

[0096] Responding to the event that marks the final conclusion within the branch session object, the main conference interface is switched from the divergent state to the convergent state.

[0097] Specifically, the meeting state machine control module is implemented using a finite state machine (FSM) model. It maintains a `meeting_state` variable in memory, initialized to 'FOCUSED'. This module subscribes to events published by other modules via an internal event bus.

[0098] 1. When the parallel topic session management module successfully creates a branch session, it will publish a BranchCreated event. After the event handler of the meeting state machine control module receives the event, it will set the meeting_state variable to 'DIVERGENT' and broadcast a WebSocket message {event:'state_change',state:'DIVERGENT'} to all clients.

[0099] 2. When a user marks a decision node as a conclusion on the client side, the client sends a request. After the database updates the node's state, a ConclusionMarked event is published. Upon receiving this, the meeting state machine control module sets the meeting_state to 'CONVERGENT' and broadcasts the corresponding state change message. The client UI dynamically adjusts its layout based on the received state; for example, displaying all branch nodes in the 'DIVERGENT' state and highlighting the conclusion to be merged in the 'CONVERGENT' state.

[0100] In summary, the multi-point collaborative meeting large-screen interactive management system provided by this invention, through the collaborative work of the above modules, realizes the process from predicting the intent of user actions to object-oriented parsing of shared content, to structured management of parallel discussions, and finally to the solidification and traceability of the decision-making process, thus solving the pain points of low meeting efficiency, chaotic processes, and difficulty in consolidating conclusions in the prior art.

[0101] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A multi-point collaborative meeting large-screen interactive management system, characterized in that: include: The identity-based interactive coordinate establishment module establishes a mapping relationship between each client terminal accessing the system and the main coordinate system of the display screen, and renders an identity-based cursor on the display screen that can uniquely identify the user's identity based on the user identity information of each client terminal. The shared content object parsing module receives content shared by any of the client terminals and parses the shared content into a content object tree composed of content objects carrying bounding box coordinate information; The user interaction intent prediction module monitors the position of the identity cursor in real time. When any of the identity cursors and the bounding box of a certain content object meet the preset association rules, it determines that the user has an interaction intent with the content object, generates an intent trigger event containing user identity information and the identity information of the content object, and generates an initial decision node as the starting point of the discussion based on the intent trigger event. The parallel topic conversation management module responds to the instruction generated based on the intent triggering event, instantiates a branch conversation object for parallel discussion, and automatically copies the content object pointed to by the intent triggering event as the initial discussion material for the branch conversation object; The multimodal decision event capture module captures specified information during the meeting based on user operations and encapsulates the specified information into standardized decision node objects; The decision process association generation module organizes all generated decision node objects into a decision graph data structure. Based on the branch session object ID and the triggered content object identity information, it automatically establishes parent-child relationships between decision nodes and generates structured meeting minutes based on the decision graph.

2. The multi-point collaborative conference large-screen interactive management system as described in claim 1, characterized in that, When establishing the mapping relationship, the identity-based interactive coordinate establishment module is specifically used to: obtain the coordinate system of the shared window on the client terminal, calculate the transformation matrix from the coordinate system to the main coordinate system, and convert the operation coordinates on the client terminal into the corresponding coordinates in the main coordinate system in real time through the transformation matrix.

3. The multi-point collaborative conference large-screen interactive management system as described in claim 1, characterized in that, The identity cursor has exclusive visual attributes that are bound to the user's identity information. These exclusive visual attributes include the user's unique color, shape, or attached name tag.

4. The multi-point collaborative conference large-screen interactive management system as described in claim 1, characterized in that, The shared content object parsing module is specifically used for: When the shared content is a document type, the document object model parser is invoked to parse the text boxes, images, charts, or tables in the document into the content object. When the shared content is a bitmap image that cannot be structured and parsed, a computer vision-based parser is invoked to generate the content object for the identified text blocks or visual elements through optical character recognition technology or object detection algorithms.

5. The multi-point collaborative conference large-screen interactive management system as described in claim 1, characterized in that, The preset association rules include: The time that the coordinates of the identified cursor remain within the bounding box of a certain content object exceeds a preset hovering time threshold. The identified cursor completes a predefined gesture trajectory within the bounding box of a certain content object.

6. The multi-point collaborative conference large-screen interactive management system as described in claim 5, characterized in that, The parallel topic conversation management module is also used to: dynamically draw a connecting line from the edge of the bounding box of the content object that triggers the branch conversation object on the main interface of the display screen, and generate an interactive node icon representing the branch conversation object at the end of the connecting line.

7. The multi-point collaborative conference large-screen interactive management system as described in claim 1, characterized in that, The branch session object is a data structure that includes: a branch session ID for uniquely identifying the parallel discussion, a parent object ID pointing to the content object that triggered the parallel discussion, a dynamic member list recording the current participants, and an independent communication channel isolated from the main venue's audio, video, and interaction.

8. The multi-point collaborative conference large-screen interactive management system as described in claim 1, characterized in that, When capturing the specified information, the multimodal decision event capture module is specifically used to perform at least one of the following operations: Copy the text content of the chat message; Allows users to select a rectangular area on the whiteboard and take a screenshot of that area; Call the speech-to-text engine to convert a specified audio dialogue into text.

9. The multi-point collaborative conference large-screen interactive management system as described in claim 1, characterized in that, The decision-making process association generation module is specifically used to: When establishing the association, the module is used to: Within the same branch session object, according to preset association rules, at least one of the following operations is performed: establishing a temporal successor association between a newly created decision node object and its directly upstream decision node object; establishing a semantic branch association between a new decision node used to supplement or refute an existing decision node and the existing decision node. After a branch session ends, the decision node object marked as the final conclusion within that session is associated with the initial decision node generated by the intent-triggered event that triggered that branch session.

10. The multi-point collaborative conferencing large-screen interactive management system as described in claim 1, characterized in that, The system further includes a conference state machine control module, which is communicatively connected to the parallel topic session management module and the decision process association generation module, and is used for: Responding to the creation event of the branch session object, the main meeting interface is switched from focused state to divergent state; Responding to the event that marks the final conclusion within the branch session object, the main conference interface is switched from the divergent state to the convergent state.