File multi-mode cooperation system and method, electronic equipment and storage medium
By playing voice information and operation information in the timeline, the problem of poor voice interference and synchronization in the existing document review system is solved, and efficient document review and annotation is achieved in multiple people, reducing the cost of real-time communication.
Patent Information
- Application Number
- CN202510453355.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-01
AI Technical Summary
It is difficult for the existing document review system to achieve the deep integration of voice discussion and dynamic document operations, and there are problems such as voice interference, poor audio and video synchronization, and limited resolution.
Through the file multi-modal collaboration system, voice information and operation information are fused into the file through the time axis to generate playback files, and play them together with time information, including voice information acquisition, recording, operation information acquisition and recording, and playback file generation module, supporting multi-person synchronous understanding and annotation of documents.
It improves the efficiency of multi-modal collaboration and review of files, realizes efficient synchronization of collaborative understanding and annotation of documents by multiple people, reduces the dependence and cost of real-time communication, and ensures high-definition voice and document operation synchronization.
Smart Images

Figure CN120407824A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electronic information technology, and relates to an information processing system and method, in particular to a file multi-modal collaboration system, method, electronic device and storage medium. Background Art
[0002] Existing document review systems are difficult to achieve a deep integration of voice discussion and dynamic document operations. Traditional solutions rely on static document sharing or low-resolution real-time recording, and there are problems such as voice interference, poor audio-visual synchronization, and limited resolution.
[0003] In view of this, there is an urgent need to design a new document review method today to overcome at least some of the above defects existing in the existing document review methods. Summary of the Invention
[0004] The present invention provides a file multi-modal collaboration system, method, electronic device and storage medium, which can fuse voice and operation information for a file into the file through a time axis, improve the efficiency of file multi-modal collaboration and review, and facilitate a comprehensive understanding of the processing and review process.
[0005] To solve the above technical problems, according to one aspect of the present invention, the following technical solution is adopted:
[0006] A file multi-modal collaboration system, the file multi-modal collaboration system includes:
[0007] A voice information acquisition module for acquiring voice information output by a set terminal device;
[0008] A voice information recording module for recording the voice information acquired by the voice information acquisition module and recording the time information corresponding to the voice information;
[0009] An operation information acquisition module for acquiring operation information for a file output by a set terminal device or / and operation information that causes the content presented on a set interface to change;
[0010] An operation information recording module for recording the operation information acquired by the operation information acquisition module and recording the time information corresponding to the operation information;
[0011] A playback file generation module for generating a playback file from the voice information recorded by the voice information recording module and the operation information recorded by the operation information recording module, and when the playback file is played, it can control the coordinated playback of the corresponding voice information and operation information based on the time information.
[0012] As an implementation of the present invention, the file multimodal collaboration system further includes a playback control module, which is used to control the collaborative playback of the voice information and operation information on the time axis.
[0013] As an implementation of the present invention, the file multimodal collaboration system further includes a time axis generation module, which is used to generate a time axis and mark the voice information and / or operation information output by the set terminal device within a set time period on the time axis.
[0014] As an implementation of the present invention, the operation information obtained by the operation information acquisition module includes at least one of the following information:
[0015] The position information of the file displayed on the set interface;
[0016] The specific operation object for file operation;
[0017] The operation position coordinates for file operation;
[0018] The operation position coordinates for operations performed in other interface areas outside the file display area without operating on the file;
[0019] The file multimodal collaboration system includes a server and at least one terminal device, and the server is respectively connected to each terminal device; the server includes the voice information recording module, the operation information recording module, and the playback file generation module.
[0020] According to another aspect of the present invention, the following technical solution is adopted: A file multimodal collaboration method, the file multimodal collaboration method includes:
[0021] Voice information acquisition step: Acquire the voice information output by the set terminal device;
[0022] Voice information recording step: Record the voice information obtained in the voice information acquisition step and record the time information corresponding to the voice information;
[0023] An operation information acquisition module, which is used to acquire the operation information for files output by the set terminal device and / or the operation information that causes the content presented on the set interface to change;
[0024] Operation information recording step: Record the operation information obtained in the operation information acquisition step and record the time information corresponding to the operation information;
[0025] Playback file generation step: Generate a playback file from the voice information recorded in the voice information recording step and the operation information recorded in the operation information recording step. When the playback file is played, it can display the voice information and operation information corresponding to different time periods on the timeline.
[0026] As an embodiment of the present invention, the file multimodal collaboration method further includes a playback control step: controlling the collaborative playback of the voice information and operation information on the timeline.
[0027] As an embodiment of the present invention, the file multimodal collaboration method further includes a timeline generation step: generating a timeline and marking the voice information and / or operation information output by a set terminal device within a set time period on the timeline.
[0028] As an embodiment of the present invention, the operation information obtained in the operation information acquisition step includes at least one of the following information:
[0029] The position information of the file displayed on the set interface;
[0030] The specific operation object for file operation;
[0031] The operation position coordinates for file operation;
[0032] The operation position coordinates for operations performed in other interface areas outside the file display area without file operation.
[0033] According to another aspect of the present invention, the following technical solution is adopted: An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0034] According to another aspect of the present invention, the following technical solution is adopted: A storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the steps of the above method are implemented.
[0035] The beneficial effects of the present invention are as follows: The file multimodal collaboration system, method, electronic device, and storage medium proposed by the present invention can fuse voice and operation information for files through a timeline, improve the efficiency of file multimodal collaboration and review, and facilitate a comprehensive understanding of the processing and review process. Brief Description of the Drawings
[0036] Figure 1 It is a schematic diagram of the composition of the file multimodal collaboration system in an embodiment of the present invention.
[0037] Figure 2 It is a flowchart of the file multimodal collaboration method in an embodiment of the present invention.
[0038] Figure 3 This is a schematic diagram of the composition of an electronic device in an embodiment of the present invention. Detailed implementation manners
[0039] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0040] To further understand the present invention, the preferred implementation manners of the present invention will be described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.
[0041] The description of this part only focuses on several typical embodiments, and the present invention is not limited to the scope described in the embodiments. The mutual replacement of the same or similar prior art means and some technical features in the embodiments is also within the scope of description and protection of the present invention.
[0042] The expression of the steps in each embodiment in the specification is only for convenience of description, and the implementation manner of the present application is not limited by the order of step implementation.
[0043] "Connection" in the specification includes both direct connection and indirect connection.
[0044] The present invention discloses a system that makes the understanding and communication of documents in any format more efficient. Figure 1 This is a schematic diagram of the composition of a file multimodal collaboration system in an embodiment of the present invention; please refer to Figure 1 , the file collaboration system includes: a voice information acquisition module 1, a voice information recording module 2, an operation information acquisition module 3, an operation information recording module 4, and a playback file generation module 5.
[0045] The voice information acquisition module 1 is used to acquire the voice information output by a set terminal device. That is, when processing a document, the voice information can be heated at the same time, so that the user who browses the document later can understand why the above processing is performed on the document, or what other modifications are still needed.
[0046] The voice information recording module 2 is used to record the voice information acquired by the voice information acquisition module, and record the time information corresponding to the voice information. The voice information recording module 2 can also match the above voice information with a time axis based on the time information; by dragging the time axis, the voice information in the corresponding area can be played.
[0047] The operation information acquisition module 3 is used to acquire the operation information for the file output by a set terminal device or / and the operation information that causes the content presented on a set interface to change.
[0048] In an embodiment of the present invention, the operation information obtained by the operation information acquisition module 3 includes at least one of the following information: the position information of the file displayed on the set interface; the specific operation object for the file operation; the operation position coordinates for the file operation; the operation position coordinates for operations performed in other interface areas outside the file display area without operating on the file. The files for collaborative operation can be documents, or can also be pictures, videos, etc. The above operations may include swiping up the document, swiping down the document, zooming in on the document, zooming out on the document, revising the document, annotating the document, modifying the picture; it is also possible to make remarks or draw pictures in areas outside the file, as well as other operations that can be presented in the display area.
[0049] In an embodiment, it is possible to obtain the area where the file is displayed, the operations performed by the user in the display area, and record them; so that the above data can be placed on the timeline and played later with the timeline as the center.
[0050] Through the operation information acquisition module 3, each user can see in real time that the document operation results are synchronized to others, enabling multiple people in the same physical room to synchronously understand and annotate any document, and also enabling remote "participants" to synchronously read and annotate the document, making it convenient to understand and review the document.
[0051] The operation information recording module 4 is used to record the operation information obtained by the operation information acquisition module, and record the time information corresponding to the operation information. The operation information recording module 4 can also match the above operation information with the timeline based on the time information; by dragging the timeline, the operation information in the corresponding area can be played.
[0052] The playback file generation module 5 is used to generate a playback file from the voice information recorded by the voice information recording module and the operation information recorded by the operation information recording module. When the playback file is played, it can control the coordinated playback of the corresponding voice information and operation information based on the time information.
[0053] In an embodiment of the present invention, the file multimodal collaboration system further includes a playback control module 6, and the playback control module 6 is used to control the coordinated playback of the voice information and operation information on the timeline.
[0054] The file multimodal collaboration system may also include a timeline generation module 7, and the timeline generation module 7 is used to generate a timeline and mark the voice information and / or operation information output by the set terminal device within the set time period on the timeline.
[0055] The described document multimodal collaboration system includes a server and at least one terminal device, and the server is respectively connected to each terminal device; the server includes the voice information recording module 2, the operation information recording module 4, and the playback file generation module 5. The terminal device may include the above-mentioned voice information acquisition module 1 and operation information acquisition module 3.
[0056] The present invention overcomes the problem that traditional document reading comprehension cannot make the collaboration of multiple people as efficient as watching a movie simultaneously. Its core is achieved through the collaborative reading and synchronous expression of documents on anyone's free device, while enabling the synchronous sound of multiple people to depend on a third party or utilize the fact that a group of people together can hear each other to achieve the synchronization of "watching" and "listening". When we can hear each other and simultaneously watch, express, and annotate the document together, it realizes turning any static document into a real-time generated animation / video stream for communication.
[0057] By independently and synchronously recording the sound in real-time communication on anyone's device, it is possible to achieve sound recording without relying on the server recording and the costs required for sound synchronous recording, and also achieve higher clarity of sound recording.
[0058] Importantly, this kind of sound recording and playback no longer requires the costs associated with initiating and recording a meeting and the dependence on an audio / video server; as a result, each "participant" records the sound locally, with a time axis established based on synchronization and the same standard; enabling the synchronous mixing of sounds to be generated after the completion of the document writing discussion, making the dynamic communication and collaboration between sound and document content infinitely high-definition without the corresponding costs of real-time communication.
[0059] The present invention further discloses a document multimodal collaboration method, Figure 2 which is a flowchart of the document multimodal collaboration method in an embodiment of the present invention; please refer to Figure 2 , and the document multimodal collaboration method includes:
[0060]
Step S1
[0061] In one embodiment, when processing a document, the voice information can be heated simultaneously to facilitate subsequent users browsing the document to understand why the above-mentioned processing is performed on the document or what further modifications are required.
[0062]
Step S2
[0063] In one embodiment, the above voice information can be coordinated with a time axis based on time information by the voice information recording module 2; by dragging the time axis, the voice information in the corresponding area can be played.
[0064]
Step S3
[0065] The files for collaborative operation can be documents, or pictures, videos, etc. The above operations can include swiping up the document, swiping down the document, zooming in on the document, zooming out on the document, revising the document, annotating the document, modifying the picture; it is also possible to make remarks or draw pictures in areas outside the file, as well as other operations that can be presented in the display area.
[0066] In one embodiment, the area where the file is displayed, the operations performed by the user in the display area, can be acquired and recorded; so that the above data can be placed on the time axis and played centered on the time axis later.
[0067] Through the operation information acquisition module 3, each user can see in real time that the operation results of the document are synchronized to others, enabling multiple people in the same physical room to synchronously understand and annotate any document, and also enabling remote "participants" to synchronously read and annotate the document, making it convenient to understand and review the document.
[0068]
Step S4
[0069] In one embodiment, the above operation information can be coordinated with the time axis based on time information by the operation information recording module 4; by dragging the time axis, the operation information in the corresponding area can be played.
[0070]
Step S5
[0071] In one embodiment of the present invention, the file multimodal collaboration method further includes a playback control step: controlling the collaborative playback of the voice information and operation information on the time axis.
[0072] In one embodiment of the present invention, the file multimodal collaboration method further includes a time axis generation step: generating a time axis and marking the voice information or / and operation information output by the set terminal device within a set time period on the time axis.
[0073] In an embodiment of the present invention, the operation information obtained in the operation information acquisition step includes at least one of the following information: the position information of a file displayed on a set interface; the specific operation object for operating on the file; the operation position coordinates for operating on the file; the operation position coordinates for operating in other interface areas outside the file display area without operating on the file.
[0074] The above method is not only limited to being carried out step by step according to the above steps. Some steps can be carried out simultaneously (such as step S1 and step S3), and some steps can adjust the execution order (such as step S2 and step S3).
[0075] The present invention also discloses an electronic device. Figure 3 It is a schematic diagram of the composition of the electronic device in an embodiment of the present invention; please refer to Figure 3 , at the hardware level, the electronic device includes a memory, a processor, and at least one communication interface; the processor can be a microprocessor, and the memory can include a memory, such as a random access memory (RAM), and can also include a non-volatile memory, etc. Of course, the electronic device can also be provided with other hardware according to needs.
[0076] The processor, the communication interface, and the memory can be interconnected through an internal bus. The memory is used to store programs (which can include an operating system program and application programs); the programs can include program codes, and the program codes can include computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0077] In an embodiment, the processor can read the corresponding program from the non-volatile memory into the memory and then run it; the processor can execute the program stored in the memory and is specifically used to perform the following operations (as Figure 2 shown):
[0078]
Step S1
[0079]
Step S2
[0080]
Step S3
[0081]
Step S4
[0082]
Step S5
[0083]
Step S6
[0084] The present invention further discloses a storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the following steps of the method of the present invention are implemented (as Figure 2 shown):
[0085]
Step S1
[0086]
Step S2
[0087]
Step S3
[0088]
Step S4
[0089]
Step S5
[0090]
Step S6
[0091] In a usage scenario of the present invention, the present invention discloses a real-time LiveDoc audio-visual method, including:
[0092] Audio-visual activity recording: Audio-visual is a creative Chinese translation of "Sync", emphasizing the "synchronization of voice (audio) and document operation intention (visual)", which is the core concept of the present invention; the audio-visual data packet contains a complete session record of voice, document operations, and timestamps.
[0093] Based on LiveDoc technology (see Chinese Patent CN2012800621218), when participants interactively browse a document, they can synchronously record voice, document animations, page navigation, and annotation activities, collectively referred to as the "audio-thought" data stream. The above activities are aligned by timestamps to achieve audio-thought playback (synchronization of voice, document operations, and animations).
[0094] Multiple participants can view the same document in real time, and their operations are synchronized in real time through the LiveDoc server, supporting the use in combination with third-party voice tools (such as phones, WeChat calls).
[0095] Local real-time audio-thought recording: Participants locally record voice (the "audio" part) to ensure high-definition sound quality and no interference; document activities (such as page turning, highlighting) are recorded in real time on the server side (the "thought" part) to form a complete audio-thought data packet.
[0096] Server-side audio-thought synthesis: The voice uploaded after the meeting and the document activities recorded on the server are aligned by a unified time standard to generate an audio-thought playback file; it supports the integrated playback of high-definition voice, infinitely resolvable documents, and animations.
[0097] Audio-thought playback enhancement: During playback, voice, document navigation, and annotations are dynamically synthesized to form an "audio-thought presentation", whose quality is better than the original session.
[0098] Advantages: Technical decoupling: Voice (audio) and document operations (thought) are recorded independently to avoid quality loss during real-time transmission. Compatibility: Adapt to any third-party voice tool (phones, conference systems, etc.).
[0099] The present invention also discloses a file multimodal collaboration system, which includes an audio-thought server and at least one application client, and the audio-thought server is connected to the application client.
[0100] Locally record "audio" data (voice) independently, supporting upload after the meeting; capture "thought" data (document operations) in real time and synchronize them to the server.
[0101] Audio-thought server: Add audio-thought timestamps to all "thought" data to generate a global operation sequence; after the meeting, align the "audio" data and the "thought" data according to the timestamps to synthesize an audio-thought playback file.
[0102] Audio-thought playback module: Render infinitely resolvable documents and high-definition voice, supporting frame-by-frame debugging (such as skipping silent paragraphs, viewing annotations separately).
[0103] In one embodiment, the present invention discloses a system for voice discussion and dynamic document collaborative review, including:
[0104] The AudioThought client is used for local voice recording (Audio) and document operations (Thought); among which, the "Audio" data is independently recorded from the real-time communication tool and uploaded after the meeting; the "Thought" data includes document page turning, annotation or animation operations.
[0105] The AudioThought server is used to synchronize the document operation timestamps and synthesize the AudioThought playback file;
[0106] The AudioThought playback module is used to render a high-definition demonstration with the voice and document operations aligned.
[0107] The present invention also discloses a method for generating an AudioThought playback, including: locally recording voice, the server synchronizing document operations, and synthesizing an AudioThought demonstration according to the timestamps.
[0108] Through the separate collection and intelligent synchronization of "Audio" and "Thought", the real-time LiveDoc AudioThought solves the problems of poor audio and video quality and dependence on dedicated servers in traditional collaborative reviews, and is applicable to scenarios such as distance education and cross-border contract negotiations.
[0109] This version strengthens the penetration of "AudioThought" as a core term, and clarifies the technical highlights by splitting the functions of "Audio" and "Thought".
[0110] The present invention also discloses a method for multi-modal collaboration of files, including:
[0111] Real-time collaboration stage: Participants communicate through a third-party tool (such as WeChat Call) and operate the document in LiveDoc at the same time; the "Audio" and "Thought" data are recorded in parallel without interference.
[0112] AudioThought synthesis stage after the meeting: Upload the local voice, and the server matches it with the document operation timeline; output an AudioThought demonstration file, supporting high-definition playback on multiple terminals.
[0113] In summary, the multi-modal collaboration system, method, electronic device and storage medium for files proposed by the present invention can fuse the voice and operation information for the file into the file through the timeline, improve the efficiency of multi-modal collaboration and review of files, and facilitate a comprehensive understanding of the processing and review process.
[0114] It should be noted that the present application can be implemented in software and / or a combination of software and hardware; for example, it can be implemented using an application specific integrated circuit (ASIC), a general purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium; for example, a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of the present application can be implemented using hardware; for example, as a circuit that cooperates with the processor to execute each step or function.
[0115] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0116] The description and application of the present invention here are illustrative and are not intended to limit the scope of the present invention to the above embodiments. The effects or advantages involved in the embodiments may not be reflected in the embodiments due to various factors. The description of the effects or advantages is not used to limit the embodiments. The deformations and changes of the embodiments disclosed here are possible, and the substitutions and equivalent components of the embodiments are well-known to those of ordinary skill in the art. Those skilled in the art should clearly understand that the present invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the present invention. Other deformations and changes can be made to the embodiments disclosed here without departing from the scope and spirit of the present invention.
Claims
1. A file multi-modal collaboration system, characterized in that, The described file multimodal collaboration system includes: A voice information acquisition module for acquiring voice information output by a set terminal device; A voice information recording module for recording the voice information acquired by the voice information acquisition module and recording the time information corresponding to the voice information; An operation information acquisition module for acquiring operation information on a file output by a set terminal device or / and operation information that causes the content presented on a set interface to change; An operation information recording module for recording the operation information acquired by the operation information acquisition module and recording the time information corresponding to the operation information; A playback file generation module for generating a playback file from the voice information recorded by the voice information recording module and the operation information recorded by the operation information recording module, and when the playback file is played, it can control the coordinated playback of the corresponding voice information and operation information based on the time information.
2. The file multimodal collaboration system according to claim 1, characterized in that: The file multimodal collaboration system further includes a playback control module, and the playback control module is used to control the coordinated playback of the voice information and operation information on the time axis.
3. The file multimodal collaboration system according to claim 1, characterized in that: The file multimodal collaboration system further includes a time axis generation module, and the time axis generation module is used to generate a time axis and mark the voice information or / and operation information output by a set terminal device within a set time period on the time axis.
4. The file multimodal collaboration system according to claim 1, characterized in that: The operation information acquired by the operation information acquisition module includes at least one of the following information: The position information of the file displayed on a set interface; The specific operation object for file operation; The operation position coordinates for file operation; The operation position coordinates for operations performed in other interface areas outside the file display area without operating on the file; The file multimodal collaboration system includes a server and at least one terminal device, and the server is respectively connected to each terminal device; the server includes the voice information recording module, the operation information recording module, and the playback file generation module.
5. A method for multi-modal collaboration of files, characterized in that, The described file multimodal collaboration method includes: A voice information acquisition step: acquiring voice information output by a set terminal device; A voice information recording step: recording the voice information acquired in the voice information acquisition step and recording the time information corresponding to the voice information; An operation information acquisition module for acquiring operation information on a file output by a set terminal device or / and operation information that causes the content presented on a set interface to change; An operation information recording step: recording the operation information acquired in the operation information acquisition step and recording the time information corresponding to the operation information; A playback file generation step: generating a playback file from the voice information recorded in the voice information recording step and the operation information recorded in the operation information recording step, and when the playback file is played, it can display the voice information and operation information corresponding to different time periods on the time axis.
6. The file multimodal collaboration method according to claim 5, characterized in that: The file multimodal collaboration method further includes a playback control step: controlling the collaborative playback of the voice information and the operation information on the timeline.
7. The file multimodal collaboration method according to claim 5, wherein: The file multimodal collaboration method further includes a timeline generation step: generating a timeline and marking the voice information and / or operation information output by a set terminal device within a set time period on the timeline.
8. The file multimodal collaboration method according to claim 5, wherein: The operation information obtained in the operation information acquisition step includes at least one of the following information: The position information of the file displayed on the set interface; The specific operation object for operating on the file; The operation position coordinates for operating on the file; The operation position coordinates for operating in other interface areas outside the file display area without operating on the file.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 5 to 8 are implemented.
10. A storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the steps of the method according to any one of claims 5 to 8 are implemented.