A method, apparatus, equipment and medium for conference playback

By generating personalized meeting replay videos, the problems of high organizational costs and low replay efficiency in traditional meeting models are solved, enabling efficient and personalized review and analysis of meeting content.

CN119653037BActive Publication Date: 2025-12-02CHINA CONSTRUCTION BANK +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411791708.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-12-02
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Traditional meeting models require all participants to gather at the same time and place, which increases the costs of organization and participation. In addition, the recording and playback of meeting content is inefficient, and participants need to spend time browsing the entire meeting to find the information they are interested in.

Method used

By acquiring participants' identifiers, operation records, and speech content, personalized meeting replay videos are generated, adding character model status changes and speech audio with model identifiers. It supports the generation of connecting content, key information tags, and summaries, and utilizes the metaverse space scene for display.

Benefits of technology

It enables efficient and personalized meeting playback, allowing participants to focus on the speeches and status changes of specific participants without having to browse lengthy transcripts, thus improving comprehension and analysis efficiency and saving time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653037B_ABST
    Figure CN119653037B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology, and more particularly to a meeting playback method, apparatus, device, and medium. The electronic device acquires playback parameters carried in the received playback request; acquires the model identifier of the selected persona model stored for the participant's identifier, each operation record, and the occurrence time; and generates a playback video showing the corresponding state changes of the persona model with the model identifier at the corresponding occurrence time. Furthermore, if the participant's identifier has stored speaking content, the audio of that speaking content is added to the playback video at the corresponding time. Therefore, when there is a need to understand certain participants, the playback video for that participant is directly generated, and the audio of that participant's speaking content is added to the playback video. This allows for a more focused attention on the participant's speaking, the state changes of the selected persona model, etc., which helps in a deeper understanding and analysis of the meeting content, greatly saving time costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device and medium for conference playback. Background Technology

[0002] With the rapid development of information technology, especially the widespread adoption of the internet and mobile communication technologies, traditional meeting models have begun to reveal their limitations. First, from a time and space perspective, traditional meetings require all participants to gather at the same time and place, which undoubtedly increases the cost of organization and participation, limiting the accessibility and flexibility of the meeting. This limitation is particularly pronounced for multinational and multi-regional teams, potentially leading to the absence of key personnel and consuming significant resources and time due to travel arrangements. Second, traditional technologies have significant shortcomings in recording and replaying meeting content. Previously, meeting minutes relied primarily on manual note-taking or recording equipment, which is not only inefficient but also prone to missing important information.

[0003] Even with the introduction of playback audio and video recording technology, it often simply involves recording the entire meeting. This playback method is inefficient and cumbersome for participants or meeting subscribers who need to review information about specific attendees. Participants or subscribers often have to spend a significant amount of time browsing through all the meeting-related information to find the content they truly care about. This not only wastes time but may also reduce their efficiency in understanding and absorbing the meeting content. Summary of the Invention

[0004] This application provides a meeting playback method, apparatus, device, and medium to address the problem in the prior art where, when there is a need to understand the information of a specific participant in a meeting, it is necessary to browse the entire meeting information, resulting in resource waste and reduced understanding of the meeting.

[0005] This application provides a meeting playback method, the method including:

[0006] Obtain the playback parameters carried in the received playback request; wherein, the playback parameters include the identifier of at least one participant;

[0007] For each participant identified by at least one identifier, obtain the model identifier of the selected character model, each operation record, and the occurrence time of the character model selected by the participant identified by that identifier; for each operation record, based on the state change information and occurrence time contained in the operation record, add video content of the character model with the model identifier undergoing the corresponding state change at the occurrence time in the playback video; if there is speech content saved for the participant's identifier, add the audio of the speech content at the corresponding time in the playback video based on the speech content and the speech time.

[0008] Furthermore, the method also includes:

[0009] If the playback parameters include a connection identifier and the corresponding connection time, then the connection content saved for the connection identifier is obtained;

[0010] Generate the video corresponding to the connected content;

[0011] The video is added at the transition time specified in the playback video.

[0012] Furthermore, generating the video corresponding to the connecting content includes:

[0013] If the connecting content is text, a preset number of video frames containing the text will be generated.

[0014] Furthermore, the step of adding audio of the speech content to the playback video at the corresponding time based on the speech content and speech time includes:

[0015] If the playback parameters include the waveform and / or pitch frequency of the timbre selected for the identifier of the at least one participant, the timbre and / or pitch of the audio containing the speech content is adjusted to the timbre of the waveform and / or the pitch of the frequency, and the adjusted audio is added to the playback video at the time corresponding to the speech time.

[0016] Furthermore, the method also includes:

[0017] If the playback parameters include key information, then obtain the key information;

[0018] A preset marking operation is performed on a preset area of ​​the video frame containing the key information in the playback video.

[0019] Furthermore, the method also includes:

[0020] If a summary generation request is received, the playback audio in the playback video is converted to text to obtain the target text. The target text and preset prompts for generating the summary are then input into the large model to obtain the target summary output by the large model.

[0021] Further, the step of inputting the target text into the large model and obtaining the target summary output by the large model includes:

[0022] If the target text exceeds a preset length, the target text is divided into a preset number of sub-texts. Each sub-text is input into the large model to obtain the sub-summary output by the large model. The sub-summaries are then concatenated according to the order of each sub-text in the target text to obtain the target summary.

[0023] Furthermore, the method also includes:

[0024] If a meeting setup instruction carrying a model identifier and a spatial scene identifier is received, then the metaverse spatial scene corresponding to the spatial scene identifier and the character model corresponding to the model identifier are constructed in the UI page. The position corresponding to the model identifier in the meeting setup instruction is obtained, the character model is added to the position in the metaverse spatial scene, and the modeling data of the metaverse spatial scene with the added character model is synchronized to the devices of each participant. This allows each participant's device to render and display the metaverse spatial scene containing the character model based on the modeling data.

[0025] Furthermore, the method also includes:

[0026] If a status change information is received from any participant's device, including behavior change information, location change information, morphological change information, and mood change information, the status of the persona model corresponding to that participant's device is adjusted based on the status change information.

[0027] Furthermore, the method also includes:

[0028] If audio or video data is received from any participant's device, the audio or video data is preprocessed and then synchronized to each participant's device.

[0029] This application embodiment also provides a conference playback device, the device comprising:

[0030] The acquisition module is used to acquire the playback parameters carried in the received playback request; wherein, the playback parameters include the identifier of at least one participant;

[0031] The processing module is configured to, for the at least one identified participant, obtain the model identifier of the selected character model, each operation record, and the occurrence time of the character model, which is stored for the identified participant; for each operation record, based on the state change information and occurrence time contained in the operation record, add video content of the character model with the model identifier undergoing the corresponding state change at the occurrence time in the playback video; if there is speech content stored for the identified participant, add the audio of the speech content to the playback video at the corresponding time based on the speech content and the speech time.

[0032] Furthermore, the processing module is also configured to, if the playback parameters include a connection identifier and a corresponding connection time, obtain the connection content saved for the connection identifier; generate the video corresponding to the connection content; and add the video at the connection time of the playback video.

[0033] Furthermore, the processing module is specifically used to generate a preset number of video frames containing the text if the connecting content is text.

[0034] Furthermore, the processing module is specifically configured to, if the playback parameters include the frequency of the waveform and / or pitch of the timbre selected for the identifier of the at least one participant, adjust the timbre and / or pitch of the audio containing the speech content to the timbre of the waveform and / or the pitch of the frequency, and add the adjusted audio to the time corresponding to the speech time in the playback video.

[0035] Furthermore, the processing module is also configured to acquire the key information if the playback parameters include key information, and to perform a preset marking operation on a preset area of ​​the video frame containing the key information in the playback video.

[0036] Furthermore, the processing module is also configured to, upon receiving a summary generation request, convert the playback audio in the playback video into text to obtain target text, input the target text and preset prompt words for prompting summary generation into the large model, and obtain the target summary output by the large model.

[0037] Furthermore, the processing module is specifically used to, if the target text exceeds a preset length, divide the target text into a preset number of sub-texts, input each sub-text into a large model, obtain the sub-summary output by the large model, and concatenate the sub-summaries according to the order of each sub-text in the target text to obtain the target summary.

[0038] Furthermore, the processing module is also configured to, upon receiving a meeting construction instruction carrying a model identifier and a spatial scene identifier, construct a metaverse spatial scene corresponding to the spatial scene identifier and a character model corresponding to the model identifier in the UI page, obtain the position corresponding to the model identifier in the meeting construction instruction, add the character model to the position in the metaverse spatial scene, and synchronize the modeling data of the metaverse spatial scene with the added character model to the devices of each participant; so that each participant's device renders and displays the metaverse spatial scene containing the character model based on the modeling data.

[0039] Furthermore, the processing module is also configured to adjust the state of the person model corresponding to the participant's device based on the state change information received from the device of any participant, the state change information including behavior change information, location change information, morphological change information and mood change information.

[0040] Furthermore, the processing module is also configured to preprocess the audio and video data received from any participant's device and synchronize the preprocessed audio and video data to each participant's device.

[0041] This application also provides an electronic device, which includes a processor for executing a computer program stored in a memory to implement the steps of any of the above-described conference playback methods.

[0042] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above-described conference playback methods.

[0043] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the steps of any of the above-described conference playback methods.

[0044] In this embodiment, the electronic device acquires playback parameters carried in the received playback request. These parameters include the identifier of at least one participant. For each participant, the device acquires the model identifier of the selected persona saved for that participant's identifier, each operation record, and the occurrence time. The device then generates a playback video showing the persona of the model identifier undergoing corresponding state changes at the corresponding occurrence time. Furthermore, if the participant's identifier contains speech content, the device adds audio of that speech content to the playback video at the corresponding time based on the speech content and time. Therefore, when there is a need to understand information about certain participants during the meeting, it is no longer necessary to spend time browsing through the lengthy records of the entire meeting. Instead, a playback video for the corresponding participant is directly generated, and the playback video includes audio of the participant's speech content. This allows for a more focused attention on the participant's speech, the state changes of the selected persona, etc., facilitating a deeper understanding and analysis of the meeting content and significantly saving time. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 A schematic diagram of a meeting playback method provided in this application embodiment;

[0047] Figure 2 A schematic diagram of the meeting playback process provided in this application embodiment;

[0048] Figure 3 A schematic diagram illustrating a meeting implementation process provided in an embodiment of this application;

[0049] Figure 4 A schematic diagram of a conference playback device provided in an embodiment of this application;

[0050] Figure 5 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, a further detailed description of this application will be provided below with reference to the accompanying drawings. Obviously, the embodiments described in this application are merely some embodiments, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0053] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0054] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0055] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still make changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features therein. Such changes or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0057] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These embodiments should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solutions of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0058] The acquisition, transmission, storage, and use of data in this application all comply with the requirements of relevant national laws and regulations.

[0059] Before introducing the meeting playback method provided in the embodiments of this application, the background technology of the embodiments of this application will be introduced first for ease of understanding.

[0060] With the rapid development of information technology, especially the widespread adoption of the internet and mobile communication technologies, traditional meeting models have begun to reveal their limitations. First, from a time and space perspective, traditional meetings require all participants to gather at the same time and place, which undoubtedly increases the cost of organization and participation, limiting the accessibility and flexibility of the meeting. This limitation is particularly pronounced for multinational and multi-regional teams, potentially leading to the absence of key personnel and consuming significant resources and time due to travel arrangements. Second, traditional technologies have significant shortcomings in recording and replaying meeting content. Previously, meeting minutes relied primarily on manual note-taking or recording equipment, which is not only inefficient but also prone to missing important information.

[0061] Even with the introduction of playback audio and video recording technology, it often simply involves recording the entire meeting. This playback method is inefficient and cumbersome for participants or meeting subscribers who need to review information about specific attendees. Participants or subscribers often have to spend a significant amount of time browsing through all the meeting-related information to find the content they truly care about. This not only wastes time but may also reduce their efficiency in understanding and absorbing the meeting content.

[0062] In view of the above, embodiments of this application provide a meeting playback method, apparatus, device, and medium. The meeting playback method includes: an electronic device acquiring playback parameters carried in a received playback request; wherein the playback parameters include the identifier of at least one participant; for each participant, acquiring the model identifier of the selected persona saved for that participant's identifier, each operation record, and the occurrence time; for each operation record, generating a persona with the model identifier undergoing a corresponding state change at the occurrence time in the playback video based on the state change information corresponding to the operation record and the occurrence time; if there is speech content saved for the participant's identifier, adding audio of the speech content at the corresponding time in the playback video based on the speech content and the speech time.

[0063] The following is a detailed description of the technical terms used in this application:

[0064] A 3D graphics engine is a software and hardware system used to create and render 3D graphics. It can handle various 3D graphics-related tasks such as 3D models, textures, lighting, shadows, and physical simulations, ultimately generating visual 3D scenes.

[0065] Real-Time Communication (RTC) technology allows any communication service between the sender and receiver to occur almost simultaneously. Compared to general communication, RTC focuses more on real-time performance, transmitting audio, video, data, and other information in an extremely short time, thus enabling real-time communication and interaction. It is suitable for large-scale, low-latency, point-to-point use cases.

[0066] Automatic Speech Recognition (ASR): ASR uses signal processing and pattern recognition technology to convert text information in speech into computer-readable input information, such as keystrokes, binary codes, or character sequences, thereby enabling human-computer interaction.

[0067] Natural Language Processing (NLP) is a machine learning technique that enables computers to interpret, process, and understand human language. It comprises two core tasks: natural language understanding and natural language generation. A pre-trained language model is a type of NLP model.

[0068] Pre-trained language models: A pre-trained language model is a machine learning technique that learns the rules and semantic information of language by training on a large amount of data, and encodes this knowledge into a model that can be used for a variety of natural language processing tasks.

[0069] Example 1:

[0070] Figure 1 A schematic diagram of a meeting playback method provided in this application embodiment, the process including the following steps:

[0071] S101: Obtain the playback parameters carried in the received playback request; wherein the playback parameters include the identifier of at least one participant.

[0072] The meeting playback method provided in this application is applied to an electronic device, which can be a smart device such as a PC or a server.

[0073] When participants or meeting subscribers wish to review the meeting content, they can select playback parameters on a designated interface on their own device or a pre-defined device, and initiate a playback request by clicking a preset button such as "Playback." Upon receiving the playback request, the electronic device can retrieve the playback parameters carried in the request. "Meeting subscriber" refers to a user who has pre-subscribed to the meeting content to be able to listen to and access the meeting and its replay. It should be noted that users with administrative privileges decide whether to allow a user to subscribe to the meeting and become a meeting subscriber.

[0074] The replay parameters include the identifier of at least one participant. This means that when a participant or meeting subscriber needs to review information from a meeting involving a specific participant or a few participants, they can select the identifier of the participant they wish to review through a designated interface on their own device or a pre-set device. Each participant's identifier is unique and can be a username, email address, or similar information to ensure that the electronic device can accurately identify the corresponding participant.

[0075] S102: For the at least one identified participant, obtain the model identifier of the selected character model, each operation record, and the occurrence time of the character model stored for the identified participant; for each operation record, according to the state change information and occurrence time contained in the operation record, add video content of the character model with the model identifier performing the corresponding state change at the occurrence time in the playback video; if there is speech content stored for the identified participant, add the audio of the speech content to the playback video at the corresponding time according to the speech content and speech time.

[0076] To enable meeting playback, the electronic device pre-stores the model identifier of the persona selected by each participant, as well as each participant's operation record and the corresponding occurrence time. The model identifier of the persona selected by each participant and the operation record and the corresponding occurrence time can be recorded in a behavior operation record file. Based on the information recorded in the operation record file, the electronic device can determine the persona selected by the participant and play back the corresponding state change video.

[0077] Specifically, the electronic device can retrieve, from the operation log file, the model identifier of the selected persona for each participant, each operation record, and the occurrence time. Each operation record includes information on the state changes of the persona; these state changes include behavioral changes, positional changes, morphological changes, and mood changes. Behavioral changes include actions such as shaking hands or clapping; positional changes refer to the participant's position within the corresponding metaverse space scene; morphological changes include changes from sitting to standing; and mood changes include focus and relaxation. For each operation record, the electronic device can determine the state change information and the occurrence time, where the occurrence time refers to the time when the corresponding state change occurs. The electronic device can replay the video at that occurrence time, create a target persona model for the persona with that model identifier after the state change corresponding to that state change information, and maintain the target persona model with that state change in the replay video until the next state change of the persona with that model identifier.

[0078] In real-world scenarios, participants often speak during meetings. To ensure accurate and effective playback, electronic devices store the spoken content for each participant, with the mapping between participant identifiers and spoken content recorded in the meeting recording file. The spoken content consists of the text data of the participants' statements during the meeting. If the electronic device stores spoken content for a participant's identifier, it retrieves the corresponding speaking time and adds the audio of that speaking content to the playback video at the corresponding time. In one possible implementation, the electronic device can play the spoken content with a preset tone and timbre to obtain the audio. The speaking time can be one or more of the following: the start time of the speech, the end time of the speech, or the duration of the speech.

[0079] During this process, the electronic devices will ensure that the status changes and speech content corresponding to each operation record are added in the order of their actual occurrence time in the meeting. Whether it is a single participant's speech or presentation, or an interactive discussion between multiple participants, it will be accurately integrated into the playback video according to the timeline.

[0080] The sorting method provided in this embodiment not only improves the logic and coherence of the playback, but also makes it easier for participants or meeting subscribers to track meeting progress and review the discussion content of specific participants during the meeting. Furthermore, the generated playback video may include additional features, such as timeline markers, to further enhance the viewing experience and review efficiency for participants or meeting subscribers.

[0081] In this embodiment, the electronic device acquires playback parameters carried in the received playback request. These parameters include the identifier of at least one participant. For each participant, the device acquires the model identifier of the selected persona saved for that participant's identifier, each operation record, and the occurrence time. The device then generates a playback video showing the persona of the model identifier undergoing corresponding state changes at the corresponding occurrence time. Furthermore, if the participant's identifier contains speech content, the device adds audio of that speech content to the playback video at the corresponding time based on the speech content and time. Therefore, when there is a need to understand information about certain participants during the meeting, it is no longer necessary to spend time browsing through the lengthy records of the entire meeting. Instead, a playback video for the corresponding participant is directly generated, and the playback video includes audio of the participant's speech content. This allows for a more focused attention on the participant's speech, the state changes of the selected persona, etc., facilitating a deeper understanding and analysis of the meeting content and significantly saving time.

[0082] Example 2:

[0083] To improve the quality of meeting playback, based on the above embodiments, the method in this application embodiment further includes:

[0084] If the playback parameters include a connection identifier and the corresponding connection time, then the connection content saved for the connection identifier is obtained;

[0085] Generate the video corresponding to the connected content;

[0086] The video is added at the transition time specified in the playback video.

[0087] To improve the quality of meeting playback, the playback videos generated by electronic devices can also include videos corresponding to connecting content, such as participant introductions or meeting introductions.

[0088] Specifically, when a participant or conference subscriber initiates a replay request, they can select a link identifier containing specific transition content as part of the replay parameters. Each transition content has a unique link identifier. When the electronic device receives replay parameters containing the link identifier, it can obtain the link identifier and its corresponding transition time. Based on a pre-saved mapping between identifiers and content, the electronic device retrieves the transition content saved for that link identifier and obtains the corresponding video. In one possible implementation, if the transition content is in video format, the video corresponding to that transition content is directly retrieved.

[0089] In order to effectively integrate the transition content into the playback video, the electronic device adds the video to the playback video at the transition time, based on the transition time included in the playback parameters.

[0090] To improve the quality of meeting playback, based on the above embodiments, in this embodiment, obtaining the video corresponding to the connecting content includes:

[0091] If the connecting content is text, a preset number of video frames containing the text will be generated.

[0092] In real-world scenarios, the connecting content may be text. To add text-based connecting content to the playback video, the electronic device can generate video frames containing that text. Furthermore, considering that a single frame of connecting text might be difficult to see clearly during fast playback, the electronic device will also generate a preset number of repeated frames of that video frame, based on actual needs, to ensure that the connecting content is fully noticed and effectively read by attendees or subscribers.

[0093] Example 3:

[0094] To improve the quality of meeting playback, based on the above embodiments, in this embodiment, adding audio of the speech content to the playback video at the corresponding time according to the speech content and speaking time includes:

[0095] If the playback parameters include the waveform and / or pitch frequency of the timbre selected for the identifier of the at least one participant, the timbre and / or pitch of the audio containing the speech content is adjusted to the timbre of the waveform and / or the pitch of the frequency, and the adjusted audio is added to the playback video at the time corresponding to the speech time.

[0096] To improve the quality of meeting playback, the playback video generated by the electronic device can also generate audio that meets the needs of the participant or meeting subscriber who initiated the playback request.

[0097] Specifically, when a participant or conference subscriber initiates a playback request, they can select a waveform and / or pitch frequency corresponding to the identifier of at least one participant. Timbre refers to the characteristic of a sound, determining its quality and style. Pitch refers to the highness or lowness of a sound's frequency; higher frequencies produce higher pitches, and lower frequencies produce lower pitches. When initiating a playback request, participants or conference subscribers can select the waveform representing the timbre and the frequency representing the pitch. Different waveforms produce different timbres, and different pitches have different frequencies. By adjusting the waveform of the timbre and / or adjusting the frequency of the pitch, different sounds can be generated. If the playback request explicitly includes the waveform and / or pitch frequency of the timbre selected for these participant identifiers, the electronic device can obtain the target waveform and / or target pitch frequency of the timbre contained in the playback parameters.

[0098] The electronic device can acquire audio containing the speech content and adjust the timbre and / or pitch of the audio to the timbre corresponding to the target waveform and / or the pitch corresponding to the target frequency contained in the playback parameters. When the timbre is not adjusted, the timbre in the audio can be the participant's own timbre or a preset timbre. When the pitch is not adjusted, the pitch in the audio can be the participant's own pitch or a preset pitch.

[0099] After acquiring the adjusted audio, the electronic device can add the audio to the time corresponding to the speech time in the playback video.

[0100] The method provided in this application ensures that the audio portion of the playback can be presented with a specified timbre and / or pitch according to the preferences of the participants or conference subscribers, thereby providing a more personalized and immersive playback experience.

[0101] In this embodiment, the playback controller can obtain playback parameters and send control commands to the spatial renderer, which then executes the playback of the spatial meeting. The playback parameters may include: participant identifiers as described in this embodiment, transitional content between meeting segments, playback speed, participant-level adjustments (such as pitch and timbre), storage method, and format. The content generated by the spatial renderer can be displayed through the playback client, and can also generate files in the corresponding format for persistent storage in a specified manner. By including the corresponding participant identifier in different playback requests, the playback perspective of the corresponding participant can be switched.

[0102] Figure 2 This is a schematic diagram illustrating the meeting playback process provided in an embodiment of this application.

[0103] Depend on Figure 2 It is known that the electronic device locally stores data such as behavior operation log files, audio clips of meeting speeches, and participant identification. The playback controller obtains the playback parameters and sends control commands to the spatial renderer. The spatial renderer executes the playback of the spatial meeting and sends the playback video to the participants' devices. Furthermore, the playback video is persistently saved.

[0104] In this embodiment, the electronic device integrates the participants' operation records and speaking content during the meeting. Simultaneously, the playback video includes information about at least one participant selected by the participant or meeting subscriber during the meeting, essentially replaying the meeting from the participant's perspective—a presentation centered on the individual participant. Furthermore, it supports outputting and storing the entire meeting in audio and video format, greatly enriching the content of the meeting records and providing strong support for subsequent meeting review, analysis, and research.

[0105] Example 4:

[0106] To improve the quality of meeting playback, based on the above embodiments, the method in this application embodiment further includes:

[0107] If the playback parameters include key information, then obtain the key information;

[0108] A preset marking operation is performed on a preset area of ​​the video frame containing the key information in the playback video.

[0109] Since participants or conference subscribers who need to replay the meeting may have a need to view certain key information during the replay process, in this embodiment of the application, the electronic device can mark the video frames containing such key information in the replay video.

[0110] Specifically, when participants or conference subscribers initiate a replay request, they can select corresponding key information. If the replay request explicitly contains key information, the electronic device can retrieve it. Next, the electronic device will convert the audio in the replay video to text, identify the key text containing that key information, and perform a preset marking operation on the video frames containing that key text. The video frames containing the key text refer to the frames that narrate that key text during playback. This preset marking operation can be highlighting subtitles or highlighting video frames.

[0111] Example 5:

[0112] To improve information acquisition efficiency and facilitate meeting review, based on the above embodiments, the method in this application embodiment further includes:

[0113] If a summary generation request is received, the playback audio data in the playback video is converted to text to obtain the target text. The target text and preset prompts for generating the summary are then input into the large model to obtain the target summary output by the large model.

[0114] When an electronic device receives a summary generation request, it can first use ASR (Automatic Language Representation) technology to perform text conversion processing on the playback video to obtain the accurate target text. The obtained target text is then input into a large model, which, with its powerful natural language processing capabilities, can efficiently analyze and extract the core information, ultimately outputting the target summary corresponding to the target text.

[0115] It should be noted that, to generate the target summary, the electronic device can also locally store preset prompts. These prompts are used to guide the large model in generating the corresponding summary. The electronic device inputs these prompts along with the target text into the large model. Leveraging the capabilities of the large model, it can be updated, supporting flexible and controllable target summary generation. The electronic device can generate meeting minutes containing the target summary. Furthermore, the electronic device can provide optional and configurable prompts, allowing it to retrieve prompts carried in playback requests and input them along with the large model, facilitating adjustments and optimizations to the content generated by the large model.

[0116] Currently, the application of metaverse in assisting with meeting content processing lacks a complete and systematic processing model. This limits the efficiency of meeting content processing and affects the final processing effect. The embodiments of this application utilize large model technology to perform processing such as summary generation on meeting content. This can greatly improve the efficiency of subsequent processing, reduce the workload of manual processing, and enable participants to quickly obtain the essence and key information of the meeting content.

[0117] Example 6:

[0118] To improve the accuracy of the summary generation, based on the above embodiments, in this embodiment, the step of inputting the target text into a large model and obtaining the target summary output by the large model includes:

[0119] If the target text exceeds a preset length, the target text is divided into a preset number of sub-texts. Each sub-text is input into a large model to obtain a sub-summary output by the large model. The sub-summaries are then concatenated according to the order of each sub-text in the target text to obtain the target summary.

[0120] If the target text exceeds a preset length, the accuracy of the summary generated by the large model may be low. To improve the accuracy of summary generation, the electronic device can segment the target text into a preset number of sub-texts. In one possible implementation, the electronic device can divide the total number of characters in the target text by the preset number of segments, i.e., the preset number described in the embodiments of this application, to obtain the number of characters that each sub-text should contain. Based on the calculated number of characters, the target text is evenly segmented into the preset number of sub-texts. In another possible implementation, the electronic device can use punctuation marks (such as periods, question marks, exclamation marks, etc.) as segmentation points. The target text is traversed to find all punctuation mark positions. Based on the punctuation mark positions and the preset number of segments, i.e., the preset number described in the embodiments of this application, the target text is segmented into corresponding sub-texts.

[0121] After acquiring each sub-text segment, the electronic device can input each segment separately into a large model to obtain the corresponding sub-summary from the output of the large model. To ensure the coherence and accuracy of the summary, the electronic device will sequentially concatenate the generated sub-summaries according to the order of these sub-texts in the original target text, ultimately forming a complete and accurate target summary.

[0122] The method provided in this application, through automatic segmentation, integration, and iteration, supports situations exceeding the single-processing context of a large model. This not only improves the efficiency of summary generation but also further enhances the accuracy and readability of the summary content.

[0123] Example 7:

[0124] To ensure accurate and effective meeting implementation, based on the above embodiments, the method in this application embodiment further includes:

[0125] If a meeting setup instruction carrying a model identifier and a spatial scene identifier is received, then the metaverse spatial scene corresponding to the spatial scene identifier and the character model corresponding to the model identifier are constructed in the UI page. The position corresponding to the model identifier in the meeting setup instruction is obtained, the character model is added to the position in the metaverse spatial scene, and the modeling data of the metaverse spatial scene with the added character model is synchronized to the devices of each participant. This allows each participant's device to render and display the metaverse spatial scene containing the character model based on the modeling data.

[0126] To enhance meeting engagement, electronic devices can be used to implement meetings based on the metaverse.

[0127] Specifically, if an electronic device receives a meeting setup instruction carrying model identifiers and spatial scene identifiers, it can retrieve these identifiers from the received instruction. In one possible implementation, only participants with administrative privileges are authorized to include spatial scene identifiers in their instructions. The electronic device can then retrieve the corresponding metaverse spatial scene resources and the corresponding character model resources from a predefined resource library based on the received spatial scene identifiers and model identifiers. These resources may include 3D model files, texture maps, animation data, etc.

[0128] Electronic devices can utilize loaded metaverse space scene resources to construct a virtual metaverse space scene on the UI page. This scene could be a conference room, an outdoor environment, or any other virtual space that meets the needs of a meeting. Furthermore, the electronic device can utilize loaded character model resources to construct a virtual character model, obtain the position of the character model corresponding to the meeting construction command, and add the constructed character model to that position in the metaverse space scene. The character model constructed by the electronic device may include features such as appearance, actions, and expressions. Additionally, the metaverse space scene constructed by the electronic device can also, depending on whether the meeting construction command includes a lighting mode and corresponding physical materials, construct a metaverse space scene containing lighting of the corresponding lighting mode and a conference table with the corresponding physical materials.

[0129] Among them, electronic devices can generate formats such as glb, fbx, and usdz for different types of content and display terminals.

[0130] The electronic devices package the rendered modeling data of the metaverse space scene containing the character models and synchronize it in real time to the devices of all participants and conference subscribers. After receiving the modeling data, the devices of participants and conference subscribers will use their own graphics rendering engines to parse and render the data, and display the rendered metaverse space scene containing the corresponding character models.

[0131] The meetings described in this application can be implemented based on professional 3D graphics engines, including but not limited to Unity, Unreal Engine, and Web Graphics Library (WebGL) engines for web browsers. Participants can participate in the meetings using various types of devices, including but not limited to mobile phones, computers, and virtual reality (VR) devices.

[0132] The methods provided in this application enable electronic devices to offer participants a richer, more interactive, and immersive metaverse meeting experience.

[0133] It's worth noting that with the rapid advancement of communication and computer technologies, global conferences and events have largely shifted to remote online formats in recent years, effectively overcoming the limitations of location and time. Metaverse, a product of technological and social progress, is leading a new trend of merging the virtual and real worlds, its development having progressed from initial conceptual nascent stages to a new phase characterized by robust technology and rich application scenarios. However, while current mainstream online conferencing systems are convenient and efficient, they struggle to fully replicate the immersive atmosphere and deep engagement of traditional offline meetings, resulting in a relatively rigid and limited interactive experience. In contrast, while metaverse digital space applications can provide participants or conference subscribers with a strong immersive experience and a realistic sense of presence, their comprehensive management of meeting content, particularly the completeness of meeting support functions, still needs improvement. Traditional conferencing systems primarily rely on voice and video communication, exhibiting significant shortcomings in providing an immersive experience and rich interactivity. Furthermore, metaverse applications lack a systematic and complete workflow for supporting meeting content processing, resulting in insufficient efficiency in post-conference processing and ultimately affecting the quality of the meeting outcomes. The method proposed in this application not only deeply integrates the immersive experience and high sense of presence of the metaverse, but also carefully constructs a complete meeting processing workflow and a powerful functional system, aiming to create a comprehensive meeting solution for participants that integrates the metaverse space meeting experience with all-round and efficient auxiliary processing functions.

[0134] Example 8:

[0135] To ensure accurate and effective meeting implementation, based on the above embodiments, the method in this application embodiment further includes:

[0136] If a status change information is received from any participant's device, including behavior change information, location change information, morphological change information, and mood change information, the status of the persona model corresponding to that participant's device is adjusted based on the status change information.

[0137] In real-world scenarios, participants may need to express changes in their state. When this is the case, they can select the desired state change information on their device's preset page and click a preset button, such as the "Change" button. The electronic device will then receive the state change information sent by the participant's device. This state change information includes changes in behavior, location, form, and mood. Specifically, behavioral changes include actions such as shaking hands or clapping; location changes indicate the participant's position within the corresponding metaverse space; form changes include sitting or standing; and mood changes include focus or relaxation.

[0138] When an electronic device receives status change information from any participant's device, it can adjust the state of the persona model corresponding to the model identifier selected by that participant. Specifically, the electronic device can store resources of the adjusted state model for that persona model. Based on these resources, it constructs the adjusted persona model to ensure that the persona model can accurately reflect the participant's real-time situation. This dynamic adjustment not only enhances the immersion and interactivity of the meeting but also allows remote participants to more intuitively and realistically experience the atmosphere and status changes of other participants, thereby promoting more effective communication and collaboration.

[0139] The spatial synchronization and recording module of the electronic device enables synchronization of meeting-related activities when participant character models enter the metaverse space scene. The synchronization of meeting-related activities described here includes: data synchronization, scene synchronization, and interaction synchronization. Data synchronization refers to the real-time synchronization of state changes of different character models within the metaverse space scene, ensuring a consistent experience for all participants. This state change information can be included in the operation log. The electronic device records operation records containing this state change information for each participant's identifier, along with the corresponding time of occurrence. Scene synchronization refers to the real-time synchronization of elements such as scenes, props, and characters within the metaverse, ensuring a consistent visual experience for all participants and meeting subscribers. Interaction synchronization refers to the real-time synchronization of interactive behaviors between different participant character models, such as greetings, handshakes, and applause, ensuring a consistent social experience for all participants and meeting subscribers within the metaverse.

[0140] Simultaneously, the electronic devices combine the meeting data of each participant with the time of occurrence to generate corresponding data records, which are then associated with the participant's identifier and stored as derivative content of the spatial meeting. Storage methods can employ a combination of time-series databases and object storage, or a document database, and support data compression and the construction of efficient indexes.

[0141] In this embodiment of the application, in order to address the lack of complete meeting content processing in the metaverse application, an auxiliary meeting content processing function is set up to improve the capability range of the metaverse space application in meeting scenarios and enhance the efficiency of meeting result output.

[0142] Traditional conferencing systems primarily rely on voice and video communication, but these methods fall short in terms of participant immersion and interactivity. This application utilizes metaverse space technology to create highly realistic virtual meeting scenarios. Participants can see each other in the virtual environment, creating a sense of presence and enhancing the feeling of being there. Combined with virtual reality technology, participants can explore the virtual meeting space more freely and engage in more natural and intuitive interactions, enabling them to participate more deeply in the meeting.

[0143] Example 9:

[0144] To ensure accurate and effective meeting implementation, based on the above embodiments, the method in this application embodiment further includes:

[0145] If audio or video data is received from any participant's device, the audio or video data is preprocessed, and the preprocessed audio or video data is synchronized to the devices of each participant and meeting subscriber.

[0146] During the meeting, participants can express their views through audio and video data. At this time, the participants' devices will send audio and video data to the electronic device. The participants' devices can perform multimedia encoding and compression on the audio content and send the audio and video data to the electronic device via streaming and file bypass transmission. If the electronic device receives audio and video data from any participant's device, it can preprocess the data. This preprocessing includes encoding and noise reduction. Specifically, the electronic device's audio acquisition and transmission module can collect audio and video data from each participant in the metaverse space scene and perform corresponding encoding and noise reduction processing. Furthermore, the electronic device will synchronize the preprocessed audio and video data to the devices of each participant and meeting subscriber via RTC real-time streaming. This synchronization can be achieved through the electronic device's real-time communication and media processing.

[0147] The devices of attendees and conference subscribers receive the audio and video data, perform corresponding decoding processing, and then play the audio and display the video on the receiving end. Due to data transmission jitter, the audio and video receiving and display module also features jitter buffering and packet loss compensation functions to provide reliable audio and video playback quality.

[0148] The audio and video recording module of the electronic device records the occurrence time of each audio and video data stream, which is saved along with the audio and video data as supporting information. Audio stream data is typically saved as multiple single-stream files independently for each participant, or it can be saved as a single mixed-stream file containing the audio of all speakers.

[0149] In actual meetings, participants often speak alternately. To reduce the size of generated audio files and facilitate flexible, individual processing of each speaker's content, meeting recordings only generate corresponding audio clip files when a participant speaks, resulting in discontinuous data in time. Simultaneously, associated metadata, including the participant's identifier, is embedded in the audio clip files for use in subsequent meeting support functions.

[0150] In this embodiment, after a meeting, the electronic device can utilize advanced post-routing processing technology to perform in-depth analysis and multi-dimensional, multimedia processing of the meeting content. This processing method allows the meeting content to be converted into various forms of information (such as text recordings, audio summaries, video clips, etc.) to meet different information presentation needs. More importantly, this processed meeting content can be intelligently driven according to different business processing logics. This means that, based on specific business needs, the system can automatically select the most suitable processing method or output format. For example, when a quick understanding of the meeting's key points is needed, the system may generate a concise text summary; while when a detailed review of the meeting is required, a complete video recording can be provided. Furthermore, the relationship between meeting content and business routing can be flexible and varied. One mode is "one-to-one," where specific meeting content directly corresponds to specific business processing logic, achieving precise matching; another mode is "one-to-many," where the same meeting content can be processed into multiple forms or output to multiple business systems according to different business needs. This flexibility ensures that meeting content can be fully utilized, providing strong support for various business activities of the enterprise.

[0151] Figure 3 This is a schematic diagram illustrating a meeting implementation process provided in an embodiment of this application.

[0152] Electronic devices can construct a metaverse space, specifically including spatial environment construction and spatial content display. Spatial environment construction refers to creating a metaverse space scene and adding character models to it. Spatial content display refers to synchronizing the modeling data of the metaverse space scene with added character models to the devices of all participants. The meeting also includes real-time communication and media processing, specifically including audio and video acquisition and transmission modules, and audio and video reception and display modules. The audio and video acquisition and transmission module refers to each participant's device collecting audio and video data and sending it to the electronic device. The audio and video reception and display module refers to the electronic device synchronizing the received audio and video data to the devices of other participants. Furthermore, the electronic devices also perform auxiliary processing of meeting content, specifically including meeting minutes generation and meeting routing to the devices of other participants.

[0153] Example 10:

[0154] Based on the same inventive concept, embodiments of this application provide a meeting playback device. Figure 4 Please refer to the schematic diagram of a conference playback device provided in this application embodiment. Figure 4 The device includes:

[0155] The acquisition module 401 is used to acquire the playback parameters carried in the received playback request; wherein, the playback parameters include the identifier of at least one participant;

[0156] The processing module 402 is configured to, for the at least one identified participant, obtain the model identifier of the selected character model, each operation record, and the occurrence time of the character model, which is stored for the identifier of the participant; for each operation record, based on the state change information and occurrence time contained in the operation record, add video content of the character model with the model identifier undergoing the corresponding state change at the occurrence time in the playback video; if there is speech content stored for the identifier of the participant, add the audio of the speech content at the corresponding time in the playback video based on the speech content and the speech time.

[0157] In one possible implementation, the processing module 402 is further configured to: if the playback parameters include a connection identifier and a corresponding connection time, obtain the connection content saved for the connection identifier; generate a video corresponding to the connection content; and add the video at the connection time of the playback video.

[0158] In one possible implementation, the processing module 402 is specifically configured to generate a preset number of video frames containing the text if the connecting content is text.

[0159] In one possible implementation, the processing module 402 is specifically configured to, if the playback parameters include the frequency of the waveform and / or pitch of the timbre selected for the identifier of the at least one participant, adjust the timbre and / or pitch of the audio containing the speech content to the timbre of the waveform and / or the pitch of the frequency, and add the adjusted audio to the time corresponding to the speech time in the playback video.

[0160] In one possible implementation, the processing module 402 is further configured to acquire the key information if the playback parameters include key information, and perform a preset marking operation on a preset area of ​​the video frame containing the key information in the playback video.

[0161] In one possible implementation, the processing module 402 is further configured to, upon receiving a summary generation request, convert the playback audio in the playback video to text to obtain target text, input the target text and preset prompt words for prompting summary generation into a large model, and obtain the target summary output by the large model.

[0162] In one possible implementation, the processing module 402 is specifically used to, if the target text exceeds a preset length, divide the target text into a preset number of sub-texts, input each sub-text into a large model, obtain the sub-summary output by the large model, and concatenate the sub-summaries according to the order of each sub-text in the target text to obtain the target summary.

[0163] In one possible implementation, the processing module 402 is further configured to, upon receiving a meeting construction instruction carrying a model identifier and a spatial scene identifier, construct a metaverse space scene corresponding to the spatial scene identifier and a character model corresponding to the model identifier in the UI page, obtain the position corresponding to the model identifier in the meeting construction instruction, add the character model to the position in the metaverse space scene, and synchronize the modeling data of the metaverse space scene with the added character model to the devices of each participant; so that each participant's device renders and displays the metaverse space scene containing the character model based on the modeling data.

[0164] In one possible implementation, the processing module 402 is further configured to, if it receives status change information sent by any participant's device, the status change information including behavior change information, location change information, morphological change information and mood change information, adjust the state of the person model corresponding to the participant's device based on the status change information.

[0165] In one possible implementation, the processing module 402 is further configured to preprocess the audio and video data received from any participant's device and synchronize the preprocessed audio and video data to each participant's device.

[0166] Example 11:

[0167] Based on the same inventive concept, embodiments of this application provide an electronic device that can implement the steps of the meeting playback method described above. Figure 5 This application provides a schematic diagram of an electronic device structure, such as... Figure 5 As shown, it includes: processor 501, communication interface 502, memory 503 and communication bus 504, wherein processor 501, communication interface 502 and memory 503 communicate with each other through communication bus 504.

[0168] The memory 503 stores a computer program. When the program is executed by the processor 501, the processor 501 performs the following steps:

[0169] Obtain the playback parameters carried in the received playback request; wherein, the playback parameters include the identifier of at least one participant;

[0170] For each participant identified by at least one identifier, obtain the model identifier of the selected character model, each operation record, and the occurrence time of the character model selected by the participant identified by that identifier; for each operation record, based on the state change information and occurrence time contained in the operation record, add video content of the character model with the model identifier undergoing the corresponding state change at the occurrence time in the playback video; if there is speech content saved for the participant's identifier, add the audio of the speech content at the corresponding time in the playback video based on the speech content and the speech time.

[0171] In one possible implementation, the method further includes:

[0172] If the playback parameters include a connection identifier and the corresponding connection time, then the connection content saved for the connection identifier is obtained;

[0173] Generate the video corresponding to the connected content;

[0174] The video is added at the transition time specified in the playback video.

[0175] In one possible implementation, generating the video corresponding to the connecting content includes:

[0176] If the connecting content is text, a preset number of video frames containing the text will be generated.

[0177] In one possible implementation, adding audio of the speech content to the playback video at the corresponding time based on the speech content and speech time includes:

[0178] If the playback parameters include the waveform and / or pitch frequency of the timbre selected for the identifier of the at least one participant, the timbre and / or pitch of the audio containing the speech content is adjusted to the timbre of the waveform and / or the pitch of the frequency, and the adjusted audio is added to the playback video at the time corresponding to the speech time.

[0179] In one possible implementation, the method further includes:

[0180] If the playback parameters include key information, then obtain the key information;

[0181] A preset marking operation is performed on a preset area of ​​the video frame containing the key information in the playback video.

[0182] In one possible implementation, the method further includes:

[0183] If a summary generation request is received, the playback audio in the playback video is converted to text to obtain the target text. The target text and preset prompts for generating the summary are then input into the large model to obtain the target summary output by the large model.

[0184] In one possible implementation, the step of inputting the target text into a large model and obtaining the target summary output by the large model includes:

[0185] If the target text exceeds a preset length, the target text is divided into a preset number of sub-texts. Each sub-text is input into the large model to obtain the sub-summary output by the large model. The sub-summaries are then concatenated according to the order of each sub-text in the target text to obtain the target summary.

[0186] In one possible implementation, the method further includes:

[0187] If a meeting setup instruction carrying a model identifier and a spatial scene identifier is received, then the metaverse spatial scene corresponding to the spatial scene identifier and the character model corresponding to the model identifier are constructed in the UI page. The position corresponding to the model identifier in the meeting setup instruction is obtained, the character model is added to the position in the metaverse spatial scene, and the modeling data of the metaverse spatial scene with the added character model is synchronized to the devices of each participant. This allows each participant's device to render and display the metaverse spatial scene containing the character model based on the modeling data.

[0188] In one possible implementation, the method further includes:

[0189] If a status change information is received from any participant's device, including behavior change information, location change information, morphological change information, and mood change information, the status of the persona model corresponding to that participant's device is adjusted based on the status change information.

[0190] In one possible implementation, the method further includes:

[0191] If audio or video data is received from any participant's device, the audio or video data is preprocessed and then synchronized to each participant's device.

[0192] Since the principle of the above-mentioned electronic device in solving the problem is similar to that of the meeting playback method, the implementation of the above-mentioned electronic device can be found in the embodiments of the method, and the repeated parts will not be described again.

[0193] The communication bus mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. Communication interface 502 is used for communication between the above-mentioned electronic device and other devices. The memory can include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.

[0194] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0195] Example 12:

[0196] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to execute any of the meeting playback methods discussed above. Since the principle by which the above-described computer-readable storage medium solves the problem is similar to that of the meeting playback method, the implementation of the above-described computer-readable storage medium can be referred to the implementation of the method, and repeated details will not be elaborated further.

[0197] Based on the same inventive concept, this application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to execute any of the meeting playback methods discussed above. Since the principle by which the above computer program product solves the problem is similar to that of any of the meeting playback methods discussed above, the implementation of the above computer program product can be referred to the implementation of the method, and repeated details will not be elaborated further.

[0198] In this embodiment, the electronic device acquires playback parameters carried in the received playback request. These parameters include the identifier of at least one participant. For each participant, the device acquires the model identifier of the selected persona saved for that participant's identifier, each operation record, and the occurrence time. The device then generates a playback video showing the persona of the model identifier undergoing corresponding state changes at the corresponding occurrence time. Furthermore, if the participant's identifier contains speech content, the device adds audio of that speech content to the playback video at the corresponding time based on the speech content and time. Therefore, when there is a need to understand information about certain participants during the meeting, it is no longer necessary to spend time browsing through the lengthy records of the entire meeting. Instead, a playback video for the corresponding participant is directly generated, and the playback video includes audio of the participant's speech content. This allows for a more focused attention on the participant's speech, the state changes of the selected persona, etc., facilitating a deeper understanding and analysis of the meeting content and significantly saving time.

[0199] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0200] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0201] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0202] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0203] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for replaying a meeting, characterized in that, The method includes: Obtain the playback parameters carried in the received playback request; wherein, the playback parameters include the identifier of at least one participant; For each participant identified by at least one identifier, obtain the model identifier of the selected character model, each operation record, and the occurrence time of the character model selected by the participant identified by that identifier; for each operation record, based on the state change information and occurrence time contained in the operation record, add video content of the character model with the model identifier undergoing the corresponding state change at the occurrence time in the playback video; if there is speech content saved for the participant's identifier, add the audio of the speech content at the corresponding time in the playback video based on the speech content and the speech time.

2. The method according to claim 1, characterized in that, The method further includes: If the playback parameters include a connection identifier and the corresponding connection time, then the connection content saved for the connection identifier is obtained; Generate the video corresponding to the connected content; The video is added at the transition time specified in the playback video.

3. The method according to claim 2, characterized in that, The generation of the video corresponding to the connecting content includes: If the connecting content is text, a preset number of video frames containing the text will be generated.

4. The method according to claim 1, characterized in that, The step of adding audio of the speech content to the playback video at the corresponding time based on the speech content and speech time includes: If the playback parameters include the waveform and / or pitch frequency of the timbre selected for the identifier of the at least one participant, the timbre and / or pitch of the audio containing the speech content is adjusted to the timbre of the waveform and / or the pitch of the frequency, and the adjusted audio is added to the playback video at the time corresponding to the speech time.

5. The method according to claim 1, characterized in that, The method further includes: If the playback parameters include key information, then obtain the key information; A preset marking operation is performed on a preset area of ​​the video frame containing the key information in the playback video.

6. The method according to claim 1, characterized in that, The method further includes: If a summary generation request is received, the playback audio in the playback video is converted to text to obtain the target text. The target text and preset prompts for generating the summary are then input into the large model to obtain the target summary output by the large model.

7. The method according to claim 6, characterized in that, The step of inputting the target text into a large model and obtaining the target summary output by the large model includes: If the target text exceeds a preset length, the target text is divided into a preset number of sub-texts. Each sub-text is input into the large model to obtain the sub-summary output by the large model. The sub-summaries are then concatenated according to the order of each sub-text in the target text to obtain the target summary.

8. The method according to claim 1, characterized in that, The method further includes: If a meeting setup instruction carrying a model identifier and a spatial scene identifier is received, then the metaverse spatial scene corresponding to the spatial scene identifier and the character model corresponding to the model identifier are constructed in the UI page. The position corresponding to the model identifier in the meeting setup instruction is obtained, the character model is added to the position in the metaverse spatial scene, and the modeling data of the metaverse spatial scene with the added character model is synchronized to the devices of each participant. This allows each participant's device to render and display the metaverse spatial scene containing the character model based on the modeling data.

9. The method according to claim 8, characterized in that, The method further includes: If a status change information is received from any participant's device, including behavior change information, location change information, morphological change information, and mood change information, the status of the persona model corresponding to that participant's device is adjusted based on the status change information.

10. The method according to claim 8, characterized in that, The method further includes: If audio or video data is received from any participant's device, the audio or video data is preprocessed and then synchronized to each participant's device.

11. A conference playback device, characterized in that, The device includes: The acquisition module is used to acquire the playback parameters carried in the received playback request; wherein, the playback parameters include the identifier of at least one participant; The processing module is configured to, for the at least one identified participant, obtain the model identifier of the selected character model, each operation record, and the occurrence time of the character model, which is stored for the identified participant; for each operation record, based on the state change information and occurrence time contained in the operation record, add video content of the character model with the model identifier undergoing the corresponding state change at the occurrence time in the playback video; if there is speech content stored for the identified participant, add the audio of the speech content to the playback video at the corresponding time based on the speech content and the speech time.

12. The apparatus according to claim 11, characterized in that, The processing module is further configured to, if the playback parameters include a connection identifier and a corresponding connection time, obtain the connection content saved for the connection identifier; generate the video corresponding to the connection content; and add the video at the connection time of the playback video.

13. The apparatus according to claim 12, characterized in that, The processing module is specifically used to generate a preset number of video frames containing the text if the connecting content is text.

14. The apparatus according to claim 11, characterized in that, The processing module is specifically configured to, if the playback parameters include the frequency of the waveform and / or pitch of the timbre selected for the identifier of the at least one participant, adjust the timbre and / or pitch of the audio containing the speech content to the timbre of the waveform and / or the pitch of the frequency, and add the adjusted audio to the time corresponding to the speech time in the playback video.

15. The apparatus according to claim 11, characterized in that, The processing module is further configured to acquire the key information if the playback parameters include key information, and to perform a preset marking operation on a preset area of ​​the video frame containing the key information in the playback video.

16. The apparatus according to claim 11, characterized in that, The processing module is further configured to, upon receiving a summary generation request, convert the playback audio in the playback video to text to obtain target text, input the target text and preset prompt words for prompting summary generation into the large model, and obtain the target summary output by the large model.

17. The apparatus according to claim 16, characterized in that, The processing module is specifically used to, if the target text exceeds a preset length, divide the target text into a preset number of sub-texts, input each sub-text into a large model, obtain the sub-summary output by the large model, and concatenate the sub-summaries according to the order of each sub-text in the target text to obtain the target summary.

18. The apparatus according to claim 11, characterized in that, The processing module is further configured to, upon receiving a meeting construction instruction carrying a model identifier and a spatial scene identifier, construct a metaverse spatial scene corresponding to the spatial scene identifier and a character model corresponding to the model identifier in the UI page, obtain the position corresponding to the model identifier in the meeting construction instruction, add the character model to the position in the metaverse spatial scene, and synchronize the modeling data of the metaverse spatial scene with the added character model to the devices of each participant; so that each participant's device renders and displays the metaverse spatial scene containing the character model based on the modeling data.

19. The apparatus according to claim 18, characterized in that, The processing module is also configured to adjust the state of the person model corresponding to the participant's device based on the state change information received from the device of any participant, the state change information including behavior change information, location change information, morphological change information and mood change information.

20. The apparatus according to claim 18, characterized in that, The processing module is also configured to preprocess the audio and video data received from any participant's device and synchronize the preprocessed audio and video data to each participant's device.

21. An electronic device, characterized in that, The electronic device includes a processor that executes a computer program stored in a memory to implement the steps of the conference playback method as described in any one of claims 1-10.

22. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the conference playback method as described in any one of claims 1-10.

23. A computer program product, characterized in that, The computer program product includes: computer program code, which, when run on a computer, causes the computer to perform the steps of the conference playback method as described in any one of claims 1-10.

Citation Information

Patent Citations

  • 3D display method for multi-person conference record playback, storage medium and terminal equipment

    CN112312062A

  • Metacosm conference hosting method, apparatus and device, storage medium and program product

    CN118200299A