Multimedia conference data processing method based on artificial intelligence model and related device
By using artificial intelligence models to filter and extract images of screen-shared content in multimedia conferences, generating descriptive information, and combining it with conference text, the problem of identifying important blocks of screen-shared content in multimedia conferences has been solved, thus enriching and enhancing the intuitiveness of conference minutes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-03-27
AI Technical Summary
In multimedia conferences, how can we effectively identify and utilize important segments of screen-shared content to enhance the richness and comprehensiveness of meeting minutes?
By acquiring image and text data from multimedia conferences, artificial intelligence models are used to filter images and extract sub-images, generating descriptive information. This information is then combined with the conference text to generate conference minutes, with added image content to enhance the relevance and intuitiveness of the minutes.
It enables efficient identification and utilization of screen-shared content, generating more comprehensive and intuitive meeting minutes, and improving the information utilization rate of multimedia conferences.
Smart Images

Figure CN120897097B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One or more embodiments of the present disclosure relate to an artificial intelligence model-based multimedia conference data processing method, an artificial intelligence model-based multimedia conference data processing apparatus, an electronic device, and a computer-readable storage medium. BACKGROUND
[0002] In a multimedia conference, a participant can share a screen, and explain in combination with screen sharing content, which can carry a large amount of information. SUMMARY
[0003] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter's scope.
[0004] At least one embodiment of the present disclosure provides an artificial intelligence model-based multimedia conference data processing method, comprising: acquiring a plurality of first images of a first multimedia conference, and acquiring conference text of the first multimedia conference, wherein the plurality of first images are derived from screen sharing content of the first multimedia conference; screening the plurality of first images according to a correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image; performing sub-image extraction on the at least one second image to obtain a plurality of sub-images, wherein each sub-image in the plurality of sub-images includes non-text data carrying an information amount satisfying a set condition; generating description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establishing an association relationship between the plurality of sub-images and the corresponding description information; analyzing the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by using a first artificial intelligence model, to determine a source text for generating a conference minutes of the first multimedia conference, and to generate the conference minutes of the first multimedia conference; in response to the source text including first description information, determining a first sub-image according to the association relationship, and adding the first sub-image in the conference minutes of the first multimedia conference, wherein the description information corresponding to the plurality of sub-images includes the first description information, and the plurality of sub-images include the first sub-image.
[0005] At least another embodiment of the present disclosure provides an artificial intelligence model-based multimedia conference data processing apparatus, comprising: an acquisition module configured to acquire a plurality of first images of a first multimedia conference, and acquire conference text of the first multimedia conference, wherein the plurality of first images are derived from screen sharing content of the first multimedia conference; a screening module configured to screen the plurality of first images according to a correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image; an extraction module configured to perform sub-image extraction on the at least one second image to obtain a plurality of sub-images, wherein each of the plurality of sub-images comprises non-text data carrying an information amount satisfying a set condition; a generation module configured to generate description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establish an association relationship between the plurality of sub-images and the corresponding description information; the generation module is further configured to analyze the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by using a first artificial intelligence model, to determine a source text used to generate a conference minutes of the first multimedia conference, and to generate the conference minutes of the first multimedia conference; in response to the source text comprising first description information, determining a first sub-image according to the association relationship, and adding the first sub-image in the conference minutes of the first multimedia conference, wherein the description information corresponding to the plurality of sub-images comprises the first description information, and the plurality of sub-images comprises the first sub-image.
[0006] At least another embodiment of the present disclosure provides an electronic device, comprising: a processing apparatus; and a storage apparatus comprising one or more computer program instructions; wherein the one or more computer program instructions are executed by the processing apparatus to perform the artificial intelligence model-based multimedia conference data processing method provided by at least one embodiment of the present disclosure.
[0007] At least another embodiment of the present disclosure provides a computer-readable storage medium, which non-transitorily stores computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the artificial intelligence model-based multimedia conference data processing method provided by at least one embodiment of the present disclosure is implemented.
[0008] At least one embodiment of the present disclosure provides a computer program product comprising a computer program, which when executed by a processor implements the artificial intelligence model-based multimedia conference data processing method provided by at least one embodiment of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above-described and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent as various embodiments of the present disclosure are described in conjunction with the following detailed description, which, together with the drawings, illustrates specific embodiments in which the principles of the present disclosure can be implemented. Throughout the drawings, like reference numerals will be understood to refer to like elements, features, and structures. It will be understood that the drawings are schematic and elements and features not be necessarily to scale.
[0010] Figure 1 A schematic diagram of an application scenario of a multimedia conference system provided by at least one embodiment of the present disclosure is shown;
[0011] Figure 2 A schematic diagram of a flow of a multimedia conference data processing method based on an artificial intelligence model provided by at least one embodiment of the present disclosure is shown;
[0012] Figure 3 A schematic diagram of a sub-image provided by at least one embodiment of the present disclosure is shown;
[0013] Figure 4 A schematic diagram of a conference summary provided by at least one embodiment of the present disclosure is shown;
[0014] Figure 5 A schematic diagram of a structure of a multimedia conference data processing apparatus based on an artificial intelligence model provided by at least one embodiment of the present disclosure is shown; and
[0015] Figure 6 A schematic diagram of a structure of an electronic device suitable for implementing an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0016] One or more embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It is understood that the drawings of the present disclosure and the embodiments are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0017] It should be understood that the various steps of the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0018] As used herein, the term "includes" and its variants are to be read to be analogous to "comprises," "comprising," "includes," "including," and "has," "having," "contains" or "containing." The term "based on" is to be interpreted as "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related terms have analogous meanings.
[0019] It should be noted that the terms "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] It should be noted that the terms "one", "multiple" mentioned in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.
[0021] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0022] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0023] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the information involved in the present disclosure should be informed to the relevant user and the authorization of the relevant user should be obtained through appropriate means, wherein the relevant user can include any type of right subject, such as an individual, an enterprise or a group.
[0024] For example, in response to receiving the active request of the user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require the information of the relevant user to be obtained and used, so that the relevant user can voluntarily choose whether to provide the information to the software or hardware such as electronic device, application program, server or storage medium performing the technical solutions of any embodiment of the present disclosure according to the prompt information.
[0025] As an optional but not limited implementation manner, in response to receiving the active request of the relevant user, the prompt information can be sent to the relevant user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.
[0026] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0027] With the rapid development of Internet technology, meetings gradually change from face-to-face offline meetings to more flexible online meetings (e.g., multimedia conferences). In a multimedia conference, participants can access the conference through the Internet (e.g., a multimedia conference system).
[0028] In a multimedia conference, participants can share screens, for example, some multimedia conference systems provide a cloud document screen projection function, and participants can use the cloud document screen projection function to share cloud documents during a multimedia conference, in which case the screen sharing content can include cloud documents; for another example, some multimedia conference systems provide a screen sharing function, and participants can use the screen sharing function to share screens during a multimedia conference, in which case the screen sharing content can include the content displayed on the shared screen.
[0029] The screen sharing content in a multimedia conference can carry a large amount of information, and the screen sharing content can be related to the conference content of the multimedia conference. How to identify important blocks in the screen sharing content becomes a problem to be solved.
[0030] To at least partially solve the at least one technical problem, at least one embodiment of the present disclosure provides a multimedia conference data processing method based on an artificial intelligence model, the method comprising: obtaining a plurality of first images of a first multimedia conference, and obtaining conference text of the first multimedia conference, the plurality of first images being derived from screen sharing content of the first multimedia conference, filtering the plurality of first images according to a degree of correlation between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image, performing sub-image extraction on the at least one second image to obtain a plurality of sub-images, each of the plurality of sub-images comprising: non-text data carrying an amount of information satisfying a set condition, generating description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establishing an association relationship between the plurality of sub-images and the corresponding description information, analyzing the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images using a first artificial intelligence model to determine source text for generating a conference summary of the first multimedia conference, and generating the conference summary of the first multimedia conference, in response to the source text including first description information, determining a first sub-image according to the association relationship, and adding the first sub-image to the conference summary of the first multimedia conference, the description information corresponding to the plurality of sub-images includes the first description information, and the plurality of sub-images includes the first sub-image.
[0031] In at least one embodiment of the present disclosure, a multimedia conference data processing method based on an artificial intelligence model is also provided. In a multimedia conference scenario, for a plurality of frame images (i.e., first images) of screen sharing content, a second image related to text content is retained by screening through the correlation between image content and conference text content, a sub-image carrying rich information of non-text data is extracted from the second image, and description information of the sub-image is generated. In this way, high-quality sub-images appearing in the screen sharing content are identified, the image information of the high-quality sub-images is understood, the rich information carried by the high-quality sub-images is converted into description information in the form of text, the utilization rate of the screen sharing content in the multimedia conference is improved, the information contained in the screen sharing content is more comprehensively and completely extracted, and in this way, the high-quality sub-images in the screen sharing content are used to generate a conference summary together, the screen sharing content in the form of images is added to the conference summary, and the correlation between the conference summary and the multimedia conference is improved.
[0032] Based on the multimedia conference data processing method based on the artificial intelligence model provided in at least one embodiment of the present disclosure, at least one embodiment of the present disclosure also provides a multimedia conference data processing apparatus based on an artificial intelligence model, an electronic device, a computer readable storage medium, and a computer program product.
[0033] One or more embodiments of the present disclosure and some examples thereof will be described in detail below with reference to the accompanying drawings.
[0034] Figure 1 An application scenario diagram of a multimedia conference system provided in at least one embodiment of the present disclosure is schematically shown.
[0035] As Figure 1 shown, the application scenario of this embodiment includes a multimedia conference system 100, which can provide services related to multimedia conferences. For example, the multimedia conference system 100 can support starting a multimedia conference (e.g., a video conference), screen sharing during the multimedia conference, generating a conference summary after the multimedia conference ends, etc.
[0036] One or more embodiments of the present disclosure do not limit the form of the multimedia conference system 100. In some embodiments, the multimedia conference system 100 can be an independent software system. For example, the multimedia conference system 100 can be an application (application, APP) providing services related to multimedia conferences, a cloud service, a plug-in, middleware, etc. In other embodiments, the multimedia conference system 100 can also be integrated into other software systems. For example, the multimedia conference system 100 can be integrated into an office collaboration system as a functional module of the office collaboration system, and provide services related to multimedia conferences in the office collaboration system.
[0037] The multimedia conference system 100 can identify and understand the content of the important block in the screen sharing content in a multimedia conference scenario. For example, the multimedia conference system 100 can obtain a plurality of first images 101 of a first multimedia conference and conference text 102 of the multimedia conference, the plurality of first images 101 are derived from the screen sharing content of the first multimedia conference, for example, are frame images of the screen sharing content of the first multimedia conference, and the plurality of first images 101 are screened according to the correlation between the plurality of first images 101 and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image 103, and the image screening is performed based on the image-text correlation.
[0038] The multimedia conference system 100 can perform sub-image extraction on the at least one second image 103 to obtain a plurality of sub-images 104, each of the plurality of sub-images 104 includes non-text data carrying an information amount satisfying a set condition, and an important block in the screen sharing content is identified.
[0039] Then, the multimedia conference system 100 can generate description information 105 corresponding to the plurality of sub-images according to the conference text 102 of the first multimedia conference, understand the important block in the screen sharing content, and convert the non-text data of the important block in the screen sharing content into description information in the form of text.
[0040] In this way, in the multimedia conference scenario, for the screen sharing content, a high-quality sub-image containing a large amount of information is identified in the screen sharing content, description information of the high-quality sub-image is generated, image understanding of the high-quality sub-image which is difficult to understand and non-text data is realized, and the image information of the high-quality sub-image is represented by using the description information.
[0041] Further, the multimedia conference system 100 can also generate a conference summary 106 of the first multimedia conference. For example, the multimedia conference system 100 can analyze the conference text 102 of the first multimedia conference and the description information 105 corresponding to the plurality of sub-images by using a first artificial intelligence model, determine a source text for generating the conference summary of the first multimedia conference, and when the source text includes first description information in the description information 105 corresponding to the plurality of sub-images, add a first sub-image corresponding to the first description information in the conference summary 106 of the first multimedia conference. In this way, on the one hand, when the conference summary is generated, the conference text and the description information of the screen sharing content are combined to enrich the content and comprehensiveness of the conference summary; on the other hand, when the conference summary is related to certain screen sharing content, the high-quality sub-image in the form of image is added in the conference summary, which is helpful to more intuitively understand the conference content of the multimedia conference.
[0042] The following will be described in combination with Figures 2 to 4A multimedia conference data processing method based on an artificial intelligence model is provided in at least one embodiment of the present disclosure.
[0043] Figure 2 A flowchart of a multimedia conference data processing method based on an artificial intelligence model is schematically shown.
[0044] As Figure 2 shown, the multimedia conference data processing method based on an artificial intelligence model of this embodiment includes steps S201-S206. In some embodiments, the execution subject of the multimedia conference data processing method based on an artificial intelligence model can be an electronic device deployed with a client, or an electronic device deployed with a server, or any electronic device connected to the client and the server, and one or more embodiments of the present disclosure do not limit this. The multimedia conference data processing method based on an artificial intelligence model includes:
[0045] Step S201: Obtain a plurality of first images of a first multimedia conference, and obtain conference text of the first multimedia conference.
[0046] In one or more embodiments of the present disclosure, the first multimedia conference can be understood as any multimedia conference with screen sharing, and the plurality of first images are derived from the screen sharing content of the first multimedia conference. That is, during the first multimedia conference, a participant performs screen sharing, and the first multimedia conference has corresponding screen sharing content, and the screen sharing content includes the plurality of first images.
[0047] In some possible implementations, the plurality of first images of the first multimedia conference are obtained by frame extraction. For example, a conference recording file of the first multimedia conference is obtained, and frame extraction processing is performed on the conference recording file to obtain the plurality of first images.
[0048] The conference recording file includes the screen sharing content of the first multimedia conference, for example, the conference recording file can be a video stream formed after recording the first multimedia conference.
[0049] By performing frame extraction processing on the conference recording file, for example, by performing frame extraction processing on the conference recording file at a fixed frame extraction interval, the conference recording file in the form of a video is converted into the plurality of first images in the form of images, and a plurality of video frames of the first multimedia conference are obtained. Since the conference recording file includes the screen sharing content of the first multimedia conference, the plurality of first images also include the screen sharing content of the first multimedia conference, for example, the plurality of first images can include screen sharing content corresponding to a plurality of frame extraction moments in the first multimedia conference.
[0050] The conference text of the first multimedia conference can be understood as corresponding speech text in the first multimedia conference process, for example, by performing automatic speech recognition (ASR) on the first multimedia conference, corresponding speech text at each moment in the first multimedia conference is obtained. The conference text record of the first multimedia conference records the speech content of each participant in the first multimedia conference, for example, the content of the participants in the first multimedia conference explaining and explaining the screen sharing content.
[0051] Step S202: According to the correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, the plurality of first images are screened to obtain at least one second image.
[0052] For each of the plurality of first images, the conference text corresponding to the first image is determined, the correlation degree between the first image and the conference text corresponding to the first image is determined, and the first image is screened.
[0053] Since the first image is derived from the screen sharing content in the first multimedia conference, the first image can be understood as the screen sharing content at a moment in the first multimedia conference, and the conference text corresponding to the first image can be understood as the conference text at the moment.
[0054] The correlation degree between the first image and the conference text corresponding to the first image can be understood as the degree of correlation between the image information of the first image and the text information of the conference text corresponding to the first image, in other words, the correlation degree between the first image and the conference text corresponding to the first image can be used to measure whether the conference text corresponding to the first image is related to the first image, for example, whether the conference text corresponding to the first image is used to describe the first image.
[0055] In some possible implementation manners, for each of the plurality of first images, the following steps are performed: determining the timestamp information of the first image, determining the conference text corresponding to the first image according to the timestamp information of the first image, determining the processing mode of the first image according to the correlation degree between the first image and the conference text corresponding to the first image, and processing the first image according to the processing mode of the first image.
[0056] The timestamp information of the first image can be understood as the occurrence moment of the first image in the first multimedia conference, for example, the first image is the screen sharing content at the Xth second in the first multimedia conference, and the timestamp information of the first image can be the Xth second, X is a natural number greater than 0.
[0057] The conference text corresponding to the first image can be conference text in the first multimedia conference, and the timestamp information of the first image corresponds to conference text in a set time range. The set time range can be Y seconds before and after the timestamp information. For example, the timestamp information of the first image is X seconds, and the conference text corresponding to the first image can be conference text from X-Y seconds to X+Y seconds in the conference text of the first multimedia conference.
[0058] By determining the correlation between the first image and the conference text corresponding to the first image, the image-text correlation of the first image and the conference text corresponding to the first image is determined, and then the processing mode of the first image is determined. The processing mode can include discarding the first image or retaining the first image.
[0059] In this way, the first image is processed according to the processing mode of the first image, and the first image is retained or discarded, realizing the screening and filtering of multiple first images, for example, retaining the first image with image-text correlation and discarding the first image with image-text irrelevance.
[0060] In this way, for the first image and the conference text corresponding to the first image, it is considered that the conference text corresponding to the first image is not explained or discussed for the first image, and the conference text corresponding to the first image is difficult to provide information related to the image information of the first image, and the conference text corresponding to the first image has low reference value for the image information of the first image. Therefore, the above-mentioned first image can be understood as a useless video frame, and the above-mentioned first image is discarded, saving computing resources.
[0061] In some embodiments, the first image and the conference text corresponding to the first image are sent to a second artificial intelligence model, and the correlation degree information returned by the second artificial intelligence model is received. In response to the correlation degree information representing that the correlation between the first image and the conference text corresponding to the first image meets the correlation condition, it is determined that the processing mode of the first image is to retain the first image. Or in response to the correlation degree information representing that the correlation between the first image and the conference text corresponding to the first image does not meet the correlation condition, it is determined that the processing mode of the first image is to discard the first image.
[0062] That is, in at least one embodiment of the present disclosure, the second artificial intelligence model is used to determine the correlation between the first image and the conference text corresponding to the first image. The correlation degree information returned by the second artificial intelligence model can be used to represent the correlation between the first image and the conference text corresponding to the first image. For example, the correlation degree information can be a value between 0 and 1. The correlation degree is 1, indicating that the first image and the conference text corresponding to the first image are related. The correlation degree is 0, indicating that the first image and the conference text corresponding to the first image are not related.
[0063] The relevant condition can be understood as a condition for determining whether to retain the first image. For example, the relevant condition can be that the degree of relevance between the first image and the conference text corresponding to the first image is greater than a degree of relevance threshold.
[0064] The second artificial intelligence model can be an artificial intelligence model with a document-image relevance determination capability. For example, the second artificial intelligence model can be a classification model, and the second artificial intelligence model can be trained using training images with labeled relevance information.
[0065] Through the document-image relevance determination capability of the second artificial intelligence model, each of the plurality of first images is screened, and only the first image with document-image relevance is retained. The first image is screened to obtain the second image.
[0066] Step S203: Sub-image extraction is performed on the at least one second image to obtain a plurality of sub-images.
[0067] In one or more embodiments of the present disclosure, sub-image extraction can be understood as an operation of extracting a partial image from a second image. Each of the plurality of sub-images can include non-text data carrying an amount of information satisfying a set condition.
[0068] That is, the sub-image satisfies two conditions: carrying an amount of information satisfying a set condition and non-text data, for example, the amount of information is greater than an information amount threshold. In this way, non-text data containing rich information is identified and extracted from the second image.
[0069] In one or more embodiments of the present disclosure, it is considered that relevant information of text data in the second image can be obtained through text processing such as character recognition, but it is difficult to directly obtain relevant information of non-text data in the second image through image processing. Therefore, an important block (i.e., a sub-image) containing rich information, including non-text data, is extracted from the second image.
[0070] One or more embodiments of the present disclosure do not limit the type of sub-image. For example, the type of sub-image can include at least one of the following: picture data, table data, video data, presentation data, or code block data.
[0071] In some possible implementations, a sub-image in a second image is identified using an artificial intelligence model. For example, a third artificial intelligence model is used to perform sub-image extraction on the at least one second image, and position information of a plurality of sub-images returned by the third artificial intelligence model is received. The position information can indicate the position of the sub-image in the at least one second image. According to the position information of the plurality of sub-images, the plurality of sub-images are determined from the at least one second image.
[0072] The third artificial intelligence model can be an artificial intelligence model with sub-image recognition capability. For example, the third artificial intelligence model can be a multi-modal model. The third artificial intelligence model can be a model constructed based on a transformer architecture, a model constructed based on a recurrent neural network, a model constructed based on an attention mechanism, or the like. Alternatively, the third artificial intelligence model can be a model improved based on a transformer architecture, such as a mixture of experts (MoE) model.
[0073] The third artificial intelligence model can be an existing, open-source general artificial intelligence model. Alternatively, the third artificial intelligence model can be an artificial intelligence model obtained by fine-tuning (e.g., full-parameter fine-tuning or partial-parameter fine-tuning) a pre-trained model using the training images and the position information of the sub-images in the training images as training data.
[0074] The position information of the sub-images is recognized from each second image by the image recognition capability of the third artificial intelligence model. For example, the position information can be coordinate information of an image frame of the sub-image. The position information is used to determine the sub-image carrying high information quantity and non-text data in each second image.
[0075] The third artificial intelligence model can extract the sub-image in the second image based on a prompt learning technique. For example, for each of the at least one second image, the following steps are performed: generating a first prompt word, sending the first prompt word to the third artificial intelligence model, and receiving the position information returned by the third artificial intelligence model.
[0076] A prompt word (prompt) can be used to guide an artificial intelligence model to perform a specific output in a generative task. By configuring the prompt word, the artificial intelligence model can understand the background and requirements of the task, and the artificial intelligence model can process different types of processing tasks without the need for retraining the artificial intelligence model, thereby increasing the scalability and flexibility of the artificial intelligence model.
[0077] The first prompt word can include: a second image, extraction criteria for sub-image extraction, and prompt information for indicating that the second image is subjected to sub-image extraction based on the extraction criteria for sub-image extraction. For example, the extraction criteria for sub-image extraction can be the type of sub-image.
[0078] In some embodiments, the first prompt word can be:
[0079] “# Character
[0080] Grounding expert;
[0081] ## Target
[0082] You will receive an image >. Your task is to identify the high-quality figures, tables, videos, PPTs, or code blocks in the image and represent the results in <bbox>x1 y1 x2 y2< / bbox> ;
[0083] ## Limitations
[0084] - The definition of high quality refers to clear and visible content with a high amount of information;
[0085] - If the image or table in > is incomplete, such as only half of the image or half of the table, it does not need to be boxed;
[0086] - There are a total of five labels that need to be identified, namely "image", "table", "video", "code block", and "PPT", and the labels and coordinates are separated by \t;
[0087] - The image in the picture, the conference room portrait, the background picture, and other parts do not belong to the information amount and do not need to be boxed, and the focus is on the relevant knowledge content in the meeting;
[0088] - For PPT type, only the complete page of PPT content needs to be boxed, and the navigation box, thumbnail, directory box, and other irrelevant information of the PPT software do not need to be boxed;
[0089] - This time, the focus is on the high-quality content in the images of documents, PDFs, and PPTs, and some software screenshots do not need to be boxed;
[0090] - The final grounding box result is represented in <bbox>x1 y1 x2 y2< / bbox> , and the coordinate range is (0, 999);
[0091] - If there are multiple boxes, output multiple lines of results, separated by \n;
[0092] - If there is no box that meets the requirements, you can directly output \n;
[0093] ## Input
[0094] * Image content:
[0095] {image}
[0096] ## Output format
[0097] Image\t <bbox>x1 y1 x2 y2< / bbox> ".
[0098] By inputting the first prompt word into the third artificial intelligence model, the third artificial intelligence model can identify the sub-image in the second image by virtue of the prompting capability of the first prompt word, output the position information of the sub-image in the second image, and realize image grounding.
[0099] Figure 3 A schematic diagram of a sub-image is schematically shown.
[0100] As Figure 3 shown, the second image 300 is the screen sharing content at a certain moment in the first multimedia conference, and in the second image 300, text data 301 and non-text data are included, and the non-text data includes picture data 302, table data 303 and code block data 304.
[0101] The third artificial intelligence model can output the position information of the picture data 302, the table data 303 and the code block data 304 by performing sub-image extraction on the second image 300 by using the third artificial intelligence model, and the sub-image in the second image 300 is extracted by using the position information of the picture data 302, the table data 303 and the code block data 304.
[0102] Step S204: According to the conference text of the first multimedia conference, description information corresponding to a plurality of sub-images is generated, and an association relationship between the plurality of sub-images and the corresponding description information is established.
[0103] Since the second image where the sub-image is located and the conference text corresponding to the second image have a correlation, image understanding is performed according to the conference text of the first multimedia conference, and description information (which can also be called image caption) corresponding to a plurality of sub-images is generated.
[0104] In some possible implementation manners, for each sub-image in the plurality of sub-images, the following steps are performed: determining the conference text corresponding to the second image where the sub-image is located, and determining the position information of the sub-image in the second image, and using the fourth artificial intelligence model, based on the conference text corresponding to the second image where the sub-image is located and the position information of the sub-image in the second image, generating description information corresponding to the sub-image.
[0105] That is, the description information of the sub-image is generated by using the fourth artificial intelligence model, the fourth artificial intelligence model is informed of which sub-image to generate the description information by using the position information of the sub-image in the second image, and the fourth artificial intelligence model is informed of the information related to the image information of the sub-image, such as the information for explaining and discussing the sub-image in the first multimedia conference, by using the conference text corresponding to the second image where the sub-image is located.
[0106] The fourth artificial intelligence model can be an artificial intelligence model with a description information generation capability. Similar to the third artificial intelligence model, the fourth artificial intelligence model can be a multi-modal model. The fourth artificial intelligence model can be a model constructed based on a transformer architecture, a model constructed based on a recurrent neural network, a model constructed based on an attention mechanism, or the like. Alternatively, the fourth artificial intelligence model can also be a model improved based on a transformer architecture, such as a mixture of experts (MoE) model.
[0107] The fourth artificial intelligence model can be an existing, open-source general artificial intelligence model. Alternatively, the fourth artificial intelligence model can also be an artificial intelligence model obtained by fine-tuning (e.g., full-parameter fine-tuning or partial-parameter fine-tuning) a pre-trained model using the training images, the position information of the sub-images in the training images, and the description information of the sub-images in the training images as training data.
[0108] In one or more embodiments of the present disclosure, the third artificial intelligence model and the fourth artificial intelligence model can be the same artificial intelligence model, or the third artificial intelligence model and the fourth artificial intelligence model can also be different artificial intelligence models.
[0109] In this way, the conference text corresponding to the second image in which the sub-image is located is taken as a context, and the context is used for image understanding to generate the description information of the sub-image, so as to convert the image information of the sub-image into description information in a text form. Since the description information in a text form is easier to process and has a higher processing accuracy, the description information of the sub-image has a stronger applicable range.
[0110] After the description information corresponding to the plurality of sub-images is generated, an association relationship between the sub-images and the description information can also be established. For example, each sub-image and the corresponding description information can be associated and stored. In this way, when the sub-image or the description information of the sub-image needs to be used subsequently, the description information corresponding to the sub-image or the sub-image corresponding to the description information can be determined based on the association relationship.
[0111] Step S205: analyzing the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by using the first artificial intelligence model, determining source text used for generating a conference summary of the first multimedia conference, and generating the conference summary of the first multimedia conference.
[0112] In one or more embodiments of the present disclosure, the description information corresponding to the plurality of sub-images can be used to generate a conference summary. For example, a conference summary of the first multimedia conference including at least one sub-image can be generated according to the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images.
[0113] That is, in the multimedia conference scenario, unlike the traditional way of generating a conference summary only by using conference text, in one or more embodiments of the present disclosure, a conference summary is generated in combination with conference text and description information corresponding to multiple sub-images.
[0114] Figure 4 A schematic diagram of a conference summary provided by at least one embodiment of the present disclosure is schematically shown.
[0115] As Figure 4 shown, in the conference summary 400, in addition to including text content, sub-images 401 (including sub-image A, sub-image B, and sub-image C) are additionally added.
[0116] In this way, on the one hand, since the description information corresponding to the multiple sub-images and the conference text are both in text form, when generating the conference summary, no additional processing capability of multi-modal information is needed, and only text processing is needed for the conference text and the description information corresponding to the multiple sub-images; on the other hand, the screen sharing content in the multimedia conference is added to generate the conference summary, which improves the utilization rate of the screen sharing content, and at the same time makes the conference summary include more comprehensive and rich conference content; on the other hand, the sub-images are added in the conference summary, achieving a picture-and-text effect in the conference summary, and a user viewing the conference summary can intuitively understand the conference content of the multimedia conference through the sub-images.
[0117] The first artificial intelligence model can be an artificial intelligence model having a conference summary generation capability, the first artificial intelligence model can be a language model, the first artificial intelligence model can be a model constructed based on a transformer architecture, a model constructed based on a recurrent neural network, a model constructed based on an attention mechanism, etc., or the first artificial intelligence model can also be a model improved on the basis of a transformer architecture, such as a mixture of experts (MoE) model, etc.
[0118] The first artificial intelligence model can be an existing, open-source general artificial intelligence model, or the first artificial intelligence model can also be an artificial intelligence model obtained by fine-tuning (such as full-parameter fine-tuning or partial-parameter fine-tuning) a pre-trained model using historical conference text and historical conference summaries as training data.
[0119] In one or more embodiments of the present disclosure, the third artificial intelligence model and the first artificial intelligence model can be the same artificial intelligence model, or the third artificial intelligence model and the first artificial intelligence model can also be different artificial intelligence models.
[0120] The conference summary generation capability of the first artificial intelligence model is used to analyze the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images. First, the source text used to generate the conference summary of the first multimedia conference is determined, that is, it is determined which part of information in the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images is used to generate the conference summary of the first multimedia conference. Then, the conference summary of the first multimedia conference is generated by using the source text of the conference summary of the first multimedia conference.
[0121] Step S206: In response to the source text including the first description information, the first sub-image is determined according to the association relationship, and the first sub-image is added in the conference summary of the first multimedia conference.
[0122] The description information corresponding to the plurality of sub-images can include the first description information, and the plurality of sub-images can include the first sub-image.
[0123] For example, by establishing the association relationship between the sub-image and the description information of the sub-image, when the first artificial intelligence model generates the conference summary by using the first description information, the first sub-image corresponding to the first description information is found according to the association relationship.
[0124] In this way, for the source text used to generate the conference summary, if some description information is included in the source text, it indicates that the conference summary of the first multimedia conference is generated based on these description information, that is, the conference summary of the first multimedia conference is related to the first description information. In this case, the corresponding sub-image is added in the conference summary of the first multimedia conference, so that the conference summary includes image information associated with the content of the conference summary, instead of only including monotonous text content.
[0125] In some possible implementations, in the conference summary of the first multimedia conference, a first position associated with the first description information is determined, and the conference summary at the first position is generated based on the first description information. The first sub-image is added at the first position of the conference summary of the first multimedia conference.
[0126] That is, in the conference summary of the first multimedia conference, the first position where the conference summary generated based on the first description information is located is located, and the first sub-image is added at the first position, so that the first position of the conference summary of the first multimedia conference includes the text content related to the first description information and the image content corresponding to the first description information, and the internal association of the conference summary of the first multimedia conference is enhanced.
[0127] Based on the multimedia conference data processing method based on the artificial intelligence model provided in at least one embodiment of the present disclosure, at least one embodiment of the present disclosure also provides a multimedia conference data processing apparatus based on an artificial intelligence model. The following will be described in combination with Figure 5The artificial intelligence model-based multimedia conference data processing apparatus is described in detail.
[0128] Figure 5 A structural schematic diagram of an artificial intelligence model-based multimedia conference data processing apparatus is shown schematically.
[0129] As Figure 5 shown, the artificial intelligence model-based multimedia conference data processing apparatus 500 of this embodiment includes an acquisition module 501, a screening module 502, an extraction module 503, and a generation module 504. For example, these units or modules can be implemented by hardware (such as a circuit) module or a software module, and the following embodiments are the same as this, and will not be described here. For example, these units or modules can be implemented by a central processing unit (CPU), a general-purpose graphics processing unit (GPGPU), an image processor (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and corresponding computer instructions.
[0130] The acquisition module 501 is configured to acquire a plurality of first images of a first multimedia conference, and acquire conference text of the first multimedia conference, wherein the plurality of first images are derived from screen sharing content of the first multimedia conference. For example, the acquisition module 501 can be configured to perform the step S201 described above, and the specific implementation principle can be referred to the related description of the step S201, which will not be described here.
[0131] The screening module 502 is configured to screen the plurality of first images according to a correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image. For example, the screening module 502 can be configured to perform the step S202 described above, and the specific implementation principle can be referred to the related description of the step S202, which will not be described here.
[0132] The extraction module 503 is configured to perform sub-image extraction on the at least one second image to obtain a plurality of sub-images, wherein each sub-image in the plurality of sub-images includes non-text data carrying an information amount satisfying a set condition. For example, the extraction module 503 can be configured to perform the step S203 described above, and the specific implementation principle can be referred to the related description of the step S203, which will not be described here.
[0133] The generating module 504 is configured to generate description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establish an association relationship between the plurality of sub-images and the corresponding description information. For example, the generating module 504 can be configured to perform step S204 described above, and the specific implementation principle can be referred to the related description of step S204, which will not be repeated here.
[0134] The generating module 504 is further configured to analyze the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by using a first artificial intelligence model, determine source text for generating a conference summary of the first multimedia conference, and generate the conference summary of the first multimedia conference; in response to the source text including first description information, determine a first sub-image according to the association relationship, and add the first sub-image in the conference summary of the first multimedia conference, wherein the description information corresponding to the plurality of sub-images includes the first description information, and the plurality of sub-images includes the first sub-image. For example, the generating module 504 can be configured to perform steps S205 and S206 described above, and the specific implementation principle can be referred to the related description of steps S205 and S206, which will not be repeated here.
[0135] In at least one embodiment of the present disclosure, the generating module 504 is further configured to determine a first position associated with the first description information in the conference summary of the first multimedia conference, wherein the conference summary of the first position is generated based on the first description information; and add the first sub-image in the first position of the conference summary of the first multimedia conference.
[0136] In at least one embodiment of the present disclosure, the obtaining module 501 is further configured to obtain a conference recording file of the first multimedia conference, wherein the conference recording file includes screen sharing content of the first multimedia conference; and perform frame extraction processing on the conference recording file to obtain the plurality of first images.
[0137] In at least one embodiment of the present disclosure, the screening module 502 is further configured to, for each first image in the plurality of first images, perform the following steps: determine timestamp information of the first image; determine conference text corresponding to the first image according to the timestamp information of the first image; determine a processing manner of the first image according to a correlation degree between the first image and the conference text corresponding to the first image, wherein the processing manner includes discarding the first image or retaining the first image; and process the first image according to the processing manner of the first image.
[0138] In at least one embodiment of the present disclosure, the screening module 502 is further configured to: send the first image and the meeting text corresponding to the first image to a second artificial intelligence model, receive the relevance degree information returned by the second artificial intelligence model; in response to the relevance degree information representing that the relevance degree between the first image and the meeting text corresponding to the first image meets a relevance condition, determining that the processing manner of the first image is to retain the first image; or in response to the relevance degree information representing that the relevance degree between the first image and the meeting text corresponding to the first image does not meet the relevance condition, determining that the processing manner of the first image is to discard the first image.
[0139] In at least one embodiment of the present disclosure, the extraction module 503 is further configured to: perform sub-image extraction on the at least one second image by using a third artificial intelligence model, and receive position information of the plurality of sub-images returned by the third artificial intelligence model, wherein the position information indicates the position of the sub-image in the at least one second image; and determine the plurality of sub-images from the at least one second image according to the position information of the plurality of sub-images.
[0140] In at least one embodiment of the present disclosure, the extraction module 503 is further configured to: for each second image in the at least one second image, perform the following steps: generating a first prompt word, wherein the first prompt word includes: the second image, extraction criteria for sub-image extraction, and prompt information for indicating sub-image extraction on the second image based on the extraction criteria for sub-image extraction; sending the first prompt word to a third artificial intelligence model, and receiving position information returned by the third artificial intelligence model.
[0141] In at least one embodiment of the present disclosure, the generation module 504 is further configured to: for each sub-image in the plurality of sub-images, perform the following steps: determining the meeting text corresponding to the second image in which the sub-image is located, and determining the position information of the sub-image in the second image; and using a fourth artificial intelligence model, generating description information corresponding to the sub-image based on the meeting text corresponding to the second image in which the sub-image is located and the position information of the sub-image in the second image.
[0142] In at least one embodiment of the present disclosure, the type of the sub-image includes at least one of: picture data, table data, video data, presentation data, or code block data.
[0143] It should be noted that, for the purpose of clearness and conciseness, the entire constituent units of the multimedia conference data processing apparatus 500 based on the artificial intelligence model are not given in the embodiments of the present disclosure. To achieve the necessary functions of the multimedia conference data processing apparatus 500 based on the artificial intelligence model, a person skilled in the art can provide and set other constituent units not shown according to specific needs, and one or more embodiments of the present disclosure do not limit this.
[0144] The electronic device in at least one embodiment of the present disclosure includes a processing apparatus and a storage apparatus including one or more computer program modules. The one or more computer program modules are stored in the storage apparatus and configured to be executed by the processing apparatus. The one or more computer program modules are used to implement the multimedia conference data processing method based on the artificial intelligence model provided by any embodiment of the present disclosure.
[0145] For example, the processing apparatus can be a processor, such as a central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), a general purpose graphics processing unit (GPGPU), or other forms of processing units with data processing and / or instruction execution capabilities, which can be general purpose processors or special purpose processors, and can control other components in the electronic device to perform desired functions.
[0146] For example, the storage apparatus can be a memory, which can include one or more computer program products including various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM), cache memory, and / or the like. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer readable storage medium, and the processing apparatus can run the program instructions to implement the functions (implemented by the processing apparatus) in the embodiments of the present disclosure and / or other desired functions. Various application programs and various data can also be stored in the computer readable storage medium, and one or more embodiments of the present disclosure do not limit this.
[0147] Reference will be made to the following description Figure 6 which shows a structural schematic diagram of an electronic device (such as a terminal device or a server) 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. Figure 6The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0148] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0149] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0150] In particular, according to one or more embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, one or more embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of the embodiments of this disclosure.
[0151] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable medium or a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.
[0152] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0153] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device and can be accessed via the electronic device.
[0154] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a plurality of first images of a first multimedia conference, and acquire conference text of the first multimedia conference; filter the plurality of first images according to a correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image; perform sub-image extraction on the at least one second image to obtain a plurality of sub-images; generate description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establish an association relationship between the plurality of sub-images and the corresponding description information; analyze the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by using a first artificial intelligence model, determine a source text for generating a conference minutes of the first multimedia conference, and generate the conference minutes of the first multimedia conference; in response to the source text including first description information, determine a first sub-image according to the association relationship, and increase the first sub-image in the conference minutes of the first multimedia conference.
[0155] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0156] One or more embodiments of the present disclosure also provide a computer program product including one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the flow or function described according to any embodiment of the present disclosure is generated in whole or in part.
[0157] The computer instructions can be stored in or transferred from one computer-readable medium to another computer-readable medium, such as from one website, computer, or data center to another website, computer, or data center, through wired (such as coaxial cable, fiber optics, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means.
[0158] The computer program product, when executed by a computer, causes the computer to execute any one of the preceding methods for processing multimedia conference data based on an artificial intelligence model. The computer program product can be a software installation package, and when any one of the preceding methods for processing multimedia conference data based on an artificial intelligence model needs to be used, the computer program product can be downloaded and executed on the computer.
[0159] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the reverse order, depending on the functionality involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0160] The units or modules described in the embodiments of the present disclosure can be implemented by means of software, or can be implemented by means of hardware. In some cases, the name of the unit or module does not constitute a limitation on the unit or module itself.
[0161] The functions described above in the present disclosure can be performed at least in part by one or more hardware logic components. For example, non-limiting examples of exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0163] According to one or more embodiments of the present disclosure, example one provides a multimedia conference data processing method based on an artificial intelligence model, comprising:
[0164] obtaining a plurality of first images of a first multimedia conference, and obtaining conference text of the first multimedia conference, wherein the plurality of first images are derived from screen sharing content of the first multimedia conference;
[0165] According to the correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, the plurality of first images are screened to obtain at least one second image;
[0166] Sub-image extraction is performed on the at least one second image to obtain a plurality of sub-images, wherein each sub-image in the plurality of sub-images includes non-text data carrying an information amount satisfying a set condition;
[0167] According to the conference text of the first multimedia conference, description information corresponding to the plurality of sub-images is generated, and an association relationship between the plurality of sub-images and the corresponding description information is established;
[0168] The conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images are analyzed by using a first artificial intelligence model to determine source text for generating a conference summary of the first multimedia conference, and the conference summary of the first multimedia conference is generated;
[0169] In response to the source text including first description information, a first sub-image is determined according to the association relationship, and the first sub-image is added in the conference summary of the first multimedia conference, wherein the description information corresponding to the plurality of sub-images includes the first description information, and the plurality of sub-images includes the first sub-image.
[0170] According to one or more embodiments of the present disclosure, example two provides that, in example one, adding the first sub-image in the meeting minutes of the first multimedia conference, includes:
[0171] In the meeting minutes of the first multimedia conference, a first position associated with the first description information is determined, wherein the meeting minutes of the first position are generated based on the first description information;
[0172] In the first position of the meeting minutes of the first multimedia conference, the first sub-image is added.
[0173] According to one or more embodiments of the present disclosure, example three provides that, in example one, obtaining a plurality of first images of a first multimedia conference, includes:
[0174] Obtaining a conference recording file of the first multimedia conference, wherein the conference recording file includes screen sharing content of the first multimedia conference;
[0175] Frame extraction processing is performed on the conference recording file to obtain the plurality of first images.
[0176] According to one or more embodiments of the present disclosure, example four provides that, in example one, according to a correlation degree between the plurality of first images and conference texts corresponding to the plurality of first images respectively, the plurality of first images are screened to obtain at least one second image, includes:
[0177] For each of the plurality of first images, the following steps are performed:
[0178] Determining timestamp information of the first image;
[0179] According to the timestamp information of the first image, determining conference texts corresponding to the first image;
[0180] According to a correlation degree between the first image and the conference texts corresponding to the first image, determining a processing mode of the first image, wherein the processing mode includes discarding the first image or retaining the first image;
[0181] According to the processing mode of the first image, processing the first image.
[0182] According to one or more embodiments of the present disclosure, example five provides that, in example four, according to a correlation degree between the first image and conference texts corresponding to the first image, determining a processing mode of the first image, includes:
[0183] sending the first image and the conference text corresponding to the first image to a second artificial intelligence model, and receiving relevance degree information returned by the second artificial intelligence model;
[0184] in response to the relevance degree information representing that the relevance degree between the first image and the conference text corresponding to the first image meets a relevance condition, determining that the processing manner of the first image is to retain the first image; or
[0185] in response to the relevance degree information representing that the relevance degree between the first image and the conference text corresponding to the first image does not meet the relevance condition, determining that the processing manner of the first image is to discard the first image.
[0186] According to one or more embodiments of the present disclosure, example six provides that the sub-image extraction of the at least one second image in example one obtains a plurality of sub-images, including:
[0187] performing sub-image extraction on the at least one second image by using a third artificial intelligence model, and receiving position information of the plurality of sub-images returned by the third artificial intelligence model, wherein the position information indicates the position of the sub-image in the at least one second image;
[0188] According to the position information of the plurality of sub-images, the plurality of sub-images are determined from the at least one second image.
[0189] According to one or more embodiments of the present disclosure, example seven provides that the sub-image extraction of the at least one second image by using a third artificial intelligence model in example six receives position information of the plurality of sub-images returned by the third artificial intelligence model, including:
[0190] for each second image in the at least one second image, the following steps are performed:
[0191] generating a first prompt word, wherein the first prompt word includes: the second image, extraction criteria of sub-image extraction, and prompt information for indicating that the second image is subjected to sub-image extraction based on the extraction criteria of sub-image extraction;
[0192] sending the first prompt word to a third artificial intelligence model, and receiving the position information returned by the third artificial intelligence model.
[0193] According to one or more embodiments of the present disclosure, example eight provides that the description information corresponding to the plurality of sub-images is generated according to the conference text of the first multimedia conference in example one, including:
[0194] for each sub-image in the plurality of sub-images, the following steps are performed:
[0195] determine conference text corresponding to a second image in which the sub-image is located, and determine location information of the sub-image in the second image;
[0196] generate, by a fourth artificial intelligence model, description information corresponding to the sub-image based on the conference text corresponding to the second image in which the sub-image is located and the location information of the sub-image in the second image.
[0197] According to one or more embodiments of the present disclosure, Example Nine provides that the type of the sub-image in any one of Examples One to Eight includes at least one of the following: picture data, table data, video data, presentation data, or code block data.
[0198] According to one or more embodiments of the present disclosure, Example Ten provides a multimedia conference data processing apparatus based on an artificial intelligence model, comprising:
[0199] an acquisition module configured to acquire a plurality of first images of a first multimedia conference, and acquire conference text of the first multimedia conference, wherein the plurality of first images are derived from screen sharing content of the first multimedia conference;
[0200] a screening module configured to screen the plurality of first images according to a correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image;
[0201] an extraction module configured to extract sub-images from the at least one second image to obtain a plurality of sub-images, wherein each of the plurality of sub-images includes non-text data carrying an amount of information satisfying a set condition;
[0202] a generation module configured to generate description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establish an association relationship between the plurality of sub-images and the corresponding description information.
[0203] The generation module is further configured to analyze the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by a first artificial intelligence model, determine source text for generating a conference summary of the first multimedia conference, and generate the conference summary of the first multimedia conference; in response to the source text including first description information, determine a first sub-image according to the association relationship, and add the first sub-image in the conference summary of the first multimedia conference, wherein the description information corresponding to the plurality of sub-images includes the first description information, and the plurality of sub-images include the first sub-image.
[0204] According to one or more embodiments of the present disclosure, example eleven provides an electronic device, comprising:
[0205] a processing device; and
[0206] a storage device comprising one or more computer program instructions;
[0207] wherein the one or more computer program instructions, when executed by the processing device, perform the artificial intelligence model-based multimedia conference data processing method provided by at least one embodiment of the present disclosure.
[0208] According to one or more embodiments of the present disclosure, example twelve provides a computer-readable storage medium, non-transitorily storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the artificial intelligence model-based multimedia conference data processing method provided by at least one embodiment of the present disclosure is implemented.
[0209] The above description is merely exemplary of the disclosure and the application of the principles thereof and it is not intended to limit the scope of the disclosure to the specific forms set forth. Rather, the scope of the disclosure is to be determined by the claims which are to be construed in accordance with the principles of patent law. For example, the various features of the disclosure are not limited to the specific embodiments described herein but can be utilized independently of, and in various permutations and combinations with, each other. Similarly, the various embodiments described herein are not limited to the specific embodiments described herein, but can be practiced with modification and alteration within the scope of the appended claims. The disclosure is to be construed as not limited only to the embodiments set forth herein but also to any other form falling within the scope of the claims and their equivalents.
[0210] Furthermore, while operations are depicted in a particular, chronological sequence in this specification, this should not be understood as requiring, or implied, that the operations be performed in that order - and that other operations be performed between.
[0211] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A multimedia conference data processing method based on an artificial intelligence model, comprising: obtaining a plurality of first images of a first multimedia conference, and obtaining conference text of the first multimedia conference, wherein the plurality of first images are derived from screen sharing content of the first multimedia conference; screening the plurality of first images according to a correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image; performing sub-image extraction on the at least one second image to obtain a plurality of sub-images, wherein each of the plurality of sub-images comprises non-text data carrying an information amount satisfying a set condition; generating description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establishing an association relationship between the plurality of sub-images and the corresponding description information; analyzing the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by using a first artificial intelligence model, to determine source text for generating a conference minutes of the first multimedia conference, and to generate the conference minutes of the first multimedia conference; in response to the source text including first description information, determining a first sub-image according to the association relationship, and adding the first sub-image in the conference minutes of the first multimedia conference, wherein the description information corresponding to the plurality of sub-images includes the first description information, and the plurality of sub-images includes the first sub-image.
2. The method of claim 1, wherein, the adding the first sub-image in the conference minutes of the first multimedia conference comprises: determining a first position associated with the first description information in the conference minutes of the first multimedia conference, wherein the conference minutes at the first position are generated based on the first description information; adding the first sub-image at the first position in the conference minutes of the first multimedia conference.
3. The method of claim 1, wherein, the obtaining a plurality of first images of a first multimedia conference comprises: obtaining a conference recording file of the first multimedia conference, wherein the conference recording file includes screen sharing content of the first multimedia conference; performing frame extraction processing on the conference recording file to obtain the plurality of first images.
4. The method of claim 1, wherein, the screening the plurality of first images according to a correlation degree between the plurality of first images and the conference text corresponding to the plurality of first images respectively, to obtain at least one second image, comprises: for each of the plurality of first images, performing the following steps: determining timestamp information of the first image; determining conference text corresponding to the first image according to the timestamp information of the first image; determining a processing mode of the first image according to a correlation degree between the first image and the conference text corresponding to the first image, wherein the processing mode comprises discarding the first image or retaining the first image; processing the first image according to the processing mode of the first image.
5. The method of claim 4, wherein, The processing manner of the first image is determined according to a correlation degree between the first image and the conference text corresponding to the first image, including: sending the first image and the conference text corresponding to the first image to a second artificial intelligence model, and receiving correlation degree information returned by the second artificial intelligence model; in response to the correlation degree information representing that the correlation degree between the first image and the conference text corresponding to the first image meets a correlation condition, determining that the processing manner of the first image is to retain the first image; or in response to the correlation degree information representing that the correlation degree between the first image and the conference text corresponding to the first image does not meet the correlation condition, determining that the processing manner of the first image is to discard the first image.
6. The method of claim 1, wherein, The sub-image extraction on the at least one second image is performed to obtain a plurality of sub-images, including: performing sub-image extraction on the at least one second image by using a third artificial intelligence model, and receiving position information of the plurality of sub-images returned by the third artificial intelligence model, wherein the position information indicates positions of the sub-images in the at least one second image; determining the plurality of sub-images from the at least one second image according to the position information of the plurality of sub-images.
7. The method of claim 6, wherein, The sub-image extraction on the at least one second image by using the third artificial intelligence model and receiving the position information of the plurality of sub-images returned by the third artificial intelligence model, includes: for each second image in the at least one second image, performing the following steps: generating a first prompt word, wherein the first prompt word includes the second image, extraction criteria for sub-image extraction, and prompt information for indicating that the sub-image extraction is performed on the second image based on the extraction criteria for sub-image extraction; sending the first prompt word to a third artificial intelligence model, and receiving the position information returned by the third artificial intelligence model.
8. The method of claim 1, wherein, The description information corresponding to the plurality of sub-images is generated according to the conference text of the first multimedia conference, including: for each sub-image in the plurality of sub-images, performing the following steps: determining conference text corresponding to a second image in which the sub-image is located, and determining position information of the sub-image in the second image; generating description information corresponding to the sub-image by using a fourth artificial intelligence model based on the conference text corresponding to the second image in which the sub-image is located and the position information of the sub-image in the second image.
9. The method according to any one of claims 1 to 8, wherein, The type of the sub-image includes at least one of the following: picture data, table data, video data, presentation data, or code block data.
10. An artificial intelligence model-based multimedia conference data processing apparatus, comprising: an acquisition module configured to acquire a plurality of first images of a first multimedia conference, and acquire conference text of the first multimedia conference, wherein the plurality of first images are derived from screen sharing content of the first multimedia conference; The screening module is configured to screen the plurality of first images according to a correlation degree between the plurality of first images and the conference texts corresponding to the plurality of first images respectively, to obtain at least one second image; The extraction module is configured to perform sub-image extraction on the at least one second image to obtain a plurality of sub-images, wherein each of the plurality of sub-images includes non-text data carrying an information amount satisfying a set condition; The generation module is configured to generate description information corresponding to the plurality of sub-images according to the conference text of the first multimedia conference, and establish an association relationship between the plurality of sub-images and the corresponding description information; The generation module is further configured to analyze the conference text of the first multimedia conference and the description information corresponding to the plurality of sub-images by using a first artificial intelligence model, to determine a source text used to generate a conference summary of the first multimedia conference, and to generate the conference summary of the first multimedia conference; in response to the source text including first description information, a first sub-image is determined according to the association relationship, and the first sub-image is added to the conference summary of the first multimedia conference, wherein the description information corresponding to the plurality of sub-images includes the first description information, and the plurality of sub-images includes the first sub-image.
11. An electronic device comprising: processing means; and storage means including one or more computer program instructions; wherein the one or more computer program instructions, when executed by the processing means, perform the method of any one of claims 1 to 9.
12. A computer-readable storage medium, non-transitorily storing computer- readable instructions, wherein, The computer readable instructions, when executed by the processor, implement the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Conference summary generation method and device based on artificial intelligence, equipment and medium
CN110866110A
Method for generating conference summary and equipment thereof
CN115240681A