Processing method and apparatus for conference data, division method and apparatus for media content, digest generation method and apparatus, electronic device and computer-readable medium

By determining the type and dividing content of the online conference data, meeting segments are obtained, and summary generation is generated based on the importance of sub-content of media content, the problem of inaccurate summary of media content in the prior art is solved, more accurate content segments and summary generation is achieved, and user experience is improved.

WO2025108378A1PCT designated stage expired Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133540
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to accurately summarize the main content of media content, which affects users' understanding of media content. Especially in the scenario of online meetings, the conference content is large and it is inconvenient for users to view it.

Method used

By obtaining the meeting data of the online meeting, determining the meeting type, and dividing the online meeting based on the meeting type and content, the meeting segments are obtained. At the same time, subcontents of media content are extracted and integrated according to their importance to generate an accurate media content summary.

Benefits of technology

It realizes more accurate segmentation of online conference content, which facilitates users to quickly understand the conference content and improve user experience; at the same time, the generated summary can more accurately summarize the main content of media content, meeting users' needs for quickly understanding media content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133540_30052025_PF_FP_ABST
    Figure CN2024133540_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure relates to a processing method and apparatus for conference data, a division method and apparatus for media content, a digest generation method and apparatus, an electronic device and a computer-readable medium. The processing method for conference data comprises: acquiring conference data of a network conference, wherein the conference data comprises voice data of the network conference; on the basis of the conference data, determining the conference type of the network conference; and on the basis of the conference type and conference content of the network conference, dividing the network conference to obtain conference segmentations of the network conference. The division method for media content comprises: acquiring data of the media content, wherein the media content comprises at least two content dimensions of content; on the basis of the content dimensions and data of the media content, determining target segmentation points of the media content; and on the basis of the target segmentation points, determining sub-content of the media content. The digest generation method comprises: acquiring content data of each sub-content comprised in media content, wherein the sub-content is obtained by dividing the media content; on the basis of the content data of each sub-content, determining a digest of the sub-content; and on the basis of a weight of each sub-content, fusing the digests of the sub-content to obtain a digest of the media content, wherein the weights of the sub-content are used for representing the degree of importance of the sub-content in the media content.
Need to check novelty before this filing date? Find Prior Art

Description

Conference data processing method and device, media content division method and device, summary generation method and device, electronic device and computer-readable medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese Patent Application No. 202311569183.4, No. 202311569948.4 and No. 202311569394.8 filed on November 22, 2023, and the contents of the above-mentioned Chinese patent application disclosures are hereby incorporated by reference in their entirety as part of this application. Technical Field

[0003] Embodiments of the present disclosure relate to a method and apparatus for processing conference data, a method and apparatus for dividing media content, a method and apparatus for generating a summary, an electronic device, and a computer-readable medium. Background Art

[0004] For some media content, users need to quickly understand the specific content. For example, in a video playback scenario, users may need to understand the general content of the video to decide whether to continue watching. Another example is a meeting scenario. After the meeting, users may need to review the meeting content to understand what was discussed.

[0005] Currently, users can quickly understand the content through summaries of media content. However, summaries often fail to accurately summarize the main points of the media content, hindering user understanding. A web conference, also known as an online meeting, is a meeting conducted over the internet. Participants can initiate and participate in a web conference online. Participants can use the client that provides the web conference service to automatically record the web conference content and generate meeting data. After the meeting, participants can use the meeting data to review the content.

[0006] However, usually, online conferences have a lot of content, which makes it inconvenient for participants to view the content. Summary of the Invention

[0007] An embodiment of the present disclosure provides a method for processing conference data, comprising: obtaining conference data of a network conference, wherein the conference data includes voice data of the network conference; determining the conference type of the network conference based on the conference data; and dividing the network conference based on the conference type and the conference content of the network conference to obtain conference segments of the network conference.

[0008] An embodiment of the present disclosure provides a method for dividing media content, comprising: obtaining data of media content, wherein the media content includes content of at least two content dimensions; determining a target segmentation point of the media content based on the content dimensions and the data of the media content; and determining sub-content of the media content based on the target segmentation point.

[0009] An embodiment of the present disclosure provides a summary generation method, comprising: obtaining content data of each sub-content included in media content, wherein the sub-content is obtained by dividing the media content; determining a summary of the sub-content based on the content data of each sub-content; and fusing the summaries of each sub-content based on the weight of each sub-content to obtain a summary of the media content, wherein the weight of the sub-content is used to represent the importance of the sub-content in the media content.

[0010] An embodiment of the present disclosure provides a device for processing conference data, including: a first acquisition unit, configured to acquire conference data of a network conference, wherein the conference data includes voice data of the network conference; a first determination unit, configured to determine the conference type of the network conference based on the conference data; and a first division unit, configured to divide the network conference based on the conference type and the conference content of the network conference to obtain conference segments of the network conference.

[0011] An embodiment of the present disclosure provides a device for dividing media content, including: a second acquisition unit, configured to obtain data of media content, wherein the media content includes content of at least two content dimensions; a second determination unit, configured to determine a target segmentation point of the media content based on the content dimensions and the data of the media content; and a second division unit, configured to determine sub-content of the media content based on the target segmentation point.

[0012] An embodiment of the present disclosure provides a summary generation device, comprising: a third acquisition unit, configured to acquire content data of each sub-content included in media content, wherein the sub-content is obtained by dividing the media content; a third determination unit, configured to determine a summary of the sub-content based on the content data of each sub-content; and a generation unit, configured to fuse the summaries of each sub-content based on the weight of each sub-content to obtain a summary of the media content, wherein the weight of the sub-content is used to represent the importance of the sub-content in the media content.

[0013] An embodiment of the present disclosure provides an electronic device, comprising: at least one processor; a storage device on which at least one program is stored, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the conference data processing method provided by any embodiment of the present disclosure, the media content division method provided by any embodiment of the present disclosure, or the summary generation method provided by any embodiment of the present disclosure.

[0014] An embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for processing conference data provided by any embodiment of the present disclosure, the method for dividing media content provided by any embodiment of the present disclosure, or the method for generating a summary provided by any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.

[0016] FIG1 is a flow chart of a method for processing conference data provided by an embodiment of the present disclosure;

[0017] FIG2 is a schematic diagram of the structure of a conference data processing device provided by an embodiment of the present disclosure;

[0018] FIG3 is a flow chart of a method for dividing media content provided by an embodiment of the present disclosure;

[0019] FIG4 is a schematic structural diagram of a device for dividing media content provided by an embodiment of the present disclosure;

[0020] FIG5 is a flowchart of a summary generation method provided by an embodiment of the present disclosure;

[0021] FIG6 is a schematic structural diagram of a summary generation device provided by an embodiment of the present disclosure; and

[0022] FIG7 is a schematic diagram of the basic structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure. After the online conference ends, conference data for recording the content of the conference can be generated. The conference data includes, for example, voice data of the online conference. Users with the authority to view the content of the conference can view the conference data and understand the content of the conference. However, for longer meetings, there is more content in the meeting. It is inconvenient for users to view a large amount of conference content.

[0024] Based on this, the embodiments of the present disclosure provide a method, apparatus, device and medium for processing conference data, in which the conference data of a network conference is obtained. The conference data includes voice data of the network conference. Based on the conference data, the conference type of the network conference is determined. The network conference is then divided based on the conference type and the conference content of the network conference to obtain conference segments of the network conference. By dividing the network conference into different sections according to the conference type and conference content, it is possible to obtain more accurate conference segments including different conference contents, thereby achieving a better segmentation effect of the network conference. In this way, the user can understand the conference content included in the conference segment by viewing the conference segment, which facilitates the user to browse the conference content and improves the user experience.

[0025] A method for processing conference data provided by an embodiment of the present disclosure can be applied to electronic devices with conference data processing capabilities. The electronic device may be, for example, a server or a terminal. The terminal includes but is not limited to a smartphone, a tablet computer, a laptop computer, a personal digital assistant (PDA) or a smart wearable device. The server may be a cloud server, such as a central server in a central computing cluster or an edge server in an edge computing cluster. Of course, the server may also be a server in a local data center. A local data center refers to a data center directly controlled by a user.

[0026] The electronic device obtains conference data from a web conference, including voice data, and determines the conference type by analyzing the conference data. The electronic conference is then divided into different conference segments based on the conference type, resulting in more reasonable conference segments containing different conference content. This allows users to understand the content of each conference segment, facilitating their browsing of the conference content and improving the user experience.

[0027] Those skilled in the art will appreciate that the above application scenario is only an example in which the embodiments of the present disclosure may be implemented, and the scope of application of the embodiments of the present disclosure is not limited in any aspect by this framework.

[0028] To facilitate understanding of the technical solution provided by the embodiment of the present disclosure, the method for processing conference data provided by the embodiment of the present disclosure is described below with reference to the accompanying drawings.

[0029] Referring to FIG. 1 , which is a flow chart of a method for processing conference data provided by an embodiment of the present disclosure, as shown in FIG. 1 , the method for processing conference data may include S101-S103:

[0030] S101: Acquire conference data of a network conference, where the conference data includes voice data of the network conference.

[0031] Web conference meeting data is data generated during a web conference and used to record the content of the meeting. The disclosed embodiments do not limit the source of web conference meeting data. In one possible implementation, web conference meeting data is generated when a web conference participant triggers the web conference recording function. In another possible implementation, web conference meeting data is data generated during a web conference and used to enable conference interaction between participants.

[0032] The conference data of an online conference includes voice data. Voice data is generated based on the voices of the participants in the online conference. In some possible implementations, the conference data of an online conference also includes other types of data, such as video data, shared content display data, and conference reservation data. Video data is generated based on the videos of the participants in the online conference. Shared content display data is generated based on the content shared by the participants in the online conference. Examples of shared content include device screens, files, videos, and images. Conference reservation data is data about the online conferences that participants have reserved to join. Conference reservation data includes relevant information about the online conference entered by the participants. The type of data included in the conference data of an online conference is determined by the interaction method used by the participants in the online conference.

[0033] S102: Determine the conference type of the network conference based on the conference data.

[0034] The conference type can be a predefined type based on the need for segmenting the web conference.

[0035] The embodiments of the present disclosure do not limit the way of dividing conference types. As an example, according to the different communication types of network conferences, network conferences are divided into interview type, sharing type, multi-topic type and general type. Among them, an interview type network conference, for example, can be a conference attended by two participants, where one party asks questions and the other party answers. A sharing type network conference, for example, can be a conference attended by at least two participants, where one participant gives an explanation. Specifically, a sharing type network conference can also be a conference attended by at least two participants, where one participant shares and explains content, and other participants ask questions. A multi-topic type network conference is a conference that discusses multiple conference topics. A multi-topic type network conference is usually attended by multiple participants. A general type network conference is a conference that does not belong to the interview type, sharing type and multi-topic type. The above method of dividing conference types is only an example and does not limit the specific method of dividing the conference types of network conferences. As another example, the conference types of network conferences can be divided according to the number of participants.

[0036] Different types of online conferences have different content characteristics. The conference data of the online conference is analyzed to determine the conference type of the online conference. In some possible implementations, the disclosed embodiments provide two specific implementations for analyzing conference data to determine the conference type of the online conference. See the following description for details.

[0037] S103: Divide the online conference based on the conference type and the conference content to obtain conference segments of the online conference.

[0038] By dividing network conferences according to conference type and conference content, we can obtain network segments with different main conference contents.

[0039] The embodiments of the present disclosure do not limit the manner of dividing network conferences based on conference type and conference content.

[0040] In one possible implementation, segmentation rules corresponding to the meeting type are pre-configured. These segmentation rules can be pre-set based on the meeting type and segmentation requirements. For example, if the meeting type is an interview, the segmentation rule for the interview type is to consider one question and one answer as one meeting segment. Alternatively, the segmentation rule for the interview type is to consider questions and answers related to the same question as one meeting segment. In other words, the first question, follow-up question, and answer all belong to the same meeting segment.

[0041] After determining the conference type, the online conference data is processed to obtain the conference content. The conference content is then segmented based on the conference content according to the segmentation rules corresponding to the conference type. For example, the conference content can be represented by conference text data obtained by performing speech recognition processing on the online conference voice data.

[0042] In another possible implementation, a segmentation model corresponding to the conference type is pre-trained. The segmentation model can identify the conference content of the online conference based on the conference data of the online conference, and segment the online conference based on the conference type. As an example, the segmentation model corresponding to the conference type can be trained using the conference data of the conference type and the segmentation results of the online conference. After determining the conference type of the online conference, the conference data of the online conference is processed using the segmentation model corresponding to the conference type to obtain the conference segments determined by the segmentation model. It should be noted that for voice data of online conferences of general types, a general segmentation model can be used to divide the conference segments.

[0043] The disclosed embodiments do not limit the method for identifying different conference segments. For example, in the voice data of a web conference, segment time points are determined. The segment time points are used to identify voice data before the segment time point as belonging to different conference segments than voice data after the segment time point.

[0044] Based on the relevant contents of S101-S103 above, it can be known that by dividing online conferences based on the conference type and conference content of the online conference, online conferences including different conference contents can be divided according to the content communication characteristics of the conference type. The conference content of the obtained conference segments is highly relevant, and the division of the conference segments is relatively accurate. Compared with segmenting online conferences according to a fixed division method, the conference data processing method provided by the embodiment of the present disclosure has a better division effect on online conferences and meets the needs of users to view conference content. By viewing the conference content included in the conference segment, users can quickly understand the content discussed in the conference segment, thereby improving the efficiency of users in viewing conference content.

[0045] The following embodiments of the present disclosure provide two possible implementation methods for determining the conference type of a network conference based on conference data.

[0046] The first method is to analyze conference data to obtain online conference communication information. Online conference communication information includes information related to the participants and the communication process among the participants within the online conference. For example, online conference communication information includes at least information representing the number of participants, information representing the communication format, and information representing the content of the communication. Information representing the number of participants indicates the number of participants, that is, the number of participants participating in the online conference. Information representing the communication format indicates the communication format, which is the form of communication between participants. Examples of communication formats include question-and-answer format, lecture format, and discussion format. Information representing the communication content indicates the specific content of the participants' communication. Information representing the communication format and content can be determined based on participant behavior information. Participant behavior information includes participant speaking time periods and the content of participants' speeches. A participant's speaking time period is the time period during which each participant speaks during the online conference. A participant's speech content is the content of each participant's speech during the online conference. Participant behavior information can be obtained based on analysis of conference data. In some possible implementations, conference data also includes shared content display data from the online conference. The communication format and content can also be determined based on the shared content display data. For example, based on the shared content display data, it can be determined that the communication information includes the explanation format. Based on the shared content display data, it can be determined that the communication content includes the shared content.

[0047] Determine the type of meeting for the web conference based on the communication information.

[0048] In some possible implementations, the conference type of the online conference is determined based on one or more of information including the number of participants, the communication format, and the content of the communication, included in the communication information. As an example, conference type determination criteria are pre-determined based on the characteristics of each conference type. The conference type determination criteria that the communication information satisfies are determined, and the conference type that satisfies the criteria is determined as the conference type of the online conference.

[0049] For example, the conditions for determining the interview type may include the number of participants being two and the communication format being question-and-answer. Based on the information representing the number of participants and the communication format included in the communication information, it is possible to determine whether the online conference is an interview type. For example, an online conference with two participants and a question-and-answer format may be determined to be an interview type.

[0050] As another example, the conditions for determining the interview type may be: the number of participants is two, the communication format is that the ratio of the speaking time of each participant to the total duration of the online meeting is less than 20%, and the communication content is that the proportion of questions in the speech of one of the two participants is greater than 80%.

[0051] For another example, a condition for determining the sharing type may be that the communication format includes a presentation format. Based on the information representing the communication format included in the communication information, it is possible to determine whether the online meeting is a sharing type. For example, an online meeting where the communication format includes a presentation format may be determined to be an interview type.

[0052] As another example, the conditions for determining the sharing type may be: the number of participants is greater than or equal to two, the communication form is that the ratio of the speaking time of the first participant among the participants to the total time of the online meeting is greater than 60%, the communication content is that the proportion of questions in the speech content of the second participant is greater than 80%, and the speech content of the first participant includes shared content.

[0053] For another example, a condition for determining a multi-topic type may be that the communication content includes multiple conference topics. Based on information representing the communication content included in the communication information, it is possible to determine whether the online conference is a multi-topic type. For example, a web conference whose communication content includes multiple conference topics may be determined to be an interview type.

[0054] As another example, the conditions for determining the multi-topic type are: the number of participants is greater than or equal to two, the communication form is that there is a participant among multiple participants whose speaking time ratio to the total time of the online meeting is greater than 30%, and the communication content is that the relevance of the participants' speech content in different speaking time periods is less than 20%.

[0055] The speaking time of a participant is the length of time included in the speaking period of the participant. The proportion of question content in the speech content of the participant is the ratio of question sentences included in the speech content.

[0056] In this way, the conference type of the network conference can be determined based on the communication information.

[0057] As another example, an artificial intelligence model is pre-trained to determine the type of a web conference. The training data for the artificial intelligence model, for example, includes training communication information obtained by analyzing training conference data from a training web conference, and labels corresponding to the training communication information. The training communication information includes one or more of information representing the number of participants, information representing the communication format, and information representing the content of the communication. The label corresponding to the training communication information is the type of the training web conference. The communication information is input into the trained artificial intelligence model, and the artificial intelligence model outputs the type of the web conference.

[0058] The second type: the conference data also includes the conference reservation data of the online conference. The conference reservation data includes the conference reservation type. The conference reservation type can be determined by the information of the scheduled conference pre-entered by the participants. For example, the conference reservation type is determined based on the title of the scheduled conference. For example, the title of the scheduled conference is "xx's sharing conference." Based on "xx's sharing conference", the conference reservation type is determined to be a sharing type. For another example, the conference reservation type is determined based on the meeting theme of the scheduled conference. For example, the meeting theme of the scheduled conference is "Discussion on business a, business b, and business c." Based on "Discussion on business a, business b, and business c", the conference reservation type is determined to be a multi-topic type. The embodiment of the present disclosure does not limit the method of determining the conference reservation type. As an example, the conference reservation type can be determined by analyzing the semantics of the conference reservation data. As another example, the conference reservation type can be determined by identifying the keywords included in the text of the conference reservation data.

[0059] The conference reservation type can be used as reference information to determine the conference type of the network conference. The conference type of the network conference is determined based on the communication information and the conference reservation type.

[0060] In one possible implementation, a determination is made as to whether the communication information satisfies the conference type determination criteria corresponding to the conference reservation type. If the communication information satisfies the conference type determination criteria, the conference type of the online conference is determined to be that conference type, i.e., the conference reservation type. If the communication information does not satisfy the conference type determination criteria, the determination is made as to other conference type determination criteria that the communication information satisfies.

[0061] In another possible implementation, an artificial intelligence model is pre-trained to determine the meeting type of a web conference. The training data for the artificial intelligence model, for example, includes training communication information, training meeting appointment types, and meeting type labels obtained by analyzing training meeting data from a training web conference. The training communication information includes one or more of information representing the number of participants, information representing the communication format, and information representing the content of the communication. The meeting type label is the meeting type of the training web conference. The communication information and meeting appointment type are input into the trained artificial intelligence model, and the artificial intelligence model outputs the meeting type of the web conference.

[0062] In some cases, a web conference may consist of sub-sessions of different types. For example, in an interview scenario, part of the web conference may consist of a question-and-answer session between the interviewer and the interviewee, while part of the session may consist of the interviewee explaining their responses to the interview test questions.

[0063] The present disclosure provides a possible implementation method for dividing online conferences based on conference type and conference content, including:

[0064] Divide the online conference into multiple sub-conferences according to the conference type;

[0065] Sub-conferences are divided based on their conference types and conference contents.

[0066] For webinars with multiple meeting types, first divide the webinars into multiple sub-meetings based on the meeting type. A sub-meeting is a segment of a webinars with the same meeting type. For example, if the first 50% of the webinars is an interview type, and the second 50% is a sharing type, then the first 50% and the second 50% of the webinars are divided into two sub-meetings.

[0067] Then, based on the conference type and content of the sub-conference, the sub-conference is divided to achieve the division of the network conference.

[0068] The method of segmenting the sub-conference is similar to the method of segmenting the network conference in S103 above, and will not be repeated here. The conference segments of each sub-conference obtained after segmentation can be used as the conference segments of the network conference.

[0069] In this way, it is possible to achieve more accurate segmentation of network conferences belonging to various conference types, and the resulting conference segmentation of different conference contents is more reasonable, making it easier for users to view the conference content.

[0070] In some possible scenarios, the divided conference segments can also be merged. For example, when the number of conference segments of a web conference is large, the divided conference segments are merged. As an example, determine whether the conference segments of a web conference exceed the quantity threshold. The quantity threshold is the maximum number of conference segments of a web conference that is preset. As an example, the quantity threshold is 12 segments per hour. That is to say, if the ratio of the number of conference segments to the number of hours of the web conference is greater than 12, it means that the number of conference segments exceeds the quantity threshold. For example, a web conference with a total duration of 2 hours has 30 conference segments, which exceeds the quantity threshold of 24.

[0071] The method for processing conference data provided by the embodiment of the present disclosure further includes the following steps:

[0072] Merge conference segments based on semantics. Semantics refers to the meaning of the main content of the conference segment.

[0073] The embodiments of the present disclosure do not limit the implementation method for determining the semantics of each conference segment. In one possible implementation method, text data is first extracted from the conference data of each conference segment, and the text data is input into a semantic extraction model to obtain the semantics of the conference segment output by the semantic extraction model. In another possible implementation method, text data is first extracted from the conference data of each conference segment, keywords are extracted from the text data, and the keywords are used to represent the semantics of the conference segment.

[0074] Determine the semantic similarity of conference segments. Merge conference segments with adjacent time periods and semantic similarity greater than a threshold. This reduces the number of conference segments while maintaining similar content, making it easier for users to view conference content.

[0075] Based on the conference data processing method provided by the above method embodiment, the embodiment of the present disclosure further provides a conference data processing device, which will be described below with reference to the accompanying drawings.

[0076] Refer to Figure 2, which is a schematic diagram of the structure of a conference data processing device provided by an embodiment of the present disclosure. As shown in Figure 2, the conference data processing device includes:

[0077] The first acquiring unit 201 is configured to acquire conference data of the network conference, where the conference data includes voice data of the network conference;

[0078] The first determining unit 202 is configured to determine the conference type of the network conference based on the conference data;

[0079] The first dividing unit 203 is configured to divide the online conference based on the conference type and the conference content of the online conference to obtain conference segments of the online conference.

[0080] In a possible implementation, the first determining unit 202 is specifically configured to obtain communication information of the network conference based on the conference data; and determine the conference type of the network conference based on the communication information.

[0081] In one possible implementation, the communication information includes one or more of the following information:

[0082] Information representing the number of participants, information representing the form of communication, and information representing the content of communication.

[0083] In a possible implementation, the conference data further includes shared content display data of the network conference, and the shared content display data is used to determine information representing the communication form and information representing the communication content.

[0084] In a possible implementation, the communication information includes information representing a communication form. The first determining unit 202 is configured to determine the conference type of the network conference based on the communication information, including:

[0085] The first determining unit 202 is configured to determine the conference type of the network conference as a sharing type if the information representing the communication form indicates that the communication form includes a lecture type.

[0086] In a possible implementation, the communication information includes information representing the number of participants and information representing the communication format. The first determining unit 202 is configured to determine the conference type of the online conference based on the communication information, including:

[0087] The first determining unit 202 is configured to determine the conference type of the network conference as an interview type if the information representing the number of participants indicates that there are two participants and the information representing the communication form indicates that the communication form is a question-and-answer form.

[0088] In a possible implementation, the communication information includes information representing the content of the communication. The first determining unit 202 is configured to determine the conference type of the network conference based on the communication information, including:

[0089] The first determining unit 202 is configured to determine the conference type of the network conference as a multi-topic type if the information representing the communication content indicates that the communication content includes multiple conference topics.

[0090] In a possible implementation, the conference data further includes a conference reservation type of the network conference, and the first determining unit 202 is specifically configured to obtain the communication information and the conference reservation type of the network conference based on the conference data;

[0091] Determine the type of online conference based on the communication information and the conference appointment type.

[0092] In a possible implementation, the first dividing unit 203 is configured to divide the online conference based on the conference type and the conference content of the online conference, including:

[0093] The first dividing unit 203 is configured to divide the network conference into multiple sub-conferences according to conference type. A sub-conference is a conference segment of the network conference belonging to the same conference type; and divide the sub-conference based on the conference type and conference content of the sub-conference.

[0094] In a possible implementation, the first dividing unit 203 is configured to divide the online conference based on the conference type and the conference content of the online conference, including:

[0095] The first dividing unit 203 is configured to adopt a segmentation rule corresponding to the conference type and divide the network conference according to the conference content.

[0096] In a possible implementation, the first dividing unit 203 is configured to divide the online conference based on the conference type and the conference content of the online conference, including:

[0097] The first segmentation unit 203 is configured to process the conference data of the online conference using a segmentation model corresponding to the conference type, where the segmentation model is used to segment the online conference.

[0098] In a possible implementation, the conference data processing device further includes:

[0099] The merging unit is configured to determine the semantics of each conference segment and merge the conference segments that are adjacent in time periods in the network conference and whose semantic similarity is greater than or equal to a similarity threshold.

[0100] Media content is presented through various dissemination methods. It can be disseminated via the internet, making it easy for users to view it. Some media content includes a lot of content. If a user is interested in a portion of a piece of media content, they may need to view other content within the content to locate the desired portion within the content and view it. This can result in a poor user experience when viewing media content.

[0101] Based on this, the embodiments of the present disclosure provide a method, apparatus, device, and medium for dividing media content. In this method, data of the media content is obtained, and the media content includes data content of at least two content dimensions; based on the content dimensions and the data of the media content, the target segmentation point of the media content is determined, and the sub-content of the media content is determined using the target segmentation point. By determining the target segmentation point of the media content based on multiple content dimensions, it is possible to divide the media content based on the characteristics of different content dimensions, and obtain more accurately divided sub-content of the media content. In this way, it is convenient for users to understand the partial content included in the media content by viewing the sub-content, thereby improving the user experience.

[0102] A method for dividing media content provided by an embodiment of the present disclosure can be applied to electronic devices with data processing capabilities. The electronic device may be, for example, a server or a terminal. The terminal includes but is not limited to a smartphone, a tablet computer, a laptop computer, a personal digital assistant (PDA) or a smart wearable device. The server may be a cloud server, such as a central server in a central computing cluster or an edge server in an edge computing cluster. Of course, the server may also be a server in a local data center. A local data center refers to a data center directly controlled by a user.

[0103] The electronic device obtains data of media content, where the media content includes data content of at least two content dimensions; determines target segmentation points of the media content based on the content dimensions and the data of the media content, and determines sub-content of the media content using the target segmentation points.

[0104] Those skilled in the art will appreciate that the above application scenario is only an example in which the embodiments of the present disclosure may be implemented, and the scope of application of the embodiments of the present disclosure is not limited in any aspect by this framework.

[0105] To facilitate understanding of the technical solution provided by the embodiment of the present disclosure, the method for dividing media content provided by the embodiment of the present disclosure is described below with reference to the accompanying drawings.

[0106] Referring to FIG. 3 , which is a flow chart of a method for dividing media content provided by an embodiment of the present disclosure, as shown in FIG. 3 , the method for dividing media content may include S301 to S303:

[0107] S301: Acquire data of media content, where the media content includes content of at least two content dimensions.

[0108] Media content is content generated based on a dissemination method. The embodiments of this disclosure do not limit the specific type of media content. As an example, the media content is film and television video content, or the media content is live video content, or the media content is conference content.

[0109] The data of the media content can include one or more types of data among video data, audio data, text data and image data. The data of the media content is specifically determined based on the type of the media content.

[0110] Media content includes content in multiple content dimensions. The disclosed embodiments do not limit the method for dividing content dimensions. As an example, different content dimensions can be divided based on different content generation methods. As another example, different content dimensions can be divided based on different data types.

[0111] As an example, taking the media content as conference content, the content dimensions included in the media content include one or more of the following content dimensions:

[0112] Voice content dimension, shared screen content dimension, shared document content dimension, stage time content dimension, and stage content dimension.

[0113] For a detailed introduction to these five content dimensions, please see below.

[0114] S302: Determine target segmentation points of the media content based on content dimensions and media content data.

[0115] Content of different content dimensions has different characteristics. According to the characteristics of the multiple content dimensions included in the media content, the data of the media content is processed separately using the division method corresponding to each content dimension to determine the segmentation points used to divide the media content. The segmentation points of the media content are used to identify the division positions of the divided media content. Taking the example that the data of the media content includes audio data or video data, the segmentation points of the media content are the moments in the timeline of the media content where the media content needs to be divided. Taking the example that the data of the media content includes text data, the segmentation points of the media content are the delimiters that divide the text in the text data.

[0116] Taking the scenario where media content is conference content as an example, the embodiments of the present disclosure respectively provide voice content dimension, shared screen content dimension, shared document content dimension, stage time content dimension and stage content dimension. These five content dimensions determine the implementation methods of the segmentation points of media content. Please see below for details.

[0117] There may be multiple segmentation points for the media content determined based on multiple content dimensions. At least one target segmentation point is determined from the multiple segmentation points as a segmentation point for dividing the media content to obtain sub-content.

[0118] The embodiments of the present disclosure do not limit possible implementation methods for determining the target segmentation point.

[0119] In a possible implementation, each segmentation point determined by multiple content dimensions is used as a target segmentation point for the media content.

[0120] In another possible implementation, all segmentation points determined based on each content dimension are used as candidate segmentation points of the media content, and a target segmentation point is selected from the candidate segmentation points of the media content.

[0121] The embodiment of the present disclosure does not limit the implementation method of selecting the target segmentation point from the candidate segmentation points.

[0122] As an example, multiple candidate segmentation points that meet the merging condition are merged into one target segmentation point, and candidate segmentation points that do not meet the merging condition are used as the target segmentation point.

[0123] The merging condition is, for example, that the distance between the division positions of the media content identified by the candidate segmentation points is less than a threshold value. The division positions of the divided media content indicated by the multiple candidate segmentation points that meet the merging condition are relatively small, which can indicate that the accuracy of the division at the division position is relatively high. Taking the case where the data of the media content includes audio data or video data, and the candidate segmentation points of the media content are the moments in the timeline of the media content where the media content needs to be divided, as an example, the merging condition is, for example, that the time interval between the candidate segmentation points is less than 5 minutes. The embodiment of the present disclosure does not limit the way of merging candidate segmentation points. As an example, any candidate segmentation point among the candidate segmentation points that meet the merging condition is selected as the target segmentation point obtained by merging the candidate segmentation points that meet the merging condition. In addition, the candidate segmentation points that do not meet the merging condition are used as target segmentation points. In this way, the number of target segmentation points can be reduced while ensuring the accuracy of dividing the media content, and the number of sub-contents can be reduced to facilitate user viewing.

[0124] As another example, the segmentation confidence of each candidate segmentation point is determined. The segmentation confidence is used to measure the segmentation accuracy of the candidate segmentation point and also indicates the reliability of the media content segmentation performed from the candidate segmentation point.

[0125] The segment confidence can be determined by a pre-set segment confidence configuration rule.

[0126] For example, the segmentation confidence is determined based on the division method of the candidate segmentation point. For example, the segmentation confidence is determined by the dimension type of the content dimension corresponding to the candidate segmentation point. Taking the above five content dimensions as an example, the segmentation confidence of the candidate segmentation point determined based on the voice content dimension and the shared content dimension, that is, the shared screen content dimension and the shared document content dimension is higher. The segmentation confidence of the candidate segmentation point determined based on the description content dimension, that is, the stage time content dimension and the stage content dimension is lower. For another example, the segmentation confidence is determined based on the accuracy of the candidate segmentation point determined by the division method. For example, the accuracy of the candidate segmentation point determined based on the description content dimension is determined by the similarity of the content before and after the candidate segmentation point. If the similarity is high, the accuracy of the candidate segmentation point is low and the segmentation confidence is low. If the similarity is low, the accuracy of the candidate segmentation point is high and the segmentation confidence is high.

[0127] For another example, the segmentation confidence is determined based on the segmentation granularity of the candidate segmentation point. Segmentation granularity refers to the granularity of dividing the media content of the candidate segmentation point. Segmentation granularity can be determined based on the type of content dimension. For example, taking the above five content dimensions as an example, based on the voice content dimension and the shared content dimension, that is, the shared screen content dimension and the shared document content dimension, the candidate segmentation points determined are fine-grained candidate segmentation points. Based on the description content dimension, that is, the stage time content dimension and the stage content dimension, the candidate segmentation points determined are coarse-grained candidate segmentation points. Coarse-grained candidate segmentation points may have the problem of inaccurate division, and the segmentation confidence of coarse-grained candidate segmentation points is low. The accuracy of the division of fine-grained candidate segmentation points is higher, and the segmentation confidence is higher.

[0128] It should be noted that the above two methods of determining segment confidence are only examples of determining segment confidence, and the embodiments of the present disclosure are not limited to this. In addition, multiple methods of determining segment confidence can be used independently or together. In the case where a candidate segment point has multiple segment confidences, the weighted value of each segment confidence of the candidate segment point can be calculated as the segment confidence of the candidate segment point. The weight of each segment confidence can be determined by the method of determining the segment confidence. For example, the weight of the segment confidence determined according to the division method is greater than the weight of the segment confidence determined according to the segment granularity.

[0129] After determining the segmentation confidence of each candidate segmentation point, sort the segmentation confidence of each candidate segmentation point from high to low, and select the candidate segmentation point with a segmentation confidence ranking before the preset sequence number as the target segmentation point. Alternatively, select the candidate segmentation point with a segmentation confidence greater than the confidence threshold as a high-priority candidate segmentation point, and select the candidate segmentation point with a segmentation confidence less than or equal to the confidence threshold as a low-priority candidate segmentation point. Select the high-priority candidate segmentation point as the target segmentation point.

[0130] S303: Determine sub-content of the media content based on the target segmentation point.

[0131] The media content is divided according to the target segmentation points to obtain sub-contents of the media content. For example, if the media content is conference content, the sub-contents are the conference segment contents.

[0132] Based on the above steps S301-S303, it can be seen that the target segmentation points for media content determined based on the content dimension can achieve segmentation of media content based on the characteristics of the content dimension. The resulting sub-content has a high degree of content relevance and is more accurately segmented. This allows users to quickly browse media content by viewing the sub-content without having to carefully review the entire media content, thereby improving the user experience.

[0133] The following introduces possible implementation methods for determining segmentation points of media content using five content dimensions, namely, voice content dimension, shared screen content dimension, shared document content dimension, stage time content dimension, and stage content dimension, provided in the embodiments of the present disclosure.

[0134] In some possible implementations, the data of the media content includes voice data, the media content includes voice content, and the at least two content dimensions involved in the media content include a voice content dimension.

[0135] The first segmentation point is determined based on the voice content dimension in the following way:

[0136] A1: Based on the voice data, obtain the speech content of the participant with the process guidance identity.

[0137] The voice data included in media content may be the voice data of people involved in the creation of the media content. For example, the media content is meeting content. The people are the meeting attendees. These people may include those with the role of process leader. People with the role of process leader are responsible for guiding the communication process of the content. For example, a person with the role of process leader may be the meeting host or organizer.

[0138] The embodiments of the present disclosure do not limit the method of determining the person whose identity information is the process guide identity.

[0139] In one possible implementation, the person pre-set as the process guide can be determined based on the information provided by each person in the media content. For example, the person pre-set as the process guide can be determined based on the person's name or person type. For example, a person with the person type "host" can be determined as the process guide.

[0140] In another possible implementation, the speech content of the person with the process guide identity has the characteristics of guiding the process. The voice data included in the media content data is analyzed to determine the person with the process guide identity.

[0141] As an example, the speech content of people in the media content is analyzed. By detecting whether the speech content includes sentences related to process guidance, or sentences with similar semantics to sentences related to process guidance, the person with the process guidance identity is determined. The sentences related to process guidance can be pre-set reference sentences. Reference sentences are, for example, "I will preside over the process of today's meeting", "First, let's discuss...", etc. In some scenarios, based on the fact that the speech order of people with the process guidance identity is earlier, the speech content of people whose speech order is within the sequence threshold is analyzed. The sequence threshold can be determined based on the number of people included in the media content. For example, for media content including 5 people, the sequence threshold is, for example, 3. In this way, the scope of detecting people with the process guidance identity can be narrowed, and costs can be reduced.

[0142] As another example, a recognition model capable of identifying people in media content is pre-trained. The recognition model can use input voice data, or text data obtained by recognizing the voice data, to determine the identity of the person in the media content. The voice data, or text data obtained by recognizing the voice data, included in the data is input into the recognition model, and the recognition model outputs a process-guided identification of the person.

[0143] A2: Determine the first segmentation point of the media content based on different process stages indicated by the speech content.

[0144] The speech content of the person with the process guidance identity includes content indicating different process stages of the media content. As an example, semantic recognition is performed on the speech content of the person with the process guidance identity to determine the sentences indicating different process stages. The sentences indicating different process stages are, for example, process sentences and summary sentences. Based on the different process stages indicated by the speech content of the person with the process guidance identity, the first segmentation point for dividing the media content into different process stages is determined. As an example, the switching moment of different process stages is used as the first segmentation point of the media content. The first segmentation point is used to determine the target segmentation point.

[0145] The above method for determining the first segmentation point based on the voice content dimension is only an example, and the embodiments of the present disclosure are not limited thereto. For another example, the first segmentation point can be determined based on the moment when a person pauses in the voice data.

[0146] In some possible implementations, the data of the media content includes shared data, the media content includes shared content, and at least two content dimensions involved in the media content include shared content dimensions.

[0147] As an example, shared data includes one or more of shared screen data and shared document data. Shared screen data, for example, is a shared device screen, or video data generated by an interface window displayed on a shared device screen. Shared screen data can reflect the screen content shared by a user. Different screen content can reflect different content stages in the media content. Shared document data is video data generated by a shared document. Shared document data can reflect the shared document content. Different document content displayed can reflect different content stages in the media content.

[0148] When the shared data includes shared screen data, the shared content dimension includes a shared screen content dimension. When the shared data includes shared document data, the shared content dimension includes a shared document content dimension.

[0149] The second segmentation point is determined for the shared screen content dimension in the following manner:

[0150] B1: Determine a time to switch the content of the shared screen based on the shared screen content.

[0151] The switching of the content of the shared screen can reflect the change of the content stage included in the media content. The embodiment of the present disclosure does not limit the method of determining the switching of the content of the shared screen. As an example, the images displayed on the shared screen in different time periods are obtained from the shared screen data. Then, by determining the similarity of the images displayed on the screens of adjacent time periods, it is determined whether the content of the shared screen has changed. The moment separating the two time periods whose image similarity is less than the similarity threshold is used as the moment of switching the content of the shared screen. As another example, based on the shared screen data, the operation of the operating cursor in the shared screen is detected to determine the moment of switching the content of the shared screen. For example, the moment when the operating cursor clicks the button for switching pages is used as the moment of switching the content of the shared screen. As another example, the moment of starting the shared screen and the moment of ending the shared screen are used as the moments of switching the content of the shared screen.

[0152] B2: Determine a second segmentation point of the media content based on the moment when the content of the shared screen is switched.

[0153] As an example, the moment when the content of the shared screen is switched is determined as the second segmentation point for dividing the media content. Alternatively, the moment when the content of the shared screen is switched is partially determined as the second segmentation point for dividing the media content. The second segmentation point is used to determine the target segmentation point.

[0154] The third segmentation point is determined for the shared document content dimension in the following way:

[0155] C1: Based on the shared document content, determine the switching time of the document content.

[0156] The shared document title can reflect the different contents involved in the process of sharing the document. The shared document title can be determined based on the processing of the shared document data. In one possible implementation, the position of the operating cursor in the display area of ​​the document is detected based on the shared document data. Based on the position of the operating cursor in the display area of ​​the document, the content of the currently shared document is determined. For example, when it is detected that the operating cursor selects the document title, or is in the display area of ​​the document title, or the operating cursor moves to the display area of ​​the content corresponding to a different document title, it is determined that the content of the document has changed, and this moment is used as the switching moment of the document content. In another possible implementation, the document content being shared is displayed in a special display manner in the shared document. Based on the shared document data, the moment when the document title to which the document content specially displayed in the document belongs changes is used as the switching moment of the document content.

[0157] C2: Determine the third segmentation point of the media content based on the switching moment of the document content.

[0158] As an example, the determined switching moments of the document content are used as the third segmentation points of the media content. Alternatively, a portion of the switching moments of the determined document content are selected as the third segmentation points of the media content. The third segmentation points are used to determine the target segmentation points.

[0159] In some possible implementations, the media content includes descriptive content. The at least two content dimensions involved in the media content include a descriptive content dimension. The descriptive content is content that describes the media content. The data of the media content includes descriptive data. The descriptive data is, for example, text data. As an example, the media content is meeting content, and the descriptive content is the meeting agenda content.

[0160] As an example, the description data includes one or more of stage time data and stage content data. The stage time data includes time information of different content stages in the media content. The stage content data includes information about the main content of different content stages in the media content.

[0161] For example, if the data describes a meeting agenda, the phase time data includes the time information for different content phases within the meeting. For example, the first half hour of the meeting discussed topic A, followed by topic B. Another example might be: Issue X was discussed from 4:00 to 5:00, and issue Y was discussed from 5:00 to 5:30. Phase content data includes information about the main content of each phase within the meeting, for example, the meeting discussed: First, topic A; Second, topic B.

[0162] When the description data includes stage time data, the description content dimension includes the stage time content dimension. When the description data includes stage content data, the description content dimension includes the stage content dimension.

[0163] The fourth segmentation point is determined in the following way for the stage time content dimension:

[0164] Based on the time indicated by the stage time content, a fourth segmentation point of the media content is determined.

[0165] In one possible implementation, each moment indicated by the stage time content is used as the fourth segmentation point of the media content. In another possible implementation, the main content of the media content in the time period before and the main content of the media content in the time period after of each moment indicated by the stage time content are determined. The duration of the front time period and the back time period can be a preset duration. If the content similarity of the main content of the media content in the time period before and the main content of the media content in the time period after a moment is less than or equal to a threshold, then the moment is used as the fourth segmentation point of the media content. If the content similarity of the main content of the media content in the time period before and the main content of the media content in the time period after a moment is greater than a threshold, then the moment is discarded.

[0166] The fifth segmentation point is determined in the following way for the stage content dimension:

[0167] The media content is clustered based on the stage content, and a fifth segmentation point of the media content is determined based on different clustered contents.

[0168] The stage content data indicates the main content of each content stage included in the media content. According to the main content indicated by the stage content data, based on the data of the media content, the media content is clustered. In one possible implementation, the data of the media content includes voice data, and the voice data is converted into text data. Semantic clustering is performed on the sentences included in the text data to achieve clustering of the media content. The clustering categories are the different main content categories indicated by the stage content data. After obtaining different cluster contents, the fifth segmentation point of the media content is determined. The fifth segmentation point is used to divide the different cluster contents included in the media content. The fifth segmentation point is used to determine the target segmentation point.

[0169] The above are five possible implementation methods for determining segmentation points based on content dimensions provided in the embodiments of the present disclosure. The above implementation methods for determining segmentation points are all examples and are not intended to be limiting methods for determining segmentation points based on content dimensions.

[0170] Based on the media content division method provided by the above method embodiment, the embodiment of the present disclosure further provides a media content division device, which will be described below with reference to the accompanying drawings.

[0171] Referring to FIG4 , which is a schematic diagram of the structure of a media content division device provided by an embodiment of the present disclosure, the media content division device includes:

[0172] The second acquisition unit 401 is configured to acquire data of media content, where the media content includes content of at least two content dimensions;

[0173] The second determining unit 402 is configured to determine a target segmentation point of the media content based on the content dimension and the data of the media content;

[0174] The second segmentation unit 403 is configured to determine sub-content of the media content based on the target segmentation point.

[0175] In one possible implementation, the second determination unit 402 is specifically configured to determine candidate segmentation points of the media content under the content dimension based on the division method corresponding to the content dimension and the data of the media content; and determine the target segmentation points of the media content based on the candidate segmentation points.

[0176] In a possible implementation, the second determining unit 402 is configured to determine a target segmentation point of the media content based on the candidate segmentation points, including:

[0177] The second determining unit 402 is configured to determine a target segmentation point of the media content based on the segmentation confidence of the candidate segmentation point. The segmentation confidence is used to measure the accuracy of the segmentation of the candidate segmentation point.

[0178] In a possible implementation, the segmentation confidence is determined based on a division method for determining candidate segmentation points.

[0179] In a possible implementation, the segmentation confidence is determined based on the segmentation granularity of the candidate segmentation point, where the segmentation granularity indicates the degree of refinement of the media content divided by the candidate segmentation point.

[0180] In a possible implementation, the media content includes voice content, the content dimension includes a voice content dimension, and the first segmentation point is determined for the voice content dimension in the following manner, where the first segmentation point is used to determine a target segmentation point:

[0181] Based on the voice data, the speech content of the person with the process guide identity included in the media content is obtained;

[0182] Based on different process stages indicated by the speech content, a first segmentation point of the media content is determined.

[0183] In one possible implementation, the media content includes shared content.

[0184] In a possible implementation, the shared content includes shared screen content, the content dimension includes a shared screen content dimension, and the second segmentation point is determined based on the shared screen content dimension in the following manner. The second segmentation point is used to determine the target segmentation point:

[0185] Determining a time to switch the content of the shared screen based on the shared screen content;

[0186] A second segmentation point of the media content is determined based on a moment when the content of the shared screen is switched.

[0187] In a possible implementation, the shared content includes shared document content, the content dimension includes a shared document content dimension, and the third segmentation point is determined with respect to the shared document content dimension in the following manner. The third segmentation point is used to determine the target segmentation point:

[0188] Determine the switching time of the document content based on the shared document content;

[0189] Based on the switching moment of the document content, a third segmentation point of the media content is determined.

[0190] In one possible implementation, the media content includes description content.

[0191] In a possible implementation, the description content includes stage time content, which is a stage time content dimension. The fourth segmentation point is determined based on the stage time content dimension in the following manner. The fourth segmentation point is used to determine the target segmentation point:

[0192] Based on the time indicated by the stage time content, a fourth segmentation point of the media content is determined.

[0193] In a possible implementation, the description content includes stage content, which is a stage content dimension. The fifth segmentation point is determined based on the stage content dimension in the following manner. The fifth segmentation point is used to determine the target segmentation point:

[0194] The media content is clustered based on the stage content, and the fifth segmentation point of the media content is determined based on different clustered contents.

[0195] In a possible implementation, the media content is conference content, and the sub-content of the media content is conference segment content.

[0196] Media content refers to content expressed through various means of dissemination. Media content includes, for example, one or more of video content, image content, audio content, and text content. Media content can be disseminated via the Internet, making it easy for users to view it. Currently, there is a large amount of media content available to users, and some media content includes relatively rich content. To facilitate user viewing, a summary describing the main content of the media content can be provided to the user. By viewing the summary, the user can quickly understand the specific content included in the media content, making it easier for the user to select the media content they need to view or to enhance their understanding of the media content. However, current media content summaries are difficult to accurately describe the main content included in the media content.

[0197] Based on this, the embodiments of the present disclosure provide a summary generation method, apparatus, device and medium. In this method, the content data of each sub-content included in the media content is obtained, and the summary of each sub-content is obtained based on the content data of each sub-content. The summary of the sub-content can describe the main content of the sub-content. Extracting the summary of the sub-content first can reduce the difficulty of generating the summary of the media content. Based on the weight of each sub-content, the summary of each sub-content is fused to obtain the summary of the media content. Among them, the weight of the sub-content can reflect the importance of the sub-content in the media content. In this way, referring to the importance of each sub-content in the media content, the summary of the sub-content is fused, which can reduce the omission of important content and avoid excessive description of non-important content. The summary of the media content obtained can more accurately describe the main content of the media content, making it easier for users to understand the media content through the summary.

[0198] An embodiment of the present disclosure provides a summary generation method that can be applied to electronic devices with data processing capabilities. The electronic device may be, for example, a server or a terminal. The terminal includes but is not limited to a smartphone, a tablet computer, a laptop computer, a personal digital assistant (PDA), or a smart wearable device. The server may be a cloud server, such as a central server in a central computing cluster, or an edge server in an edge computing cluster. Of course, the server may also be a server in a local data center. A local data center refers to a data center that is directly controlled by the user.

[0199] The electronic device obtains the content data of each sub-content included in the media content, obtains the summary of each sub-content based on the content data of each sub-content, and then fuses the summaries of each sub-content based on the weight of each sub-content to obtain a summary of the media content. The weight of the sub-content is used to represent the importance of the sub-content in the media content. Firstly extracting the summary of the sub-content and then fusing the summary of the sub-content to obtain the summary of the media content can reduce the difficulty of generating the summary of the media content and facilitate the generation of the summary. Referring to the importance of each sub-content in the media content, fusing the summary of the sub-content can reduce the omission of important content and avoid excessive description of non-important content. The obtained summary of the media content can more accurately summarize the main content of the media content and meet the needs of users to view the summary of the media content.

[0200] Those skilled in the art will appreciate that the above application scenario is only an example in which the embodiments of the present disclosure may be implemented, and the scope of application of the embodiments of the present disclosure is not limited in any aspect by this framework.

[0201] To facilitate understanding of the technical solution provided by the embodiments of the present disclosure, the summary generation method provided by the embodiments of the present disclosure is described below with reference to the accompanying drawings.

[0202] Referring to FIG. 5 , which is a flowchart of a summary generation method provided by an embodiment of the present disclosure, as shown in FIG. 5 , the summary generation method may include S501 to S503:

[0203] S501: Acquire content data of each sub-content included in the media content.

[0204] The media content includes, for example, one or more of video content, image content, audio content, and text content, which is not limited in the embodiments of the present disclosure.

[0205] Sub-content is a portion of media content obtained by dividing the media content. The media content includes at least two sub-contents.

[0206] The embodiments of the present disclosure do not limit the manner of dividing sub-contents. In one possible implementation, the media content is divided evenly according to the duration to obtain various sub-contents. For example, if the media content is video content, the video content is divided into 20-minute durations to obtain a plurality of sub-contents. In another possible implementation, the content included in the media content is analyzed, the content is clustered, and the media content is divided based on the same type of content. In this way, part of the content with a high degree of content relevance can be divided into a sub-content. The sub-content obtained by such division conforms to the actual content structure of the media content, and can extract a more accurate summary of the sub-content, thereby generating a summary that summarizes the media content more accurately.

[0207] In a possible implementation, the media content is the content of a conference, for example, and the sub-content is the content of a sub-conference obtained by dividing the conference.

[0208] The embodiment of the present disclosure does not limit the method of dividing a conference into sub-conferences.

[0209] As an example, a meeting can be divided into sub-conferences based on its meeting type and content. The meeting type can be determined based on the meeting's communication information. The communication information may include, for example, one or more of information representing the number of attendees, information representing the communication format, and information representing the communication content. The meeting can be divided into sub-conferences using a division model corresponding to the meeting type. Alternatively, the meeting can be divided into sub-conferences using division rules corresponding to the meeting type.

[0210] As another example, the content of the meeting includes content of at least two content dimensions. The meeting is divided based on the content dimension to determine the sub-conference. In one possible implementation, a division method corresponding to the content dimension is used to determine the candidate segmentation points of the meeting and the segmentation confidence of the candidate segmentation points. The target segmentation point is determined based on the segmentation confidence, and the target segmentation point is used to divide the meeting to obtain sub-conferences. In another possible implementation, the meeting is divided using a division method corresponding to the content dimension to obtain meeting segments. The meeting segments that meet the conditions are then merged to obtain sub-conferences. Meeting segments that meet the conditions are, for example, meeting segments that are adjacent in time periods in the meeting and whose content similarity is greater than the similarity threshold.

[0211] Alternatively, if the meeting is a recurring scheduled meeting, the sub-content is the content of at least one scheduled meeting included in the recurring scheduled meeting. For example, if the meeting is a regular meeting held every Monday afternoon, the sub-content is the content of at least one Monday afternoon meeting included in the regular meeting.

[0212] The content data of a sub-content is data related to the sub-content. For example, if the sub-content is a sub-conference, the content data of the sub-conference is conference data. The present disclosure does not limit the specific type of the content data of the sub-content. For example, the content data of the sub-content can be one or more of audio data, video data, image data, and text data.

[0213] S502: Determine a summary of each sub-content based on the content data of each sub-content.

[0214] The summary of the sub-content is used to describe the main content of the sub-content. The embodiments of the present disclosure do not limit the implementation method of determining the summary of each sub-content based on the content data of each sub-content. In one possible implementation method, the keywords of the sub-content are determined by analyzing the content data of the sub-content. The summary of the sub-content is generated based on the keywords. In another possible implementation method, the content data of the sub-content is processed using a second language processing model to obtain a summary of the sub-content. The second language processing model has a natural language processing function. The second language processing model can analyze the input content data and output a summary.

[0215] The amount of content data for each sub-content included in the media content is relatively small, making it easier to process the sub-content content data to obtain a summary of the sub-content. This reduces the cost of generating a summary of the sub-content and increases the accuracy of the summary of the sub-content, thereby improving the accuracy of the resulting media content summary.

[0216] S503: The summaries of the sub-contents are integrated based on the weights of the sub-contents to obtain a summary of the media content.

[0217] Different sub-contents have different importance in the media content, and the importance of the sub-content in the media content is represented by the weight of the sub-content.

[0218] The embodiment of the present disclosure does not limit the method for determining the weight of the sub-content.

[0219] In one possible implementation, the weight of each sub-content may be set by, for example, the creator of the media content. For example, if the media content is generated based on a web conference, the weight of each sub-content in the media content may be set by, for example, the organizer, host, or other person with conference management authority of the web conference.

[0220] In another possible implementation, the weight of the sub-content is determined based on the content information of the sub-content. As an example, the content information of the sub-content is determined by the time information of the sub-content in the media content, and the content data. For example, the content information of the sub-content includes the time information of the sub-content in the media content, and the relevant information of the specific content determined by the content data. The time information of the sub-content in the media content can reflect the importance of the sub-content in the media content to a certain extent. For example, the sub-content at the beginning of the media content is usually the introduction part and has a lower importance. For another example, the sub-content at the end of the media content is usually the summary part and has a higher importance. Analyzing the content data can obtain the specific content of the sub-content, and then determine the importance of the sub-content and determine the weight of the sub-content.

[0221] As an example, an artificial intelligence model for determining the weight of sub-content is pre-trained. For example, the artificial intelligence model is trained using training data including training content information and labels for the training content information. The labels for the training content information are the weights of the training content information. The trained artificial intelligence model is capable of determining the weights of sub-content based on the input content information of the sub-content. As another example, the labels for the training content information are the importance values ​​of the training content information. The importance values ​​are used to measure the importance of the training content information. The trained artificial intelligence model is capable of determining the importance values ​​of sub-content based on the input content information of the sub-content. The weights of the sub-content are then determined based on the weights corresponding to the importance values ​​of the sub-content.

[0222] As another example, the content information of a sub-content is analyzed to obtain sub-information of at least one dimension capable of determining the weight of the sub-content. The presently disclosed embodiments do not limit the division method of the dimensions included in the content information. As an example, the sub-information included in the content information may include one or more of the following: the duration of the sub-content, the number of characters involved in the sub-content, and the position of the sub-content time period within the time period of the media content. The duration of the sub-content and the position of the sub-content time period within the time period of the media content can be determined by analyzing the time information of the sub-content within the media content. The number of characters involved in the sub-content can be obtained by analyzing the content data of the sub-content. In another possible implementation, for a scenario where the media content is about a meeting, the characters involved in the sub-content are attendees, and the sub-information also includes the attendee's status within the meeting. The status can be determined based on the attendee's speaking style, speaking frequency, and other speech characteristics included in the sub-content's content data. For example, if the speaking style is summarizing the event and the speaking frequency is high, the participant is determined to be the speaker. If the speaking style is promoting the process and the speaking frequency is high, the participant is determined to be the moderator.

[0223] Each sub-information has a corresponding sub-weight. The sub-weight of the sub-information can be determined based on a pre-set sub-weight determination rule. For example, the sub-weight of the sub-content at the initial stage of the media content may be lower, while the sub-weight of the sub-content at the final stage of the media content may be higher.

[0224] When conference information includes multiple sub-information items of different dimensions, the weight of the sub-content is determined based on the sub-weights of each sub-information item. For example, a statistical value of the sub-weights of each sub-information item included in the sub-content is calculated as the weight of the sub-content. The statistical value can be a value obtained using statistical methods, such as an average, weighted average, or median.

[0225] The weight of the sub-content can affect the proportion of the summary of the sub-content included in the summary of the media content. Based on the weight of each sub-content, the summaries of each sub-content are integrated to obtain the summary of the media content.

[0226] The embodiments of the present disclosure do not limit the implementation method of obtaining a summary of the media content by fusing the summaries of the sub-contents based on the weights of the sub-contents. In one possible implementation method, the weights of the sub-contents and the summaries of the sub-contents are processed by a first language processing model to obtain a summary of the media content. The first language processing model has the ability to process natural language. The first language processing model can be used to fuse the summaries based on the summaries and the weights of the summaries and output a summary. In another possible implementation method, rules for fusing summaries are pre-set. The rules for fusing summaries are used to define a method for fusing summaries of different weights. Based on the rules for fusing summaries, the summaries of the sub-contents are processed to obtain a summary of the media content.

[0227] Based on the relevant content of S501-S503 above, it can be seen that first extracting sub-content summaries and then fusing the sub-content summaries to obtain a media content summary can reduce the difficulty of generating a media content summary and facilitate summary generation. By fusing the sub-content summaries based on the importance of each sub-content within the media content, it can reduce the omission of important content and avoid excessive description of non-important content. The resulting media content summary can more accurately summarize the main content of the media content, meet the user's need for viewing media content summaries, and improve the user experience.

[0228] Based on the summary generation method provided by the above method embodiment, the embodiment of the present disclosure further provides a summary generation device, which will be described below with reference to the accompanying drawings.

[0229] See FIG6 , which is a schematic diagram of the structure of a summary generation device provided by an embodiment of the present disclosure. As shown in FIG6 , the summary generation device includes:

[0230] The third acquisition unit 601 is configured to acquire content data of each sub-content included in the media content, where the sub-content is obtained by dividing the media content;

[0231] The third determining unit 602 is configured to determine a summary of the sub-content based on the content data of each sub-content;

[0232] The generating unit 603 is configured to merge the summaries of the sub-contents based on the weights of the sub-contents to obtain a summary of the media content. The weights of the sub-contents are used to indicate the importance of the sub-contents in the media content.

[0233] In a possible implementation, the weight of the sub-content is determined according to content information of the sub-content, and the content information is determined based on time information of the sub-content in the media content and content data of the sub-content.

[0234] In a possible implementation, the weight of the sub-content is determined based on the sub-weight corresponding to the sub-information included in the content information of the sub-content.

[0235] In a possible implementation, the content information includes one or more of the following sub-information: duration of the sub-content, number of characters involved in the sub-content, and position of the time period of the sub-content in the time period of the media content.

[0236] In one possible implementation, the weight of the sub-content is determined based on an artificial intelligence model, and the artificial intelligence model is used to output the weight based on input content information.

[0237] In a possible implementation, the generating unit 603 is specifically configured to process the weights of the sub-contents and the summaries of the sub-contents based on the first language processing model to obtain a summary of the media content.

[0238] In a possible implementation, the content data is text data, and the third determining unit 602 is configured to process the content data of each sub-content separately based on the second language processing model to obtain a summary of each sub-content.

[0239] In a possible implementation, the media content is the content of a meeting, the sub-content is the content of sub-meetings obtained by dividing the meeting, or the meeting is a recurring scheduled meeting, and the sub-content is the content of at least one scheduled meeting included in the recurring scheduled meeting.

[0240] In a possible implementation, the sub-conferences are obtained by dividing the conference into sub-conferences by using the following method: dividing the conference into sub-conferences based on the conference type and conference content of the conference.

[0241] In a possible implementation, the sub-conferences are obtained by dividing the conference in the following manner: dividing the conference based on at least two content dimensions of the conference to obtain sub-conferences.

[0242] Based on the conference data processing method, media content division method or summary generation method provided in the above embodiments, the present disclosure also provides an electronic device, including: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the conference data processing method, media content division method or summary generation method provided in any of the above embodiments.

[0243] Reference is now made to FIG7 , which illustrates a schematic diagram of the structure of an electronic device 700 suitable for implementing an embodiment of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable Android devices), PMPs (Portable Media Players), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital televisions (TVs) and desktop computers. The electronic device illustrated in FIG7 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.

[0244] As shown in Figure 7, electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of electronic device 700 are also stored in RAM 703. Processing device 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0245] Typically, the following devices may be connected to the I / O interface 705: an input device 708 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 shows the electronic device 700 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0246] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowcharts (e.g., Figures 1, 3, and 5) can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method shown in the above flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 709, or installed from a storage device 708, or installed from a ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the conference data processing method, media content division method, or summary generation method provided in the embodiment of the present disclosure are executed.

[0247] The electronic device provided by the embodiment of the present disclosure and the conference data processing method, media content division method or summary generation method provided by the above-mentioned embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above-mentioned embodiment, and this embodiment has the same beneficial effects as the above-mentioned embodiment.

[0248] Based on the conference data processing method, media content division method or summary generation method provided in the above-mentioned method embodiments, the embodiments of the present disclosure provide a computer storage medium on which a computer program is stored, wherein when the program is executed by a processor, it implements the conference data processing method, media content division method or summary generation method provided in any of the above-mentioned embodiments.

[0249] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0250] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0251] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0252] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the conference data processing method, media content division method or summary generation method provided in any of the above embodiments.

[0253] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0254] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0255] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. The name of a unit / module does not necessarily limit the unit itself. For example, a voice data acquisition module may also be described as a "data acquisition module."

[0256] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0257] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0258] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to in detail. For the systems or devices disclosed in the embodiments, since they correspond to the conference data processing method, media content segmentation method, or summary generation method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method section.

[0259] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0260] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0261] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present disclosure. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is intended to be construed in the widest manner consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing conference data, comprising: Acquire conference data of a network conference, wherein the conference data includes voice data of the network conference; Determining the conference type of the network conference based on the conference data; The network conference is divided based on the conference type and the conference content of the network conference to obtain conference segments of the network conference.

2. The method according to claim 1, wherein: The determining the conference type of the network conference based on the conference data includes: Based on the conference data, obtaining communication information of the network conference; Based on the communication information, a conference type of the network conference is determined.

3. The method according to claim 2, wherein: The communication information includes one or more of the following information: Information representing the number of participants, information representing the form of communication, and information representing the content of communication.

4. The method according to claim 3, wherein: The conference data also includes shared content display data of the network conference, and the shared content display data is used to determine the information representing the communication form and the information representing the communication content.

5. The method according to claim 3, wherein: The communication information includes information representing a communication form, and determining the conference type of the network conference based on the communication information includes: If the information representing the communication form indicates that the communication form includes an explanation form, the conference type of the network conference is determined to be a sharing type.

6. The method according to claim 3, wherein: The communication information includes the information representing the number of participants and the information representing the communication form, and determining the conference type of the network conference based on the communication information includes: If the information representing the number of participants indicates that the number of participants is two, and the information representing the communication form indicates that the communication form is a question-and-answer form, the conference type of the network conference is determined to be an interview type.

7. The method according to claim 3, wherein: The communication information includes the information representing the communication content, and the determining the conference type of the network conference based on the communication information includes: If the information representing the communication content indicates that the communication content includes multiple conference topics, the conference type of the network conference is determined to be a multi-topic type.

8. The method according to claim 1, wherein: The conference data also includes a conference reservation type of the network conference, and determining the conference type of the network conference based on the conference data includes: Based on the conference data, obtaining communication information of the network conference and the conference reservation type; Based on the communication information and the conference reservation type, a conference type of the network conference is determined.

9. The method according to any one of claims 1 to 8, wherein: The dividing the network conference based on the conference type and the conference content of the network conference includes: Dividing the network conference into a plurality of sub-conferences according to the conference type, wherein the sub-conferences are conference segments of the network conference belonging to the same conference type; The sub-conferences are divided based on the conference types of the sub-conferences and the conference contents of the sub-conferences.

10. The method according to any one of claims 1 to 8, wherein: The dividing the network conference based on the conference type and the conference content of the network conference includes: The network conference is divided based on a segmentation rule corresponding to the conference type and the conference content of the network conference.

11. The method according to any one of claims 1 to 8, wherein: The network conference is divided based on the conference type and the conference content of the network conference, including: The conference data of the network conference is processed using the segmentation model corresponding to the conference type, wherein the segmentation model is used to divide the network conference into sections.

12. The method according to any one of claims 1 to 8, wherein: The method further comprises: determining semantics of each of said conference segments; The conference segments in the network conference that are adjacent in time period and whose semantic similarity is greater than or equal to a similarity threshold are merged.

13. A method for dividing media content, comprising: Acquire data of media content, wherein the media content includes content of at least two content dimensions; Determining a target segmentation point of the media content based on the content dimension and the data of the media content; Sub-content of the media content is determined based on the target segmentation point.

14. The method according to claim 13, wherein: The determining, based on the content dimension and the data of the media content, a target segmentation point of the media content includes: Determining candidate segmentation points of the media content under the content dimension based on a division method corresponding to the content dimension and data of the media content; Based on the candidate segmentation points, a target segmentation point of the media content is determined.

15. The method according to claim 14, wherein: The determining, based on the candidate segmentation points, a target segmentation point of the media content includes: Based on the segmentation confidence of the candidate segmentation point, a target segmentation point of the media content is determined, and the segmentation confidence is used to measure the accuracy of the segmentation of the candidate segmentation point.

16. The method according to claim 15, wherein: The segmentation confidence is determined based on a division method of determining the candidate segmentation points.

17. The method according to claim 15, wherein: The segmentation confidence is determined based on the segmentation granularity of the candidate segmentation point, where the segmentation granularity indicates a refinement degree of the media content divided by the candidate segmentation point.

18. The method according to claim 13, wherein: The media content includes voice content, the content dimension includes a voice content dimension, and the first segmentation point is determined in the following manner for the voice content dimension, where the first segmentation point is used to determine the target segmentation point: Based on the voice data, obtaining the speech content of the person with the process guide identity included in the media content; Based on different process stages indicated by the speech content, a first segmentation point of the media content is determined.

19. The method according to claim 13, wherein: The media content includes shared content.

20. The method according to claim 19, wherein: The shared content includes shared screen content, the content dimension includes a shared screen content dimension, and the second segmentation point is determined in the following manner for the shared screen content dimension, where the second segmentation point is used to determine the target segmentation point: Based on the shared screen content, determining a time to switch the content of the shared screen; Based on a time point when the content of the shared screen is switched, a second segmentation point of the media content is determined.

21. The method according to claim 19, wherein: The shared content includes shared document content, the content dimension includes a shared document content dimension, and the third segmentation point is determined in the following manner for the shared document content dimension, and the third segmentation point is used to determine the target segmentation point: Based on the shared document content, determining a switching time of the document content; Based on the switching moment of the document content, a third segmentation point of the media content is determined.

22. The method according to claim 13, wherein: The media content includes descriptive content.

23. The method according to claim 22, wherein: The description content includes stage time content, the content dimension includes stage time content dimension, and the fourth segmentation point is determined in the following manner for the stage time content dimension, and the fourth segmentation point is used to determine the target segmentation point: Based on the time indicated by the stage time content, a fourth segmentation point of the media content is determined.

24. The method according to claim 22, wherein: The description content includes stage content, the content dimension includes stage content dimension, and the fifth segmentation point is determined in the following manner for the stage content dimension, and the fifth segmentation point is used to determine the target segmentation point: The media content is clustered based on the stage content, and the fifth segmentation point of the media content is determined based on different clustered contents.

25. The method according to any one of claims 13 to 24, wherein: The media content is conference content, and the sub-content of the media content is conference segment content.

26. A method for generating a summary, comprising: Acquire content data of each sub-content included in the media content, wherein the sub-content is obtained by dividing the media content; Determining a summary of the sub-content based on the content data of each of the sub-contents; The summaries of the sub-contents are merged based on the weights of the sub-contents to obtain a summary of the media content, wherein the weights of the sub-contents are used to indicate the importance of the sub-contents in the media content.

27. The method according to claim 26, wherein: The weight of the sub-content is determined according to content information of the sub-content, and the content information is determined based on time information of the sub-content in the media content and content data of the sub-content.

28. The method according to claim 27, wherein: The weight of the sub-content is determined based on the sub-weight corresponding to the sub-information included in the content information of the sub-content.

29. The method according to claim 28, wherein: The content information includes one or more of the following sub-information: The duration of the sub-content, the number of characters involved in the sub-content, and the position of the time period of the sub-content in the time period of the media content.

30. The method of claim 27, wherein: The weight of the sub-content is determined based on an artificial intelligence model, and the artificial intelligence model is used to output the weight based on input content information.

31. The method of claim 26, wherein: The step of fusing the summaries of the sub-contents based on the weights of the sub-contents to obtain the summary of the media content includes: The weights of the sub-contents and the summaries of the sub-contents are processed based on the first language processing model to obtain a summary of the media content.

32. The method of claim 26, wherein: The content data is text data, and determining the summary of the sub-content based on the content data of each sub-content includes: The content data of each sub-content is processed respectively based on the second language processing model to obtain a summary of each sub-content.

33. The method according to any one of claims 26 to 32, wherein: The media content is the content of a meeting, and the sub-content is the content of a sub-meeting obtained by dividing the meeting, or the meeting is a recurring scheduled meeting, and the sub-content is the content of at least one scheduled meeting included in the recurring scheduled meeting.

34. The method of claim 33, wherein: The sub-conference is obtained by dividing the conference in the following way: The conference is divided into the sub-conferences based on the conference type and the conference content.

35. The method of claim 33, wherein: The sub-conference is obtained by dividing the conference in the following way: The conference is divided based on at least two content dimensions of the conference to obtain the sub-conferences.

36. A conference data processing device, comprising: A first acquisition unit is configured to acquire conference data of the network conference, wherein the conference data includes voice data of the network conference; A first determining unit is configured to determine a conference type of the network conference based on the conference data; and The first division unit is configured to divide the network conference based on the conference type and the conference content of the network conference to obtain conference segments of the network conference.

37. A device for dividing media content, comprising: A second acquisition unit is configured to acquire data of media content, wherein the media content includes content of at least two content dimensions; A second determination unit is configured to determine a target segmentation point of the media content based on the content dimension and the data of the media content; and The second segmentation unit is configured to determine the sub-content of the media content based on the target segmentation point.

38. A summary generation device, comprising: A third acquisition unit is configured to acquire content data of each sub-content included in the media content, wherein the sub-content is obtained by dividing the media content; A third determining unit configured to determine a summary of the sub-content based on content data of each of the sub-contents; and The generating unit is configured to merge the summaries of the sub-contents based on the weights of the sub-contents to obtain the summary of the media content, wherein the weights of the sub-contents are used to indicate the importance of the sub-contents in the media content.

39. An electronic device comprising: at least one processor; as well as A storage device having at least one program stored thereon, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the conference data processing method described in any one of claims 1 to 12, the media content division method described in any one of claims 13 to 25, or the summary generation method described in any one of claims 26 to 35.

40. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method for processing conference data according to any one of claims 1 to 12, the method for dividing media content according to any one of claims 13 to 25, or the method for generating a summary according to any one of claims 26 to 35 is implemented.

Citation Information

Patent Citations

  • Text abstract generation method and apparatus

    CN106021226A

  • Conference segmentation based on conversational dynamics

    CN107211058A

  • Video segmentation method and video segmentation device

    CN111918145A

  • Audio and video processing method and device, computer equipment and storage medium

    CN113096687A

  • Online conference implementation method and device, electronic equipment and storage medium

    CN116319697A