Voice interaction information display method and device for conference summary and electronic equipment
By combining the operation information of the interactive documents and voice signals, the interactive segments are determined and displayed, which solves the problem of inaccurate interactive segmentation in the existing technology and achieves more efficient acquisition of interactive information.
Patent Information
- Application Number
- CN202511113863.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2025-10-17
AI Technical Summary
In existing technologies, the segmented recording of real-time voice interaction is not accurate enough, resulting in low efficiency for users when viewing the recorded results, requiring them to manually drag the progress bar to find the relevant parts.
By combining operation information based on interactive documents and sound signals from real-time voice interaction with speech recognition results, interaction segments are determined, segment information is displayed, and the hierarchical relationship of interaction segments is constructed.
It improves the accuracy of interactive segmentation, helps users quickly find and locate interactive content, and improves the efficiency of obtaining interactive information.
Smart Images

Figure CN120808782A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese Patent Application No. CN202211351762.7, filed on October 31, 2022, entitled "Information Display Method, Device and Electronic Equipment Based on Voice Interaction", the priority of which is hereby affirmed. TECHNICAL FIELD
[0002] The present disclosure relates to the technical field of Internet, and in particular, to a voice interaction information display method, device and electronic equipment for meeting minutes. BACKGROUND
[0003] With the development of the Internet, users use the functions of terminal devices more and more, making work and life more convenient. For example, users can start real-time voice interaction with other users online through terminal devices. Users can realize remote interaction through online real-time voice interaction, and can also start interaction without gathering together. Real-time voice interaction largely avoids the limitations of location and venue in traditional face-to-face interaction. SUMMARY
[0004] This part of the disclosure is provided to briefly introduce the concepts, which will be described in detail in the specific embodiments part. This part of the disclosure is not intended to identify key features or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0005] In a first aspect, the embodiments of the present disclosure provide a voice interaction-based information display method, which comprises: determining an interaction segment of a real-time voice interaction based on operation information of an interaction-related document of the real-time voice interaction; and displaying segment information of the determined interaction segment.
[0006] In a second aspect, the embodiments of the present disclosure provide a voice interaction-based information display method, which comprises: obtaining a voice recognition result by performing voice recognition on a voice signal period in the real-time voice interaction; determining an interaction segment of the real-time voice interaction according to the voice recognition result; and displaying segment information of the determined interaction segment.
[0007] In a third aspect, the embodiments of the present disclosure provide a voice interaction-based information display device, which comprises: a recognition module configured to obtain a voice recognition result by performing voice recognition on a voice signal period in the real-time voice interaction; a determination module configured to determine an interaction segment of the real-time voice interaction according to the voice recognition result; and a display module configured to display segment information of the determined interaction segment.
[0008] In a fourth aspect, the embodiments of the present disclosure provide an information display device based on voice interaction, comprising: a determination unit configured to determine an interaction segment of a real-time voice interaction based on operation information of an interaction-related document of the real-time voice interaction; and a display unit configured to display segment information of the determined interaction segment.
[0009] In a fifth aspect, the embodiments of the present disclosure provide an electronic device, comprising: one or more processors; and a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the information display method based on voice interaction according to the first aspect.
[0010] In a sixth aspect, the embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the information display method based on voice interaction according to the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features, aspects, and advantages of the embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals indicate the same or similar elements. It should be understood that the drawings are schematic and elements and features are not necessarily drawn to scale.
[0012] Figure 1 is a flowchart of one embodiment of the information display method based on voice interaction according to the present disclosure;
[0013] Figure 2 is a flowchart of an optional implementation according to the present disclosure;
[0014] Figure 3 is a flowchart of an optional implementation according to the present disclosure;
[0015] Figure 4 is a schematic diagram of one application scenario of the information display method based on voice interaction according to the present disclosure;
[0016] Figure 5 is a schematic diagram of one application scenario of the information display method based on voice interaction according to the present disclosure;
[0017] Figure 6 is a schematic diagram of one application scenario of the information display method based on voice interaction according to the present disclosure;
[0018] Figure 7A is a schematic diagram of one application scenario of the information display method based on voice interaction according to the present disclosure;
[0019] Figure 7Bis a schematic diagram of one application scenario of the information display method based on voice interaction according to the present disclosure;
[0020] Figure 7C is a schematic diagram of one application scenario of the information display method based on voice interaction according to the present disclosure;
[0021] Figure 8 is a flow chart of one embodiment of the information display method based on voice interaction according to the present disclosure; Figure 9 is a structural schematic diagram of one embodiment of the information display device based on voice interaction according to the present disclosure;
[0022] Figure 10 is a structural schematic diagram of one embodiment of the information display device based on voice interaction according to the present disclosure; Figure 11 is an exemplary system architecture in which the information display method based on voice interaction of one embodiment of the present disclosure can be applied;
[0023] Figure 12 is a schematic diagram of the basic structure of an electronic device provided according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are merely for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0025] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this respect.
[0026] The term "comprising" and variations thereof as used in the present disclosure are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions of other terms will be given in the description below.
[0027] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0028] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0029] The name of the message or information exchanged between the plurality of devices in the embodiments of the present disclosure is only for illustrative purposes, and is not used to limit the scope of the message or information.
[0030] Reference is made to Figure 1 , which shows a flow of one embodiment of the voice interaction based information display method according to the present disclosure. As Figure 1 shown, the voice interaction based information display method includes the following steps:
[0031] Step 101, determining the interaction segment of the real-time voice interaction based on the operation information of the interaction related document in the real-time voice interaction.
[0032] In this embodiment, the execution subject (such as a server and / or a terminal device) of the voice interaction based information display method can determine the interaction segment of the real-time voice interaction based on the operation information of the interaction related document in the real-time voice interaction. It can be understood that the real-time voice interaction can be understood as a voice interaction, and the segment of the real-time voice interaction can be called an interaction segment.
[0033] In this embodiment, the real-time voice interaction can be a voice interaction performed by an electronic device in real time, which can include, for example, an online interaction performed by a multimedia. The multimedia can include, but is not limited to, at least one of audio and video. The real-time voice interaction interface can be a related interface of the real-time voice interaction.
[0034] In this embodiment, the application that starts the real-time voice interaction can be any kind of application, which is not limited here. For example, the application can be an instant video interaction application, a communication application, a video playing application, and a mail application, etc.
[0035] Here, the interaction segment of the real-time voice interaction can be bound to the interaction time point, and the time period between two interaction time points can be taken as an interaction segment.
[0036] Here, the interaction related document can include a document related to the interaction. As an example, the interaction related document can include, but is not limited to, at least one of the following: a shared document bound to the interaction, a document displayed during the sharing of the screen. The shared document bound to the interaction can be bound to the interaction before the interaction, or can be bound to the interaction during the interaction (i.e. a document shared in the meeting).
[0037] Here, the operation information of the interaction-related document can indicate the operation on the interaction-related document.
[0038] As an example, the operation on the interaction-related document can include, but is not limited to, at least one of the following: switching the document, opening the document, closing the document, browsing the document, selecting the document title, annotating the document.
[0039] As an example, the interaction segment can be determined according to the switching operation of the user on different interaction-related documents. The time point at which the user switches between different interaction-related documents can be taken as the interaction segment.
[0040] As an example, the interaction segment can be determined according to the operation of the user switching the title of the interaction-related document. The time point of the operation of the user switching the title of the interaction-related document each time can be taken as the demarcation point of the interaction segment.
[0041] As an example, the interaction segment can be determined according to the operation of the user browsing the interaction-related document. The time point of the page turning operation on the interaction-related document can be taken as the demarcation point of the interaction segment.
[0042] Step 102: display the segment information of the determined interaction segment.
[0043] In this embodiment, the above execution subject can display the segment information of the determined interaction segment.
[0044] In this embodiment, the segment information can indicate the relevant situation of the interaction segment. The segment information can include, but is not limited to, at least one of the following: segment time, segment theme.
[0045] In this embodiment, the display position of the segment information can be determined according to the actual application scenario, which is not limited here.
[0046] As an example, the segment information can be displayed in the interaction summary area.
[0047] As an example, the segment information can include the text converted from the interaction voice.
[0048] In some embodiments, the scheme of the present application can be implemented offline, or can be performed in real time for real-time voice interaction. The segmentation of the record of the real-time multimedia conference is essentially an offline processing.
[0049] It should be noted that the information display method based on voice interaction provided in this embodiment can determine the interaction segment of real-time voice interaction based on the operation information of the interaction-related document, and can provide a new way of determining the interaction segment, so that the determined interaction segment can refer to the document display process of the interaction-related document. It can be understood that in real-time voice interaction, the participating user can carry out interaction in the form of the display process of the interaction-related document. Therefore, based on the operation information of the interaction-related document to determine the interaction segment, and display the segment information, the interaction segment that is more accurate and consistent with the real-time voice interaction process can be determined, and the accuracy of determining the interaction segment and the interaction information can be improved.
[0050] By comparison, in some related technologies, there is no good record of the interaction segment, and the user is inefficient when watching the interaction recording result, and needs to manually drag the progress bar to find the part related to himself. The interaction video segment based on the document can be implemented in the embodiment of the application, which can effectively structure the interaction and help the user to find and locate the interaction content.
[0051] In some embodiments, the above step 101 can include determining the interaction segment of the real-time voice interaction according to the operation information of the interaction-related document and the sound signal of the real-time voice interaction.
[0052] Here, the sound signal of the real-time voice interaction can be classified into different categories according to different classification criteria.
[0053] As an example, if classified according to whether it includes voice, it can include voice signal and non-voice signal; if classified according to sound intensity, the sound signal can include sound signal greater than a preset intensity threshold and sound signal not greater than the preset intensity threshold.
[0054] In some embodiments, the part of the sound signal with a sound intensity greater than the preset intensity threshold can be detected first, and then the voice signal in this part can be detected. In this way, the sound signal can be divided into a voice signal period and a period not including voice signal.
[0055] In some embodiments, the period of the sound signal not including voice signal can be used as the boundary of the interaction segment. In addition, the voice signal period in the sound signal can be segmented according to the operation information of the interaction-related document, for example, the operation of switching the interaction-related document in the voice signal period can be used as the boundary point of the voice signal period to segment the voice signal period.
[0056] It should be noted that during a real-time voice interaction, a participant may stop talking to switch topics. This period of cessation may indicate the demarcation point between segments of the real-time voice interaction. Therefore, by combining operational information about interaction-related documents and the sound signal of the real-time voice interaction to segment the real-time voice interaction, more accurate segmentation can be determined by referring to both operational information and sound signals, which can characterize the demarcation points of the interaction segments.
[0057] In some embodiments, the above steps may include determining the interaction segment of the real-time voice interaction based on the operation information of the interaction-related document and the sound signal of the real-time voice interaction. Figure 2 The process shown. Figure 2 The process shown may include step 201 , step 202 and step 203 .
[0058] Step 201 : Perform speech recognition on a speech signal period in real-time speech interaction to obtain a speech recognition result.
[0059] Here, the sound signal of the real-time voice interaction may include a voice signal. Thus, based on whether the time period includes a voice signal that lasts for a preset duration, a voice signal period can be determined from the real-time voice interaction. A threshold for interruption duration can be set in determining the duration of the preset duration. If the interruption duration between two voice signal segments is less than the threshold, the two voice signal segments can be considered continuous, with no interruption between them.
[0060] Here, voice recognition can be performed on the voice signal period in the real-time voice interaction to obtain a voice recognition result, which can include text information.
[0061] Step 202 : Segment the real-time voice interaction according to the semantic segmentation result of the voice recognition result to obtain candidate segments.
[0062] Here, the speech recognition results can be semantically segmented to obtain corresponding segments of text information. Each segmented text information can be mapped to the time point of real-time voice interaction.
[0063] In some embodiments, the above step 202 may include: performing semantic segmentation on the speech recognition results, dividing the speech recognition results into at least two segments; determining the dividing point of the real-time voice interaction segment based on the time dividing point between two adjacent speech recognition results, and obtaining two adjacent candidate segments of the real-time voice interaction.
[0064] As an example, the speech recognition result is semantically divided, and the speech recognition result is divided into two segments. Thus, the time point corresponding to the speech recognition result divided into two segments can be taken as a division point of the real-time voice interaction segmentation, and two candidate segments of the real-time voice interaction are obtained. Thus, the multimedia can be segmented, and the candidate segments are obtained.
[0065] In step 203, the candidate segments are adjusted according to the operation information of the interaction-related document, and interaction segments are obtained.
[0066] In some embodiments, at least one of the following operations can be performed according to the operation information: merging two candidate segments into one interaction segment, adjusting the time point of an existing candidate segment, and dividing one candidate segment into at least two interaction segments.
[0067] It should be noted that, by Figure 2 According to the corresponding implementation, the speech recognition result can be semantically divided to obtain candidate segments, and then the candidate segments are adjusted according to the operation information of the interaction-related document. Thus, the accuracy of the interaction segmentation can be improved.
[0068] In some embodiments, the step of determining the interaction segments of the real-time voice interaction according to the operation information of the interaction-related document and the sound signal of the real-time voice interaction can include: if the duration of a time period in which the sound signal does not include the voice signal is greater than a preset first time threshold, the part is determined as a first type of interaction segment.
[0069] As an example, the specific value of the preset first time threshold can be set according to the actual application scenario, for example, it can be 30 seconds.
[0070] In some embodiments, if the duration of a time period in which the sound signal does not include the voice signal is not greater than the preset first time threshold, the time period can be merged into the previous or subsequent time period, or the time period can be split and part of it is merged into the previous time period and part of it is merged into the subsequent time period.
[0071] It should be noted that, by judging the duration of the time period in which the sound signal does not include the voice signal, the silent period in the interaction can be accurately found. Specifically, the sound signal in the interaction can include voice signal or non-voice signal; for the time period including non-voice signal, even if the time period includes sound signal, through the division of the present implementation, the time period can not participate in the segmentation of the voice signal, thereby improving the accuracy of the interaction segmentation.
[0072] In some embodiments, the step 203 can include: determining the title switching time of the interaction-related document according to the presentation position information of the interaction-related document; and adjusting the start and end times of the candidate segments according to the title switching time.
[0073] Here, the title switching time is used to indicate the time of switching different sub-sections of the interaction-related document.
[0074] Here, the presentation position information can include time-bound document position information. The presentation position information can be the position to which the document is presented.
[0075] Here, the presentation position information can be determined in various ways.
[0076] In some embodiments, the presentation position information can be determined based on at least one of the following: title switching operation, document focus, or comment corresponding document topic information.
[0077] Here, the title switching operation can include user-triggered determination of different entries in the title, and can also include user-triggered determination of different levels of titles in the interaction-related document.
[0078] Here, the title switching time of the interaction-related document can indicate the switching time of different entries in the title. As an example, the title switching time between the first section and the second section can indicate the time at which the user switches the first section of the document to the second section.
[0079] Here, the candidate segment is adjusted in time based on the title switching time.
[0080] Here, the user can trigger a comment on the interaction-related document, and the comment corresponding document topic information can be used to determine the title switching time. For example, changing from displaying comments of the first section to displaying comments of the second section can be determined as the title switching time.
[0081] It should be noted that by adjusting the start and end times of the candidate segment based on the title switching time, the capture of the interaction focus in the interaction can be used to adjust the candidate segment, and in combination with the sound signal, the accuracy of the segmentation can be improved from both sound and visual aspects.
[0082] Please refer to Figure 3 , Figure 3 An optional implementation of step 102 is shown. Figure 3 The flow shown can include step 1011 and step 1012.
[0083] Step 1021, based on the voice signal in the real-time voice interaction and / or the document switching operation, constructing a hierarchical relationship of the interaction segments.
[0084] Step 1022, presenting the segment information with the hierarchical relationship.
[0085] Here, the document switching operation can be used to switch the interaction-related document of the real-time voice interaction. As an example, the number of the interaction-related document of the real-time voice interaction is two, numbered as a first document and a second document, the document switching operation can switch from the first document to the second document.
[0086] Here, the segment hierarchy of the interaction segment is constructed, which can display the interaction segments at different levels, embodying the relationship between the segments of the interaction segments.
[0087] For example, the number of the obtained interaction segments is three, numbered as a first segment, a second segment and a third segment; the hierarchy of the interaction segments is constructed, which determines the interaction segments as two levels, the first segment and the third segment belong to the first level, and the second segment belongs to the sub-level of the first segment; accordingly, the segment information of the first segment and the third segment is the interaction first-level title, and the segment information of the second segment is the interaction second-level title, which is under the interaction first-level title corresponding to the first segment.
[0088] Here, the display of the segment information with the hierarchical relationship can include displaying the relationship between the interaction segments in various forms. For example, the segment information of the first segment and the third segment of the first level is displayed in the top row, and the segment information of the second segment is displayed under the segment information of the first level, which is indented.
[0089] It should be noted that, by constructing the hierarchical relationship of the interaction segments based on the voice signal of the real-time voice interaction and the document switching operation, and displaying the segment information with the hierarchical relationship, the user can clearly know the hierarchical relationship between the interaction segments, and it is convenient for the user to understand the interaction structure of the real-time voice interaction.
[0090] In some embodiments, the step 1021 can include: in response to no document switching operation being detected in the real-time voice interaction, determining the interaction first-level title of the real-time voice interaction based on the document first-level title of the interaction-related document.
[0091] Here, the document directory can include multiple levels of titles, and the first-level title in the document directory can be referred to as a document first-level title.
[0092] Here, the interaction can include multiple levels of segments, as an example, the interaction first-level segment can include an interaction second-level segment, and the interaction second-level segment can include an interaction third-level segment. The interaction directory can include multiple levels of titles of the interaction, and the first-level title in the interaction directory can be referred to as an interaction first-level title, which indicates the interaction first-level segment. Each level of title in the interaction directory can indicate an interaction segment. Optionally, each level of title in the interaction can be the segment topic of the interaction segment, and the hierarchical relationship of the interaction titles is consistent with the hierarchical relationship of the interaction segments.
[0093] As an example, please refer to Figure 4 ,Figure 4 The scenario of taking the document first title of the interaction related document as the interaction first title is shown.
[0094] Figure 4 In the scenario, the playing area 401 can play the interaction video of the real-time voice interaction. The interaction first title 402 can be the document first title of the A document. The interaction second title 403 can be the document second title in the A document, and the interaction second title 403 is the second first title of the interaction first title 402, and the section 1.1 belongs to the first chapter in the A document. The interaction first title 403 can be the document first title in the A document.
[0095] It should be noted that when there is no document switching in the real-time voice interaction, i.e., there is only one interaction related document, the interaction titles and the interaction titles of the real-time voice interaction are determined according to the document titles of the interaction related document, so that the interaction process can be quickly and accurately determined in the interaction with the interaction related document as the main line.
[0096] In some embodiments, if there is document switching in the real-time voice interaction, the document identifier can be taken as the first title, and the interaction first title of the interaction related document can be taken as the second title of the interaction segment.
[0097] In some embodiments, the step 1021 comprises: in response to detecting a document switching operation in the real-time voice interaction, determining the interaction first title of the real-time voice interaction based on the document identifier of the interaction related document; and determining the interaction N-level title of the real-time voice interaction based on the document title of the interaction related document, where N≥2.
[0098] As an example, please refer to Figure 5 , Figure 5 The scenario of taking the document identifier of the interaction related document as the interaction first title is shown.
[0099] Figure 5 In the scenario, the playing area 501 can play the interaction video of the real-time voice interaction. The interaction first title 502 can be the document identifier of the A document. The interaction second title 503 can be the document first title in the A document, and the interaction second title 503 is the second first title of the interaction first title 502, and the interaction third title 504 is the second first title of the interaction second title 503, and the section 1.1 belongs to the first chapter in the A document. The interaction first title 505 can be the document identifier of the B document.
[0100] It should be noted that in the real-time voice interaction with multiple interaction-related documents, the document identifier is taken as the interaction first-level title, and the interaction period of the real-time voice interaction can be divided based on the document as the main node; the interaction N-level title (N≥2) of the real-time voice interaction is determined according to the document title, and some interaction segments corresponding to the same interaction-related document can be quickly determined by the document segment.
[0101] In some embodiments, the step 1012 can include displaying the segment information of the second-type interaction segment as the interaction first-level title.
[0102] Here, the real-time voice interaction does not show the interaction-related document during the first-type interaction segment.
[0103] For example, for an interaction segment in the real-time voice interaction without sharing the interaction-related document, and the interaction segment is greater than a preset second time threshold, the interaction segment is determined as the second-type interaction segment.
[0104] For example, please refer to the interaction first-level title 506 in Figure 5 , Figure 5 , which can indicate the second-type interaction segment. As shown in Figure 5 , the interaction first-level title 506 can be marked with the word “discussion”, indicating that the interaction segment is a discussion between the participants.
[0105] In some embodiments, the first-type interaction segment and the document identifier of the interaction-related document are at the same level.
[0106] It should be noted that the interaction segment without showing the interaction-related document is taken as the interaction first-level title, so that the user discussion segment and the interaction period of the document sharing are parallel, which can make the hierarchical relationship of the interaction segment more accurate and reasonable, and the interaction structure can be quickly determined when the interaction process needs to be reviewed.
[0107] In some embodiments, the step 1021 can include: for a target interaction period in the real-time voice interaction, determining whether the target interaction period includes a third-type interaction segment according to the voice signal in the target interaction period; if the target interaction period includes the third-type segment, adding a preset indication identifier to the interaction segment after the third-type interaction segment in the target interaction period.
[0108] In some embodiments, if the target interaction period includes a third type segment, a segment information level downgrading is performed for the interaction segment following the third type interaction segment in the target interaction period, and a preset indication mark is displayed before the downgraded segment information. Here, the interaction segments in the target interaction period correspond to the same interaction related document. As an example, the segment topic of the interaction segment in the target interaction period belongs to the same interaction related document.
[0109] As an example, the target interaction period includes three interaction segments, the first interaction segment is a communication between users, the second interaction segment is a third type interaction segment, and the third interaction segment is a discussion on the A document.
[0110] As an example, the first interaction segment has at least one of the following characteristics: at the beginning of the recording of the real-time voice interaction, multiple people frequently alternate speaking, and the voice recognition is a short sentence (<30 words / sentence), and the duration is >1 minute, then segment dotting is performed at the beginning of the video, and the segment title is a guide reading.
[0111] As an example, the second interaction segment has at least one of the following characteristics: after the first interaction segment, a >5 minute silent segment appears, which can be considered as a document reading stage (i.e., the third interaction segment). After the document reading stage ends, the interaction segment is in accordance with the document title, and a first level title "over-comment" can be added before the document title.
[0112] Here, the determination condition of the third type interaction segment includes a voice silence duration greater than a third duration threshold (e.g., 5 minutes).
[0113] As an example, please refer to Figure 6 , Figure 6 An exemplary scenario is shown in which the target interaction period includes a third type interaction segment.
[0114] In Figure 6 , the playing area 601 can play the interaction video of the real-time voice interaction. The preset indication mark 602 (e.g., indicating the over-comment word) can be used as an interaction first level title, which is displayed before the interaction second level title 603, which is the document first level title of the A document (i.e., the first chapter). The interaction third level title 604 is a sub-first level title of the interaction second level title 603, and in the A document, the 1.1 section belongs to the first chapter. The preset indication mark 605 is displayed as an interaction first level title, which is displayed before the interaction second level title 606, which is the document first level title of the A document (i.e., the second chapter).
[0115] It should be noted that by identifying the third type of interaction segment in the target interaction period, it is possible to accurately determine whether the silent period is associated with the interaction-related documents for real-time voice interaction in which participants first read the interaction-related documents and then focus on the discussion, thereby determining the main content of the user interaction after the silent period, and indicating the user interaction content after the silent period with preset indication information.
[0116] In some embodiments, the segment information may include segment topics.
[0117] The above-mentioned step 1022 may include: determining the segmentation topics of the interaction segment according to the document content of the interaction-related document in the real-time voice interaction; and displaying the segmentation topics.
[0118] Please refer to Figure 7A , Figure 7A A related scenario for displaying segment information is shown.
[0119] exist Figure 7A In the video playback area 701, the interactive video of the real-time voice interaction can be played. The document content of the interactive related document can include lunch and dinner options. The real-time voice interaction segment can include two segments. The first segment corresponds to the title or document content summary in the document (i.e., what to eat for lunch), i.e., segment title 702, and the sub-title 703 of segment title 702 (marked with noodles); the second segment corresponds to another title or document content summary in the document (i.e., what to eat for dinner), i.e., segment title 704.
[0120] Therefore, the segmented topics of the determined interaction segments can refer to the interaction-related documents. It can be understood that in real-time voice interactions with interaction-related documents, by making full use of the characteristics of the interaction-related documents and the interaction, the segmented topics of the interaction segments can be accurately determined, so that users can quickly understand the interaction process according to the segmented topics and improve the efficiency of obtaining interaction-related information.
[0121] In some embodiments, displaying the segmented topics may include: displaying the segmented topics with a hierarchical relationship in the interaction minutes.
[0122] As an example, see Figure 7A , Figure 7A An interaction minutes display area 705 is shown. In the interaction minutes display area 705, segmented topics can be displayed, and there is a hierarchical relationship between the segmented topics.
[0123] In some embodiments, the method further includes: in response to a trigger operation on the displayed segmented topic, jumping the recorded interaction video to the triggered interaction segment, and playing the triggered interaction segment.
[0124] As an example, when a user triggers a segment title 704 in the user interface 700, the playing area 701 can play the interactive segment indicated by the segment title 704. Figure 4
[0125] Thus, the user can quickly understand the interactive process by comparing the segment topics, and if the user wants to watch the segment in the real-time voice interaction, the user can trigger the segment title to quickly jump to the segment corresponding to the triggered title.
[0126] In some embodiments, the displaying the segment information of the determined interactive segment includes at least one of, but is not limited to: displaying the segment information of the determined interactive segment during the real-time voice interaction; and / or displaying the segment information of the determined interactive segment in the voice recognition result corresponding to the real-time voice interaction.
[0127] It should be noted that, during the voice interaction, displaying the interactive segment information can facilitate the user in the interaction to view the previous interactive structure in time, and facilitate the user in the interaction to recall the interactive content that has been communicated.
[0128] It should be noted that, displaying the segment information of the interactive segment in the voice recognition result can intuitively obtain the interactive structure when the user recalls the content by means of the voice recognition result, and can make the user further understand the voice recognition result by means of the interactive structure, thereby helping the user to quickly obtain the interactive content.
[0129] In some embodiments, the displaying the segment information of the determined interactive segment can include at least one of, but is not limited to: displaying the segment information corresponding to a time point on a time axis corresponding to the real-time voice interaction; displaying the segment information of the interactive segment in association with document content information; and / or displaying the segment information of the interactive segment in association with a document structure.
[0130] As an example, please refer to the playing area 701 in the user interface 700 in Figure 7B , Figure 7B which can play the interactive video of the real-time voice interaction, Figure 7B and display a time axis 706 corresponding to the real-time voice interaction. The document content of the interactive related document can include the selection of lunch and dinner. The real-time voice interaction segment can include two segments, which are what to eat for lunch starting from the 30th minute of the interaction and what to eat for dinner starting from the 60th minute of the interaction. On the time axis 706, what to eat for lunch corresponding to the 30th minute can be displayed, and what to eat for dinner corresponding to the 60th minute can be displayed.
[0131] In some embodiments, the document content information can be used to indicate the document content. As an example, the document content information can include the document body, the document title. As an example, the respective segment time corresponding to each body part can be shown in the document body.
[0132] In some embodiments, the document structure can be used to indicate the structure of the document. As an example, the document structure can interact with the structure of the document.
[0133] Please refer to Figure 7C , Figure 7C The document structure display area 707 of the document structure of the document structure display area 707 can display the document structure, which includes the noon what to eat indicating the first part of the document and the evening what to eat indicating the second part of the document. With the noon what to eat of the first part of the document, the segment time of the interactive segment can be associated with the display (i.e. 00:00-30:00); with the evening what to eat of the first part of the document, the segment time of the interactive segment can be associated with the display (i.e. 30:01-60:00).
[0134] Please refer to Figure 8 , which shows the flow of one embodiment of the voice interaction based information display method according to the present disclosure. As Figure 1 The voice interaction based information display method shown in the figure includes the following steps:
[0135] Step 801, according to the voice recognition of the voice signal period in the real-time voice interaction, the voice recognition result is obtained.
[0136] Step 802, according to the voice recognition result, the interactive segment of the real-time voice interaction is determined.
[0137] Step 803, the segment information of the determined interactive segment is displayed.
[0138] It should be noted that through Figure 8 The embodiments provided can determine the interactive segment of the real-time voice interaction according to the voice recognition result, so that the voice recognition result can be used to indicate the difference between different interactive segments, and the interactive segment can be determined. Thus, the accuracy of the interactive segment can be improved.
[0139] In some embodiments, the voice interaction based information display method includes: according to the voice recognition result, the interactive segment of the real-time voice interaction is determined, including: the voice recognition of the voice signal period in the real-time voice interaction, the voice recognition result is obtained; according to the semantic division result of the voice recognition result, the real-time voice interaction is segmented, and the candidate segment is obtained; according to the candidate segment time, the interactive segment is determined.
[0140] In some embodiments, the segmenting the real-time voice interaction according to the semantic division result of the voice recognition result to obtain candidate segments comprises: performing semantic division on the voice recognition result to divide the voice recognition result into at least two segments; determining a segment boundary of the real-time voice interaction according to a time boundary between two adjacent voice recognition results to obtain two adjacent candidate segments of the real-time voice interaction. It should be noted that Figure 8 The technical features of the corresponding embodiments can be combined with any technical features or technical solutions in other embodiments of the present application. Further reference Figure 9 , as an implementation of the method shown in the above figures, the present disclosure provides an embodiment of an information display device based on voice interaction. The device embodiment corresponds to the method embodiment shown in Figure 1 , and the device can be applied to various electronic devices.
[0141] As shown in Figure 9 , the information display device based on voice interaction in the embodiment comprises a determination unit 901 and a display unit 902. The determination unit is configured to determine interaction segments of a real-time voice interaction based on operation information of an interaction-related document of the real-time voice interaction. The display unit is configured to display segment information of the determined interaction segments.
[0142] In the embodiment, the specific processing of the determination unit 901 and the display unit 902 of the information display device based on voice interaction and the technical effects brought by the specific processing can be respectively referred to Figure 1 the related description of step 101 and step 102 in the corresponding embodiment, which will not be repeated here.
[0143] In some embodiments, the determining the interaction segments of the real-time voice interaction based on the operation information of the interaction-related document comprises: determining the interaction segments of the real-time voice interaction according to the operation information of the interaction-related document and a voice signal of the real-time voice interaction.
[0144] In some embodiments, the voice signal of the real-time voice interaction comprises a voice signal; and the determining the interaction segments of the real-time voice interaction according to the operation information of the interaction-related document and the voice signal of the real-time voice interaction comprises: performing voice recognition on a voice signal period in the real-time voice interaction to obtain a voice recognition result; segmenting the real-time voice interaction according to a semantic division result of the voice recognition result to obtain candidate segments; and adjusting a candidate segment time according to the operation information of the interaction-related document to obtain the interaction segments.
[0145] In some embodiments, the segmenting the real-time voice interaction according to the semantic division result of the voice recognition result to obtain the candidate segments comprises: performing semantic division on the voice recognition result to divide the voice recognition result into at least two segments; determining a segment boundary point of the real-time voice interaction according to a time boundary point between two adjacent voice recognition results to obtain two adjacent candidate segments of the real-time voice interaction.
[0146] In some embodiments, the adjusting the candidate segment time according to the operation information of the interaction-related document to obtain the interaction segment comprises: determining a title switching time of the interaction-related document according to presentation position information of the interaction-related document; and adjusting the start and end time of the candidate segment according to the title switching time, the title switching time being used to indicate a time of switching different subparts of the interaction-related document.
[0147] In some embodiments, the presentation position information is determined according to at least one of the following: a title switching operation, document theme information corresponding to a document focus, and document theme information corresponding to a current display comment.
[0148] In some embodiments, the adjusting the candidate segment time according to the operation information of the interaction-related document to obtain the interaction segment comprises: in response to a time interval between start time points of two candidate segments being less than a preset first time threshold, merging the two candidate segments.
[0149] In some embodiments, the determining the interaction segment of the real-time voice interaction according to the operation information of the interaction-related document and the sound signal of the real-time voice interaction comprises: if a duration of a time period in which no voice signal is included in the sound signal is greater than a preset first time threshold, determining the time period as a first type of interaction segment.
[0150] In some embodiments, the displaying the segment information of the determined interaction segment comprises: constructing a hierarchical relationship of the interaction segment based on the voice signal in the real-time voice interaction and / or a document switching operation; and displaying the segment information with the hierarchical relationship.
[0151] In some embodiments, the constructing the hierarchical relationship of the interaction segment based on the voice signal in the real-time voice interaction and / or the document switching operation comprises: in response to no document switching operation being detected in the real-time voice interaction, determining an interaction first-level title of the real-time voice interaction based on a document first-level title of the interaction-related document.
[0152] In some embodiments, the constructing the hierarchical relationship of the interaction segments based on the voice signal and / or the document switching operation in the real-time voice interaction comprises: in response to detecting a document switching operation in the real-time voice interaction, determining an interaction first-level title of the real-time voice interaction based on a document identifier of the interaction-related document; and determining an interaction N-level title of the real-time voice interaction based on a document title of the interaction-related document, where N≥2.
[0153] In some embodiments, the displaying the segmented information with the hierarchical relationship comprises: displaying the segmented information of the second type of interaction segment as an interaction first-level title, where the real-time voice interaction does not display the interaction-related document during the second type of interaction segment, and the duration of the second type of interaction segment is greater than a preset second duration threshold.
[0154] In some embodiments, the constructing the hierarchical relationship of the interaction segments based on the voice signal and / or the document switching operation in the real-time voice interaction comprises: for a target interaction period in the real-time voice interaction, determining whether a third type of interaction segment is included in the target interaction period according to a voice signal in the target interaction period, where the target interaction period corresponds to the same interaction-related document, and the determination condition of the third type of interaction segment includes that a voice silence duration is greater than a third duration threshold; if the third type of segment is included in the target interaction period, downgrading the level of the segmented information for an interaction segment after the third type of interaction segment in the target interaction period, and displaying a preset indication identifier in front of the downgraded segmented information.
[0155] In some embodiments, the segmented information includes a segmented topic; and the displaying the segmented information with the hierarchical relationship comprises: determining a segmented topic of the interaction segment according to a document content of the interaction-related document in the real-time voice interaction; and displaying the segmented topic.
[0156] In some embodiments, the displaying the segmented topic comprises: displaying the segmented topic with the hierarchical relationship in the interaction summary.
[0157] In some embodiments, the apparatus is further configured to: in response to a triggering operation on the segmented topic, jump to the triggered interaction segment in the recorded interaction video, and play the triggered interaction segment.
[0158] Further reference Figure 10 , as an implementation of the method shown in the above figures, the present disclosure provides an embodiment of an information display device based on voice interaction. The device embodiment corresponds to the method embodiment shown in Figure 8 , and the device can be applied to various electronic devices.
[0159] AsFigure 10 As shown in the figure, the information display device based on voice interaction in the embodiment includes an identification module 1001, a determination module 1002, and a display unit 1003. The identification module is configured to perform voice recognition on a voice signal period in the real-time voice interaction to obtain a voice recognition result. The determination module is configured to determine an interaction segment of the real-time voice interaction according to the voice recognition result. The display module is configured to display segment information of the determined interaction segment.
[0160] In some embodiments, the determination of the interaction segment of the real-time voice interaction according to the voice recognition result includes: performing voice recognition on a voice signal period in the real-time voice interaction to obtain a voice recognition result; segmenting the real-time voice interaction according to a semantic division result of the voice recognition result to obtain a candidate segment; and determining the interaction segment according to a candidate segment time.
[0161] In some embodiments, the segmentation of the real-time voice interaction according to the semantic division result of the voice recognition result to obtain a candidate segment includes: performing semantic division on the voice recognition result to divide the voice recognition result into at least two segments; determining a time demarcation point of the real-time voice interaction segment according to a time demarcation point between two adjacent voice recognition results to obtain two adjacent candidate segments of the real-time voice interaction. Figure 11 , Figure 11 An exemplary system architecture in which the information display method based on voice interaction of one embodiment of the present disclosure can be applied is shown.
[0162] As shown in the figure, the system architecture can include terminal devices 1101, 1102, and 1103, a network 1104, and a server 1105. The network 1104 is a medium for providing a communication link between the terminal devices 1101, 1102, and 1103 and the server 1105. The network 1104 can include various connection types, such as wired, wireless communication links, or optical fiber cables, etc. Figure 11 The terminal devices 1101, 1102, and 1103 can interact with the server 1105 through the network 1104 to receive or send messages, etc. Various client applications, such as web browser applications, search applications, news information applications, etc. can be installed on the terminal devices 1101, 1102, and 1103. The client applications in the terminal devices 1101, 1102, and 1103 can receive instructions of a user and complete corresponding functions according to the instructions of the user, such as adding corresponding information in information according to the instructions of the user.
[0163]
[0164] The terminal devices 1101, 1102, and 1103 can be hardware or software. When the terminal devices 1101, 1102, and 1103 are hardware, they can be various electronic devices having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a desktop computer, and the like. When the terminal devices 1101, 1102, and 1103 are software, they can be installed in the above-listed electronic devices. They can be implemented as a plurality of software or software modules (for example, software or software modules for providing a distributed service) or as a single software or software module. No specific limitation is made herein.
[0165] The server 1105 can be a server providing various services, for example, receiving an information acquisition request sent by the terminal device 1101, 1102, or 1103, and acquiring, in various manners, display information corresponding to the information acquisition request according to the information acquisition request, and sending data related to the display information to the terminal device 1101, 1102, or 1103.
[0166] It should be noted that the information display method based on voice interaction provided by the embodiments of the present disclosure can be executed by a terminal device, and accordingly, the information display apparatus based on voice interaction can be arranged in the terminal device 1101, 1102, or 1103. In addition, the information display method based on voice interaction provided by the embodiments of the present disclosure can also be executed by the server 1105, and accordingly, the information display apparatus based on voice interaction can be arranged in the server 1105.
[0167] It should be understood that Figure 11 The number of terminal devices, networks, and servers in
[0168] Reference is made below to Figure 12 which shows a structural schematic diagram of an electronic device (for example, a terminal device or a server in Figure 11 The terminal device in the embodiments of the present disclosure can include but is not limited to a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle-mounted terminal (for example, a vehicle-mounted navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like.Figure 12 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0169] like Figure 12 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage device 1208 into a random access memory (RAM) 1203. Various programs and data required for the operation of the electronic device are also stored in the RAM 1203. The processing device 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0170] Typically, the following devices may be connected to the I / O interface 1205: an input device 1206 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1207 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1208 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1209. The communication device 1209 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 12 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0171] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1209, or installed from the storage device 1208, or installed from the ROM 1202. When the computer program is executed by the processing device 1201, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0172] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0173] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0174] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.
[0175] The computer readable medium carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: determine an interaction segment of the real-time voice interaction based on operation information of an interaction-related document for the real-time voice interaction; and present segment information of the determined interaction segment.
[0176] In some embodiments, the determining the interaction segment of the real-time voice interaction based on the operation information of the interaction-related document comprises: determining the interaction segment of the real-time voice interaction according to the operation information of the interaction-related document and a sound signal of the real-time voice interaction.
[0177] In some embodiments, the sound signal of the real-time voice interaction comprises a voice signal; and the determining the interaction segment of the real-time voice interaction according to the operation information of the interaction-related document and the sound signal of the real-time voice interaction comprises: performing voice recognition on a voice signal period in the real-time voice interaction to obtain a voice recognition result; segmenting the real-time voice interaction according to a semantic division result of the voice recognition result to obtain a candidate segment; and adjusting a time of the candidate segment according to the operation information of the interaction-related document to obtain the interaction segment.
[0178] In some embodiments, the adjusting the time of the candidate segment according to the operation information of the interaction-related document to obtain the interaction segment comprises: determining a title switching time of the interaction-related document according to presentation position information of the interaction-related document; and adjusting start and end times of the candidate segment according to the title switching time.
[0179] In some embodiments, the presentation position information is determined according to at least one of the following: a title switching operation, document theme information corresponding to a document focus, and document theme information corresponding to a currently presented comment.
[0180] In some embodiments, the adjusting the time of the candidate segment according to the operation information of the interaction-related document to obtain the interaction segment comprises: in response to a time interval between start time points of two candidate segments being less than a preset first time threshold, merging the two candidate segments.
[0181] In some embodiments, the determining the interaction segment of the real-time voice interaction according to the operation information of the interaction-related document and the sound signal of the real-time voice interaction comprises: if a duration of a period in which no voice signal is included in the sound signal is greater than a preset first time threshold, determining the period as a first type interaction segment.
[0182] In some embodiments, the displaying the segment information of the determined interaction segment comprises: constructing a hierarchical relationship of the interaction segment based on the voice signal and / or the document switching operation in the real-time voice interaction; and displaying the segment information with the hierarchical relationship.
[0183] In some embodiments, the constructing the hierarchical relationship of the interaction segment based on the voice signal and / or the document switching operation in the real-time voice interaction comprises: in response to no document switching operation being detected in the real-time voice interaction, determining an interaction first-level title of the real-time voice interaction based on a document first-level title of the interaction-related document.
[0184] In some embodiments, the constructing the hierarchical relationship of the interaction segment based on the voice signal and / or the document switching operation in the real-time voice interaction comprises: in response to a document switching operation being detected in the real-time voice interaction, determining an interaction first-level title of the real-time voice interaction based on a document identification of the interaction-related document; and determining an interaction N-level title of the real-time voice interaction based on a document title of the interaction-related document, where N≥2.
[0185] In some embodiments, the displaying the segment information with the hierarchical relationship comprises: displaying segment information of a second type of interaction segment as an interaction first-level title, where the real-time voice interaction does not display the interaction-related document during the second type of interaction segment, and a duration of the second type of interaction segment is greater than a preset second duration threshold.
[0186] In some embodiments, the constructing the hierarchical relationship of the interaction segment based on the voice signal and / or the document switching operation in the real-time voice interaction comprises: for a target interaction period in the real-time voice interaction, determining whether a third type of interaction segment is included in the target interaction period according to a voice signal in the target interaction period, where the target interaction period corresponds to a same interaction-related document, and a determination condition of the third type of interaction segment includes a voice silence duration being greater than a third duration threshold; and if the third type of segment is included in the target interaction period, performing a segment information level downshift for an interaction segment after the third type of interaction segment in the target interaction period, and displaying a preset indication identifier in front of the downshifted segment information.
[0187] In some embodiments, the segment information comprises a segment topic; and the displaying the segment information with the hierarchical relationship comprises: determining a segment topic of the interaction segment according to a document content of the interaction-related document in the real-time voice interaction; and displaying the segment topic.
[0188] In some embodiments, the displaying the segment topic comprises: displaying the segment topic with the hierarchical relationship in an interaction summary.
[0189] In some embodiments, the electronic device is further configured to, in response to a triggering operation for the segmented topic, jump the recorded interactive video to the triggered interactive segment, and play the triggered interactive segment.
[0190] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: obtain a speech recognition result according to speech recognition on a speech signal period in the real-time voice interaction; determine an interactive segment of the real-time voice interaction according to the speech recognition result; and display segmented information of the determined interactive segment.
[0191] Computer program code for carrying out operations of the present disclosure can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0192] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a program segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in some alternative implementations, in fact be executed substantially concurrently or in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0193] The units described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the names of the units do not constitute a limitation on the units themselves. For example, the selecting unit can also be described as a unit for selecting a first type of pixel.
[0194] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that can be used include: Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0195] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0196] The above description is merely the preferred embodiments of the present disclosure and the explanation of the principles of the applied technology. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above technical features can be replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0197] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0198] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for displaying voice interaction information for conference minutes, characterized in that: include: Determining an interaction segment of the real-time voice interaction based on operation information of the interaction-related document for the real-time voice interaction; Displaying segment information for the determined interaction segment; The segment information of the interaction segment determined by presenting includes: Building a hierarchical relationship of interaction segments based on the voice signal and / or document switching operation in the real-time voice interaction; Displays segmented information with hierarchical relationships.
2. The method according to claim 1, characterized in that The determining of the interaction segmentation of the real-time voice interaction based on the operation information of the interaction-related document for the real-time voice interaction includes: An interaction segment of the real-time voice interaction is determined according to the operation information on the interaction-related document and the sound signal of the real-time voice interaction.
3. The method according to claim 2, characterized in that The sound signal of the real-time voice interaction includes a voice signal; as well as The determining, based on the operation information on the interaction-related document and the sound signal of the real-time voice interaction, the interaction segment of the real-time voice interaction includes: Performing speech recognition on a speech signal period in the real-time speech interaction to obtain a speech recognition result; Based on the semantic segmentation results of the speech recognition results, the real-time speech interaction is segmented to obtain candidate segments; The candidate segment time is adjusted according to the operation information on the interaction-related document to obtain the interaction segment.
4. The method according to claim 3, characterized in that The real-time voice interaction is segmented according to the semantic segmentation result of the voice recognition result to obtain candidate segments, including: Performing semantic segmentation on the speech recognition result, dividing the speech recognition result into at least two segments; According to the time dividing point between two adjacent speech recognition results, the dividing point of the real-time speech interaction segment is determined to obtain two adjacent candidate segments of the real-time speech interaction.
5. The method according to claim 2, characterized in that The step of adjusting the candidate segment time according to the operation information on the interaction-related document to obtain the interaction segment includes: determining a title switching time of the interaction-related document according to presentation position information of the interaction-related document, wherein the title switching time is used to indicate a time for switching between different sub-sections of the interaction-related document; According to the title switching time, the start and end times of the candidate segments are adjusted.
6. The method according to claim 5, characterized in that The demonstration position information is determined according to at least one of the following: a title switching operation, document subject information corresponding to the document focus, and document subject information corresponding to the currently displayed comment.
7. The method according to claim 3, characterized in that The step of adjusting the candidate segment time according to the operation information on the interaction-related document to obtain the interaction segment includes: In response to a time interval between start time points of two candidate segments being less than a preset first duration threshold, the two candidate segments are merged.
8. The method according to claim 2, characterized in that The determining, based on the operation information on the interaction-related document and the sound signal of the real-time voice interaction, the interaction segment of the real-time voice interaction includes: If the duration of a period in the sound signal that does not include a voice signal is greater than a preset first duration threshold, the period is determined as a first type of interaction segment.
9. The method according to claim 1, characterized in that The step of constructing a hierarchical relationship of interaction segments based on the voice signal and / or document switching operation in the real-time voice interaction includes: In response to no document switching operation being detected in the real-time voice interaction, an interaction first-level title of the real-time voice interaction is determined based on the document first-level title of the interaction-related document, wherein each level of interaction title corresponds to each level of interaction segmentation.
10. The method according to claim 1, characterized in that The step of constructing a hierarchical relationship of interaction segments based on the voice signal and / or document switching operation in the real-time voice interaction includes: In response to detecting a document switching operation in the real-time voice interaction, determining an interaction first-level title of the real-time voice interaction based on a document identifier of the interaction-related document; Based on the document titles of the interaction-related documents, determine the N-level interaction titles of the real-time voice interaction, where N is greater than or equal to 2, and each level of interaction title corresponds to each level of interaction segmentation.
11. The method according to claim 1, characterized in that The display of segmented information with a hierarchical relationship includes: The segmentation information of the second type of interaction segment is displayed as the first-level interaction title, wherein the real-time voice interaction during the second type of interaction segment does not display interaction-related documents, and the duration of the second type of interaction segment is greater than a preset second duration threshold.
12. The method according to claim 1, characterized in that The step of constructing a hierarchical relationship of interaction segments based on the voice signal and / or document switching operation in the real-time voice interaction includes: For a target interaction period in the real-time voice interaction, determining, based on a voice signal in the target interaction period, whether the target interaction period includes a third type of interaction segment, wherein segment topics of the interaction segments in the target interaction period belong to the same interaction-related document, and a determination condition for the third type of interaction segment includes a voice silence duration greater than a third duration threshold; If the target interaction period includes a third type of segment, a preset indication mark is added to the interaction segment after the third type of interaction segment in the target interaction period.
13. The method according to claim 1, wherein Segment information includes segment topics; and The display of segmented information with a hierarchical relationship includes: Determine segment topics of interaction segments based on document contents of interaction-related documents in real-time voice interaction; Present the segmented topics.
14. The method according to claim 13, characterized in that The presenting of the segmented topics includes: In the interaction minutes, segmented topics with hierarchical relationships are displayed.
15. The method according to claim 13, characterized in that The method further comprises: In response to a triggering operation on the segmented topic, the recorded interaction video is redirected to the triggered interaction segment, and the triggered interaction segment is played.
16. The method according to claim 1, wherein The segment information of the interaction segment determined by presenting the interaction segment includes at least one of the following: During real-time voice interaction, display segment information of the determined interaction segment; During the real-time voice interaction process and / or after the real-time voice interaction ends, the segmentation information of the determined interaction segment is displayed in the voice recognition result corresponding to the real-time voice interaction.
17. The method according to claim 1, wherein The segment information of the interaction segment determined by presenting the interaction segment includes at least one of the following: On the timeline corresponding to the real-time voice interaction, display the segment information corresponding to the time point; Displaying segment information of the interaction segment in association with document content information; Segment information of the interaction segment is displayed in association with the document structure, wherein the segment information includes segment time.
18. The method according to claim 1, wherein The interaction-related documents include documents shared during the real-time voice interaction process.
19. A method for displaying voice interaction information for conference minutes, characterized in that: include: Obtaining a speech recognition result based on performing speech recognition on a speech signal period in the real-time speech interaction; Determining interaction segments of the real-time voice interaction according to the voice recognition result; Displaying segment information for the determined interaction segment; The segment information of the interaction segment determined by presenting includes: Building a hierarchical relationship of interaction segments based on the voice signal and / or document switching operation in the real-time voice interaction; Display segmented information with hierarchical relationships.
20. The method according to claim 19, wherein The step of determining the interaction segment of the real-time voice interaction according to the voice recognition result includes: Performing speech recognition on a speech signal period in the real-time speech interaction to obtain a speech recognition result; Based on the semantic segmentation results of the speech recognition results, the real-time speech interaction is segmented to obtain candidate segments; The interaction segment is determined according to the candidate segment time.
21. The method according to claim 20, characterized in that The real-time voice interaction is segmented according to the semantic segmentation result of the voice recognition result to obtain candidate segments, including: Performing semantic segmentation on the speech recognition result, dividing the speech recognition result into at least two segments; According to the time dividing point between two adjacent speech recognition results, the dividing point of the real-time speech interaction segment is determined to obtain two adjacent candidate segments of the real-time speech interaction.
22. A voice interactive information display device for conference minutes, characterized in that: include: a determining unit, configured to determine an interaction segment of the real-time voice interaction based on operation information of an interaction-related document for the real-time voice interaction; A display unit, configured to display segment information of the determined interaction segment; The display unit is further configured to: construct a hierarchical relationship of interaction segments based on the voice signal and / or document switching operation in the real-time voice interaction; and display segment information with a hierarchical relationship.
23. A voice interactive information display device for conference minutes, characterized in that: include: A recognition module, configured to perform speech recognition on a speech signal period in the real-time speech interaction to obtain a speech recognition result; A determination module, configured to determine an interaction segment of the real-time voice interaction according to the voice recognition result; A display module, configured to display segment information of the determined interaction segment; The display module is further used to: construct a hierarchical relationship of interaction segments based on the voice signal and / or document switching operation in the real-time voice interaction; and display segment information with a hierarchical relationship.
24. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 21.
25. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 21 is implemented.