Method and system for editing digital audio-visual or audio recording content of an oral presentation
The method of editing digital audio-visual content by transcribing and tagging oral presentations addresses the inefficiencies in current technologies, enabling effective indexing and search, and improving the use of digital content in professional and educational contexts.
Patent Information
- Application Number
- JP2022518373
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-19
- Filing Date
- 2020-09-18
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2040-09-18
AI Technical Summary
Current technologies for editing and indexing digital audio and video content are inefficient, requiring users to search through lengthy recordings to find specific information, which hinders effective use in professional and educational settings.
A method and system for editing digital audio-visual content by transcribing oral presentations, analyzing and marking them with tags and timecodes, and generating enhanced digital files that can be indexed and searched more effectively.
This solution enables efficient editing and indexing of digital audio-visual content, reducing production costs and lead times, and facilitating the creation of high-quality digital training platforms and knowledge databases.
Smart Images

Figure 0007700104000002 
Figure 0007700104000003 
Figure 0007700104000004
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method of editing digital aurality, the purpose of which is to improve the professional and contextual use of the content of digital audio and audiovisual recordings of knowledge and communication. It also relates to an editing system for implementing this method.
Background Art
[0002] While there has been an explosive growth in online video content, it can be seen that the viewing of this content remains linear and "monolithic". Indexing of video content constitutes a list of keywords attached to the video file. Accounts of conversations such as lecture videos and conference trainings are not strictly tagged, nor are they tagged by their meaning, and are only in a rough order.
[0003] Therefore, when conducting an online search for video content related to a specific theme, it becomes necessary to examine complete videos ranging from several minutes to several hours. From such observations, the everyday operational use in work is virtually excluded.
[0004] European Patent No. 1582067 B1 discloses a method for creating a multimedia publication based on an oral presentation, which, on the one hand, starts from the information contained in previously recorded audio and / or video flows, and on the other hand, starts from supplementary data or information and references associated with the audio and / or video information, and from its structure, indexing of the audio and / or video flow, and indexing of the supplementary data and references so that the supplementary data or information and the reference documents can be displayed on a display device in parallel with the audio and / or video flow, including providing a plurality of modes of presentation of the audio and / or video flow, and / or associated supplementary data or information and reference documents, each mode of presentation resulting from a specific combination of indexing of the audio and / or video flow and indexing of the associated supplementary data or information and reference documents, the information contained in the previously recorded audio and / or video flow and the reference documents and related data being organized, on the one hand, in the time axis of the video and, on the other hand, in an information structure having the nature of a matrix with various types of data, this information structure being centered around the reference documents, these reference documents providing both (i) the associated audio and / or video sequences by indexing of the audio and / or video flow, and (ii) access to the entire reference documents by indexing of these documents.
[0005] French Patent No. 2910987 A discloses a method for constructing a knowledge database from a product for publishing an oral presentation integrating previously recorded audio and / or video flows, with additional data and reference documents being listed in chronological order by synchronization data providing selective access to fragments within the audio and / or video flow, indexed and associated in synchronization with the oral flow. This method includes - collecting synchronization for each publishing product, and - To generate global synchronization data for all publishing products, aggregating previously collected synchronization data, and - Generating a graphical interface for selectively searching within a knowledge database to input search instructions.
[0006] The global synchronization data is organized to provide a search engine for processing search instructions.
[0007] International Publication No. 2006 / 053993 A2 discloses a method for providing training in the use of software applications in information devices and / or communication devices, the software application generating a graphical user interface displayed on a display member of the information device and / or communication device, the graphical user interface - A specific graphic frame of the application including a selection area such as a selection icon and / or tab and a content area, generating a training graphical user interface displayed on the display member, and selecting a function of the software application, the training graphical user interface including - an area for transmitting an audio and / or video flow corresponding to a training lecture, and - At least a partial replication of a specific graphic frame of the software application, an area including icons and / or tabs indexed to the audio and / or video flow for controlling the training lecture, and - an area provided for displaying an operation or sequence of operations required to execute the selected function, the operation including an area that is wholly or partly pre-indexed with the audio and / or video flow.
[0008] While video has contributed to the active development of e-learning, due to the limitations as described here, the training needs have not increased so rapidly, and despite being recognized as such, it has not been deployed rapidly and effectively in a professional environment.
[0009] The object of the present invention is to provide an operational solution to this need and a means to effectively support companies and organizations aiming at a true training structure. The stakes are high. For the medium- and long-term survival of its ecosystem, the acceleration of the learning process is necessarily required.
Summary of the Invention
[0010] This object is achieved by a method of editing digital audio-visual or audio-recorded content of an oral presentation made by a speaker using presentation support, the oral presentation consisting of a sequence of words and being recorded in the form of a digital audio or audio-visual file, the presentation support being enhanced with tags.
[0011] According to the present invention, this method comprises - transcribing the oral presentation with the display of timecodes for all or part of the words - automatically analyzing the comparison of the transcription and the tagged presentation support, - replacing the timecodes from the transcription to the tagged presentation support, - identifying the tags and the timecodes of the presentation support, - marking the digital audio-visual file with the tags and timecodes to generate an enhanced digital audio-visual file.
[0012] In the first embodiment of the present invention, the editing method of the present invention implements a list of documents intended to be verbally cited during an oral presentation, and the presentation support is enhanced with tags associated with the documents before the oral presentation.
[0013] In the second embodiment of the present invention, there is no prior reference list, and tagging is done live by the presenter from the broadcast of their audio or video presentation.
[0014] It should be noted that another version of the editing method according to the present invention may be provided. In this version, the presentation support pre-enhanced with tags associated with the references is live corrected by the presenter during their audio or video presentation by adding tags to the support.
[0015] In the context of the editing method according to the present invention, it is also possible to provide a step of excluding predetermined words for which time codes are not shown during the transcription stage of the oral presentation from the time codes of all the words included in the oral presentation. These predetermined words may be, for example, pronouns or articles.
[0016] The tags marked on the digital audio-visual file are preferably composed of a code including a structural tag and information related to the citation.
[0017] Also, the editing method according to the present invention may include restoring the information included in the code and using this information.
[0018] In a specific embodiment of the present invention, the editing method further includes a phase of aggregating digital audio-visual files enhanced in a plurality of knowledge environments into a knowledge database structured to be accessible through a search engine.
[0019] Furthermore, it may include a phase indicating the traceability of copyrights or licenses attached to the content within enhanced presentation support, and / or within enhanced digital audio-visual files, and / or within a documented aurality database, and / or within the knowledge environment.
[0020] The tag identification step may further be designed to detect simple tag(s) before and / or after the title of a detailed presentation support outline, and / or tags (W) before and / or after keywords or concepts within the presentation support, although it is also possible to detect the code within the presentation support at the position of the orally cited document.
[0021] In a particular embodiment of the present invention, the editing method implements a graphical interface that provides the user of the method with digital tools for generating a digital presentation support tagged and enhanced with codes from an oral presentation, and the graphical interface includes - a list for selecting the main areas of knowledge (K), and - a table (C) of orally cited documents organized by pre-defined category data, and - a table (S) of tags for the detailed presentation support outline, and - a table (W) of tags for keywords and concepts.
[0022] The editing method may advantageously be implemented in the form of an application that can be downloaded to a computer or smartphone or any digital medium.
[0023] The step of tagging the digital audio-visual file may include emitting a signal corresponding to a simple structure tag that implements a method for using the digital audio-visual file, which is designed to transcribe the sound flow of the words corresponding to the tagged segment from the tagged audio-visual file.
[0024] This step of transmitting a signal corresponding to a simple structure tag may include transmitting the string of characters enclosed by the structure tag, and these titles corresponding to the string of characters enclosed in this way are written directly into the digital audio-visual file in the time stream corresponding to its call.
[0025] The method of using a digital audio-visual file implemented in the marking step may be designed to read the character information written in the digital audio-visual file.
[0026] The step of transmitting a signal corresponding to the enhanced tag of type (C) may advantageously include transmitting a signal containing all the information included in the code by associating it with the time code during the presenter's oral conversation in response to the identification of the code in the oral presentation.
[0027] The code (C) and the time code associated therewith are reported in a digital file provided for the structured time scheduling of the oral presentation.
[0028] The step of transmitting a signal corresponding to the keyword tag may include transmitting a signal providing the time stamp of the digital audio-visual file and transmitting the string of characters enclosed by the tag (W), and the string of characters enclosed in this way is written into the digital audio-visual file in the time stream corresponding to the call of the keyword during the oral presentation.
[0029] Also, the editing method according to the present invention may include automatically constructing a concordance index of the keywords mentioned in the oral presentation. It can be implemented from a computer designed to read the digital audio-visual recording file when the audio-visual recording of the oral presentation is completed.
[0030] In another embodiment of the present invention, the editing method is - Extracting the tags and code data included in the tagged audio-visual file; - Incrementing the tags and code data according to the time reference (time code) in the matrix table; - Generating a collection of data tables corresponding to predefined categories from the matrix table, the data tables including a structure table and a category table; - Creating an on-the-fly exchangeable database from the structure table and the category table; It may also include a phase of processing a tagged audio-visual file of an oral presentation, including:
[0031] The method may further include a phase of creating one or more knowledge environments associated with the oral presentation, the creating phase including generating a time database. This creating phase may further include editing oral and document resources extracted from the tense database.
[0032] According to another aspect of the present invention, there is proposed a system for editing digital audio-visual or audio recording content of an oral presentation made by a speaker, the oral presentation implementing presentation support and a list of documents intended to be orally cited during the oral presentation, the oral presentation including a series of words, recorded in the form of a digital audio-visual file, the presentation support being pre-enhanced with tags, - Means for transcribing the presentation by indicating a time code for all or part of the words; - Means for performing a comparative analysis of the transcription and the tagged presentation support; - Means for transferring the time code of the transcription to the tagged presentation support; - Means for identifying the tags and the time codes of the presentation support; - means for marking the digital audio - visual file with the tag and time code to generate an enhanced digital audio - visual file, characterized by including the above.
[0033] The editing system according to the present invention may further include means for the traceability of copyrights or licenses attached to the content within the enhanced presentation support and / or within the enhanced digital audio - visual file and / or within the time database of documented oral data and / or within the knowledge environment.
[0034] The editing method according to the present invention thus provides automation of the transition from the written digital support, of the structure of the main speech, the quotations and the content of the quotations, to the audio or audio - visual support that takes up the same content in the oral audio or audio - visual version, whatever the form of this explanation: story, explanation, discussion, demonstration, description, etc.
[0035] As a result, it brings about a reduction in production costs and lead times for the mass processing of audio and video content, and the possibility of large - scale automatic generation of digital training and training platforms, high - quality digital information, etc.
[0036] Also, the editing method according to the present invention provides for automatically generating a link between digital training content and the documentary resources of the organization using the training, contributing to a significant improvement in the accuracy of the survey results for consultations and navigations within audio and video media and multimedia knowledge spaces.
[0037] (Definition) Main speech: Hereinafter, the main speech means the same explanation in the spoken oral version recorded in the digital text version and video version created by the author, which is the main object of processing by the method according to the present invention. This explanation can be of any type, such as narrative, explanatory, argumentative, illustrative, descriptive, etc.
[0038] Editing: Structuring and architecting knowledge and know-how in a digital context.
[0039] Digital orality: Oral characters in the digital dimension.
[0040] Tag: A spatial marking that can be read by a digital system, such as a logo type or a tabulation form.
[0041] Knowledge environment: A major theme unit composed of edited digital orality or orality, which responds to consultations through a user interface.
Brief Description of the Drawings
[0042] Other advantages and particularities of the present invention will become apparent from embodiments and detailed descriptions that are by no means limiting, as well as the following attached drawings.
[0043]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Embodiments for Carrying Out the Invention
[0044] The embodiments described below are not in any way limiting. In particular, it is possible to consider variations of the present invention consisting only of the selection of features described separately from other described features (even if this selection is separated within a passage containing these other features), and this selection is possible when this selection of features is sufficient to technically advantageously distinguish the present invention from the prior art or to differentiate it. This selection consists of at least one preferably functional feature without structural details or with only some of the structural details, and when this part alone is sufficient to technically advantageously distinguish the present invention from the prior art or to differentiate it, it is accompanied by only this part.
[0045] Regarding different phases of an implementation example of the editing method according to the present invention, it will be described with reference to FIG. 1. Each phase P1 - P5 is associated with software LO1 - LO5.
[0046] Phase P1 is the preparation stage for an oral presentation and generates enhanced presentation support. Phase P2 corresponds to the oral presentation and its video recording. This generates an enhanced digital video file. Phase P3 corresponds to the processing of the enhanced video file and generates a time database of documented, imaged, and videoed oral data. Phase P4 corresponds to the development of a documented, videoed, or imaged oral knowledge environment. Phase P5 includes the tracing, management, and royalty payment of presenters and submitters, and in particular creates an integrated mechanism for managing royalties.
[0047] One possible option is to eliminate phase P2 by directly using the results of phase P1 from the video recording in the operating software LO3.
[0048] Another possible option is to integrate P1 and P2 and perform them on a specific remote. On this premise, for this purpose, the author tags the live presentation using LO2 configured for a specific remote mode.
[0049] As preprocessing for presentation support, the objects that the author tags are as follows. - Structural elements of the oral presentation: All titles (word sequences) of the detailed outline of the oral presentation - Quoted data or references: Elements orally quoted in the references of explanations, discussions, or descriptions. - Keywords or concepts: All words or concepts that the author considers important, specific, or relevant in the context of the presentation and that do not necessarily appear in the title's wording or character string.
[0050] To create structure tags, the character string of each title is enclosed in a frame. As an example, in presentation support, a rule could be provided to place tags (S) with a 10 - level indent before and after titles, subtitles, and sub - subtitles.
[0051] For the end - user, the structure tags provide the following. - Simple and structured access following the author's speech by the author himself - Structural stratification of knowledge to accelerate the process of understanding and assimilation - A solution that substitutes for completely transcribing the detailed account of the knowledge conversation (presumably more effective for learning).
[0052] The editing method according to the present invention involves complementary cognitive reflection of the synthetic interpretation of the conversation architecture (fixation by visual memory of the logical structure elements of the conversation), related to listening to the explanatory and argumentative speech of the outline structure.
[0053] At a technical level, this structure tag provides the following. - Industrialization of the automatic generation of an access table for ultra-fine-grained video content - Specification of the duration of each part - Improvement of reliability and relevance (the content is composed by the author or the publisher has it as a job). - Tracing of each sequence fragment of the speech. - Depending on the linearity of the author's outline, several possible negotiation keys are required, independently and / or in combination with other digital orality fragments associated with the same word, concept or document. - Creating a theme-based logical consultation chain by associating and making searchable the cited data, associating the orality fragment with its complete call context. This provides a new extended listening function that accelerates access to the audio content of important concepts.
[0054] Here, the process of creating reference tags and oral citations is explained. The first rule is code creation, where the entire character string that constitutes the document orally cited during presentation is selected from the list of orally cited references created by the author. The author associates the "C" icon corresponding to the category of that character string. By clicking on the "C" icon, a reference code is created. At the end of this step, a list (2) of "codes" is created. Using this list of codes, each code can be inserted at its (oral) citation location in the document or presentation support.
[0055] Referring to Figure 2, it should be noted that the code has already been parameterized by the parameterization of software LO1 in the knowledge field selected by the author. All the information selected by the author is associated with each code created. Therefore, the "code" category already contains all the meta-information of the references used in the context of the author's knowledge. For this reason, "K" is at least partially conditional on the data of the "C" code.
[0056] The second rule relates to the insertion of code into presentation support. From the code list, the author inserts each code at its oral citation position within the presentation support. These codes contain all the information of the cited data. These, along with the "S" tag and in addition to it, enhance the "presentation support" document. It should be noted that the "C" tag itself is in an advantageous position in terms of the logical structure of the speech due to the insertion of the "S" tag.
[0057] For the end user, this method provides the following. - Creation of hyper-contextualization of the cited data. - When referring to the knowledge environment and when listening to the conversation, there is no need to search for the cited literature. Most of the references cited can be accessed in full while watching the video and, where applicable, are traceable within the framework of a blockchain system that includes the management of copyright and / or publisher rights. - Definitively attach (comment on, explain, illustrate ideas or discussions, make analogies, etc.) to the presentation that verbally refers to the data information unit. At a technical level, this method provides the following. - Context processing of third-party data that exists in other databases and is intended to be used or utilized in the context of the construction of analysis and discussion by the user or the author. Optimize the processing of third-party data from the perspectives of relevance, traceability, and completeness, and save the labor of manual input. All that the upstream author has to do is to have only one piece of information selected from the beginning to the end throughout the editing process.
[0058] Next, an embodiment of the keyword tag will be described. This implements a framework of character strings that form keywords or key concepts. The implementation of the editing method according to the present invention provides the ability to automatically trace one or more related sentences for the understanding of words or concepts contained therein.
[0059] For the end user, this method provides the following. - Generation of a list of keywords: An index of keywords that provides targeted access to the transition of oral calls of keywords or concepts by accessing at the moment of the immediately preceding citation that is sufficiently relevant to the agreement for article creation. - The targeted agreement of digital orality is improved, thereby improving its handling and facilitating its operation and utilization in the work context.
[0060] At the technical level, the following is provided by this method. - Reinforcement of the referentiality of content by a list of keywords and concepts in addition to a universal classification system or a business classification system. - Automation of the contextualization, prioritization, and relevance of keywords or concepts that depends on the hyperdifferentiation of conversations.
[0061] A rule may be defined to default to referring to the tag "S" closest to the keyword. This ensures that the agreement of "keywords or concepts" is appropriate in its hierarchical and citation logical environment without having to manually "go back" during the listening time to find the meaning of the agreed word or concept.
[0062] Regarding the "W" type of tags, the author places the "W" tags before and after important "keywords or concepts".
[0063] This first preparatory stage P1 leads to creating support for a digital presentation that is tagged and enhanced with reinforcement codes in text or any format.
[0064] In the first option, this presentation support that is tagged and enhanced with codes can be placed on a technical device immediately before the presentation starts.
[0065] In the second option, this presentation support that is tagged and enhanced with codes is placed in software LO3 after the presentation, in combination with the digital audio-visual file of the oral presentation. The tagged and enhanced codes are transferred to the digital audio-visual file while strictly adhering to the conversational timing of the presentation.
[0066] Referring to Figure 3, the graphical user interface is given the ability to view on the same screen page a set of lists of tag categories implemented in the preparatory text P1. Thus, this graphical interface represents a list K for selecting a knowledge area, a table C of orally cited references including general category data C1, C2, C3 for the knowledge area of the presentation, a structural tag S of the detailed outline of the oral presentation, and a table W of keywords and key concepts.
[0067] Referring to FIG. 4, the graphical interface provides the user, on the one hand, with a graphical representation of a code list of an oral presentation showing various code tags C1, C2, C3 of references cited orally, and on the other hand, with a representation of a document of enhanced presentation support tagged and enhanced by the author, having a graphical connection between the C code and the position of the reference cited in the presentation support.
[0068] Phase 2 of the editing method according to the present invention consists in concentrating on an oral presentation and on recording in the form of a digital audio-visual file.
[0069] As a first option, the method according to the present invention is implemented in synchronization with an oral presentation.
[0070] In practice, this method is implemented in the form of a software application LO2 downloaded to a telephone or a personal computer. The operator is the presenter himself or an assistant. The application is launched immediately before the start of the oral presentation.
[0071] As a second option, this Phase 2 is executed asynchronously after the oral presentation. The software LO2 combines a presentation support file and a digital audio-visual recording file.
[0072] The software LO2 integrates a speech recognition software module and an automatic text reading software module and is designed to emit a signal for marking with tags or codes on a video tape or real-time digital recording (synchronous mode).
[0073] In the first option, before starting the presentation, the author inserts the tagged presentation support file into the tagging software. Speech recognition is performed in real time and character recognition is commanded in the tagged presentation support.
[0074] Subsequently, in synchronization with the oral statement of the presentation, the tags identified by the software LO2 are issued to the presentation support. The digital audio-visual file is marked with structural elements and codes, and all information of the orally cited references, quotations, and knowledge objects is integrated in the temporal unfolding.
[0075] In a second option, after the oral presentation, the video file of the digital presentation without tags is retrieved. At that time, a time code is displayed for each word.
[0076] Next, a comparative analysis of the presentation transcription and the tagged presentation support is performed. Then, the time code of the transcription is transcribed to the tagged presentation support. The tags and time codes of the presentation support are read and transferred to the presentation video file according to the time codes.
[0077] Every time the software LO2 identifies a tag or a citation code, it needs to emit a signal that can mark the digital audio-visual file recording the speaker's speech.
[0078] The emitted signal is used to tag the digital audio-visual file: - A simple time tag for the structural tag. - Or a "code" tag containing all information related to the citation.
[0079] The software LO2 can read the code and the information contained therein from the audio-visual digital file and increment it in other ways of using the content, such as a database or a structured spreadsheet document.
[0080] There are several types of signal emissions: - Emissions specific to the structure tag (S). - Transmission (C) specialized for code.
[0081] In the first option, a simple transmission of a signal enabling simple temporary marking of an audiovisual file is performed. This requires post - operation software LO3 that can transcribe the sound flow of words corresponding to the tagged segment from the tagged audiovisual file.
[0082] In the second option, transmission of a character string enclosed by tags (S) is performed: The title is directly recorded in the digital audiovisual file within the time flow corresponding to its call. It requires software for operating the digital audiovisual file that can read the enhanced tag (S) of character data, which is the text information written in the audiovisual file (this written information is not visible to the naked eye of the person watching the video).
[0083] This signal generation automatically constructs a detailed outline of the oral presentation and the listening access thereto, enabling a targeted and segmented agreement following the author's logic.
[0084] Next, the mode of generating a signal related to C code will be described.
[0085] Software LO2 outputs C code containing pre - input structured information of orally cited references with the complete text scanned by the author / presenter / publisher, associating it with the "time code" orally called during the presentation.
[0086] After identifying the code, software LO3 can emit a signal containing all the information the code includes by systematically and deterministically associating the code with its citation time code during the presenter's oral presentation. This provides a way for complete contextualization of the data cited within the context of a digital orality call.
[0087] References are, in most cases, the subject of conversations and explanations, and cognitively, the messing of this cited reference is decisive and systematic.
[0088] The results of this Phase 2 are tagged digital audio-visual files that contain all the codes related to the aforementioned citations. This tagged and enhanced digital audio-visual file is similar to a primary "database" that contains all the data to be edited. Therefore, it is in line with the first prerequisite to enable the formalization of a complete knowledge environment with structured, hierarchical, and contextualized knowledge through their classification and reorganization according to clearly defined principles and models.
[0089] The method according to the present invention thus provides for automatically constructing a concordance index of oral references and citations.
[0090] Time codes and C codes are automatically reported in a document or technical device that enables structured time scheduling of the speech and the reference content contained therein.
[0091] Regarding the signal related to the keyword tag W, in the first option, a simple transmission of a signal that enables simple temporary marking of the audio-visual file is performed. This requires the operating software LO3 that can copy the sound stream of the words corresponding to the tagged segment from the tagged audio-visual file.
[0092] As a second option, the transmission of a character string enclosed by a tag(s) is performed: keywords and key concepts are directly recorded in the digital audio-visual file within the time stream corresponding to their call. This requires the digital audio-visual file operating software LO3 that can read the character information written in the audio-visual file (this character information is not visible to the naked eye of the person watching the video). This method provides for automatically constructing a concordance index of the oral keywords.
[0093] Referring to FIG. 5 showing the spatial image of the steps of phase 2 in synchronous mode, the speech recognition system is provided to a system for reading the speech flow of the oral presentation, which is captured, tagged, and enhanced presentation support. Subsequently, the tag signal and the code of its call in the oral presentation are transmitted in real time, and the sequence of tags and codes is visualized.
[0094] Referring to FIG. 6 showing the temporal image of the steps of phase 2 in synchronous mode, as soon as the speech recognition system becomes corresponding to either a structure tag (S) or a code (C), a signal containing tagged or coded information is oscillated. The tagged information and / or the coded information are received in real time and printed on the digital audio-visual file. A definitive time code is associated with them.
[0095] Referring to FIGS. 7 and 8 respectively showing the spatial image and the temporal image of the stage of phase 2 in option 2, the video file of the oral presentation is combined with the tagged and enhanced presentation support. The audio signal of the video of the oral presentation is read, and after it is combined with the reading of the presentation support, the tag and code signals are transmitted to the presentation video file, and the tagged digital video file is distributed.
[0096] The result of this phase 2 is a tagged digital audio-visual file containing all the codes related to the reference, including the temporal mesh of the document data and the oral speech.
[0097] Continuing the description of the steps of phase 3 for processing the enhanced digital audio-visual file while referring to FIG. 9 and simultaneously with the operation of software LO3.
[0098] This phase 3 is - As shown in FIG. 10, extracting the tag and code data included in the tagged audio-visual file of the oral presentation; - In the matrix table, incrementing all data according to a time reference (time code); - As shown in FIG. 11, generating a collection of data tables corresponding to categories predefined in the preparation stage from the matrix table; - Creating an on-the-fly exchangeable database from the structure table and the category table.
[0099] The processing of the audio-visual file includes the following. - Restoring the time data and character strings of the tag(s). - Restoring the decrypted time data of the code (C), restoring the encrypted information of the code.
[0100] It should be noted that the extracted and decrypted data is associated with their time codes, transmitted and incremented in a file or a structured table specified for this purpose. For this, a blockchain may be implemented.
[0101] Also provided is to identify objective structure data specific to each field of knowledge and increment the citation (code C) according to predefined categories.
[0102] Thereafter, classification according to the hierarchy specific to these category data is performed. The data included in the code associated with the data related to the reference is used. This also applies to the data in order to restore and make accessible the entire conversation within the framework of the environment predefined at the audio and visual levels, and especially to temporarily access the targeted and always contextualized speech and citation data so as to meet the relevance requirements.
[0103] The structure table generated from the matrix table is given the task of reconstructing the detailed content table of a presentation (oral presentation) from the title list (S). By expressing it in text on the screen, one is given the opportunity to view very detailed content in line with the era of the presenter's speech.
[0104] The structure table serves as a contextual marker for other cited category data. The category data is displayed according to the citation time in a visual space separated by two titles, and any arbitrary text transcription recorded in the video is excluded.
[0105] Through a regular calculation that compares the last time code (title timeout) of each title with the title's time-in, it is possible to determine the exact time in seconds for each part that makes up the detailed content table. In terms of usability, this information is essential for enabling the user to control their own negotiation.
[0106] This structural information is used to generate the detailed content table and further functions as an access table for the sequence of the oral presentation of the end-user interface. Up to 10 levels of indentation are possible, and in both the negotiation mode and the search mode, sufficient and appropriate granularity of knowledge is achievable.
[0107] This structural information also provides the corpus composition of the keywords of the oral presentation with specific keywords indicated by the code W by the author or presenter, and allows for the contextualization of the category data (C) located within the time value interval (between two time codes) that permits recontextualization at any time, regardless of the viewing angle selected.
[0108] The category elements are used, for example, to generate an index of the data in Category 1, such as compounds. The sort index is not temporary but is specific to Category 1.
[0109] This index not only gives the ability to reference data by category logic, but also gives the ability to reference on-the-fly the data of category 2 within the context of the oral presentation that explains or calls it.
[0110] Regarding phase 4 of the development of the knowledge environment, within the framework of the editing method according to the invention, it is advantageously possible to implement the functions disclosed in French Patent No. 2849564, French Patent No. 2878345 and French Patent No. 2910987.
[0111] Phase 5 can be achieved by implementing an oral presentation and reference temporal database, the principles of navigation and cognitive design, and a library of design templates.
[0112] In Table 1, the inputs, functions, and outputs of three software modules L1, L2, L3, L4 of the types of software LO1, LO2, LO3, which are explained by non-limiting exemplary embodiments of the editing method according to the invention, are set out.
[0113] [Table 1]
[0114] The version text V2 output of software module L1 includes the following. - Presentation text - Structure tags (detailed outline tags) - Content tags (encapsulating the content information cited in the main speech)
[0115] What is important here is that when the author writes normal text or an outline, software L1 can automatically tag the text for digital support. In this case, functions such as "title", "bullet points", "numbering" etc. of any text software are linked to software L1, and tags are generated simultaneously when the function of the "title" tip is activated. The author can maintain control and add tags around the keywords.
[0116] The software module L2 indicates the time code of each word and it is explicitly stated that it is always possible to set the exclusion of words, for example, the exclusion of pronouns and articles (personal, definite, indefinite, possessive, etc.).
[0117] In the software module L3, at the end of the transfer operation of the time code from the retroscript - V1 to the text - V2, the two text versions contain similar time codes. The transfer of tags from the text - V2 to the retroscript - V1 is performed based on the correspondence of the time codes existing in the two versions as a result of the previous transfer step.
[0118] The retroscript - V2 output from the software module L3 includes the following. - The transcribed version of the audio - video version of the main speech. - Time codes arranged line - by - line, word - by - word, and paragraph - by - paragraph. - Tags (S), (C), (W).
[0119] The audio - video output V2 of the software module 4 includes the following. - The audio - video version of the main speech - Structure tags - Citation tags and their contents - Keyword tags - Other types of tags predefined at the time of setting LO2, or in remote mode, other types of tags defined when the author directly tags using LO2 used in the simple mode). - Time code (TC) specific to the audio - visual file.
[0120] The editing method according to the present invention can advantageously be used for any kind of support and format for an investigator or a document creator to conduct a cross-sectional investigation in a series of fields of knowledge. Therefore, it is possible to perform a new combination of listening to fragments of orality based solely on the criterion of relevance obtained from keyword searches and reference searches, and moreover, it is possible to systematically maintain the context specific to each oral segment listened to.
[0121] Also, the editing method according to the present invention can also be used by learners in training and can greatly contribute to the transmission of knowledge, especially knowledge from experience.
[0122] In particular, in an industrial knowledge environment that generates a large amount of technical data, the editing method according to the present invention can contribute to problem-solving in knowledge transmission. By processing a large number of oral presentations for explanation attached to technical documents, it becomes possible to accelerate operation education in a digital context and to preserve and enhance know-how, especially the know-how of experts who retire and leave the company.
[0123] As can be easily understood, the present invention is not limited to the examples described, and numerous changes can be made to these examples without departing from the scope of the present invention. Furthermore, the various features, forms, variations, and embodiments of the present invention can be grouped together in various combinations as long as they are not incompatible or mutually exclusive.
Claims
**Claim 1**: A method implemented by a computer for editing digital audio-visual or audio recording content of an oral presentation given by a speaker, using presentation support including a detailed plan of the oral presentation, wherein the oral presentation is recorded in the form of a sequence of words and a digital audio-visual file, and the presentation support is enhanced by being tagged, implementing speech recognition applied to the digital audio-visual file, transcribing the oral presentation with time code display for all or part of the words, and replacing the time code from the transcription into the tagged presentation support, and identifying the tags and the time code of the presentation support, and marking the digital audio-visual file with the identified tags and time code to generate an enhanced digital audio-visual file, comprising, wherein the tags attached to the presentation support comprise codes containing information related to structural elements of the oral presentation, cited data or references, and keywords or concepts, wherein the codes are pre-parameterized in the field of knowledge selected by the author of the oral presentation, characterized by the method. **Claim 2** Implementing a list of references intended to be orally cited during the oral presentation, and enhancing the presentation support with the tags related to the references before the oral presentation, characterized by the method according to claim 1. **Claim 3** The method according to any one of claims 1 or 2, characterized in that the presentation support is created live by the presenter during the broadcast of the presenter's oral presentation. **Claim 4** The method according to claim 3, characterized in that the creation of the presentation support comprises the addition of tags by the presenter onto the support. **Claim 5** The method according to claim 1, further comprising restoring the information contained in the code and using the information. **Claim 6** A method according to any one of claims 1 to 5, further comprising a phase (P5) of aggregating the digital audio-visual files enhanced in a plurality of knowledge environments into a knowledge database structured to be accessible through a search engine.
7. A method according to any one of claims 1 to 6, further comprising a phase (P5) of traceability of copyright or usage rights attached to the content in the enhanced presentation support and / or in the enhanced digital audio-visual files and / or in the temporal database of documented orality data.
8. The method according to claim 6, further comprising a phase (P5) of traceability of copyright or usage rights attached to the content within the knowledge environment.
9. The method according to any one of claims 1 to 8, wherein the step of identifying the tags is designed to detect simple structure tags (S) before and / or after the title of the outline of the detailed presentation support, and / or keyword tags (W) before and / or after keywords or concepts within the presentation support.
10. The method according to any one of claims 1 to 9, wherein the step of identifying the tags is designed to detect codes (C) in the presentation support at the position of the references cited orally.
11. Implement a graphical interface that provides the user of the method with a digital tool for generating a digital presentation support enhanced with tagged codes from an oral presentation, the graphical interface comprising: a list for the selection of the main areas of knowledge (K); a table of references cited orally, organized by predefined category data (C); a table of the simple structure tags (S) of the outline of the detailed presentation support; - a table of the keyword tags (W) of keywords or concepts; The method according to any one of claims 1 to 10, characterized by including the above.
12. The method according to any one of claims 1 to 11, characterized in that it is implemented in the form of an application that can be downloaded to a computer, a smartphone, or any digital medium.
13. The step of attaching the tag to the digital audio-visual file comprises transmitting a signal corresponding to a simple structure tag (S) that is designed to implement a method for using the digital audio-visual file to transcribe the sound flow of words corresponding to the tagged segment from the tagged audio-visual file. The method according to any one of claims 1 to 12.
14. The step of transmitting a signal corresponding to the simple structure tag (S) comprises transmitting a character string enclosed by the simple structure tag (S), and a title corresponding to the enclosed character string is directly written into the digital audio-visual file in a time stream corresponding to its call. The method according to claim 13.
15. The method for using the digital audio-visual file is designed to read text information written on the digital audio-visual file. The method according to any one of claims 13 or 14.
16. The step of transmitting a signal corresponding to an enhanced tag of the code (C) comprises transmitting a signal containing all the information included in the code by associating it with the time code during the presenter's oral presentation in response to the identification of the code (C) in the oral presentation. The method according to any one of claims 13 to 15.
17. The code (C) and its associated time code are reported in a digital file provided for the structured temporary configuration of the oral presentation. The method according to claim 16.
18. The step of transmitting a signal corresponding to the keyword tag (W) comprises transmitting a signal providing a time stamp of the digital audio-visual file. The method according to any one of claims 13 to 17.
19. The step of transmitting a signal corresponding to the keyword tag (W) comprises transmitting a character string surrounded by the keyword tag (W), and the surrounded character string is then written into the digital audio-visual file in a time stream corresponding to the invocation of the keyword during the oral presentation. The method according to any one of claims 13 to 18, characterized in that.
20. The method according to any one of claims 18 or 19, further comprising automatically constructing a negotiation index of the keywords mentioned in the oral presentation.
21. The method according to any one of claims 18 to 20, characterized in that it is implemented from a computer designed to read a digital audio-visual recording file when the audio-visual recording of the oral presentation is completed.
22. Further comprising a phase (P3) of processing the tagged audio-visual file of the oral presentation, the phase (P3) comprising: - extracting data of tags and codes included in the tagged audio-visual file; - incrementing the data of the tags and codes according to a time reference (time code) in a matrix table; - generating a collection of data tables corresponding to predefined categories from the matrix table, the data tables including a structure table and a category table; - creating an exchangeable database on the fly from the structure table and the category table; The method according to any one of claims 1 to 21, characterized in that it comprises.
23. Further comprising a phase (P4) of creating one or more knowledge environments related to the oral presentation, the creating phase (P4) comprising generating a time database (LO4). The method according to any one of claims 1 to 22, characterized in that.
24. The method according to claim 23, characterized in that the creating phase (P4) further comprises editing (LO5) oral and document resources extracted from the tense database.
25. The method according to any one of claims 1 to 24, further comprising the step of excluding predetermined words for which no time code is indicated during the transcription stage of the oral presentation from the time coding of all words included in the oral presentation. [
26. ] A system for editing digital audio-visual or audio recording content of an oral presentation given by a speaker, wherein the oral presentation implements presentation support comprising a detailed plan of the oral presentation and a list (L1) of references intended to be orally cited during the oral presentation, the oral presentation being recorded in the form of a sequence of words and a digital audio-visual file, and the presentation support being enhanced by being tagged, means for implementing a speech recognition method applied to the digital audio-visual file to transcribe the presentation with the display of time codes for all or some of the words; means for replacing the time codes of the transcription with the tagged presentation support; means for identifying the tags and the time codes of the presentation support; means for marking the digital audio-visual file with the identified tags and time codes so as to generate an enhanced digital audio-visual file; comprising the tags attached to the presentation support comprise codes containing information related to structural elements of the oral presentation, cited data or references, and keywords or concepts; the codes are pre-parameterized in the field of knowledge selected by the author of the oral presentation; A system characterized by the above. [
27. ] The system according to claim 26, further comprising means for the traceability of copyright or usage rights attached to the content within the enhanced presentation support and / or within the enhanced digital audio-visual file and / or within a documented temporality database of oral data and / or within a knowledge environment.
Citation Information
Patent Citations
Multimedia information viewing device
JP2000099524A
Index creation device, index creation method and index creation program
JP2004326404A
System and method for processing content and computer program
JP2005260513A
Device, method, program, and recording medium for remotely controlling application for presentation
JP2006208696A
Method and apparatus for producing a multimedia version based on oral presentation.
JP2006512845A