Media navigation
The method addresses inefficiencies in navigating media data by using semantically linked text to automatically identify relevant segments, enhancing user access to specific parts of media content through vector representations and machine learning models.
Patent Information
- Application Number
- GB2023006454
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-05-02
- Publication Date
- 2025-06-25
AI Technical Summary
Existing methods for navigating media data, such as videos and audio recordings, are inefficient and impractical, especially for lengthy content, as they often require manual labeling or are domain-specific and computationally complex, failing to provide personalized and efficient access to relevant segments.
A computer-implemented method that utilizes a semantically linked text to media data, employing vector representations and machine learning models to automatically navigate to relevant segments of media based on user-selected text portions, enabling efficient access to specific parts of media content.
Enables efficient navigation to relevant segments of media data by leveraging related text documents, improving user experience and reducing the time required to find specific information in lengthy recordings.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Field The present disclosure relates to a computer-implemented methods of navigating media data. The present disclosure also relates to corresponding systems and computer-readable storage media. Background Enormous amounts of audio and video data (referred to herein generally as “media data” or “media”) is generated each day. In addition to professionally produced media (i.e. television and radio programs, films, podcasts and the like), the readily available nature of modern digital recording equipment has enabled high-quality recordings in a variety of other settings. This may include recordings captured for the purpose of online streaming, videos uploaded to video sharing platforms (e.g. YouTube®), and the use of recording equipment in organisations to document internal meetings. This proliferation of media presents challenges in effectively searching the generated media, to find relevant parts within it. One approach to addressing this issue is to manually label the media with bookmarks or chapters (e.g. such as the chapters in YouTube® videos). This is of course time consuming, and may be impractical for lengthy videos, and does not allow personalisation. Other approaches may involve automatically detecting salient parts of the media data - for example for the purpose of automatically generating highlights from videos of sporting events. Such approaches may take cues from increased volume (e.g. crowd volume or laughter), key words or phrases, or apply computer vision techniques to analyse the videos to extract salient portions thereof. These approaches may disadvantageously be domain-specific (e.g. for sporting events only), and computationally complex. It is an aim of the present disclosure to address these issues, and any other issues that would be apparent to the skilled reader from the disclosure herein. Summary Examples of the present disclosure addresses some of the issues discussed above by providing a means of automatically navigating to relevant segments of media data based on a selection of a part of a text related to the media data. The related text is notably not a transcript of the media data file, but instead is semantically linked to the content of the media file. For example, the media file may be a recording of a court hearing, and the related text a judgement issued following the hearing. Another example is a recording of a meeting and meeting notes or an agenda of the meeting. According to a first aspect of the disclosure, there is provided a computer-implemented method, comprising: receiving media data comprising at least an audio track; receiving a transcription of the audio track, wherein the transcription comprises a plurality of transcription segments, each transcription segment associated with timing metadata representative of a time of the transcription segment in the media data; receiving, via a user interface, a user selection of a portion of text data, wherein the text data is semantically linked to the media data; determining a transcription segment corresponding to the selected portion of the text data; outputting via the user interface, based on the timing metadata associated with the determined transcription segment, a part of the media data corresponding to the determined transcription segment. The media data may comprise video data and the audio track. Outputting via the user interface may comprise displaying the part of the media data. The media data may comprise only the audio track. The text data may not be a transcript of the audio track. The text data and media data may be in different linguistic registers. The text data may at least partially summarise the audio track. The audio track may be a recording of legal proceedings and the text data may be a text issued in legal proceedings. The audio track may be a recording of a lecture and the text data may be lecture notes. The audio track may be a recording of a meeting and the text data may be meeting notes. The audio track may be a presentation at an academic conference and the text data may be conference proceedings. The audio track may be a recording of a discussion in a healthcare setting and the text data may be a treatment plan. The audio track may be a speech and the text data may be a news article related to the speech. The method may comprise obtaining the transcription from a speech-to-text service. The method may comprise generating the transcription using a locally-stored speech-to-text model. The method may comprise segmenting the transcription to generate the plurality of transcription segments. A boundary between consecutive transcription segments may comprise a change in speaker in the audio track. Each transcription segment may comprise an utterance or sentence. The method may comprise: generating a first vector from the selected portion of the text data, and generating a plurality of second vectors from the transcription segments. The method may comprise determining the transcription segment corresponding to the selected portion of the text data based on the first vector and plurality of second vectors. The method may comprise determining a similarity between the first vector and each of the plurality of second vectors, and determining the second vector with a greatest similarity as the transcription segment corresponding to the selected portion of the text data. Generating the first vector and / or plurality of second vectors may comprise inputting the selected portion of the text data and the transcription segments to a word embedding model and receiving the first vector and / or plurality of second vectors from the word embedding model. The word embedding model may be one of: a GloVe embedding model; a large language model, suitably GPT-3, embedding model; an entailment model; an asymmetric similarity sentence embedding model. The word embedding model may be a domain-specific embedding model. Generating the first vector and / or plurality of second vectors may comprise inputting the selected portion of the text data and the transcription segments to a recurrent neural network (RNN) and receiving the first vector and / or plurality of second vectors as output from a hidden layer, suitably the final hidden layer, of the RNN. Generating the first vector and / or plurality of second vectors may comprise generating frequency vectors, suitably tf-idf (term frequency inverse document frequency) vectors. Determining the similarity may comprise determining a cosine similarity or a dot product between the first vector and each of the plurality of second vectors. The method may comprise constructing a prompt for a large language model comprising the selected portion of the text data and at least some of the transcription segments, inputting the prompt to the large language model and determining the transcription segment corresponding to the selected portion of the text data based on an output of the large language model. The method may comprise inputting the selected portion of the text data and at least some of the transcription segments to a trained machine learning model, and determining the transcription segment corresponding to the selected portion of the text data based on receiving output from the trained machine learning model. The method may comprise displaying a media player for playing the media data. Displaying the part of the media data corresponding to the determined transcription segment may comprise seeking the media player based on the timing metadata associated with the determined transcription segment. The method may comprise determining a plurality of transcription segments corresponding to the selected portion of the text data. The method may comprise displaying a media player for playing the media data, wherein the media player comprises a progress bar. The method may comprise displaying a plurality of indicators of the progress bar, each indicator corresponding to a respective part of the media data corresponding to a respective one of the plurality of transcription segments. The method may comprise, in response to receiving user input selecting one of the plurality of indicators, seeking the media player to the respective part of the media data. According to a second aspect of the disclosure, there is provided a computer system, comprising: a processor; a user interface comprising: a display, and user input means configured to detect user input; a storage storing: media data comprising at least an audio track; a plurality of transcription segments corresponding to the audio track, each transcription segment associated with timing metadata representative of a time of the transcription segment in the media data; and text data related to the media data; and computer-readable instructions, which when executed by the processor, cause the device to: render, on the display, at least some of the text data; receive a selection of a section of the text data via the user input means; determine a transcription segment corresponding to the selected section of the text data; output on the user interface, based on the timing metadata associated with the determined transcription segment, a part of the media data corresponding to the determined transcription segment. Further optional features of the second aspect are as defined above in relation to the first aspect, and may be combined in any combination. The disclosure also extends to a computer-readable medium storing instructions that, when executed by the processor, cause the processor to carry out any of the methods discussed herein. The disclosure extends to a computer system comprising a processor and a memory, the memory storing instructions that, when executed by the processor, cause the processor to carry out any of the methods discussed herein. Brief description of the drawings For a better understanding of the present invention and to show how the same may be carried into effect, reference will now be made by way of example only to the accompanying drawings, in which: Figure 1 is a schematic block diagram of an example system. Figure 2 is a schematic flowchart illustrating a process of determining a related segment. Figure 3 is an example user interface. Figure 4 shows a media player of the example user interface of Figure 3 in more detail. Figure 5 is a schematic flowchart of an example method. Figure 6 is a bar graph illustrating average precisions of example models. Figure 7 is a schematic illustration of a selected text portion and corresponding segment. In the drawings, corresponding reference characters indicate corresponding components. The skilled person will appreciate that elements in the figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. For example, the dimensions of some of the elements in the figures may be exaggerated relative to other elements to help to improve understanding of various example embodiments. Also, common but well-understood elements that are useful or necessary in a commercially feasible embodiment are often not depicted in order to facilitate a less obstructed view of these various example embodiments. Detailed Description In overview, examples of the disclosure provide techniques for navigating media (videos or audio), based on a related text. The user selects a section of the text, and seeks to a segment of the media based on the selected portion of the text. This provides an efficient means of automatically navigating lengthy and unlabelled video data to find relevant parts of the content. This is particularly helpful when dealing with very long videos such as long court hearings spanning several days, where manually seeking through the video to find a relevant part thereof is an arduous task. Figure 1 is a schematic illustration of an example computer system 100. The system 100 comprises a controller 101 and storage 102. The controller 101 comprises one or more processors or other compute elements such as central processing units (CPUs), graphics processing units (GPUs) or field-programmable gate arrays (FPGAs). The storage 102 is configured to store transiently or permanently any suitable data required for the operation of the system, including machine readable instructions which when executed cause the controller 101 to carry out the methods discussed herein. The storage 102 may comprise any combination of random-access memory, read-only memory and storage devices such as solid-state drives or hard disk drives. The system 100 receives media data 111 and text data 112. The media data 111 may take the form of a suitable video including an audio track. The video comprises a plurality of still frames that are sequentially displayed to form a moving image. Video in this context includes footage captured by cameras and computer-generated imagery. Alternatively, the media data may be audio data alone. Either way, as will be apparent from the discussion herein, the audio includes speech data. In other words, the media data comprises a recording of utterances spoken by one or more speakers. The media data is encoded using a suitable encoding format. Example video encoding formats include: H.264, HEVC, VP9. Example audio encoding formats include: MP3, AAC, AC-3. In examples where the media data is a video file, it may include video and audio data, contained in a suitable container file format. Example container file formats include: MPEG-4 (Moving Pictures Expert Group-4), HTTP Live Streaming (HLS), Audio Video Interleave (AVI), WebM. The text data 112 may employ any suitable encoding scheme and / or file format. Examples include ASCII, UTF-8, RTF, .docx, .txt, .doc and the like. The media data 111 and text data 112 are related. In other words, the media data 111 and text data 112 share some semantic link. A wide variety of examples of media data 111 and related text data 112 are within the scope of the present disclosure. In one example, which forms the basis of the discussion below, the media data 111 is a recording of legal proceedings (e.g. a court hearing), and the text data 112 is a text issued as part of the legal proceedings (e.g. a judgement). This is a particularly prescient example, in that it illustrates the differences in length between the media data 111 and text data 112. For example, the media data 111 may be a recording of a hearing lasting several days (i.e. tens of hours), whereas the resulting judgement in the text data 112 may only amount a few hundred words. In general, the media data 111 and text data 112 are in different registers. A “register” in linguistics refers to a variety of language used for a particular purpose or in a particular communicative situation. For example, the register of spoken language in a court hearing is different to the register of a written judgement. Consequently, it is not the case that the text data 112 is a transcription of the media data 111. This does not exclude the possibility that in some circumstances the text data 112 includes parts that are transcriptions of some portion of the media data - for example a judgement that includes excerpts of a transcript of the hearing. Other examples of related media data 111 and text data 112 include: • lectures and accompanying lecture notes; • meetings (e.g. business meetings) and accompanying meeting notes; • presentations at academic conferences and related published conference proceedings; • discussions in a healthcare setting (e.g. a consultation with a medical professional) and a treatment plan resulting from the consultation; • a speech (e.g. by a politician) and a news article covering the speech; • financial earnings calls and corresponding press releases; • interviews and summaries thereof • presentations and related published content such as articles, blogs and slideshows; • e-learning video or audio content and associated written content. These examples are intended to be merely illustrative of the breadth of applications of the techniques discussed herein, and are not in any sense an exhaustive list. The system 100 comprises a transcription module 120, configured to generate a transcription of speech comprised in the media data 111. In one example, the media data 111 (or the audio thereof in the case of a video) is transmitted to a suitable cloud speech-to-text service 20. Examples of such services include those provided by Amazon Web Services, Google Cloud, Microsoft Azure, Otter.ai, OpenAI Whisper or IBM Watson. In response to the transmission of the audio data, the system 100 receives corresponding text data comprising the transcription. In other examples, the speech-to-text conversion may be carried out locally to generate the transcription. For example, the system 100 comprises a suitable trained machine learning model stored in storage 102, configured to receive the audio and generate the transcription. It will be appreciated that for many languages, pre-trained speech-to-text models are readily available, for example as part of open-source libraries. The transcription is segmented by the module 120 into a plurality of segments 121. For example, each segment 121 may correspond to the contiguous speech of a particular speaker. That is to say, the boundary between each segment 121 effectively corresponds to a change in speaker. However, in other examples, each segment 121 may correspond to a smaller part of the speech of a particular speaker. For example, the boundary between each segment can be a pause in speech of the speaker. The service 20 may be configured to carry out this step, or the module 120 may apply some other logic (e.g. breaking on punctuation included in the transcription). In other examples, the segments 121 may correspond to particular units of the transcription. The segments 121 may for example each correspond to an independently meaningful portion of the transcription. In other words, each segment corresponds to an utterance, or a whole sentence. There may be upper and / or lower bounds applied to the length of the segments 121, (e.g. a maximum and minimum number of words) to avoid excessively long or short segments 121. Beyond the actual text of the transcription itself, the transcription module 120 is also configured to generate timing metadata, linking the segments 121 of the transcription to the time at which they occur in the audio track of the media file 111. This comprises at least a start time of the segment 121 and may also comprise either an end time or a duration. Again, this may be provided by the service 20, or readily calculated from data provided by the service 20. For example, it may be the case that the service 20 provides timing data for each word, from which the timing metadata for the segment 121 can be calculated. Once generated, the segments 121 and the timing metadata are stored in the storage 102. The received media data 111 and text data 112 are also stored in the storage 102. The system 100 also comprises a suitable user interface 103. This may for example be a web interface, accessible by a user U over the internet or another suitable network connection, for example on a client device such as a personal computer, laptop computer, tablet computer, or mobile device. The interface 103 may take the form of an embedded application (i.e. an “app”) installed or accessible from the client device. Alternatively or additionally, the user interface 103 may comprise display devices (e.g. monitors, touch screens etc) and input means such as mice, keyboards, touch screens and the like. The interface 103 may also include audio output means, including speakers or headphones connectable to the system 100. These may form part of the system 100 or the client device. In general, the system 100 is configured to render data for display to the user U via the user interface 103 and receive input from a user via the user interface 103. The user interface 103 is configured to output the media data 111. For example, the user interface 103 may include a media player (e.g. a video player or audio player). The media player may include suitable controls for controlling the playback of the media data 111 - e.g. including play / pause, fast forward and rewind buttons and a progress bar etc. The user interface 103 is also configured to receive a selection of a section of the text data 112. Accordingly, the user interface 103 displays the text data 112, along with input means that allow a selection of a portion thereof. This may comprise detecting a user clicking on a part of the text (e.g. a paragraph) or highlighting a part of the text. It may also comprise interacting with some other user interface element (e.g. a button positioned beside each portion) to make a selection. In some circumstances, not all of the text data 112 is displayed concurrently. For example, the interface 103 may include some means of scrolling or navigating the text, which may be particularly appropriate for longer texts where it is impractical or undesirable to display all of the text data 112 at once. An example of a user interface of the system 100 will be discussed in more detail hereinbelow. The system 100 furthermore comprises a segment determination module 130. The segment determination module 130 is configured to receive the selected portion of the text data 112 and determine at least one segment 121 of the transcription related to the selected portion of the text data 112. In this context, the segment 121 and portion are related in the sense that the text portion has a semantic link or relationship to the segment 121 (and thus the corresponding part of the media data 111). For example, consider the above-discussed situation in which the case that the text data 112 is a judgement and the media data 111 is a recording of a court hearing. In this example, the portion of the text data 112 may correspond to a part (e.g. a paragraph) of the judgement, and the segment 121 may be a part of the hearing in which arguments were presented in relation to the issue decided in the selected paragraph. Figure 2 schematically illustrates techniques for determining the related segment 121 of the transcription based on the selected portion of the text. In the techniques, each segment 121 is converted into a respective vector representation 121 v. Similarly, the selected text portion 112a is converted into a vector representation 111v. The vectors 121 v and 111v are in the same vector space - i.e. they represent a common embedding of the input segment 121 and text 112a. Subsequently, as indicated by box 202, a comparison is made between the vectors 121 v and the vector 112v. The comparison measures the proximity of the two vectors 121v, 112v in multidimensional vector space. In other words, each vector 121v is compared to the vector 112v, to determine the relevance of the respective segments 121 to the text portion 112a. Subsequently, as indicated by box 204, the top k most relevant vectors 121v - or rather the corresponding segments 121 from which the vectors 121v are derived - are returned. The conversion to vectors 121v / 111v may be carried out by employing a pretrained word embedding model. In a first example, the pretrained GloVe embeddings are used (https: / / nip.stanford.edu / projects / gtave / ). For each word in a given segment 121, a word embedding vector is obtained using the GloVe model. A vector representative of the whole of the segment (i.e. a vector 121v) is then generated based on the word vectors. For example, each entry in the segment vector may be the mean of the corresponding entries in the word vectors. A corresponding process is carried out to generate the vector 111v. The respective vectors are then compared, to generate a metric representative of their similarity. In this example, cosine similarity is employed. Instead of the mean, each entry in the segment vector may be the minimum, maximum, median or some other value calculated from the corresponding entries in the word vectors. It may also be the case that multiple vectors are calculated - e.g. one from the mean, one from the minimum values and one from the maximum values. Respective vectors are compared (i.e. mean with mean, minimum with minimum etc), and a maximum or average similarity is taken as the output similarity. In a second example, a neural network is used as a means of generating the vectors. For example, a recurrent neural network (RNN) is trained on both the segments 121 and the text data 111. To generate a vector, the segment 121 or text portion 112a is input to the RNN, and the output (i.e. the activations) of the final hidden layer of the network forms the vector. Again, cosine similarity may be used to compare the vectors. In a third example, embeddings are used from a pretrained model for entailment. Textual entailment is the task of finding sentence pair relations, i.e. one sentence entails the other or contradicts the other. There is various work in generating word embeddings that are useful for entailment. This task is conceptually similar to the present task, in terms of the potential similarity in the relationship between a pair of sentences in an entailment task and the segments 121 and the text portion 112a. Again, cosine similarity is used as the distance metric between the vectors. In one example, the entailment model is trained on the Microsoft® dataset MiniLM-L6-H384- fine-tuned on 1B sentence pairs (available at https: / / huggingface.co / blog / 1b-sentence- embeddings ). This is further discussed in Wang et al, MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers, 2020, https: / / arxiv.org / abs / 2002.10957). In a fourth example, an asymmetric similarity technique is applied. Asymmetric similarity refers to finding similarity between unequal spans of text, which may be particularly applicable to the examples of the text 112 and media 111 discussed herein. The embeddings are created from pretrained SBERT Sentence embeddings (see https: / / www.sbert.net / ). SBERT embeds a whole segment 121 or text portion 112a as a vector, rather than requiring operations to generate a vector from word vectors. In this example, the dot product of vectors is used as a distance metric. In a fifth example, a question-answer linking approach is adopted. In this example, the selected text portion 112a is treated as a question, and the segments 121 as potential answers. Pretrained embeddings obtained from OpenAI®’s GPT-3 model are used to generate the vectors 112v and 121 v. This may be the text-embedding-ada-002 model. The vectors are then compared using dot product. In other examples, a prompt is constructed for input to GPT-3’s completion endpoint, which includes the selected text portion 112a as a question, and at least some of the segments 121 as potential answers. The embeddings may be used to select a subset of relevant segments 121 to include in the prompt. In response to the prompt, GPT-3 may return an indication of the most similar segments 121. For example, GPT-3 may indicate 10 segments as correct links. In any of the examples above, the embedding models may be domain-specific. That is to say, the models may be trained or fine-tuned based on text specific to the relevant underlying domain (e.g. the legal domain). In other examples, frequency-based method may be applied instead of the embeddingbased approaches discussed above. An example technique is tf-idf (term frequency inverse document frequency), tf-idf is a measure of the relevance of an input word with respect to a document in a corpus. The tf (term frequency) part of the metric corresponds to frequency of the input word in a document. The tf may be a raw count, or a relative count, or scaled or a weighted count. The idf (inverse document frequency) part of the metric is a measure of the rarity of the input word across the corpus. The idf is a logarithmically scaled inverse fraction of the documents in the corpus containing the word. Herein, the documents of the corpus may be the segments 121 and the selectable portions 112a of the text data 112. A tf-idf vector can be calculated for each document in the corpus, wherein the vector includes a tf-idf score for every word in the corpus relative to that document. To determine similarity, the tf-idf vector of the selected portion 112v is compared to the vectors 121 v of each segment 121, for example using cosine similarity. In another example, a similar approach may be employed using BM25 (see https: / / en.wikipedia.org / wiki / Okapi BM25 the contents of which are incorporated herein in their entirety), which is effectively a modified version of tf-idf. Various aspects of the techniques discussed above in relation to Figure 2 can be combined in various different ways. For example, any of the techniques above can be employed with any suitable vector similarity measures (e.g. Euclidian distance, Manhattan distance, Minkowski distance etc). Various other word embedding models may be available and employed. In some examples, a plurality of the techniques above may be employed, with the highest or average score selected. GPT-3 is an example of a suitable large language model (e.g. a pretrained transformer model) which can be used - other large language models may be employed, including future iterations of GPT such as GPT-4. It will be understood that in many examples, the vectors will be calculated and stored in advance in storage 102, such that at runtime in response to the user input selecting a text portion, only the comparison step occurs. The techniques discussed above are effectively unsupervised machine learning approaches to retrieving the relevant segments 121. In other examples, supervised machine learning approaches may be employed to predict whether a text portion 112a and segment 121 are related. For example, a model may be trained based on manually labelled text portion / segment pairs. In some examples, rather than training a model from scratch, an existing model (e.g. a pretrained large language model such as GPT-3) may be fine-tuned based on labelled data. In yet further examples, online learning techniques may be employed, in which the model is further trained after deployment based on user feedback (e.g. via the UI 103). Labelled data may also be used to select thresholds for similarity metrics for use in the unsupervised techniques. As discussed above, the techniques may generate a list of matching segments 121, which may be ranked by their relevance. The number kof matching segments 121 is a parameter of the system 100 that can be set according to preference. For example, k may equal 10. In other examples, only a single matching segment 121 may be returned. Returning now to Figure 1, the system 100 retrieves the timing data associated with the segment 121 identified by the segment (or segments) determination module 130. The media 111 is then displayed on the U1103 at the time indicated by the timing data. In other words, the system 100 navigates the media displayed on the UI 103 to the time matching the segment 121. Accordingly, the selection of a particular text portion 112a by a user U controls the navigation of the media file 111. Whilst the system 100 may carry out the process shown in Figure 2 in real time in response to user input, it is equally contemplated that these steps can be carried out in advance (i.e. “offline”). For example, the system 100 may iterate through each portion of text 112a of the text data 112. For each portion 112a, one or more related segments 121 are identified using the techniques discussed herein. The identified segments 121 associated with each portion 112a may then be stored in a suitable database or other suitable data structure in storage 102. Accordingly, at runtime, the system 100 receives a selection of a text portion 112a and accesses the database to retrieve the associated stored segments 121. Such examples may improve the responsiveness of the user interface 103. In some examples, some or all of these steps may be carried out on a different system to system 100. Figure 6 is a chart 600 illustrating the performance of some of the techniques discussed above. Particularly, the chart 600 illustrates the average precision of different models on the task of linking the judgement text of a UK Supreme Court case with segments of the video transcript of the hearing of the same case. The models are evaluated with reference to human annotation. The chart 600 illustrates the average precision of the first 5 (i.e. the 5 highest-ranked) segments returned by the technique. Precision in this context refers to the total number of transcript segments retrieved that are relevant to the judgement compared to the total number of transcript segments that are retrieved by each model. The GPT model referenced is according to the dot product comparison approach discussed above, rather than the prompt engineering approach. As the chart 600 illustrates, each of the approaches discussed herein return highly relevant results. Figure 7 illustrates a text portion 701 and a matching segment 702, from the dataset upon which the results of Figure 6 are derived. The part 701a of the text portion includes a relevant legal point - the parts 702a of the matching segment 702 reflect that legal point. Turning now to Figure 3, an example interface screen 300 rendered by UI 103 is illustrated. The example interface screen 300 pertains to the example discussed above in which the media file 111 is a video of a court hearing and the text data 112 is the judgement. The screen 300 comprises a menu 301 comprising a plurality of selectable tabs 301a-d. The tab 301c comprises the means of navigating the media file 111 discussed herein. The screen 300 comprises a first panel 302 for displaying the text data 112. The text data 112 is broken into paragraphs 112-1 to 112-3, of which paragraph 112-2 is selected as indicated by the dashed line. The screen comprises a second panel 303 including a media player 304. The media player 304 may include a progress bar 305. Upon selection of paragraph 112-2, the system 100 causes media player 304 to seek to the part of the media 111 related to the selected paragraph 112-2. The screen 300 may also display the segment 121 of the transcript corresponding to that part of the media 111, in region 306. Figure 4 illustrates an example of the media player 304 in more detail. In this example, the module 130 returns a plurality of matching segments 121. On the progress bar 305, each matching segment 121 is indicated by an indicator 311. Each of the indicators 311 may be selected. In response to the selection of an indicator 311, the media file 111 is navigated to the time corresponding to the indicator 311. A selected indicator 311a may be displayed more prominently than the other indicators 311, for example by being larger than the other indicators 311. The indicators 311 are effectively bookmarks in the video. Figure 5 is a schematic flowchart of an example method 500. The method comprises a first step S501 of receiving media data comprising at least an audio track. The method comprises a second step S502 of receiving a transcription of the audio track, wherein the transcription comprises a plurality of transcription segments, each transcription segment associated with timing metadata representative of a time of the transcription segment in the media data. The method comprises a third step S503 of receiving, via a user interface, a user selection of a portion of text data, wherein the text data is semantically linked to the media data. The method comprises a fourth step S504 of determining a transcription segment corresponding to the selected portion of the text data. The method comprises a fifth step S505 of outputting on the user interface, based on the timing metadata associated with the determined transcription segment, a part of the media data corresponding to the determined transcription segment. The method may comprise further steps, as discussed herein. Various modifications and alterations may be made to the examples discussed above, within the scope of the disclosure. Although the discussion above involves automated transcription of the input media, it will be understood that the transcription may in some examples be wholly or partly human generated. In some instances, the system may receive the transcript, possibly already segmented, as input. In the above examples, the media player seeks (i.e. skips) to the relevant part of the media. This then permits the user to carry on watching / listening to the media beyond the relevant segment, and may allow them to manually navigate to earlier in the media, for example to better understand the context of the relevant part. However, this need not be the case, and in other examples it may be that only the relevant part is displayed - for example by extracting a clip from the media corresponding to the identified segment. In addition, for examples where the media data is audio only, the outputting of the relevant part of the audio need not involve rendering data on the display, but instead may merely involve outputting the relevant audio via suitable audio output means. In some examples, not all of the text data 112 is displayed on the user interface. For example, as noted above, it may be the case that not all of the text 112 is able to fit on the user interface, and thus means of scrolling or otherwise navigating the text may be provided. In other cases, some portions of the text data 112 (e.g. those having above a threshold number of associated segments) may be preselected for display. In some examples, the user may be able to move between portions (preselected or otherwise) of the text data 112 by providing appropriate navigation input. This may comprise selecting “next” and / or “previous” buttons. Advantageously, the examples herein provide an improved means of navigating lengthy, unlabelled video and audio recordings, by leveraging related text documents that may have already been generated. It will be understood that the controller or processor or processing system or circuitry referred to herein may in practice be provided by a single chip or integrated circuit or plural chips or integrated circuits, optionally provided as a chipset, an applicationspecific integrated circuit (ASIC), field-programmable gate array (FPGA), digital signal processor (DSP), graphics processing units (GPUs), etc. The chip or chips may comprise circuitry (as well as possibly firmware) for embodying at least one or more of a data processor or processors, a digital signal processor or processors, which are configurable so as to operate in accordance with the exemplary embodiments. In this regard, the exemplary embodiments may be implemented at least in part by computer software stored in (non-transitory) memory and executable by the processor, or by hardware, or by a combination of tangibly stored software and hardware (and tangibly stored firmware Although at least some aspects of the embodiments described herein with reference to the drawings comprise computer processes performed in processing systems or processors, the invention also extends to computer programs, particularly computer programs on or in a carrier, adapted for putting the invention into practice. The program may be in the form of non-transitory source code, object code, a code intermediate source and object code such as in partially compiled form, or in any other non-transitory form suitable for use in the implementation of processes according to the invention. The carrier may be any entity or device capable of carrying the program. For example, the carrier may comprise a storage medium, such as a solid-state drive (SSD) or other semiconductor-based RAM; a ROM, for example a CD ROM or a semiconductor ROM; a magnetic recording medium, for example a floppy disk or hard disk; optical memory devices in general; etc. Terms such as ‘component’, ‘module’ or ‘unit’ used herein may include, but are not limited to, a hardware device, such as circuitry in the form of discrete or integrated components, a Field Programmable Gate Array (FPGA) or Application Specific Integrated Circuit (ASIC), which performs certain tasks or provides the associated functionality. In some examples, the described elements may be configured to reside on a tangible, persistent, addressable storage medium and may be configured to execute on one or more processors. These functional elements may in some examples include, by way of example, components, such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. Although the example examples have been described with reference to the components, modules and units discussed herein, such functional elements may be combined into fewer elements or separated into additional elements. Various combinations of optional features have been described herein, and it will be appreciated that described features may be combined in any suitable combination. In particular, the features of any one example may be combined with features of any other example, as appropriate, except where such combinations are mutually exclusive. The examples described herein are to be understood as illustrative examples of embodiments of the invention. Further embodiments and examples are envisaged. Any feature described in relation to any one example or embodiment may be used alone or in combination with other features. In addition, any feature described in relation to any one example or embodiment may also be used in combination with one or more features of any other of the examples or embodiments, or any combination of any other of the examples or embodiments. Furthermore, equivalents and modifications not described herein may also be employed within the scope of the invention, which is defined in the claims.
Claims
1. A computer-implemented method, comprising:receiving media data comprising at least an audio track;receiving a transcription of the audio track, wherein the transcription comprises a plurality of transcription segments, each transcription segment associated with timing metadata representative of a time of the transcription segment in the media data;receiving, via a user interface, a user selection of a portion of text data, wherein the text data is semantically linked to the media data;determining a transcription segment corresponding to the selected portion of the text data;outputting via the user interface, based on the timing metadata associated with the determined transcription segment, a part of the media data corresponding to the determined transcription segment.
2. The method of claim 1, wherein:the media data comprises video data and the audio track, andoutputting via the user interface comprises displaying the part of the media data.
3. The method of claim 1 or 2, wherein the text data and media data are in different linguistic registers.
4. The method of any preceding claim wherein: the audio track is a recording of legal proceedings and the text data is a text issued in legal proceedings; or the audio track is a recording of a lecture and the text data is lecture notes; or the audio track is a recording of a meeting and the text data is meeting notes; or the audio track is a presentation at an academic conference and the text data is conference proceedings; or the audio track is a recording of a discussion in a healthcare setting and the textdata is a treatment plan; or the audio track is a speech and the text data is a news article related to the speech.
5. The method of any preceding claim, wherein a boundary between consecutive transcription segments comprises a change in speaker in the audio track.
6. The method of any preceding claim, comprising:generating a first vector from the selected portion of the text data, and generating a plurality of second vectors from the transcription segments; anddetermining the transcription segment corresponding to the selected portion of the text data based on the first vector and plurality of second vectors.
7. The method of claim 6, comprising:determining a similarity between the first vector and each of the plurality of second vectors, anddetermining the second vector with a greatest similarity as the transcription segment corresponding to the selected portion of the text data.
8. The method of claim 6 or 7, wherein:generating the first vector and / or plurality of second vectors comprises inputting the selected portion of the text data and the transcription segments to a word embedding model and receiving the first vector and / or plurality of second vectors from the word embedding model.
9. The method of claim 6 or 7, wherein generating the first vector and / or plurality of second vectors comprises generating frequency vectors.
10. The method of any preceding claim, comprising:constructing a prompt for a large language model comprising the selected portion of the text data and at least some of the transcription segments;inputting the prompt to the large language model, anddetermining the transcription segment corresponding to the selected portion of the text data based on an output of the large language model.
11. The method of any preceding claim, comprising:inputting the selected portion of the text data and at least some of the transcription segments to a trained machine learning model, anddetermining the transcription segment corresponding to the selected portion of the text data based on receiving output from the trained machine learning model.
12. The method of any preceding claim, comprising:displaying a media player for playing the media data;wherein outputting the part of the media data corresponding to the determined transcription segment comprises seeking the media player based on the timing metadata associated with the determined transcription segment.
13. The method of any preceding claim, comprising:determining a plurality of transcription segments corresponding to the selected portion of the text data;displaying a media player for playing the media data, wherein the media player comprises a progress bar;displaying a plurality of indicators of the progress bar, each indicator corresponding a respective part of the media data corresponding to a respective one of the plurality of transcription segments; andin response to receiving user input selecting one of the plurality of indicators, seeking the media player to the respective part of the media data.
14. A computer system, comprising:a processor;a user interface comprising:a display, anduser input means configured to detect user input;a storage storing:media data comprising at least an audio track;a plurality of transcription segments corresponding to the audio track, each transcription segment associated with timing metadata representative of a time of the transcription segment in the media data; andtext data related to the media data; andcomputer-readable instructions, which when executed by the processor, cause the device to:render, on the display, at least some of the text data;receive a selection of a section of the text data via the user input means;determine a transcription segment corresponding to the selected section of the text data; andoutput on the user interface, based on the timing metadata associated with the determined transcription segment, a part of the media data corresponding to the determined transcription segment.
15. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to perform the method of any of claims 1 to 13.