Multi-modal Content Switching Method, Apparatus, Device, and Storage Medium
By detecting the playback status and switching to the associated modal content, the video stuttering problem caused by extremely weak network signals is solved, and smooth playback and plot coherence are achieved in a weak network environment.
Patent Information
- Application Number
- CN202210403998.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-04-18
AI Technical Summary
In an environment where the network signal is extremely weak or the network is not available, the terminal device cannot load the video, affecting the user's video viewing experience.
By detecting the playback status and obtaining associated modal content, such as audio or text content when the switching conditions are met, and modal switching is performed, the multimodal theme model can achieve switching and coherence of different modal content.
When the network signal is poor, by switching to the associated modal content, the smoothness and user experience of video playback are improved, ensuring plot consistency.
Smart Images

Figure CN114979540B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and particularly relates to a multimodal content switching method, device, equipment and storage medium. Background Art
[0002] When a user watches film and television works such as movies and TV series on a video website using a mobile phone, the video playback is usually affected by fluctuations in the communication network environment, such as video playback stuttering and being unsmooth, which affects the user's video viewing experience. For example, when a high-speed train passes through a tunnel, in some elevators with weak signals, in the basement, etc. Currently, in order to improve the user's viewing effect, multiple videos with different resolutions are usually pre-set. When the user's network fluctuates, the user is prompted to switch to a video with a lower resolution to continue playing, so as to solve the problem that the user can continue to play the video when switching from a strong network environment to a weak network environment. However, when the user is in a scenario where the network signal is usually extremely weak or even there is no network, the terminal device cannot load the video, which further affects the user's video viewing experience. Summary of the Invention
[0003] The main purpose of the present application is to provide a multimodal content switching method, device, equipment and storage medium, aiming to solve the technical problem in the prior art that the terminal device cannot load the video due to extremely weak network signals, which affects the user experience.
[0004] To achieve the above purpose, the present application provides a multimodal content switching method, and the multimodal content switching method includes:
[0005] During the playback of the first file, detecting the playback state of the first file; the first file is an audio file or a video file;
[0006] When it is determined that the playback state meets the switching condition, obtaining the associated modal content of the first file, wherein the file type of the associated modal content is different from the file type of the first file;
[0007] Playing the associated modal content.
[0008] The present application also provides a multimodal content switching device, and the multimodal content switching device is a virtual device. The multimodal content switching device includes:
[0009] A detection module, configured to detect the playback state of the first file during the playback of the first file; the first file is an audio file or a video file;
[0010] An obtaining module, configured to obtain the associated modal content of the first file when it is determined that the playback state meets the switching condition, wherein the file type of the associated modal content is different from the file type of the first file;
[0011] A playback module for playing the associated modal content.
[0012] This application also provides a multi-modal content switching device, which is a physical device. The multi-modal content switching device includes: a memory, a processor, and a multi-modal content switching program stored on the memory. The multi-modal content switching program is executed by the processor to implement the steps of the multi-modal content switching method as described above.
[0013] This application also provides a storage medium, which is a computer-readable storage medium. A multi-modal content switching program is stored on the computer-readable storage medium. The multi-modal content switching program is executed by a processor to implement the steps of the multi-modal content switching method as described above.
[0014] This application provides a multi-modal content switching method, device, equipment, and storage medium. Compared with the prior art that can only switch between different video resolutions and the video can only be stuck when the user cannot load the video under a weak network, this application first detects the playback state of the first file during the playback of the first file; the first file is an audio file or a video file. Then, when it is determined that the playback state meets the switching condition, the associated modal content of the first file is obtained, where the file type of the associated modal content is different from the file type of the first file. Then, the associated modal content is played, realizing the switching to the corresponding associated modal content based on the playback state of the first file. That is, the currently played video or audio can be switched to other associated modal content, so that when the user watches a movie or TV show (video) and it gets stuck and the user cannot watch the video, it can also jump to other associated modal content, thereby improving the fluency of the user's reading of the work. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of the first embodiment of the multi-modal content switching method of this application;
[0018] Figure 2Schematic structural diagram for video clarity switching in this application;
[0019] Figure 3 Schematic structural diagram for modal content switching in this application;
[0020] Figure 4 Schematic flowchart of the second embodiment of the multi-modal content switching method in this application;
[0021] Figure 5 Schematic flowchart of the third embodiment of the multi-modal content switching method in this application;
[0022] Figure 6 Schematic flowchart for training a topic model in this application;
[0023] Figure 7 Schematic flowchart of the fourth embodiment of the multi-modal content switching method in this application;
[0024] Figure 8 Schematic structural diagram for querying the topic sequence corresponding to the playback position in this application;
[0025] Figure 9 Schematic flowchart of the fifth embodiment of the multi-modal content switching method in this application;
[0026] Figure 10 Schematic structural diagram for matching candidate subsequences in this application;
[0027] Figure 11 Schematic structural diagram for determining the target subsequence in this application;
[0028] Figure 12 Schematic structural diagram of the multi-modal content switching device in the hardware operating environment involved in the solution of the embodiment of this application;
[0029] Figure 13 Schematic diagram of the functional modules of the multi-modal content switching device in this application.
[0030] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0031] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0032] The embodiment of this application provides a multi-modal content switching method. In the first embodiment of the multi-modal content switching method of this application, specifically, referring to Figure 1 , the multi-modal content switching method includes:
[0033] Step S10, during the playback of the first file, detect the playback status of the first file; the first file is an audio file or a video file.
[0034] In this embodiment, specifically, during the playback of the first file, the current network status is detected in real time, so as to determine whether to switch to other associated modal content according to the current network status.
[0035] Additionally, in step S10: during the playback of the first file, detect the playback status of the first file, and then, it further includes:
[0036] Step a1, determine whether there is associated modal content for the first file;
[0037] Step a2, if there is, return to execute step S20: when it is determined that the playback status meets the switching condition, obtain the associated modal content of the first file;
[0038] Step a3, if not, prompt the target user to perform a clarity switching operation.
[0039] In this implementation, it should be noted that since not all works are configured with multiple modal contents, when it is detected that there is no associated modal content for the first file, the user is prompted that the current network is poor, so as to switch the clarity of the current playback, and thus continue to play the first file by reducing the clarity. Specifically, refer to Figure 2 , Figure 2 which is a schematic structural diagram for video clarity switching in this application. For example, if the video is played in 720P ultra-high definition clarity, it can be reduced to 280P high definition or 270P standard definition clarity for playback.
[0040] Step S20, when it is determined that the playback status meets the switching condition, obtain the associated modal content of the first file, where the file type of the associated modal content is different from the file type of the first file.
[0041] In this embodiment, it should be noted that currently, many literary works are being adapted into video content such as movies and TV dramas, as well as made into audiobooks (audio playback content) and e-books (text content). The same literary work will have three types of modal content: video, audio, and text. Further, the file of the associated modal content and the first file belong to the same work. That is, in this embodiment, the first file is switched to other modal content belonging to the same work. The switching conditions may include when the playback state is in a stuck state, or when it is detected that the current network is sufficient to stably play video or audio, and may also include detecting that the target user issues a click-switch command on the terminal device, or identifying the switching instruction corresponding to the voice command generated by the target user, etc.
[0042] As an implementable manner, specifically, when, during playback, due to network signal fluctuations, the first file currently being played on the user's mobile terminal is in a stuck state, obtain other associated modal content of the same work corresponding to the first file. Specifically, different types of content in each work can be pre-associated and stored in the work database. For example, for the work "Water Margin", the corresponding video file, audio file, and text file of "Water Margin" are associated and stored. Further, during the playback of the first file, when it is determined that the first file needs to be switched, based on the work to which the first file belongs, query the work database for other associated modal content corresponding to this work, so that even when the network signal is poor, the work can still be continuously enjoyed by switching to other modal content that can be played at a low network speed.
[0043] In addition, to improve the user experience, if the current first file is being played at a low network speed, the current playback state is detected in real time. When it is detected that the network state is good, the associated modal content of the current first file is automatically obtained to perform content switching, so that when the network state is better, it is possible to automatically switch to a higher-level content. For example, switch from a text file to a video file for playback.
[0044] As another implementable manner, to make the plot playback coherent, when switching content, it can be switched to a position aligned with the playback position based on the playback position of the first file. Specifically, when it is determined that the playback state meets the switching condition, the playback position of the first file is obtained, and then based on the preset multi-modal theme library, the complete theme sequence corresponding to the first file is queried. It should be noted that the preset multi-modal theme library stores the multi-modal theme sequences corresponding to each work. For example, the complete theme sequences corresponding to videos, audios, and texts. Among them, the multi-modal theme sequence is a multi-modal theme sequence generated by a multi-modal theme model based on the text information of different modal contents in each work. The multi-modal theme model is an LDA theme model. LDA (Latent Dirichlet Allocation) is a document topic generation model, also known as a three-layer Bayesian probability model, including three layers: words, topics, and documents. The generation process of a piece of text has the following rules: each word is selected from a certain topic with a certain probability and a certain word is selected from this topic with a certain probability. The LDA theme model can obtain the "document-topic" information and "topic-word" information through training the corpus. In this embodiment, based on the text information of different modal contents in the same work, different modal theme sequences of the same work are generated by the multi-modal theme model. Among them, the different modal theme sequences include video theme sequences, text theme sequences, and audio theme sequences. The multi-modal theme sequences corresponding to different work information can be stored in the preset multi-modal theme library, and then based on the playback position, the theme sequence corresponding to the playback position is determined from the complete theme sequence corresponding to the first file. Thus, based on the theme sequence corresponding to the playback position, the target content that matches the theme sequence is determined as the associated modal content among the multiple second files corresponding to the same work, so that the plot playback is also coherent after the content is switched.
[0045] Additionally, during the switching process, if it is detected that there are multiple associated modal contents in the currently played first file, the user's mobile terminal can display the multiple associated modal contents for the user to select, which can be referred to Figure 3 , Figure 3 which is the structural schematic diagram for the modal content switching of this application. If the user selects the e-book mode, then it is switched to the modal content corresponding to the e-book.
[0046] Step S30, play the associated modal content.
[0047] The embodiment of the present application provides a multimodal content switching method, device, equipment and storage medium. Compared with the prior art that can only switch between different video resolutions and the video can only be in a stuck state when the user cannot load the video under a weak network, the embodiment of the present application first detects the playing state of the first file during the playing process of the first file; the first file is an audio file or a video file. Then, when it is determined that the playing state meets the switching condition, the associated modal content of the first file is obtained, wherein the file type of the associated modal content is different from the file type of the first file. Then, the associated modal content is played, realizing the switching to the corresponding associated modal content based on the playing state of the first file. That is, the currently played video or audio can be switched to other associated modal content, so that when the user watches a video of a film and television work and the video is stuck and the user cannot watch the video, it is also possible to jump to other associated modal content, thereby improving the fluency of the user's reading of the work.
[0048] Further, referring to Figure 4 , based on the second embodiment in the present application, the above step S20: obtaining the associated modal content of the first file specifically includes:
[0049] Step S21, obtaining the theme sequence of the first file;
[0050] Step S22, determining the target content that matches the theme sequence from multiple second files as the associated modal content. The file types of the multiple second files are different from the file type of the first file, and the first file and the second file belong to the same work.
[0051] In this embodiment, specifically, the playback position of the first file is obtained, and based on a preset multi-modal theme library, the first complete theme sequence corresponding to the current first file is queried. It should be noted that the preset multi-modal theme library includes multi-modal theme sequences corresponding to each work information. The multi-modal theme sequence is formed by converting text information of different modalities of the work information through the multi-modal theme sequence. The text information of different modality contents includes audio reading text information, e-book text information, script text information corresponding to the video, etc. Correspondingly, the multi-modal theme sequence includes an audio theme sequence, a text theme sequence, a video theme sequence, etc. Further, based on the first complete theme sequence, according to a preset theme search method, the theme sequence corresponding to the playback position is searched. For example, the theme corresponding to the current playback is determined, and then a preset number of themes are searched forward and backward from the first complete theme sequence to form the theme sequence. Further still, in the preset multi-modal theme library, the second complete theme sequences corresponding to multiple second files belonging to the same work as the first file are searched, and then the associated theme sequences matching the theme sequence are searched from each of the second complete theme sequences respectively. Further, the associated modality content corresponding to each associated theme sequence is determined from the multiple second files.
[0052] Through the above solution, the embodiment of the present application realizes a multi-modal theme model trained based on different modality contents, so as to convert the text information of different modality contents in each work information into corresponding multi-modal theme sequences through the multi-modal theme model, and then store the multi-modal theme sequences of each work information into a preset modality theme library. Furthermore, based on the preset modality theme library, the theme sequence of the first file and the modality theme sequences corresponding to other multiple second files belonging to the same work can be quickly searched, improving the efficiency of modality content switching.
[0053] Further, referring to Figure 5 , before the step of determining the theme sequence corresponding to the playback position based on the playback position and the preset multi-modal theme library in the second embodiment of the present application, the following is further included:
[0054] Step A10, obtaining the text information of different modality contents corresponding to each work respectively;
[0055] In this embodiment, it should be noted that for the same work, the text information of different modalities of the work is obtained. For example, the video script text information, audio text information, and e-book text information of the work information are obtained.
[0056] Step A20, respectively performing paragraph division on the text information of different modality contents of each work to obtain a multi-modal text segment sequence of each work;
[0057] In this embodiment, it should be noted that the multi-modal text segment sequence is a sequence of text segments containing multiple different modalities. Specifically, the text information of different modalities is divided into paragraphs, so that the text information of each modality is divided into several text segments, and the text segment sequence corresponding to this modality is formed based on each text segment. For example, the e-book text information is divided into several text segments according to paragraphs to obtain the text segment sequence of the e-book, which is defined as D w ={d w1 , d w2 ,..., d wn}. Similarly, the audio text segment sequence of the audiobook reading script can be obtained: D m ={d m1 , d m2 ,..., d mn} and the text segment sequence of the video script: D v ={d v1 , d v2 ,..., d vn}., where D represents the text segment sequence and d represents the text segment.
[0058] Step A30: Merge the multi-modal text segment sequences in the same work to obtain each merged text segment sequence;
[0059] In this embodiment, it should be noted that in order to enable the model to recognize the narrative methods of text segments of different modalities, the multi-modal text segment sequences in the same work are merged to obtain each merged text segment sequence. For example, for "Water Margin", the original work may be more biased towards classical Chinese, but the script after being adapted into a movie is in vernacular Chinese. To enable the model to recognize these two narrative methods, the original work and the script adapted into a movie are merged for training, so as to improve the accuracy of the model's theme recognition.
[0060] Step A40: Based on each of the merged text segment sequences, perform iterative training on the to-be-trained theme model to obtain the multi-modal theme model, and output the multi-modal theme sequences of each work, where the multi-modal theme sequences include video theme sequences, text theme sequences, and audio theme sequences;
[0061] In this embodiment, it should be noted that the multi-modal topic model is the LDA topic model. LDA (Latent Dirichlet Allocation) is a document-topic generation model, also known as a three-layer Bayesian probability model, which includes three layers: words, topics, and documents. The generation process of a text has the following rules: each word is selected from a certain topic with a certain probability and a certain word is selected from this topic with a certain probability. Through training the corpus, the LDA topic model can obtain the "document-topic" information and the "topic-word" information. In this embodiment, based on the text information of different modal contents in the same work, different modal topic sequences of the same work are generated through the multi-modal topic model.
[0062] Specifically, each of the merged text segment sequences is input into the to-be-trained topic model, so as to perform topic prediction on each text segment in the multi-modal text segment sequence through the to-be-trained topic model, and obtain the topic probability set of each text segment in the multi-modal text segment sequence. Among them, the topic probability set of each text segment in the multi-modal text segment sequence is as follows:
[0063] P(z|d) = {P(z1|d), P(z2|d),..., P(z k |d)}, where,
[0064] where d represents the text segment, z represents the topic, and P(z|d) represents the topic probability corresponding to this text segment. It means that the sum of all topic probabilities in the topic probability set corresponding to each text segment is equal to 1. Further, for each text segment in the multi-modal text segment sequence, the topic with the largest topic probability is selected as the representative topic of this text segment: z i = max({P(z1|d), P(z2|d),..., P(z k|d)}, where zi represents the topic with the highest topic probability, and max() represents selecting the topic with the highest topic probability from the set of topic probabilities as the predicted topic. Then, based on the predicted topic and the pre-set topic labels for each text segment, the model loss between the predicted topic and the topic labels is calculated. Further, based on the model loss, the to-be-trained topic model is iteratively optimized, and it is determined whether the optimized to-be-trained topic model meets the training end condition, where the training end condition includes conditions such as iterative convergence and the number of iterations reaching a preset training times threshold. If not, return to execute the step: Based on each of the merged text segment sequences, the to-be-trained topic model is iteratively trained to obtain the multi-modal topic model to continue training the model. If so, the optimized to-be-trained topic model is used as the multi-modal topic model. Then, based on the multi-modal topic model, the text information of different modal contents corresponding to the same work is subject predicted to obtain the multi-modal topic sequence corresponding to the work. The multi-modal topic sequence includes a video topic sequence, a text topic sequence, and an audio topic sequence. For example, following the example of step A20 above, the text segment sequences D w 、D m 、D v are converted into a topic sequence, denoted as Z w 、Z m 、Z v , where the difference between the topic sequence and the text segment sequence is that different text segments can belong to the same topic to obtain the optimal multi-modal topic model.
[0065] Step A50, based on the multi-modal topic sequences of each of the works, form the preset multi-modal topic library.
[0066] In this embodiment, specifically, the multi-modal topic sequences of the same work are associated and stored in the preset multi-modal topic library, so that through the preset multi-modal topic library, the topic sequence corresponding to the currently played file can be queried, and it can be determined whether there are topic sequences of other modalities for the currently played file.
[0067] Further, referring to Figure 6 , Figure 6The figure is a schematic flow chart for training a topic model of this application. Specifically, the e-book text information, the audiobook reading script text information (audio text information), and the video script text information of the same work information are obtained. Then, the e-book text information, the audiobook reading script text information (audio text information), and the video script text information are respectively divided into several text segments according to paragraphs, obtaining an e-book text segment sequence, an audiobook text segment sequence, and a video text segment sequence. Further, the e-book text segment sequence, the audiobook text segment sequence, and the video text segment sequence are merged into a text segment sequence, and the optimal multi-modal topic model is trained. Then, based on the multi-modal topic model, the e-book text segment sequence is transformed into an e-book topic sequence, the audiobook text segment sequence is transformed into an audiobook topic sequence (audio topic sequence), and the video text segment sequence is transformed into a video topic sequence.
[0068] Through the above solution in the embodiment of this application, that is, obtaining the text information of different modal contents respectively corresponding to each work; respectively performing paragraph division on the text information of different modal contents of each work to obtain the multi-modal text segment sequences of each work; merging the multi-modal text segment sequences in the same work to obtain each merged text segment sequence; based on each merged text segment sequence, performing iterative training on the to-be-trained topic model to obtain the multi-modal topic model, and outputting the multi-modal topic sequences of each work, where the multi-modal topic sequences include a video topic sequence, a text topic sequence, and an audio topic sequence; based on the multi-modal topic sequences of each work, forming the preset multi-modal topic library, realizing training the topic model with the multi-modal text segment sequences of the same work, enabling the model to learn the text information of different modalities of the same work, and then transforming the multi-modal text segment sequences into multi-modal topic sequences through the topic model, and then storing the multi-modal topic sequences of the same work information in the preset multi-modal topic library, so that the video topic sequence corresponding to the currently played file can be quickly queried through the preset multi-modal topic library, and it can be determined whether there are topic sequences of other modalities for the currently played file.
[0069] Further, referring to Figure 7 , based on the second embodiment in this application, in another embodiment of this application, where the above step S21: obtaining the topic sequence of the first file includes:
[0070] Step S211, obtaining the playing position of the first file;
[0071] Step S212: Based on the playback position and a preset multi-modal theme library, determine the theme sequence corresponding to the playback position; wherein, the preset multi-modal theme library stores multi-modal theme sequences generated by a multi-modal theme model based on the text information of different modal contents in each work, and the multi-modal theme model is obtained through iterative training based on the text information of different modal contents in each work collected in advance.
[0072] Among them, in the above step S212, based on the playback position and the preset multi-modal theme library, determining the theme sequence corresponding to the playback position specifically includes:
[0073] Step S2121: Query the first complete theme sequence corresponding to the first file in the preset multi-modal theme library;
[0074] Step S2122: Based on the first complete theme sequence, determine the theme position corresponding to the playback position;
[0075] Step S2123: Based on the theme position, inversely query a preset number of target themes in the first complete theme sequence;
[0076] Step S2124: Based on each of the target themes, form the theme sequence.
[0077] In this embodiment, specifically, query the first complete theme sequence corresponding to the first file in the preset multi-modal theme library, and then determine the progress position corresponding to the playback position in the first complete theme sequence. Further, after determining the progress position corresponding to the target video theme sequence based on the playback position, based on a preset query algorithm, query the theme sequence corresponding to the progress position in the target video theme sequence, where the preset query algorithm includes querying a preset number of themes forward or backward in the theme sequence based on the progress position. For example, refer to Figure 8 , Figure 8 is a schematic structural diagram for querying the theme sequence corresponding to the playback position in this application. Specifically, the playback position of the video is R1, and according to the playback position R1, determine the progress position R2 of the target video theme sequence (the first complete theme sequence). Then, from R2 in the reverse direction of the first complete theme sequence (search forward), obtain the nearest 3 themes to form a theme sequence, and the theme sequence is {Z1, Z2, Z2}.
[0078] Through the above solution, the embodiment of the present application realizes quickly querying the first complete theme sequence corresponding to the first file in the preset multi-modal theme library, and then finding the theme sequence corresponding to the playing position in the first complete theme sequence, so that the playing position of the second file can be aligned with that of the first file according to the theme sequence, thereby making the playing progress of the work coherent after content switching.
[0079] Further, referring to Figure 9 , based on the second embodiment of the present application, in another embodiment of the present application, the above step S22: determining the target content matching the theme sequence from multiple second files as the associated modal content includes:
[0080] Step S221, querying the second complete theme sequences respectively corresponding to the multiple second files based on the preset multi-modal theme library;
[0081] Step S222, respectively querying the associated theme sequences matching the theme sequence in each of the second complete theme sequences;
[0082] Among them, the above step S222, respectively querying the associated theme sequences matching the theme sequence in each of the second complete theme sequences, specifically includes:
[0083] Step S2221, respectively querying each candidate subsequence identical to the theme sequence in each of the second complete theme sequences;
[0084] Step S2222, determining the progress values of the candidate subsequences in each of the second complete theme sequences;
[0085] Step S2223, for each of the second complete theme sequences, respectively comparing the progress values of the candidate subsequences with the progress value of the theme sequence, and determining the associated theme sequence corresponding to each of the second complete theme sequences based on the comparison result.
[0086] Step S223, determining the associated modal content corresponding to the associated theme sequence from the multiple second files.
[0087] In this embodiment, specifically, first, based on the preset multi-modal theme library, query the second complete theme sequences respectively corresponding to the multiple second files, and then for each of the second complete theme sequences, the following steps are executed:
[0088] By means of the string search method, query the subsequences identical to the theme sequence in the second complete theme sequence to obtain at least one candidate subsequence, and mark the progress value of each candidate subsequence. For example, referring to Figure 10 , Figure 10Structural schematic diagram for matching candidate subsequences for this application. The subject sequence is {Z1, Z2, Z2}. All subsequences identical to {Z1, Z2, Z2} are queried from the e - book subject sequence (the second complete subject sequence) to obtain each candidate subsequence.
[0089] It should be further noted that since the complete subject sequence is formed by the subjects corresponding to each paragraph in the entire work, each subject has its corresponding progress value relative to the entire subject sequence. For example, if the complete subject sequence is set to 100%, each subject in the complete subject sequence is set with a corresponding value range [0%, 100%] according to its sequence position.
[0090] After obtaining each candidate subsequence, determine the progress values corresponding to each candidate subsequence of the second complete subject sequence, and then calculate the difference between the progress value of each candidate subsequence and the progress value of the subject sequence respectively. Further, select the candidate subsequence with the smallest difference as the target subsequence of the modal subject sequence, and use the progress value of the target subsequence as the target progress position. Among them, the distance calculation method is: Dis = min(|R1 - R n |), where R1 - R n represents the distance between the progress value corresponding to the subject sequence and the progress values corresponding to each candidate subsequence in the second complete subject sequence. R1 represents the progress value corresponding to the subject sequence of the first complete subject sequence, and R n represents the progress values corresponding to each candidate subsequence in the second complete subject sequence. min represents selecting the candidate subsequence with the smallest absolute value of R1 - R n . For example, please refer to Figure 11 , Figure 11 Structural schematic diagram for determining the associated subject sequence for this application. The progress value R1 of the video freeze frame is 20%. The progress values R2 = 25% and R3 = 40% of the e - book subject sequence that meets the conditions are matched in the e - book. Finally, it is calculated that: R2 - R1 = 5% is less than R3 - R1 = 20%. Therefore, the candidate subsequence with the progress value of R2 is selected as the associated subject sequence, realizing the alignment of the progress position of the second complete subject sequence with the playback position of the video freeze frame. Thus, when jumping between multi - modal contents, it is possible to start playing at the nearest progress position to the video playback position, so as to achieve the alignment of the playback progress of switching to other modalities with the playback progress of the first file, making the plot coherent for the user to continue enjoying the work.
[0091] Through the above solution, the embodiment of the present application realizes matching the same candidate subsequence in the second complete theme sequence based on the theme sequence corresponding to the progress position, so as to align the progress position of the second complete theme sequence with the playback position of the video freeze. Therefore, when jumping between multi-modal contents, it is possible to start playing at the nearest progress position from the playback position of the first file, making the playback plot coherent for the user to continue enjoying the work and improving the user's viewing experience.
[0092] Referring to Figure 12 , Figure 12 FIG. is a schematic structural diagram of a multi-modal content switching device in the hardware operating environment involved in the solution of the embodiment of the present application.
[0093] As Figure 12 shown, the multi-modal content switching device may include: a processor 1001, such as a CPU, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between the processor 1001 and the memory 1005. The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0094] Optionally, the multi-modal content switching device may further include a rectangular user interface, a network interface, a camera, an RF (Radio Frequency) circuit, sensors, an audio circuit, a WiFi module, etc. The rectangular user interface may include a display screen (Display) and an input sub-module such as a keyboard (Keyboard). Optionally, the rectangular user interface may further include a standard wired interface and a wireless interface. The network interface may optionally include a standard wired interface and a wireless interface (such as a WIFI interface).
[0095] Those skilled in the art can understand that Figure 12 the structural diagram of the multi-modal content switching device shown in
[0096] does not limit the multi-modal content switching device, and may include more or fewer components than shown, or combine some components, or have different component arrangements. Figure 12 As shown, the memory 1005, as a computer storage medium, may include an operating device, a network communication module, and a multi-modal content switching program. The operating device is a program for managing and controlling the hardware and software resources of the multi-modal content switching device, supporting the operation of the multi-modal content switching program and other software and / or programs. The network communication module is used to realize the communication between the components inside the memory 1005, as well as the communication between the memory 1005 and other hardware and software in the multi-modal content switching device.
[0097] In Figure 12 In the multimodal content switching device shown, the processor 1001 is used to execute the multimodal content switching program stored in the memory 1005 to implement the steps of the multimodal content switching method described in any one of the above.
[0098] The specific implementation manner of the multimodal content switching device of this application is basically the same as that of each embodiment of the above multimodal content switching method, and will not be elaborated here.
[0099] In addition, please refer to Figure 13 , Figure 13 which is a schematic diagram of the functional modules of the multimodal content switching device of this application. This application also provides a multimodal content switching device, and the multimodal content switching device includes:
[0100] A detection module, configured to detect the playing state of the first file during the playing process of the first file; the first file is an audio file or a video file;
[0101] An acquisition module, configured to acquire the associated modal content of the first file when it is determined that the playing state meets the switching condition, wherein the file type of the associated modal content is different from the file type of the first file;
[0102] A playing module, configured to play the associated modal content.
[0103] Optionally, the acquisition module is further configured to:
[0104] Acquire the theme sequence of the first file;
[0105] Determine the target content that matches the theme sequence from multiple second files as the associated modal content, the file types of the multiple second files are different from the file type of the first file, and the first file and the second file belong to the same work.
[0106] Optionally, the acquisition module is further configured to:
[0107] Acquire the playing position of the first file;
[0108] Based on the playing position and a preset multimodal theme library, determine the theme sequence corresponding to the playing position;
[0109] Wherein, the preset multimodal theme library stores multimodal theme sequences generated by a multimodal theme model based on the text information of different modal contents in each work, and the multimodal theme model is obtained through iterative training based on the text information of different modal contents in each work collected in advance.
[0110] Optionally, the multimodal content switching device is further configured to:
[0111] Obtain the text information of different modal content corresponding to each work respectively;
[0112] Perform paragraph division on the text information of different modal content of each said work respectively to obtain the multi-modal text segment sequences of each said work;
[0113] Merge the multi-modal text segment sequences in the same work to obtain each merged text segment sequence;
[0114] Based on each said merged text segment sequence, perform iterative training on the to-be-trained topic model to obtain the multi-modal topic model, and output the multi-modal topic sequences of each said work, where the multi-modal topic sequences include video topic sequences, text topic sequences, and audio topic sequences;
[0115] Based on the multi-modal topic sequences of each said work, form the preset multi-modal topic library.
[0116] Optionally, the multi-modal content switching device is further configured to:
[0117] Query the first complete topic sequence corresponding to the first file in the preset multi-modal topic library;
[0118] Based on the first complete topic sequence, determine the topic position corresponding to the playback position;
[0119] Based on the topic position, perform reverse query of a preset number of target topics in the first complete topic sequence;
[0120] Based on each said target topic, form the topic sequence.
[0121] Optionally, the multi-modal content switching device is further configured to:
[0122] Based on the preset multi-modal topic library, query the second complete topic sequences corresponding to the multiple second files respectively;
[0123] In each said second complete topic sequence, query the associated topic sequences that match the topic sequence respectively;
[0124] Determine the associated modal content corresponding to the associated topic sequence from the multiple second files.
[0125] Optionally, the multi-modal content switching device is further configured to:
[0126] Query the candidate subsequences that are the same as the topic sequence in each said second complete topic sequence respectively;
[0127] Determine the progress values of the candidate subsequences in each said second complete topic sequence;
[0128] For each of the second complete theme sequences, the progress value of each candidate subsequence is compared with the progress value of the theme sequence respectively, and based on the comparison result, the associated theme sequence corresponding to each second complete theme sequence is determined.
[0129] The specific implementation manner of the multi-modal content switching device of the present application is basically the same as that of the above-mentioned embodiments of the multi-modal content switching method, and will not be elaborated here.
[0130] The embodiments of the present application provide a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores one or more programs, and the one or more programs can also be executed by one or more processors to implement the steps of the multi-modal content switching method described in any one of the above.
[0131] The specific implementation manner of the computer-readable storage medium of the present application is basically the same as that of the above-mentioned embodiments of the multi-modal content switching method, and will not be elaborated here.
[0132] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent scope of the present application by the same token.
Claims
1. A multimodal content switching method, characterized in that, The multi-modal content switching method includes: During the playback of the first file, detecting the playback state of the first file; the first file is an audio file or a video file; When it is determined that the playback state meets the switching condition, obtaining the associated modal content of the first file, where the file type of the associated modal content is different from the file type of the first file, and the associated modal content includes video, audio, or text modal content of the work corresponding to the first file; In the case where the first file has the associated modal content, playing the associated modal content; When the first file does not have the associated modal content and the first file is a video file, prompting the user that the current network is poor and reducing the playback clarity of the first file; The switching condition includes at least one of the following: the playback state of the first file is in a stuck state, the current network is sufficient to stably play video or audio, and a click-switching instruction of the target user is detected.
2. The multimodal content switching method according to claim 1, wherein The obtaining of the associated modal content of the first file includes: Obtaining the theme sequence of the first file; Determining the target content that matches the theme sequence from multiple second files as the associated modal content, where the file types of the multiple second files are different from the file type of the first file, and the first file and the second file belong to the same work.
3. The multimodal content switching method according to claim 2, wherein, The obtaining of the theme sequence of the first file includes: Obtaining the playback position of the first file; Based on the playback position and a preset multi-modal theme library, determining the theme sequence corresponding to the playback position; Among them, the preset multi-modal theme library stores multi-modal theme sequences generated by a multi-modal theme model based on the text information of different modal contents in each work, and the multi-modal theme model is obtained by iterative training based on the text information of different modal contents in each work collected in advance.
4. The multimodal content switching method according to claim 3, wherein Before the step of determining the theme sequence corresponding to the playback position based on the playback position and the preset multi-modal theme library, it further includes: Obtaining the text information of different modal contents corresponding to each work; Respectively performing paragraph division on the text information of different modal contents of each work to obtain the multi-modal text segment sequences of each work; Merging the multi-modal text segment sequences in the same work to obtain each merged text segment sequence; Based on each of the merged text segment sequences, performing iterative training on the to-be-trained theme model to obtain the multi-modal theme model, and outputting the multi-modal theme sequences of each work, where the multi-modal theme sequences include video theme sequences, text theme sequences, and audio theme sequences; Based on the multi-modal theme sequences of each work, forming the preset multi-modal theme library.
5. The multimodal content switching method according to claim 3, wherein The determining of the theme sequence corresponding to the playback position based on the playback position and the preset multi-modal theme library includes: Querying the first complete theme sequence corresponding to the first file in the preset multi-modal theme library; Based on the first complete theme sequence, determining the theme position corresponding to the playback position; Based on the theme position, performing reverse query in the first complete theme sequence for a preset number of target themes; Form the theme sequence based on each of the target themes.
6. The multimodal content switching method according to claim 2, wherein, Determining the target content that matches the theme sequence from multiple second documents as the associated modal content includes: Query the second complete theme sequences respectively corresponding to the multiple second documents based on a preset multimodal theme library; In each of the second complete theme sequences, query the associated theme sequences that match the theme sequence; Determine the associated modal content corresponding to the associated theme sequence from the multiple second documents.
7. The multimodal content switching method according to claim 6, wherein The querying, in each of the second complete theme sequences, of the associated theme sequences that match the theme sequence includes: Query, in each of the second complete theme sequences, each candidate subsequence that is the same as the theme sequence; Determine the progress value of each candidate subsequence in each of the second complete theme sequences; For each of the second complete theme sequences, compare the progress value of each candidate subsequence with the progress value of the theme sequence respectively, and based on the comparison result, determine the associated theme sequence corresponding to each of the second complete theme sequences.
8. A multimodal content switching device, characterized in that, The multimodal content switching device includes: A detection module, configured to detect the playing state of a first document during the playing of the first document; the first document is an audio document or a video document; An acquisition module, configured to acquire the associated modal content of the first document when it is determined that the playing state meets the switching condition, where the file type of the associated modal content is different from the file type of the first document, and the associated modal content includes video, audio, or text modal content of the work corresponding to the first document; A playing module, configured to play the associated modal content when there is the associated modal content for the first document; The playing module is further configured to, when there is no associated modal content for the first document and the first document is a video document, prompt the user that the current network is poor and reduce the playing clarity of the first document; The switching condition includes at least one of the following: the playing state of the first document is in a stuck state, the current network is sufficient to stably play video or audio, and a click switching instruction of a target user is detected.
9. A multimodal content switching device, characterized in that, The multimodal content switching device includes: a memory, a processor, and a multimodal content switching program stored on the memory, The multimodal content switching program is executed by the processor to implement the steps of the multimodal content switching method according to any one of claims 1 to 7.
10. A storage medium, the storage medium being a computer-readable storage medium, characterized in that, A multimodal content switching program is stored on the computer-readable storage medium, and the multimodal content switching program is executed by the processor to implement the steps of the multimodal content switching method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method, device and terminal for switching play modes
CN104581320A