Information display method, information generation method, device, equipment and storage medium
By acquiring target objects and their words in the video, video notes are generated and displayed, solving the problem that users find it difficult to quickly take video notes and achieving efficient note generation and viewing.
Patent Information
- Application Number
- CN202210157319.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-02-21
AI Technical Summary
Users cannot quickly and easily take or view video notes while watching videos; existing technologies are complex and inefficient.
By acquiring the target objects associated with the video and their corresponding target words, the text fragments of the target objects in the video are determined. Note information is generated by combining text information and audio-visual information, and video notes are automatically generated by combining geographical location information.
It reduces the complexity of users taking video notes, improves the efficiency of generating and viewing video notes, and optimizes the user experience.
Smart Images

Figure CN116662607B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an information display method, information generation method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of computer technology, users can watch videos through video playback software installed on their terminals or by visiting video playback websites. Users with learning purposes can take study notes while watching videos.
[0003] In related technologies, users can take notes on video content in their notebooks or office software while watching videos, or create notes using interfaces provided by the video playback page. However, these technologies do not allow users to quickly and easily take or view video notes while watching videos; users need to create and edit video notes themselves, resulting in high complexity and low efficiency. Summary of the Invention
[0004] This application provides an information display method, information generation method, apparatus, device, and storage medium, which can reduce the operational complexity of users taking video notes while watching videos, improve the generation and viewing efficiency of video notes, and optimize the user experience of taking notes while watching videos.
[0005] According to one aspect of the embodiments of this application, an information display method is provided, the method comprising:
[0006] Display a first page, which includes a first video;
[0007] In response to the instruction to display note information, the video note information corresponding to the first video is displayed on the first page;
[0008] The video notes information includes geographic location information and notes information corresponding to at least one target object associated with the first video. The notes information is generated based on text information and audio / video information corresponding to text segments of the at least one target object in the first video. The text segments are determined based on target words corresponding to the at least one target object in the first video.
[0009] According to one aspect of the embodiments of this application, an information generation method is provided, the method comprising:
[0010] Acquire a first video, wherein the first video includes at least one target word corresponding to a target object;
[0011] Based on the target words, determine the text segment corresponding to the at least one target object in the first video;
[0012] Based on the text fragment, determine the text information and audio / video information corresponding to the at least one target object;
[0013] The text information and the audio / video information are fused together to generate note information corresponding to the at least one target object;
[0014] Obtain the geographical location information corresponding to the at least one target object;
[0015] Based on the geographic location information and the note information, video note information corresponding to the first video is generated, and the video note information is used to display on the first page.
[0016] According to one aspect of the embodiments of this application, an information display device is provided, the device comprising:
[0017] A page display module is used to display a first page, which includes a first video.
[0018] The note display module is used to display video note information corresponding to the first video on the first page in response to the note information display instruction;
[0019] The video notes information includes geographic location information and notes information corresponding to at least one target object associated with the first video. The notes information is generated based on text information and audio / video information corresponding to text segments of the at least one target object in the first video. The text segments are determined based on target words corresponding to the at least one target object in the first video.
[0020] According to one aspect of the embodiments of this application, an information generation apparatus is provided, the apparatus comprising:
[0021] The video acquisition module is used to acquire a first video, wherein the first video includes at least one target word corresponding to a target object;
[0022] A text segment determination module is used to determine, based on the target words, the text segment corresponding to the at least one target object in the first video;
[0023] The media information determination module is used to determine the text information and audio / video information corresponding to the at least one target object based on the text fragment;
[0024] The note information generation module is used to fuse the text information and the audio and video information to generate note information corresponding to the at least one target object;
[0025] A location information acquisition module is used to acquire the geographical location information corresponding to the at least one target object;
[0026] The video note generation module is used to generate video note information corresponding to the first video based on the geographical location information and the note information, and the video note information is used to display on the first page.
[0027] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the above-described information display method or the above-described information generation method.
[0028] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, wherein the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the above-described information display method or the above-described information generation method.
[0029] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform an action to implement the above-described information display method or the above-described information generation method.
[0030] The technical solution provided in this application can bring the following beneficial effects:
[0031] By obtaining the target object associated with the video and its corresponding target words, the text segment corresponding to the target object in the video can be determined. Based on the text information and audio-visual information corresponding to the text segment, note information corresponding to the target object can be generated. By combining the note information with the geographical location information corresponding to the target object, video notes information displayed on the page can be automatically generated. This increases the amount of information displayed on the page, reduces the operational complexity of users taking video notes while watching videos, improves the efficiency of video note generation and viewing, and optimizes the user experience of taking notes while watching and learning videos. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a schematic diagram of an application runtime environment provided in one embodiment of this application;
[0034] Figure 2 This is a flowchart of an information display method provided in one embodiment of this application. Figure 1 ;
[0035] Figure 3 This is a flowchart of an information display method provided in one embodiment of this application. Figure 2 ;
[0036] Figure 4 An exemplary diagram of a note display control is shown;
[0037] Figure 5 An exemplary diagram of a map display control is shown;
[0038] Figure 6 This is a flowchart of an information generation method provided in one embodiment of this application. Figure 1 ;
[0039] Figure 7 This is a flowchart of an information generation method provided in one embodiment of this application. Figure 2 ;
[0040] Figure 8 This is an interactive flowchart of an information display method provided in one embodiment of this application;
[0041] Figure 9 This is a block diagram of an information display device provided in one embodiment of this application;
[0042] Figure 10 This is a block diagram of an information generation apparatus provided in one embodiment of this application;
[0043] Figure 11 This is a structural block diagram of a computer device provided in one embodiment of this application. Figure 1 ;
[0044] Figure 12 This is a structural block diagram of a computer device provided in one embodiment of this application. Figure 2 . Detailed Implementation
[0045] The information display method and information generation method provided in the embodiments of this application involve artificial intelligence technology, which will be briefly described below to facilitate understanding by those skilled in the art.
[0046] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0047] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and intelligent transportation.
[0048] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0049] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0050] In this embodiment of the application, the text corresponding to the video can be processed based on the above-mentioned natural language processing technology and machine learning technology. For example, the above-mentioned artificial intelligence technology can be applied to text processing related to text, such as word recognition, text segmentation, and summary generation.
[0051] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0052] Please refer to Figure 1 This diagram illustrates an application runtime environment provided in one embodiment of this application. The application runtime environment may include: terminal 10 and server 20.
[0053] Terminal 10 includes, but is not limited to, electronic devices such as mobile phones, computers, smart voice interaction devices, smart home appliances, in-vehicle terminals, game consoles, e-book readers, multimedia playback devices, and wearable devices. Application clients can be installed on terminal 10.
[0054] In this embodiment, the application described above can be any application capable of providing video services. Typically, the application is a video application. Of course, other types of applications besides video applications can also provide video services. For example, browser applications, news applications, social applications, interactive entertainment applications, shopping applications, content sharing applications, virtual reality (VR) applications, augmented reality (AR) applications, etc., are not limited in this embodiment. In addition, the videos pushed by different applications will be different, and the corresponding functions will also be different. These can be pre-configured according to actual needs, and are not limited in this embodiment. Optionally, the terminal 10 runs a client of the above-mentioned application.
[0055] Server 20 provides background services to clients of applications in terminal 10. For example, server 20 can be a background server for the aforementioned applications. Server 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, server 20 can simultaneously provide background services to applications in multiple terminals 10.
[0056] Optionally, terminal 10 and server 20 can communicate with each other via network 30. Terminal 10 and server 20 can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0057] Please refer to Figure 2 This illustrates the flow of an information display method provided in one embodiment of this application. Figure 1 This method can be applied to computer devices, which refer to electronic devices with data computing and processing capabilities. For example, the entity executing each step can be... Figure 1 Terminal 10 in the application runtime environment shown. The method may include the following steps (210-220).
[0058] Step 210: Display the first page.
[0059] The first page includes the first video. Optionally, the first page is a video playback page, which can be used to play the aforementioned first video. The aforementioned video playback page can be any page capable of playing videos, including but not limited to client-side pages and browser pages.
[0060] In an exemplary embodiment, the first video mentioned above is a content explanation video, including but not limited to science videos, learning videos, and lecture videos. For example, the first video mentioned above is a geographical documentary video.
[0061] Geographical documentaries and other explanatory videos differ from ordinary films and television programs in that they are objective and rational. The videos have a strong overall structure and use formal language, with direct meaning and rigorous grammar, making them very suitable for text analysis.
[0062] Step 220: In response to the instruction to display note information, display the video note information corresponding to the first video on the first page.
[0063] The video notes information includes geographic location information and notes information corresponding to at least one target object associated with the first video. The notes information is generated based on the text information and audio-visual information corresponding to the text segment of at least one target object in the first video. The text segment is determined based on the target words corresponding to at least one target object in the first video.
[0064] The above-mentioned note information display command is a command used to trigger the display of note information. This application embodiment does not limit the triggering method of the note information display command, including but not limited to operation control triggering, gesture triggering, remote control triggering, voice command triggering, etc.
[0065] In an exemplary embodiment, in response to a note information display instruction, video note information corresponding to the first video is obtained; and the video note information is displayed on the first page.
[0066] In one possible implementation, the process of obtaining video note information includes: sending note acquisition request information to the server, the note acquisition request information including the video identifier of the first video, and the server generating video note information for the first video based on the video identifier of the first video.
[0067] In another possible implementation, in response to the instruction to display note information, video note information corresponding to the first video is generated; the video note information is then displayed on the first page.
[0068] The process of generating the above video notes information can be found in the embodiments of the information generation methods described below.
[0069] In an exemplary embodiment, the first page is a video playback page, which includes notes display controls. Correspondingly, as... Figure 3 As shown, the implementation process of step 220 above includes the following steps (221-223). Figure 3 The flowchart of an information display method provided in one embodiment of this application is shown. Figure 2 .
[0070] Step 221: In response to the selection operation for the note display control, display map display information on the video playback page.
[0071] In an exemplary embodiment, the selection operation is used to trigger the note information display instruction. The map display information includes a location identifier corresponding to at least one target object. The location identifier is used to represent the geographical location information. The at least one target object includes a first target object.
[0072] Optionally, the map display information includes a map display control, which is a page control used to display video notes information on the target map. Optionally, the note display control is displayed in the operation bar of the video player component, and the map display control can display the target map and the aforementioned location markers. Optionally, the map display control is an interactive map control.
[0073] The first target object mentioned above is a target object among at least one of the aforementioned target objects.
[0074] In one example, such as Figure 4 As shown, it exemplifies a schematic diagram of a note display control. Figure 4 The note display control 41 shown can be displayed in the operation area 40 of the video player. The user can display the video notes after clicking the note display control 41.
[0075] In one possible implementation, at least one location identifier corresponding to the target object can be displayed in the map display control described above.
[0076] In an exemplary embodiment, geographic location information corresponding to at least one target object is obtained, including location coordinates of the target object; based on the geographic location information corresponding to at least one target object, the target display position corresponding to the location identifier of the at least one target object in the map display control is determined; and the location identifier corresponding to the at least one target object is displayed at the target display position. Optionally, based on the geographic location information corresponding to at least one target object, a target map area corresponding to the first video is determined, and the target map area is displayed in the map display control in a first form. For example, the border corresponding to the target map area is displayed, and the target map area is displayed with target color and target transparency. Optionally, the target map area includes the location identifier corresponding to the at least one target object.
[0077] If at least one target object is a location name, obtain the corresponding geographic location information, including location coordinates such as latitude and longitude. Then, display a location marker at the target location.
[0078] In one possible implementation, the map display control may also display the location marker corresponding to the target object in the second video. The second video is an associated video of the first video; for example, if the first video is an episode of a geographical documentary, the second video is another episode of that geographical documentary besides the first video, or another documentary video.
[0079] Step 222: In response to the trigger operation of the location identifier corresponding to the first target object, display the information display control corresponding to the first target object.
[0080] The first location identifier is the location identifier corresponding to the first target object, and the first target object is any one of at least one target object.
[0081] The aforementioned triggering operation is used to trigger the display of the aforementioned information display control. This application embodiment does not limit the method of triggering the operation. Optionally, the triggering operation includes, but is not limited to, click operations, double-click operations, hover operations, etc.
[0082] Step 223: Display the note information corresponding to the first target object based on the information display control.
[0083] Optionally, based on the information display control, audio / video information and text information corresponding to the first target object are displayed. Optionally, the text information includes, but is not limited to, title information and summary information corresponding to the text fragment.
[0084] In an exemplary embodiment, the information display control described above may display at least one type of media information, including title information, summary information, and audio / video information corresponding to the first target object. The audio / video information includes, but is not limited to, image information, video information, and audio information corresponding to the first target object.
[0085] In one example, such as Figure 5 As shown, it exemplifies a schematic diagram of a map display control. Figure 5 The map display control 50 shown allows users to mark video notes corresponding to target locations on the map. Each target location corresponds to a location marker 51. When a user hovers the mouse over a location marker 51 or clicks on a location marker 51, the map display control 50 displays a corresponding information display control 52. This information display control shows the video notes corresponding to the location and allows for editing. Figure 5 In the middle, the information display control 52 displays the note information corresponding to the target object "Mount A", with the title "Mount A" and the corresponding summary information "Mount A is located in Province B". The above summary information "Mount A is located in Province B" is the summary information extracted from the text segment corresponding to the target object "Mount A" in the video text.
[0086] By displaying location markers for each target object in the map display control on the page, and triggering an information display control to show the corresponding notes for that target object by interacting with the location markers, users can quickly and intuitively understand the video content, such as the locations mentioned in the video. Furthermore, for users with learning purposes, this method can quickly help them organize video notes, reducing the complexity of note-taking and saving them time.
[0087] In an exemplary embodiment, such as Figure 3 As shown, after step 223 above, the above method may further include the following step 230.
[0088] Step 230: In response to a selection operation on the information display control, display note information corresponding to at least one target object.
[0089] In one possible implementation, in response to a selection operation on the information display control, a second page is displayed; the second page displays text information and audio / video information corresponding to at least one target object.
[0090] Optionally, the second page described above is a video notes details display page. In some embodiments, the first page and the second page may also be the same page, and this application embodiment does not limit this.
[0091] In an exemplary embodiment, the second page may simultaneously display text information and audio / video information corresponding to at least one target object.
[0092] Optionally, the second page may display text information and audio / video information corresponding to the target object of at least one video. The at least one video includes the first video. For a second video other than the first video, the second page may display text information and audio / video information corresponding to the target object of the second video.
[0093] By receiving selection operations on the information display controls on the page, a second page is triggered. The second page can display text and audio / video information corresponding to each target object, which can increase the amount of information displayed on the page, make it easier for users to view more detailed note information, enhance the richness of video note information, and increase the flexibility of viewing video notes. Displaying note details after operating the information display controls can reduce page rendering pressure and improve page rendering efficiency.
[0094] In an exemplary embodiment, such as Figure 3 As shown, the above method may also include the following steps (240-250).
[0095] Step 240: In response to the editing operation on the video notes information, generate note correction information corresponding to the video notes information.
[0096] In an exemplary embodiment, an interface for users to correct or edit video notes can be provided on the first or second page. Users can perform editing operations on any note information or geolocation information in the video notes to generate note correction information corresponding to the video notes.
[0097] Optionally, the above note correction information can be sent to the server.
[0098] Step 250: Update the video notes information based on the note correction information.
[0099] By receiving user edits and modifications to automatically generated video notes, error correction information can be obtained, improving the flexibility and accuracy of video note-taking.
[0100] In an exemplary embodiment, such as Figure 3 As shown, the above method may also include the following steps (260-270).
[0101] Step 260: Display at least one note correction message corresponding to the video notes.
[0102] In an exemplary embodiment, while displaying video note information, at least one note correction message may also be displayed for any item of information in the video note information (any item of information refers to the geographical location information or note information corresponding to any target object). The note correction message is note correction information published and uploaded by the target account object. The target account object includes, but is not limited to, the user corresponding to the target account, electronic device, AI customer service, AI robot, etc.
[0103] Step 270: In response to a preset operation for a target note correction information in at least one note correction information, send feedback information corresponding to the preset operation.
[0104] Feedback information is used to characterize the accuracy of the target note correction information.
[0105] When at least one note correction information corresponding to the video notes is displayed on the first or second page, the user can choose to agree, refuse, or ignore the note correction information through the options of agreeing to the modification, refusing to modify, and ignoring the modification displayed on the page.
[0106] In one possible implementation, in response to a selection operation for the "agree to modification" option, an "agree to modification" message is sent to the server; in response to a selection operation for the "reject modification" option, a "reject modification" message is sent to the server; and in response to a selection operation for the "ignore modification" option, an "ignore modification" message is sent to the server. The aforementioned preset operations include selection operations for the "agree to modification" option, the "reject modification" option, and the "ignore modification" option.
[0107] If the number of "agree to modification" messages received by the server exceeds a preset number, the target note information can be updated based on the aforementioned error correction information. The updated target note information is then returned to the terminal, which can then display the updated target note information on the first or second page. Similarly, if the number of "reject to modification" messages received by the server exceeds a preset number, the target correction note can be deleted, preserving the original target note information.
[0108] By receiving feedback on note corrections, the system can automatically update existing notes, improving the accuracy of video notes and the efficiency of note management.
[0109] In an exemplary embodiment, the aforementioned page and its various controls can be implemented based on a web (web page) canvas.
[0110] In summary, the technical solution provided in this application, by obtaining the target object associated with the video and its corresponding target words, can determine the text segment corresponding to the target object in the video. Based on the text information and audio / video information corresponding to the text segment, note information corresponding to the target object can be generated. By combining the note information with the geographical location information corresponding to the target object, video note information displayed on the page can be automatically generated. This increases the amount of information displayed on the page, reduces the operational complexity of users taking video notes while watching videos, improves the efficiency of video note generation and viewing, and optimizes the user experience of taking notes while watching and learning videos.
[0111] The generation of notes for geographical documentaries is a typical application scenario of the technical solution provided in this application. Using the technical solution provided in this application to generate video notes for displaying geographical documentaries can help users better save and find relevant information in the videos. The above-mentioned note extraction scheme based on text analysis can intelligently display relevant notes corresponding to the documentary to users in the form of interactive maps and text, and supports users to modify the notes generated initially and save them as their own notes, thereby improving the efficiency of video note production and viewing, and optimizing the user experience.
[0112] Please refer to Figure 6 It illustrates the flow of an information generation method provided in one embodiment of this application. Figure 1 This method can be applied to computer devices, which refer to electronic devices with data computing and processing capabilities. For example, the entity executing each step can be... Figure 1 The application runtime environment shown is a terminal 10 or a server 20. The method may include the following steps (601-606).
[0113] Step 601: Obtain the first video.
[0114] The aforementioned first video includes at least one target word corresponding to a target object. The target word is a word in the video text corresponding to the first video, which includes, but is not limited to, narration text, subtitle text, and audio transcript text. Optionally, the video text includes at least one target word corresponding to at least one target object.
[0115] In an exemplary embodiment, the first video mentioned above is a content explanation video, including but not limited to science videos, learning videos, and lecture videos. For example, the first video mentioned above is a video from a geographical documentary.
[0116] Geographical documentaries and other explanatory videos differ from ordinary films and television programs in that they are objective and rational. The videos have a strong overall structure and use very formal scripts. The video text is direct in meaning and grammatically rigorous, making it very suitable for text analysis.
[0117] In an exemplary embodiment, such as Figure 7 As shown, the above information generation method further includes the following steps (607-609), Figure 7 The flowchart of an information generation method provided in one embodiment of this application is shown. Figure 2 .
[0118] Step 607: Perform word recognition processing on the video text corresponding to the first video to obtain a word sequence.
[0119] The above word sequence includes at least one word in the video text that meets the preset conditions.
[0120] Optionally, speech recognition processing is performed on the audio of the first video to obtain the aforementioned video text, which may be narration text. Optionally, the video text corresponding to the first video is extracted, which may be subtitle text. In some possible scenarios, such as when the video is a content-based explanatory video, the aforementioned narration text and subtitle text may be the same.
[0121] In an exemplary embodiment, all words in the video text that meet the aforementioned preset conditions are sequentially acquired. For example, all location nouns in the video text are sequentially acquired; correspondingly, the aforementioned preset conditions include that the word type belongs to the location noun type.
[0122] Optionally, based on a preset word segmentation method, the words in the video text are compared with words in a preset word library to extract all words in the text that meet preset conditions, forming the aforementioned word sequence. If a word in the video text matches a word in the preset word library, that word can be identified as meeting the preset conditions. In one example, the word sequence can be represented as [A1, A2, A3, ...], where A1, A2, A3, etc., represent words that meet the preset conditions. Illustratively, the aforementioned preset word library is a geographic word library. The word sequence includes at least one word in the video text that meets the preset conditions.
[0123] The aforementioned pre-defined word segmentation methods include, but are not limited to, shortest path word segmentation, n-gram word segmentation, word segmentation based on character formation, and word segmentation using neural network models. The main neural network models include Long Short-Term Memory (LSTM) network models and Transformer (self-attention transformation network) models.
[0124] Step 608: Obtain the word level information corresponding to at least one word.
[0125] In an exemplary embodiment, word level information corresponding to a word is retrieved from a preset database using words as query terms.
[0126] The aforementioned word and phrase hierarchy information includes, but is not limited to, administrative division hierarchy information, affiliation hierarchy information, and level hierarchy information. By determining the word and phrase hierarchy information corresponding to at least one word in a word and phrase sequence, words and phrases with related relationships in the word and phrase sequence can be identified.
[0127] Step 609: Based on the word level information, segment at least one word to obtain a word segment sequence corresponding to at least one target object.
[0128] The above word segmentation sequence includes the above target words. Optionally, the word segmentation sequence includes the target words corresponding to the target.
[0129] Based on the word level information corresponding to at least one word in the above word sequence, the corresponding association between adjacent words in the word sequence is determined; based on the above association, at least one word can be segmented to obtain at least one word segment sequence.
[0130] In one example, at least one of the above words is a place noun. The place nouns can be divided into levels from high to low, such as country, province / city, prefecture, county, township, village, and specific location. There is a subordinate relationship between the levels. The relationship between words in the above word segment sequence can be determined based on the level information.
[0131] Accordingly, based on the word level information corresponding to at least one word and the sorting position of at least one word in the word sequence, at least one word in the word sequence is segmented to obtain at least one word segment sequence.
[0132] In an exemplary embodiment, the segmentation process for at least one word follows word segmentation rules. These word segmentation rules include, but are not limited to, the following:
[0133] a. The nouns in the segments are consecutive, and each segment should contain as many nouns as possible, allowing a noun to be divided into two segments;
[0134] b. No different nouns of the same level can appear in a paragraph, and the lower-level administrative region is subordinate to the higher-level administrative region;
[0135] c. Locations that appear only once are excluded as distractors.
[0136] Here's an example to illustrate this. Suppose the device obtains a word sequence of geographical terms from the video text corresponding to the first video, which is ["Mount SiA", "Province SiB", "Autonomous Prefecture C", "Province SiB", "Xiang D La", "Grassland E", "Plateau F"]. Since Mount SiA is located in Autonomous Prefecture C of Province SiB, and Xiang D La is located in Province SiB, we can extract three segments: ["Mount SiA", "Province SiB", "Autonomous Prefecture C", "Province SiB"], ["Province SiB", "Xiang D La", "Grassland E"], and ["Plateau F"]. After filtering out "Plateau F", the word sequence is finally divided into two segments, thus achieving the effect of dividing the first video into two segments.
[0137] After determining the above word segmentation sequence, the target object corresponding to at least one word segmentation sequence can be determined based on the word level information of each word in each word segmentation sequence.
[0138] In an exemplary embodiment, words whose word level information in a word segmentation sequence meets preset level conditions can be identified as target objects corresponding to the word segmentation sequence. The preset level conditions include, but are not limited to, conditions such as word level being greater than or equal to a preset threshold, word level being less than a preset threshold, and word level being a target level. This application embodiment does not limit these conditions.
[0139] In one example, the place name with the lowest administrative division level in each word segment sequence is identified as the target object corresponding to the word segment sequence.
[0140] By identifying words in the video text that meet preset conditions, a word sequence corresponding to the video text can be obtained. Then, the relationship between words can be determined according to the level of words in the word sequence. Subsequently, the word sequence can be segmented to obtain a word segment sequence. Based on the level information of words in the word segment sequence, the target objects that need to be recorded can be accurately determined, thus improving the accuracy of generating video notes.
[0141] Step 602: Based on the target words, determine at least one text segment corresponding to the target object in the first video.
[0142] In an exemplary embodiment, one of the major differences between content-explanation videos, such as documentaries, and other videos lies in the video text. While segmenting a typical video requires processing the video frames, in content-explanation videos, text plays the primary role in explanation; almost every scene containing key information is accompanied by corresponding text. Furthermore, the video text for content-explanation videos is significantly longer than that of typical videos. Therefore, this embodiment uses text segmentation to achieve the effect of segmenting the video.
[0143] In some possible implementations, the video text can be segmented according to the above-mentioned word segmentation sequence to obtain text segments corresponding to at least one word segmentation sequence.
[0144] Optionally, the starting sentence corresponding to at least one word segmentation sequence in the video text is determined, and the video text is segmented according to each starting sentence to obtain a text segment corresponding to at least one word segmentation sequence. Optionally, the starting sentence of any text segment is the starting sentence corresponding to the corresponding word segmentation sequence, and the sentence preceding the starting sentence of the next word segmentation sequence corresponding to that text segment is the ending sentence of that text segment.
[0145] Optionally, the process of determining the starting sentence includes: determining the first word in at least one word segment sequence; and determining the sentence corresponding to the first word in the video text as the starting sentence corresponding to at least one word segment sequence in the video text.
[0146] After obtaining the above text fragments, the text fragments corresponding to at least one word segmentation sequence can be identified as the text fragments corresponding to at least one target object. The above at least one word segmentation sequence corresponds to at least one target object.
[0147] The text segments corresponding to the word segmentation sequences determined in the preceding steps are the text segments corresponding to the target words in the word segmentation sequences. The content of these text segments is related to the target object, such as text segments in video text used to describe the target object.
[0148] In an exemplary embodiment, the target words mentioned above are place nouns, such as... Figure 7 As shown, the implementation process of step 602 above includes the following steps (6021 to 6023).
[0149] Step 6021: Determine the timestamp corresponding to the location term in the first video.
[0150] The timestamps mentioned above include, but are not limited to, the video frame identifier corresponding to the location term in the first video and the time identifier corresponding to the location term on the timeline of the first video.
[0151] Step 6022: Determine the start and end timestamps corresponding to the word segmentation sequence based on the timestamps.
[0152] Optionally, the statement corresponding to the first word in the word segmentation sequence is obtained, and the start time of the statement is determined as the above-mentioned start timestamp; the statement corresponding to the last word in the word segmentation sequence is obtained, and the end time of the statement is determined as the above-mentioned end timestamp.
[0153] Step 6023: Based on the start and end timestamps, the video text is segmented to obtain text fragments.
[0154] Based on the start and end timestamps corresponding to each word segment sequence, the video text can be segmented to obtain the text segments corresponding to each word segment sequence.
[0155] By segmenting the video text using word segmentation sequences, we can obtain the text fragments corresponding to each word segmentation sequence. Then, based on the correspondence between the word segmentation sequences and the target words, we can identify the text fragments corresponding to the word segmentation sequences as the text fragments corresponding to the target words. This improves the accuracy and comprehensiveness of identifying the text fragments corresponding to the target words, avoids the loss of text information corresponding to the target words, and thus ensures the accuracy of the generated video notes.
[0156] Step 603: Based on the text fragment, determine at least one text information and audio / video information corresponding to the target object.
[0157] Optionally, the name of at least one target object is determined as the title information corresponding to at least one target word. The title of each text segment can be determined as the name of the lowest-level administrative region in the word segmentation sequence, which is the name of the aforementioned target object. Therefore, in this embodiment, the name of the aforementioned target object can be directly determined as the corresponding title information. The aforementioned title information is the title information of the note information corresponding to the target object.
[0158] Optionally, the text fragments are subjected to summary information extraction processing to obtain summary information corresponding to at least one target object.
[0159] The above-mentioned summary information includes, but is not limited to, extractive summary information and generative summary information.
[0160] The process of generating extractive summaries is based on the assumption that the core idea of a document can be summarized in one or a few sentences from the text. Therefore, extractive summarization involves finding the most important sentences in a document, which is essentially a ranking problem. Commonly used methods include TextRank and R2N2 (Residual Recurrent Neural Networks). Among them, TextRank is a graph-based ranking algorithm for text.
[0161] Generative summary information is generated by learning from a large amount of data through encoding and decoding to produce abstract summary content. The source of the summary content is not limited to the original text. Commonly used methods include BERT (Bidirectional Encoder Representations).
[0162] In one possible implementation, text fragments are processed using a deep learning algorithm to obtain generative summary information. Optionally, the generative summary information is typically a summary of a single sentence.
[0163] In another possible implementation, the text fragment is processed by extracting summary information based on the TextRank algorithm to obtain summary information corresponding to at least one target word. The specific process is as follows:
[0164] The process involves segmenting the text fragment into sentences, obtaining at least two corresponding sentences, then performing word segmentation on these sentences to obtain word segmentation results, and finally filtering out stop words from the segmentation results to obtain the words corresponding to the at least two sentences. In simpler terms, this process involves dividing the text into sentences, segmenting the sentences, and removing stop words.
[0165] Based on the words corresponding to at least two statements, determine the statement similarity between each pair of statements in at least two statements.
[0166] In an exemplary embodiment, the TextRank model can be represented by an undirected weighted graph G = (V, E), where V represents the set of nodes (here, the set of sentences) and E represents the set of edges. If two sentences are similar, a weighted edge is considered to exist between the corresponding two nodes, with the weight representing the sentence similarity. In one example, the sentence similarity can be determined by the following formula:
[0167]
[0168] Among them, S i Let S represent the i-th sentence. j Let |S| represent the j-th sentence. i | represents the number of words in the i-th sentence, |S j | represents the number of words in the j-th sentence, |{w k |w k ∈S i &w k ∈S j}| indicates that in sentence S i And in sentence S j The word w in k The quantity of Similarity(S) mentioned above. i ,S j ) represents sentence S i With sentence S j The similarity between corresponding statements can be denoted as W. ji This makes it easier to use later.
[0169] Based on statement similarity, the weights corresponding to at least two statements are determined. In one example, an arbitrary initial value WS is first assigned to each statement in the undirected weighted graph, typically set to 1. The formula for calculating the weights corresponding to statements is as follows:
[0170]
[0171] Among them, V i Indicates the sentence i, V j Let WS(Vi) represent sentence j, and let WS(V) represent the weight of sentence i. j The weight of sentence j is represented by ), where a higher weight indicates greater importance. d is the damping coefficient, ranging from 0 to 1, typically 0.85. The summation on the right-hand side of the equation calculates the contribution of each sentence adjacent to sentence i to sentence i. In(Vi) is the set of sentences pointing to sentence i, Out(Vi) is the set of sentences pointed to by sentence i, and V... k It is sentence k, W in Out(Vi). jk W represents the sentence similarity between corresponding sentences k and j. ji This represents the sentence similarity between corresponding sentences i and j.
[0172] Based on weights, at least two statements are sorted in descending order to obtain the statement sorting result; the top T statements in the statement sorting result are determined as candidate statements, where T is an integer greater than 0; the candidate statements are filtered to obtain the target statement.
[0173] Optionally, the abstract word count condition information and / or abstract sentence quantity condition information are obtained. Based on these conditions, candidate sentences are filtered to obtain target sentences that meet the abstract word count and / or abstract sentence quantity conditions. The abstract word count condition includes a first interval corresponding to the number of abstract words, and the abstract sentence quantity condition includes a second interval corresponding to the number of sentences. If the abstract word count is within the first interval, the target sentence is determined to meet the abstract word count condition; if the number of sentences is within the second interval, the target sentence is determined to meet the abstract sentence quantity condition.
[0174] The above summary information can be generated based on the target statement.
[0175] It should be noted that the embodiments of this application do not limit the method for generating extractive summaries. In addition to the TextRank-based method listed above, other methods such as the PACSUM (Principal Component Analysis Summarization) algorithm, which is improved based on it, can also generate the above-mentioned summary information. In the implementation process, the appropriate summary generation method can be selected according to the specific use case.
[0176] In an exemplary embodiment, the aforementioned text fragment is a narration text fragment, and the text information includes summary information corresponding to the narration text fragment. Accordingly, such as... Figure 7 As shown, the implementation process of step 603 above includes the following steps (6031 to 6032).
[0177] Step 6031: Extract summary information from the narration text fragment to obtain summary information.
[0178] The process for extracting summary information from the aforementioned narration text fragments can be followed as described above, and will not be repeated here. After extracting the summary information from the narration text fragments, the corresponding summary information can be obtained.
[0179] By identifying target words as note titles, the intuitiveness of the note information can be improved. Conversely, by identifying the summary information corresponding to text fragments as note content, the conciseness of the note information can be improved, making it easier for users to view.
[0180] Step 6032: Determine the video segment corresponding to the narration text fragment in the first video.
[0181] Optionally, the start and end times corresponding to the aforementioned narration text fragments are obtained; based on the start and end times, the corresponding video segment in the first video can be determined. The occurrence times of the video frames in the aforementioned video segment fall within the time intervals corresponding to the aforementioned start and end times.
[0182] Step 6033: Determine audio and video information based on video clips.
[0183] In an exemplary embodiment, the aforementioned audio and video information includes at least one type of media information, including but not limited to video information, image information, and audio information associated with the target object in the first video.
[0184] In one possible implementation, after determining the video segment corresponding to the aforementioned text fragment in the first video, the aforementioned video segment can be used as the aforementioned video information.
[0185] In another possible implementation, after determining the aforementioned video segment, a set of video frames associated with the target object within the video segment can be determined; the video frames in the aforementioned set of video frames can serve as the aforementioned image information. Optionally, the video frames in the set of video frames can be filtered to obtain at least one target video frame corresponding to the target object, and the target video frame can serve as the aforementioned image information.
[0186] In one example, when the name of the lowest-level administrative division is mentioned in each video clip, the video frame at the corresponding time point can be identified as the target video frame. If the same name of the lowest-level administrative division is mentioned multiple times in a video clip, the first video frame corresponding to that name can be selected as the target video frame.
[0187] In another possible implementation, the audio data in the video segment can be separated to obtain at least one audio segment corresponding to the target object, and the audio segment can be used as the audio information.
[0188] Step 604: Perform fusion processing on the text information and audio / video information to generate note information corresponding to at least one target object.
[0189] The aforementioned text information includes title information and summary information, and the aforementioned audio and video information includes at least one type of media information. By fusing the title information, summary information, and at least one type of media information corresponding to each target object, note information corresponding to at least one target object can be obtained.
[0190] Accordingly, the aforementioned note information includes the title information, summary information, and at least one type of media information corresponding to the target object in the first video.
[0191] Step 605: Obtain the geographical location information corresponding to at least one target object.
[0192] The aforementioned geographic location information includes the location information of at least one of the aforementioned target objects on the target map.
[0193] Step 606: Generate video notes information corresponding to the first video based on the geographical location information and note information.
[0194] The video notes information is used for display on the first page. This video notes information includes the geographic location information and notes corresponding to at least one of the target objects.
[0195] Optionally, the first page is a video playback page, which can be used to play the aforementioned first video. The aforementioned video playback page can be any page capable of playing videos, including but not limited to client-side pages and browser pages.
[0196] In an exemplary embodiment, after step 606, the video note information can be sent to the corresponding terminal so that it can be displayed on the first page. In some embodiments, to reduce the response time on the terminal side, the video note information can be pre-generated and stored in a database in the server background, so that the server can quickly return the video note information when the terminal requests it. In addition, after the terminal makes its first request, the corresponding map can be cached locally on the terminal to reduce the bandwidth and time of the request.
[0197] In an exemplary embodiment, the target account object can perform error correction and modification processing on the aforementioned first page to generate note correction information corresponding to the note information; the terminal can send the aforementioned note correction information to the server.
[0198] Accordingly, the above method also includes the following process: receiving note correction information published by at least one target account object, and sending the note correction information to the terminal corresponding to the video viewing object, so that the note correction information is displayed in the terminal.
[0199] In an exemplary embodiment, the target account object can perform a preset operation on the note correction information, thereby sending feedback information corresponding to the preset operation to the server; the feedback information includes information indicating agreement to modification, refusal to modify, and ignoring modification of the note correction information. Accordingly, the method further includes the following process:
[0200] Upon receiving a preset number of feedback messages indicating that the target note's error correction information is correct, the server updates the target note information corresponding to the target note's error correction information based on the target note's error correction information. The server then sends the updated target note information to the terminal.
[0201] In one possible implementation, if at least a preset number of consent messages for modifying the target note's error correction information are received, the target note's error correction information can be determined to be correct; if a preset number of rejection messages for modifying the target note's error correction information are received, the target note's error correction information can be determined to be incorrect.
[0202] If at least a preset number of feedback messages indicate that the target note's correction information is incorrect, delete the target note's correction information.
[0203] In summary, the technical solution provided in this application, by obtaining the target object associated with the video and its corresponding target words, can determine the text segment corresponding to the target object in the video. Based on the text information and audio / video information corresponding to the text segment, note information corresponding to the target object can be generated. By combining the note information with the geographical location information corresponding to the target object, video note information displayed on the page can be automatically generated. This increases the amount of information displayed on the page, reduces the operational complexity of users taking video notes while watching videos, improves the efficiency of video note generation and viewing, and optimizes the user experience of taking notes while watching and learning videos.
[0204] Please refer to Figure 8 The diagram illustrates an interactive flowchart of an information display method provided in one embodiment of this application. The method may include the following steps (801-826).
[0205] Step 801: The terminal displays the first page.
[0206] The first page includes the first video. Optionally, the first page is a video playback page used to play the aforementioned first video.
[0207] In an exemplary embodiment, the first page further includes a note display control, and the note information includes text information and audio / video information.
[0208] In step 802, the terminal responds to the selection operation of the note display control by sending a note retrieval request to the server.
[0209] The note retrieval request information includes the video identifier of the first video.
[0210] Accordingly, the server receives the aforementioned note retrieval request information.
[0211] Step 803: The server obtains the video text corresponding to the first video based on the video identifier of the first video.
[0212] Step 804: The server performs word recognition processing on the video text corresponding to the first video to obtain a word sequence.
[0213] The word sequence includes at least one word in the video text that meets preset conditions.
[0214] Step 805: The server obtains word level information corresponding to at least one word.
[0215] Step 806: The server segments at least one word based on word level information to obtain a word segment sequence corresponding to at least one target object.
[0216] The above word segmentation sequence includes the target words. Optionally, the target words are place nouns.
[0217] Step 807: The server determines the timestamp corresponding to the location term in the first video.
[0218] Step 808: The server determines the start and end timestamps corresponding to the word segmentation sequence based on the timestamps.
[0219] Step 809: The server segments the video text based on the start and end timestamps to obtain text fragments.
[0220] Optionally, the above text fragment is a narration text fragment; the above text information includes summary information and title information corresponding to the narration text fragment.
[0221] Step 810: The server extracts summary information from the narration text fragments to obtain summary information.
[0222] Step 811: The server determines the words in the word segment sequence that meet the preset word level conditions as title information.
[0223] Step 812: The server determines the video segment corresponding to the narration text fragment in the first video.
[0224] Step 813: The server determines the audio and video information based on the video clip.
[0225] Step 814: The server performs fusion processing on the title information, summary information, and audio / video information to obtain note information corresponding to at least one target object.
[0226] Step 815: The server obtains the geographical location information corresponding to at least one target object.
[0227] Step 816: The server generates video note information corresponding to the first video based on the geographical location information and note information.
[0228] Step 817: The server sends video note information to the terminal.
[0229] Correspondingly, the terminal receives video notes sent by the server.
[0230] In step 818, the terminal responds to the selection operation for the note display control by displaying map information on the video playback page.
[0231] The selection operation is used to trigger the instruction to display note information. The map display information includes a location identifier corresponding to at least one target object. The location identifier is used to represent geographical location information. The at least one target object includes a first target object.
[0232] Step 819: In response to the trigger operation of the location identifier corresponding to the first target object, the terminal displays the information display control corresponding to the first target object.
[0233] Step 820: The terminal displays the audio and video information and text information corresponding to the first target object based on the information display control.
[0234] Step 821: In response to the selection operation of the information display control, the terminal displays note information corresponding to at least one target object.
[0235] Step 822: The terminal displays at least one note correction message corresponding to the video note information.
[0236] Step 823: In response to a preset operation for the target note correction information in at least one note correction information, the terminal sends feedback information corresponding to the preset operation to the server.
[0237] Accordingly, the server receives the aforementioned feedback information.
[0238] Step 824: When the server receives a preset number of feedback messages indicating that the target note correction information is correct, it updates the target note information corresponding to the target note correction information based on the target note correction information.
[0239] Step 825: The server sends the updated target note information to the terminal.
[0240] Accordingly, the terminal receives the updated target note information.
[0241] Step 826: The terminal displays the updated target note information.
[0242] For a description of each step in this embodiment, please refer to the description in the previous embodiments, which will not be repeated here.
[0243] In summary, the technical solution provided in this application receives a selection operation that triggers note display through a note display control on the video playback page, thereby initiating a note retrieval request to the server. The server, based on the request, can identify and segment the word sequence in the video text. Then, based on the segmentation results of the word sequence, it further segments the video text to obtain the text fragment corresponding to the target object and extracts the summary information of the text fragment. This summary information is then combined with the audio and video information corresponding to the target object in the video to obtain the note information corresponding to the target object. Finally, the note information is combined with the geographical location information of the target object to automatically generate video note information and return it to the terminal. This improves the efficiency of video note generation. The terminal can display an interactive map control on the video playback page and mark the note information of each target word on the map, facilitating user viewing and improving the efficiency of video note viewing. This reduces the operational complexity of recording video notes while watching videos and optimizes the user experience of taking notes while watching and learning videos.
[0244] The following are embodiments of the apparatus of this application, which can be used to execute embodiments of the method of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method of this application.
[0245] Please refer to Figure 9 This diagram illustrates a block diagram of an information display device according to an embodiment of this application. The device has the function of implementing the above-described information display method; this function can be implemented in hardware or by hardware executing corresponding software. The device can be a computer device or can be installed within a computer device. The device 900 may include: a page display module 910 and a note display module 920.
[0246] The page display module 910 is used to display a first page, which includes a first video.
[0247] The note display module 920 is used to display video note information corresponding to the first video on the first page in response to the note information display instruction.
[0248] The video notes information includes geographic location information and notes information corresponding to at least one target object associated with the first video. The notes information is generated based on text information and audio / video information corresponding to text segments of the at least one target object in the first video. The text segments are determined based on target words corresponding to the at least one target object in the first video.
[0249] In an exemplary embodiment, the first page is a video playback page, which includes a note display control; the note display module 920 includes: a map information display unit, an information control display unit, and a note information display unit.
[0250] A map information display unit is configured to display map information on the video playback page in response to a selection operation on the note display control. The selection operation triggers the note information display instruction, and the map display information includes a location identifier corresponding to the at least one target object. The location identifier represents the geographical location information, and the at least one target object includes a first target object.
[0251] The information control display unit is used to display the information display control corresponding to the first target object in response to a trigger operation on the location identifier corresponding to the first target object.
[0252] The note information display unit is used to display the note information corresponding to the first target object based on the information display control.
[0253] In an exemplary embodiment, the note display module 920 is further configured to display note information corresponding to the at least one target object in response to a selection operation on the information display control.
[0254] In an exemplary embodiment, the device 900 further includes: an error correction information generation module and a note information update module.
[0255] The error correction information generation module is used to generate error correction information corresponding to the video note information in response to the editing operation on the video note information.
[0256] The note information update module is used to update the video note information based on the note correction information.
[0257] In an exemplary embodiment, the device 900 further includes: an error correction information display module and a feedback information sending module.
[0258] The error correction information display module is used to display at least one error correction information corresponding to the video note information.
[0259] The feedback information sending module is used to respond to a preset operation for the target note correction information in the at least one note correction information, and send feedback information corresponding to the preset operation, wherein the feedback information is used to characterize the correctness of the target note correction information.
[0260] In summary, the technical solution provided in this application, by obtaining the target object associated with the video and its corresponding target words, can determine the text segment corresponding to the target object in the video. Based on the text information and audio / video information corresponding to the text segment, note information corresponding to the target object can be generated. By combining the note information with the geographical location information corresponding to the target object, video note information displayed on the page can be automatically generated. This increases the amount of information displayed on the page, reduces the operational complexity of users taking video notes while watching videos, improves the efficiency of video note generation and viewing, and optimizes the user experience of taking notes while watching and learning videos.
[0261] Please refer to Figure 10 This diagram illustrates a block diagram of an information generation apparatus according to an embodiment of this application. The apparatus has the function of implementing the aforementioned information generation method; this function can be implemented in hardware or by hardware executing corresponding software. The apparatus can be a computer device or can be installed within a computer device. The apparatus 1000 may include: a video acquisition module 1010, a text fragment determination module 1020, a media information determination module 1030, a note information generation module 1040, a location information acquisition module 1050, and a video note generation module 1060.
[0262] The video acquisition module 1010 is used to acquire a first video, wherein the first video includes at least one target word corresponding to a target object;
[0263] The text segment determination module 1020 is used to determine the text segment corresponding to the at least one target object in the first video based on the target words;
[0264] The media information determination module 1030 is used to determine the text information and audio / video information corresponding to the at least one target object based on the text fragment;
[0265] The note information generation module 1040 is used to fuse the text information and the audio and video information to generate note information corresponding to the at least one target object;
[0266] Location information acquisition module 1050 is used to acquire geographical location information corresponding to the at least one target object;
[0267] The video note generation module 1060 is used to generate video note information corresponding to the first video based on the geographical location information and the note information, and the video note information is used to display on the first page.
[0268] In an exemplary embodiment, the device 1000 further includes: a word sequence generation module, a word level acquisition module, and a word sequence segmentation module.
[0269] The word sequence generation module is used to perform word recognition processing on the video text corresponding to the first video to obtain a word sequence, wherein the word sequence includes at least one word in the video text that meets preset conditions.
[0270] The word / phrase level acquisition module is used to acquire word / phrase level information corresponding to the at least one word / phrase.
[0271] The word sequence segmentation module is used to segment the at least one word according to the word level information to obtain the word segmentation sequence corresponding to the at least one target object, wherein the word segmentation sequence includes the target word.
[0272] In an exemplary embodiment, the text segment determination module 1020 includes: a first timestamp determination unit, a second timestamp determination unit, and a text segmentation unit.
[0273] The first timestamp determination unit is used to determine the timestamp corresponding to the location name in the first video.
[0274] The second timestamp determination unit is used to determine the start timestamp and end timestamp corresponding to the word segmentation sequence based on the timestamp.
[0275] The text segmentation unit is used to segment the video text based on the start timestamp and the end timestamp to obtain the text segments.
[0276] In an exemplary embodiment, the text fragment is a narration text fragment, and the text information includes summary information corresponding to the narration text fragment. The media information determination module 1030 includes: a summary information determination unit, a video fragment determination unit, and an audio and video information determination unit.
[0277] The summary information determination unit is used to extract summary information from the narration text fragment to obtain the summary information.
[0278] A video segment determination unit is used to determine the video segment corresponding to the narration text segment in the first video.
[0279] The audio and video information determination unit is used to determine the audio and video information based on the video segment.
[0280] In summary, the technical solution provided in this application, by obtaining the target object associated with the video and its corresponding target words, can determine the text segment corresponding to the target object in the video. Based on the text information and audio / video information corresponding to the text segment, note information corresponding to the target object can be generated. By combining the note information with the geographical location information corresponding to the target object, video note information displayed on the page can be automatically generated. This increases the amount of information displayed on the page, reduces the operational complexity of users taking video notes while watching videos, improves the efficiency of video note generation and viewing, and optimizes the user experience of taking notes while watching and learning videos.
[0281] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0282] Please refer to Figure 11 It illustrates the structural block of a computer device provided in one embodiment of this application. Figure 1 The computer device can be a terminal. This computer device is used to implement the information display method provided in the above embodiments. Specifically:
[0283] Typically, computer device 1100 includes a processor 1101 and a memory 1102.
[0284] Processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1101 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1101 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1101 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0285] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 is used to store at least one instruction, at least one program, code set, or instruction set, configured to be executed by one or more processors to implement the above-described information display method.
[0286] In some embodiments, the computer device 1100 may optionally include a peripheral device interface 1103 and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface 1103 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1103 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1104, a touch display screen 1105, a camera assembly 1106, an audio circuit 1107, a positioning assembly 1108, and a power supply 1109.
[0287] Those skilled in the art will understand that Figure 11The structure shown does not constitute a limitation on the computer device 1100 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0288] Please refer to Figure 12 It illustrates the structural block of a computer device provided in one embodiment of this application. Figure 2 The computer device can be a server used to execute the information generation method described above. Specifically:
[0289] Computer device 1200 includes a central processing unit (CPU) 1201, a system memory 1204 including random access memory (RAM) 1202 and read-only memory (ROM) 1203, and a system bus 1205 connecting the system memory 1204 and the CPU 1201. Computer device 1200 also includes a basic input / output system (I / O system) 1206 to facilitate information transfer between various devices within the computer, and a mass storage device 1207 for storing the operating system 1213, application programs 1214, and other program modules 1215.
[0290] The basic input / output system 1206 includes a display 1208 for displaying information and an input device 1209 for user input, such as a mouse or keyboard. Both the display 1208 and the input device 1209 are connected to the central processing unit 1201 via an input / output controller 1210 connected to the system bus 1205. The basic input / output system 1206 may also include the input / output controller 1210 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1210 also provides output to a display screen, printer, or other types of output devices.
[0291] Mass storage device 1207 is connected to central processing unit 1201 via a mass storage controller (not shown) connected to system bus 1205. Mass storage device 1207 and its associated computer-readable media provide non-volatile storage for computer device 1200. That is, mass storage device 1207 may include computer-readable media (not shown) such as hard disk or CD-ROM (Compact Disc Read-Only Memory) drive.
[0292] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1204 and mass storage device 1207 described above can be collectively referred to as memory.
[0293] According to various embodiments of this application, the computer device 1200 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1200 can be connected to the network 1212 via the network interface unit 1211 connected to the system bus 1205, or the network interface unit 1211 can be used to connect to other types of networks or remote computer systems (not shown).
[0294] The memory also includes a computer program stored in the memory and configured to be executed by one or more processors to implement the above-described information generation method.
[0295] In an exemplary embodiment, a computer-readable storage medium is also provided, the storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set, when executed by a processor, implements the above-described information display method or information generation method.
[0296] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0297] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned information display method or information generation method.
[0298] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0299] In addition, in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0300] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An information display method, characterized in that, The method comprises: Display a first page, which includes a first video; In response to the instruction to display note information, the video note information corresponding to the first video is displayed on the first page; The video notes information is generated using the following method: Obtain word level information corresponding to words in the video text of the first video, wherein the word level information includes administrative division level information, affiliation level information, and grade level information; The association between adjacent words in the word sequence is determined based on the word level information; based on the association, at least one word is segmented to obtain at least one word segment sequence; The words in the word segmentation sequence that meet the preset level conditions are identified as at least one target object; The process involves determining the text segment corresponding to the at least one target object in the first video; determining the text information and audio / video information corresponding to the at least one target object based on the text segment; and performing a fusion process on the text information and the audio / video information to generate note information corresponding to the at least one target object. Obtain the geographical location information corresponding to the at least one target object; Based on the geographic location information and the note information, video note information corresponding to the first video is generated.
2. The method according to claim 1, characterized in that, The first page is a video playback page, which includes notes display controls; The step of responding to the instruction to display note information and displaying the video note information corresponding to the first video on the first page includes: In response to a selection operation on the note display control, map display information is displayed on the video playback page; wherein the selection operation is used to trigger the note information display instruction, the map display information includes a location identifier corresponding to the at least one target object, the location identifier is used to represent the geographical location information, and the at least one target object includes a first target object; In response to a trigger operation on the location identifier corresponding to the first target object, an information display control corresponding to the first target object is displayed; Based on the information display control, the note information corresponding to the first target object is displayed.
3. The method according to claim 2, characterized in that, The method further includes: In response to a selection operation on the information display control, note information corresponding to the at least one target object is displayed.
4. The method according to claim 1, characterized in that, The method further includes: In response to the editing operation on the video note information, note correction information corresponding to the video note information is generated; The video notes information is updated based on the error correction information in the notes.
5. The method according to claim 1, characterized in that, The method further includes: Display at least one note correction information corresponding to the video note information; In response to a preset operation on the target note correction information in the at least one note correction information, feedback information corresponding to the preset operation is sent, the feedback information being used to characterize the correctness of the target note correction information.
6. An information generation method, characterized in that, The method comprises: Get the first video; Obtain word level information corresponding to words in the video text of the first video, wherein the word level information includes administrative division level information, affiliation level information, and grade level information; The association between adjacent words in the word sequence is determined based on the word level information; based on the association, at least one word is segmented to obtain at least one word segment sequence; The words in the word segmentation sequence that meet the preset level conditions are identified as at least one target object; Determine the text segment corresponding to the at least one target object in the first video; Based on the text fragment, determine the text information and audio / video information corresponding to the at least one target object; The text information and the audio / video information are fused together to generate note information corresponding to the at least one target object; Obtain the geographical location information corresponding to the at least one target object; Based on the geographic location information and the note information, video note information corresponding to the first video is generated, and the video note information is used to display on the first page.
7. The method according to claim 6, characterized in that, The method further includes: The video text corresponding to the first video is subjected to word recognition processing to obtain the word sequence, which includes at least one word in the video text that meets preset conditions.
8. The method according to claim 7, characterized in that, Determining the text segment corresponding to the at least one target object in the first video includes: Determine the timestamp corresponding to the location name in the first video; Based on the timestamps, determine the start and end timestamps corresponding to the word segmentation sequence; Based on the start timestamp and the end timestamp, the video text is segmented to obtain the text fragments.
9. The method according to claim 6, characterized in that, The text fragment is a narration text fragment, and the text information includes summary information corresponding to the narration text fragment. The step of determining the text information and audio / video information corresponding to the at least one target object based on the text fragment includes: The narration text fragment is processed to extract summary information, thereby obtaining the summary information; Determine the video segment in the first video that corresponds to the narration text fragment; Based on the video segment, the audio and video information is determined.
10. An information display device, characterized in that, The device includes: A page display module is used to display a first page, which includes a first video. The note display module is used to display video note information corresponding to the first video on the first page in response to the note information display instruction; The video notes information is generated using the following method: Obtain word level information corresponding to words in the video text of the first video, wherein the word level information includes administrative division level information, affiliation level information, and grade level information; The association between adjacent words in the word sequence is determined based on the word level information; based on the association, at least one word is segmented to obtain at least one word segment sequence; The words in the word segmentation sequence that meet the preset level conditions are identified as at least one target object; The process involves determining the text segment corresponding to the at least one target object in the first video; determining the text information and audio / video information corresponding to the at least one target object based on the text segment; and performing a fusion process on the text information and the audio / video information to generate note information corresponding to the at least one target object. Obtain the geographical location information corresponding to the at least one target object; Based on the geographic location information and the note information, video note information corresponding to the first video is generated.
11. The apparatus according to claim 10, characterized in that, The first page is a video playback page, which includes notes display controls; The note display module includes: a map information display unit, an information control display unit, and a note information display unit; A map information display unit is configured to display map display information on the video playback page in response to a selection operation on the note display control; wherein the selection operation is used to trigger the note information display instruction, the map display information includes a location identifier corresponding to the at least one target object, the location identifier is used to represent the geographical location information, and the at least one target object includes a first target object; The information control display unit is used to display the information display control corresponding to the first target object in response to a trigger operation on the location identifier corresponding to the first target object. The note information display unit is used to display the note information corresponding to the first target object based on the information display control.
12. The apparatus according to claim 11, characterized in that, The note display module is also configured to display note information corresponding to the at least one target object in response to a selection operation on the information display control.
13. The apparatus according to claim 10, characterized in that, The device further includes: an error correction information generation module and a note information update module; The error correction information generation module is used to generate error correction information corresponding to the video note information in response to the editing operation on the video note information; The note information update module is used to update the video note information based on the note correction information.
14. The apparatus according to claim 10, characterized in that, The device further includes: an error correction information display module and a feedback information sending module; The error correction information display module is used to display at least one note error correction information corresponding to the video note information; The feedback information sending module is used to respond to a preset operation for the target note correction information in the at least one note correction information, and send feedback information corresponding to the preset operation, wherein the feedback information is used to characterize the correctness of the target note correction information.
15. An information generation device, characterized in that, The device comprises: The video acquisition module is used to acquire the first video. The word level acquisition module is used to acquire word level information corresponding to words in the video text of the first video. The word level information includes administrative division level information, affiliation level information, and grade level information. The word sequence segmentation module is used to determine the corresponding relationship between adjacent words in the word sequence based on the word level information; to segment at least one word based on the relationship to obtain at least one word segmentation sequence; and to identify at least one target object as a word whose word level information in the word segmentation sequence meets the preset level conditions. A text fragment determination module is used to determine the text fragment corresponding to the at least one target object in the first video; The media information determination module is used to determine the text information and audio / video information corresponding to the at least one target object based on the text fragment; The note information generation module is used to fuse the text information and the audio and video information to generate note information corresponding to the at least one target object; A location information acquisition module is used to acquire the geographical location information corresponding to the at least one target object; The video note generation module is used to generate video note information corresponding to the first video based on the geographical location information and the note information, and the video note information is used to display on the first page.
16. The apparatus according to claim 15, characterized in that, The device further includes a word sequence generation module, used to perform word recognition processing on the video text corresponding to the first video to obtain the word sequence.
17. The apparatus according to claim 16, characterized in that, The text segment determination module includes: a first timestamp determination unit, a second timestamp determination unit, and a text segmentation unit; The first timestamp determination unit is used to determine the timestamp corresponding to the location term in the first video. The second timestamp determination unit is used to determine the start timestamp and end timestamp corresponding to the word segmentation sequence based on the timestamp. The text segmentation unit is used to segment the video text based on the start timestamp and the end timestamp to obtain the text segments.
18. The apparatus according to claim 16, characterized in that, The text fragment is a narration text fragment, and the text information includes summary information corresponding to the narration text fragment. The media information determination module includes: a summary information determination unit, a video fragment determination unit, and an audio and video information determination unit. The summary information determination unit is used to extract summary information from the narration text fragment to obtain the summary information. A video segment determination unit is used to determine the video segment corresponding to the narration text segment in the first video; The audio and video information determination unit is used to determine the audio and video information based on the video segment.
19. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the information display method as described in any one of claims 1 to 5, or the information generation method as described in any one of claims 6 to 9.
20. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the information display method as described in any one of claims 1 to 5, or the information generation method as described in any one of claims 6 to 9.
21. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform an information display method as described in any one of claims 1 to 5, or an information generation method as described in any one of claims 6 to 9.
Citation Information
Patent Citations
Text labeling method of streetscape video, terminal equipment and storage medium
CN108108443A
Method and device for extracting video theme text, equipment and storage medium
CN113395578A