A method and system for multimedia control
Patent Information
- Application Number
- CN202310579085.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-22
- Filing Date
- 2023-05-22
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-05-22
AI Technical Summary
[0003]有很多广泛需求的场景,比如老师在视频教学时,需要反复播放核心内容,则这个过程中,老师需要反复在视频中找片段播放起始位置,这样播放并不精准高效,容易导致学生注意力不能集中,影响教学效果;如在警员现场处置时,对一段证据视频在检出时需要反复看时,如果不能高效的从核心内容开始,自然导致工作效率低,特别是路面事故现场处置时,效率低,直接影响交通恶化;此外在警方调查时,往往需要从多媒体文件中获得证据信息,以现在的播放控制技术逻辑,反复从视频中找到信息的效率低;此外在视频编辑中,特别广泛使用的智能终端编辑视频时,也存在快速、精确地控制的诉求不能达成的问题,而基于当前常态的多媒体播放控制,无疑导致效率低下,体验不佳,特别在多点触控的智能终端情形下,更是效率低下,上述播放控制在学习如英语、学习唱歌等环节,问题如上所述,也是广泛存在
[0016]结合第二及第三方面,采用视频微调方法,则进一步提升了标记或者进度条的精度与执行时的效率问题。
Smart Images

Figure CN117241096B_ABST
Abstract
Description
Technical Field
[0001] Multimedia control, playback, editing, human-computer interaction, multimedia processing technologies and systems, computer software or smart terminal apps, video fine-tuning, precise positioning, multimedia player control technologies, or systems incorporating the aforementioned technologies. Background Technology
[0002] In today's fast-paced life, traditional methods of video and audio playback and editing have not kept up with the times. They are merely simple digitizations based on traditional tape playback control logic, resulting in efficiency similar to that of the traditional tape era, poor user experience, and low efficiency. The technology and embodiments disclosed herein illustrate and solve problems that have not been discovered or solved for decades, thereby improving user control efficiency and experience.
[0003] There are many scenarios where this is needed. For example, when teachers are teaching via video, they need to repeatedly play the core content. However, this requires repeatedly finding the starting point of each segment in the video, which is inaccurate and inefficient, easily causing students to lose focus and affecting teaching effectiveness. Similarly, when police officers are handling on-site incidents, repeatedly reviewing evidence videos is inefficient if they cannot efficiently start from the core content, especially in handling road accident scenes, where inefficiency directly worsens traffic conditions. Furthermore, during police investigations, evidence information is often obtained from multimedia files, and current playback control technology makes repeatedly finding information from videos inefficient. In video editing, especially with the widespread use of smart terminals, the need for fast and precise control cannot be met. Current multimedia playback control methods undoubtedly lead to inefficiency and a poor user experience, especially with multi-touch smart terminals. These playback control issues are also widespread in learning activities such as English and singing.
[0004] Although the prior patent application (patent number: 2023104357360) included fine-tuning multimedia control technology, the inventors still found the above-mentioned problems to be solved during engineering implementation. Therefore, in conjunction with the partial disclosure of the prior application, this disclosure is submitted. In addition, the progress bars of traditional video, audio and other multimedia playback are not optimized and do not meet the above requirements. For example, they are simply dragged or have certain positioning points, but in reality, they are far from the effect achieved in solving the above problems. This disclosure is used to solve the above-mentioned multiple problems, so that the problems that have not been solved for many years have a concrete and implementable technical solution. The multimedia files in this application include audio files and video files. Summary of the Invention
[0005] In a first aspect, this disclosure provides a method for pausing multimedia playback, characterized by comprising: if a system containing multimedia playback control detects a trigger for pausing playback, reading the position data of the currently playing multimedia file; calculating a rollback value based on a delay value and playback speed; calculating a pause position based on the rollback value, playback direction, and the position data; and pausing playback at the pause position.
[0006] In conjunction with the method in the first aspect, different methods are used to trigger pause in different embodiments, such as triggering pause by any one of the following: physical buttons, control keys in the graphical user interface, menus, or gestures.
[0007] In conjunction with the method of the first aspect, in some embodiments, the delay value is calculated based on the total duration including human eye and ear recognition of multimedia, brain reaction, brain-controlled hand operation, and the system's recognition of pause triggers.
[0008] In conjunction with the method of the first aspect, in some embodiments, the delay value is preset by the system or input by the user from a human-computer interaction interface.
[0009] In conjunction with the method of the first aspect, in some embodiments, the location data includes at least one of time values or frame sequence values (video files include frames); while for audio files, time values are required.
[0010] Secondly, this disclosure provides a method for multimedia tagging, applicable to video or audio multimedia, in which users generate tags by triggering instructions or functions for multimedia tagging; The generation of the markers is accomplished by the following steps: Get the current position of the multimedia file being played; input or select playback control parameters in the human-computer interaction interface; generate a marker in the human-computer interaction interface; If the generated marker is subsequently triggered or clicked by the user, the playback system will play the game according to the location value and playback control parameters.
[0011] In some embodiments of the second aspect, if a fine-tuning method is applied during position marking, wherein the fine-tuning method includes continuous fine-tuning and discontinuous fine-tuning.
[0012] In some embodiments of the second aspect, the pause method described in the first aspect and the techniques in conjunction with embodiments of the first aspect are applied to achieve precise pause, thereby improving the accuracy of the marking.
[0013] In conjunction with some embodiments of the second aspect, by clicking two markers in a set instruction window, a marker for a multimedia segment between the two markers is generated, and a human-computer interaction window is displayed for the user to select or input playback control parameters; when the system listens to the playback of the multimedia segment, it plays the multimedia of the segment according to the selected or input playback control parameters, that is, a segment marker is formed, thereby making it more convenient for the user to control the media to meet different usage purposes.
[0014] In conjunction with the second aspect of the embodiment, in the case of multiple tags, image tags are used to reduce confusion. Specifically, the image is extracted from the position value at which the multimedia file is played.
[0015] Thirdly, this disclosure provides a multimedia playback progress bar that applies the marking method of the second aspect, so that users can use the progress bar more flexibly. The progress bar is no longer limited to traditional playback control and playback information such as the display of the current position. The progress bar provided in the third aspect includes the second aspect and the embodiments combined with the second aspect, such as image markers, paragraph markers, etc. When the user clicks on the marker, the effect of quick control is achieved.
[0016] Combining the second and third aspects, the video fine-tuning method further improves the accuracy of markers or progress bars and the efficiency of execution.
[0017] In addition, by combining the methods from the first aspect and the real-time example, a precise pause button is added to the progress bar, thus enabling the progress bar to have the ability to pause precisely.
[0018] Fourthly, this disclosure provides multimedia fine-tuning techniques, including continuous fine-tuning, non-continuous fine-tuning, and differential fine-tuning methods. When combined with the methods or embodiments provided in the first aspect, the fine-tuning steps are reduced and efficiency is improved.
[0019] In conjunction with the fourth aspect, in some embodiments, when marking, i.e., the second aspect and related embodiments, the user experience is improved, such as by directly marking keyframes, thereby marking media segments more accurately and quickly, and playing according to tags.
[0020] Fifthly, a multimedia control system is provided, which includes one or more of the above aspects in combination. The system is included in computer programs, smartphone apps or terminal apps, display systems of smart electronic devices such as smart TVs, XR display devices, smart XR helmets, automobiles, and other vehicles, as well as various electronic devices and devices combining electronics and mechanics. These systems include a multimedia control system. Attached Figure Description
[0021] Figure 1A diagram illustrating the method for marking multimedia playback. Figure 2 This is a diagram illustrating continuous adjustment during video playback. Figure 3 This is a schematic diagram of a multimedia playback progress bar with markers. Figure 4 A schematic diagram of the video playback program interface (including markings); Figure 5 A diagram illustrating the method for pausing; Figure 6 This is a diagram illustrating the method of fine-tuning video control. Figure 7 This is a schematic diagram of the video fine-tuning interface; Detailed Implementation
[0022] The specific implementation methods, codes, quantities, values, durations, triggering methods, diagrams, etc., described in the following disclosure, description, or exemplary embodiments do not represent all implementation methods or embodiments consistent with those provided in this application. On the contrary, they are merely some methods, systems, apps, etc., consistent with this disclosure as described in the appended claims.
[0023] The technology and methods of this disclosure will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the embodiments described below are only some embodiments of this disclosure, and not all embodiments. In addition, in the objective of multimedia control, on the one hand, it is based on the inventor's artistic and technical subject to reflect the modern human beings in real life to use innovative technologies to complete tasks that could only be completed in static environments and under tool-dependent conditions in the past. At the same time, it is also to improve the efficiency of existing playback control and user experience, so as to enable multimedia applications in various scenarios, such as using XR display equipment, such as playing multimedia information provided by others in sports, combat and other situations. It is also the inventor's continued innovation based on previous technical inventions, so as to further improve the efficiency of media playback and control, use faster and more accurate marking, use fine-tuning technology for marking, and innovative use of marking methods and pausing methods in multimedia applications, thereby improving playback and control efficiency. In fact, it also solves the problems that have been neglected by the industry for many years. In addition, the combination of fine-tuning, marking and pausing methods helps each other in terms of efficiency and accuracy, so as to improve the overall effect and the user experience.
[0024] In multimedia applications such as teaching, song and music imitation learning, martial arts movement learning, craft content video learning, magic trick solving, video evidence finding, VR, MR (XR), etc., users need to repeatedly play one or more segments of content. Based on the characteristics of the content, playback needs to be adjusted according to the user's needs at different speeds. However, traditional technology leads users to roughly remember the position and time on the progress bar, control the progress bar to roughly reach that position, and then play it repeatedly.
[0025] However, in reality, for magic videos, the key information is only about one or two frames. Therefore, the precision of a single point on the progress bar can be several minutes or even seconds (depending on the length of the video and the progress bar). This makes it difficult to pinpoint the exact playback point when replaying, resulting in a poor user experience. Constantly controlling the multimedia progress bar is also a waste of time. In addition, the playback requirements vary for different multimedia content. For example, key information may need to be played slowly, while evidence information may need to be played at 0.1x speed. In teaching, for example, if a teacher wants to improve students' listening skills, they may play a segment of English listening content at 1.25x speed, or after reaching a certain level at 1.25x speed, play a selected segment of the key multimedia file at an even faster speed to improve students' listening skills.
[0026] Current multimedia playback tools can only be played by manually adjusting the progress bar to the approximate target position and then selecting the player's very limited speed function. Therefore, they are inefficient and involve many steps. The problems are obvious in scenarios such as work, teaching, and entertainment, but they have not been changed or effectively solved by the industry, or they may have been completely ignored by the industry (due to the bias caused by the technical logic of the magnetic tape era).
[0027] To address this long-standing problem, this disclosure provides a method for multimedia tagging, such as... Figure 1 As shown, Figure 1 In this context, the M01 step triggers a multimedia tag command or function. The triggering method varies across different systems. For example, on traditional computers, the most common way to trigger this function is to click a "key" on the human-computer interaction interface. This key is usually defined as a "button" in the program, or a menu item, clickable visual element such as a "Label," etc. Users can activate the M01 command or the corresponding program or system function by clicking with a mouse or touchscreen. In addition, in traditional computer programs, frequently used functions are activated by keyboard input, such as using shortcut function keys or specially defined key combinations, such as Ctrl+A, etc., using a set keyboard combination to activate the M01 command or function.
[0028] In smartphones and other smart terminals, the corresponding method is usually to use touch or multi-touch, and click on the control keys defined on the human-computer interaction interface / graphical user interface (GUI), such as clickable "button" or "label", to activate the corresponding M01 function or command; of course, the embodiments of this disclosure also include the use of gesture triggering method or trigger coding method to activate M01 command or function.
[0029] For dedicated multimedia systems, in addition to the methods mentioned above such as mouse, keyboard, and touch screen, specially defined keys may also be used, such as the editing console of a professional video editing system, which uses a corresponding key to start.
[0030] For XR systems such as MR and AR, it can be defined as a gesture triggering method or a sensor triggering method, thereby activating and starting the marking function in the video file displayed on such display devices; of course, it can also be defined as "control keys" such as clickable buttons, labels and other visual resources, which can be triggered by gestures, fingers, etc.
[0031] When the user clicks on the physical keyboard or the virtual "control key", the M01 instruction or function is triggered. After M01 is executed, the mark generation step M02 is executed, thus completing the mark generation.
[0032] The M02 marker generation step comprises three steps, which combine to generate the marker. Step M021 involves obtaining the time value or frame sequence value of the current playback position in the multimedia file (audio multimedia files do not contain frame sequence values, so time values are used). When a user plays a multimedia file, a time value is generated after playback begins, indicating the current position (usually in milliseconds) of the media file. In video playback, this time value corresponds to a frame sequence value. Obtaining the time value or frame sequence value of the current playback position allows the user to determine if they have selected a position in the multimedia file through playback, progress bar control, or thumbnail control (referred to as a thumbnail in video processing). Marking the multimedia file requires this value; therefore, when the user confirms the mark, they are essentially obtaining the specific position of the mark point in the media file in terms of time or frame sequence. Figure 2In this context, assuming tx is a time value for a location to be marked, and if tx is a video file, it can have a corresponding frame sequence value such as fx, then for the user, the choice between time marking or frame marking depends on the multimedia file type or the convenience of program execution. After marking is completed, when the user clicks the mark again, the time or frame sequence value obtained in this step is the starting point for playback control in the media file after the marked point is clicked. The control command executed at that point is the next action, step M022, which is the user's selection or input control parameter.
[0033] M022 is what appears on the human-machine interface (HMI) after the user executes the M01 command or function. It can appear as a pop-up menu or window. The HMI contains fixed selection options, such as playback speed (e.g., 0.2x, 0.75x, 1x, 1.25x, 2x), or playback control functions like pause. For example, in teaching, a teacher might explain key points before playing a video, and then the user can click the play button to resume playback at that point. The solution for this type of requirement is to reach the marked point and pause. In addition, there are users who, based on their needs, may not have fixed options, such as a video requiring a very slow speed, such as 0.1x or 0.05x playback. In this case, when marking, the user can enter 0.1x speed in the input field or input human-computer interaction method of the marked point in the human-computer interface of step M022. Therefore, in order to improve the functionality and effect of marking, in addition to fixed options, there is also the input of user-defined content, such as the user entering their desired speed in the input field as mentioned above.
[0034] In addition, it also includes deletion, such as when the marker is accidentally selected, the user can delete it in the human-computer interface, that is, delete the marker.
[0035] Furthermore, unlike traditional methods, such as a progress bar in multimedia playback that can only mark a position at one point, this disclosure allows for multiple markings at a single position. For example, a segment can be marked for playback at 0.8x speed, while simultaneously marked for playback at 1x and 1.2x speed. In English listening training, teachers can control playback from easy to difficult using different speeds. Teachers only need to click on different markings. Therefore, from the same position, clicking on different markings with different selected input parameters allows for playback at different speeds, thus improving training efficiency.
[0036] There are many forms of human-computer interaction interfaces, and this embodiment is not limited to any particular form. For example, pop-up windows, pop-up menus, and display layouts (called layouts in the program) are all possible. Users select the clickable, selectable, and inputtable content and parameters on these human-computer interaction interfaces to determine the instruction to be executed after the mark is clicked. For convenience, human-computer AI interaction methods, such as voice recognition, can also be used. The user inputs voice, which is recognized as specific execution parameters, such as the speed value being determined as 1.11 times the speed.
[0037] Through steps M021 and M022, we obtained the location of the marker in the multimedia file and the playback command parameters to be executed after clicking the marker. We then discovered an indispensable step, M023, which is the generation of visual markers. Users can click on this visual marker to execute parameters, commands, or functions selected or entered at the marked location, such as playing backwards at 0.2x speed at the tx location.
[0038] Specifically, such as in Figure 1 As described above, the user of the multimedia system triggers a pre-set multimedia playback marker (step M03), and then the multimedia system executes step M04, which is to play according to the time value or frame sequence value (position value or location value) recorded at the time of marking, and the selected or input parameters. That is, it starts from the position value recorded at the time of marking and plays according to the input parameters, thereby realizing the linkage between marking and playback, thus improving the efficiency of multimedia playback control.
[0039] Because a mark is a visual element, its forms of expression are very diverse, and this disclosure does not limit them; it is merely an example. Figure 3 In the diagram, M10 is a progress bar commonly used in multimedia playback, and M11 is the current position of the media being played (time value or frame sequence value, usually time value). To clearly indicate the playback progress to the user, the left and right sides of M11 are usually represented by different colors. For example, the left side of M11, which has already been played, is in a light color, while the right side of M11 is in a dark or light color.
[0040] In the diagram, M101 to M105 are all marker styles on the progress bar. For example, M101 is the position of the media file when the user triggers the M01 command or function. This triggers the marker, retrieves the position value of the media file (such as time value, frame sequence value, and the corresponding frame image in the multimedia file), and extracts this image at a set quality value, such as 20%, and displays it in the square box of M101. This way, when there are multiple markers, the user can determine the marker to process or click by looking at the image in M101 (the video image extracted from the corresponding position in the video), thus avoiding confusion when there are multiple markers. As mentioned earlier, at the same marker position... Different control requirements necessitate different playback methods. For example, M101 is a marker that should play at normal speed. However, as students reach the teaching requirements, the teacher clicks on a marker that plays at 1.2x speed, such as M102. The frame displays the image of the frame at that position when the marker is set, along with the speed information. Alternatively, P1 can be added to indicate the first positioning marker point, allowing users to accurately distinguish between them. When multiple markers are set at the same position, traditional progress bars cannot be used for marking, nor can parameters be confirmed or input based on the markers, nor can playback be performed according to parameters or instructions. Therefore, the above method overcomes the efficiency and functional deficiencies of traditional logic.
[0041] M103 is another form of marking, such as using icons and marking order, like P1 representing Position 1, which marks position 1, and the Nth mark can be Pn. This method does not require the image information of the corresponding frame, such as in music and other multimedia files. M104 is a simpler marking method, which uses text descriptions, such as the playback control parameters of the point, where the text "0.25X" represents playback at 0.25x speed.
[0042] The visual elements of the marker can have many styles, and this disclosure is not limited to any particular style. The examples above are simply used to illustrate the method and the new ability to control the progress bar. Specifically, a visual element follows the marker; when the user clicks (triggers) this element, the M04 function is executed, and the trigger is step M03. Furthermore, in some applications, it may not be directly associated with the progress bar but displayed in a separate area, such as... Figure 4 The S301 in the video uses larger image information extracted from the video as markers (left side), and can also add text labels (right side) so that users will not accidentally click on these markers (images, text, etc.) with prompts. That is, the image in the video is extracted according to the position value of the marker. Since the space is relatively large, the marked image is generally larger than the progress bar, so the recognition effect can be clearer.
[0043] Throughout the above description, it has been repeatedly emphasized that these markers are clickable. Clicking such a marker triggers the playback control program to execute the position used when the marker was generated, along with the parameters selected and input by the user (such as playback speed). The media player then moves from its current position to the position indicated by the marker and plays the content according to the selected or input parameters. This avoids the shortcomings of existing playback technologies that require repeatedly searching for the correct position and allows playback at a set speed, thus reducing the number of operation steps. For teaching, this can reduce children's attention span issues (currently, teachers frequently operate the system, making it easy for children to lose focus). For children learning activities such as origami, they can also control the video playback themselves, whereas otherwise, adult assistance would be needed.
[0044] For multiple markers, which essentially correspond to a certain range of a media file, such as M101 and M105, the area between them represents multimedia information within a multimedia file. If a user wants to mark this multimedia segment, they can click M101, and within the command window period (e.g., setting the window value to 800ms), click M105. Therefore, the system recognition program will process the content between the two markers, such as the multimedia segment marker. When both markers are clicked within the command window period, the human-computer interaction interface, windows, menus, etc., for processing the multimedia segment will be displayed, including options such as loop playback, speed value, and SKIP (skip, meaning the player skips this segment). The system offers various control functions, such as skipping the information between two markers. For example, if a segment in a multimedia file contains an advertisement, useless information, or unsuitable content, the first user uses segment markers (marking the start and end positions, and selecting / inputting the content and parameters for the segment). In subsequent playback, if the segment is shown to a child, it will be skipped. When using segments as markers, basic speed selection and parameter input are available, such as looping a segment at 0.5x speed. This is especially important for protecting a child's enthusiasm when learning to imitate more complex steps, such as origami.
[0045] In multimedia playback programs, to facilitate the multiple uses of a file for segment processing as described above, additional functionalities are needed, such as: recording tags and segment tags, including position values and playback parameter values; recording information about the tagged file, such as file name, file length (time value), and file media information (such as information in the media file's encapsulation). This allows the playback system to first determine whether the file has been previously tagged based on one or more of the above information when the user opens the video file again. If it has been tagged, the system retrieves the previous tagging. During playback, the media playback system monitors the current playback position and compares it with the segment data based on the position. The system marks the start and end positions and then plays the video based on the input parameters. If these marks are set in a pre-defined location and format in a multimedia file and recorded in the video file, it becomes a universal video player function. Therefore, filtering is performed during playback. For example, if a movie contains violent content, it may be marked by others as unsuitable for viewers under 13 years old when the source is provided. Therefore, during playback, the player will prompt the user whether they are under 13 years old. If they select yes, the marked content segment will be automatically skipped by the player. This allows the universal player to combine with the definition file in the video file or the file accompanying the video file to perform mark-based playback.
[0046] In one example, the marked segment is from 10000ms to 20000ms. When the player starts playing and reaches 10000ms, it determines that the next segment from 10000ms to 20000ms has been marked. Therefore, it checks the mark parameter and content in the record. If the playback parameter is skip, the player jumps directly to the 20000ms position to continue the previous playback. If the marked segment is slow playback with a speed parameter of 0.2x, it plays at that speed. If the segment is a loop, after playing to the 20000ms position, it automatically returns to 10000ms to continue playback until the user pauses or commands to exit the loop.
[0047] Therefore, whether the marked data is stored locally for playback control or written in a multimedia file, it is an implementation of this disclosure. This disclosure does not specify the content, but in terms of effect, based on the marked records, playback control is performed according to the parameters and position values of the marked data.
[0048] In addition, users can also trigger markers to execute steps such as M03 and M04. This playback control and recording method is particularly effective for teachers preparing and teaching multimedia videos, greatly improving efficiency. Since teachers need to teach multiple classes, a single marker can be used multiple times in the same teaching manner, making it far more efficient than current teaching multimedia playback systems.
[0049] The following explanation is based on triggering the M01 command. During multimedia playback, whether it is video or audio, the file is being played, whether it is local or remote, and is played in the form of a stream on the terminal (the file is localized or stored in memory or cache as a segment or the whole file). To trigger the M01 command, as mentioned above, you can use the mouse to click the "button", the keyboard, or the touch screen to click virtual buttons or control keys. However, the precision of the marking can be divided into two main categories: precise and imprecise. For example, if the user triggers the M01 command and starts marking at the approximate position during multimedia playback, such as when the user feels that the position is about right for the video being played, then the marking using any of the above or other methods is considered imprecise marking. For precise marking, such as repeatedly viewing evidence segments or selecting content during video editing, it needs to be implemented in the following way so that the marking precision can locate the appropriate frame of the specific multimedia file or the specific syllable of the audio file. Therefore, in order to accurately mark the pause, we will first explain a method for providing an efficient and precise pause, which is an improvement over current playback control technologies such as marking and fine-tuning.
[0050] Specifically, as described below, the current playback pause function, whether activated by pressing a button, clicking a button in the GUI, or touching the playback interface on the touchscreen, pauses playback until the user inputs the next command. However, this pause method is essentially lagging behind the user's perceived pause position, resulting in a low-precision pause. This leads to the problem that in image editing and trimming, it is necessary to use commands in the opposite direction of playback to retrieve several frames or milliseconds, undoubtedly reducing efficiency, especially when precisely handling multimedia content or markers.
[0051] For example, if a video is played at 1.2x speed and has a frame rate of 30 frames per second, when a user watches the video and realizes that the content needs to be paused (the image travels from the eyes to the brain, or the audio travels from the ears to the brain, then the brain recognizes and determines the action), and then the brain controls the hand to press the pause button (whether physical or virtual), there is approximately a delay of 50-100ms for the eyes, a brain reaction time of about 200-400ms (seeing the video information, changes in brain waves), plus the time for the brain-controlled hand to trigger the button (100-300ms) for the system to recognize and complete the command, which is about 70-100ms. Therefore, in video control, with current technology, if the delay is between 500ms and 1000ms, at normal speed (… At 1x speed, there is a deviation of approximately 15 to 30 frames from the point where the pause is needed. At 1.2x speed, there is lag and a deviation of 18 to 36 frames (normal frame rate * 1.2). Therefore, in order to make the pause closer to the point where the user wants to pause under different playback speeds, and to reduce the duration and number of fine-tuning triggers, the following pause method is adopted to improve the efficiency of video fine-tuning.
[0052] First, the program listens for playback pause triggers. As mentioned earlier, playback pause can be triggered in various ways, such as step M201. After M201 is detected, the system program reads the current playback position of the multimedia file. This position can be a time value or a frame sequence value. For audio files, it is usually a time value, while for video files, it can be either a time value or a frame sequence value. This application does not limit this; all are position values. Based on the current playback speed and the set human reaction deviation time, such as the 500-1000ms hysteresis and deviation time mentioned above, the frame rate of the multimedia needs to be obtained during video playback to calculate the backoff value when using the frame sequence as the calculation method. In the following embodiments, time values are used for description: Suppose a media file is 40 minutes long and plays at a normal speed (1x). Suppose the current position m1 is 10 minutes 30 seconds 300 milliseconds. Step M201 detects a pause trigger and therefore reads this value as m1 (i.e., step M202, obtains the current position value of the multimedia file). Of course, we know that when a user triggers a pause, it is usually a delayed action. It is a process in which the human body perceives, the brain recognizes, the brain judges, and then the hand is activated to operate. This process usually lags by several hundred milliseconds. Let's assume this time is set to 500ms. Then, the relatively accurate pause position should be m1 time minus 500ms. So when the playback speed is 1x, the pause can be set to the current pause position value m1 minus the delay value (the delay value includes the lag value of perception, brain recognition, brain judgment, brain control of hand movements, and system program).
[0053] For playback at different speeds, this latency value needs to be further calculated based on the playback speed. For example, when the playback speed is 1.2x, a latency of 500ms means that 500 * 1.2 = 600ms has actually been played for a multimedia file. If the playback speed is 1.5x, then 500ms * 1.5 = 750ms. Therefore, when calculating based on time, it is necessary to consider the playback speed of the media and the latency or hysteresis value to obtain a relatively accurate backoff value, which corresponds to the calculation position in the multimedia file. For playback speeds slower than normal, such as 0.5x speed, the latency value is assumed to be 500ms. In reality, the user should pause at the pause value in the media file minus 500ms * 0.5 = 250ms (rollback value). Therefore, the rollback value in step M203 is 150ms.
[0054] The above time values illustrate how the backoff value is calculated and determined based on the delay or hysteresis value and playback speed. M204 calculates the pause position and pauses at that position. It calculates the actual pause position based on the obtained m1 position value, adding or subtracting a backlash value according to the playback direction (usually forward playback, but reverse playback is also common in digital playback in the future). For example, in forward playback, the backlash value is subtracted from the m1 value, while in reverse playback, the backlash value is added to the m1 value, thus forming a new pause position value. The playback system program then adjusts to this pause position value and pauses (that is, returns to the visual and auditory recognition position that the user initially identified and believed needed to pause).
[0055] In the above explanation using time values as an example, the pause position is adjusted as close as possible to the actual position when the user wants to pause, based on the delay value of human operation and playback speed. This reduces the need for subsequent fine-tuning operations and improves efficiency.
[0056] This precise pause method is adopted because when marking multimedia files, especially when marking precision, it is necessary to locate certain frames or a precise point in the content. In addition, during video editing, because the actual position always deviates by several frames after triggering the pause, the user needs to reposition it using editing tools, which is inefficient. Therefore, a new pause method was proposed.
[0057] Users can customize the delay or hysteresis value for precise pause. For example, the elderly react relatively slowly, so it can be defined as 800ms, while children react relatively quickly, so it can be defined as about 400ms. Therefore, the system program needs a human-computer interaction part to provide input for user customization, such as a parameter setting page, such as a long press of the pause "control key" that pops up an input interface, allowing users to input their customized parameters, thus better matching the user's pause precision.
[0058] In video playback, if the frame sequence value is used for positioning, the calculation method requires: 1. Obtaining the current frame sequence value when the pause is triggered; 2. Obtaining the frame rate and playback rate; 3. Calculating the backoff value based on the delay value, playback rate, and frame rate, i.e., how many frames to backoff, and then calculating the frame sequence value at the pause position and pausing with this value. For example, taking 30 frames as an example, the frame sequence value is fx, and the set delay value is still 500ms. Then 500ms * frame rate * 30 frames / 1000ms equals 9 frames. Therefore, when using the frame sequence, the current frame sequence value needs to be added or subtracted by 9 frames before determining the pause position.
[0059] After the aforementioned pause position is corrected, the image on the main display screen should also show the corrected image, such as in... Figure 4 In S203, the image after correction is displayed, which means the pause position is calculated and paused at the calculated position value. The above-mentioned pause method solves the problem of inaccurate pause for users. In fact, some users often trigger the pause in advance or place their finger on the corresponding button and wait for the video image to appear in order to pause accurately, and then make multiple fine adjustments. This method solves the problem of wanting to pause accurately but not being able to pause accurately. At the same time, it also improves the accuracy of pause effect matching control in priority application, thereby reducing the steps of fine adjustment and improving work efficiency.
[0060] The above-mentioned pausing method is closely related to the above-mentioned marking method. In fact, accurate marking requires the above-mentioned pausing method, as well as the method of prioritizing fine-tuning, so as to pause at a relatively accurate position and then mark precisely.
[0061] Coarse markings can be triggered during multimedia playback. Although they lack the effect of precise markings, they still provide marking and playback control. This is especially beneficial for repeated playback, learning, and evidence collection, improving work and learning efficiency and saving time. Fine markings, on the other hand, can save users' time, improve marking efficiency, and enhance content accuracy during playback control.
[0062] The precise video control method is as follows, which has been described in the prior application. This application will further explain the fine-tuning control in the prior application in conjunction with this application, so as to make media playback more efficient and precise.
[0063] Video fine-tuning involves two different methods: continuous fine-tuning and non-continuous fine-tuning. While continuous fine-tuning can quickly adjust the video to the target position, if there is a deviation of several frames, non-continuous fine-tuning is needed to complement it. This allows for accurate and efficient location of the target image or content, enabling various operations such as image extraction, cropping, editing, and marking at that video location. Alternatively, combining fine-tuning with marking at the target video location can eliminate the need for repeated continuous or non-continuous fine-tuning, allowing for clearer viewing or learning of video content, such as magic tricks, complex origami, and handicrafts. Of course, fine-tuning can also more accurately extract law enforcement evidence and key content images from the video.
[0064] The following is an explanation of continuous fine-tuning. Figure 2 In this context, G100 represents short-trigger and long-trigger pulses. T1 is a short-trigger, defined as a trigger duration greater than 0 and less than a set value. For example, in one embodiment of a touchscreen, the upper limit of the short-trigger setting can be set to 300ms, meaning that the duration of contact with the touchscreen after triggering must not exceed 300ms (the triggering of other types of sensors is similar, such as fabric fiber sensors or time-triggered sensors composed of buttons and clocks). T2 is a long-trigger, defined as a pulse width greater than the short-trigger setting. For example, if the upper limit of the short-trigger setting in this embodiment is 300ms, then the lower limit of the long-trigger setting must be greater than 300ms. The upper limit setting, for example, is set to 1200ms, meaning that the pulse width of the long-trigger is greater than 300ms and less than 1200ms (corresponding to the pulse diagram, the pulse width is greater than 300ms and less than 1200ms). In actual triggering, such as touching the touchscreen and immediately leaving, the trigger time is generally between 90ms and 300ms. However, if there is a slight hesitation after touching before leaving, the trigger time will be greater than 300ms but less than a certain value, such as 1200ms. Therefore, for ordinary people, it is easy to release the trigger upon touching, so short triggering is easy for anyone to master. Long triggering, on the other hand, involves a slight delay, and how long is considered long is a matter of perception. Therefore, in order to avoid false triggering, the setting range is relatively wide, such as 1200ms or 1500ms. Of course, the setting value is selected according to the specific implementation environment and object. For example, in a certain embodiment of multi-touch screen, it is set to 1600ms.
[0065] In G200, it is described that when a trigger lasts longer than T3, it begins counting according to a set time interval value, such as from V1 to Vn, assuming the end of the trigger is at Vn. During continuous fine-tuning of the video, the interval value depends on the user's reaction. For example, if the interval value is set to V=500ms, it means that every 500ms, it enters the next interval. For example, if T3 is set to 1600ms, after triggering and holding the trigger for more than 1600ms, it enters the V1 interval area. After holding for more than 2100ms, it enters the next interval area, as shown in the V2 interval area. When the holding time is more than 2600ms, it enters the time interval area such as V3. For example, if the trigger release time is 2800ms, it is exactly in the V3 time interval area, so the V3 interval area is selected. In the patent CN201910281922 (A Method for Controlling Volume in an Intelligent Electronic Device) invented and applied for by the inventor, the working logic of G200 is adopted. Combined with the short and long triggers to define the change of volume value, it solves the long-standing technical defect that single-touch devices such as true wireless earphones cannot adjust the volume with a single earphone. It also allows multiple functions to be implemented on single-touch devices, such as earphones. At the same time, it solves the legacy problem of adjusting volume through touch screen as a general function without the need for buttons on touch screen devices. The embodiments in this disclosure also adopt the technology of patent CN201910281922 (hereinafter referred to as patent 2), which can eliminate the dependence on control buttons or mobile phone volume buttons for video volume. At the same time, this patent is applied to adjusting video brightness.
[0066] In video fine-tuning, since the video is displayed on at least one two-dimensional plane (of course, it can be multi-dimensional), directional triggers can be used to control the input. Therefore, the G300 has two directional triggers with different trigger durations, for example, triggering to the right (similarly, in addition to the direction of the trigger, the duration from when the user touches the screen to when they leave the screen is recorded for any directional trigger).
[0067] In environments requiring both duration-based and directional triggering, any directional trigger (regardless of left, right, up, down, etc.) will generate duration feedback on the sensor (such as the duration pulse under the directional arrow in G300). If a user slowly moves to the right, and the touch duration exceeds the T3 setting, if the directional trigger hasn't ended, it will be mistakenly identified as a G200 trigger with a duration greater than T3. Therefore, a tolerance value is typically set for directional triggering. A trigger displacement greater than the touchscreen's set value of 10 points (the tolerance value, and within the corresponding trigger time) is considered a directional trigger. If it's less than this value, it's not recognized as a directional trigger. Thus, a value greater than the tolerance value can be identified as a directional trigger, not a trigger greater than T3 in G200. (The tolerance value needs to be tested and set based on whether the direction is mistakenly identified as a duration-based trigger during implementation. Alternatively, it can be set larger, such as 30 points, which requires a larger range of movement on the screen during the user's directional trigger.)
[0068] Therefore, in order to implement the fine-tuning method of this disclosure on sensor groups capable of sensing trigger direction and trigger duration, including but not limited to touchscreens, capacitive screens, and computer touchpads, the setting of the T3 duration, if the trigger code includes short and long duration triggers in G100, requires consideration of both the pulse width of T2 and the duration of the tolerance value for directional triggering. However, for simplicity, the T3 duration is usually set to a triggering value greater than the upper limit of the T2 duration setting. The upper limit of T2 is also set considering the tolerance value for directional triggering; for example, it is set to 1200ms in one embodiment, while in another embodiment, due to touchscreen issues, it is set to 1... The key factor in setting T3 is to prevent accidental triggering of any other trigger that exceeds the T3 limit before starting to work according to the time interval. Therefore, the T3 setting is based on trigger coding and aims to avoid accidental triggering. When it's only a combination of directional triggering and G200 triggering, the T3 duration should be greater than the directional trigger tolerance duration. For example, the tolerance value for directional triggering can be set to 10 points, and the tolerance time can be set to 800ms (this can be understood as a trigger touching the screen but displacing less than 10 points, which is considered a non-directional trigger; the tolerance value can be used to determine if a trigger exceeds T3). The trigger value is triggered, but because it is twisted on the touchscreen and there is displacement, but the displacement does not meet the judgment condition for directional triggering, the setting value of T3 should be determined based on the duration of the trigger code to avoid misidentification. Generally, when short, long and directional triggers are included, the setting value should be greater than the duration of long-duration triggers, that is, greater than the duration range defined by T2, such as greater than 1600ms. In this application, the setting duration is used as a suggestion that if the triggering system includes duration triggers, the setting duration should be greater than the upper limit duration of T2. If there is no duration trigger pre-set, there is no possibility of misoperation, and the setting duration can be set with the tolerance duration of directional triggers as the lower limit, that is, greater than the tolerance duration.
[0069] G401 and G501 are two ways of representing a video. G401 represents the video according to the structure of the video frame sequence, while G501 describes the video according to the length (such as milliseconds). This is because once the frame rate of a video file is determined, such as 60 frames / second or 30 frames / second, its frames and video time are also corresponding. Therefore, a specific frame image can be located from the time value, and a specific frame can also be located from the frame sequence. Key image information and target image information are in the frame. From the perspective of duration, it is only necessary to read the frame corresponding to the set duration (the video time usually refers to the time value or duration value of the video from the beginning to the current position. For example, tx = 500ms means that the duration of the video from the beginning to the current position is 500ms or the time at the current position is 500ms).
[0070] In G401 and G501, there are corresponding fx or tx (fx represents a certain frame in the video, and tx represents a certain time in the video), which are used to express the current position of the video. G402 can be a pointer or a parameter used as a marker in a program. When using the G402 frame sequence to locate the position of a certain frame in the video, the frame sequence value is used, while when using the duration as the location, the time value is used, usually a millisecond value.
[0071] The following explains the working mode of directional triggering and triggering with a duration greater than T2. In the diagram, G701 is a trigger combination of directional triggering and triggering with a duration greater than long or greater than T2. The tw (interval value between triggers) between the two triggers must be less than the instruction window value, for example, the instruction window value is set to 800ms. This ensures that the system recognizes that the subsequent trigger is a subsequent trigger of the preceding trigger, that is, there is a trigger within 800ms after the previous trigger. Otherwise, the previous trigger code is judged to have ended the input, the instruction window is closed, and the system enters the instruction recognition and execution stage instead of waiting for the subsequent trigger.
[0072] In one embodiment, when a directional trigger is detected and there is a subsequent trigger within the window period, and the subsequent trigger is a trigger with a duration longer than the long trigger duration (a trigger with a duration longer than T2), after the trigger duration reaches the time interval, the image of the paused position read from the video or the image of the positioned frame is displayed in the image display area on the interface, according to the video time interval or frame interval.
[0073] To further explain, if the pause frame is an fx frame or the pause duration is tx, and the user wants to continuously fine-tune towards the end of the video to find and locate the target image, then a right-trigger is used (for ease of memorization, a right-trigger is defined as towards the end of the video). Then, in the trigger command window, hold triggers and triggers that are longer than T2 duration, longer duration, or a set duration are used. When the system or app determines that a subsequent trigger exceeds the set duration and continues to hold the trigger, it is determined that continuous adjustment is needed. Therefore, based on the set video frame interval or video duration interval, the system reads frames or extracts images at the duration position according to the frame interval or video duration interval, and displays them in G601. For example, assuming the video file is 30 frames per second, if the trigger for right-triggered and triggers exceeding the long duration are defined as displaying each frame in the G601 towards the end of the video, then after the system detects a right-triggered event, another trigger occurs in the command window, and the trigger duration is greater than the long duration (set duration). Since the trigger is not deactivated, the command is determined to display each frame towards the end, and the displayed image changes every time interval. That is, after exceeding the set duration, fx+1 frames are displayed, and in the second time interval, fx+2 frames are displayed, and so on. When the trigger is maintained until the nth time interval, the displayed frame is fx+n. If the trigger command defines two right-triggered events and triggers exceeding the set duration within the command window, and if the video frame interval is set to 5 frames, then the first time interval displays fx+5*1, the second time interval displays fx+5*2, and so on, until the nth time interval displays fx+5*n (displayed in the G601 display area).If you make continuous fine adjustments towards the beginning of the video, you can use a leftward trigger, triggering a duration longer than the set time within the window. In this case, if the video frame interval is A, then in the first time interval, fx - A * 1; if it is the second time interval, then fx - A * 2, and so on until fx - A * n deactivates the trigger. For example, if the frame interval is 3, the image displayed in G601 is the frame number of the paused position, such as fx-3*1, and so on, until the target frame. The user looks at the current image in the display area, and if it is in the correct position, the trigger is released. This fine-tuning method is more accurate than the mouse GUI method because releasing the trigger and pausing the image at the target position is more precise than the mouse or computer method, where even if the position is reached, the user still needs to click the mouse, so the accuracy is usually lower than this method. Of course, in the program, fx is usually assigned the value of the frame currently displayed in G601. So at each interval, such as fx=fx+A towards the end of the video or fx=fx-A towards the beginning of the file, the pointer is automatically changed according to the frame interval value every time interval, thus completing the pointer movement or fx assignment. Therefore, there is a dashed pointer in G402, which is used to indicate that the pointer is moving every time interval until it reaches the target position. The dashed pointer indicates the position that has been moved.
[0074] For a video using video duration, assuming a 30-frame video, each frame is 1000ms / 30 = 33.333ms. Following the same method described above, assuming B is the video time interval, if it's frame-by-frame, B = 33.333ms; if the video time interval is 3 frames, it's 100ms. Therefore, the calculation for each time interval is tx = tx + B (end direction) or tx = tx - B (start direction). For example, if the current time interval (tx) is 1000ms and the video time interval is 100ms (3 frames), assuming the direction is towards the end of the video, the time to extract the video image in the first time interval triggered by a duration greater than the set duration is 1000ms + 100ms. The second time interval is the current pointer time (tx) of 1100ms plus 100ms, which is 1200ms. In other words, as the pointer moves, the video of the next video time interval is the current time (tx) plus the video time interval value. So, simply put, the video duration of the image displayed in the previous time interval is the video duration corresponding to the current time interval, which is the video duration based on the direction of continuous fine-tuning plus or minus the video duration interval. The image in the video is then read out based on this time and displayed on G601.
[0075] It is important to note that triggers combined with triggers of duration greater than T3 or a set duration usually contain pre-encoding. That is, the meaning of the instruction is determined by the pre-encoding combined with the trigger of duration greater than the set duration. The pre-encoding not only expresses the direction, but also determines the video time interval value or frame interval value (the mapping relationship of the pre-encoding). The instruction recognition system displays the corresponding image in the display area according to the continuous display of the frame interval or video duration interval defined by the trigger instruction, based on the video status value, trigger encoding, and the current frame number or time.
[0076] Therefore, by using the above method, we are actually simulating the knob on a professional video editing console. This allows for the continuous display of images near the target location. By deactivating the trigger, the key target image and content can be selected (essentially a knob with a constant speed, whereas a knob on a workbench can be adjusted according to the user's needs).
[0077] Generally, directional triggers are more representative of video direction. Therefore, in the real-time example, a right trigger plus a trigger with a duration greater than T3 is used, with each frame of the video displayed on G601 at intervals. A combination of right triggers and right triggers plus a trigger with a duration greater than T3 displays the image of the fifth frame from the current frame towards the end in G601 at intervals (the frame interval is assumed to be 5, or the video duration interval is the corresponding 33.333ms*5). Therefore, by using frame intervals or video duration intervals, and based on the pre-triggered encoding and triggers with a duration greater than T3, the target position can be reached more quickly and accurately. This is equivalent to creatively using trigger encoding greater than the set duration to simulate the knob for quickly locating images on a video editing console.
[0078] As for the pre-triggering of triggers with a duration greater than T3, directional triggering can be used, or short-duration, long-duration triggering, or a combination of directional triggering can be used. This invention is not limited to any particular method. However, since there are many video operation instructions and even more micro-operations, it is best to use user-derivable encoding logic to reduce the cost and difficulty for users to remember the trigger codes. In addition, it should be considered that if a general volume adjustment function is introduced, the trigger codes should not conflict.
[0079] In the aforementioned continuous fine-tuning, if the frame interval or video duration interval is not frame-by-frame, especially in videos such as magic tricks, martial arts action, motor vehicle accidents, and crime scenes, the target image is often not displayed in the previous interval, but has already occurred in the next interval. Therefore, non-continuous fine-tuning should be adopted. Specifically, an encoded triggering method is used, such as triggering to the left or right in the video pause state, representing one frame or the corresponding video time interval. For example, triggering to the left means moving one frame to the left, i.e., fx-1 frame, or tx-33.33ms. For example, in continuous fine-tuning, if 4 frames are missed in the direction of ending, the user triggers to the left once. After execution, similar triggering 3 times will find the key target image. In magic tricks, the useful and recognizable images are often just 1 or 2 frames. In the embodiment, for non-continuous fine-tuning, such as trigger codes ->, ->->, or ->->-> (note that -> represents triggering in the right direction, and <- represents triggering in the left direction), which respectively represent triggering one frame, three frames, or ten frames to the right, or <-, <-<-, or <-<-<-, which respectively represent triggering one frame, three frames, or ten frames to the left, the user can quickly find the target image by combining continuous fine-tuning with their position in the current frame and the possible target frame image (by using the trigger code and the corresponding video frame interval or time interval value, the image is non-continuously displayed in the display area according to the direction and interval value defined by the trigger code, and the image is positioned after adjustment according to the direction and interval value defined by the trigger code).
[0080] The reason for introducing both continuous and non-continuous fine-tuning in the design is that finding the target image in a video is not easy. This is why professional video editing consoles have been developed and remain irreplaceable for decades. However, implementing this on touchscreens and other sensor components is not as convenient and fast as using rotary knobs and large screens. Therefore, after adopting continuous fine-tuning, if you still use continuous fine-tuning after missing the target image, it is usually not as good as non-continuous fine-tuning. For example, if you miss 1 or 2 frames, non-continuous fine-tuning is actually more convenient in finding the target image. If you only use non-continuous fine-tuning, in a long video, the user has to trigger the screen too many times, which is not as time-saving and improves the user experience as continuous fine-tuning. Furthermore, watching videos with a certain degree of continuity makes it easier to predict the target position. When the video is not continuous, the continuity decreases due to the longer image display time, making it difficult to predict. In terms of the image effect of continuous fine-tuning, it is similar to a continuous video played at an extremely slow speed, allowing you to rewind or advance the video. Currently, this experience can only be achieved by manually rotating the knob on the video editing console. In addition, it is worth noting that in order to achieve better accuracy and efficiency, a differential method can be used. For example, continuous fine-tuning can use a slightly coarser interval, such as 3 frames or 100ms, while non-continuous fine-tuning can use 1 frame or 33.33ms. The purpose is that when users use continuous fine-tuning for efficiency and to quickly reach the target position, they may miss key target frames or images with each time interval. In this case, by deactivating the trigger that exceeds the set time and using non-continuous fine-tuning, the target can be reached more quickly. It is not necessary to set the same video interval in continuous fine-tuning as in non-continuous fine-tuning. Therefore, the combination of the two types of fine-tuning, with differential intervals, can achieve both efficiency and accuracy.
[0081] The fine-tuning features include the aforementioned continuous video fine-tuning method and non-continuous video fine-tuning; the non-continuous fine-tuning is performed in a paused video state, using triggering or triggering combinations, according to the video frame interval or video duration interval defined by the triggering or triggering encoding, and the triggering direction, to display an image in the display area that is repositioned and read from the video according to the direction and interval defined by the triggering, and the current image position is repositioned in the video.
[0082] like Figure 6 This describes the above-mentioned fine-tuning method. In step S102, the above-mentioned pause method is used. If the above-mentioned pause method is used in the fine-tuning method, the pause position will be more precise. This will reduce the number of operations required in steps S104 or S105, which means reducing the number of fine-tuning steps. Therefore, the efficiency of users is improved, and the user experience is also improved.
[0083] In conjunction with video fine-tuning, in some implementation examples, the video fine-tuning method is used in the video editing function to fine-tune both sides of the video selection box, and the start and end positions of the selection box in the video are determined by the content of the image in the display area.
[0084] In conjunction with video fine-tuning, in some implementation examples, the fine-tuning method is applied to fine-tune the video playback or video editing functions to a target image, including using the image as the preview image that best reflects the video content. Specifically, for example, when a smart terminal records video, it is generally stored on the smart terminal, so each video has its own preview screen. However, the preview screen usually does not reflect the video content, and over time, it is not easy to find previously recorded content among multiple video files. This requires the user to find the target image in the recorded video that reflects the video content or features and set it as the preview image. Therefore, the fine-tuning method of this disclosure is needed, and the video control method... This method allows end users to quickly locate target content after recording a video and set the located target image as a preview image of the recorded video for later use and retrieval. Therefore, both fine-tuning and precise pausing improve the efficiency of finding frames containing key content, extracting images, and setting them as preview images. In magic trick videos, for example, by finding the magic trick decryption frame and marking and pausing parameters, users or other users can directly click on the mark to decrypt or learn the magic trick. Therefore, fine-tuning is an essential function in both marking and image editing, but traditional players do not include it, which leads to many user demands not being met. This disclosure provides corresponding technologies to improve the effect.
[0085] Video fine-tuning includes any type of fine-tuning, and the input area triggered can be in the main image display area. This allows for tasks such as finding target positions and editing without an editing interface in full-screen mode, because traditional editing requires multiple windows, especially on smartphones, where the screen cannot work in the large-screen manner of a video editing console.
[0086] For video editing functions, such as Figure 7 As shown, f1 is a GUI interface for a smart terminal or mobile phone, f2 is the display area for images or videos, f3 is an information bar containing thumbnails of videos used for video editing / trimming functions, which includes multiple images extracted in video order, and f4 and f5 are video cropping or selection boxes, with f4 being the left edge and f5 the right edge. As is well known, current touch technology only allows users to select and frame the video to be cropped or selected by moving the left or right edge of the multi-touch screen with their fingers. Typically, video thumbnails are displayed in the f3 bar, making it very difficult to select the desired video information. Furthermore, if the video is long, the thumbnail interval becomes long. In a video workbench, a knob is needed for precise cropping. Therefore, to achieve the capabilities of a professional video workbench on smart terminals and other devices, a fine-tuning method is used, thus achieving the effect of a professional video workbench.
[0087] Here's a specific real-time example: When you tap or press F4, the left edge of the selection box moves to the target position (the actual selection box is like a video progress bar, the only difference being that a progress bar can only select one point, while a selection box can select two time points (or the positions of two frame sequences) in the video, so the video between the two selected points can be cropped or extracted). At this point, there is still a distance to the target point in the video. Therefore, if the precision of your finger cannot be adjusted any closer, you need to use continuous fine-tuning and non-continuous fine-tuning methods. For example, first move closer to the target position relatively quickly (such as continuous fine-tuning), then use non-continuous fine-tuning, and finally move F4 frame by frame to the target frame, target content, or corresponding video time point.
[0088] However, given that the thumbnail of f3 is too small (its layout logic is based on traditional workstation video editing software or systems, which cannot be laid out in the small display space of smart terminals, so the thumbnail is too small to be used for judging the accuracy of the editing), in this disclosure, the positioned display image can be displayed in S203 and S202, where S203 is the video display area and S202 is the thumbnail display area. Even if a particularly small continuous thumbnail like f3 is used, the positioned image still needs to be displayed in a larger display area.
[0089] The display is best when done in S203, especially when fine-tuning the left and right edges of the selection box (f5 adjusts the right edge in the same way as the left edge).
[0090] Therefore, as described in the above embodiments, this disclosure achieves the same level of precision in video editing as professional video editing systems on a smart terminal touchscreen, with the precision required to select a video range. Traditional multi-touch technology, on the other hand, cannot efficiently fulfill the capability of fine-tuning and selecting images, including in terms of interface, controls, complexity of implementing touch commands, and efficiency.
[0091] In addition, it should be noted that in this disclosure, the touch area can be located anywhere on the screen, such as in the S203 video display area or in a designated area. However, if the video is horizontally full-screen, the advantage of touch in the S203 display area is very strong. Thus, the desired result can be achieved without any controls or screen buttons that occupy or cover the screen. This is an advantage that traditional video playback and editing software do not have, as they require a lot of space to display controls.
[0092] However, it should be noted that when adding a precise pause function, since the fine-tuning method uses a combination of duration, direction, and trigger coding, for the sake of accuracy and time saving, the precise pause needs to be displayed as a "control key" so that users can precisely trigger the pause, thereby reducing the steps of fine-tuning. Therefore, in engineering, new innovative technologies are used to better improve the effect. So in actual engineering, an additional pause key can be added and displayed. This pause can be called a precise pause key to distinguish it from the traditional pause. As mentioned above, it can be placed in the position of the progress control bar, such as the right side of the progress bar.
[0093] Smart electronic devices such as smartphones, tablets, computers, and laptops typically have touchscreens or touchpads. These sensors or sensor groups usually consist of sensor components such as capacitive touchscreens and clock systems. The clock system times the trigger duration and converts the trigger into a time value. Triggers such as direction and gestures are recognized by specialized programs. For system programs, if it's for video fine-tuning, it might be the system program of the smart electronic device. For computers and smart terminals that can install apps, an app can be used to monitor the trigger information of the aforementioned sensors or sensor groups. Of course, in video fine-tuning, continuous fine-tuning requires sustained triggering. However, as mentioned above, in an environment where duration triggering and direction triggering are used together, to ensure accurate trigger recognition, the time values of long-duration triggering and direction triggering must be determined. Therefore, to solve this problem, a tolerance value for direction triggering is introduced. So, in the case of mixed use of duration triggering and direction triggering, the upper limit of long-duration triggering is usually set to the "set duration." That is, a trigger exceeding the upper limit of long-duration triggering, such as 1600ms, is recognized as a trigger exceeding the set duration. After setting the duration, the system enters the set time interval judgment process, such as a time interval of 500ms (adjusted according to human reaction speed, such as 800ms). This means that after each interval, when entering the next interval, the image needs to be displayed as a different image. Therefore, the display direction and frame or time interval need to be predefined in the trigger code. For example, in the case of right trigger + duration longer than the set duration trigger mentioned above, the right trigger is the pre-trigger code. The pre-trigger code not only determines the trigger video display direction (towards the beginning or end of the video), but also needs to set the adjusted video frame interval or video duration interval. In this way, in each time interval, according to the trigger code settings, in the next time interval, at the position where the video was displayed in the previous time interval, according to the defined adjustment direction and adjustment interval (as described above), the video image is read and displayed.
[0094] Of course, when there is only directional trigger pre-coding, the set duration can be greater than the minimum tolerable duration value of directional trigger. Therefore, the so-called set duration should be determined based on the characteristics of the trigger encoding and the required duration, with the aim of avoiding misidentification of the trigger.
[0095] When users want to learn, see the details of a video, or extract information on the road like traffic police to gain the approval of those involved in an accident, they need to find the target frames or images that contain key information in the video content. Therefore, they need to use a fine-tuning method. In order to improve the efficiency of fine-tuning and make precise pausing methods, the above-mentioned marking method can be used to quickly and repeatedly watch the video, even one frame per second (such as playing at 1 / 30 speed), or to crop and extract image and video information.
[0096] Even with a playback method, when a video is played (whether at normal speed or abnormal speeds such as 2x or 0.5x), it takes hundreds of milliseconds for the user to see key information, for the brain to process it, and for the body to react by touching the screen to pause. For images, even at 0.25x speed, several frames have passed. Considering the time it takes for the human brain to recognize image content and for feedback, it is obvious that key target positions may be missed. Since progress bar control (S201) cannot handle this situation, fine-tuning must be used to solve the problem, and precision pausing further improves the efficiency of fine-tuning.
[0097] In one embodiment, downward triggering (directional triggering) is used for marking. For example, when the video is playing or paused, downward triggering is used in the touch area to perform the above-mentioned marking steps and functions. As mentioned above, pausing can be precise pausing, or precise pausing can be included in downward triggering. This way, when the video is playing, not only does marking begin, but it also rewinds to the position where it should be paused, rather than the position that is delayed today. The time of directional triggering is also added to the deviation to calculate and locate a relatively accurate position and pause. Precision pausing is just a pausing step in marking triggering.
[0098] Having explained the methods and necessity of video fine-tuning, let's now briefly explain video control methods, such as... Figure 6 As shown, in this disclosure, playback means, but is not limited to, playback at normal speed or abnormal speed.
[0099] Specifically, the technical methods used in this disclosure rely on sensors or sensor groups that can sense trigger direction and trigger duration. These are not limited to touch screens of smart terminals, computer touch screens, touchpads, or even sensor groups composed of fabric fibers and circuits that can identify and receive directional and duration triggers. Such sensors or sensor groups can be built into smart electronic devices, such as touch screens of smartphones and smart terminals, touchpads of ordinary laptops used for mouse and function input, or external touch electronic devices that can sense duration and trigger direction and are connected to smart electronic devices via wired or wireless connections.
[0100] The video control method disclosed herein is as follows: the user of the video determines the position of the target image or content by controlling the video's progress bar and based on the content displayed on the video's interface. Figure 6 (Step S101 in the middle); Generally speaking, since the video progress bar is based on the video length, such as 120 minutes, and the video length is proportional to the progress bar length, the lines on the video progress bar that indicate the progress (such as different colored lines indicating the progress that has been played and the progress that has not been played) or the indicators that control the movement, such as... Figure 4 The black circular indicator on the S201. Therefore, when selecting a video using the progress bar, the progress bar will move forward or backward one position on the screen depending on the length of the video, causing the video to deviate by a few seconds or minutes (depending on the video length). As a result, video player users can usually only control the video to a certain approximate position, but this is still several seconds or frames away from the precise target position.
[0101] The video user controls the video using a progress bar and determines the viewpoint based on the content displayed on the video interface. Figure 4 In the middle, the display interface contains S202 or S203, where S202 is a thumbnail, that is, a video thumbnail corresponding to the current progress bar position. By controlling the indicator of moving the progress bar, the information in the thumbnail can be used to determine whether the general target position has been reached.
[0102] S202 can be Figure 4 Any location on the left side of the smart terminal interface S204, but not superimposed on or obstructing the main video display interface (image display area) of S203; it can also be... Figure 4 There are other ways to overlay or obstruct the S203 video main display interface on the right side of the smart terminal interface, such as not using the S202 thumbnail, but directly displaying the S203, and determining the target position by dragging the progress bar according to the content displayed by the S203.
[0103] As mentioned above, the progress bar corresponds proportionally to the video length. Therefore, the longer the video, the longer the progress bar moves (one point). Since one point represents a longer video length, whether you use a mouse to click or multi-touch to control the position of the progress bar marker, you can only reach a certain range of the target position. To solve this problem, professional video processing software and its associated control consoles use knobs to fine-tune to the required precise target position, or display each frame in sequence on a large screen computer so that the operator can search frame by frame. However, smart terminals obviously do not have such screen space.
[0104] Therefore, for ordinary video playback users, traffic police on duty, and editing users, it is obviously impossible to carry external devices (knob video control devices) to control the relatively precise positioning of the video.
[0105] In addition, users can also adjust the displayed content during video playback, that is... Figure 4 In the case where only the main display interface S203 is displayed without S202, the user playing the video can pause the video according to the content to proceed to the next step. That is, the video playback is paused at the target position according to the content displayed on the video playback display interface. Figure 6 (Step S102), which, as described above, may include the precise pause playback method.
[0106] Therefore, video users can control the video progress bar to reach the target position of the video content and then play the video to reach the target position, or they can control the video progress bar to reach the target position of the video content and then play the video to pause when they are closer to the target position (so precise pausing is required to reduce the number of fine-tuning steps).
[0107] One important point to note is that the video playback described in this application includes normal speed playback, which can typically be set to a speed value of 1, meaning 30 frames per second for a 30-frame-per-second image, and 60 frames per second for a 60-frame-per-second image. It also includes non-normal speed playback, such as setting a speed value of 2 or 1.5, which is 2x or 1.5x speed playback, commonly referred to as fast playback. Setting the playback speed to 0.25, 0.5, 0.75, etc., represents video playback slower than normal speed playback, typically slow playback. 0.25 is equivalent to a playback speed of 1 / 4 of normal speed playback. Therefore, for precise pauses under non-normal speed playback, as mentioned earlier, the playback speed needs to be included in the calculation.
[0108] Imagine a user watching a long video and using a progress bar to get to a position near the target location. However, the user is still some distance away from the target location. Due to the touchscreen and multi-touch features, no matter how the user moves the control indicators on the progress bar, they cannot quickly and accurately reach the target location. Therefore, the user can use video playback to get closer to the target location. However, it should be noted that playback can be forward at the user's selected speed or rewind at the selected speed. Therefore, in the current video player mode, control space is needed to display these control buttons, which will further reduce or cover the main display space.
[0109] Therefore, in this disclosure, the technology of patent CN201811253657 (a method for operating and controlling a smart terminal or smart electronic device, hereinafter referred to as patent 1) is applied. The method of forming instructions is based on the trigger code and the function and state of the controlled target (APP or electronic device). Combined with the logic of the code, the user does not need to use the control method, but can use a simple touch method to control the video.
[0110] For example, in the settings such as playback mode, short, short, right trigger means 2x speed to the right, short, right trigger means 1.5x speed to the right, long trigger, right trigger means 0.5x speed forward, and long trigger, long trigger, right trigger means 0.25x speed forward (video end direction).
[0111] In human consciousness, short triggers are also fast triggers, while long triggers result in slightly slower, more hesitant behavior. Therefore, a short, short, right trigger can be perceived as "quickly moving forward," for example, executing a 2x speed command. Conversely, a long, long, right trigger command is perceived as "slowly moving forward," meaning executing a 0.25x speed playback command. Similarly, a short trigger combined with a right trigger is "quickly moving forward," and so on. "Quickly moving backward" corresponds to a short, short, left trigger code. This avoids the need for several conventional controls, allowing users to easily memorize or deduce multiple playback control trigger codes after a single repetition.
[0112] Therefore, for users who need to repeatedly view details of a certain location, touch commands are more efficient and convenient than controls. Marking methods are more convenient than touch commands. Of course, touch methods do not take up video content space (which is squeezed or covered by the control interface). Moreover, the logic of triggering encoding and corresponding commands can be deduced after one or two uses. Therefore, for operations like video that require repeated searching for key information, the coded touch method is faster and more direct. At least users will not have to repeatedly bring up the playback control interface.
[0113] Regardless of the playback method, the reason for adopting the human-computer interaction technology of Patent 1 is to overcome many existing technical defects or obstacles in video playback control, such as the control covering the video display area.
[0114] When users want to learn, see the details of a video, or extract information on the road like traffic police to gain the approval of those involved in an accident, they need to find the target frames or images that contain key information in the video content. Therefore, they need to use fine-tuning methods, marking methods, and watch the video repeatedly, or even frame by frame, especially in the process of extracting evidence or studying criminal behavior.
[0115] Even with a playback method, when a video plays (whether fast or slow), it takes hundreds of milliseconds for the user to see key information, for the brain to process it, and for the body to react by touching the screen to pause. For images, even at 0.25x speed, several frames have passed. Adding the time it takes for the human brain to recognize image content and for feedback, it is obvious that key target locations may be missed. Since progress bar control cannot handle this situation, fine-tuning methods must be used to solve the problem, while precise pausing reduces the steps and time required for fine-tuning.
[0116] In smartphone touch technology, especially multi-touch technology, fine-tuning is not possible. Therefore, to utilize the touch screen and enable fine-tuning, it is necessary to simulate the working logic of a rotary knob in a video editing console. However, users cannot install knobs on smart terminals. Therefore, in this application, two methods are used for fine-tuning: continuous fine-tuning and non-continuous fine-tuning. By using two types of fine-tuning or a combination of the two, users can accurately locate the target image in the video, just like using a video editing console. The combination of the two types of adjustment improves both efficiency and accuracy.
[0117] Therefore for Figure 6 In the S103 step (video control) method, it is usually necessary to control the video progress bar, determine the target position according to the image content in the video display interface, or pause playback after reaching the target position according to the content in the video playback display interface. Then, the video is controlled by the continuous fine adjustment of video S104 method or the non-continuous fine adjustment of video S105 method.
[0118] For example, video cropping Figure 7 In the video progress bar, F4 and F5 are actually special types of video progress bars. They allow you to select the range of the video (the progress in the video, but not in two directions) from two directions. As explained in the fine-tuning section above, by using fine-tuning, you can adjust the left or right edge to determine the position of the left and right edges in the video based on the image in the video display area, thereby enabling more precise cropping or video content extraction.
[0119] The methods described above can be applied to video target localization, marking, and precise cropping, selection, or accurate locating of a target image, marking an image, or marking a segment, thereby enabling learning and serving as evidence, which becomes possible on devices such as smart terminals.
[0120] Through the description of the above methods and embodiments, we know that the marking, precise pausing, and fine-tuning of a multimedia file are interconnected. Through innovation in multiple aspects, efficiency has undoubtedly been systematically improved from the perspective of multimedia control, and the user experience has been enhanced. At the same time, the efficiency of functions such as media players, video editing, evidence extraction, and learning has also been improved.
[0121] Those skilled in the art can implement the above-described technologies, thereby significantly improving the capabilities and efficiency of multimedia control technology.
[0122] The above-mentioned methods and playback control progress bars, one or more in combination, can be applied to systems that include multimedia playback and editing. These systems are not limited to computer programs, smart terminal apps, professional video editing systems, or multimedia functions of electronic devices or electromechanical equipment, such as vehicles like cars and ships, treadmills, and smart XR helmets, as functional applications of the multimedia component.
Claims
1. A method for multimedia control, characterized in that... include: After the playback system detects a pause trigger, it reads the current position data of the multimedia file being played. The backoff value is calculated based on a preset or user-defined delay value and the playback speed of the multimedia. The pause position is calculated based on the rewind value, playback direction, and position data; Playback will be paused at the specified pause position.
2. The multimedia control method according to claim 1, characterized in that, include: The pause trigger can be any one of the following: physical button trigger, control key trigger in the graphical user interface, menu trigger, or gesture trigger.
3. The multimedia control method according to claim 1, characterized in that, include: The delay value includes the time for the human eye or ear to recognize the multimedia, the brain reaction time, the brain-controlled hand operation time, and the time for the playback system to recognize the pause trigger.
4. The multimedia control method according to claim 1, characterized in that, include: The location data includes at least one of time values or frame sequence values; When playing an audio file, use the time value.
5. The multimedia control method according to claim 1, characterized in that, include: If the file being played is a video file, the main display screen shows an image of the current paused position of the multimedia file.
6. The multimedia control method according to claim 1, characterized in that, It also includes a tag generation step: Users generate tags by triggering instructions or functions related to multimedia tags; The generation of the markers is accomplished by the following steps: Get the current position value of the multimedia file; Enter or select playback control parameters in the human-computer interaction interface; Generate markers in the human-computer interaction interface; If the generated marker is subsequently triggered or clicked by the user, the playback system will play the game according to the location value and playback control parameters.
7. The multimedia control method according to claim 6, characterized in that, include: During the marking process, a fine-tuning method was also applied, which included both continuous and discontinuous fine-tuning.
8. The multimedia control method according to claim 6, characterized in that, include: By clicking on the two markers in the set instruction window, a human-computer interaction window is displayed, allowing the user to select or input playback control parameters, thereby generating a marker for the multimedia segment between the two markers.
9. The multimedia control method according to claim 6, characterized in that, include: The marker includes an image marker, and the image is extracted from the position at which the marker is located.
10. The multimedia control method according to claim 6, characterized in that, include: The playback progress bar includes generated markers. When the user clicks a marker, the playback will proceed according to the marker's playback parameters.
11. A multimedia playback system, characterized in that... include: The multimedia control method according to any one of claims 1-10 is used for multimedia playback control.
Citation Information
Patent Citations
A method for controlling smart terminals or smart electronic devices
CN109462690B
A method for controlling volume in an intelligent electronic device, the electronic device, and headphones.
CN110087160B
Frame playing method and client end for game application
CN104461718A
A streaming media data playing method and device
CN109819310A