Video fine adjustment, control and evidence obtaining method

By employing a combination of duration-triggered and direction-triggered encoding on smart terminals, combined with video frame intervals or duration intervals, continuous and non-continuous micro-adjustments are achieved, solving the problem of insufficient precision in video micro-adjustment on multi-touch devices and improving video editing efficiency and accuracy.

CN121815019APending Publication Date: 2026-04-07CROSSOVER FREEDOM TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-31
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve rapid and precise fine-tuning and positioning of videos on multi-touch devices, especially on smart terminals. This results in low video editing efficiency, an inability to accurately extract key content, and hinders the dissemination and application of videos.

Method used

By employing a combination of duration-triggered and direction-triggered encoding, combined with video frame intervals or duration intervals, and using sensors on smart devices to detect and identify triggers, continuous and non-continuous micro-adjustments are achieved, simulating the knob function of a professional video editing console and improving positioning accuracy.

Benefits of technology

It enables rapid and precise fine-tuning of video frames on smart terminals, improving the efficiency and accuracy of video editing and solving the problem of insufficient video control precision in existing technologies, especially on devices such as smartphones, tablets, and computers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815019A_ABST
    Figure CN121815019A_ABST
Patent Text Reader

Abstract

The invention provides a video fine adjustment method, a video control method and an on-site quick disposal method, so that on one hand, the well-known problem that a target frame and a position in a video are difficult to quickly and accurately position by adopting an intelligent terminal of a traditional multi-point touch control technology is solved; therefore, no matter how the evidence video is expected to be quickly processed and extracted in video clipping depending on video learning and precision and on-site law enforcement of law enforcement officers, the technical restriction problem exists, and the man-machine interaction technology, the video fine adjustment method and the video control method provided in the disclosure are just used for solving the restriction problem. In addition, the invention also provides an on-site rapid disposal method, and solves the difficult problem of rapid evidence obtaining and video of current on-site event disposal such as rapid traffic accident disposal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical fields involved in this disclosure include human-computer interaction technology for quickly and accurately fine-tuning targets in an APP under multi-touch conditions, as well as video fine-tuning, precise positioning and video control, and precise extraction of key video content or images and videos; in addition, it also involves methods for rapid on-site processing and methods for quickly and accurately locating and extracting evidence from video files. Background Technology

[0002] As we all know, when making fine adjustments to videos using multi-touch technology, such as quickly and accurately locating the target image frame, it is a very difficult operation. In fact, it is also difficult to achieve on a GUI workbench with a mouse. Therefore, professional video editing workstations are equipped with dedicated editing keyboards with knobs to achieve the requirements of precision and efficiency.

[0003] However, given the widespread use of videos today, whether it's individual media editors, children repeatedly learning from videos such as origami or solving magic tricks, or law enforcement officers extracting evidence from videos during on-site enforcement, they all encounter difficulties due to the low efficiency of video playback control. This problem has been persistent and has troubled a large number of video users, regardless of video application scenarios such as editing, learning, or evidence collection.

[0004] Furthermore, traditional editing methods on smart terminals result in small main image screens and an inability to edit in landscape mode (the main image display area is even smaller, so it is generally only cropped in portrait mode). When encountering vertically shot videos, the main image is also very small, making it difficult for users to observe details and thus make precise editing. Existing editing methods apply the logic of computers, so in portrait mode, there are multiple editing function areas at the bottom of the screen, but the images inside are very small, and the number of frames that can be displayed is limited, making it difficult to clearly see the image and make accurate cropping.

[0005] When video users typically publish video files online, the best image that reflects the video content cannot be accurately extracted and used as a preview image for video retrieval. Therefore, in the current situation where video content is retrieved based on the preview image, the dissemination rate of the video is affected.

[0006] In video editing, precise displacement and rotation of targets are difficult to control and adjust precisely with existing multi-touch technology, including movement, rotation, and boundary adjustments. Similar problems exist in other apps, where controlled targets, especially small ones, cannot be moved, rotated, or have their boundaries finely adjusted. These apps include image editing, video editing, and design software, as well as those that require precise movement and rotation of targets in games.

[0007] All of the aforementioned apps and applications, including those running on smartphones, smart terminals, or computers with multi-touch screens or touchpads and capable of sensing duration and direction triggering, face the challenge of precisely controlling targets, especially tiny targets; and especially XR (VR, AR, MR) devices, which suffer from a lack of precision in control.

[0008] The aforementioned problems have been a widespread and long-standing problem for users of smart terminals who are unable to use peripherals such as mice and knobs.

[0009] The inventors of this disclosure previously invented and obtained authorized patents CN201811253657 (a method for operating and controlling a smart terminal or smart electronic device) or CN201910281922 (a method for controlling the volume of a smart electronic device). The methods and embodiments provided in this disclosure apply the above-mentioned human-computer interaction technology for all scenarios to achieve the goals and effects. Summary of the Invention

[0010] In a first aspect, a method for continuous fine-tuning of video is provided, characterized by comprising: Applications are made in intelligent electronic devices containing sensors or sensor arrays that can sense both duration and direction triggering. Triggered by detection and identification through system programs or apps running on the smart electronic device; If the trigger includes a duration longer than the set duration, then at each time interval, according to the adjustment direction and video frame interval or video duration interval preset by the identified pre-triggered code of the duration trigger, the image of the corresponding frame or video time after adjustment is read according to the adjustment direction and the video frame interval or video duration interval in the video of the currently displayed frame or the image of the previous time interval, and displayed in the display area. The user can then select to deactivate the duration trigger based on the image content.

[0011] In conjunction with the first aspect, in some implementation examples, the pre-encoding of the trigger includes duration triggering, direction triggering, or a combination of duration triggering and direction triggering; the duration triggering encoding also includes short duration triggering and long duration triggering; the pre-encoding and the duration triggering containing duration triggering greater than the set duration constitute the triggering encoding for continuous fine-tuning of the video.

[0012] In conjunction with the first aspect, in some implementation examples, the position of the video at the previous time interval is the frame position of the displayed image in the video or the time position in the video at the previous time interval.

[0013] In conjunction with the first aspect, in some implementation examples, the video frame interval or video duration interval is a granular value in the continuous fine-tuning of the video, and the granular value is used to represent the precision of the fine-tuning; the pre-trigger determines the granular value and direction of the continuous fine-tuning.

[0014] In conjunction with the first aspect, in some implementation examples: the display area includes a thumbnail display area or a main display area for the video.

[0015] Secondly, a method for video micro-adjustment is provided, characterized by comprising: In conjunction with the first aspect and various embodiments combined with the first aspect, the video fine-tuning method includes a method for continuous video fine-tuning and a method for non-continuous video fine-tuning; the non-continuous fine-tuning is performed in a video paused state, using a trigger or trigger code combination, according to the video frame interval or video duration interval and trigger direction defined by the trigger or trigger code, to display in the display area an image that is repositioned and read in the video according to the direction and interval defined by the trigger, and the current image position is repositioned in the video.

[0016] In conjunction with the second aspect, in some implementation examples, in the video editing function, the video fine-tuning method is used to fine-tune both sides of the video selection box, and the start and end positions of the selection box in the video are determined by the content of the image in the display area.

[0017] In conjunction with the second aspect, in some implementation examples, the aforementioned fine-tuning method is applied to fine-tune the target image in video playback or video editing functions.

[0018] In some implementation examples combining the first and second aspects and their combinations, the triggered input area can be in the main image display area.

[0019] A third aspect provides a video control method, characterized in that it includes: By controlling the video progress bar, the target position is determined based on the image content in the video display interface, or the playback is paused after reaching the target position based on the content in the video playback display interface. Then, the video is controlled by the video continuous micro-adjustment method as described in any one of the embodiments of the first aspect and the embodiments combined with the first aspect, or the video micro-adjustment method as described in any one of the embodiments of the second aspect and its related aspects.

[0020] In conjunction with the third aspect, some implementation examples include: the playback of the video includes playback at a regular speed or at an unconventional speed; and the use of a differential interval method for continuous and non-continuous fine-tuning to make the combination of continuous and non-continuous fine-tuning more efficient.

[0021] The fourth aspect provides a method for rapid event processing and evidence collection, characterized by including: When installing a camera, or by surveying an already installed camera, record information about the camera, including its installation location, camera orientation, and the area covered by the camera's image. Enter data, or annotate it on the GIS, or generate a GIS layer for the camera; The user of the camera video can query the entered data based on the location of the handheld smart terminal, or display the image captured by the camera covering the location and its coverage area on the GIS layer based on the location data of the user's handheld smart terminal. The caller of the video reads images captured by the queried camera or image, covering the camera at the location, according to the time range of the event.

[0022] In conjunction with some embodiments of the fourth aspect, the feature includes: The user of the video, on the handheld smart terminal, uses some embodiments of the first, second, third, and fourth aspects to locate the target image or target video segment in the video.

[0023] In conjunction with some embodiments of the fourth aspect, it is characterized in that The video from the camera can be retrieved directly from the camera or from the backend management system or storage system.

[0024] The above-mentioned aspects are examples of implementation. After finding the target frame, the user uses a fine-tuning method (continuous, non-continuous, or differential fine-tuning) and sets the target frame as the preview image of the video, the first frame, or the first segment of the video.

[0025] Based on the above examples, the user uses a fine-tuning method (continuous, discontinuous, or differential fine-tuning) to find the target frame, and then extracts the target frame for subsequent use, including evidence images, preview images, the first frame of a video, the first segment of a video, or the target frame being edited into an image to serve as the first frame, preview image, or first segment of video. Attached Figure Description

[0026] Figure 1 A diagram illustrating continuous fine-tuning; Figure 2 A schematic diagram of the video playback interface; Figure 3 A diagram illustrating the video control method; Figure 4 A diagram illustrating methods for rapid on-site response; Figure 5 Site diagram; Figure 6 Video editing function interface diagram Figure 1 ; Figure 7Video editing function interface diagram Figure 2 Detailed implementation method:

[0027] The specific implementation methods, codes, quantities, values, durations, triggering methods, diagrams, etc., described in the following exemplary embodiments do not represent all implementation methods and embodiments consistent with those provided in this application. On the contrary, they are merely some methods, systems, apps, etc., consistent with this disclosure as described in the appended claims.

[0028] The technology and methods of this disclosure will be clearly and completely described below with reference to the accompanying drawings and embodiments. Obviously, the embodiments described below are only some embodiments of this disclosure, and not all embodiments. In addition, in addressing the goal of video fine-tuning and image editing target fine-tuning, on the one hand, it is based on the inventor's artistic and technical subject matter, reflecting the creativity of modern people in real life to use innovative technologies to complete tasks that could only be completed in static environments and under tool-dependent conditions in the past, such as video fine-tuning, editing and refining. On the other hand, based on the inventor's previous accumulation of innovative technologies (such as the two patented technologies mentioned above), in the long-standing difficult problem of video fine-tuning and target fine-tuning in APPs, continuous fine-tuning has been further creatively created and combined with non-continuous fine-tuning to propose differential fine-tuning, thereby solving a problem that has long existed in the industry and cannot be effectively solved by current popular human-computer interaction technologies such as touch. In addition, based on the difficulties of rapid disposal, law enforcement and video evidence collection, the technology of this disclosure solves the problems that plague on-site law enforcement.

[0029] The video micro-adjustment disclosed herein includes two different micro-adjustment methods: continuous micro-adjustment and non-continuous micro-adjustment. When continuous micro-adjustment can quickly adjust the video to the target position, but there may be deviations of several frames, non-continuous micro-adjustment is required to complement continuous micro-adjustment. Therefore, the target image or content can be found accurately and efficiently, allowing for various operations such as image extraction, cropping, and editing at that video position. Alternatively, continuous or non-continuous micro-adjustment can be repeatedly performed at the target video position to clearly see or learn the content in the video, such as magic tricks or origami, or to quickly and accurately find and extract evidence from the video at the scene.

[0030] The following is an explanation of continuous fine-tuning. Figure 1In this context, G100 represents short-trigger and long-trigger pulses. T1 is a short-trigger, defined as a trigger duration greater than 0 and less than a set value. For example, in one embodiment of a touchscreen, the upper limit of the short-trigger setting can be set to 300ms, meaning that the duration of touchscreen contact after triggering must not exceed 300ms. T2 is a long-trigger, defined as a pulse width greater than the short-trigger setting. If the upper limit of the short-trigger setting in this embodiment is 300ms, then the lower limit of the long-trigger setting must be greater than 300ms. The upper limit setting, for example, is set to 1200ms, meaning that the pulse width of the long-trigger is greater than 300ms and less than 1200ms (corresponding to a pulse diagram, the pulse width is greater than 300ms and less than 1200ms). In actual triggering, such as touching the touchscreen and immediately leaving, the trigger time is generally between 90ms and 300ms. However, if there is a slight hesitation after touching before leaving, the trigger time will be greater than 300ms but less than a certain value, such as 1200ms. Therefore, for ordinary people, it is easy to release the trigger upon touching, so short triggering is easy for anyone to master. Long triggering, on the other hand, involves a slight delay, and how long is considered long is a matter of perception. Therefore, in order to avoid false triggering, the setting range is relatively wide, such as 1200ms or 1500ms. Of course, the setting value is selected according to the specific implementation environment and object. For example, in a certain embodiment of multi-touch screen, it is set to 1600ms.

[0031] In G200, it is described that when a trigger lasts longer than T3, it begins counting according to a set time interval value, such as from V1 to Vn, assuming the end of the trigger is at Vn. During continuous fine-tuning of the video, the interval value depends on the user's reaction. For example, if the interval value is set to V=500ms, it means that every 500ms, it enters the next interval. For example, if T3 is set to 1600ms, after triggering and holding the trigger for more than 1600ms, it enters the V1 interval area. After holding for more than 2100ms, it enters the next interval area, as shown in the V2 interval area. When the holding time is more than 2600ms, it enters the time interval area such as V3. For example, if the trigger release time is 2800ms, it is exactly in the V3 time interval area, so the V3 interval area is selected. In the patent CN201910281922 (A Method for Controlling Volume in an Intelligent Electronic Device) invented and applied for by the inventor, the working logic of G200 is adopted. Combined with the short and long triggers to define the change of volume value, it solves the long-standing technical defect that single-touch devices, such as true wireless earphones, cannot adjust the volume with a single earphone. It also allows multiple functions to be implemented on single-touch devices, such as earphones. At the same time, it solves the legacy problem of adjusting volume through touch screen as a general function without the need for buttons on touch screen devices. The embodiments in this disclosure also adopt the technology of patent CN201910281922 (hereinafter referred to as patent 2), which can eliminate the dependence on control buttons or mobile phone volume buttons for video volume. At the same time, this patent is applied to adjusting video brightness.

[0032] In video fine-tuning, since the video is displayed on at least one two-dimensional plane, it can be controlled by directional triggering. Therefore, the G300 has two directional triggers with different trigger durations, for example, triggering to the right (similarly, in addition to the direction of the trigger, the duration from when the user touches the screen to when they leave the screen can be obtained by computing devices, including the computer's touch pad).

[0033] In environments requiring both duration-based and directional triggering, any directional trigger (regardless of left, right, up, down, etc.) will generate duration feedback on the sensor (such as the duration pulse under the directional arrow in G300). If a user slowly moves to the right, and the touch duration exceeds the T3 setting, if the directional trigger hasn't ended, it will be mistakenly identified as a G200 trigger with a duration greater than T3. Therefore, a tolerance value is typically set for directional triggering. A trigger displacement greater than the touchscreen's set value of 10 points (the tolerance value, and within the corresponding trigger time) is considered a directional trigger. If it's less than this value, it's not recognized as a directional trigger. Thus, a value greater than the tolerance value can be identified as a directional trigger, not a trigger greater than T3 in G200. (The tolerance value needs to be tested and set based on whether the direction is mistakenly identified as a duration-based trigger during implementation. Alternatively, it can be set larger, such as 30 points, which requires a larger range of movement on the screen during the user's directional trigger.)

[0034] Therefore, in order to implement the method of this disclosure on sensor groups capable of sensing trigger direction and trigger duration, including but not limited to touchscreens, capacitive screens, and computer touchpads, the setting of the T3 duration, if the trigger code includes short and long duration triggers in G100, requires consideration of both the pulse width of T2 and the duration of the tolerance value for direction triggering. However, for simplicity, the T3 duration is usually set to a trigger duration greater than the upper limit of the T2 duration setting. The upper limit of T2 is also set considering the directional trigger tolerance value. For example, in one embodiment, it is set to 1200ms, while in another embodiment, due to touch screen issues, it is set to 1600ms. The key factor in setting this value is to prevent any other type of trigger from mistakenly triggering a trigger that only starts working according to the time interval after exceeding the T3 upper limit. Therefore, the T3 setting is based on trigger coding and aims to avoid mis-triggering. When it is only a combination of directional trigger and G200 trigger, the duration of T3 only needs to be greater than the directional trigger tolerance duration value. For example, the tolerance value for directional trigger is set to 10 points, and the tolerance time is set to 800ms (which can be understood as a trigger touching the screen but displacing less than 10 points, which is judged as a non-directional trigger; the tolerance value can be used to judge a trigger that exceeds T3). The trigger value is triggered, but because it is twisted on the touchscreen and there is displacement, but the displacement does not meet the judgment condition for directional triggering, the setting value of T3 should be determined based on the duration of the trigger code to avoid misidentification. Generally, when short, long and directional triggers are included, the setting value should be greater than the duration of long-duration triggers, that is, greater than the duration range defined by T2, such as greater than 1600ms. In this application, the setting duration is used as a suggestion that if the triggering system includes duration triggers, the setting duration should be greater than the upper limit duration of T2. If there is no duration trigger pre-set, there is no possibility of misoperation, and the setting duration can be set with the tolerance duration of directional triggers as the lower limit, that is, greater than the tolerance duration.

[0035] G401 and G501 are two ways of representing a video. G401 represents the video according to the structure of the video frame sequence, while G501 describes the video according to the length (such as milliseconds). This is because once the frame rate of a video file is determined, such as 60 frames / second or 30 frames / second, its frames and video time are also corresponding. Therefore, a specific frame image can be located from the time value, and a specific frame can also be located from the frame sequence. Key image information and target image information are in the frame. From the perspective of duration, it is only necessary to read the frame corresponding to the set duration (the video time usually refers to the time value or duration value of the video from the beginning to the current position. For example, tx = 500ms means that the duration of the video from the beginning to the current position is 500ms or the time at the current position is 500ms).

[0036] In G401 and G501, there are corresponding fx or tx (fx represents a certain frame in the video, and tx represents a certain time in the video), which are used to express the current position of the video. G402 can be a pointer or a parameter used as a marker in a program. When using the G402 frame sequence to locate the position of a certain frame in the video, the frame sequence value is used, while when using the duration as the location, the time value is used, usually a millisecond value.

[0037] The following explains the working mode of directional triggering and triggering with a duration greater than T2. ​​In the diagram, G701 is a trigger combination of directional triggering and triggering with a duration greater than long or greater than T2. ​​The tw (interval value between triggers) between the two triggers must be less than the instruction window value, for example, setting the instruction window value to 800ms. This ensures that the system recognizes that the subsequent trigger is a subsequent trigger of the preceding trigger, that is, there is a trigger within 800ms after the previous trigger. Otherwise, the previous trigger code is judged to have ended the input, the instruction window is closed, and the system enters the instruction recognition and execution stage instead of waiting for the subsequent trigger.

[0038] In one embodiment, when a directional trigger is detected and there is a subsequent trigger within the window period, and the subsequent trigger is a trigger with a duration longer than the long trigger duration (a trigger with a duration longer than T2), after the trigger duration reaches the time interval, the image of the paused position read from the video or the image of the positioned frame is displayed in the image display area on the interface, according to the video time interval or frame interval.

[0039] To further explain, if the pause frame is an fx frame or the pause duration is tx, and the user wants to continuously fine-tune towards the end of the video to find and locate the target image, then a right-trigger is used (for ease of memorization, a right-trigger is defined as towards the end of the video). Then, in the trigger command window, hold triggers and triggers that are longer than T2 duration, longer duration, or a set duration are used. When the system or app determines that a subsequent trigger exceeds the set duration and continues to hold the trigger, it is determined that continuous adjustment is needed. Therefore, based on the set video frame interval or video duration interval, the system reads frames or extracts images at the duration position according to the frame interval or video duration interval, and displays them in G601. For example, assuming the video file is 30 frames per second, if the trigger for right-triggered and triggers exceeding the long duration are defined as displaying each frame in the G601 towards the end of the video, then after the system detects a right-triggered event, another trigger occurs in the command window, and the trigger duration is greater than the long duration (set duration). Since the trigger is not deactivated, the command is determined to display each frame towards the end, and the displayed image changes every time interval. That is, after exceeding the set duration, fx+1 frames are displayed, and in the second time interval, fx+2 frames are displayed, and so on. When the trigger is maintained until the nth time interval, the displayed frame is fx+n. If the trigger command defines two right-triggered events and triggers exceeding the set duration within the command window, and if the video frame interval is set to 5 frames, then the first time interval displays fx+5*1, the second time interval displays fx+5*2, and so on, until the nth time interval displays fx+5*n (displayed in the G601 display area).If you make continuous fine adjustments towards the beginning of the video, you can use a leftward trigger, triggering a duration longer than the set time within the window. In this case, if the video frame interval is A, then in the first time interval, fx - A * 1; if it is the second time interval, then fx - A * 2, and so on until fx - A * n deactivates the trigger. For example, if the frame interval is 3, the image displayed in G601 is the frame number of the paused position, such as fx-3*1, and so on, until the target frame. The user looks at the current image in the display area, and if it is in the correct position, the trigger is released. This fine-tuning method is more accurate than the mouse GUI method because releasing the trigger and pausing the image at the target position is more precise than the mouse or computer method, where even if the position is reached, the user still needs to click the mouse, so the accuracy is usually lower than this method. Of course, in the program, fx is usually assigned the value of the frame currently displayed in G601. So at each interval, such as fx=fx+A towards the end of the video or fx=fx-A towards the beginning of the file, the pointer is automatically changed according to the frame interval value every time interval, thus completing the pointer movement or fx assignment. Therefore, there is a dashed pointer in G402, which is used to indicate that the pointer is moving every time interval until it reaches the target position. The dashed pointer indicates the position that has been moved.

[0040] For a video duration-based method, assuming a 30-frame video, each frame is 1000ms / 30 = 33.333ms. Following the same method described above, assuming B is the video time interval, if it's frame-by-frame, B = 33.333ms. If the video time interval is 3 frames, it's 100ms. Therefore, each time interval is calculated as tx = tx + B (end direction) or tx = tx - B (start direction). For example, if the current tx is 1000ms, and the video time interval is 100ms (3 frames), assuming the direction is towards the end of the video, the time to extract the video image for the first time interval exceeding the set duration is 1000ms + 100ms. The second time interval is the current pointer tx time, 1100ms plus 100ms, which is 1200ms. In other words, as the pointer moves, the video of the next video time interval is the current time tx plus the video time interval value; so simply put, it is the video duration of the image displayed in the previous time interval (the currently displayed image, but not yet switched to the image corresponding to the next time interval). The video duration interval is the current time interval based on the direction of continuous fine-tuning plus or minus the video duration interval, and the image in the video is read out based on this time and displayed on G601.

[0041] It is important to note that triggers combined with triggers of duration greater than T3 or a set duration usually contain pre-encoding. That is, the meaning of the instruction is determined by the pre-encoding combined with the trigger of duration greater than the set duration. The pre-encoding not only expresses the direction, but also determines the video time interval value or frame interval value (the mapping relationship of the pre-encoding). The instruction recognition system displays the corresponding image in the display area according to the continuous display of the frame interval or video duration interval defined by the trigger instruction, based on the video status value, trigger encoding, and the current frame number or time.

[0042] Therefore, the above method essentially simulates the knobs on a professional video editing console. It continuously displays images near the target location, and by deactivating the trigger, selects the key target image and content (essentially resembling a constant-speed knob, whereas a console knob can adjust speed according to user needs).

[0043] Generally, directional triggers are more representative of video direction. Therefore, in the real-time example, a right trigger plus a trigger with a duration greater than T3 is used, with each frame of the video displayed on G601 at intervals. A right trigger plus a trigger with a duration greater than T3 displays the image of the fifth frame from the current frame towards the end in G601 at intervals (the frame interval is assumed to be 5, or the video duration interval is the corresponding 33.333ms*5). Therefore, by using frame intervals or video duration intervals, and based on the pre-triggered encoding and triggers with a duration greater than T3, the target position can be reached more quickly and accurately. This is equivalent to creatively using a trigger encoding greater than the set duration to simulate the knob for quickly locating images on a video editing console.

[0044] As for the pre-triggering of triggers with a duration greater than T3, directional triggering can be used, or short-duration, long-duration triggering, or a combination of directional triggering can be used. This invention is not limited to any particular method. However, since there are many video operation instructions and even more micro-operations, it is best to use user-derivable encoding logic to reduce the cost and difficulty for users to remember the trigger codes. In addition, it should be considered that if a general volume adjustment function is introduced, the trigger codes should not conflict.

[0045] In the aforementioned continuous fine-tuning, if the frame interval or video duration interval is not frame-by-frame, especially in videos such as magic tricks, martial arts action, motor vehicle accidents, and crime scenes, the target image is often not displayed in the previous interval, but has already occurred in the next interval. Therefore, non-continuous fine-tuning should be adopted. Specifically, an encoding method is used, such as triggering to the left or right in the video pause state, representing one frame or the corresponding video time interval. For example, triggering to the left means moving one frame to the left, i.e., fx-1 frame, or tx-33.33ms. For example, in continuous fine-tuning, if four frames are missed in the direction of ending, the user triggers to the left once. After execution, similar triggering three times will find the key target image. In magic tricks, the useful and recognizable images are often just one or two frames. In the embodiment, for non-continuous fine-tuning, such as trigger codes ->, ->->, or ->->-> (note that -> represents triggering in the right direction, and <- represents triggering in the left direction), which respectively represent triggering one frame, three frames, or ten frames to the right, or <-, <-<-, or <-<-<-, which respectively represent triggering one frame, three frames, or ten frames to the left, the user can quickly find the target image by combining continuous fine-tuning with their position in the current frame and the possible target frame image (by using the trigger code and the corresponding video frame interval or time interval value, the image is non-continuously displayed in the display area according to the direction and interval value defined by the trigger code, and the image is positioned after adjustment according to the direction and interval value defined by the trigger code).

[0046] The reason for introducing both continuous and non-continuous fine-tuning in the design is that finding the target image in a video is not easy. This is why professional video editing consoles have been developed and remain irreplaceable for decades. However, implementing this on touchscreens and other sensor components is not as convenient and fast as using rotary knobs and large screens. Therefore, after adopting continuous fine-tuning, if you still use continuous fine-tuning after missing the target image, it is usually not as good as non-continuous fine-tuning. For example, if you miss 1 or 2 frames, non-continuous fine-tuning is actually more convenient in finding the target image. If you only use non-continuous fine-tuning, in a long video, the user has to trigger the screen too many times, which is not as time-saving and improves the user experience as continuous fine-tuning. Furthermore, watching videos with a certain degree of continuity makes it easier to predict the target's location. When the video is discontinuous, the longer display time reduces continuity, making prediction more difficult. In terms of the image effect of continuous fine-tuning, it's similar to watching a continuous video at an extremely slow playback rate, allowing you to rewind or forward. Currently, this experience can only be achieved by manually adjusting knobs on a video editing console. It's also worth noting that to better achieve both precision and efficiency, differential fine-tuning can be used. For example, continuous fine-tuning uses slightly coarser intervals (larger interval values), such as 3 frames or 100ms, while discontinuous fine-tuning uses finer intervals (smaller interval values). During video frame finding, a minimum interval value can be set, such as 1 frame. This allows for continuous fine-tuning to get very close to the target position, rather than continuous fine-tuning that adjusts frame by frame to accurately find the target frame. For example, 1 frame or 33.33ms. The purpose is that when users use continuous fine-tuning for efficiency and to quickly reach the target position, they may miss key target frames or images every time interval. In this case, by deactivating triggers that exceed the set duration and using non-continuous fine-tuning, the target position can be accurately reached. It is not necessary to set the same video interval as in continuous fine-tuning. Therefore, the two types of fine-tuning, with interval difference, can be combined to achieve both efficiency and accuracy.

[0047] The fine-tuning features include the aforementioned continuous video fine-tuning method and non-continuous video fine-tuning; the non-continuous fine-tuning is performed in a paused video state, using triggering or triggering combinations, according to the video frame interval or video duration interval defined by the triggering or triggering encoding, and the triggering direction, to display an image in the display area that is repositioned and read from the video according to the direction and interval defined by the triggering, and the current image position is repositioned in the video.

[0048] In conjunction with video fine-tuning, in some implementation examples, the video fine-tuning method is used in the video editing function to fine-tune both sides of the video selection box, and the start and end positions of the selection box in the video are determined by the content of the image in the display area.

[0049] In conjunction with video fine-tuning, in some implementation examples, the fine-tuning method is applied to fine-tune to the target image (target frame) in video playback, video editing, or video shooting functions. This includes using the image of the target frame as the preview image that best reflects the video content. Specifically, for example, when a smart terminal records video, it is generally stored on the smart terminal, so each video has its own preview screen. However, the preview screen usually cannot reflect the video content, and over time, it is not easy to find previously recorded content among multiple video files. This requires the user to find the target image in the recorded video that reflects the video content or features (such as the most exciting moment or the most beautiful moment) and set it as the preview image. Therefore, the fine-tuning method and video control method of this disclosure are needed to quickly locate the target content after the terminal user shoots the video and set the located target image as the preview image of the recorded video for later use and retrieval. A specific implementation example is to set continuous fine-tuning, such as using a frame interval of 2 or 3 frames (differential fine-tuning, of course, efficiency can be reduced to only Continuous fine-tuning is used, while non-continuous fine-tuning can use a one-frame interval. When the user sees the video playing to the target position, the video is paused, and continuous fine-tuning is used to fine-tune it to a position closer to the specific target frame with high efficiency and less time. Then, if the displayed frame image is not the most accurate and clear target frame, the user uses non-continuous fine-tuning to fine-tune the frame forward or backward, thus finding the target frame with extremely high efficiency. Because the display is all in the main display area, it solves the problem of finding the target frame in the thumbnail, which is unclear and inefficient, as in traditional video editing (especially inefficient on small-screen electronic devices such as mobile phones). As for the target frame in the video, the image of the target frame, as mentioned above, can be set as the preview image of the video, or as the first frame of the video, or even as the video for a specific duration at the beginning of the video, such as the first 500ms or 1 second. This frame is used as the video within the set time period at the beginning of the video (e.g., 0 to 800ms from the beginning of the video). In this way, others see the frame that best reflects the video, found by continuous and non-continuous fine-tuning, as the preview or the first frame. Then, people become interested in or remember the content of the video. For others watching the video, the quality of the first few hundred milliseconds or seconds directly affects whether they watch the video. Therefore, in this application, the found target frame can be inserted into the original video as the first frame or as a video segment of a set duration at the beginning of the video. The specific frame insertion method is to form a video from the found target frame according to the set duration, and then combine the formed video with the original video into one video. In this way, the video segment generated by the found target frame is the beginning part, and the original video is the video segment after the segment formed by the target frame. When playing, what is seen first is the image most representative of this video for a few hundred milliseconds (set duration), and then the subsequent original video.Video compositing and image-to-video generation are conventional techniques in this field and will not be described in detail here.

[0050] Video fine-tuning includes any type of fine-tuning, and the triggering input area can be in the main image display area. This allows for tasks such as finding target positions in full-screen mode and editing without an editing interface.

[0051] For video editing functions, such as Figure 6 As shown, f1 is a GUI interface for a smart terminal or mobile phone, f2 is the image or video display area, f3 is a thumbnail information bar for video editing / trimming functions, containing multiple images extracted in video order, and f4 and f5 are video cropping or selection boxes, with f4 being the left edge and f5 the right edge. As is well known, current touch technology only allows users to select and frame the video to be cropped or selected by moving the left or right edge of the multi-touch screen with their fingers. Typically, the video thumbnail is displayed in the f3 bar, making it very difficult to select the desired video information. Furthermore, a knob is needed for precise cropping in a video workbench. Therefore, to achieve the capabilities of a professional video workbench on smart terminals and other devices, the above fine-tuning method is used, thus achieving the effect of a professional video workbench.

[0052] Here's a specific real-time example: When you tap or press F4, the left edge of the selection box moves to the target position (the actual selection box is like a video progress bar; the only difference is that a progress bar can only select one point, while a selection box can select two time points (or points in two frame sequences) in the video, so the video between the two selected points can be cropped or extracted). At this point, there's still a distance to the target point in the video. Therefore, if the finger's precision cannot make any closer adjustments, you need to use continuous fine-tuning and non-continuous fine-tuning methods. For example, first move closer to the target position relatively quickly (e.g., continuous fine-tuning), then use non-continuous fine-tuning, and finally move F4 frame by frame to the target frame, target content, or corresponding video time point.

[0053] However, given that the thumbnail of f3 is too small (its layout logic is based on traditional workstation video editing software or systems, which cannot be laid out in the small display space of smart terminals, so the thumbnail is too small to be used for judging the accuracy of the editing), in this disclosure, the positioned display image can be displayed in S203 and S202, where S203 is the video display area and S202 is the thumbnail display area. Even if a particularly small continuous thumbnail like f3 is used, the positioned image still needs to be displayed in a larger display area.

[0054] The display is best when done in S203, especially when fine-tuning the left and right edges of the selection box (f5 adjusts the right edge in the same way as the left edge).

[0055] Therefore, as described in the above embodiments, this disclosure achieves the same level of precision in video editing as professional video editing systems on a smart terminal touchscreen, with the precision required to select a specific video range. Traditional multi-touch technology, on the other hand, cannot efficiently fulfill the capability of fine-tuning and selecting images, including the complexity of the interface, controls, and the implementation of touch commands.

[0056] In addition, it should be noted that in this disclosure, the touch area can be located anywhere on the screen, such as in the S203 video display area or in a designated area. However, if the video is horizontally full-screen, the advantage of touch in the S203 display area is very strong. Thus, the desired result can be achieved without any controls or screen buttons that occupy or cover the screen. This is an advantage that traditional video playback and editing software do not have, as they require a lot of space to display controls.

[0057] Smart electronic devices such as smartphones, tablets, computers, and laptops typically have touchscreens or touchpads. These sensors or sensor groups usually consist of sensor components such as capacitive touchscreens and clock systems. The clock system times the trigger duration and converts the trigger into a time value. Triggers such as direction and gestures are recognized by specialized programs. For system programs, if it's for video fine-tuning, it might be the system program of the smart electronic device. For computers and smart terminals that can install apps, an app can be used to monitor the trigger information of the aforementioned sensors or sensor groups. Of course, in video fine-tuning, continuous fine-tuning requires sustained triggering. However, as mentioned above, in an environment where duration triggering and direction triggering are used together, to ensure accurate trigger recognition, the time values ​​of long-duration triggering and direction triggering must be determined. Therefore, to solve this problem, a tolerance value for direction triggering is introduced. So, in the case of mixed use of duration triggering and direction triggering, the upper limit of long-duration triggering is usually set to the "set duration." That is, a trigger exceeding the upper limit of long-duration triggering, such as 1600ms, is recognized as a trigger exceeding the set duration. After setting the duration, the system enters the set time interval judgment process, such as a time interval of 500ms (adjusted according to human reaction speed, such as 800ms). This means that after each interval, when entering the next interval, the image needs to be displayed as a different image. Therefore, the display direction and frame or time interval need to be predefined in the trigger code. For example, in the case of right trigger + duration longer than the set duration trigger mentioned above, the right trigger is the pre-trigger code. The pre-trigger code not only determines the trigger video display direction (towards the beginning or end of the video), but also needs to set the adjusted video frame interval or video duration interval. In this way, in each time interval, according to the trigger code settings, in the next time interval, at the position where the video was displayed in the previous time interval, according to the defined adjustment direction and adjustment interval (as described above), the video image is read and displayed.

[0058] Of course, when there is only directional trigger pre-coding, the set duration can be greater than the minimum tolerable duration value of directional trigger. Therefore, the so-called set duration should be determined based on the characteristics of the trigger encoding and the required duration, with the aim of avoiding misidentification of the trigger.

[0059] When users want to learn, see the details of a video, or extract information on the road like traffic police, so as to gain the consent of the parties involved in an accident, they need to find the target frame or image of the key information in the corresponding video content. Therefore, they need to use fine-tuning methods, watch repeatedly or even frame by frame, or crop and extract image and video information.

[0060] Even with a playback method, when a video is played (whether at normal speed or abnormal speeds such as 2x or 0.5x), it takes hundreds of milliseconds for the user to see key information, for the brain to process it, and for the body to react by touching the screen to pause. For images, even at 0.25x speed, several frames have passed. Considering the time it takes for the human brain to recognize image content and for feedback, it is obvious that key target locations may be missed. Since progress bar control (S201) cannot handle this situation, fine-tuning methods must be used to solve the problem.

[0061] Having explained the methods and necessity of video fine-tuning, let's now briefly explain video control methods, such as... Figure 3 As shown, in this disclosure, playback means, but is not limited to, playback at normal speed or abnormal speed.

[0062] Specifically, the technical methods used in this disclosure rely on sensors or sensor groups that can sense trigger direction and trigger duration. These are not limited to touch screens of smart terminals, computer touch screens, touchpads, or even sensor groups composed of fabric fibers and circuits that can identify and receive directional and duration triggers. Such sensors or sensor groups can be built into smart electronic devices, such as touch screens of smartphones and smart terminals, touchpads of ordinary laptops used for mouse and function input, or external touch electronic devices that can sense duration and trigger direction and are connected to smart electronic devices via wired or wireless connections.

[0063] The video control method disclosed herein is as follows: the user of the video determines the position of the target image or content by controlling the video's progress bar and based on the content displayed on the video's interface. Figure 3 (Step S101 in the middle); Generally speaking, since the video progress bar is based on the video length, such as 120 minutes, and the video length is proportional to the progress bar length, the lines on the video progress bar that indicate the progress (such as different colored lines indicating the progress that has been played and the progress that has not been played) or the indicators that control the movement, such as... Figure 2The black circular indicator on the S201. Therefore, when selecting a video using the progress bar, the progress bar will move forward or backward one position on the screen depending on the length of the video, causing the video to deviate by a few seconds or minutes (depending on the video length). As a result, video player users can usually only control the video to a certain approximate position, but this is still several seconds or frames away from the precise target position.

[0064] The video user controls the video using a progress bar and determines the viewpoint based on the content displayed on the video interface. Figure 2 In the middle, the display interface contains S202 or S203, where S202 is a thumbnail, that is, a video thumbnail corresponding to the current progress bar position. By controlling the indicator of moving the progress bar, the information in the thumbnail can be used to determine whether the general target position has been reached.

[0065] S202 can be Figure 2 Any location on the left side of the smart terminal interface S204, but not superimposed on or obstructing the main video display interface (image display area) of S203; it can also be... Figure 2 There are other ways to overlay or block the S203 video main display interface on the right side of the smart terminal interface. For example, instead of using the S202 thumbnail, the S203 can be displayed directly. The target position can be determined by dragging the progress bar and based on the content displayed by the S203.

[0066] As mentioned above, the progress bar corresponds proportionally to the video length. Therefore, the longer the video, the longer the progress bar moves (one point). Since one point represents a longer video length, whether you use a mouse to click or multi-touch to control the position of the progress bar marker, you can only reach a certain range of the target position. To solve this problem, professional video processing software and its associated control consoles use knobs to fine-tune to the required precise target position, or display each frame in sequence on a large screen computer so that the operator can search frame by frame. However, smart terminals obviously do not have such screen space.

[0067] Therefore, for ordinary video playback users, traffic police on duty, and editing users, it is obviously impossible to carry external devices (knob video control devices) to control the relatively precise positioning of the video.

[0068] In addition, users can also adjust the displayed content during video playback, that is... Figure 2 In the case where only the main display interface S203 is displayed without S202, the user playing the video can pause the video according to the content and then proceed to the next step. That is, the video playback is paused at the target position according to the content displayed on the video playback display interface. Figure 3 (Step S102).

[0069] Therefore, video users can control the video progress bar to reach the target position of the video content, and then play the video to reach the target position further, or they can control the video progress bar to reach the target position of the video content, and then play the video to pause when they are closer to the target position.

[0070] It is important to note that the video playback described in this application includes normal speed playback, which can typically be set to a speed value of 1, meaning 30 frames per second for a 30-frame-per-second image, and 60 frames per second for a 60-frame-per-second image. It also includes non-normal speed playback, such as setting a speed value of 2 or 1.5, which means 2x or 1.5x playback, commonly referred to as fast playback. Setting a playback speed of 0.25, 0.5, 0.75, etc., means video playback that is slower than normal speed, usually slow playback. 0.25 is equivalent to a playback speed of 1 / 4 of normal speed.

[0071] Imagine a user watching a long video and using a progress bar to get to a position near the target location. However, the user is actually some distance away from the target. Due to the touchscreen and multi-touch features, no matter how the user moves the control indicators on the progress bar, they cannot accurately reach the target location. Therefore, the user can play the video to get closer to the target location. However, it should be noted that playback can be forward at the user's selected speed or backward at the selected speed. Therefore, in the current video player mode, control space is needed to display these control buttons, which will further reduce or cover the main display space.

[0072] Therefore, in this disclosure, the technology of patent CN201811253657 (a method for operating and controlling a smart terminal or smart electronic device, hereinafter referred to as patent 1) is applied. The method of forming instructions is based on the trigger code and the function and state of the controlled target (APP or electronic device). Combined with the logic of the code, the user does not need to use the control method, but can use a simple touch method to control the video.

[0073] For example, in the settings such as playback mode, short, short, right trigger means 2x speed to the right, short, right trigger means 1.5x speed to the right, long trigger, right trigger means 0.5x speed forward, and long trigger, long trigger, right trigger means 0.25x speed forward (video end direction).

[0074] In human consciousness, short triggers are also fast triggers, while long triggers result in slightly slower, more hesitant behavior. Therefore, a short, short, right trigger can be perceived as "quickly moving forward," for example, executing a 2x speed command. Conversely, a long, long, right trigger command is perceived as "slowly moving forward," meaning executing a 0.25x speed playback command. Similarly, a short trigger combined with a right trigger is "quickly moving forward," and so on. "Quickly moving backward" corresponds to a short, short, left trigger code. This avoids the need for several conventional controls, allowing users to easily memorize or deduce multiple playback control trigger codes after a single repetition.

[0075] Therefore, for users who need to repeatedly view details of a certain location, touch commands are more convenient than controls and do not take up video content space (as they are squeezed or covered by the control interface). Moreover, the logic of triggering encoding and corresponding commands can be deduced after one or two uses. So for operations like video that require repeated searching for key information, the coded touch method is faster and more direct. At least users will not have to repeatedly bring up the playback control interface.

[0076] Regardless of the playback method, the reason for adopting the human-computer interaction technology of Patent 1 is to overcome many existing technical defects or obstacles in video playback control, such as the control covering the video display area.

[0077] When users want to learn, see the details of a video, or extract information on the road like traffic police to gain the approval of those involved in an accident, they need to find the target frames or images that contain key information in the video content. Therefore, they need to use fine-tuning methods to watch the video repeatedly, or even frame by frame, especially in the process of extracting evidence or studying criminal behavior.

[0078] Even with a playback method, when a video plays (whether fast or slow), it takes hundreds of milliseconds for the user to see key information, for the brain to process it, and for the body to react by touching the screen, such as to pause. For images, even at 0.25x speed, several frames have passed. Adding the time for the human brain to recognize image content and for feedback, it is obvious that key target locations may be missed. Since progress bar control cannot handle this situation, fine-tuning methods must be used to solve the problem.

[0079] In smartphone touch technology, especially multi-touch technology, fine-tuning is not possible. Therefore, to utilize the touch screen and enable fine-tuning, it is necessary to simulate the working logic of a rotary knob in a video editing console. However, users cannot install knobs on smart terminals. Therefore, in this application, two methods are used for fine-tuning: continuous fine-tuning and non-continuous fine-tuning. By using two types of fine-tuning or a combination of the two, users can accurately locate the target image in the video, just like using a video editing console. The combination of the two types of adjustment improves both efficiency and accuracy.

[0080] Therefore for Figure 3 In the S103 step (video control) method, it is usually necessary to control the video progress bar, determine the target position according to the image content in the video display interface, or pause playback after reaching the target position according to the content in the video playback display interface. Then, the video is controlled by the continuous fine adjustment of video S104 method or the non-continuous fine adjustment of video S105 method.

[0081] For example, video cropping Figure 6 In the video, F4 and F5 are actually special video progress bars that allow you to select the video range from two directions. As explained in the fine-tuning section above, by using fine-tuning to adjust the left or right edge, you can determine the positioning of the left and right edges in the video based on the image in the video display area, thereby enabling more precise cropping or video content extraction.

[0082] The methods described above can be applied to video target localization, precision cropping, selection, or precise location of a target image frame for learning, evidence, etc., making it possible on devices such as smart terminals.

[0083] Furthermore, it's important to note that in video editing, adding content to a video image at a predetermined target location—such as adding a display image, text, or effect—is difficult due to the limitations of multi-touch screens and finger thickness, making precise adjustments challenging, such as length, width, and angle (often resulting in unsatisfactory length, width, or angle). The fine-tuning method disclosed herein addresses this issue. Specifically, similar to the aforementioned video clipping frame, the target image location is located through fine-tuning or video control. Then, video editing functions are used to add content to the image, each piece of content marked with a corresponding frame. Based on the aforementioned method of fine-tuning the video clipping frame, clicking on each side frame of the added content, using direction and triggers longer than a set duration, and adjusting the granularity at intervals, allows for precise adjustments. The selected border moves according to the direction defined by the trigger. For example, clicking the top border triggers an upward trigger and a trigger for a duration longer than the set time, causing the top border to move upward according to the set pixel granularity until it reaches the height required by the user (the side border lines automatically connect, and the content within the border automatically stretches). If the sides also need adjustment, clicking on one side border and using a trigger for a duration longer than the set time allows for continuous fine-tuning. Of course, fine-tuning can be continuous or discontinuous. Furthermore, if the entire content needs to be moved, the content can be selected by tapping with a finger, and the entire content can be fine-tuned to the target position in the image using directional triggers and triggers for a duration longer than the set time. This addresses the problem of precisely adding content to a target video image using a finger, which is not easily achieved on multi-touch screens. Figure 7 As shown.

[0084] Figure 7In the video display area, if a region B is added, this region is the content area added to the video image. As shown in the figure, the content of this region is a picture of love. However, the user wants the region B to be larger to match the overall content of the video. So, after initially adjusting the size with their finger, they are still not satisfied. However, traditional multi-touch has limited the precision of its adjustment. Therefore, the above description is used. For example, the top border is B-UP, the bottom border is B-BT, the left border is BL, and the right border is BR. The user clicks on the border and then uses direction and continuous fine adjustment for a duration greater than the set time, or non-continuous fine adjustment, to move the border in set increments. When the entire region B is selected, the entire region B is finely adjusted on the screen. In addition, for rotation, if it is necessary to select the region B, for example, to trigger the rotation indicator with a certain trigger (such as long-pressing the region B until the rotation indicator appears), then use direction trigger (front trigger) combined with trigger for a duration greater than the set time to make it rotate right or left according to the set angle increment value, such as 1 degree. Based on the above description, those skilled in the art can address the issue of fine-tuning added content in video editing; however, it should be noted that when adding content to a video, additional control keys and menu areas are still required. For example, to add an image to the current frame, the file needs to be opened via the menu or control bar and pasted onto the current video display. Furthermore, to simulate classic video editing software, one can... Figure 7 As shown in f21, a rotary knob control is set up. After the previous trigger is triggered, if the user triggers the f21 control again within the window period, the selected target (such as borders, videos, or added content in the view image, such as the image or information in area B) can be continuously fine-tuned according to the previous trigger code definition. This makes the editing interface more classic and makes it easier for traditional professional editors to accept the application of new technologies under touch technology, rather than relying solely on traditional video editing consoles.

[0085] In the above description, for example in Figure 6 or Figure 7 In this disclosure, the aforementioned continuous fine-tuning, non-continuous fine-tuning, and differential fine-tuning techniques no longer process frame images (video), but are used for moving, rotating, zooming in, or zooming out of visible objects in programs or apps. As is well known, since its invention and widespread application at the beginning of this century, multi-touch technology has suffered from inherent limitations in all electronic devices based on it. Specifically, it cannot accurately, or precisely, adjust the size of small objects, such as moving or rotating them. Furthermore, for CAD, image editing, text editing, and text-and-image mixed applications, as well as game software, precise control of objects is impossible on multi-touch screens.

[0086] In the following embodiments, a jigsaw puzzle game is used as an example to describe the precise movement of a visually controlled target in the program. This example is similar to the movement, rotation, scaling up, and scaling down of controlled targets in CAD software, such as 2D targets, and 3D targets. It is also similar in graphics or the video editing mentioned above. That is, under various fine-tuning techniques, the visually controlled target can be moved, scaled up, and scaled down quickly and accurately.

[0087] Assumption Figure 7 In this context, B represents the target to be controlled in the various apps mentioned above. B may also serve as the border of the target content, which can be a two-dimensional or three-dimensional visual object. To zoom in, zoom out, move, and rotate this target within the app, we need to understand the limitations of current technology. Currently, when handling such targets, the finger needs to be pressed on the target, making it impossible to see the target beneath the finger, especially for small targets. Secondly, due to the limitations of multi-touch technology, stretching one side cannot be accurate and precise; usually, the entire target is stretched first, and then the finger is pressed to move it. If the target is obscured by the finger at this point, it cannot be seen or controlled. Therefore, current technology has obvious shortcomings, resulting in low efficiency and accuracy. To address this problem, firstly, the finger controlling the target should not be pressed on the target itself; secondly, regardless of size, the target should be fine-tuned, not just able to handle large targets. Therefore, based on this, [the following is a proposed solution]. Figure 7For example, the user first clicks on target B to select it. Once selected, a rotation box appears, assuming it has four sides: BL, BR, B-UP, and B-BT. If target B is small, a user's finger typically cannot select a specific side. This disclosure uses the technology from patent 1, employing a trigger duration longer than specified, for switching and rotation. Specifically, in the touch input area on the screen, such as area f2 or f1 (program-defined input area for touch input), after target B is selected, a trigger duration longer than specified is activated. After this duration, at intervals such as 600ms or 800ms (depending on the programmer's chosen interval), the specific side (BL, B-UP, BR, BT) is switched. The border of B is then color-coded to indicate the border corresponding to the specified time interval. When the selected border is the one chosen by the user, the trigger is deactivated, indicating that the side has been selected. For large targets, clicking on the target side is possible, but for small targets, clicking on the target side would challenge existing technology. Therefore, a click is used to select the target. Triggering every time interval for a duration longer than the set time, the color of the target B's border changes. The user can then select a specific border by deactivating the trigger. When choosing between selecting a specific edge or rotating the target, a rotation indicator needs to be added to the corresponding function or item for each time interval. When the user triggers every time interval for a duration longer than the set time and displays the current edge or rotation, the user sees the change in the edge or rotation on target B. Deactivating the trigger when the target is reached selects the content to be fine-tuned (the specific edge or target rotation). Alternatively, the user can click on the target, select it, and then long-press it. If the press duration exceeds a set value, such as 1500ms, a selection menu is displayed, allowing the user to choose a specific edge or rotation function.

[0088] The advantage of using a time interval that varies and triggers the object to rotate is that when an object rotates, its left and right positions may have changed. However, menus can only correspond to left and right values. For targets that have already been rotated, menu selection obviously has problems.

[0089] Therefore, the triggering method of File 1, which is longer than the set duration, is used to select borders, rotation, etc., which can be used to select the correct edge very intuitively. When the program is set, it only cares about the current edge and the edge corresponding to the current time interval or the rotation option after the triggering of the duration in a counterclockwise or clockwise direction. Then, the user can accurately select the specific target edge or select the rotation target based on the actual angle of the target.

[0090] Once the edge of the target object is selected, directional triggering and triggering for a duration longer than the set time are used in the touch area. The specific set movement pixels are fine-tuned at each interval, such as 1 pixel, 2 pixels, or multiple pixels (set by the user or the app). In other words, after the combination of directional triggering and triggering for a duration longer than the set time, the selected border is continuously moved according to the direction of the trigger at each interval.

[0091] If the border is moved to the correct position or is still a few pixels short, it is corrected with non-continuous fine-tuning, such as moving only one pixel at a time in the trigger direction. If it is desired to select other borders, a trigger is triggered in the touch area for a duration longer than the set duration. Starting from the current border (e.g., clockwise), the border corresponding to the current interval is displayed at intervals. After the user releases the trigger, the border corresponding to the current interval can be selected.

[0092] If the user selects to rotate, the trigger will be deactivated after the rotation prompt appears, and then the user can select to rotate. The rotation also has a direction, such as left rotation or right rotation. After the direction trigger is combined with the trigger for a duration greater than the set time, it will continuously make fine adjustments according to the set degree at intervals. For example, the set degree is 0.5 degrees or 1 degree, so the user can continuously watch the angle rotate.

[0093] To achieve faster adjustment and fine-tuning, a differential fine-tuning method can be used. Larger granularity values ​​are used for continuous fine-tuning, while smaller granularity values ​​are used for discontinuous fine-tuning. For example, when an edge is selected, each interval for continuous fine-tuning is 3 pixels, while for discontinuous fine-tuning it is 1 pixel. For instance, when moving the edge (BL), if continuous fine-tuning goes too far, discontinuous fine-tuning can be used to readjust the position.

[0094] If the touch area commands only include directional triggers longer than the set duration for edge or rotation selection prompts at intervals, or continuous fine-tuning triggers for movement or rotation per time interval, then the set duration only needs to be longer than the duration for distinguishing directional triggers. However, if the trigger area and trigger commands also include short and long triggers, then the set duration needs to be set to be longer than the long duration trigger. This allows for richer trigger commands, such as rapid changes, like short triggers combined with directional triggers combined with continuous changes longer than the set duration.

[0095] In the techniques described above, by employing direction triggering and any combination of one or more of continuous fine-tuning, discontinuous fine-tuning, or differential fine-tuning, precise control of the boundaries and angles of tiny targets in applications can be achieved.

[0096] As for the overall displacement of the target, after clicking to select the target, if the edge is not selected in the touch area, the whole is selected by default. You can use directional triggering combined with triggering for a duration longer than set, and make continuous fine adjustments at each time interval, such as rotating the target as a whole.

[0097] The above description places the trigger area outside the target. However, in games, targets are generally not too small. If the area inside the target is large enough, the target itself can be used as a touchable trigger area. This allows you to directly use the above methods on the target, such as fine-tuning the angle or making small movements, which were previously impossible to achieve effectively.

[0098] In the above description, we naturally think of the cursor technology created in the first generation of human-computer interaction technology, especially the use of a mouse and touch pad to move the cursor in graphical interfaces. However, in multi-touch or XR devices, due to the shortcomings of existing multi-touch technology, the cursor cannot be used accurately. In this disclosure, the cursor can also be used as a controlled target. For example, if a user taps the screen of an image editing software with their finger, the cursor will follow the corresponding tap point, but it will be a few to a dozen pixels away from the ideal position. With the technology of this disclosure, when the cursor is the selected target (usually a crosshair in image editing, and a timeline in video and music editing), a fine-tuning with a duration greater than a long time is triggered in the touch area. This forms continuous fine-tuning, and we can see the cursor moving on the screen according to the triggering direction. If the cursor reaches the position according to the triggering direction, the user releases the trigger, and the cursor stops. If it is misaligned because it is very close to the target point, non-continuous fine-tuning is used, and the cursor is moved by directional triggering to reach the accurate position. In this way, the shortcomings of multi-touch technology can be solved. When accurately moving the cursor on an XR (AR, MR, VR) device using a sensor on the controller, the technology disclosed herein can be used to control the cursor or other controlled targets. MR devices, which contain cameras to capture gestures, can determine that a target has been selected when a finger is pointing at it. If, after a directional trigger, the finger points back at that target (or a specific gesture such as a clenched fist or an index finger pointing forward, representing a fine-tuning trigger at intervals exceeding a set duration), this can be interpreted as continuous fine-tuning. The directional trigger is considered a pre-triggered trigger. If such a trigger code occurs within a set window period, the target can be continuously fine-tuned. Once the target is in position, the user's gesture is released, which the device interprets as de-triggering, thus stopping the fine-tuning movement. If the target is over- or under-triggered, non-continuous fine-tuning is used, such as directional triggering. When XR devices use video technology for gesture recognition, they are essentially determining the trigger direction and duration, only using image recognition to perceive the trigger.

[0099] In Invention Document 1, the inventor originally intended to solve the difficult problem of human-computer interaction in all scenarios, especially to overcome the shortcomings of first- and second-generation human-computer technologies in non-static scenarios that could not work in all scenarios. However, in the fine-tuning of human-computer interaction, combined with the technical foundation of Documents 1 and 2 and its application in specific fields, such as video, CAD, image, video editing, and document editing, the inventor uses the continuous fine-tuning, non-continuous fine-tuning, and differential fine-tuning provided in this disclosure, combined with the characteristics of the controlled target, to achieve unexpected and long-standing unsolved problems, such as fine-tuning the target under multi-touch technology, and the defect that APPs based on multi-touch technology cannot fine-tune the target; at the same time, it also solves the defect that XR devices cannot fine-tune the target.

[0100] In video control, the technology of Patent 2 can also be introduced. This technology uses duration triggering as a general function to adjust the volume of electronic devices. If this technology is used, the brightness can be adjusted in the video. For example, setting upward triggering and triggering with a duration greater than a long time will adjust the brightness to brighten it, and downward triggering and triggering with a duration greater than a long time will adjust the brightness to dim it. Using the technology of Patent 2, brightness adjustment can be added to video control.

[0101] Brightness adjustment is also very important in practice. For example, in outdoor environments under sunlight, the video must be brightened to see the screen clearly. When projecting from a mobile phone to a TV or projector, the brightness of the two devices often differs, so adjustments are necessary. The traditional approach is to exit the video application (or switch it to the background) and adjust the brightness using the system's brightness adjustment function on the smart terminal. However, this adjustment does not represent the actual brightness of the video, so multiple adjustments are needed to achieve the appropriate brightness (existing technology cannot be based on the brightness of the video itself). In this disclosure, a brightness adjustment function is added to the embodiments, so the brightness adjustment is based on the brightness display of the video, making it more efficient and accurate.

[0102] The following describes the application of the above methods to rapid response and video evidence collection in on-site law enforcement, while also introducing an evidence collection method to facilitate rapid response.

[0103] Firstly, for traffic police, there are numerous S401 cameras deployed on the road, some in fixed positions and others in non-fixed positions. Therefore, when deploying new cameras, it is necessary to record the installation location, the camera's corresponding angle (shooting angle), and the area and range covered by the captured images. For example, cameras at level intersections have fixed shooting positions, so the information can be entered into the database based on the camera's number, orientation, shooting angle, and coverage area. In addition, for previously installed cameras, a comprehensive survey can be conducted to obtain the camera's location and its corresponding number, position, road, orientation (shooting angle), etc. For example, the orientation can be an absolute direction, such as expressed in terms of east, south, west, north, or south, or a relative direction, such as a road from east to west, or a road leading out of the city. Therefore, this orientation can be expressed using the conventional or general methods used by accident handling personnel, or according to relevant standards if applicable. Of course, it would be even better if the range could be marked on a GIS based on the camera images. Through step S403, the location, shooting angle or direction, and approximate coverage of cameras in the city can be obtained. For coverage, in step S404, the data of the deployed cameras or the cameras obtained through surveys are entered into the database, marked on the GIS, or a camera layer is generated on the GIS, which includes the camera angle and coverage. For coverage, the area covered and included by the camera can be directly marked on the GIS, thus forming the data for each camera.

[0104] For law enforcement officers, such as traffic police, who are on-site users of camera footage, if a party is dissatisfied with or objects to the enforcement action, they need to retrieve video footage as evidence. Currently, the industry practice involves the police requesting the footage, and the officers searching for possible cameras at the accident site based on the parties involved or the incident record. They then painstakingly search for video evidence from each camera, one by one, to obtain evidence and persuade the party to accept punishment or have erroneous penalties revoked. This process is extremely inefficient and costly. Furthermore, the process of finding cameras usually relies on the user's understanding of the incident, rather than a definitive knowledge that a particular camera's recording contains the incident video, making the workload immense.

[0105] In cases of vehicles blocking the road, the inability to retrieve camera information in real time during rapid on-site handling often leads to disputes and dissatisfaction with the traffic police's actions, resulting in vehicles blocking the road and causing traffic congestion. Information that the parties involved can accept is usually stored in nearby cameras or in the backend system. Traffic police handling the incident also need to check each camera on the road to obtain video. In addition, existing multi-touch police terminals cannot allow traffic police to quickly locate the target image in the video. Therefore, current technology in the industry limits the use of camera resources for rapid handling.

[0106] Therefore, in this method, if law enforcement officers use a smart terminal S405 to query (S406) the camera information, camera angle, and coverage information of the current location, and if the camera information forms a GIS layer, then the law enforcement officers can directly or selectively obtain the camera information layer based on the location of the terminal they are using on the GIS display page of the terminal. If the current location is covered by several cameras, the law enforcement officers can directly retrieve the video of the camera in the covered event area based on the time range of the event.

[0107] However, some cameras store video at the front end, and the back end retrieves it via the network when needed. Other cameras store video in a back-end storage system. Therefore, when the S405 retrieves an image, it may read it directly from the S401 or retrieve it from the back-end management and storage system.

[0108] In fact, obtaining data such as the location, angle of view, and coverage of cameras, and effectively turning it into queryable or GIS layered information, is a very important step (efficiency). With layered camera information, law enforcement officers can find the corresponding cameras and videos more quickly.

[0109] Once video information is obtained, because some traffic accidents occur very quickly, law enforcement officers need to rapidly locate the target position on the road after acquiring video data from S401 or S407 highways. This allows for swift action and ensures that the parties involved are processed in the presence of video evidence. Currently, this is done using a computer with a mouse or a workbench with knobs, with the parties handling the situation from the law enforcement office. One reason for this is that multi-touch technology on smart terminals makes it difficult for on-site law enforcement officers to quickly and accurately extract key image information, thus hindering rapid on-site response. Therefore, to overcome the shortcomings of current multi-touch functionality in smart terminals, the S408 procedure was adopted. The terminal program uses the video control method described above, allowing law enforcement officers to quickly locate key images in the video. If multiple cameras cover a location, the videos from other cameras can be extracted according to the approximate time of the event, and the key content can be located in each video and displayed to the parties involved, thereby reducing disputes and improving efficiency, such as quickly clearing the road for accident vehicles. If the parties involved need images of key information from the video, or a key segment of video, the video can be cropped or the key target content extracted, watermarked, and sent to the parties involved, or information associated with a case number can be generated and made public on the law enforcement agency's external information platform. The parties involved can obtain the information on the public information platform based on the associated information such as the event number, the generated QR code, and the event information link. The above work on police terminals requires the technology disclosed in this paper to extract images accurately down to the frame level, because many accidents only provide a few frames of direct evidence, and traditional video control technology cannot accurately and quickly locate the image position.

[0110] When enforcing the law on the road, the video is often not bright enough under sunlight, so the brightness adjustment method mentioned in s408 can be used.

[0111] Of course, when fine-tuning video, combining continuous and non-continuous fine-tuning can quickly locate the target position in the video. Therefore, the fine-tuning or video control methods mentioned above are all applicable in rapid processing.

[0112] It is worth noting that if the parties involved need key images or cropped complete videos containing the incident extracted by law enforcement officers, the officers can download them through interactive software or links on the law enforcement agency's specific website. If the video needs to be transmitted to the parties involved on the spot, it needs to be watermarked with the law enforcement agency's information, date, and the information of the officers handling the incident. This is achieved through the specific functions of the Quick Handling APP, so that the parties involved have on-site evidence such as images and continuous videos available in insurance or subsequent legal proceedings, as well as process proof from the judicial agency such as watermarked information such as time, incident number, and officer number.

[0113] Of course, law enforcement officers use a wireless network to retrieve images captured by cameras at the scene, but they need a secure channel such as a VPN to access the cameras of law enforcement agencies in order to securely obtain data, while others cannot access the cameras of law enforcement agencies at will.

[0114] exist Figure 5 The diagram illustrates a simple, rapid response scenario. C1 represents the camera on the GIS layer, along with its camera angle and coverage area. C2 is a bird's-eye view camera with a larger coverage area. A traffic accident occurs within the shared coverage area of ​​both cameras. Law enforcement officers, using mobile devices, rely on their location data and the GIS layer containing the camera's image, angle, and coverage area. While C1 should theoretically capture video of the accident, a vehicle carrying excessively tall cargo obstructs its recording. Therefore, the officers use the less ideal C2 camera on the GIS layer, retrieving video from it based on the approximate timeframe of the accident. Then, employing S408 (video control method), they quickly and accurately locate the accident. The key image information and relevant video footage were obtained, so that both law enforcement officers and the parties involved could agree on the accident handling and quickly clear the road. Instead of the current on-site handling by law enforcement officers, the parties involved are dissatisfied and there is controversy, resulting in long road closures and congestion. In addition, because there is no on-site video evidence to make everyone accept the handling, the relevant legal procedures are followed, which also requires the back-end processing staff to review and retrieve a large amount of video. This work is more difficult than the on-site location based on smart terminals. Since there are always accident parties who are not satisfied with the on-site handling, there are always people going through the legal procedures, resulting in a backlog of cases and a long queue of people waiting for reconsideration. Therefore, the overall efficiency is not high. Both law enforcement officers and accident personnel have complaints about the current state of rapid handling.

[0115] The aforementioned rapid response method can effectively solve the problem of rapid response to current events and address a long-standing issue that appears to be limited by existing terminal touch technology, especially video control technology.

[0116] Furthermore, GIS layers can be integrated with camera management systems. Because outdoor equipment is exposed to harsh environments and is easily damaged, when it is damaged, the GIS layers can display the information in real time and in a linked manner. This saves on-site personnel from wasting time trying to extract images from broken equipment. On-site law enforcement officers and video retrieval personnel are usually not equipment maintenance personnel, so the more real-time and accurate the information, the more efficient the operation will be. And the rapid response of traffic police is directly related to the traffic congestion status of a city.

[0117] To facilitate image access for law enforcement personnel, a unified management interface could be implemented on the S407. When on-site personnel access the video via their terminals, they would only need to enter the camera number covering the current accident location and the approximate time range. The S407 would then handle the entire process, regardless of whether the video is stored at the headend or the backend, making it easier for on-site personnel to access the video.

[0118] When processing or collecting evidence on-site, the on-site evidence collectors must obtain the target frame. Therefore, the fine-tuning method disclosed herein is needed to quickly find the target evidence. In addition, a large number of other videos are useless, so the evidence portion needs to be extracted. Therefore, the most critical frame (target frame) also needs to serve as a preview image or first frame of the evidence. Then, the evidence collectors can quickly find the video file based on the preview image among multiple videos on the terminal. However, whether the evidence frame serves as the starting video segment of the evidence collection video, such as 500ms or 1s as mentioned above, depends on the legal requirements related to the evidence.

[0119] This disclosure also includes some useful functions. For example, during full-screen playback, users can trigger commands on the video display screen to control the video, such as sharing the current image with others. To make the trigger code easy to remember, an upward trigger is used. When remembering the code, the user remembers it as "flying over the net," meaning that when the video is paused, the upward trigger is the function to share the current image. For saving locally, a two-finger downward trigger can be used. Since two fingers need to be bent to ensure simultaneous triggering, similar to a chicken claw digging downwards for food, "digging in" makes it easy to remember as saving the current image locally. In addition, if it is necessary to pause and fine-tune to a certain image during full-screen playback to use as the starting point for cropping, a downward trigger is used, which is remembered as "cutting a slice." A "cutting a slice" at a subsequent or previous position is also a downward trigger, which is the next cropping position (this application is not limited to only starting from the previous position). (The following text appears to be a fragmented and nonsensical collection of phrases and sentences, making a coherent translation impossible. It includes phrases like "cutting backwards, and the need for a cropping window," "the embodiment using the technology of document 1," "the innovative method of this disclosure," "breaking the conventional video editing," "improving accuracy and efficiency," "avoiding the problem of screens full of controls and editing windows," "the main image that should be clearly displayed being compressed or covered by these controls," "the existing technology is not ideal for extracting evidence or valuable target images," "especially on horizontal screens," "traditional video cropping techniques on smart terminals make the main image space very small," "so it can only be implemented in a vertical layout," "the fine-tuning and editing techniques in this disclosure can be implemented in horizontal full-screen playback," "for keys or buttons, this disclosure obviously does not require keys (controls), and can control the video with touch using the most ideal screen utilization, including fine-tuning, which is not achievable with existing touch technology," "for keys or buttons without buttons," "the ideal screen utilization using touch," "the need for touch control ...

[0120] This disclosure describes methods for continuous video fine-tuning, video control, and rapid response. Video control is currently widely used, but there are still unresolved problems. This disclosure provides a technical method to solve these long-standing problems. Rapid response is also a difficult problem that has plagued law enforcement for many years. On-site personnel are well aware that existing human-computer interaction technologies cannot solve the problem of quickly locating the target image position using image control technology. The method described in this application provides a solution.

[0121] All of the above methods can be implemented via an app. The app applies these methods and is installed on a smart terminal, utilizing the multi-touch screen of the smart terminal. Using the technical methods disclosed herein and those in patent documents 1 and 2, video control and rapid processing can be achieved. Furthermore, televisions are becoming increasingly intelligent, and touchscreens are likely to become widespread, similar to interactive displays in schools. Therefore, the method disclosed herein can solve many problems affecting teaching, such as quickly locating key positions in courseware and repeatedly showing students short pieces of key content. Currently, relying on teachers operating a few traditional video control buttons is inefficient, time-consuming, and distracting for students. In addition, the extraction of the target frame image can be used for multiple purposes, such as evidence images, video preview images, or editing (e.g., graphic design, AI editing) and then used as the beginning of the original video. The target frame has a crucial impact on video dissemination and is vital in evidence; therefore, depending on the application requirements, the target frame is extremely important.

[0122] The above method is also applicable to computers such as laptops. Laptops usually contain a touch pad, which can be used to detect duration and direction triggering. However, in the absence of a mouse, it is not possible to accurately edit images or fine-tune targets. Using the technology disclosed herein, a touch pad, touch bar, etc., can also solve many of the problems that the original technologies could not achieve accurately and quickly.

[0123] Those skilled in the art can implement the above-mentioned methods or use the above-mentioned technologies in APP or system programs based on the methods disclosed in this disclosure and the technologies disclosed in Patent Documents 1 and 2, thereby solving the long-standing and unresolved problems. Obviously, some embodiments in this application, such as the continuous and non-continuous fine-tuning technology of file time, can be applied to audio precision trimming in situations where multi-touch touch screens and touchpads are used as inputs.

[0124] Of course, for dedicated smart electronic devices, the method disclosed in this application can also be used to implement the method directly as a system program of the smart electronic device, thereby achieving the goal that the method can achieve on the dedicated smart electronic device. In addition, computers or laptops with touch pads can also implement the application of this method on laptop computers by writing programs, calling hardware resources, and using the control method disclosed in this application. This allows precise editing of video control to be used on a wider range of smart terminals and computers without relying on peripherals required for fine editing. Furthermore, this method is also applicable to a large number of apps that require precise control. By using the above-mentioned continuous fine-tuning, non-continuous fine-tuning, and differential fine-tuning methods, precise and fast control can be achieved using non-precise touch means, which is difficult to achieve even with traditional human-computer interaction technology.

Claims

1. A method for video fine-tuning, characterized in that... include: It includes continuous fine-tuning, which is applied to intelligent electronic devices containing sensors or sensor groups that can sense duration triggering and orientation triggering. Triggered by detection and identification through system programs or apps running on the smart electronic device; If the trigger includes a duration longer than the set duration, then at each time interval, according to the adjustment direction and video frame interval or video duration interval preset by the identified pre-triggered code of the duration trigger, the image of the corresponding frame or video time after adjustment is read according to the adjustment direction and the video frame interval or video duration interval, and displayed in the display area. The user can then deactivate the duration trigger based on the displayed image.

2. The video fine-tuning method according to claim 1, characterized in that... include: It includes non-continuous fine-tuning, which is performed when the video is paused, using triggers or trigger combinations, according to the video frame interval or video duration interval defined by the trigger or trigger encoding, and the trigger direction, to display the current image's position in the video in the display area, and to reposition and read the image in the video according to the direction and interval defined by the trigger.

3. The video fine-tuning method according to claim 1, characterized in that... include: It includes differential fine-tuning, which combines continuous fine-tuning and non-continuous fine-tuning using different set video frame intervals or video duration intervals; wherein the continuous fine-tuning uses the coarse interval, and the non-continuous fine-tuning uses the fine interval; the user first approaches the target position through the continuous fine-tuning, and then adjusts to the target position through non-continuous fine-tuning.

4. The video fine-tuning method according to claim 1, characterized in that... ,Include: The pre-encoding for triggering includes duration triggering, direction triggering, or a combination of duration triggering and direction triggering; the duration triggering encoding also includes short duration triggering and long duration triggering; the pre-encoding and the duration triggering with duration triggering greater than the set duration constitute the triggering encoding for continuous fine-tuning of the video.

5. The video fine-tuning method according to claim 1, characterized in that... ,Include: The position of the video in the previous time interval refers to the frame position of the displayed image in the video or the position of time in the video during the previous time interval. The video frame interval or video duration interval is the granular value in the video fine-tuning, and the granular value is set by the program. The pre-triggered state determines the particle value and direction for continuous fine-tuning.

6. The video fine-tuning method according to claim 1, characterized in that, include: In the video editing function, the video fine-tuning method is used to fine-tune both sides of the video selection box. The start and end positions of the selection box in the video are determined by the content of the image in the display area.

7. The video fine-tuning method according to claim 1, characterized in that, include: The display area includes a thumbnail display area or a main video display area; the triggered input area can be in the main image display area.

8. A method for video control, characterized in that, include: Users can control the video's progress bar to determine the target position based on the image content in the video display interface or pause playback after reaching the target position based on the content in the video playback display interface. Then, the video can be controlled using the video adjustment method described in any one of claims 1-7. If the target frame is found, the user can set that frame as the preview frame, first frame, or first segment of the video.

9. A method for video forensics, characterized in that, include: When installing a camera, or by investigating an already installed camera, record information about the camera, including its installation location, camera orientation, and the area covered by the camera's image. Enter the aforementioned information data; Alternatively, the information data can be labeled on a GIS. Alternatively, a GIS layer for the camera can be generated based on the aforementioned information data; The user of the camera video can query the entered information data based on the location of the handheld smart terminal, or display the image captured by the camera covering the location and its coverage area on the GIS layer based on the location data of the user's handheld smart terminal. The caller of the video reads video images captured by the queried camera or image, covering the camera at the location, according to the time range of the event.

10. A method for video forensics according to claim 9, characterized in that... include: The caller of the video, on the handheld smart terminal that applies the method of any one of claims 1-7, locates the target video in the video image of the queried camera.

Citation Information

Patent Citations

  • A method for controlling smart terminals or smart electronic devices

    CN109462690B

  • A method for controlling volume in an intelligent electronic device, the electronic device, and headphones.

    CN110087160B