Video application method and system
By introducing target data generation and video dissemination modules into video applications, the problem that existing video technology cannot effectively display target content in videos is solved, the in-depth display and interaction of the targets are achieved, and the video dissemination effect is enhanced.
Patent Information
- Application Number
- CN202411757479.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-17
AI Technical Summary
Due to the defects of unidirectional propagation and imaging technology, the existing video technology has poor video propagation effect and cannot effectively display the target content in the video, affecting the viewing experience and information transmission.
By introducing a target data generation module and a video propagation module in a video application, the user terminal program can calculate and obtain the target position data in the video frame, judge the target to which the triggered position belongs, and read the target preset information to display or execute the preset program.
It realizes in-depth display and interaction of the goals in the video, enhances the effect of video dissemination and user experience, and can clearly display detailed information that cannot be displayed in the video.
Smart Images

Figure CN120166261A_ABST
Abstract
Description
Technical Field
[0001] Video interaction technology, video dissemination ability enhancement technology, which also includes supporting technologies such as image recognition, feature comparison, image segmentation, image target analysis, and data communication, and is also the next-generation Internet application technology. Background Art
[0002] The current video dissemination continues based on the idea of traditional video broadcasting, that is, one-way video dissemination; in addition, the content of the targets in the video (PPTs, film and television content, textured commodities shown in the video) is limited by the video shooting equipment, lenses, encoding, transmission mode, and the modes of the user terminal and the player. As a result, for targets in the video such as planar targets or three-dimensional targets, their clarity or the detail of the content is limited by video encoding and video composition. Therefore, detailed and clear content cannot be effectively transmitted to video viewers or players.
[0003] For example, a recorded screen video of a teacher, which involves planar information of the targets in the recorded screen such as tables, pictures, code areas, etc. Due to the pixel of the recorded screen, such as 1080P, the content still cannot be clearly displayed, thus affecting the learning effect of the students on the viewing side (such as showing the molecular structure or a large amount of program code, the details of a classical realistic oil painting).
[0004] For example, a video shown in a museum, which contains multiple famous statues (three-dimensional targets) and oil paintings (planar targets). However, when viewers hope to clearly view any target, the existing video methods cannot clearly display the content of the target, which affects the viewers' mastery of the content in the video and the viewing experience at the same time (such as an oil painting in the video background, the focus of the video is on the narrator, while the viewer wants to carefully view the background oil painting).
[0005] For video dissemination in the forms of online sales, marketing, or speeches such as live broadcasts, the targets (including commodities, background projections) in the video are either foreground or background composition elements, so they cannot be shown in detail. For example, the display of materials and textures. Therefore, in most cases, it is the influence of the live broadcaster's words on consumers, rather than the information that the target commodity itself should display, that cannot be clearly shown (such as a book shown, what the book buyer is interested in is certain content in the book, rather than the seller's own perspective explanation).
[0006] Given the mobility of the terminal and the mobile phone screen, it is impossible to clearly and fully display the targets in the video based on the current mode. Therefore, it is normal that the effect is not good on the terminal, and the depth that the dissemination needs to achieve cannot be realized, and the sales effect that effect advertisements and the like need to achieve cannot be realized.
[0007] In film and television technology, film and television dramas require interaction, which means they need to interact with the audience based on specific goals. However, the current A and B drama model, which is derived from game ideas, hinders user viewing and flexible interaction, which in turn limits the realization of deeply interactive dramas.
[0008] Traditional video conferencing uses T.120 to share files, but in reality there are a large number of non-real-time videos and group live broadcasts, and traditional video conferencing ideas cannot cope with the dynamic and personalized demands of the user side.
[0009] In fact, the video technology currently widely used on the Internet has a large number of defects and cannot adapt to scene requirements. The above description is only a small part of the defects. The following disclosure contains several embodiments for solving specific problems. It is worth mentioning that the method of using pictures such as Google's patent to identify pictures on the Internet, retrieve information and display it to users is completely contrary to the problem to be solved by this disclosure.
[0010] In fact, to sum up, the current video technology, due to the technical bias of one-way communication and the defects of imaging technology, has led to the limitation of communication effect, not limited to content, goods, transactions, etc., and the video that is finally seen is limited to the video itself, and cannot bring the influence, extension and multi-dimensional transmission of information. Therefore, the present invention uses a new technical method to solve the above-mentioned technical defects of the current video. Summary of the invention
[0011] In a first aspect, the present disclosure provides a video application method, comprising the following contents: The terminal user watching the video triggers the target in the video being played on the terminal; The user terminal program calculates and obtains the position data triggered in the video frame; The user terminal program or the video transmission module determines the target to which the position belongs at the current video position according to the target information input by the video generator or the circled target, and generates target data; The user terminal program reads the preset information of the target and displays it on the user terminal, or reads the preset information of the target and executes the preset program or instruction.
[0012] In combination with the method disclosed in the first aspect, in some implementation examples, the targets in the video may include two-dimensional targets and three-dimensional targets; the two-dimensional targets include information in the background, such as projections, photos, pictures, and two-dimensional artworks; and the three-dimensional targets include three-dimensional goods and objects; in addition, the targets also include people in the video.
[0013] In combination with the method disclosed in the first aspect, in some implementation examples, the information of the target input by the video generator includes a picture of the target, and the picture is used for feature extraction and comparison.
[0014] Combined with the method disclosed in the first aspect, in some embodiments, the generating personnel demarcate a target in the video display interface of the target data generation module; The target data generation module generates target data of the demarcated target based on the time value or frame sequence value of the video.
[0015] Combined with the method disclosed in the first aspect, in some embodiments, the generating personnel of the video uploads the preset information of the target to the video dissemination module through the video generation end program or directly.
[0016] Combined with the method disclosed in the first aspect, in some embodiments, the target in the displayed video is demarcated. The demarcation method includes using a human-computer interaction method, such as using a mouse or a touch screen, using a geometric frame to select and demarcate the target. The content of the target includes planar targets and three-dimensional targets. Among the planar targets, there are targets with dynamic content, including playing PPTs, PPTS, and videos.
[0017] Combined with the method disclosed in the first aspect, in some embodiments, the generation of target data includes using methods involving artificial intelligence and image processing, and generating target data based on the timeline of the video, the position and trajectory of the target in the video.
[0018] Combined with the method disclosed in the first aspect, the preset information of the target by the generating personnel includes files or resource links; the forms of the files include pictures, PDFs, PPTs, vector files, and video files; and the resource links include web links and video links. The video links also include video streams.
[0019] Combined with the method disclosed in the first aspect, the generation end program or the video transmission module generates the feature information of the target based on the target set by the video generating personnel (such as directly selecting on the video) or the imported picture of the target.
[0020] Combined with the method disclosed in the first aspect, if the user of the terminal watching the video triggers a target in the video and the corresponding target is determined, the information interface related to the target displayed on the user terminal may also include an interaction area, including interaction content with service personnel; the interaction area of the audience; the service personnel also include robotic service personnel. The service personnel further describe, provide, and communicate with the terminal user about the information of the target. The interaction area can be in the form of a web page, a window, or other conventional forms on the program side in the terminal, and the AI robot can be in the form of a window or an animation, combined with TTS, to enhance the user experience.
[0021] Combined with the method disclosed in the first aspect, when a user watching a video clicks on the face of a person in the video, if it is determined that the target is triggered, according to the preset next-step instruction and preset information of the target, the instruction is executed, and the executed instruction includes adding attention or establishing communication.
[0022] Combined with the method disclosed in the first aspect, if there are multiple targets in the video and the target type attributes are different, the target data generation method and the feature comparison method are combined and applied, and the target data generation method is used for targets with dynamic content changes.
[0023] Combined with the first aspect, the present disclosure provides a method for video application. Specifically, it includes the following content: The video editing, recording or generating personnel (video generating personnel) import the video file or the video being captured into the target data generation module; The generating personnel circle the target in the video display interface of the target data generation module; The target data generation module generates the target data of the circled target based on the time value or frame sequence value of the video; The generating personnel import the preset information of the target into the target data generation module or the video dissemination module; The end user uses the end user program, accesses the video dissemination module via the network, and obtains the video file or the target data of the circled target in the video; When the end user is watching the video, the target is triggered; The end user program or the video dissemination module determines the triggered target according to the target data and the position of the trigger point in the video, and displays the preset information of the target, or reads the preset information of the target and executes the preset next-step instruction.
[0024] Combined with the above method, in some specific implementation examples, the personnel use a human-computer interaction method to circle the target in the video displayed in the target data generation module. The circling method includes using a human-computer interaction method, including using a mouse and a touch screen, using a geometric frame to select and circle the target. The content of the target includes planar targets and three-dimensional targets, and among the planar targets, there are targets with dynamic content, such as playing PPT, PPTS, videos, etc.
[0025] Combined with the above method, in some specific implementation examples, the target data generation module further includes a target detection module. The target detection module uses methods including artificial intelligence and image processing, and based on the time line of the video and the position and trajectory of the target in the video, generates target data (this data usually tracks the target in the video with a rectangular frame).
[0026] Combined with the above method, in some specific implementation examples, the target data and the file corresponding to the target that can clearly reflect the target are generated in the same file; while in other examples, the generated target data, the detailed file of the target, and the video file are separated.
[0027] Combined with the method described above, in some specific implementation examples, the video dissemination module can disseminate videos in the form of video streams or in the form of video files, such as using HLS, DASH, RTP / RTSP, multicast, etc.; in addition, the video dissemination module will also disseminate the preset information corresponding to the target separated from the video, including the file corresponding to the target, digital links; when the video is uniquely encoded in the video dissemination module, the program of the end user, according to the unique encoding of the video, if the user clicks to trigger the video and determines the target to which the click position belongs, reads the preset information of the target in the video and displays it, or executes the preset next instruction.
[0028] Combined with the method described above, in some specific implementation examples, the target can include planar targets and three-dimensional targets; planar targets include information in the background, including projections, photos, pictures, planar artworks; while three-dimensional targets include three-dimensional commodities, items, etc.
[0029] Combined with the method described above, in some specific implementation examples, the generation of target data is based on the size of the original imported video to generate target data, and the value of the target box of the target data can be an absolute value or a relative value.
[0030] Combined with the above method, in some specific implementation examples, when the program of the user terminal displays the target in the video, it can include a target box or not; of course, it can also generate a three-dimensional target based on the feature information of the target to enhance the user experience. For example, the terminal program segments the target based on the video and target data to form a three-dimensional target, thus increasing the sense of technology.
[0031] Combined with the above method, in some specific implementation examples, when the user sees the target on the video display interface of the terminal and clicks or triggers the target, the terminal program determines the triggered or clicked target according to the feedback information of the triggered position, the current frame or time information of the video, and the generated data of the target; of course, the video dissemination module can also be used for judgment. After the user terminal receives the trigger, it feeds back the position clicked in the video to the video dissemination module, and the video dissemination module judges the specific target based on the target data and feeds it back to the terminal program.
[0032] Combined with the first aspect, the present disclosure provides a method for video applications, specifically as follows: The editor, recorder, or generator of the video (the video generator) imports, through the video editing / generation terminal program, a video file or a video stream being captured, and the preset information of the target in the video (including the file or digital connection corresponding to the target, the preset connection, including video, file, web page connection, etc.). The video editing / generation terminal program uploads or imports the video and the preset information of the target in the video to the video dissemination module. In addition, the program or the video transmission module generates the feature information of the target based on the target set by the video generator (such as directly circling on the video) or the imported target picture. The end user uses the end user program to obtain the video file through the network via the video dissemination module. When the end user watches the video, the end user triggers and clicks on the target. The end user program obtains the position of the clicked target in the video. The end program or the video dissemination module determines the target to which the clicked position belongs based on the feature data of the target. If the set target is matched, the preset information of the target is displayed on the end user program, or the preset information is read and the preset next instruction or program is executed.
[0033] Combined with the above method, the user of the terminal clicks or triggers the target, and the software of the terminal reads and displays the preset information of the target (the corresponding file information or resource link); the resource link includes web page links and also includes live video streams.
[0034] Combined with the above method, the end user program of the terminal can display according to the format of the file in the preset information of the target. The format of the file can include file forms such as tables, pictures, vectors, etc., and can also include video files; the user can adjust the preset display area to clearly view the details that cannot be displayed in the video, such as zooming in or out on the preset information in the display area (such as vector pictures, file pictures, PPT or PDF page turning, etc.).
[0035] Combined with the above method, after the user of the terminal clicks or triggers the target in the video, the information interface related to the target displayed can also include interaction content and interaction areas for service personnel. The service personnel also include machine attendants, and the service personnel further describe, provide, and communicate with the end user about the information of the target.
[0036] Combined with the above method, after the end user clicks or triggers the target in the video, the messages of other users in the information interface corresponding to the target can be displayed, and other people can also see the message information.
[0037] Combined with the method described above, after the end user clicks on the target in the video, they can purchase the target product, receive red envelopes, and discount cards in the information interface that displays the target.
[0038] Combined with the above method, the attribution determination of the clicked area can be carried out in the following manner: The target picture is matched with the triggered frame. If the matched range contains the triggered position, the triggered target is determined. Alternatively, in the video frame, extract the image of the trigger point position, extract features, and compare them with the features of the target to determine the clicked target. Among them, it is further subdivided into extracting the picture of the clicked position according to the set size or segmenting the image of the clicked point target, and then comparing it with the features of the target to judge the clicked target.
[0039] Combined with the above method, in some embodiments, when there are multiple targets in the video and the target type attributes are different, the method of generating target data is combined with the method of feature comparison. The method of generating target data is used for targets with dynamic content changes.
[0040] Combined with the above method, if the user clicks on the face of a person in the video and it is determined to be the target, then according to the preset next-step instructions for this target, execute the next-step instructions such as following the official account and communication.
[0041] Combined with the above method, after the user clicks on the target item in the video and it is confirmed that the set target is clicked, the next execution instruction includes putting the item into the shopping basket. Description of the Drawings
[0042] Figure 1 、Method based on target detection Figure 1 ; Figure 2 、Schematic Figure 1 ; Figure 3 、Schematic Figure 2 ; Figure 4 、Method based on feature comparison Figure 2 ; Figure 5 、Method diagram; Detailed Implementation Manner
[0043] The specific implementation manners, coding, quantities, numerical values, durations, triggering methods, schematic diagrams, next-step instructions, programs, etc. described in the following exemplary embodiments do not represent all the implementation manners and embodiments consistent with the present disclosure. On the contrary, they are only some application methods and systems of videos consistent with the present disclosure described in the appended claims.
[0044] The following clearly and completely describes the technology and method of the present disclosure in combination with the accompanying drawings and embodiments in the present disclosure. Obviously, the following described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments.
[0045] The technical logic of the present disclosure is to use video as the basis for dissemination, and the targets in the video, as the contact points for interaction. Based on the preset of the video producer, the targets are triggered in the video, and then displayed or executed according to the preset for the next instruction. Thus, with multi-dimensional information, the dissemination efficiency of the video is improved. This itself is also an effective and dynamic change to the mechanical setting control modes such as Flash and H5, thereby improving both the dissemination efficiency and the visual experience of the video.
[0046] When the video exists in the form of a file, such as a video teaching material recorded by a teacher for students to watch and learn remotely (video on demand, which can be in the form of a file or a slice, and the present disclosure does not limit the form of video on demand), it contains planar target information, such as a classical realistic oil painting picture used by the teacher to interpret techniques, features, etc. Although on the teacher's computer screen, this oil painting may be a very clear image file with a size of 10 megabytes, after screen recording, the information of the oil painting in the screen recording video (defined as the target in the video in the present disclosure), that is, the content information of the target, far fails to meet the minimum standards required for professional learning after becoming a video. And if the student happens to watch the video on a mobile phone, the dissemination effect that the video should achieve is even far lower than the target dissemination effect. Of course, in the above process, the teaching video can be a live broadcast. Then, the oil painting described by the teacher in the live broadcast video also fails to meet the professional requirements. Thus, the existing technical methods of video restrict the dissemination of knowledge and video content. For example, in the above-mentioned situation where the teacher explains the details of the oil painting, because the terminal cannot see the details clearly, the teacher's explanation of techniques and details cannot be clearly seen by the students, and it is impossible for the students to understand the techniques.
[0047] For another example, a doctor shares his medical experience using a video and includes information pictures of nuclear magnetic resonance imaging to illustrate how to diagnose diseases. However, after these nuclear magnetic resonance pictures become videos, they have no value for medical learning and judgment. Therefore, the viewers cannot form a logical closed loop of cognition and recognition based on the information of the nuclear magnetic resonance pictures in the video (the information of the target in the video). For example, the nodules in the nuclear magnetic resonance pictures cannot be identified at all in the video. Thus, the current form of video dissemination substantially hinders the effective dissemination of information in many aspects.
[0048] Therefore, to solve these problems, A01, that is, the video editing or recording personnel, the generating personnel (defined as video generating personnel in this disclosure), must first have a video file V01. Assume it is a teacher's screen recording video (or a video recorded by a doctor, this disclosure is only for illustration with examples). Another file is F01, that is, the information corresponding to the target in the video (in this disclosure, the files, digital links, etc. corresponding to the target are all defined as the preset information of the target). And the information can be files, such as medical picture files, clear photos of oil paintings, vector files of designs, etc. For example, the planar target - classical oil painting, the picture information of nuclear magnetic resonance imaging in the above example (of course, there can be multiple targets in the video, and each target can have corresponding different forms of files, resource link chains, or even video streams). Each target has its corresponding preset information, and the corresponding information includes both file form and digital link form.
[0049] To achieve good results, a functional module D01, the target data generation module, is required. This module can be included in video editing software or be the terminal software of a video playback platform, becoming a functional module; of course, it can also be a function of a video application software. When A01, that is, the video generating personnel, imports the video into the D01 module, such as opening the video with a program containing the D01 module, A01 can use the D02 target bounding module or function. The video generating personnel set and select to bound the target in the display interface of the imported video. For example, in the above example of the oil painting, A01 uses geometric shapes, such as circles, ellipses, rectangles, polygons, etc. (usually rectangles) to bound the target, enclosing the target within the closed geometric shape; of course, it can also be an automatic method. For example, A01 clicks or triggers the target in the video image (through a human-computer interaction method, so that the program knows that an image whole containing the triggered position in the video is the target, usually using image segmentation to bound the target), that is, the program segments based on the click position and the image containing the click position. Under the setting of A01 personnel, the function of bounding the target is set, and a marked target box awaits the user's confirmation or fine adjustment; therefore, in target bounding, the target can be bounded by a combination of manual and artificial intelligence or image processing methods.
[0050] For situations where the content of the video does not change much, such as video recording based on the layout of the various contents on the screen, the information does not change much over a period of time, so manual methods are completely sufficient; for live recorded videos, such as an economics teacher giving an economics lecture, the chart information in the video background (usually projected information) may move, change size, or become out of focus as the teacher moves (the teacher is the focus of the video), but in fact the chart in the background information may be the information that the viewer pays more attention to (detailed information), which will be blurred and out of focus when the camera is shooting. Therefore, video editors and producers select the chart display area of this background as the target area. Since it always changes in the video image, machine target detection can be used to detect and obtain the target's position data in the image, and mark the target box according to the data.
[0051] Video object detection technology has always been a key technical direction in the field of AI and computer vision in the industry. It is also a field where multiple implementation methods have emerged, but it is also a field that is always looking for the best effect. The present disclosure does not limit the specific method. On the contrary, as shown in the attached letter of claim, the technical method described in the letter of claim is a specific embodiment of the present disclosure.
[0052] In addition, if a video focuses on expressing a product, such as a video containing advertisements, although the product is placed at the core of the video, the texture of many commodities cannot be expressed in the video, especially videos without professional video shooting and processing, which cannot arouse consumers' desire to pay attention, so video viewers are indifferent; in the current live video broadcasts, one of the reasons why a large number of high-quality commodities have poor sales is the live broadcast method, which makes high-quality commodities lose their texture, and products that need to show details cannot show details through live broadcasts, thus hindering consumers' willingness to purchase.
[0053] In a video, three-dimensional objects such as commodities usually move in the image. For example, an advertisement for a pulley is usually expressed in a moving video, so it is necessary to use the target detection module D03. D03 detects and frames the target in the video in real time based on the target selected by D02 and the timeline of the video, and forms a target frame and a marking frame (reducing the complexity of manual work).
[0054] The technology of D03 has been developed for many years, and there are already various technical paths, including artificial intelligence, computer vision, or a combination of both. The specific implementation path is not limited in this disclosure. However, it should be noted that the method of manually selecting the target, that is, the D02 method, is adopted in this application because traditional target detection is to find the target in the entire video, while the problem to be solved in this disclosure is to target specific objects rather than finding various moving objects in the video and ignoring the background in the traditional target detection method, that is, the target selected by D02. Along the video timeline, the target box is calibrated. Therefore, during the change of the video, the size and position of the target box change with the change of the video, and the size is also adjusted. Therefore, based on the video timeline, the position data of the target can be obtained (and based on the position data of the target in the video image, a rectangular box is usually returned to label the target), which is the target data generation that the D01 functional module needs to achieve. When generating data, there can be multiple targets in a video, and each target can correspond to a target number. There can be multiple targets and their uniquely corresponding numbers in a video, and each target will have its own generated data (the position data of the target is required in this disclosure as the basis for judging whether the target is clicked later).
[0055] If there is no D03, personnel A01 can handle it manually. For example, based on video frames, I-frames, and the timeline, manually circle the target range and generate data. After adopting D03, personnel A01 can rely on an automated method. For example, in the case of live broadcasts, when manual operation cannot keep up with the image changes, D03 can be adopted, and artificial intelligence or computer vision technology can be used to automatically generate target data.
[0056] In a video, there can be multiple targets, which are used in the D01 process. For example, in a video of an exhibition hall, there are multiple targets (such as multiple oil paintings and sculptures) in the video corresponding to a certain duration of a single shot. Therefore, the D02 selection and the D03 automatic target detection process can be adopted to generate their respective target data. For a documentary on knowledge topics, there may be many targets, such as sculptures, murals, historical documents, cultural relics, etc. When presented in video form, they cannot be detailed one by one and cannot be clearly shown (such as the content in the document). Therefore, according to the D01 process, target data is generated in the video. When the viewer of the video clicks on a target of their interest at the user end, such as a document in the video; since the user-end program determines the clicked position and whether this position is within the specific target range of the current video position, it is then determined that the user has clicked on the target. Therefore, at the user terminal side, the pre-set information corresponding to the target (such as the content of a document displayed in PDF format) is read and displayed on the user-end interface. The user at the user end can view the content of the document instead of just watching the video. Of course, for the oil painting in the previous example, the user can zoom in to view the specific details (assuming the user clicks on the oil painting in the video).
[0057] In the foregoing examples, video files are used as input and import for illustration. Below, the scenario of on-site recording will be used for illustration, such as a live video format.
[0058] When the video recording personnel use a camera recording device such as the V02 video acquisition module for acquisition, the digital video stream can be input into the D01 module. At this time, the A01 personnel, according to the target to be marked, use the D02 function to select and circle the target. Then, under the D03 function, as the lens of the V02 acquisition device changes over time, the target is detected, moved, and changed to generate target data, and the framed area (target box) is determined. At this time, if the terminal watching the video also receives the generated target data, after clicking on the target, the terminal program will, according to the clicked position, within the data range of the target, then determine the clicked target. The terminal can read the preset information of the target and display it on the user terminal, so that the user can clearly understand the clear information of the target of interest, or information different from that shown in the video. For example, for a book in the video, after being clicked on the user side, the PDF file page of the book is displayed.
[0059] The V02 video acquisition module usually includes an optical imaging function module, a digital imaging module, etc., and then converts the optical signal into a digital signal, and then imports it into D01 in the form of video encoding (common acquisition modules such as mobile phones, video cameras, etc. encode the optical signal into a video stream).
[0060] When the A01, video editing personnel, or recording and generating personnel are operating in step D01, that is, generating target data, they should prepare the corresponding target files (preset information of the target) in advance. For example, the graphs and tables used by an economics teacher, or the classical realistic oil painting photos used by an artist. So it can be a clear picture file, a vector file, or a digitally connected file. Then, on the video playback side, after the user clicks on the target, the terminal program or the background program determines that the user has clicked on the target, and after determining the clicked target, displays the preset information of these targets. The end user on the terminal side can magnify the details of this information as if viewing it with a magnifying glass on-site, and then achieve the on-site perception of the real object.
[0061] For three-dimensional objects such as commodities, the preset information of the object can be the corresponding photo of the object, a vector diagram, a design drawing, or of course, a close-up of a video. It can even include a video with a close-up captured by a dedicated imaging device, such as a live jewelry sale, a sculpture, etc. For network shooting equipment and lighting settings, the object moves regularly (such as rotating under the light). If a viewer clicks on the object in the host's video, they will enter the video resource (URL resource link) for the special display of the object. If it is an online video, essentially, another video stream is called and displayed on the terminal. When the terminal plays this video stream, the detailed video enhances the communication effect of the original video. That is, the first video is for dissemination, and the second video is played when the user who has received the dissemination becomes interested in the object, clicks on it, and then gains a better understanding and feeling of the object. Of course, on the user's terminal side, the two videos can be displayed in a picture-in-picture format.
[0062] In the F01 video, the files of the object can be one or multiple (the preset information of one object can be multiple. For example, in a video of a stack of books, which are different volumes of a set of books. After clicking on the object, the PDFs of each volume are read by the terminal, so that the person who clicks can preview the PDFs, such as PDFs, and decide whether to buy this set of books) to meet the goals of the A01 personnel for the video dissemination effect and the demands of the video viewers, that is, the terminal users U01, for the video, especially for obtaining the object information in the video. Of course, some frames or a certain segment of a video can contain multiple objects, and different segments of a video can also contain different objects. Each object can correspond to one file or multiple files (the respective preset information of each object). After the object data is generated, according to the timeline on the user's terminal, when the user clicks and the terminal determines that the object has been clicked, the user's terminal reads and obtains the preset information corresponding to the object, that is, a more detailed information file / or the user's terminal actively reads it. And all these preset information are generated and imported or uploaded to the video dissemination module.
[0063] The object data generation module can of course also form a new format of video file (the current video coding standard cannot apply the video so flexibly) from the F01 such files in a certain coding form, and then transmit and import it as a complete file or data stream to T01, that is, the video dissemination module.
[0064] Since today's video standards do not support the above flexible method, in the traditional case, it is still common to transfer the files of F01 to the T01 video dissemination module for video users to retrieve and view these files after clicking on the target; of course, for URL resource links (digital links), in the T01 module, the video, the target, and the URL can be associated. Because the video has a unique encoding, the target also has a unique encoding in the video, and the target is uniquely associated with the URL or the file (each target has its own preset information), so in a video, when the user clicks on a certain target displayed on the terminal, the terminal program determines that the target has been triggered based on the click position within the range of the target. Therefore, the terminal program reads the file or URL (preset information) associated with the target and displays it. And because target detection is adopted, that is, the circled targets in the current video all have position range data, so it can be determined that the click position is within the current range of the target in the video (usually within the target box), and it can be determined that the target has been triggered.
[0065] In this application, there is no limitation on the specific format in which the file corresponding to the target forms a new regular video file or remains in a separated state, that is, the preset information of the video and the target is separated, and the terminal retrieves and reads the preset information according to the user's request.
[0066] In addition, the data generated after the target data is generated can be stored in the video file in a set format to form a new video file, or the data can be transmitted to the database corresponding to the T01 module (in the form of a database, such as a real-time database during live broadcasts) or in the form of a file (such as a non-live video). When the terminal program for viewing the video reads the video file, it also reads the generated data of the target at the same time (so as long as based on the ID of the video, when the terminal reads the video and obtains the generated data of the target, it is a specific embodiment of the present disclosure).
[0067] The T01 module usually exists in the form of a video file server or a streaming media server. The system runs on the server, and the user terminal accesses the network service of the server to obtain the link of the video file; in the present disclosure, the T01 module may also include a database. In the case where the video file, the target file, and the generated target data are separated, the user terminal not only reads the video file, but also reads the data generated by D01 associated with the video file and the file or preset information corresponding to the target. Of course, for unstructured data, an unstructured database can be used, or a combination of a structured database and a file service form. In this application, there is no limitation on the specific implementation manner of the T01 module, but any manner consistent with the present disclosure is a specific embodiment of the present disclosure.
[0068] In addition, for real-time video stream playback such as live streaming and broadcasting, such as the video captured by V02, the data generated by the D01 part is associated with the unique encoding of the stream (such as the video ID or the video serial number specially generated based on the video, so as to ensure the uniqueness of the video encoding). When the user terminal calls the stream, the data can be extracted from the database on T01.
[0069] The T01 module is usually interconnected with external user terminals through the network and usually includes servers, storage, databases, etc., and is accessed by end users to obtain video files, video streams or related data and the files and information corresponding to the target (preset information of the target); in the T01 module, usually the video will be uniquely encoded (video ID). When the video playback terminal plays the video during the above separation, based on the unique encoding, the generated data associated with the unique encoding or the preset information of the target of the video or the generated data of the target can be read.
[0070] For the video user side, usually users of certain smart TVs or smart terminals use the programs in the user terminal to watch videos. Specifically, they read the information of the video from the T01 video dissemination module and use file stream (including slices) or video stream technology to display and play the video file or video stream on the display device / module of the local terminal for the user to watch.
[0071] Since U02 can read the target generated data and the file corresponding to the target (preset information of the target) from T01, when the video is played on the terminal, based on the video and the corresponding target generated data, the target will also be displayed during display. For example, the oil painting, table, three-dimensional target, etc. in the above example. If the played video contains the data of the target, the terminal can display the target box (or not display it). For example, if the user understands this method or is prompted to click or trigger the area or target within the target box, the user terminal program calls the F01 file corresponding to the T01 video, that is, the preset information of the triggered target, and reads and displays the corresponding file on the terminal according to the clicked target (the prompting methods include voice, subtitles in the video or other explicit or implicit prompts).
[0072] For example, in the video of the screen exam, there is an oil painting, and the oil painting is circled by the generator as the target to generate the target data. When viewed on the user side, when the terminal program reads the video, it also reads the associated target data. Then, when the video is displayed, according to the time line of the data and the position data of the target, and in accordance with the size of the displayed image, the target will be displayed on the terminal and the target box will be marked (of course, it can also not be marked). When the user clicks on the area or the target within the target box, the information of the target will be opened and displayed on the display interface of the user terminal.
[0073] For example, if the target is a design drawing and it is transmitted via video, such as when a designer describes it through a video, the details of the design drawing are simply not visible (e.g., a large building). However, when implemented through the above method, when the end user is listening to the explanation, they can click on the target, the design drawing area, in the designer's video on the user terminal. Then, according to the above method, the vector design drawing can be opened. Subsequently, the user can zoom in arbitrarily, so as to have a better grasp of the details, and the designer's explanation will have a better effect.
[0074] As a special case, if the video is a live broadcast, and at this time the person in the live broadcast hopes to adjust the target, such as zooming in or focusing on a certain point. When the person in the live broadcast operates, the magnified data and the display position of the file can be read proportionally by each viewing terminal. Subsequently, what is displayed on the user side is not only a clear file but also the content that the person in the live broadcast is focusing on and describing. Thus, a higher communication effect is achieved, that is, the adjustment target of the video generation side personnel. And the preset information of the target opened by the video user side is also magnified as the generation side zooms in and moves with the movement of the key points. Consequently, the communication efficiency is greatly improved in terms of breadth and effect. In this case, a specific embodiment can be that the video generation personnel, on the generation side, click on the target. Then, the terminal interface of the user watching the video reads the preset information of the target and then displays it on the screen. That is, on the generation side, on the one hand, the target can be circled. Also, at a specific time period, the personnel on the video generation side can click on the target in the live broadcast. Since the target on the generation side is circled and target data is generated, and the trigger or click position is within the range of the currently generated data of a certain target (usually a rectangle), T01 can send data to each viewing terminal. This data drives the programs of each viewing terminal to execute reading the preset information of the specific triggered target and display it proportionally to the generation side. In this way, the current video live broadcast personnel can, according to the situation of the live broadcast, control the remote users in a clearer manner to view the preset information of the target they are describing, thereby improving their description effect (in implementation, usually the size, display ratio, and display position data of the target content on the generation side are received by the terminals that have clicked and displayed the preset information of the target through T01. These terminals display the information of the target corresponding one-to-one according to these data. For example, if the preset information of the target being watched on the generation side is, assume, the second paragraph on page 5 of a PDF file, then the user side will also automatically turn to page 5, second paragraph because it receives the data from the generation side).
[0075] Of course, there is another situation. That is, after the user reads and plays a video, the user of the video is passively shown the file information corresponding to the target, that is, the player and editor of the video are already very clear about the defects of traditional video dissemination. Subsequently, after the target appears, rules (preset instructions) are set. That is, when the terminal playback system reads and plays the video, according to the rules, the terminal automatically reads the file corresponding to the target and displays it. The user who plays the video passively views the information preset by the target. In addition, the way in which the user views passively can also be set. According to the display path designed by the A01 personnel, such as from macroscopic to microscopic, from the boundary to the core, when the user-side software sees the target period, the preset information is displayed. Therefore, this method can quickly improve the effect of video dissemination (this implementation effect is equivalent to the generation side synchronizing with the user side, and the data of the preset information controlled by the user side is received by the viewing-side user and the preset information is displayed according to the control of the remote end).
[0076] Of course, for video dissemination, the viewers of the video only purposefully select the information of the target content. So there is the following form. There is no mark indicating the target such as a target box on the video viewed by the user side. However, the voice information and prompt information in the video prompt the viewer to click on the target area. After clicking on the target area, the preset information associated with the clicked target can be seen.
[0077] For the program of the user terminal, since the target generation data associated with the video can also be read when reading the video, based on the target data corresponding to the clicked frame (it can also be the data corresponding to the time of this frame), if the click position is within the range included in the target data, it means that the user wants to view the detailed information of the target. Subsequently, the terminal displays the preset information corresponding to the target, etc.
[0078] That is, after a video goes through the above steps and methods, in the user-side program, the user can click on the target according to the prompt, and then can see the information that the video itself cannot express in detail. Of course, in a specific situation (special case), when the video generation personnel, such as a teacher, talk about the content of the target and trigger and click on the target on the generation or live side, the program judges the clicked target and notifies the user programs of each viewer through T01 to read the preset information of the clicked target and display it on each end.
[0079] The above description is the way when the target corresponds to a file. If the target corresponds to a resource link such as a web page, the logic is the same. That is, the target in the video corresponds to a web page or a file corresponding to a special file URL when the data is generated. After the user clicks on the target, the user terminal APP accesses the web page corresponding to the URL. There can also be descriptions, exchanges, evaluations, etc. of the target on the web page. This can play a role in directly attracting traffic and selling for commodities.
[0080] For example, in a movie or a TV series, if a product (the target) is implanted, although the director will try to make it eye-catching (through filming techniques) when implanting the product (the target) in the video, it is still just a small area shown in the video. However, if the user clicks on the product (the target), the webpage or the page in the terminal program directly associated with the target will open, presenting the product details and promotions, thus achieving the dual realization of product advertising and its effects.
[0081] Therefore, based on the above, the preset information associated with the target, F01, can be a file corresponding to the target, such as a video, a picture, a vector, a table, etc., or a network link, or a page of the terminal software (the preset information corresponding to the target).
[0082] Next, taking Figure 2 as an example, the embodiments will be used to illustrate. Figure 2 In, s01 is an image of the nth frame in a video. In the image, there is a person S110 holding an item S103, and behind is the projected information S102. Assume that S110 is an economist giving a lecture on the economy, and S102 is a projected data table. We know that after this video is transmitted to the terminal, due to encoding and terminal size, information such as that in S102 cannot be clearly seen. Moreover, this S110 also endorses products such as S103 during his speech, such as a microphone of a certain brand or a microphone with the logo of the organizing committee on it.
[0083] We know that a video is composed of consecutive images such as Figure 2 . If the camera position changes, S110 will move at different times, and the S103 in his hand will also move accordingly. S103 is a three-dimensional target, while the S102 behind is a two-dimensional target. When this video is imported and needs to be processed by the D01 function, video editing, creation, or recording personnel will use the target frame 122 to frame or enclose the target S102, and S123 is used to enclose the target S103.
[0084] Since each frame in the video has a corresponding frame sequence number or corresponding time, for each target box at the frame value or time value, under the corresponding target encoding (such as S102 for target 001, S103 for target 002), each enclosed target box has corresponding position parameters (the generated data of the aforementioned target). Corresponding to the video, there are corresponding X and Y values. For example, the S122 rectangle has the position values of four corners. As the video progresses, the camera shooting angle may change and the lens may stretch, and the values of S122 will all change. However, since D01 generates the data of the target, that is, corresponding to the time or frame value of the video, the target box of target 001 will generate corresponding values; by the same token, S103 will also generate the position values of its target box corresponding to the video frame or video time; taking the rectangle as an example, it has four corners of up, down, left, and right, and each corner has a corresponding value. For example, in Figure 2 for the 122 rectangle, the four corner values are (x1 = 100, y1 = 40), (X2 = 500, y2 = 40), (x3 = 100, y3 = 800), (x4 = 500, y4 = 800) in the order of up and down, left and right, that is, in the way of real pixel values (absolute values), or it can also be the ratio value to the video pixels. For example, when the video is 1080P, the vertical direction is 1080 pixels and the horizontal direction is 1920. When using the ratio relationship, it is (x1 / 1080, y1 / 1920); when the original video is transcoded to 720P, if in absolute values, the correct position of the target box cannot be corresponded, while in relative positions such as ratio values, as long as the image ratio relationship remains unchanged during the transcoding of the video, on the user side, relative values can be used to determine the position of the target in the transcoded video. Then, after the target position is clicked, the clicked target can be correctly recognized and matched (such as the conversion relationship x1 / 1920 * 1080,y1 / 1080 * 720). Of course, relative position is a normal means in video processing. In this disclosure, it is only to remind that in the transcoding scenario, it is necessary to consider that the data generation of the target can use absolute values or relative values, and relative values are more convenient for various dissemination qualities to be used in the transcoding environment (the original video is transcoded).
[0085] When a user watching the video views the video on a very small mobile phone screen, the video is actually scaled down proportionally according to the screen size during display. Therefore, using ratio values, that is, relative values, will reduce the computational complexity of conversion when clicking on the target. Otherwise, it is also necessary to know the format during the production of the original video and then calculate according to the transcoded video that may be changed on the terminal side.
[0086] In Figure 3 assuming that according to the above method, the video software in the user terminal is playing this video, and the S211 part is corresponding to Figure 2The content, as described above, the target can include a target box and be displayed to the end user, or it can be without a target box and be normally displayed to the viewers of the video. Displaying the target box has the drawback of interfering with the normal viewing experience of the video. For example, the viewer is just killing time and has no interest in any target. So for them, the target box is an indication of a degraded video effect; in Figure 3 such as a handheld target without a target box identifier; of course, in the future when the terminal computing power is strong, a three-dimensional graphic can be formed on the target image, and then the human-machine interaction method is more ideal, which also indicates that the target has been set and can be triggered; of course, the target box can also be shown only during pauses; but if the computing power and image processing ability are very powerful, it can also be shown in a three-dimensional and protruding manner during real-time playback. For example, in current image games, the set target is presented in a more user-friendly way than the target box throughout the video playback, and then the user can trigger the target to obtain the preset information of the target or execute the preset next instruction.
[0087] However, as described above, Figure 1 the little person S110 in Figure 2 can tell (prompt) the viewer in words, and then the video viewer is prompted to click on the three-dimensional target in his hand or click on the chart, and more detailed preset information will be seen. Or there are subtitles on the video interface to prompt the user, such as please click
[0088] the chart on the right in
[0089] to see clear chart information. Figure 2 And Figure 3, from the terminal examples on the creator or editor side and the video viewer side, we can know that through the above method, the information that may have been impossible to see clearly before can be displayed very clearly in this way, and one of the defects of the original video transmission, that is, no matter what kind of target information is uniformly encoded, the problem of reduced clarity has been solved.
[0090] if Figure 2 and Figure 3 The current live broadcast is described in the following example. The item in the little man's hand is the item he is promoting. Under normal circumstances, the user cannot see it clearly and can only imagine the details of the product. However, through the method disclosed in the present invention, the user clicks Figure 3 In the target corresponding to S103 in FIG. 2 , the user terminal displays a clear photo or design drawing of S103 or a live video or recorded video rotating under good light and photographic equipment facing S103, and then the user at the video playback end can carefully watch the target product.
[0091] Of course, video users in this situation can also interact with the video producer or service team. For example, by overlaying a social network or instant messaging system in the present disclosure, when users enter the details for viewing, the producer or marketer will also clearly know that people of interest are already paying attention to the product or information, and the service team can then provide more information and interact with the video users. Of course, the service team can also be an AI robot, such as an AI robot based on a large model, which can provide more knowledge and information to individuals, thereby increasing the efficiency of communication (of course, these systems and terminal software need to be interconnected with the corresponding service system).
[0092] A specific example is that a video user, after clicking on the target, enters the display Figure 2 When the user sees the table interface, the system will notify the AI robot, and the AI robot will explain to the user in detail the purpose of the table, data formation, conclusion, etc., and these statements can be the information that the previous narrator trained the AI robot with, and then the AI robot deeply influences the purpose of the video audience, so that the viewer can understand the information more deeply than in the video.
[0093] And if Figure 2 The content in the advertisement shows the comparison of product parameters, etc. When users are viewing the advertisement, the product guide, such as an AI robot or a real person, can persuade the users to read the details in a concise manner, and then further complete the purchase.
[0094] In the current situation, on the video editing or live streaming side, the method of using artificial intelligence or computer vision technology, that is, the object detection technology or function of D03, once the object has been set, the program can track the object along with the video timeline and generate object data; while on the user side, since the frame number or the time value of the video has been determined when clicking on the video object, combined with the generated object data read, by clicking on the position displayed in the video, the specific object that the viewer wants to see can be calculated; because the adopted method has no difference from normal video playback in the case of not displaying the object frame, it does not affect the user's viewing, and for interested users, they can deeply understand the information of the object content.
[0095] In addition, the terminal program of the present disclosure is also applicable to terminals such as XR (VR, AR, MR), etc. When the user is watching a video and clicks on an object through a human-computer interaction method, detailed information of the object will also be displayed on these terminals. At the same time, it can also be served by customer service, and messages, discussions, etc. from other people can also be seen. Thus, regardless of the terminal, it can make unrelated video viewers become active with each other without affecting video viewing. And corresponding to the bullet screen, there will be no defect that the bullet screen affects video viewing and is not conducive to users' direct and relatively in-depth expressions.
[0096] In the above description, by clicking and triggering an object in the video to obtain the information of the object. In human-computer interaction, it can be clicked with a mouse, or clicked with a finger on the touch screen. Of course, it can also be other agreed methods, such as double-clicking on the object. In this application, it is not limited to specific triggering methods such as clicking or double-clicking (triggering methods). As long as the user uses an agreed method to trigger an object in the video, judge, and read the preset information corresponding to the object, it is a specific implementation example of the method described in the present disclosure.
[0097] The above method usually needs to be implemented on both the video source side and the video playback side. On the video source side, object data and files corresponding to the object need to be generated, and on the user side, the generated object data can be obtained when reading the video. After triggering the object and the system makes a judgment, the terminal program displays the file corresponding to the object, and thus the effect of clearly displaying the object is achieved.
[0098] The above method is actually not reasonable for some objects. For example, the content of an oil painting does not change, and using the object detection method, the efficiency is relatively low and unnecessary data is generated; while for a PPT that may change (the person presenting the PPT is turning the page), there is also a video played in the video, and this video is set as the object, so the content of the object in the video is changing, so the object detection method can be adopted.
[0099] In the above manner, after the user clicks the screen due to triggering a target, the judgment can also occur in the background. For example, after the user terminal clicks a target in a video, the position data of the click and the information of the clicked time or frame are fed back to the server. The server determines whether the clicked area is included in a specific target box based on the video, the frame or time information, and the click information, and then determines the target. Then, the information is fed back to the terminal where the user clicks the target. Then, the terminal reads the preset information of the target, that is, the work of judging the clicked target is given to the background system. In the present disclosure, it is not limited that each client software obtains the data generated by the target. The client program judges the clicked target or the client clicks, and the server judges the clicked target. However, both methods can achieve the problem to be solved by the present disclosure, only in different ways.
[0100] The above-described embodiment undoubtedly solves the problem that the target information in the original video cannot be effectively displayed in the video. However, there is a problem, that is, target data will be generated. For example, if a video is 120 minutes long and assuming 30 frames per second, a lot of target data (computing resources) may be generated. Next, we provide the second method adopted to achieve the same purpose. In the second method, it relies on the fast processing ability of the terminal, specifically such as Figure 4, video editors or recorders (video generators) use the program on the U10 video editing or generation terminal to import the digital video stream collected by the video capture module of video file V01 or V02. In addition, the preset information corresponding to the target in the video is also imported, such as files and digital links. Corresponding to the target, there is also the picture F02 of the target. Since F02 is used to extract the image features of the target and does not require the preset information (which may be various types of files or digital links for display) consistent with the requirements of F01, and F02 is only a picture for comparison. Of course, if a target has different angles in the video, but each angle is different or has a large difference, a target can have multiple corresponding feature pictures. After the user clicks on the target in the video, it is compared with the image information of the target or area at the clicked position, so as to determine the clicked target; and F11 is the feature extraction of the target image. This part can be placed in the function of the T01 module or in the U10 module. The present disclosure does not limit this; through U10, the video is finally imported into T01, including the files corresponding to the targets in the video (included in the preset information), including the features of the targets; the user accesses T01 with the terminal program to obtain the video information and play it. Since the video coding is unique, the feature information of the corresponding target of the video (which is also the uniquely coded information) can of course be extracted; the user plays and watches the video. When seeing the target in the video, the user clicks and triggers the target (substantially triggering the area covered by the target). At this time, the terminal software extracts the image information and features of the position (such as x value, y value) triggered and clicked in the video, and based on the comparison with the feature information of one or more targets read from T01, then judges the target corresponding to the clicked position. If there is a matching target, it is judged that the target has been clicked, and the terminal program displays the preset information of the target triggered in the video. Of course, it is also possible to find the matching area for the features of the target in the current video frame. If the extracted x and y values are points in the matching area of a certain target, it is judged that the target is the target triggered by the end user, and then the preset information of the target is displayed on the user terminal screen.
[0101] The above method based on feature comparison saves unnecessary calculations compared with the method of object detection. This is because the previous method 1 frames the target and generates target data in theory for each frame, which requires a certain amount of computing power. However, for the method based on feature comparison, it only starts to execute the comparison after the target is triggered, and then judges the triggered target. Therefore, the amount of calculation is relatively small, but the real-time performance is worse than that of the object detection method.
[0102] After the user uses a finger to trigger a target in the video through a triggering method such as clicking or double-clicking, the video-playing program on the user side obtains the triggered position of the touch screen on the one hand, and then calculates the clicked position in the video image according to the relative relationship between the layout shown in the video and the screen (since the display screen resolution is different from the video resolution, proportional conversion is required, and conversion based on the layout size ratio is also necessary). Then, the position of the clicked position in the video frame is calculated. The image segmentation method can be used to segment the target in the clicked area as a whole, extract the image information of the clicked and segmented target, and compare it with the features of the pictures of multiple pre-input targets (such as SURF, SIFT). Subsequently, it is determined which specific one of the multiple pre-set targets the clicked target is, and then, according to the pre-set information of the target, it is displayed on the program of the terminal. For the specific pre-set information, these information can be, as described above, the files, video streams, etc. defined by the video producer for the target, such as a high-definition picture of Mona Lisa, or it can be a PPT, PDF corresponding to the target, or a real-time video stream corresponding to the target, design vector graphics, etc.; for example, a user clicks on a red envelope in the hand of the protagonist (person) in the video (such as Figure 2If the target held by the person in S103 is a red envelope, the program of the terminal extracts the image of the red envelope at the target position clicked by the user (such as segmenting or extracting the image of this area), and then compares it with the features extracted from multiple pre-input target images at the terminal program end or on the background server, and then compares the target at the clicked position, for example, determines that the target at the clicked position is a red envelope; of course, a video can contain different types of red envelopes, such as red envelopes with the character "Fu" or "Cai", and then according to the corresponding identified and compared targets, and display according to the information preset for different targets. For example, in a video, a person holds a red envelope (the envelope cover contains the character "Cai") and says a paragraph such as a New Year's greeting. Then it is prompted that clicking on the red envelope will give everyone a share. When a person watching the video clicks on the red envelope held by the person in the video and it is identified as a red envelope with the character "Cai", the terminal program displays the red envelope (it can be a photo file of the preset information of this red envelope), or the amount of the red envelope (the preset information of this target), and automatically or by clicking a control, the user manually transfers the cash of this amount into the account of the clicker (click to confirm receiving the red envelope on the video terminal program); of course, if the person in the video is a live e-commerce seller and holds a red envelope with the character "Fu", and prompts the viewers watching the video to click on this red envelope, then similarly, the user program extracts the image and features of the target at the clicked position, and generates a comparison with the feature image of the target pre-input by the user at the terminal program or the background server, and then determines and compares it to be a red envelope with the character "Fu"; if the information preset by the video generator corresponding to this red envelope is a discount card for goods, the user video interface displays the discount card and the discount amount, and the user who clicks on this red envelope manually clicks to confirm receipt or the discount card is automatically transferred to the electronic card package. In this embodiment, instead of displaying the preset information, the next step instruction is completed, and the cash or discount card is automatically or manually transferred to the card package or wallet.
[0103] In the above embodiment, after clicking or triggering the target, the terminal program segments or extracts the target image or extracts the image at the clicked position (including the target or part of the target) corresponding to the video image, and compares it with the features of the picture of the target pre-input by the generator, and after judging the target, then displays the preset information corresponding to the preset target set by the video generator, and executes the next program or instruction, such as automatically transferring the cash to the digital wallet after 5 seconds or transferring the discount card to the card package, or to meet the requirements of certain standard processes, a human-computer interaction such as a button with a confirmation meaning appears on the screen of the terminal, and then the user needs to click to transfer to the electronic account or card package, such as transferring the money in the red envelope to the digital wallet of the clicker, or transferring the discount card to the digital card package of the clicker.
[0104] In the above disclosure, in addition to the user being able to click on the target in the video to clearly display the content preset for certain targets, for the content that cannot be shown in the video to make up for the deficiencies of the video, it is also possible to prompt the viewer of the video to click on specific targets in the video. The front-end program obtains the clicked position, and based on the layout, pixels, and screen pixels of the video display, extracts the target in the video frame, and compares the features with the features of the target images pre-entered at the front-end or back-end to determine the clicked target, and then displays the information preset for the target. Even according to the tasks to be completed by the preset target, execute programs or instructions, such as transferring the cash in the red envelope to the digital wallet of the clicker, or transferring the discount card to the digital card pack of the clicker. Of course, clicking on the goods held by the person in the video can also be set to be displayed and manually or automatically transferred to the shopping cart.
[0105] Of course, segmenting and extracting the target image according to the position of the click position in the video image is a relatively computationally intensive method (although there are ready-made algorithms and functions). A simpler approach is to extract the image according to the click position in accordance with specific dimensions, such as setting a width of 150 pixels and a height of 90 pixels, or extracting the image as 90 pixels * 90 pixels. Then extract the features and compare them with the features of the pictures of one or more pre-entered targets, and then determine the clicked target. Specifically, taking the clicked position in the video as the center, for example, when it is 150x90 pixels, assuming the clicked position is x = 300, y = 200, then when taking a rectangle as an example, the two diagonals of the rectangle are (300 - 150 / 2, 200 - 90 / 2), (300 + 150 / 2, 200 + 90 / 2); in the image processing function, usually with these two diagonals, the image in the rectangle can be extracted; then extract the features of the image in the region, match them with the features of each target pre-entered by the video generator, calculate the matched target, and display the information preset for the target, or display the preset information and execute the program, such as transferring the cash in the red envelope to the electronic wallet of the clicker in the terminal program (the above numbers are only for illustration, and in actual applications, different sizes of the extraction area image can be set, including other geometric shapes such as circles, with the contact point as the center of the circle and a radius set such as 90 pixels to extract the image. The embodiments are only for illustration).
[0106] In the present disclosure, multiple methods are adopted to determine the clicked target. Since the method paths for determining the clicked target in a video are not unique, and as prior art, this application will not elaborate on each one. However, as required by the claims in the present disclosure, any method or technology that extracts the target in the clicked area, compares it with the features of the pictures of one or more targets pre-input by the video generator, determines the matching target, and displays the preset information of the matching target, or reads the preset information and executes the next program, or displays a human-computer interaction interface for the user to confirm or select to execute the next program, is the technology consistent with the present disclosure.
[0107] In the above embodiment, to execute the next program, such as the transfer of red packet cash, usually requires the user to click to confirm. Therefore, after the user clicks on the red packet and it is recognized that the user has clicked on the red packet, the amount of cash preset by the video generator is then displayed. The user who clicks on the red packet confirms on the terminal program interface that they want to receive this cash. After the click, the cash is transferred. In a video of selling goods, in addition to clearly seeing the preset image by clicking on the target in the video as described in the present disclosure above, it can also be that after clicking on the commodity, the clicked commodity is directly added to the shopper's shopping basket. That is, after recognition and matching, the next program or instruction is set to add the target goods to the shopping basket. This is the same as the method in the red packet click embodiment, except that the shopping basket has a different manifestation form from the e-wallet and the e-card package.
[0108] In an interactive drama, the technology of the present disclosure is also used for interactions within the video. For example, in the plot, a detective places a crime scene photo on the table and hopes that everyone can help analyze it (including the characters in the drama and the audience). At this time, the audience watching the drama clicks on the photo in the video. In the same way as described above, the photo is clearly displayed on the audience's terminal program (the digital photo is the preset information, and the photo in the drama video is the target). Then the audience can analyze the photo, such as discussing in the interactive group of the drama, and then participate in the crime-solving interaction based on the photo. Then, in the next episode of the drama, according to the interactive data, the director modifies the plot and shoots based on the interactive information, thus forming a deeply interactive film and television work.
[0109] Similarly, in the example of film and television dramas, a character in a film or television drama picks up a book that contains the decryption of a password. Usually, the decryption in such film and television dramas has nothing to do with the audience. However, according to the present disclosure, assuming that the book is placed on the table in the video and the video user clicks on the book, after comparing with the above-mentioned technology, it is known that the user has clicked on the target, and then the preset information - the content of the book will be presented on the user side. The user can search for and crack the password in the book in the form of a password, thus turning the one-way infusion of the film and television drama into a two-way deep interaction. In the community of the film and television drama, the audience can interact and cooperate to crack the password. Therefore, for film and television works, the present disclosure endows new deep interaction capabilities. So in the case of film and television dramas, in addition to displaying the preset information on the user terminal interface, there is also an interactive area, which is the interactive area of the audience. They can discuss in text, leave photo descriptions for analysis, etc., so as to analyze the preset information and express their own opinions and supporting logic, thus obtaining the support of more people. Therefore, the opinions of the audience with views and the support of the crowd can influence the direction or turning point of the plot.
[0110] In film and television dramas, if the set target is a storage disk containing images, in the traditional way, only the characters in the movie can try to read it. But after adopting the technology of the present disclosure, the viewer can click on the storage disk in close-up in the video. After judging that the user has clicked on the storage disk according to the technology of the present disclosure, on the user's terminal, a file list in the storage disk can be displayed. The user can participate in the process of finding key files based on the list and looking for clues in the files. The information clicked and seen on the user terminal is also the information preset by the video producer based on the target, but the interactivity is more simulated in reality (and these contents can be a set web page, and the stored contents can be real files, and the user can look for clues among a pile of real files). Therefore, for the traditional film industry, the technology of the present disclosure undoubtedly improves the interactivity and dissemination of film and television dramas (the intense discussion among the crowd will undoubtedly attract the media, and the intervention of the crowd will then increase the popularity and dissemination).
[0111] In each of the above embodiments, what is clicked by the user is the target. Under the feature comparison method, the target needs to preset pictures for extracting features and making comparisons, and then judge the clicked target. The displayed information is also preset for the display during the dissemination effect of the video or the direction of the video content. The next procedure is that after different attributes of the target are clicked, the set purpose needs to be achieved, and the set procedures or instructions.
[0112] This disclosure believes that there are many defects in existing video technologies. Therefore, in the comparison method, when generating a video, the pictures of the objects to appear in the video are used as the input for feature extraction, comparison, and matching. When the user watches the video, after being prompted or getting into the habit, the user triggers the objects that appear in the video. The terminal program extracts the image and / or features of the clicked position and compares them with the pictures of the aforementioned objects to determine the clicked object. Then, according to the preset information, it is displayed on the display interface of the terminal, or displayed and the next program is executed.
[0113] In movies and TV shows, such as New Year's films, when users watch on the terminal, the characters in the video can also congratulate the audience with red envelopes (such as in New Year's films). When the audience clicks on the red envelope, the platform that plays the video gives the audience a certain amount of cash, tokens, discount cards, etc., or discounts or samples of some sponsored products. On the user terminal program that clicks on the target, the gains after clicking are displayed, and the gains are transferred to the corresponding digital wallet, card package, etc. This not only makes the audience happy but also realizes the conversion of some products.
[0114] One type of object is a person. For example, in the video of a video blogger, the video blogger takes himself / herself as the target. When the user triggers the face of the video blogger, the terminal program extracts the image of the clicked position in the video (the face of the person) and compares it with the target picture (face photo) pre-entered by the blogger. Then the program determines that the video viewer has triggered the face of the video blogger. The next program, such as automatically adding the set information, such as the communication contact information of the video blogger to the viewer's contact list, or following the blogger, etc. Of course, it can also be that when clicking on the face position of the video blogger, the program is set to communicate with him / her by clicking on his / her face (the contact information can be displayed or not). The communication can include voice, video, or entering the voice or text or half-duplex discussion group of the video blogger. In this way, whether it is a live video or a video, after the video viewer watches the video and is prompted by the information in the video (including voice, text displayed on the screen, or icons with specific meanings), after clicking on the target, the information preset by the video generator for the clicked target and the next program to be executed can be displayed, thus turning the one-way video into an effective tool for dissemination.
[0115] In the above-described embodiment of clicking on a person in a video, for a person engaged in promoting and selling a product, tourism, etc., if their video is seen by a video viewer and the viewer can click on their face in the video, thereby generating a commercial effect, such as direct communication, consultation, and then attracting customers, it is a great improvement to the video as a communication tool. For example, when a tourist sees a video of a scenic spot blogger, in the traditional way, a series of means such as querying are required to contact the blogger. However, by using the technology disclosed in this application, by clicking on the face area of the blogger in the video and following the above-disclosed technology, communication can be directly established with the blogger. Moreover, this method does not require the blogger to disclose contact information. Before communication, the blogger can clearly see that the caller is a trusted user with the above-disclosed technology (the caller sends the trusted information of the caller when making a call, so that the called party can distinguish that it is not an irrelevant person, and then confirm whether to answer or reject. For details, see the authorized patent 201810357591.6, "A Mobile Communication System Oriented to Scenarios and Communication Content"), thus avoiding interference from irrelevant people and serving its customer group. Of course, after triggering and identifying the target, the next step program is executed, and establishing communication is one of the next step programs to be executed.
[0116] Of course, for a video containing public figures, if the video producer sets the person as the target and sets adding attention as the next step to be executed, then if a viewer of a video sees a favorite public figure and clicks on the favorite public figure in the video, the next step program can be set as the viewer follows the public figure's official account, such as official accounts on WeChat, Twitter, etc., to increase the attention to the public figure, which will have a very good effect, thus making up for the defects of existing videos and enhancing the dissemination effect of the video.
[0117] Of course, in the previous example, for a character in a movie or TV drama, if the viewer likes him and clicks on the character in the video, then the viewer of the video follows his official account, which is also the execution of the next instruction or the next step program.
[0118] Regarding following the official account, it can be achieved by calling the API, as well as the official account of the target person (pre-set information) and the user number of the video viewer. However, all of these are executed according to the pre-set next step program after matching as described in this disclosure. And calling the API to implement this function is a typical example of reading information and executing the next instruction after confirming the comparison target.
[0119] The technology disclosed in this application is used to solve multiple defects in video technology, thereby making the video a more effective communication tool, whether in terms of information dimension, product sales, commercial activity coupons (red envelopes, discount cards, shopping, etc.) or communication and interaction with the people in the video.
[0120] The technology of the present disclosure is not used to extract picture content from videos and then match certain information from Internet data (essentially searching for pictures by pictures); the present disclosure is used to solve the deficiencies of video technology and further enhance video as a carrier, so that videos have stronger dissemination capabilities and interaction capabilities.
[0121] The above methods are all solutions proposed to solve the problem that the targets in current videos cannot be effectively expressed and displayed in videos, thus affecting the dissemination of videos. The difference is that, for example, the comparison method relies more on the image processing capabilities of the terminal. For example, in the current frame, based on the features of the target, it quickly matches, and determines the specific target based on the matching position and the position of the triggering target; if a video is long and contains multiple targets, such as a 90-minute movie containing 50 targets, it is obviously very time-consuming to match all 50 targets once; while for the first method, although the generation of target data starts after importing the video, when the user clicks, it is only the calculation based on the range of the current frame between the click position and the generated target data. Therefore, when the target is finally matched, the time efficiency is very high; during live broadcasts, the target data generated by the editing or generating side needs to be grasped by the user side in real time, that is, data channels need to be provided by T01 to ensure efficiency.
[0122] In order to simplify the second comparison method, in u10, the video generator or editor can associate time with the target according to the video timeline. For example, only targets 1, 2, and 3 exist in time period A, and only target 2 exists in time period B. Therefore, based on time, the calculation amount and time during comparison can be reduced; in a certain video segment, there are only certain targets, so only the features of these targets are used for matching during matching. These settings can all be used as preset auxiliary information. After the user triggers the target, the user terminal program or the background server, based on these preset auxiliary information, only compares the features of the targets in that specific time period, which will then reduce the feedback delay (such as the duration of finding the target and displaying the preset information of the target); however, during live broadcasts, when there are multiple targets, the recording personnel may be overwhelmed because live broadcasts have strong time constraints. In method 1, the system generates target data (the range value of the target in the current frame), but corresponding data will be generated and theoretically calculated for each frame. If no one clicks on the target, it is purely a waste of computing resources.
[0123] In method 2, the picture of F11 can also be implemented in U10 in a human-computer interaction manner. For example, the generating personnel select and circle the target, and then extract the features of the image from the circled target, but the information of the image background may be extracted; while inputting a clean picture without a background sometimes results in higher comparison efficiency.
[0124] In Method 1 or 2, when the target is determined and the information of the target file or URL is displayed, the content displayed is the same for Method 2 and 1; including, but not limited to: displaying the preset information of the target, and it is also possible to access customer service resources, and the customer service resources include human or machine resources. Specifically, for example, if a video is about a product and a user interested in the product clicks on the target product, the preset information of the product is clearly displayed on the user side, including details that cannot be shown by video technology. At this time, the machine customer service or customer service staff actively provides more information to users who enter the details for viewing, so as to better guide from the perspectives of brand advertising and performance advertising; Naturally, the information displayed for the product also includes product purchase information, such as a shopping basket, etc.; For resource links, in addition to web pages, it can of course directly be a video. For example, a video of jewelry taken with a macro lens, with the jewelry rotating under the light, and the end user can place an order for consumption when being infected by the video. Of course, at this time, the customer service can also intervene for shopping guidance and further provide more detailed information, and this information itself can also be seen by other entering users, and they can leave messages and interact with each other; For the information in pictures and vector files, the user side can zoom in, zoom out, etc., and adjust according to the user's needs. For example, a doctor views a nuclear magnetic resonance film and needs to zoom in on the position of the nodule for evaluation.
[0125] In Method 2, feature extraction is also a common-sense technology in image processing. Like object detection, there is a large amount of technical literature on these technologies, and the common-sense technologies will not be elaborated in this application.
[0126] In addition, regarding the display in the user interface, it can be displayed on the same page as the video or on a separate page, and this is not limited in this disclosure. For example, on the same page, a bottom sheet, a pop-up window, etc. can be used, and for different pages, such as opening an interface within a program or calling other general programs such as a browser or using a Deeplink or a protocol of the same nature to call a program already installed on the terminal.
[0127] For current and future technologies, the terminal is no longer a one-way receiving device like a TV. Therefore, based on the functionality of the receiving terminal, the user can read target files in various forms such as pictures, vector pictures, videos, tables, texts, etc. Consequently, the defect that the information conveyed by a video only depends on the frame content information is completely broken. Through the dissemination of the video, users can obtain information that is included in the video but cannot be clearly expressed based on their own demands. And this kind of traditional information does not have the dissemination power of the video. Therefore, this disclosure attempts to use the breadth of the video to overlay the depth of the information, so that the dissemination effect of the video is both wide and deep.
[0128] Based on the above, it actually includes video editing and generation software, playback software on the user terminal, and a video dissemination module. The above software or module runs on a computing device, which includes a CPU, storage, and network module. The above method is applied for video editing and generation personnel to edit and generate videos, and end-users to watch videos.
[0129] Among them, the software on the video editing and generation side runs on a computer or intelligent terminal; while the user-side software runs on a computer or terminal, and the video dissemination module runs on server resources; because it involves image processing and computing, it all includes a CPU, and all require storage resources including memory and external storage, and all require networking and interconnection with the video dissemination module; and the video dissemination module is usually a server or a server cluster, which provides video access services externally with network support, and also includes a database, etc.
[0130] In the above disclosure, if the target in a video includes the PPT presented by the speaker and also includes other products recommended by the speaker, such as Product A, then Method 1 and Method 2 can be used in combination. That is, the video generator demarcates the PPT display area in the video and uses Method 1, while Product A uses the image feature comparison method. This is because the content in the PPT area is dynamic (such as the PPT or video played in the projection area). Using the image comparison method, the effect is not as good as using the area as the target and adopting the target detection method. Therefore, when Method 1 and Method 2 are combined for different situations, it can not only meet content transmission but also reduce the computing resources of the system. At the same time, the user experience on the client side will be stronger. After all, image processing and comparison consume computing resources and time, and dynamic content is not very suitable for feature comparison without special processing.
[0131] The technology disclosed above essentially discloses a method for video application, that is: the user watching the video triggers the target in the video being played on the terminal, the user terminal program calculates the position of the triggered position in the video frame, and the terminal program or the background system determines the target to which the triggered position belongs according to the information of the target input by the video generation personnel. The terminal program reads the preset information of the target to which it belongs and displays it on the terminal, or reads the preset information of the target to which it belongs and executes the preset next instruction or program.
[0132] Actually, whether it is the first method or the image feature matching method, the position of the triggered point is converted into the position in the video frame, and the contact point is usually the coordinate value of (X, Y). And no matter which method, in fact, this value is within the target, so it is a certain point within the target. For example, in Figure 2In [the method], when the projection area is clicked, the clicked position is within the projection area and belongs to the point position in the target. In the method of image feature comparison, whether it is image segmentation or extracting the image in the video frame according to the triggered point based on the geometric frame, the triggered point (trigger position) is within the area of the target. Therefore, the trigger position (the position in the video frame) is the input for determining the target to which the clicked position belongs. So, whether it is Method 1 or Method 2, it is to determine the target to which the clicked position in the image belongs.
[0133] In addition, it should be noted that in image processing and comparison, such as comparing books with human faces, different algorithm models are usually adopted for better results. Therefore, in order to better achieve the goal, when setting the target, the video generator can, based on classification logic, such as human faces, items, books, designs, and the preset information, set the display content type, such as pictures, PDF files, video streams, etc.; the next program or instruction after the target is triggered can be set in various forms, such as adding to the shopping basket, adding the official account, adding attention, or cash transfer, placing an order, etc.; specific examples are as follows: When setting the video, the video generator sets Target A in the setting interface, with the type being the human face, and uploads the face photo of Target 1; the preset information is the official account of this person, and the next program executed after triggering this target is: the user who clicks adds the official account of this target. Then, when the user of the video sees this video and is interested, triggers (clicks) the face of the person in the video, the system recognizes and determines that the user has clicked Target A, and then, based on the preset information, executes the next instruction and adds the official account of this target; in the same video, if Target B is classified as a book and the preset information is the PDF version of the book, the generator needs to set Target B, upload the picture of Target B, and upload the PDF file of Target B; when the user is watching the video and clicks on this book in the video, that is, Target B, the system determines that the triggered area belongs to the position of this book in the current frame of the video, that is, Target B, then the user terminal reads this PDF file and displays it on the user terminal where the target is triggered, and the display method can be any human-computer interaction method, including windows, display areas, etc.; when executing the instruction in the above manner, it actually takes advantage of the dissemination and vividness of the video, and at the same time combines specific functions and instruction execution, thus replacing some functions of the traditional business-needed APP or small program with video interaction methods, thereby directly enhancing the business effect of the video.
[0134] In the above manner, for any television, display screen, smart terminal, or XR (VR, MR) device that can be touched (although such devices are usually triggered in a virtual manner, it is still a trigger), by applying the technology of the present disclosure, various contents in the video can be defined as targets. According to the classification, attributes, preset information, and next-step instructions of each target, and then triggering can display the preset information or display the preset information and execute the corresponding program or instruction, thus completely reversing the long-existing defects of traditional videos. In addition, different from XR, video producers do not need programs, so ordinary people can produce content, while the content of XR is currently limited to program producers, resulting in very limited content.
[0135] In the above disclosure, the program of the user terminal, such as U02, can have different manifestations on different terminals. For example, on a computer, it can be in the form of Web access, while on a mobile phone, such as Android, Apple, etc., it is usually in the form of an APP. In fact, it can also be in the form of Web. By obtaining the clicked position of the mouse or multi-touch screen, and then extracting the position in the displayed video frame, and feeding back the data to the server, that is, the video dissemination module, in a data communication manner. The video dissemination module makes a judgment and then, according to the preset, displays the pre-information or executes the preset instruction on the user side. Similarly, the program of the video generation end can also be in the form of Web access, but if it is in the form of live broadcast, the usually simple and easy way is to generate the program required by the user side, such as the program running on a computer or the APP running on a mobile terminal such as a mobile phone.
[0136] In the present disclosure, in order to achieve the above technical objectives and based on the different characteristics and attributes of the targets in the video, the above-mentioned various methods or combinations of methods are adopted, which will further improve the dissemination effect of the video. Although in implementation, according to the target, any of the above paths can be adopted, as described in the present disclosure, they are all specific embodiments of the present disclosure.
[0137] Through the above description, those skilled in the art can, based on the disclosed technology, implement the steps and processes to solve the problems of poor dissemination effect and failure to achieve the dissemination purpose that have long existed in videos.
Claims
1. A video application method, characterized in that: include: The terminal user watching the video triggers the target in the video being played on the terminal; The user terminal program calculates and obtains the triggered position in the video frame; The user terminal program or the video transmission module determines the target to which the position belongs at the current video position according to the target information input by the video generator or the target data generated by the circled target; The user terminal program reads the preset information of the attributed target and displays it on the user terminal program, or reads the preset information of the attributed target and executes the preset next step instruction.
2. A video application method according to claim 1, characterized in that: include: The objects in the video include two-dimensional objects and three-dimensional objects; Planar targets include at least one of projections, photographs, pictures, and flat artworks; while three-dimensional targets include non-planar objects; In addition, the target may also include a person in the video.
3. A video application method according to claim 1, characterized in that: include: The target information input by the video generator includes a picture of the target, and the picture is used for feature extraction and comparison.
4. A video application method according to claim 1, characterized in that: include: The video generation personnel circles the target in the video display interface of the target data generation module; The target data generation module generates the target data of the circled target based on the time value or frame sequence value of the video; Methods for delineating targets include image segmentation or geometric frame delineation.
5. A video application method according to claim 1, characterized in that: include: The video generation personnel uploads the preset information of the target to the video dissemination module through the video generation terminal program or directly.
6. A video application method according to claim 1, characterized in that: include: The generation of the target data includes using artificial intelligence and / or computer vision technology to generate the target data based on the timeline of the video and the position and trajectory of the target in the video.
7. A video application method according to claim 1, characterized in that: include: The information preset by the video generator for the target includes a file or resource link; The resource link includes a web link and a video link, and the video link includes a video stream.
8. A video application method according to claim 1, characterized in that: include: If the user of the terminal watching the video triggers and is determined to have triggered the target in the video, the video user terminal may display the preset information and also include an interactive area; The interactive area includes: the service staff interactive area or the audience interactive area; The service personnel also include machine attendants; The service personnel further describe, provide and communicate the target information with the end user.
9. A video application method according to claim 1, characterized in that: include: If the terminal user watching the video triggers the face of a person in the video, the next step instruction preset by the video generator according to the target, the preset information, read the preset information and execute the instruction; the instruction includes adding attention or establishing communication.
10. A video application method according to claim 1, characterized in that: include: If the video contains multiple targets and the target type attributes are different, the target data generation method can be combined with the feature comparison method. The target data generation method is used for targets with dynamic content changes, including targets with planar target area content changes.
Citation Information
Patent Citations
Communication scene and content-oriented new-type mobile communication system
CN108600536A