Systems and Methods for Enhancing the Live Audience Experience on Electronic Devices

Through deep learning technology, the problem of the inaccurate direction of advertising in the existing technology is solved, and efficient advertising customization and audience experience improvement are achieved.

CN112840377BActive Publication Date: 2025-06-13ZYETRIC TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980066926.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-30
Filing Date
2019-10-24
Publication Date
2025-06-13
Estimated Expiration
2039-10-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and display customized advertisements in live video streaming or broadcasting, especially in the absence of a predetermined identifier, resulting in the inability to accurately target the target audience.

Method used

By using deep learning techniques, especially trained deep neural networks, target objects and non-target objects in video frames are identified and the surface area of ​​the target objects is defined based on the identified group of pixels, covering a predetermined graphic image to form the processed live video.

Benefits of technology

It realizes accurate identification and coverage of content on billboards without relying on pre-determined identifiers, improving the customization and purpose of advertisements, and enhancing the live experience of the audience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112840377B_ABST
    Figure CN112840377B_ABST
Patent Text Reader

Abstract

The present disclosure describes methods and systems for receiving a plurality of live video frames; identifying one or more target objects and one or more non-target objects in a first live video frame of the plurality of live video frames by at least one trained deep neural network; identifying one or more sets of pixels belonging to the one or more target objects; identifying regions on surfaces of the one or more target objects based on the identified one or more sets of pixels belonging to the one or more target objects; overlaying one or more predetermined graphic images on the regions on the surfaces of the one or more target objects in the plurality of live video frames; and overlaying the one or more non-target objects on the one or more predetermined graphic images in the plurality of live video frames to form processed live video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to live video streaming or broadcasting, and more particularly, to a live audience experience in video streaming or dissemination via an electronic device. Background Art

[0002] In live sports game video streaming or broadcasting, not only the players and the game itself being streamed / broadcasted by the players, but also other static objects are shown in the video scene, such as seats, stadiums, billboards / banners. Some of these static objects carry information, but this information is not relevant to the audience / viewers. For example, the billboards / banners around a football field during a football game display advertisements. In the case of different demographics and different background technologies, the advertisements are not limited to / targeted at the audience / viewers who may come from all over the world. For example, in a live World Cup football game, one of the billboards shows an advertisement related to Deloitte (a public accounting firm) in the UK. However, this advertisement is not relevant to a Brazilian high school boy who is watching the live football game, so he is not interested in it. Moreover, the high school boy may not understand English, and as a result, the information / message of the advertisement cannot be conveyed to the target audience / viewers (in other words, the advertisement is wasted on non-target audience / viewers). There is a need to customize the content of the advertisement so that the information / message can be successfully conveyed to the target audience / viewers.

[0003] According to the known technology, during a football game, audiences in different countries watch different advertisements shown on the billboards around the edge of the football field. For example, a video of a football game played in Germany is broadcast to audiences in different countries. The advertisements (alternate advertisements) watched by audiences in China and Australia are different from those watched by German audiences. However, there are some limitations in applying alternate advertisements to videos based on the known technology. In one instance, the billboard suitable for displaying alternate advertisements has at least one identifier. A computing system (e.g., provided by a broadcasting organization) can identify the billboard as a target object based on the identifier to display an alternate advertisement on the target object. The identifier is regarded as a predetermined criterion for the computing system to identify the billboard.

[0004] For example, the identifier is the green screen / surface of the billboard. When the computing system identifies the billboard as a target object based on the green screen / surface, the alternate advertisement is configured to be displayed on the target object. In another instance, the identifier is an infrared emitter. The billboard contains an infrared emitter that transmits an infrared signal to a camera. Based on the infrared signal, the camera identifies the billboard as a target object, and the computing system then arranges to display the alternate advertisement on the billboard.

[0005] In the absence of an identifier, the computing system is unable to determine the target object, and as a result, the viewer cannot watch the alternative advertisement. The present invention can identify the target object through deep learning without adhering to any predetermined criteria. For example, a video contains a billboard that does not contain predetermined criteria, and the alternative broadcast cannot be applied to the billboard. For example, the 1998 World Cup final video (recorded video) is available on an online video sharing platform. The video contains multiple billboards around the edge of a football field. However, none of the billboards are green (predetermined criteria), and as a result, the mentioned advertisement cannot be applied to those billboards during the user's video streaming.

[0006] The present invention relates to improved techniques for enhancing the live viewer experience and provides related advantages. Summary of the Invention

[0007] Example methods are disclosed herein. The examples include: receiving, at an electronic device, a plurality of live video frames by the electronic device; identifying, by at least one trained deep neural network, one or more target objects and one or more non-target objects in a first live video frame of the plurality of live video frames; identifying a set or sets of pixels belonging to the one or more target objects; defining, based on the identified set or sets of pixels belonging to the one or more target objects, a region on the surface of the one or more target objects; overlaying, in the plurality of live video frames, one or more predetermined graphic images on the region on the surface of the one or more target objects; and overlaying, in the plurality of live video frames, one or more non-target objects on the one or more predetermined graphic images to form a processed live video, wherein the processed live video includes the one or more non-target objects and the one or more predetermined graphic images overlaid on the one or more target objects.

[0008] In some examples, the one or more target objects include one or more static objects, and the one or more non-target objects include one or more objects in front of the one or more static objects, wherein the one or more objects obscure the one or more static objects.

[0009] In some examples, the one or more static objects include one or more billboards.

[0010] In some embodiments, a computer-readable storage medium stores one or more programs, and the one or more programs include instructions that, when executed by the electronic device, cause the electronic device to perform any of the methods described above and herein.

[0011] In some embodiments, an electronic device includes one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods described above and herein.

[0012] For the foregoing reasons, there is a need for a computing system that can effectively display customized advertisements without requiring billboards to follow any predetermined guidelines. There is also a need for a computing system for customizing a live broadcast of an event in real time or near real time according to various advertising requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A screenshot depicting an example of a live soccer game video displayed on an electronic device according to various embodiments of the present invention.

[0014] Figure 2A and Figure 2B A schematic diagram depicting the determination of the true boundary of a target object using a delimiting member according to various embodiments of the present invention.

[0015] Figures 3A to 3D A schematic diagram depicting a line generated based on pixels identified as extreme points according to various embodiments of the present invention.

[0016] Figure 4 A screenshot depicting an example of a processed live soccer game video displayed on an electronic device based on first viewer personal information according to various embodiments of the present invention.

[0017] Figure 5 A screenshot depicting an example of a processed live soccer game video displayed on an electronic device based on second viewer personal information according to various embodiments of the present invention.

[0018] Figure 6 A screenshot depicting an example of a processed live soccer game video displayed on an electronic device in a country according to various embodiments of the present invention.

[0019] Figure 7 An example flowchart depicting a process of generating processed live soccer game video frames according to various embodiments of the present invention.

[0020] Figure 8 An example flowchart depicting a process of training an electronic device to identify target objects and non-target objects according to various embodiments of the present invention.

[0021] Figures 9A to 9BSchematic diagram depicting a processed live video displayed on an electronic device based on first viewer personal information according to various embodiments of the present invention.

[0022] Figures 10A to 10C Schematic diagram depicting a processed live video displayed on an electronic device according to various embodiments of the present invention.

[0023] Figure 11 Computing system that can be used to implement various embodiments of the present invention.

[0024] Figure 12 Example flowchart depicting the process of generating processed live football game video frames at a server according to various embodiments of the present invention.

[0025] Figure 13 Alternative example flowchart depicting the process of generating processed live football game video frames at a server according to various embodiments of the present invention. Detailed Description

[0026] The following description is presented to enable a person having ordinary skill in the art to make and use various embodiments. The description of specific devices, techniques, and applications is provided only as an example. Those skilled in the art will readily appreciate various modifications to the examples described herein, and the general principles defined herein can be applied to other examples and applications without departing from the spirit and scope of the present invention. Thus, the disclosed invention is not intended to be limited to the examples described and shown herein, but is accorded the scope consistent with the claims.

[0027] Today, people can watch live videos (i.e., for example, live sports game videos) through various platforms. Some platforms are free, while some are paid monthly or annually. Live sports games can be football games, tennis games, hockey games, basketball games, baseball games, or any sports game. For example, the World Cup is the largest global sports event, with billions of people watching the month-long, quadrennial event. The football game period is a valuable time for various enterprises and units to promote their products or services. Multiple billboards / banners are located around the football field / football stadium. Multiple billboards are dedicated to displaying advertisements promoting various products / services. The advertisements can carry information in different languages.

[0028] Figure 1A screenshot depicting an example of live football game video streaming or broadcasting on an electronic device. In some examples, viewers / audiences like to watch live football game video streaming / broadcasting on an electronic device such as smart device 100. Smart device 100 can be a desktop computer, a laptop computer, a smart phone, a tablet computer, a wearable device, or goggles. Smart device 100 is similar to and includes all or some components of computing system 1100 described below in FIG. 9. In some embodiments, smart device 100 includes a touch-sensitive display 102, a front camera 120, and a speaker 122. In other examples, the electronic device can be a television, a monitor, or other video display device.

[0029] Streaming / broadcasting live football game video to viewers via a video recording device located at a football field / football stadium. The live football game video streaming / broadcasting includes a plurality of live football game video frames. In some examples, viewers are allowed to watch the live football game video on smart device 100 via a website, application software, or software program. The website, application software, or software program can be free or chargeable.

[0030] As Figure 1 depicted, view 160 includes but is not limited to football field 162, players 164A, 164B, 164C, and 164D, football 166, goal 168, audience 170, first billboard 182, and second billboard 184. In view 160, players 164A, 164B, 164C, and 164D and goal 168 in the live football game video streaming / broadcasting are objects in front of first billboard 182 and second billboard 184 and also block first billboard 182 and second billboard 184 when a viewer watches the live football game video on smart device 100.

[0031] There is no limitation on the objects shown in the live football game video frames. For example, the video frames can include ten billboards, two goals, one football, one referee, and twenty-two players, can include three billboards, two footballs, one goal, and two players, can include two billboards and one goal, or can include two billboards. There is no limitation on the objects in front of and also blocking the billboards. For example, the objects can include players 164A and 164B, football 166, and goal 168, can include football 166 and goal 168, or can include players 164C and 164D and football 166.

[0032] The first billboard 182 and the second billboard 184 are static objects in the live football game video. In view 160, players 164A - 164D and the goal 168 are in front of the first billboard 182 and the second billboard 184. Players 164A - 164D and the goal 168 occlude the first billboard 182 and the second billboard 184. The first billboard 182 and the second billboard 184 are determined as target objects by at least one trained deep neural network. Players 164A - 164D and the goal 168 are determined as non - target objects by the trained deep neural network. There is no restriction on the position of the billboards. The billboards can be located at any position around the football field.

[0033] The trained deep neural network is obtained by feeding multiple photos and / or videos of the football game as training data into a training module, where a process of running a deep learning algorithm is performed. The training module can be located in the smart device 100 or the server. In some examples, the trained deep neural network includes a first trained deep neural network suitable for identifying one or more target objects and a second trained deep neural network suitable for identifying one or more non - target objects.

[0034] In some examples, the first advertisement content and the second advertisement content are respectively displayed on the surfaces of the first billboard 182 and the second billboard 184. The first advertisement content relates to a Chinese car brand and the second advertisement content relates to a UK power tool brand (which are respectively displayed on the first billboard 182 and the second billboard 184 in the live football game streamed or broadcast in real - time or near - real - time). Billions of viewers from different countries watch the live football game video. However, for non - Chinese viewers, they may not understand the first advertisement content. Additionally, not every viewer is interested in power tools (the second advertisement content). Based on viewer preferences, viewer background, or other information associated with the viewer, the first and second advertisement contents need to be suitable for the viewer.

[0035] Figure 2A and 2B An example of using delimiting members to determine the true boundaries of target objects so that a predetermined graphic image can be overlaid thereon is depicted. In some examples, the smart device 100 receives the live football game video. The live football game video includes multiple live football game video frames. When the smart device 100 identifies one or more target objects in the first live football game video frame of the multiple live football game frames through at least one deep neural network trained by deep learning, one or more predetermined graphic images are configured to overlay the one or more target objects. However, since the true boundaries of the one or more targets cannot be determined, the predetermined graphic images may not be aligned with the one or more target objects.

[0036] AsFigure 2A As depicted, for simplicity, the first billboard 182 is described herein as the target object. View 260A is shown on the touch-sensitive display 102 and includes a first delimiting member 290 generated as a range around the billboard 182. A similar delimiting member is also applied to the second billboard 184. The first delimiting member 290 can be circular, box-shaped, or of any shape. The first delimiting member 290 is generated in a conventional manner without applying any mathematical functions thereto (e.g., linear regression), so the first delimiting member 290 does not align with the true boundary of the billboard 184, and when a predetermined graphic image is overlaid on the billboard 384, the predetermined graphic image cannot be aligned with the billboard.

[0037] To optimize the accuracy of the delimiting member, by way of example only, the smart device 100 is configured to scan received live soccer game video frames to identify one or more groups of pixels belonging to the billboard 182 through a trained deep neural network. Based on the identified one or more groups of pixels, a second delimiting member 292 is formed. View 260B includes the second delimiting member 292 that substantially aligns with the true boundary of the first billboard 182, as Figure 2B depicted (substantially matching the contour / shape of the first billboard 182). For example, the smart device 100 scans the first live soccer game video frame among multiple live soccer game video frames in a predetermined order, such as from left to right, from top to bottom, from right to left, and from bottom to top. The smart device 100 scans the first live soccer game video frame from top to bottom to determine a first group of pixels belonging to the first billboard 182 through a trained deep neural network.

[0038] There is no limitation on the predetermined scanning order. For example, the predetermined order can be from right to left, from top to bottom, from bottom to top, from left to right. There is no limitation on the scanning area. For example, the smart device 100 can partially scan the first live soccer game video frame, that is, the smart device 100 can scan the area of the first live soccer game video frame containing the target object. One benefit of partial scanning is to reduce the computational cost when scanning fewer pixels.

[0039] Among the first group of pixels, the smart device 100 then identifies one or more pixels in the first group of pixels as extreme points 302A (based on 2D coordinates) by scanning from left to right, as Figure 3A depicted. An extreme point is a pixel in a prominent position among adjacent pixels. Subsequently, at least one mathematical function is applied to the extreme points 302A to obtain a line 304A. The mathematical function can take one of multiple forms, including but not limited to linear regression. The line 304A will correspond to the top boundary line of the second delimiting member 292.

[0040] The intelligent device 100 will then scan the first live football game video frame from top to bottom, from right to left, and from bottom to right to obtain the extreme points 302B, 302C, and 302D depicted in 2D, Figure 3B , 3C , and 3D respectively. Linear regression will be applied to each of the extreme points 302B, 302C, and 302D, thereby forming lines 304B, 304C, and 304D. Lines 304B, 304C, and 304D correspond to the left boundary line, bottom boundary line, and right boundary line of the second delimiting member 292 respectively. Figure 3B , 3C and the extreme points 302B, 302C, and 302D depicted in 3D. Linear regression will be applied to each of the extreme points 302B, 302C, and 302D, thereby forming lines 304B, 304C, and 304D. Lines 304B, 304C, and 304D correspond to the left boundary line, bottom boundary line, and right boundary line of the second delimiting member 292 respectively.

[0041] Based on the second delimiting member 292, the true boundary of the first billboard 182 is determined. The second delimiting member 292 delimits an area 294 on the surface of the first billboard 182. The intelligent device 100 will determine the 3D visual features of the first billboard 182 in the original live football game video frame, such as perspective projection shape, lighting, or any other features. A predetermined graphic image is appropriately overlaid on the said area. The predetermined graphic image may contain the 3D visual features of the first billboard 182. In order to make the predetermined graphic image feel real (as if it should have been in the appropriate position in the real environment), the 3D visual features of the target object (the first billboard 182) are applied to the predetermined graphic image. 3D features are extracted from the target object. The 3D features include but are not limited to brightness, resolution, aspect ratio, viewing angle. Taking the viewing angle and aspect ratio as examples, since the 3D object is projected onto a 2D screen, the 3D regular object may become trapezoidal, and the angles and side lengths of the trapezoid are measured. The predetermined graphic image is transformed with the same angles and side lengths, that is, the predetermined graphic image is transformed into the same trapezoid and then appropriately overlaid on the target object. Taking brightness as another example, the target object is divided into smaller regions of equal size. The smaller the region, the higher the resolution of the brightness, but the higher the computing power required. For each region, the brightness is estimated. One estimation method is to use OpenCV to test the β value of the specific region. Subsequently, the same β value is applied to the corresponding region of the predetermined graphic image.

[0042] The shape of the second delimiting member 292 depends on the actual shape of the target object (the billboard 182). There is no restriction on the shape of the target object. Determining the extreme points from a set of one or more groups of pixels of the target object and the linear regression applied thereto can be used to determine the true boundary of the target object of any shape.

[0043] Figure 4 A screenshot depicting an example of the processed live football game video displayed on the electronic device based on the first viewer's personal information is shown. By way of example only, a live football game video is received by the electronic device used (such as the intelligent device 400).

[0044] A live football game video includes multiple live football game video frames. A first viewer is allowed to view the live football game video via a smart device 400. The received live football game video frames will be processed at the smart device 400 by displaying advertisement content that may be suitable for the first viewer or that the viewer may be interested in.

[0045] In a first live football game video frame among the multiple live football game video frames, the smart device 400 will identify one or more target objects (static objects in the first live football game video frame) and one or more non-target objects (objects in front of the static objects and that can also occlude the static objects in the first live football game video frame) through at least one deep neural network trained by deep learning. In this case, the smart device 400 determines the first billboard 182 and the second billboard 184 as target objects and the players 164A, 164B, 164C, and 164D and the goal 168 as non-target objects through the trained deep neural network.

[0046] As Figure 4 depicted, view 460 is displayed on the touch-sensitive display 402 of the smart device 400. View 460 includes a football field 162, players 164A, 164B, 164C, and 164D, a football 166, a goal 168, spectators 170, and a first billboard 182 and a second billboard 184. In this case, based on the first viewer's personal information, the first advertisement content related to a Chinese car brand and the second advertisement content related to a British power tool brand are replaced by first and second predetermined advertisement content.

[0047] The smart device 400 identifies the first billboard 182 and the second billboard 184 as target objects. A second delimiting member 292 will be generated to surround each extent of the billboards 182 and 184. The second delimiting member 292 is configured to determine the true boundaries of the first billboard 182 and the second billboard 184 and to delimit regions 294 on each surface of the first billboard 182 and the second billboard 184.

[0048] When the regions 294 are defined on each surface of the first billboard 182 and the second billboard 184, the first predetermined graphic image 486 and the second predetermined graphic image 488 are respectively and appropriately overlaid on the surfaces of the first billboard 182 and the second billboard 184. The first predetermined graphic image 486 and the second predetermined graphic image 488 belong to a plurality of predetermined graphic images stored in the memory of the smart device 400 or the server. Based on the first viewer's personal information, the first predetermined graphic image 486 and the second predetermined graphic image 488 respectively show the first predetermined advertisement content and the second predetermined advertisement content. The first predetermined graphic image 486 and the second predetermined graphic image 488 may respectively include 3D visual features of the first billboard 182 and the second billboard 184 in the original live football game video frame, such as perspective projection shape, lighting, or any other features.

[0049] Once the first predetermined graphic image 486 and the second predetermined graphic image 488 are respectively laid flat on the first billboard 182 and the second billboard 184, non-target objects are subsequently overlaid in front of the first billboard 182 and the second billboard 184 at positions that are the same as or substantially similar to those in the original live football game video frame. In subsequent live football game video frames among the plurality of live football game video frames, the predetermined graphic images 486 and 488 are overlaid on the billboards 182 and 184, and then non-target objects are overlaid in front of the billboards 182 and 184. In this way, any graphic image laid flat on the billboard looks natural and feels as if these graphic images should be on the billboard in the real world.

[0050] Once a target object in the first football game video frame (e.g., view 460) of the plurality of live football game video frames is recognized by the trained deep neural network, the target object is tracked by using a video object tracking algorithm. For subsequent live football game video frames among the plurality of live football game video frames, the tracked target object is recognized by using the video object tracking algorithm. When a new target object appears in a subsequent live football game video frame, the trained deep neural network keeps recognizing the new target object. Video object tracking algorithms are known to those skilled in the art. Known video object tracking algorithms such as MedianFlow, MOSS (Minimum Output Sum of Squares) can be used.

[0051] One benefit of using a video object tracking algorithm is saving the neural network training cost, which is based on a series of huge training datasets and computing power. The trained deep neural network may not be able to recognize the target object in each of multiple live soccer game video frames. If tracking is not performed, in some of the multiple live soccer game video frames, when the target object cannot be recognized by the trained deep neural network, a predetermined graphic image will not be overlaid on the target object. In this case, a highly accurate trained deep neural network is required, which requires huge training datasets and powerful computing power. Additionally, if tracking is not performed, it is necessary to determine the true boundary of the target object in each of the multiple live soccer game video frames (with the target object), which requires powerful computing power and more processing time.

[0052] In some examples, the first viewer is allowed to pre-enter their personal information at the user interface or any platform / media. The user interface can be provided by a website, application software, or a software program implementing the present invention. The personal information can include age, gender, education level, address, nationality, religion, occupation, marital status, family members, preferred language, geographical location, salary, hobbies, or any other information related to the first viewer.

[0053] In other examples, the personal information of the first viewer can also be obtained through the first viewer's other online activities rather than pre-entry. For example, based on their online shopping records, their preferences for certain products and their interests and hobbies can be inferred.

[0054] For example, the first personal information of the first viewer is male, married, has a child, 35 years old, lives in San Francisco, has English as the mother tongue, is a lawyer, a movie lover, and a traveler. Based on their personal information, the predetermined graphic image can include advertising content related to high-end HIFI / home theater equipment, luxury watches, luxury cars, household products, health products, airlines, and / or travel agencies. The language used in most of the predetermined advertising content is English. It is desired to display the predetermined advertising content closely related to the daily life of the first viewer on the first billboard 182 and the second billboard 184. For example, the first predetermined graphic image 486 can include the first predetermined advertising content related to a luxury watch brand, and the second predetermined graphic image 488 can include the second predetermined advertising content related to a luxury car brand. Both the first and second predetermined information are in English. Now, the first viewer can watch the advertising content during the live soccer game video streaming / broadcasting, and the advertising content may attract their attention (through the processed live soccer game video frames).

[0055] Alternatively, it is allowed to process a live football game video in an electronic device such as a server. The server receives the live football game video from a video recording device. The live football game video includes a plurality of live football game video frames. The server will identify one or more target objects and one or more non-target objects in the received live football game video frames through a trained deep neural network stored in the server. In this case, the server determines the billboards 182 and 184 as target objects, and determines the players 164A, 164B, 164C, and 164D and the goal 268 as non-target objects.

[0056] Based on the first user personal information, the first and second advertisement contents in the original live football game video frame will be replaced by the first and second predetermined advertisement contents respectively displayed on the first predetermined graphic image 486 and the second predetermined graphic image 488. The first predetermined graphic image 486 is appropriately overlaid on the surface of the first billboard 182. The second graphic image 488 is appropriately overlaid on the surface of the second billboard 184, and then the non-target objects are overlaid in front of the first billboard 182 and the second billboard 184, where the positions are the same as or substantially similar to those in the original live football game video frame. Subsequently, the processed live football game video image is transmitted to the smart device 400. The first viewer can view the processed live football game video on the touch-sensitive display 402 of the smart device 400.

[0057] In a variant, the server receives a live football game video from a video recording device. The live football game video includes a plurality of live football game video frames. The server will identify one or more target objects and one or more non-target objects in the received plurality of live football game video frames by using a trained deep neural network. The trained deep neural network is stored in the server. The server determines the true boundaries of the target objects, determines the 3D visual features of the target objects and tracks the target objects.

[0058] Subsequently, the server takes all this information as metadata of the live football video frame, and then sends the original live football video frame with the metadata object to the viewer device (smart device 400). The smart device 400 reads the metadata object, and places the predetermined graphic image stored in the smart device 400 on the target objects (the first billboard 182 and the second billboard 184) according to the information provided by the metadata object to form a processed video. The processed video will then be displayed on the smart device 400.

[0059] Figure 5A screenshot depicting an example of a processed live football game video displayed on an electronic device based on second viewer personal information. In some examples, the second viewer is a single male, lives in Tokyo, is 25 years old, has Japanese as a native language, is a salesperson, and is a sports enthusiast. The live football game video will be processed in the electronic device used by the second viewer to view the live football game video, such as the smart device 500, or other electronic devices such as a server (as mentioned above). The smart device 500 receives the live football game video from a video recording device. The live football game video includes a plurality of live football game video frames.

[0060] In a first live football game video frame among the plurality of live football game video frames, the smart device 500 will identify one or more target objects (static objects in the first live football game video frame) and one or more non-target objects (objects in front of the static object and also obscuring the static object in the first live football game video frame) through at least one deep neural network trained by deep learning. In this case, the smart device 500 determines the billboards 182 and 184 as target objects and the players 164A, 164B, 164C, and 164D and the goal 168 as non-target objects through the trained deep neural network.

[0061] As Figure 5 depicted, the view 560 is displayed on the touch-sensitive display 502 of the smart device 500. The view 560 includes a football field 162, players 164A, 164B, 164C, and 164D, a football 166, a goal 168, spectators 170, and a first billboard 182 and a second billboard 184. In this case, based on the second viewer personal information, the first advertisement content related to a Chinese car brand and the second advertisement content related to a British power tool brand are replaced by the first and second predetermined advertisement contents.

[0062] The smart device 500 identifies the billboards 182 and 184 as target objects. A second delimiting member 292 will be generated to surround each extent of the billboards 182 and 184. The second delimiting member 292 is applicable to determine the true boundaries of the first billboard 182 and the second billboard 184, and defines regions 294 on each surface of the first billboard 182 and the second billboard 184.

[0063] When the regions 294 are defined on each surface of the first billboard 182 and the second billboard 184, the first predetermined graphic image 586 and the second predetermined graphic image 588 are respectively and suitably overlaid on the surfaces of the first billboard 182 and the second billboard 184. The first predetermined graphic image 586 and the second predetermined graphic image 588 belong to a plurality of predetermined graphic images stored in the memory of the smart device 500 or the server. Based on the second viewer personal information, the first predetermined graphic image 586 and the second predetermined graphic image 588 respectively show the first predetermined advertisement content and the second predetermined advertisement content. The first predetermined graphic image 586 and the second predetermined graphic image 588 may respectively include 3D visual features of the first billboard 182 and the second billboard 184 in the original live football game video frame, such as perspective projection shape, lighting, or any other features. In this way, any predetermined graphic image laid flat on the billboard looks natural and feels as if these predetermined graphic images should be on the billboard in the real world.

[0064] Once the first predetermined graphic image 586 and the second predetermined graphic image 588 are respectively laid flat on the first billboard 182 and the second billboard 184, non-target objects are then overlaid in front of the first billboard 182 and the second billboard 184, where the positions are the same as or substantially similar to those in the original live football game video frame. In subsequent live football game video frames among a plurality of live football game video frames, the predetermined graphic images 586 and 588 are overlaid on the billboards 182 and 184, and then non-target objects are overlaid in front of the billboards 182 and 184.

[0065] Based on the second viewer personal information, the predetermined graphic image may include sports equipment, computers, wearable devices, entry-level cars, travel agencies, and / or social media. The language used in most of the advertisement content is Japanese. It is desired to display advertisement content closely related to the daily life of the second viewer on the first billboard 182 and the second billboard 184. For example, the first predetermined graphic image 586 may include advertisement content related to a Japanese video game brand, and the second predetermined graphic image 588 may include advertisement content related to a Japanese sports equipment brand. Now, the second viewer can watch the advertisement content during the live football game video streaming / broadcasting, and the advertisement content may attract their attention (through the processed live football game video frame).

[0066] Figure 6A screenshot depicting an example of a processed live football game video displayed on an electronic device based on geographical location. In some examples, a third viewer uses a smart device 600 to view the live football game video. The smart device 600 is located in the United States. The smart device 600 receives the live football game video from a video recording device. The received live football game video will be processed in the smart device 600. Alternatively, it is also allowed to process the live football game video in a server.

[0067] As Figure 6 depicted, view 660 is shown on the touch-sensitive display 602 of the smart device 600. View 660 includes a football field 162, players 164A, 164B, 164C, and 164D, a football 166, a goal 168, spectators 170, and a first billboard 182 and a second billboard 184.

[0068] The smart device 600 will identify one or more target objects (static objects in the original live football game video frames) and one or more non-target objects (objects in front of the static objects and obscuring the static objects in the original live football game video frames) through at least one deep neural network trained by deep learning. In this case, the smart device 600 determines the billboards 182 and 184 as target objects and the players 164A, 164B, 164C, and 164D and the goal 168 as non-target objects through the trained deep neural network.

[0069] In this case, a first predetermined graphic image 686 is configured to appropriately cover the surface of the first billboard 182. A second graphic image 688 is configured to appropriately cover the surface of the second billboard 184. The first predetermined graphic image 686 contains first predetermined advertisement content, and the second predetermined graphic image 688 contains second predetermined advertisement content. For example, the first predetermined graphic image 686 may contain first predetermined advertisement content related to UK sports equipment, and the second predetermined may contain second predetermined advertisement content related to a UK car brand.

[0070] There is no limitation on what predetermined advertisement content is included in the predetermined graphic images 686 and 688. For example, the predetermined graphic image may contain advertisement content related to household products, professional services, fashion products, food and beverage products, electronic products, or any product / service in the UK.

[0071] Now refer to Figure 7, an example process 700 for generating and providing a process live video on an electronic device is shown. In some examples, process 700 is implemented in real-time or near real-time at an electronic device (e.g., smart device 400) having a display, one or more image sensors. Process 700 includes receiving a live video, e.g., a live football game video (block 701). The live football game video is received from a video recording device located at a football field. The live football game video includes a plurality of live football game video frames (original live football game video frames).

[0072] The smart device 400 then determines the target objects and non-target objects in the first live football game video frame among the plurality of live football game video frames. For example, the first live football game video frame includes a football field 162, players 164A, 164B, 164C, and 164D, a football 166, a goal 168, spectators 170, and a first billboard 182 and a second billboard 184. The first billboard 182 and the second billboard 184 are static objects in the original live football game video frame. The players 164A, 164B, 164C, and 164D and the goal 168 are objects in front of the static objects and also occlude the static objects.

[0073] The smart device 400 determines the first billboard 182 and the second billboard 184 as target objects and the players 164A, 164B, 164C, and 164D and the goal 168 as non-target objects through at least one trained deep neural network (block 702).

[0074] The smart device 400 scans the first live football game video frame in a predetermined order, e.g., from left to right, from top to bottom, from right to left, and from bottom to top, to identify a group of pixels belonging to the target object through the trained deep neural network (block 703). For simplicity, the first billboard 182 is described as the target object herein. The same process also applies to the second billboard 184.

[0075] Based on the left-to-right scan, the smart device 400 identifies a first group of pixels belonging to the first billboard 182 through the trained deep neural network. Among the first group of pixels, the smart device 400 then identifies one or more pixels in the first group of pixels as extreme points 302A based on the Y coordinate value of the pixels. For example, as Figure 3AAs depicted, when scanning from left to right, the position of pixel 312A is higher than the positions of pixels 310A and 314A (pixel 312A has a larger Y coordinate value than pixels 310A and 314A). Thus, pixel 312A is identified as extreme point 302A. Subsequently, pixel 318A is identified as another extreme point 302A because its position is higher than both its adjacent right and left pixels (pixels 316A and 320A). In the same manner, pixels 322A and pixel 328A are identified as other extreme points 302A. By way of counterexample, pixel 324A is not considered an extreme point 302A. Although pixel 324A is higher than pixel 326A (pixel 324A has a larger Y coordinate value than 326A), pixel 324A is lower than pixel 322A (pixel 324A has a smaller Y coordinate value than 322A). To be identified as an extreme point, a pixel must be higher than its two immediately adjacent pixels. Linear regression is then applied to extreme points 302A to obtain a first line 304A (block 704). For a regular shape or a straight line, linear regression can include the formula y = b + ax, where a and b are constants estimated according to the linear regression process. x and y are coordinates on the image frame, i.e., coordinates on the screen of the smart device or any other video player. For an irregular shape or a curve, linear regression can include the formula By adjusting the value of n, the curve can be aligned as closely as possible with the boundary of the target object, a i is a constant estimated according to the linear regression process.

[0076] Based on scanning from top to bottom, the smart device 400 identifies a second group of pixels belonging to the billboard 182 through a trained deep neural network. Among the second group of pixels, the smart device 400 then identifies one or more pixels in the second group of pixels as extreme points 302B based on the X coordinate values of the pixels. For example, as Figure 3BAs depicted, when scanning from top to bottom, the position of pixel 312B is more to the left than the positions of pixels 310B and 314B (pixel 312B has an X coordinate value smaller than those of pixels 310B and 314B). Therefore, pixel 312B is identified as the extreme point 302B. Subsequently, pixel 318B is identified as the extreme point 302B because its position is more to the left than both of its adjacent upper and lower pixels (pixels 316B and 320B). Using the same method, pixels 322B and pixel 328B are identified as other extreme points 302B. By way of counterexample, pixel 316B is not regarded as the extreme point 302B. Although pixel 316B is more to the left than pixel 314B (pixel 316B has an X coordinate value smaller than 314B), pixel 316B is more to the right than pixel 318B (pixel 316B has an X coordinate value larger than 318B). To be identified as an extreme point, a pixel must be more to the left than its two immediately adjacent pixels. Subsequently, linear regression is applied to the extreme points 302B to obtain the second line 304B (block 704).

[0077] Based on scanning from right to left, the smart device 400 identifies a third group of pixels belonging to the billboard 182. In the third group of pixels, the smart device 100 subsequently identifies one or more pixels in the third group of pixels as extreme points 302C based on the Y coordinate values of the pixels. For example, as Figure 3C depicted, when scanning from right to left, the position of pixel 312C is lower than the positions of pixels 310C and 314C (pixel 312C has a Y coordinate value smaller than those of pixels 310C and 314C). Therefore, pixel 312C is identified as the extreme point 302C. Subsequently, pixel 318C is identified as another extreme point 302A because its position is lower than both of its adjacent right and left pixels (pixels 316C and 320C). Using the same method, pixels 322C and pixel 328C are identified as other extreme points 302C. By way of counterexample, pixel 324C is not regarded as the extreme point 302C. Although pixel 324C is lower than pixel 326C (pixel 324C has a Y coordinate value smaller than 326C), pixel 324C is higher than pixel 322C (pixel 324C has a Y coordinate value larger than 322C). To be identified as an extreme point, a pixel must be lower than its two immediately adjacent pixels. Subsequently, linear regression is applied to the extreme points 302C to obtain the third line 304C (block 704).

[0078] Based on scanning from bottom to top, the smart device 400 identifies a fourth group of pixels belonging to the billboard 182. In the third group of pixels, the smart device 400 subsequently identifies one or more pixels in the third group of pixels as extreme points 302D based on the X coordinate values of the pixels. For example, as Figure 3DAs depicted, when scanned from bottom to top, the position of pixel 312D is further to the right than the positions of pixels 310D and 314D (pixel 312D has a larger X coordinate value than pixels 310D and 314D). Thus, pixel 312D is identified as extreme point 302B. Subsequently, pixel 318D is identified as extreme point 302D because its position is further to the right than both of its adjacent upper and lower pixels (pixels 316D and 320D). In the same manner, pixel 322D and pixel 328D are identified as other extreme points 302D. By way of counterexample, pixel 316D is not considered an extreme point 302D. Although pixel 316D is further to the right than pixel 314D (pixel 316D has a larger X coordinate value than 314D), pixel 316D is further to the left than pixel 318B (pixel 316D has a smaller X coordinate value than 318D). To be identified as an extreme point, a pixel must be further to the right than two pixels adjacent to it. Subsequently, linear regression is applied to extreme points 302D to obtain the fourth line 304D (block 704).

[0079] Based on lines 304A - 304D, a second delimiting member 292 is formed (block 704). Lines 304A and 304C correspond to the top boundary line of the second delimiting member 292 and the bottom boundary line of the second delimiting member 292, respectively. Lines 304B and 304D correspond to the left boundary line of the second delimiting member 292 and the right boundary line of the second delimiting member 292, respectively. The second delimiting member 292 is substantially aligned with the true boundary of the first billboard 182 (substantially matching the contour / shape of the first billboard 182). The second delimiting member 292 delimits an area 294 on the surface of the first billboard 182. The smart device 100 will determine the 3D visual features of the first billboard 182 in the original live soccer game video frame, e.g., perspective projection shape, illumination, or any other feature (block 705).

[0080] Once the target object (the first billboard 182) in the live soccer game video frame is identified by the trained deep neural network, the target object is tracked by using a video object tracking algorithm (block 706). For subsequent live soccer game video frames in a plurality of live soccer game video frames, the tracked target object is identified using the video object tracking algorithm. When a new target object appears in a subsequent live soccer game video frame, the trained deep neural network continues to identify the new target object.

[0081] Based on the personal information of the first viewer, a predetermined graphic image is appropriately overlaid on region 294 (block 707). In one example, a first graphic image layer containing a first predetermined graphic image 486 is overlaid on a first target object layer containing a first billboard 182, resulting in the first predetermined graphic image 486 being appropriately overlaid on region 294 of the first billboard 182. The first predetermined graphic image 486 includes the 3D visual features of the first billboard 182 in the original live soccer game video frame. In this way, the first predetermined graphic image 486 lying flat on the first billboard 182 looks natural and feels as if the first predetermined graphic image 486 should be on the first billboard 182 in the real world. When the target object and its true boundaries are determined, block 707 will be applied to subsequent frames in multiple live soccer game video frames.

[0082] Once the first graphic image layer is overlaid on the first target object layer, a first non-target object layer containing non-target objects is overlaid on the graphic image layer. The non-target objects are then placed in front of the first billboard 182, where the positions are the same as or substantially similar to those in the original live soccer game video frame (block 708). When the target object and its true boundaries are determined, block 708 will be applied to subsequent frames in multiple live soccer game video frames.

[0083] When blocks 707 and 708 are applied to multiple live soccer game video frames, a processed live soccer game video is formed, including the first predetermined graphic image 486 lying flat on the first billboard 182 and the second predetermined graphic image 488 lying flat on the second billboard 184. The first viewer is allowed to watch the processed live soccer game video on the touch-sensitive display 402 of the smart device 400 in real time or near real time, as if the first viewer is watching a live soccer game including the first billboard 182 displaying a luxury watch brand advertisement and the second billboard 184 displaying a luxury car brand advertisement in the real world.

[0084] In a variant, the electronic device can be a server. The server executes as Figure 12The process 1200 described therein. For example, the server is allowed to execute box 1201 to box 1208 (which is equivalent to executing box 701 to box 708 of process 700). At box 1209, the server will generate a processed live video by overlaying one or more predetermined graphic images (the first predetermined graphic image 486) on one or more target objects (the first billboard 182) and overlaying one or more non-target objects on one or more predetermined graphic images in subsequent frames of multiple live football game video frames. The server will then transmit the processed live football game video to one or more other electronic devices (e.g., a desktop computer, a laptop computer, a smart device, a monitor, a television, or any other video display device) at box 1210 for display thereon.

[0085] In a variant, the server executes box 1301 to box 1306 of process 1300 as shown in Figure 13 which is equivalent to executing box 701 to box 706 of process 700. The server takes all the information (generated from box 1301 to box 1306) as metadata of the live football game video frames at box 1307, and then sends the live football game video frames with metadata to the viewer device (e.g., the smart device 400) at box 1308. The smart device 400 then applies box 707 to box 708 to the live football game video frames. The processed video will then be displayed on the touch-sensitive display 402 of the smart device 400.

[0086] The smart device 100 or the server is pre-trained to identify one or more target objects and one or more non-target objects through at least one deep neural network trained by deep learning. Figure 8 Depicts an example process 800 for training at least one deep neural network, which resides in, for example, the smart device 100 or the server, to identify target objects and non-target objects in a live video (e.g., a live football game video). The smart device 100 or the server includes at least one training module. At box 801, the training module receives multiple photos and / or videos of a football game as training data, and trains at least one deep neural network at the training module. The deep neural network can be a convolutional neural network (CNN), or a variant that combines CNN with a recurrent neural network (RNN), or any other form of deep neural network. The photos and / or videos of the football game may include multiple video frames in which players and goalposts are in front of a billboard and also block the billboard. Photos and / or videos of the football game for training data need to be acquired at different perspectives with different backgrounds or lighting. The multiple photos and / or videos of the football game include but are not limited to footballs, players, referees, goalposts, billboards / banners, audiences, football fields.

[0087] At block 802, data augmentation is applied to the received photos and / or videos of a soccer game (training data). Data augmentation can refer to any processing of the received photos and / or videos of a soccer game immediately following, in order to increase the diversity of the training data. For example, the training data can be flipped to obtain a mirror image, noise may be added to the training data, or the brightness of the training data can be changed. At block 803, the training data is then applied to a process running a deep learning algorithm in order to train a deep neural network at a training module.

[0088] At block 804, at least one trained deep neural network is formed. The trained deep neural network is suitable for identifying one or more target objects and one or more non-target objects respectively. One or more target objects are static objects in a live soccer game video (e.g., a billboard). One or more non-target objects are objects in front of one or more target objects in a live soccer game video (e.g., players and / or a goal). One or more non-target objects also occlude one or more target objects in a live soccer game video frame. In other embodiments, the training process can also produce a first trained deep neural network and a second trained deep neural network. The first trained deep neural network is suitable for identifying one or more target objects, and the first trained deep neural network is suitable for identifying one or more non-target objects.

[0089] The trained deep neural network will be stored in the memory of the smart device 100 and will be used together with the application software or software program installed in the smart device 100. When the application software or software program receives a live soccer game video, the trained deep neural network is applied to the received live soccer game video in order to identify one or more target objects and one or more non-target objects in real time or near real time.

[0090] Alternatively, the server can execute process 800 entirely or can execute process 800 partially. For example, the server is allowed to execute blocks 801 to 804. The server then transmits the trained deep neural network to one or more other electronic devices (e.g., a desktop computer, a laptop computer, a smart device, or a television) to identify target objects and non-target objects.

[0091] For exemplary purposes only, video streaming or broadcasting contains some content that may not be suitable for every viewer, may not be understood by every viewer, or may not appeal to every viewer. Figure 9A A screenshot depicting an example of a video stream or broadcast displayed on an electronic device. In some examples, Figure 4A first user watches a video on the touch-sensitive display 402 of the smart device 400. The video can be a live video or a recorded video. There is no restriction on the source of the video. The video can be provided by a TV company, an online video sharing platform, an online social media network, or any other video producer / video sharing platform. For example, the first user watches a video from an online video sharing platform. The video includes multiple video frames. As Figure 9A depicted, view 960A is shown on the touch-sensitive display 402 and includes one or more target objects in the multiple video frames that the smart device 400 is trained to identify through deep learning. In some examples, a sign / billboard located at a building is regarded as a target object. The smart device 400 includes at least one training module where at least one deep neural network (for identifying the sign / billboard) is trained by feeding multiple photos and videos containing the sign / billboard located at the building. The trained deep neural network will be stored in the smart device 400. Based on the trained deep neural network, the smart device 400 is capable of identifying a first sign 982 and a second sign 984 located at the building as target objects. For objects other than the target objects, the smart device 400 will regard them as non-target objects.

[0092] View 960A includes target objects (e.g., the first sign 982 and the second sign 984) and non-target objects (e.g., buildings 962 and 964 and vehicles 966 and 968). The first sign 982 contains advertisement content associated with a Japanese electronics manufacturer, and the second sign 984 contains advertisement content associated with a Japanese bookstore. The smart device 400 includes a trained deep neural network, and the smart device 400 is capable of identifying the sign / billboard (target object) in the multiple video frames through the trained deep neural network. The smart device 400 will then perform one or more of the above processes.

[0093] Figure 9B Depicts a screenshot of an example of a processed video generated by a user covering a predetermined image on the Figure 9A video frames based on personal information. As Figure 9B depicted, by performing the above process, view 960B is shown on the display 402 and includes a first predetermined graphic image 986 and a second predetermined graphic image 988 appropriately covered on signs 982 and 984 respectively based on the personal information of the first user.

[0094] The first predetermined graphic image 986 includes first predetermined advertising content related to a luxury car brand, and the second predetermined graphic image 988 includes second predetermined advertising content related to a luxury watch brand. A second graphic image layer containing the first predetermined graphic image 986 and the second predetermined graphic image 988 is overlaid on a second target object layer containing billboards 982 and 984. A second non-target layer containing non-target objects (such as buildings 962 and 964 and vehicles 966 and 968) is overlaid on the second graphic image layer. A processed video is formed by overlaying multiple layers in multiple video frames in real time or near real time.

[0095] Figure 10A It is a screenshot of another example of video streaming or broadcasting containing one or more target objects. In one embodiment, the intelligent device 1000 is trained to identify one or more target objects through deep learning. The target object is an airplane 1090 (in Airline A) in a video (the video can be a live video or a recorded video). The intelligent device 400 includes at least one trained deep neural network associated with the target object in the memory. Figure 4 A first user of [] uses the intelligent device 400 to enjoy video streaming or broadcasting. For example, the first user watches a video from an online video sharing platform. The video includes multiple video frames. As Figure 10A depicted, view 1060A includes a target object (airplane 1090) and other non-target objects, such as buildings 1062 and 1064, vehicles 1066 and 1068, billboards / advertising boards 1082 and 1084. In some examples, the airplane is regarded as a target object. The intelligent device 400 includes at least one training module, at which at least one deep neural network (for identifying airplanes) is trained by feeding multiple photos and multiple videos containing airplanes. The trained deep neural network will be stored in the intelligent device 400. Based on the trained deep neural network, the intelligent device 400 can identify the airplane 1090 in the sky as a target object. For objects other than the target object, the intelligent device 400 will regard them as non-target objects.

[0096] The intelligent device 400 includes a trained deep neural network, and the intelligent device 400 can identify the airplane 1090 in multiple live video frames through the trained deep neural network. The intelligent device 400 will then perform one or more of the above processes.

[0097] Figure 10B Depicts a screenshot of an example of a processed video generated by overlaying a predetermined image on Figure 10A the live video frames of Figure 10BAs depicted, view 1060B includes a predetermined graphic image 1092 that is overlaid on a target object (airplane 1090) and non-target objects by performing the above-described process. The predetermined graphic image 1092 includes first predetermined advertisement content related to Airline B. A third graphic image layer containing the predetermined graphic image 1092 is overlaid on a third target object layer containing the airplane 1090. A third non-target layer containing non-target objects (such as buildings 1062 and 1064, vehicles 1066 and 1068, billboards / signs 1082 and 1084) is overlaid on the third graphic image layer. A processed video is formed by overlaying the multiple layers in multiple video frames in real-time or near real-time.

[0098] In a variant, the target object is replaced by a predetermined graphic image having the same nature as the target object. Figure 10C Depicting a screenshot of an example of a processed video generated from appropriately overlaying a predetermined image on Figure 10A live video frames. As Figure 10C depicted, view 1060C includes a predetermined graphic image 1094 (including an airplane of Airline B) that is appropriately overlaid on a target object (airplane 1090 of Airline A) and non-target objects by performing the above-described process. A fourth graphic image layer containing the predetermined graphic image 1094 is overlaid on a fourth target object layer containing the airplane 1090. A fourth non-target layer containing non-target objects (such as buildings 1062 and 1064, vehicles 1066 and 1068, billboards / signs 1082 and 1084) is overlaid on the fourth graphic image layer. A processed video is formed (as if an airplane of Airline B appears in the video stream / broadcast) by overlaying the multiple layers in multiple video frames in real-time or near real-time.

[0099] Now referring to Figure 11 , components of an exemplary computing system 1100 configured to perform any one of the above-described processes and / or operations are depicted. For example, the computing system 1100 can be used to implement the above-described intelligent device 100, which implements any combination of the above embodiments or the processes 700 and 800 described with respect to Figure 7 and Figure 8 . The computing system 1100 can include, for example, a processor, a memory, a storage device, and input / output peripheral devices (such as a display, a keyboard, a stylus, a drawing device, a disk drive, an Internet connection, a camera / scanner, a microphone, a speaker, etc.). However, the computing system 1100 can include circuitry or other dedicated hardware for performing some or all aspects of the process.

[0100] In computing system 1100, the main system 1102 may include a motherboard 1104, such as a printed circuit board with components mounted thereon, having a bus connecting an input / output (I / O) section 1106, one or more microprocessors 1108, and a memory section 1110, and the memory section may have a flash card 1138 associated therewith. The memory section 1110 may contain computer-executable instructions and / or data for executing any of processes 700 and 800 or other processes described herein. The I / O section 1106 may be connected to a display 1112 (e.g., to display views), a touch-sensitive surface 1114 (to receive touch inputs and in some cases may be combined with the display), a microphone 1116 (e.g., to obtain audio recordings), a speaker 1118 (e.g., to play audio recordings), a disk storage unit 1120, a media drive unit 1122. The media drive unit 1122 may read / write a non-transitory computer-readable storage medium 1124, which may contain a program 1126 and / or data for implementing processes 700 and 800 or any of the other processes described above.

[0101] Additionally, a non-transitory computer-readable storage medium may be used to store (e.g., tangibly embody) one or more computer programs for execution by a computer of any of the processes described above. The computer programs may be written, for example, in a general-purpose programming language (e.g., Pascal, C, C++, Java, etc.) or some dedicated application-specific language.

[0102] The computing system 1100 may include various sensors, such as a front camera 1128 and a rear camera 1130. These cameras may be configured to capture various types of light, such as visible light, infrared light, and / or ultraviolet light. Additionally, the cameras may be configured to capture or generate depth information based on the light they receive. In some cases, the depth information may be generated from a sensor different from the cameras, but may still be combined or integrated with the image data from the cameras. Other sensors or input devices included in the computing system 1100 include a digital compass 972, an accelerometer 1134, and a gyroscope 1136. Other sensors and / or output devices (e.g., a dot projector, an IR sensor, a photodiode sensor, a time-of-flight sensor, etc.) may also be included.

[0103] Although the various components of the computing system 1100 are depicted separately in FIG. 9, the various components may be combined together. For example, the display 1112 and the touch-sensitive surface 1114 may be combined together into a touch-sensitive display.

[0104] In one variant, the computing system 1100 may be used to implement the above-described server, which implements the above embodiments or with respect to Figure 7 andFigure 8 Any combination of the described processes 700 and 800. The server may include, for example, a processor, processors, storage devices, and input / output peripheral devices. In the server, the main system 1102 may include a motherboard 1104, such as a printed circuit board with components mounted thereon, having a bus connecting an input / output (I / O) section 1106, one or more microprocessors 1108, and a memory section 1110, and the memory section may have a flash card 1138 associated therewith. The memory section 1110 may contain computer-executable instructions and / or data for executing any of the processes 700 and 800 or other processes described herein. The media drive unit 1122 may read / write a non-transitory computer-readable storage medium 1124, which may contain a program 1126 and / or data for implementing the processes 700 and 800 or any of the other processes described above.

[0105] Additionally, a non-transitory computer-readable storage medium may be used to store (e.g., tangibly embody) one or more computer programs for performing any of the processes described above by a computer. The computer programs may be written, for example, in a general-purpose programming language (e.g., Pascal, C, C++, Java, etc.) or some dedicated application-specific language.

[0106] Various exemplary embodiments are described herein. These examples are referred to in a non-limiting sense. They are provided to illustrate more broadly applicable aspects of the disclosed invention. Various changes can be made and equivalents can be substituted without departing from the true spirit and scope of the various embodiments. Additionally, many modifications can be made to adapt a particular situation, material, composition of matter, process, process act, or step to the objectives, spirit, or scope of the various embodiments. Further, as will be understood by those skilled in the art, each individual variation described and illustrated herein has discrete components and features that can be readily separated from or combined with the features of any other several embodiments without departing from the scope or spirit of the various embodiments.

[0107] It should also be noted that embodiments may be described as a process, and the process is depicted as a flowchart, job diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many operations may be performed in parallel or simultaneously. Additionally, the order of the operations may be rearranged. The process terminates when its operations are completed, but may have additional steps not included in the figure. The process may correspond to a method, function, procedure, subroutine, subprogram, etc. When the process corresponds to a function, its termination corresponds to the function returning to the calling function or the main function.

Claims

1. A method for enhancing the live audience experience on an electronic device, which comprises: Receiving, by the electronic device, a plurality of live video frames; Identifying, by at least one trained deep neural network, one or more target objects and one or more non-target objects in a first live video frame among the plurality of live video frames; Identifying one or more sets of pixels belonging to the one or more target objects; Defining, based on the identified one or more sets of pixels belonging to the one or more target objects, a region on the surface of the one or more target objects; Overlaying, in the plurality of live video frames, one or more predetermined graphic images on the region on the surface of the one or more target objects; Overlaying, in the plurality of live video frames, the one or more non-target objects on the one or more predetermined graphic images to form a processed live video, wherein the processed live video includes the one or more non-target objects and the one or more predetermined graphic images overlaid on the one or more target objects; Tracking the one or more target objects by a video object tracking algorithm; Wherein, once a target object in a live video frame is identified by a trained deep neural network, the target object is tracked by using a video object tracking algorithm. For subsequent live video frames among the plurality of live video frames, the video object tracking algorithm is used to identify the tracked target object. And when a new target object appears in a subsequent live football game video frame, the trained deep neural network keeps identifying the new target object.

2. The method according to claim 1, wherein the one or more target objects include one or more static objects.

3. The method according to claim 2, wherein the one or more non-target objects include one or more objects in front of the one or more static objects, and the one or more objects occlude the one or more static objects.

4. The method according to claim 3, wherein the one or more static objects include one or more billboards.

5. The method according to claim 1, which further comprises: Scanning the first live video frame among the plurality of live video frames in a predetermined order to identify the one or more sets of pixels belonging to the one or more target objects.

6. The method according to claim 5, which further comprises: Identifying one or more extreme points corresponding to each of the identified one or more sets of pixels belonging to the one or more target objects; Applying at least one mathematical function to the identified one or more extreme points to form one or more lines.

7. The method according to claim 6, which further generates a delimiting member based on the one or more lines generated from the at least one mathematical function, wherein the delimiting member is substantially aligned with the true boundary of the one or more target objects and defines the region.

8. The method according to claim 6, wherein the at least one mathematical function is linear regression.

9. The method according to claim 1, further comprising determining 3D visual features of the one or more target objects.

10. The method according to claim 1, further comprising displaying the processed live video on a display of the electronic device or a display of another electronic device in real time or near real time.

11. The method according to claim 1, wherein the at least one trained deep neural network comprises a convolutional neural network (CNN) or a variant of a CNN, and / or is combined with a recurrent neural network (RNN).

12. A computer-readable storage medium storing one or more programs, the one or more programs comprising instructions which, when executed by an electronic device having a display, cause the device to perform any one of the methods according to claims 1 to 11.

13. An electronic device, which comprises: one or more processors; at least one display; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising instructions for performing any one of the methods according to claims 1 to 11.

Citation Information

Patent Citations

  • Interactive Video Insertions, And Applications Thereof

    US20100050082A1

  • Identifying visual objects depicted in video data using video fingerprinting

    US20180082125A1