Video screening using a machine learning video screening model trained using self-supervised learning

A self-supervised learning-based machine learning video screening model enhances the detection of restricted content in uploaded videos by expanding and processing training data, addressing the limitations of existing systems in accuracy and efficiency.

JP7813365B2Active Publication Date: 2026-02-12GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024532294
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-29
Filing Date
2022-08-23
Publication Date
2026-02-12
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

Existing content storage and distribution systems face challenges in accurately and efficiently screening uploaded videos for restricted or access-controlled content, as the accuracy and efficiency of video screening using machine learning models can be compromised by the size and quality of training data.

Method used

A machine learning video screening model trained using self-supervised learning is employed, which automatically generates training data by expanding screening data temporally and processing it to include false negatives, improving the training process and enhancing the model's ability to detect restricted content.

Benefits of technology

The self-supervised training approach significantly increases the cardinality of the training dataset, leading to improved accuracy and efficiency in detecting and flagging videos that embed access-controlled content, thereby enhancing the system's ability to manage restricted content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813365000001
    Figure 0007813365000001
  • Figure 0007813365000002
    Figure 0007813365000002
  • Figure 0007813365000003
    Figure 0007813365000003
Patent Text Reader

Abstract

A method for screening video content using a trained video screening model trained using self-supervised training includes automatically generating a training dataset by obtaining predicate screening data indicative of a predicate time segment in a training video and a corresponding reference time segment in a reference video, and obtaining candidate screening data for an extended time segment from the training video, the extended time segment including the predicate time segment and at least one frame from the training video adjacent to the predicate time segment. The candidate screening data indicates a similarity between a screening frame from the reference video and a spatial portion of a candidate frame from the extended time segment. In response to determining the determined similarity between the candidate subframes, the automatically generated training dataset is trained to include example data indicating the similarity between the candidate subframes and the screening frame.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Digital images and videos can be hosted or stored, for example, on a server or content storage and distribution system using electronic communication over an electronic communication network such as the Internet. Client devices that can be operated by users can upload images and videos to the content storage and distribution system and can access images and videos stored by the content storage and distribution system. Summary of the Invention

[0002] Disclosed herein are aspects of systems, methods, and apparatus for video screening using machine learning video screening models trained using self-supervised learning.

[0003] One aspect is a method for video screening using a machine learning video screening model trained using self-supervised training. Video screening using the machine learning screening model trained using self-supervised training may include screening a current video in response to automatically identified screening data obtained from a trained video screening model trained using self-supervised training, the screening data indicating similarity between the current video and a reference video. The self-supervised training includes obtaining automatically generated predicate screening data indicating a predicate temporal segment in a training video and a corresponding reference temporal segment in a reference video; obtaining candidate screening data for an extended temporal segment from the training video, wherein the extended temporal segment includes the predicate temporal segment and at least one frame from the training video adjacent to the predicate time segment, the candidate screening data indicating similarity between a screening frame from the reference video and a candidate sub-frame, the candidate sub-frame being a spatial portion of the candidate frame from the extended temporal segment; and obtaining a trained video screening model by training an untrained video screening model using an automatically generated training dataset by including training example data indicating similarity between the candidate sub-frame and the screening frame in the automatically generated training dataset in response to determining that the determined similarity between the candidate sub-frame and the screening frame is equal to or greater than a defined similarity threshold.

[0004] Another aspect is a method for video screening using a machine learning video screening model trained using self-supervised training. Video screening using the machine learning video screening model trained using self-supervised training includes obtaining an input video, obtaining screening data from a trained video screening model trained using self-supervised training, the screening data indicating automatically identified associations between the input video and a reference video, and identifying the input video as a screened video in response to obtaining the screening data. The self-supervised training includes obtaining an automatically generated training dataset. Obtaining the automatically generated training dataset may include obtaining a training video, obtaining a reference video, obtaining predicate screening data generated using a first previously trained video screening model with respect to the training video and the reference video, wherein the predicate screening data indicates a predicate time segment in the training video and a corresponding reference time segment in the reference video, and obtaining candidate screening data for an extended time segment from the training video from a second previously trained video screening model, wherein the extended time segment includes the predicate time segment and at least one of a frame from the training video preceding the predicate time segment or a frame from the training video following the predicate time segment, wherein the candidate screening data indicates an association between a screening frame from the reference video and a candidate sub-frame, and the candidate sub-frame is a spatial portion of a candidate frame from the extended time segment.Obtaining the automatically generated training dataset includes including training example data indicating an association between the candidate subframe and the screening frame in the automatically generated training dataset in response to determining that the determined similarity value for the candidate subframe and the screening frame is equal to or greater than a first predefined similarity threshold, in response to determining that data indicating an association between the candidate subframe and the screening frame is not present in the filtered screening data obtained from the first previously trained video screening model for the training video and the reference video, and in response to determining that the similarity value for the screening frame and the spatial portion of the screening frame is less than a second predefined similarity threshold. Self-supervised training includes debiasing the automatically generated training dataset, obtaining an untrained video screening model, and obtaining a trained video screening model by training the untrained screening model using the automatically generated training dataset.

[0005] Another aspect is a system for video screening using a machine learning video screening model trained using self-supervised training. The system may include a non-transitory computer-readable storage medium storing instructions for self-supervised training, and a processor configured to execute the instructions stored on the non-transitory computer-readable storage medium to obtain a trained video screening model, the processor executing the instructions to train an untrained video screening model using a training dataset to obtain the trained video screening model. To automatically generate the training data set, the processor executes instructions for obtaining automatically generated predicate screening data indicating a predicate time segment in the training video and a corresponding reference time segment in the reference video; instructions for obtaining candidate screening data for an extended time segment from the training video, wherein the extended time segment includes the predicate time segment and at least one frame from the training video adjacent to the predicate time segment, the candidate screening data indicating similarity between a screening frame from the reference video and a candidate sub-frame, the candidate sub-frame being a spatial portion of a candidate frame from the extended time segment; and instructions for including, in response to determining that the determined similarity between the candidate sub-frame and the screening frame is equal to or greater than a defined similarity threshold, in the automatically generated training data set.

[0006] Variations on these and other aspects are described in more detail below.

[0007] This description refers to the accompanying drawings, in which like reference numerals refer to like parts throughout the several views unless otherwise stated or apparent from the context. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram of a computing device. [Figure 2] FIG. 1 is a diagram of a computing and communication system. [Figure 3] FIG. 1 is a diagram of a video stream. [Figure 4] FIG. 1 is a block diagram of a video hosting system. [Figure 5] FIG. 1 is a block diagram of an example video similarity engine. [Figure 6] FIG. 2 is a block diagram illustrating an example of a video frame including spatial sub-frames that include restricted content. [Figure 7] FIG. 1 is a block diagram of an example method for video screening using a machine learning video screening model trained using self-supervised training. [Figure 8] FIG. 1 is a block diagram of an example of a method for obtaining a trained machine learning video screening model trained using self-supervised training. [Figure 9] FIG. 1 is a block diagram of an example method for obtaining an automatically generated training data set. [Figure 10] 10 is a diagram of an example of a graphical representation of predicate screening data 1000 for a reference video and an input video or probe video. [Figure 11] FIG. 1 is a block diagram of another example method of video screening using a machine learning video screening model trained using self-supervised training. DETAILED DESCRIPTION OF THE INVENTION

[0009] A server or content storage and distribution system may host thousands, millions, or billions of videos uploaded or otherwise provided to the server or content storage and distribution system. The server or content storage and distribution system may include or have access to a repository, database, data store, or other collection of videos defined or described as restricted, protected, or access-controlled videos, including content associated with copyright protection or other content for which access restrictions are defined. The server or content storage and distribution system may restrict, limit, or otherwise control access to videos defined or described as restricted, protected, or access-controlled videos. The server or content storage and distribution system may screen uploaded videos to detect restricted content contained in the uploaded videos. For example, the server or content storage and distribution system may generate fingerprint data representative of each restricted-access video, and in response to determining the fingerprint data representative of each uploaded video, the server or content storage and distribution system may generate fingerprint data representative of each uploaded video.

[0010] A portion of a video uploaded to or hosted by a server or content storage and distribution system may include one or more portions of restricted, protected, or access-controlled content. For example, a video uploaded to a server or content storage and distribution system may include an access-controlled video or portion thereof embedded or otherwise included in a spatial portion of the uploaded video, and may include other content in other spatial portions of the uploaded video. In some content storage and distribution systems, the accuracy, efficiency, or both of controlling access to the access-restricted content may be circumvented or reduced by including the access-restricted content in another video, such as the uploaded video.

[0011] The content storage and distribution systems described herein may screen uploaded or hosted videos using a trained machine learning video screening model that may detect and flag uploaded or hosted videos that embed access-controlled content. The accuracy, efficiency, or both of video screening using a trained machine learning video screening model may be correlated with the size and quality of the training data used to train the machine learning video screening model. Video screening using a machine learning video screening model trained using self-supervised training described herein improves the training used for other models by automatically generating training data. For example, the cardinality, or number of training examples, of an automatically generated training dataset may be significantly larger than a manually generated training dataset. The automatic or self-supervised generation of training data is based on a previously trained machine learning video screening model. To further improve video screening using a machine learning video screening model trained using self-supervised training, the screening data generated by the previously trained machine learning video screening model is temporally expanded and further processed as described herein so that training examples that are false negatives for the previously trained machine learning video screening model are included as training examples.

[0012] 1 is a block diagram of an example computing device 100. The illustrated computing device 100 includes memory 110, a processor 120, a user interface (UI) 130, an electronic communication unit 140, sensors 150, a power supply 160, and a bus 170. As used herein, the term "computing device" includes any unit or combination of units capable of performing any of the methods disclosed herein, or any one or more portions thereof.

[0013] Computing device 100 may be a fixed computing device, such as a personal computer (PC), server, workstation, minicomputer, or mainframe computer, or a mobile computing device, such as a mobile phone, personal digital assistant (PDA), laptop, or tablet PC. Although shown as a single unit, any one or more elements of computing device 100 may be integrated into any number of separate physical units. For example, user interface 130 and processor 120 may be integrated into a first physical unit, and memory 110 may be integrated into a second physical unit.

[0014] Memory 110 may include any non-transitory computer-usable or computer-readable medium, such as any tangible device that can contain, store, communicate, or transfer, for example, data 112, instructions 114, operating system 116, or any information associated therewith, for use by or in connection with other components of computing device 100. The non-transitory computer-usable or computer-readable medium may be, for example, a solid-state drive, a memory card, removable media, read-only memory (ROM), random-access memory (RAM), a hard disk, a floppy disk, an optical disk, any type of disk including a magnetic or optical card, an application-specific integrated circuit (ASIC), or any type of non-transitory memory or medium suitable for storing electronic information, or any combination thereof.

[0015] Although shown as a single unit, memory 110 may include multiple physical units, such as one or more primary memory units, such as random access memory units, one or more secondary data storage units, such as disks, or a combination thereof. For example, data 112 or portions thereof, instructions 114 or portions thereof, or both, may be stored in a secondary storage unit and loaded or otherwise transferred to a primary storage unit in conjunction with processing the respective data 112, executing the respective instructions 114, or both. In some embodiments, memory 110, or portions thereof, may be removable memory.

[0016] Data 112 may be or include input data, encoded data, decoded data, etc. Instructions 114 may include instructions, such as code, for performing the methods disclosed herein or any part or parts thereof. Instructions 114 may be realized in hardware, software, or any combination thereof. For example, instructions 114 may be implemented as information stored in memory 110, such as a computer program or application, that may be executed by processor 120 to perform any of the respective methods, algorithms, aspects, or combinations thereof described herein.

[0017] Although shown as contained in memory 110, in some implementations, the instructions 114, or portions thereof, may be implemented as a special-purpose processor or circuitry, which may include dedicated hardware for executing any of the methods, algorithms, aspects, or combinations thereof described herein. Portions of the instructions 114 may be distributed across multiple processors on the same machine or different machines, or distributed across a network, such as a local area network, a wide area network, the Internet, or combinations thereof.

[0018] Processor 120 may include any device or system now existing or later developed that is capable of manipulating or processing digital signals or other electronic information, including an optical processor, a quantum processor, a molecular processor, or a combination thereof. For example, processor 120 may include a special-purpose processor, a central processing unit (CPU), a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic array, a programmable logic controller, microcode, firmware, any type of integrated circuit (IC), a state machine, and / or any combination thereof. As used herein, the term "processor" includes a single processor or multiple processors.

[0019] User interface 130 may include any unit capable of interfacing with a user, such as a virtual or physical keypad, a touchpad, a display, a touch display, a speaker, a microphone, a video camera, a sensor, or any combination thereof. For example, user interface 130 may be an audiovisual display device, and computing device 100 may use user interface 130, the audiovisual display device, to present audio, such as decoded audio, e.g., in conjunction with the display of video, e.g., decoded video. Although shown as a single unit, user interface 130 may include one or more physical units. For example, user interface 130 may include an audio interface for performing audio communication with the user and a touch display for performing visual and touch-based communication with the user.

[0020] The electronic communication unit 140 can transmit, receive, or transmit and receive signals via a wired or wireless electronic communication medium 180, such as a radio frequency (RF) communication medium, an ultraviolet (UV) communication medium, a visible light communication medium, an optical fiber communication medium, a wired communication medium, or a combination thereof. For example, as shown, the electronic communication unit 140 is operably connected to an electronic communication interface 142, such as an antenna, configured to communicate via wireless signals.

[0021] 1 as a wireless antenna, electronic communication interface 142 may be a wireless antenna as shown, a wired communication port such as an Ethernet port, an infrared port, a serial port, or any other wired or wireless unit capable of interfacing with a wired or wireless electronic communication medium 180. Although FIG. 1 shows a single electronic communication unit 140 and a single electronic communication interface 142, any number of electronic communication units and any number of electronic communication interfaces may be used.

[0022] Sensor 150 may include, for example, an audio sensing device, a visible light sensing device, a motion sensing device, or a combination thereof. For example, sensor 150 may include a sound sensing device such as a microphone, or any other now-existing or later-developed sound sensing device that can detect sound in the vicinity of computing device 100, such as speech or other utterances made by a user operating computing device 100. In another example, sensor 150 may include a camera or other now-existing or later-developed image sensing device that can detect images, such as images of a user operating a computing device. Although a single sensor 150 is shown, computing device 100 may include several sensors 150. For example, computing device 100 may include a first camera oriented with a field of view directed toward the user of computing device 100 and a second camera oriented with a field of view directed away from the user of computing device 100.

[0023] Power supply 160 may be any suitable device for providing power to computing device 100. For example, power supply 160 may include a wired external power interface, one or more dry batteries, such as nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc., a solar cell, a fuel cell, or any other device capable of providing power to computing device 100. Although a single power supply 160 is shown in FIG. 1 , computing device 100 may include multiple power sources 160, such as batteries and wired external power interfaces.

[0024] Although shown as separate units, electronics communication unit 140, electronics communication interface 142, user interface 130, power supply 160, or portions thereof, may be configured as a combined unit. For example, electronics communication unit 140, electronics communication interface 142, user interface 130, and power supply 160 may be implemented as a communication port that can interface with an external display device and provide communication, power, or both.

[0025] One or more of memory 110, processor 120, user interface 130, electronic communication unit 140, sensor 150, or power supply 160 may be operably coupled via bus 170. Although a single bus 170 is shown in FIG. 1 , computing device 100 may include multiple buses. For example, memory 110, processor 120, user interface 130, electronic communication unit 140, sensor 150, and bus 170 may receive power from power supply 160 via bus 170. In another example, memory 110, processor 120, user interface 130, electronic communication unit 140, sensor 150, power supply 160, or combinations thereof may communicate data, such as by sending and receiving electronic signals, via bus 170.

[0026] 1, one or more of the processor 120, the user interface 130, the electronic communication unit 140, the sensor 150, or the power supply 160 may include internal memory, such as an internal buffer or register. For example, the processor 120 may include internal memory (not shown) and may read data 112 from the memory 110 into the internal memory (not shown) for processing.

[0027] Although shown as separate elements, the memory 110, processor 120, user interface 130, electronic communication unit 140, sensors 150, power supply 160, and bus 170, or any combination thereof, may be integrated into one or more electronic units, circuits, or chips.

[0028] 2 is a diagram of a computing and communication system 200. The illustrated computing and communication system 200 includes computing and communication devices 100A, 100B, and 100C, access points 210A and 210B, and a network 220. For example, the computing and communication system 200 may be a multiple-access system that provides communications, such as voice, audio, data, video, messaging, broadcast, or combinations thereof, to one or more wired or wireless communication devices, such as computing and communication devices 100A, 100B, and 100C. For simplicity, FIG. 2 depicts three computing and communication devices 100A, 100B, and 100C, two access points 210A and 210B, and one network 220, although any number of computing and communication devices, access points, and networks may be used.

[0029] Computing and communication devices 100A, 100B, 100C may be computing devices such as computing device 100 shown in FIG. 1. For example, computing and communication devices 100A, 100B may be user devices such as mobile computing devices, laptops, thin clients, or smartphones, and computing and communication device 100C may be a server such as a mainframe or cluster. Although computing and communication devices 100A and 100B are described as user devices and computing and communication device 100C is described as a server, either computing and communication device may perform some or all of the functions of a server, some or all of the functions of a user device, or some or all of the functions of a server and a user device. For example, server computing and communications device 100C may receive data such as audio or video data and may process it, such as by encoding, processing, storing, transmitting, or a combination thereof, and one or both of computing and communications device 100A and computing and communications device 100B may receive data and may process it, such as by decoding, processing, storing, presenting, or a combination thereof.

[0030] Each computing and communication device 100A, 100B, 100C, which may include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a personal computer, a tablet computer, a server, a consumer electronics appliance, or any similar device, may be configured to perform wired or wireless communication, such as over network 220. For example, computing and communication device 100A, 100B, 100C may be configured to transmit or receive wired or wireless communication signals. Although each computing and communication device 100A, 100B, 100C is shown as a single unit, the computing and communication device may include any number of interconnected elements.

[0031] Each access point 210A, 210B may be any type of device configured to communicate with computing and communication devices 100A, 100B, 100C, network 220, or both via wired or wireless communication links 180A, 180B, 180C. Access points 210A, 210B may include a base station, a base transceiver station (BTS), a Node B, an enhanced Node B (eNode-B), a Home Node B (HNode-B), a wireless router, a wired router, a hub, a repeater, a switch, or any similar wired or wireless device. For example, although each access point 210A, 210B is shown as a single unit, an access point may include any number of interconnected elements.

[0032] Network 220 may be any type of network configured to provide services such as voice, data, applications, Voice over Internet Protocol (VoIP), or any other communication protocol or combination of communication protocols over wired or wireless communication links. For example, network 220 may be a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a mobile or cellular telephone network, the Internet, or any other means of electronic communication. The network may use communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Internet Protocol (IP), Real-time Transport Protocol (RTP), Hypertext Transport Protocol (HTTP), or a combination thereof.

[0033] Computing and communication devices 100A, 100B, 100C can communicate with each other over network 220 using one or more wired or wireless communication links, or a combination of wired and wireless communication links. For example, as shown, computing and communication devices 100A, 100B can communicate over wireless communication links 180A, 180B, and computing and communication device 100C can communicate over wired communication link 180C. Any of computing and communication devices 100A, 100B, 100C may communicate using any one or more wired or wireless communication links. For example, a first computing and communication device 100A can communicate over a first access point 210A using a first type of communication link, a second computing and communication device 100B can communicate over a second access point 210B using a second type of communication link, and a third computing and communication device 100C can communicate over a third access point (not shown) using a third type of communication link. Similarly, the access points 210A, 210B can communicate with the network 220 via one or more types of wired or wireless communication links 230A, 230B. Although Figure 2 shows the computing and communication devices 100A, 100B, 100C communicating via the network 220, the computing and communication devices 100A, 100B, 100C can communicate with each other via any number of communication links, such as direct wired or wireless communication links.

[0034] In some implementations, communication between one or more of computing and communication devices 100A, 100B, 100C may omit communicating over network 220 and may include transferring data via another medium (not shown), such as a data storage device. For example, server computing and communication device 100C may store data, such as encoded data, on a data storage device, such as a portable data storage unit, and one or both of computing and communication device 100A or computing and communication device 100B may access, read, or retrieve the stored audio data from the data storage unit by physically disconnecting the data storage device from server computing and communication device 100C and physically connecting the data storage device to computing and communication device 100A or computing and communication device 100B.

[0035] Other implementations of the computing and communications device 200 are possible. For example, in an implementation, the network 220 may be an ad-hoc network, and one or more of the access points 210A, 210B may be omitted. The computing and communications device 200 may include devices, units, or elements not shown in FIG. 2. For example, the computing and communications device 200 may include more communications devices, networks, and access points.

[0036] 3 is a diagram of a video stream 300. The video stream 300, such as a video stream captured by a video camera or a video stream generated by a computing device, may include a video sequence 310. The video sequence 310 may include a sequence of adjacent frames 320. Although three adjacent frames 320 are shown, the video sequence 310 may include any number of adjacent frames 320.

[0037] Each frame 330 from adjacent frames 320 may represent a single image from the video stream. Although not shown in Figure 3, frame 330 may include one or more segments, tiles, or planes that may be coded or otherwise processed independently, such as in parallel. Although not shown in Figure 3, a frame may include pixels. A frame, a portion of a frame, such as a block, a pixel, or a combination thereof may include display information, such as luminance information, chrominance information, or any other information that can be used to store, modify, communicate, or display a video stream or portion thereof.

[0038] Figure 4 is a block diagram of a video hosting system 400. The video hosting system 400 may be or may include a computing device such as the computing device 100 shown in Figure 1 or one or more of the computing and communication devices 100A, 100B, and 100C shown in Figure 2. As shown in Figure 4, the video hosting system 400 includes a front-end server 410, an ingest server 420, a video search server 430, a video similarity engine 440, a video access server 450, a video data store 460, and a fingerprint data store 470. In some embodiments, the video hosting system 400 may include other components not shown in Figure 4, such as a firewall, a load balancer, an application server, a failover server, and a site management tool. In some embodiments, one or more of the front-end server 410, the ingest server 420, the video search server 430, the video similarity engine 440, the video access server 450, the video data store 460, or the fingerprint data store 470 may be omitted or may not be present from the video hosting system 400.

[0039] One or more computing devices, such as the computing device 100 shown in FIG. 1, or one or more computing and communication devices, such as one or more of the computing and communication devices 100A, 100B, and 100C shown in FIG. 2, may access the video hosting system 400 to retrieve, review, or both video content. For example, the video hosting system 400 may implement a video search interface, a video browsing interface, or both, which may be a user interface, a programmatic interface, or both. The video hosting system 400 may retrieve videos, video content, or video files from uploading videos, searching or crawling other websites or video databases, etc., or any combination thereof. The video hosting system 400 can be configured to allow uploading of content (e.g., user-generated content (UGC)). The video hosting system 400 can be configured to retrieve videos from other sources by crawling such sources or searching such sources in real time.

[0040] The video hosting system 400 may be a website or may be available at a website. As used herein, the term "website" may refer to a computing device, such as the computing device 100 shown in FIG. 1, a computing and communication device, such as one or more of the computing and communication devices 100A, 100B, 100C shown in FIG. 2, or a computing system including one or more computing devices adapted to provide content using one or more electronic communication protocols, which may include, but are not limited to, content uploaded or downloaded over the Internet using the HTTP protocol.

[0041] The front-end server 410 can communicate with other computing devices, such as user devices or client devices, using electronic communications protocols over an electronic communications link, such as the wired or wireless electronic communications medium 180 shown in FIG. 1, or one or more of the wired or wireless communications links 180A, 180B, 180C shown in FIG. 2, which may include communicating over a network, such as the network 220 shown in FIG. 2, which may include the Internet.

[0042] The front-end server 410 can receive requests from external computing devices (with respect to the video hosting system 400), such as user devices. For simplicity, external computing devices in communication with the video hosting system 400 are referred to herein as client devices. The front-end server 410 can communicate with other components of the video hosting system 400, such as to process requests.

[0043] The front-end server 410 can monitor client device interactions with the video hosting device 400. For example, a client device may access a web page, upload a video, watch a video, make a purchase, or fill out a web-based form, and the front-end server 410 can monitor these interactions. The front-end server 410 can send a video and associated video links to the client device for presentation, such as on a web page. The requested video can be streamed to the client device by the front-end server 410. One or more associated video links can be presented on the web page on which the requested video is playing, so that the client can select an associated video link to watch the associated video.

[0044] Content received by the video hosting system 400, such as from a client device for posting to the video hosting system 400, is sent to or otherwise made available to an ingest server 420 for processing. Processing of video files includes assigning identifiers to received video files. Other aspects of processing video files may include formatting, transcoding, compression, metadata tagging, content analysis, or other video data processing techniques, or combinations thereof.

[0045] In one embodiment, a client device transmits metadata, such as data entered into a user interface form, in conjunction with uploading a video file to the video hosting system 400. The transmitted metadata may include data describing the video, such as a title, a description, tag data, or a combination thereof. The transmitted metadata may include an indication of the media type of the content, such as a "video" type. The ingest server 420 stores the processed video files in the video data store 460. The ingest server 420 stores the metadata for the video files.

[0046] The video data store 460 stores video files, such as video files sent to the video hosting system 400. Storing video files may include storing one or more icons or thumbnail views. Storing video files may include storing associated metadata, such as title, author, tags, description, comments, ratings, or combinations thereof. In some embodiments, the ingest server 420 may send or otherwise make available received videos to the video similarity engine 440 for analysis.

[0047] The video search server 430 can process requests received by the front-end server 410 and can identify videos relevant to the requests. The requests, which may be responsive to input, such as user input, provided to the front-end server 410 via a client device, may include a search query specifying one or more search terms. The video search server 430 may use the search terms to query metadata of video files stored in the video data store 460. The search results may include videos whose associated metadata is related to at least one of the search terms. The search results, or a subset thereof, may be sent to the front-end server 410. The front-end server 410 may send the search results, or a subset thereof, to the client device for presentation to the user.

[0048] The video access server 450 receives requests for the videos to be shown from the front-end server 410, such as from respective client devices. The requests for the videos may be received from user devices in response to browsing for videos, such as through respective categories in the video hosting system 400, or in response to receiving input, such as user input clicking a link to the video from a search results webpage. The request sent by the client device may include an identifier for the video. The video access server 450 can use the identifier to locate the video within the video data store 460. The video access server 450 can provide the requested video to the front-end server 410. The front-end server 410 transmits or makes available the video to the client device, such as via streaming.

[0049] The video similarity engine 440 can determine whether a video, such as an uploaded video, contains video content of one or more other videos, such as other videos that are copyrighted, have restricted access, etc. (protected videos). In response to determining that the uploaded video is similar to the protected video, such as within a defined similarity threshold, the video similarity engine 440 can flag the video as unauthorized, etc., or remove the video from the video hosting system 400. The video similarity engine 440 can process the video substantially simultaneously as the video is uploaded to the video hosting system 400. The video similarity engine 440 can process the video substantially simultaneously with the ingest server 420 processing the video.

[0050] To determine similarity, the video similarity engine 440 may create one or more fingerprints, one or more sub-fingerprints, or a combination thereof for the video. In one example, the sub-fingerprints can be generated using video content that includes motion. The sub-fingerprints represent respective portions of video content included in the video. The sub-fingerprints can be used to determine whether a video contains video content that has been copied or partially copied from another video. The video similarity engine 440 can compare the sub-fingerprints to fingerprints stored in the fingerprint data store 470.

[0051] In response to determining that the sub-fingerprint of the video sufficiently matches a fingerprint stored in fingerprint data store 470 derived from another video, video similarity engine 440 determines that the video contains video content copied from another video. A video stored in video hosting system 400 that is identified as containing content copied from another video may be deleted from video hosting system 400. In response to determining that a video to be uploaded to video hosting system 400 contains content copied from another video, the upload of the video may be terminated.

[0052] Fingerprint data store 470 stores fingerprints derived from videos corresponding to video files stored in video data store 460. The fingerprints stored in fingerprint data store 470 can be used as reference data for video similarity engine 440 to determine whether a video contains video content from one or more other videos.

[0053] Figure 5 is a block diagram of an example video similarity engine 500. The video similarity engine 500 is similar to the video similarity engine 440 shown in Figure 4, except as described herein or otherwise apparent from the context. The video similarity engine 500 may be included in a video hosting device, such as the video hosting system 400 shown in Figure 4.

[0054] 2, the video similarity engine 500 includes a fingerprint generation module 510, a sub-image generation module 520, a shot detection module 530, a sub-fingerprint generation module 540, a composite fingerprint generation module 550, and a fingerprint matching module 560. The video similarity engine 500 may include other components not shown in FIG. 5. In some embodiments, one or more of the fingerprint generation module 510, the shot detection module 530, the composite fingerprint generation module 550, the sub-fingerprint generation module 540, the fingerprint matching module 560, and the sub-image generation module 520 may be combined. In some embodiments, although shown as a unit in FIG. 4, the video similarity engine 500 may be implemented as two or more different units.

[0055] The fingerprint generation module 510 generates one or more fingerprints for an image, a sequence of images, or a video. The fingerprint generation module 510 generates a fingerprint for a time interval of the video using video frames of the video. The fingerprint can be generated based on a video frame or an uninterrupted sequence of video frames with image content continuity.

[0056] A fingerprint may be represented as a bit vector that represents spatial characteristics of a video frame, temporal characteristics of a video frame, structural characteristics of a video frame, or a combination thereof. Because a fingerprint is an identifier for a video frame based on the content of the video frame, small variations due to compression, decompression, noise, frame rate, start and stop times, resolution, etc. do not significantly affect the fingerprint.

[0057] The fingerprint generation module 510 may receive the video, or one or more frames thereof, from a component of the video hosting system, such as the front-end server 410 shown in Figure 4, the ingest server 420 shown in Figure 4, or the video data store 460 shown in Figure 4. In some embodiments, the fingerprint generation module 510 generates a fingerprint for the video simultaneously or substantially simultaneously with the processing of the video by the ingest server.

[0058] The sub-image generation module 520 generates sub-images using video frames of the video. For brevity and clarity, a video from which one or more sub-images can be generated is referred to herein as the input video, and frames of the input video are referred to herein as input frames. The sub-images are used to generate sub-fingerprints, which are used to detect whether the input video contains unauthorized content. A sub-image is a portion or region of an input frame, such as a rectangular region, that contains motion relative to one or more other frames of the video, while other portions of the input frame surrounding or adjacent to the rectangular region are static or semi-static. Video content containing motion within a sub-image has a relatively high likelihood of containing unauthorized content. The sub-image generation module 520 identifies video content containing motion and corresponding regions of each input video frame. The sub-image generation module 520 extracts, copies, or identifies image content, such as pixel values, of the identified sub-image regions from each input frame to generate sub-images.

[0059] In some embodiments, an input frame or input video may include static, semi-static content (disguised content) spatially proximate to a rectangular sub-image region, such as along a boundary or edge of the rectangular region, adjacent to or partially overlapping the rectangular sub-image region, and the corresponding sub-image may include the content of the rectangular region and omit or exclude the static or semi-static content. For example, the disguised content may be displayed as a visual frame or border around the rectangular sub-image region.

[0060] In one example, the input video may include a first portion and a second portion, where the first portion is a rectangular sub-image region that includes or may include unauthorized content, and the second portion includes static or semi-static content with respect to the input frames of the video, such as background content. For each input frame of the video, the sub-image generation module 520 may generate a respective sub-image that includes the rectangular sub-image region of the respective input frame that may include unauthorized content.

[0061] In some embodiments, an input video may contain unauthorized or potentially unauthorized content from two or more other videos, and the sub-image generation module 520 may generate respective separate sub-images corresponding to each of the other videos. For example, the input video may include a first rectangular sub-image portion corresponding to a first unauthorized video, a second rectangular sub-image portion corresponding to a second unauthorized video, and a third portion that may contain static or semi-static content; for an input frame, the sub-image generation module 520 may generate a first sub-image corresponding to the first unauthorized video and a second sub-image corresponding to a frame of the second unauthorized video. Multiple non-overlapping sub-images may be generated from an input frame. The sub-images preserve the temporal characteristics of the corresponding input frames from which they were generated.

[0062] To identify sub-images, the sub-image generation module 520 tracks the motion of the video content across multiple input frames. The sub-image generation module 520 performs motion analysis to determine relative motion between frames. Motion analysis may include comparing pixels, such as brightness, of a first video frame with spatially corresponding pixels in a subsequent video frame. Pixels whose difference between frames equals or exceeds a predefined motion threshold are identified as having motion (motion pixels). Pixels whose difference between frames is less than a predefined threshold are identified as still or static pixels. In some implementations, the motion threshold can be evaluated across multiple consecutive frames corresponding to a defined time window.

[0063] The sub-image generation module 520 generates a binary image or motion pixel map for each frame, with pixels or pixel locations in the motion pixel map having a value of one (1) for motion pixels and a value of zero (0) for static pixels.

[0064] For example, for an input video that includes regions containing unauthorized content and regions that otherwise contain static or semi-static content, such as outside the region, the corresponding motion pixel map will be a substantially rectangular region, with each pixel or pixel location, or substantially the majority of them, being a motion pixel, and other pixels or pixel locations outside the rectangular region being static pixels.

[0065] The sub-image generation module 520 uses rectangular regions where each pixel or pixel location, or a substantial majority of which, is a motion pixel to identify and extract the region as a sub-image. The sub-image generation module 520 may form the region by fitting a rectangle around the identified motion pixel such that the rectangle encompasses the identified motion pixel. In some implementations, one or more static pixels spatially proximate to the motion pixel may be included in the rectangular sub-image region. In some embodiments, a rotating caliper algorithm may be used to determine a minimum-area rectangle for the sub-image region to maximize the number, density, or percentage of motion pixels within the rectangle, minimize the number, density, or percentage of static pixels within the region, or both.

[0066] In some embodiments, a portion or region of an input frame in which a small number of pixels or pixel locations are identified as motion pixels and a large number of pixels or pixel locations are identified as static pixels may be identified as containing static or semi-static content, and the sub-image generation module 520 may omit extracting the portion as a sub-image.

[0067] For example, the sub-image generation module 520 may determine a ratio of still pixels to motion pixels in a defined portion of the input frame and compare the determined ratio to a threshold ratio to determine whether to identify a region for extraction as a sub-image. The determined ratio may be equal to or greater than the threshold ratio, and the sub-image generation module 520 may identify the region as a region to be extracted as a sub-image. The determined ratio may be less than the threshold ratio, and the sub-image generation module 520 may omit identifying the region as a region to be extracted as a sub-image. In some embodiments, the sub-image generation module 520 may omit identifying regions as regions to be extracted as a sub-image that have a size, such as the number or density of pixels, below a defined minimum size.

[0068] The sub-image generation module 520 determines a sub-image identifier for each sub-image. The sub-image identifiers, or portions thereof, may be generated from successive input video frames and assigned to multiple consecutive sub-images corresponding to regions that are spatially simultaneous or substantially simultaneous, e.g., with respect to location and size. The location and size of the regions used to generate the sub-images may be determined using matrix or Cartesian notation, for example, based on the locations of pixels on the boundaries of the regions, e.g., the top left and bottom right pixels relative to the input frames.

[0069] The sub-image generation module 520 determines whether a sub-image identifier can be assigned to each of a plurality of regions, such as consecutive input frames, based on, for example, comparing the position and size of a first region in a first input frame with the second position and size of a second region in a second input frame. In response to determining that the difference in position between the first region in the first input frame and the second region in the second input frame is within a position difference threshold, such as less than a position difference threshold, and determining that the difference in size between the first region in the first input frame and the second region in the second input frame is within a size difference threshold, such as less than a size difference threshold, the sub-image generation module 520 determines that the first region in the first input frame and the second region in the second input frame have the same or substantially the same position and size, and assigns a sub-image identifier, or a portion thereof, to the sub-image generated from the first input frame and the sub-image generated from the second input frame. In response to determining that the difference in position between the first region of the first input frame and the second region of the second input frame is greater than or equal to the position difference threshold, or determining that the difference in size between the first region of the first input frame and the second region of the second input frame is greater than or equal to the size difference threshold, the sub-image generation module 520 determines that the first region of the first input frame and the second region of the second input frame have different positions or sizes and assigns a sub-image identifier to the sub-image generated from the first input frame and assigns another, different sub-image identifier to the sub-image generated from the second input frame.

[0070] The sub-image generation module 520 may generate sub-images of the input video simultaneously or substantially simultaneously with the processing of the input video by the ingest server.

[0071] The shot detection module 530 identifies a sequence of consecutive sub-images as a shot that can be used to generate a sub-fingerprint. The shot detection module 530 analyzes characteristics of consecutive sub-images to determine the temporal location of discontinuities within the video content of the sub-images. A discontinuity may be an abrupt change from one frame to the next in sequence, which may correspond to a scene change or camera change, or a scene transition such as a fade or dissolve. A discontinuity may be identified based on one or more sub-image characteristics that are discernible from the content of consecutive sub-images. A discontinuity may be identified based on a change in sub-image identifier between sub-images. In some embodiments, the shot detection module 530 may generate shots based on the input video.

[0072] The set of sub-image shots is sent or otherwise made available to the sub-fingerprint generation module 540 for generation of sub-fingerprints. The generated sub-image shots are used to create a set of sub-fingerprints for each time interval of the video. For example, a sub-fingerprint may be generated for each time interval (T) having a defined temporal length, such as one second of the input video, from the beginning of the input video (T=0). For a temporal span from a first temporal position (nT, where n is an integer) to a subsequent second temporal position ((n+1)T) of the input video, the shot detection module 530 determines one or more shots having a first frame at or corresponding to a temporal position after the first temporal position (nT) and a last frame at or corresponding to a temporal position before the second temporal position ((n+1)T) to generate a sub-fingerprint.

[0073] If a shot is available for the time interval, an empty sub-fingerprint for the time interval may be generated, and the shot detection module 530 may notify the sub-fingerprint generation module 540 accordingly.

[0074] In another implementation, the shot detection module 530 organizes the generated shots before providing them to the sub-fingerprint generation module 540. The shot detection module 530 may group the shots by sub-image IDs associated with the sub-images included in each one or more shots. One or more shots with matching sub-image IDs are organized as a group. A sub-fingerprint can be generated using the group of shots with their respective sub-image IDs.

[0075] The subfingerprint generation module 540 generates subfingerprints for time intervals of the video using subimages generated for the video. Subfingerprints are generated using one or more subimages, or subframes, shots of subimages, or groups of shots of the time interval, for each time interval T, starting from the beginning of the video (T=0). In some implementations, for a time interval of the video, subfingerprints are generated using one or more shots of the video whose start time is at or after the start of the time interval. Each shot may be unavailable, and an empty subfingerprint may be generated for the time interval of the video. Shots may overlap multiple time intervals of the video, and a subfingerprint generated using one shot for one time interval of the video may represent the video content of a subsequent time interval of the video. An empty subfingerprint is generated for the video content of those time intervals of the represented video.

[0076] Composite fingerprint generation module 550 generates a composite fingerprint for each time interval T of the video, starting from the beginning of the video (T=0). For a time interval T of the video, the composite fingerprint is a data structure that includes or references one or more fingerprints generated for time interval T of the video and one or more sub-fingerprints generated for time interval T of the video. A composite fingerprint of a video may represent a portion of the “motion” video content for time interval T of the video. Composite fingerprint generation module 550 receives the fingerprint generated by fingerprint generation module 510 and the sub-fingerprints generated by sub-fingerprint generation module 540. The sub-fingerprints may be empty sub-fingerprints.

[0077] The fingerprint and sub-fingerprints each represent different aspects of the substantive content of the video, and the composite fingerprint represents, in a condensed form, the substantive characteristics of the video from the fingerprint and the characteristics of sub-images extracted from the video from the sub-fingerprints. The composite fingerprint can be used to determine a video that includes video content from another video, such as a video that in turn embeds content from one or more other videos.

[0078] The fingerprint matching module 560 receives the composite fingerprint and matches it with reference fingerprints from a data store associated with the reference video. The fingerprint matching module 560 matches the fingerprint of the video and sub-fingerprints of the sub-images of the video included in the composite fingerprint with the reference fingerprints. The matching result indicates that the video under consideration contains video content from one of the reference videos. The fingerprint matching module 560 may perform the matching simultaneously or partially in parallel with an ingest server that processes the video.

[0079] 6 is a block diagram illustrating an example of a video frame 600 that includes a spatial sub-frame 610 that includes restricted content 620. Video frame 600 is a frame or image, such as frame 330 shown in FIG. 3, of a video, such as video stream 300 shown in FIG.

[0080] Frame 600 may be represented, represented, encoded, or stored as a matrix, such as a two-dimensional matrix or Cartesian plane, of pixel values, and each pixel location within frame 600 may be indicated using Cartesian coordinates or other matrix notation. While described herein with reference to a matrix or Cartesian representation of a frame for clarity, the frame may be stored, transmitted, processed, or any combination thereof in any data structure that can efficiently represent the pixel values ​​for the frame or image. For example, the frame may be stored, transmitted, processed, or any combination thereof in a two-dimensional data structure, such as a matrix as shown, or in a one-dimensional data structure, such as a vector array. In implementation, a representation of a frame, such as the two-dimensional representation shown, may correspond to a physical location in a rendering of the frame as an image. For example, the location of the upper left corner of a block in the upper left corner of the frame may correspond to the physical location of the upper left corner of the rendering of the frame as an image.

[0081] The frame 600 includes spatial sub-frames, i.e., a sub-image 610, unrestricted content 630, and an obstruction 640. The spatial sub-frame 610 is a spatial portion, such as a rectangular portion, of the frame 600. The location of the spatial sub-frame 610 may be within the frame 600 as shown or another location within the frame 600. The size, e.g., height in pixels and width in pixels, of the spatial sub-frame 610 may be greater than a pixel and smaller than the frame 600, such as the size shown or another size. The spatial sub-frame 610 includes the restricted content 620. The obstruction 640 may be superimposed on or displayed over a respective portion of the unrestricted content 630, the spatial sub-frame 610, or both, as shown. Portions of the boundaries of the spatial sub-frame 610 are shown using dashed lines to indicate that the respective portions of the spatial sub-frame 610 are occluded by the obstruction 640. The size, shape, orientation, location, and number or concentration of the obstructions 640 may differ from the example shown in FIG. 6 . In some implementations, the obstructions 640 may not be present or may be omitted from the frame 600. The unrestricted content 630 differs from the restricted content 620. While one subframe 610 is shown in FIG. 6 , a frame may include multiple spatial subframes, each of which may include restricted content. For example, a frame may include a first spatial subframe including a first restricted content and a second spatial subframe including a second restricted content. While described herein as rectangular subframes, non-rectangular subframes may also be used.

[0082] In some implementations, frame 600 may be a frame from an unscreened video, and data explicitly identifying spatial sub-frame 610, data indicating that frame 600 or the unscreened video, or portions thereof, contain restricted content, or both may not be available. In some implementations, frame 600 may be a frame from a screened video, and data explicitly identifying spatial sub-frame 610, data indicating that frame 600 or the unscreened video, or portions thereof, contain restricted content, or both are available, e.g., screening data generated by video screening using a machine learning video screening model trained using self-supervised training, such as video screening using a machine learning video screening model trained using self-supervised training, is shown in FIG.

[0083] 7 is a block diagram of an example method of video screening using a machine learning video screening model trained using self-supervised training 700. Video screening using the machine learning video screening model trained using self-supervised training 700, or a portion or portions thereof, is implemented by a computing device such as computing device 100 shown in FIG. 1, one or more of computing and communication devices 100A, 100B, 100C shown in FIG. 2, a video hosting system such as video hosting system 400 shown in FIG. 4, or a component thereof, or a video similarity engine such as video similarity engine 500 shown in FIG. 5.

[0084] Video screening using machine learning video screening trained using self-supervised training 700 automatically identifies videos containing restricted content, or a portion or portions thereof, using a machine learning video screening model trained using self-supervised training. Video screening using machine learning video screening trained using self-supervised training 700 includes obtaining a trained machine learning video screening model trained using self-supervised training at 710, obtaining a current video at 720, obtaining screening data for the current video from the trained machine learning video screening model trained using self-supervised training at 730, and screening the current video at 740.

[0085] A trained machine learning video screening model, which may be an object detection model trained using self-supervised training, is obtained at 710. An example of obtaining a trained machine learning video screening model trained using self-supervised training is shown in FIG.

[0086] A current video, i.e., input video, is obtained at 720. For example, an unscreened video may be uploaded or otherwise made available to a computing device or a system including a computing device, and the current video may be obtained from the unscreened video. Although not separately shown in FIG. 7 , video screening using machine learning video screening trained using self-supervised training 700 may include obtaining an unscreened video, such as a set of unscreened videos, and obtaining a current video may include identifying an unscreened video from the unscreened videos as the current video. As used herein, the term “unscreened video” refers to a video designated or identified as unscreened, such that screening data for the unscreened video, other than screening data obtained by video screening using machine learning video screening trained using self-supervised training 700, is unavailable or unused prior to video screening using machine learning video screening trained using self-supervised training 700. The current video includes an image, i.e., a sequence of frames, such as frame 600 shown in FIG. 6 .

[0087] At 730, screening data is obtained from a trained machine learning video screening model trained using self-supervised training. The screening data includes data identifying or describing similarities between a current video and a reference video detected by the trained machine learning video screening model. The reference video is obtained from a repository, database, data store, or other collection of videos defined or described as reference videos. For example, a reference video or protected video may include content associated with copyright protection or other content for which access restrictions are defined. The repository of reference videos may include or otherwise be associated with previously generated fingerprint data for each reference video. Although not explicitly shown in FIG. 7 , one or more reference videos may be obtained for which previously generated fingerprint data is not available, and video screening using machine learning video screening trained using self-supervised training 700 may include automatically generating corresponding fingerprint data.

[0088] Obtaining screening data includes inputting or otherwise making available a current video to the trained machine learning video screening model, and receiving or otherwise accessing screening data from the trained machine learning video screening model in response to inputting or otherwise making available the current video to the trained machine learning video screening model. The screening data obtained at 730 may be similar to the predicate screening data obtained as shown at 920 in FIG. 9 , except as described herein or otherwise apparent from the context. For example, the predicate screening data obtained as shown at 920 in FIG. 9 may be obtained using a previously trained machine learning video screening model prior to training the current trained machine learning screening model at 710 using self-supervised training, and the screening data obtained at 730 may be obtained using the current trained machine learning video screening model trained at 710 using self-supervised training.

[0089] In response to obtaining the screening data at 730, the current video is automatically screened at 740. Screening the current video automatically generates, stores, or both screened video data indicating that the current video is being screened, which may include associating, e.g., by storing, the screened video data, which may include the screened data or portions thereof, with the current video. For example, the screening data obtained at 730 may include data that identifies or describes portions, e.g., subframes, of the current video detected by the trained machine learning video screening model obtained at 710, as candidate similar portions having similarity to, or portions thereof, of the reference video, and automatically generating the screened video data at 740 may include automatically generating fingerprint data for each frame from the portions of the current video indicated in the screening data. The fingerprint data generated for each frame from the portion of the current video shown in the screening data may be automatically compared to the reference video shown in the screening data, or to other reference videos, or to both the reference video shown in the screening data and other reference videos, to determine whether the fingerprinted portion or portions of the current video are similar to the respective portions of the respective reference videos, e.g., have a similarity value greater than a defined similarity threshold, such as 0.9, i.e., 90 percent.

[0090] In some implementations, screening the current video at 740 may include automatically restricting the current video, for example, in response to determining that one or more fingerprinted portions of the current video are similar to one or more respective portions of one or more respective reference videos. Restricting the current video may include storing restricted video data indicating that the current video is a restricted video. Restricting the current video may include preventing access to the restricted video. In some implementations, restricting the current video at 730 may be omitted or replaced or combined with other processing.

[0091] Obtaining the current video at 720, obtaining screening data from a trained machine learning video screening model trained using self-supervised training at 730, and screening the current video at 740 may be performed on other videos, such as other videos from unscreened videos, as indicated by the dashed arrow at 750.

[0092] In some implementations, in response to inputting or otherwise making available the current video to the trained machine learning video screening model, screening data may not be available from the trained machine learning video screening model, or in response to inputting or otherwise making available the current video to the trained machine learning video screening model, screening data obtained from the trained machine learning video screening model for the current video may indicate that no identified similarities exist, and screening the current video at 740 may be omitted.

[0093] Other implementations of video screening using machine learning video screening trained using self-supervised training 700 are available. For example, other screening techniques may be used in addition to screening using machine learning video screening trained using self-supervised training 700. In some implementations, additional elements of video screening using machine learning video screening trained using self-supervised training 700 may be added, certain elements may be combined, and / or certain elements may be removed.

[0094] 8 is a block diagram of an example method for obtaining a trained machine learning video screening model trained using self-supervised training 800. Obtaining a trained machine learning video screening model trained using self-supervised training 800, or a part or parts thereof, may be implemented by a computing device such as computing device 100 shown in FIG. 1 , one or more of computing and communication devices 100A, 100B, 100C shown in FIG. 2 , a video hosting system such as video hosting system 400 shown in FIG. 4 or a component thereof, or a video similarity engine such as video similarity engine 500 shown in FIG. 5. Obtaining a trained machine learning video screening model trained using self-supervised training 800 may be similar to obtaining a trained machine learning video screening model trained using self-supervised training as shown at 710 in FIG. 7 , except as described herein or otherwise apparent from the context.

[0095] Obtaining a trained machine learning video screening model trained using self-supervised training 800 includes obtaining a training dataset at 810, obtaining an untrained machine learning video screening model at 820, and training the untrained machine learning video screening model using the training dataset at 830.

[0096] A training dataset, which is an automatically generated training dataset, is obtained at 810. The training dataset includes training examples. The usefulness, such as accuracy and efficiency, of a trained machine learning model, such as the trained machine learning video screening model described herein, correlates with the number, or density, and quality of the training examples in the training dataset used to train the model. Manually generating data, such as by a human, to accurately annotate or label the training examples utilizes substantial resources, including human and computing resources, and is impractical for generating a relatively large number, quantity, or density of training examples, such as 100,000 or more training examples. The automatically generated training dataset described herein includes automatically generated training examples that are generated without the presence of manual, i.e., human-generated, annotation or labeling data. Although not explicitly described herein, training examples that include manual, i.e., human-generated, annotation or labeling data, unlike the training examples included in the automatically generated training dataset described herein, may be used in combination with the automatically generated training dataset described herein. Although described as a training data set, multiple separate or non-overlapping portions or subsets of the automatically generated training data set may be identified, which may include identifying a subset of the training data and identifying another subset as validation data. An example of obtaining a training data set is shown in FIG.

[0097] Each training example includes a training video or its identifier, a reference video or its identifier, and data indicating the similarity between the training video and the reference video, such as automatically generated annotations or labels, fingerprint similarity, etc. A training example may include recall examples, precision examples, or both. A training dataset may include multiple training samples, such as thousands of training samples. For example, a training dataset may include 600,000 high-confidence training examples, 400,000 low-confidence training examples, and 150,000 precision examples.

[0098] The example accuracy may be training examples associated with the accuracy of the trained machine learning video screening model, which may be expressed as a ratio of true positive screening data generated by the trained machine learning video screening model to the sum of true positive screening data generated by the trained machine learning video screening model and false positive screening data generated by the trained machine learning video screening model. For example, the example accuracy may match false positive screening data generated using another, different machine learning video screening model, such as a machine learning video screening model trained before training the current machine learning video screening model.

[0099] Recall examples may be training examples associated with the recall or sensitivity of the trained machine learning video screening model, which may be expressed as a ratio of true-positive screening data generated by the trained machine learning video screening model to the sum of true-positive screening data generated by the trained machine learning video screening model and false-positive screening data generated by the trained machine learning video screening model. Recall examples may include high-confidence training examples, low-confidence training examples, or both. For example, high-confidence recall examples may match true-positive screening data generated using another, different machine learning video screening model, such as a machine learning video screening model trained before training the current machine learning video screening model. In another example, low-confidence recall examples may match false-negative screening data generated using another, different machine learning video screening model, such as a machine learning video screening model trained before training the current machine learning video screening model. As used herein, the terms “confidence” and variations and forms thereof, such as “high confidence” and “low confidence,” refer to one or more values, such as floating-point values ​​in the range of zero (0) to one (1), that indicate the probability or likelihood that a corresponding output, such as corresponding output data indicating a sub-frame portion of an input frame that matches a reference frame, automatically generated by a machine learning model and determined by the machine learning model is accurate.

[0100] An untrained machine learning video screening model is obtained at 820. The untrained machine learning video screening model may be a self-supervised object detection model. In some implementations, the untrained machine learning video screening model may be a partially trained untrained machine learning video screening model. The machine learning video screening model may be a mathematical model for evaluating an input video or a probe video to detect similarity with one or more reference videos, for example, based on fingerprint similarity that determines similarity between an automatically generated fingerprint for the input video and an automatically generated fingerprint for the reference video. Obtaining the untrained machine learning video screening model may include reading or otherwise accessing data that defines or describes the untrained machine learning video screening model, such as from a file, a database, or another data source.

[0101] In some implementations, the machine learning or non-linear video screening model is an artificial neural network model. As used herein, the term "neural network" refers to an artificial neural network.

[0102] A neural network model includes layers, each containing connected units (nodes, perceptrons, or neurons) followed by a nonlinearity. As used herein, the term "neuron" refers to an artificial neuron. A layer is a set of nodes or neurons in a neural network that processes a set of input functions or the output of those neurons. An artificial neural network model describes layers for organizing and arranging nodes or neurons within an artificial neural network, including input layers, output layers, hidden layers, inner layers, or hidden layers.

[0103] An artificial neural network model describes nodes or artificial neurons. A node or neuron in an artificial neural network may receive or otherwise access input values ​​and generate an output value. For example, a node or neuron may calculate an output value by applying an activation function (a nonlinear transformation) to a weighted sum of input values. A node in an artificial neural network may be represented as a mathematical function, which may include describing or defining one or more parameters or thresholds for the node. A node in an artificial neural network may receive one or more input signals, determine an internal state following or in response to receiving the input signals (activation), and output an output signal based on (e.g., using or responding to) the input signals and the internal state. The input signals may be associated with respective weighting values. The artificial neural network model may describe or define the weighting values. For example, determining the internal state may include determining a weighted sum of the input signals, transforming the sum using an activation function or transformation function, which may be, for example, a nonlinear function, and outputting the transformed result or a function thereof (an output function).

[0104] The input layer, or first layer, may receive or otherwise access input data (functions) for the neural network. Nodes in the input layer (input nodes) of an artificial neural network may receive input data for the artificial neural network. Input data for layers other than the input layer may be output data from another adjacent layer of the neural network. Nodes in adjacent layers may be interconnected along edges. An artificial neural network model may describe or define weighting values ​​associated with each edge. A hidden layer may be a composite layer within a neural network between the input layer and the output layer. The hidden layer may include an activation function, such as for training. The output layer, or final layer, may output data indicating the answer or prediction of the neural network, for example, in response to input data accessed by the input layer. The activation function may be a function that may be nonlinear, using a weighted sum of inputs from previous layers, to generate data that may be output (output values) to subsequent layers, etc. An output node in the output layer of an artificial neural network may output a predicted value based on (using or responding to, etc.) the received input values.

[0105] Video screening using a machine learning video screening model trained using self-supervised training may include using a convolutional neural network (CNN) model. The convolutional neural network may be a neural network whose layers are convolutional layers. The convolutional neural network may include multiple convolutional layers. In some embodiments, the convolutional neural network may include one or more convolutional layers, one or more pooling layers, one or more dense or fully connected layers, or a combination thereof. The convolutional neural network may be a deep neural network, which may be a neural network including multiple hidden layers.

[0106] A convolutional layer may be a layer of a neural network that applies a convolutional filter to an input matrix, which may involve performing one or more convolution operations. A convolutional filter is a matrix that has the rank (order) of the input matrix but has a smaller shape (element dimension). Each element, or cell, of the convolutional filter matrix may be a single-digit binary value, such as zero or one, that may be initialized, such as randomly, and trained to optimize. A convolutional operation may involve element-by-element multiplication of the convolutional filter with a portion, or slice, of the input matrix that has the same rank and size as the convolutional filter. A convolutional operation may involve summation of the matrix resulting from the element-by-element multiplication. A convolutional layer may perform a respective convolutional operation on each portion, or slice, of the input matrix.

[0107] A pooling layer may be a layer of a neural network that reduces one or more matrices output by a previous convolutional layer into smaller matrices. For example, the pooling layer may determine the maximum or average value of the pooled region (pooling operation). The pooling operation may divide the matrix (convolution output) into portions that may overlap, such as partially overlapping, and the difference in the positions of the matrices of each adjacent portion may be called a stride.

[0108] A dense or fully connected layer may be a layer of a neural network, such as a hidden layer, in which each node is connected to a node in the subsequent hidden layer. The convolutional neural network may be a multi-layer convolutional neural network with a K×K weight matrix (kernel), which may include spatial processing such as downsampling, upsampling, or modulation.

[0109] In some implementations, the machine learning video screening model is a deep neural network and a single-shot multi-box detector that may use depthwise separable convolutions, which provides improved speed and efficiency compared to other models.

[0110] The untrained machine learning video screening model is trained using the training dataset to generate a trained machine learning video screening model at 830. Training the machine learning video screening model includes inputting each pair of a training video and a reference video (training pair) into the machine learning video screening model, where annotation data is not available, to obtain output data indicating whether a similarity, such as fingerprint similarity, between the training video or a spatial portion thereof and the reference video is detected, comparing the output data with the annotation data for each pair, and updating one or more parameters of the machine learning video screening model based on the comparison so that the accuracy of the machine learning video screening model improves, which may be performed iteratively, etc., so that the output of the trained machine learning video screening model matches the annotation data.

[0111] In some implementations, obtaining a trained machine learning video screening model trained using self-supervised training 800 may be performed automatically according to a defined time period, such as daily, or may be performed in response to a defined detection event, such as detecting a request to obtain a trained machine learning video screening model trained using self-supervised training.

[0112] Other implementations of obtaining a trained machine learning video screening trained using self-supervised training 800 are available. In some embodiments, additional elements can be added, certain elements can be combined, and / or certain elements can be removed from obtaining a trained machine learning video screening trained using self-supervised training 800.

[0113] 9 is a block diagram of an example method for obtaining an automatically generated training dataset 900. Obtaining the automatically generated training dataset 900, or a portion or portions thereof, may be implemented by a computing device, such as computing device 100 shown in FIG. 1, one or more of computing and communication devices 100A, 100B, 100C shown in FIG. 2, a video hosting system, such as video hosting system 400 shown in FIG. 4, or a component thereof, or a video similarity engine, such as video similarity engine 500 shown in FIG. 5.

[0114] Obtaining the automatically generated training dataset 900 may be similar to obtaining a training dataset as shown at 810 in Figure 8, except as described herein or otherwise apparent from the context. For example, video screening using a machine learning video screening model trained using self-supervised training, such as video screening using a machine learning video screening model trained using self-supervised training 700 shown in Figure 7, may include obtaining a trained machine learning video screening model trained using self-supervised training, such as obtaining a trained machine learning video screening model trained using self-supervised training as shown at 710 in Figure 7 or obtaining a trained machine learning video screening model trained using self-supervised training 800 as shown in Figure 8, which may include obtaining a training set using self-supervised training as shown at 810 in Figure 8 or obtaining an automatically generated training dataset 900 as shown in Figure 9.

[0115] Obtaining the automatically generated training dataset 900 includes obtaining unannotated training pairs at 910, obtaining predicate screening data at 920, obtaining candidate screening data at 930, obtaining candidate pairs at 940, determining similarity at 950, determining confidence at 960, determining complexity at 970, including the pairs in the training dataset at 980, and removing bias from the training dataset at 990.

[0116] Although not separately explicitly shown in FIG. 9 , obtaining the automatically generated training dataset 900 may include obtaining or identifying a set, group, or other collection of input or probe videos for use as training videos. The input or probe videos used as training videos may be videos uploaded to or otherwise made available to one or more computing devices, or a system including one or more computing devices, that implements the self-supervised training, or a portion thereof. For example, the set, group, or other collection of input or probe videos identified for use as training videos may include thousands of videos, such as over 100,000 videos. Although not separately explicitly shown in FIG. 9 , obtaining the automatically generated training dataset 900 may include obtaining or identifying reference videos, for example, by obtaining or accessing a repository of reference videos described herein.

[0117] Unannotated training pairs including a training video and a reference video are obtained at 910. The training videos are obtained from a set, group, or other collection of input or probe videos identified for use as training videos. The reference videos are obtained from a repository or other collection of videos defined as reference videos. Manually generated, e.g., human-generated, annotation data, such as frame-by-frame annotation data, or training label data, is not available for the unannotated training pairs obtained at 910, i.e., is omitted or does not exist from the unannotated training pairs.

[0118] Predicate screening data is obtained at 920 for the unannotated training pairs. The predicate screening data is data, such as log data, generated by automatically screening one or more of the input videos or probe videos with respect to a reference video by a previously trained machine learning video screening model (a predicate machine learning video screening model or a predicate video screening model) that was trained before training the current machine learning video screening model. In some implementations, generating the predicate screening data may include using automatic video screening other than video screening using a predicate video screening model. An exemplary graphical representation of predicate screening data is shown in FIG. 10.

[0119] Prior to training the current machine learning video screening model, such as within a defined time range such as one day (24 hours), one or more of the input videos or probe videos are screened by the predicate video screening model relative to the reference videos using a predefined high-confidence confidence threshold such as 0.6, and the corresponding screening data output or generated by the predicate video screening model is stored or otherwise made available to the current machine learning video screening model, such as in a screening data log. To train the current machine learning video screening model, previously generated log data corresponding to the training videos is obtained as predicate screening data.

[0120] The predicate screening data includes screening data generated by the predicate video screening model regarding similarities, such as fingerprint similarities, between spatial sub-frames from each time sequence from each input video and corresponding frames from each reference video, identified by the predicate video screening model.

[0121] Each record, row, or entry in the predicate screening data includes an identifier for an input video or a probe video. Each record, row, or entry in the predicate screening data includes an identifier for a reference video. Each record, row, or entry in the predicate screening data includes a segment start position, such as the temporal or sequential position of a frame in the input video or the probe video. Each record, row, or entry in the predicate screening data includes a segment end position, such as the temporal or sequential position of a frame in the input video or the probe video that follows the frame corresponding to the segment start position. Each record, row, or entry in the predicate screening data includes a reference start position, such as the temporal or sequential position of a frame in the reference video. Each record, row, or entry in the predicate screening data includes a reference end position, such as the temporal or sequential position of a frame in the reference video that follows the frame corresponding to the reference start position. The segment start position and segment end position describe an interval or segment (predicate time segment) of the input video or probe video that has been identified by the predicate video screening model as being similar to a corresponding segment (reference time segment) of the reference video, for example, based on fingerprint similarity, with a confidence level equal to or above a predefined minimum high-confidence confidence threshold, and the reference start position and reference end position describe the corresponding reference time segment of the reference video.

[0122] Each record, row, or entry in the predicate screening data includes a center value indicating the spatial center of a sub-frame in the input video corresponding to the identified predicate time segment, which may be an aggregate, such as an average, of the sub-frame positions for each frame from each frame corresponding to the identified predicate time segment. Each record, row, or entry in the predicate screening data includes an area, such as an area for each frame, of a sub-frame in the input video corresponding to the identified predicate time segment, which may be an aggregate, such as an average, of the sub-frame areas for each frame from each frame corresponding to the identified predicate time segment.

[0123] Multiple spatial portions, temporal portions, or both of each input video may be represented by respective records, rows, or entries in the predicate screening data. The predicate screening data includes respective records, rows, or entries indicating similarities between the input video, or portions thereof, and each of the respective reference videos, for example, based on fingerprint similarity. Data corresponding to multiple spatial portions, temporal portions, or both of the input video that are not identified by the predicate video screening model or that are identified with a confidence below a minimum high-confidence confidence threshold may be omitted or absent from the predicate screening data. Data for each frame, such as frames other than the first and last frames, within the segment represented by the predicate screening data may be omitted or absent from the predicate screening data. Confidence data and similarity data may be omitted or absent from the predicate screening data.

[0124] For unannotated training pairs identified at 910, obtaining predicate screening data at 920 includes obtaining, such as reading or otherwise accessing, each row, record, or entry from the predicate screening data that indicates a similarity, such as a fingerprint similarity, between the training video and the reference video identified by the predicate video screening model. The training video includes a predicate time segment indicated in each row, record, or entry from the predicate screening data. The training video may include one or more frames or images, in temporal or sequential order, before the predicate time segment indicated in each row, record, or entry from the predicate screening data. The training video may include one or more frames or images, in temporal or sequential order, after the predicate time segment indicated in each row, record, or entry from the predicate screening data.

[0125] Candidate screening data is obtained at 930. Obtaining the candidate screening data includes identifying an extended time segment from the training video. The extended time segment includes the predicate time segment of the training video identified at 920 based on the predicate screening data. The extended time segment includes one or more frames from the training video other than the frame corresponding to the predicate time segment, such as a frame preceding the predicate time segment, a frame following the predicate time segment, or both. The extended time segment may include one or more frames from a defined time range, such as 2 seconds, or a defined concentration of frames, such as 60 frames, immediately before the predicate time segment, immediately after the predicate time segment, or both. For example, a sampling of frames, such as a uniform sampling of frames, such as a sampling of frames including a defined concentration of frames, such as 3 frames, from a defined time range immediately before the predicate time segment, or from a defined concentration of frames, may be included in the extended time segment. In another example, a sampling of frames may be included in the extended time segment, such as a uniform sampling of frames, such as a sampling of frames including a defined concentration of frames, such as three frames, from a defined time range immediately following the predicate time segment, or a defined concentration of frames. In another example, the extended time segment may include three frames from two seconds of the current video immediately preceding the predicate time segment and three frames from two seconds of the current video immediately following the predicate time segment.

[0126] Obtaining the candidate screening data includes obtaining a second previously trained video screening model. For example, the second previously trained video screening model may be similar to the predicate video screening model obtained in 920, except that the second previously trained video screening model may be used with a predefined low-confidence confidence threshold that is lower than the predefined high-confidence confidence threshold used by the predicate video screening model. For example, the predefined high-confidence confidence threshold may be 0.6, or 60 percent, and the predefined low-confidence confidence threshold may be 0.4, or 40 percent. The predefined high-confidence confidence threshold may be referred to as a first predefined confidence level, and the predefined low-confidence confidence threshold may be referred to as a second predefined confidence threshold. Alternatively, the predefined high-confidence confidence threshold may be referred to as a second predefined confidence level, and the predefined low-confidence confidence threshold may be referred to as a first predefined confidence threshold. Stated differently, the use of first, second, etc. herein is used to distinguish elements herein and does not imply any order unless otherwise stated or clear from the context.

[0127] Obtaining candidate screening data includes generating candidate screening data for the extended time segment with respect to the reference video identified at 910 using a second previously trained video screening model on a frame-by-frame basis, such as for the extended time segment. The candidate screening data indicates associations that may be based on similarity between respective screening frames from the reference video and spatial portions of candidate frames (candidate sub-frames) from the extended time segment, for example, based on fingerprint similarity. Fingerprint similarity may indicate, for example, similarity between automatically generated fingerprints for the screening frames and automatically generated fingerprints for the candidate sub-frames.

[0128] The candidate subframes may be rectangular spatial portions of respective candidate frames from the extended time segment. For example, referring to FIG. 6, the candidate frame from the extended time segment may be video frame 600 shown in FIG. 6, and the candidate subframes may correspond to spatial subframes 610 shown in FIG. 6. Obtaining the candidate screening data includes identifying candidate subframe pairs of the candidate subframes shown in the candidate screening data. Each candidate subframe pair includes a candidate subframe and a corresponding screening frame from a reference video identified from the candidate screening data.

[0129] A candidate subframe pair is obtained at 940. Obtaining the candidate subframe pair includes identifying a candidate subframe from the candidate screening data obtained at 930, extracting a candidate subframe from a corresponding frame from the extended time segment from the training video, and using the extracted candidate subframe as a frame for comparison with a corresponding screening frame from the reference video. The candidate subframe pair includes the extracted candidate subframe as a frame and the corresponding screening frame from the reference video.

[0130] A similarity score or value is determined at 950 for the candidate subframe pair, e.g., based on fingerprint similarity. The similarity score for the candidate subframe pair indicates a measure of similarity between the candidate subframe extracted as a frame and a corresponding screening frame from the reference video, e.g., based on fingerprint similarity. In some implementations, determining the similarity score for the candidate subframe pair includes generating fingerprints representing the extracted candidate subframes and comparing the fingerprints representing the extracted candidate subframes with fingerprints of the corresponding screening frames from the reference video. The fingerprints of the corresponding screening frames from the reference video may be fingerprints previously generated for the corresponding screening frames from the reference video. Determining the similarity at 950 may include determining whether the similarity score for the candidate subframe pair is greater than or equal to a predefined minimum similarity threshold. Candidate subframe pairs identified as having similarity scores that are greater than or equal to a predefined minimum similarity threshold may be identified as similar, and candidate subframe pairs identified as having similarity scores that are less than a predefined minimum similarity threshold (first predefined similarity threshold) may be identified as dissimilar. Obtaining the automatically generated training dataset 900 includes determining whether the similarity scores of the candidate subframe pairs are within a range, such as less than the predefined minimum similarity threshold, or greater than or equal to the predefined minimum similarity threshold.

[0131] In response to determining that the similarity score of the candidate subframe pair is below a predefined minimum similarity threshold, the candidate subframe pair is omitted or excluded from the training data set. For example, a candidate subframe pair identified as dissimilar has a similarity value that may be zero or greater than zero and is below a predefined minimum similarity threshold, indicating insufficient similarity to be used as a training example.

[0132] Obtaining the automatically generated training data set 900 includes filtering the candidate subframe pairs to obtain low-confidence training examples at 960. In response to determining that the similarity scores of the candidate subframe pairs are equal to or greater than a predefined minimum similarity threshold, frame-by-frame filtered screening data is obtained at 960 from the predicate video screening model for the training video corresponding to the candidate subframe pair and the reference video using a predefined high-confidence confidence threshold. Obtaining the automatically generated training data set 900 includes determining whether the filtered screening data includes data indicative of an association between the candidate subframes and the screening frames indicated in the candidate subframe pairs.

[0133] In response to determining that the filtered screening data contains data indicating an association between the candidate subframe and the screening frame shown in the candidate subframe pair, if a predicate video screening model using a predefined high-confidence confidence threshold indicates that the candidate subframe pair is identified as similar, the candidate subframe pair is omitted or excluded from the training data set or is included in the training data set as a high-confidence training example.

[0134] In response to determining that data indicating an association between the candidate subframe and the screening frame shown in the candidate subframe pair is absent or omitted from the filtered screening data, if the predicate video screening model using a predefined high-confidence confidence threshold indicates that the candidate subframe pair is identified as a mismatch or as a match with a confidence value below the predefined high-confidence confidence threshold, the candidate subframe pair is included in the training dataset as a low-confidence training example.

[0135] Obtaining the automatically generated training dataset 900 includes, at 970, complexity-based filtering of candidate subframe pairs to omit or exclude low-complexity training examples. The complexity-based filtering includes identifying a spatial portion of the screening frame, such as an upper quarter quadrant of the screening frame, from the reference video; generating fingerprint data for the spatial portion of the screening frame; and comparing the fingerprint data for the spatial portion of the screening frame with the fingerprint data of the screening frame to obtain a similarity value indicating a measure of similarity between the spatial portion of the screening frame and the screening frame, e.g., based on fingerprint similarity. A similarity value greater than a predefined minimum similarity complexity threshold (second predefined similarity threshold) indicates that the screening frame lacks sufficient complexity or image detail, e.g., the color of the screening frame is substantially uniform, e.g., substantially uniformly black.

[0136] In response to determining that the similarity value for the screening frame and the spatial portion of the screening frame is equal to or greater than a predefined minimum similarity complexity threshold (a second predefined similarity threshold), indicating that the screening frame lacks sufficient complexity or image detail, the corresponding training pair is omitted or excluded from the training data set.

[0137] In response to determining that the similarity values ​​for the screening frame and the spatial portion of the screening frame are less than a predefined minimum similarity complexity threshold (a second predefined similarity threshold), the corresponding training pair is included in the training data set at 980. Other complexity metrics may also be evaluated.

[0138] Although not separately shown in FIG. 9, obtaining the automatically generated training dataset 900 may be performed, for example, iteratively, for each combination of training videos from input videos or probe videos and reference videos from a repository of reference videos.

[0139] In some implementations, obtaining candidate pairs at 940, determining similarity at 950, determining confidence at 960, determining complexity at 970, and including pairs in training data at 980 may be performed with respect to two or more candidate subframe pairs, which may differ spatially, temporally, or both, as indicated by dashed arrows at 982.

[0140] In some implementations, obtaining unannotated training pairs at 910, obtaining predicate screening data at 920, obtaining candidate screening data at 930, obtaining candidate pairs at 940, determining similarity at 950, determining confidence at 960, determining complexity at 970, and including pairs in training data at 980 may be performed in connection with two or more training pairs, as indicated by dashed arrows at 984.

[0141] The training dataset is debiased at 990. Debiasing the training dataset may improve the efficiency, accuracy, or both of the trained video screening model by reducing or eliminating bias introduced by the automated generation of the training dataset. Debiasing may detect bias with respect to spatial location, temporal location, or both, of the candidate subframes. Debiasing may detect bias with respect to spatial size, temporal size, or both, of the candidate subframes. Debiasing may detect bias with respect to spatial orientation, such as horizontal or vertical orientation, of the candidate subframes.

[0142] Other implementations of video screening using machine learning video screening trained using self-supervised training 700 are available. In some embodiments, additional elements of video screening using machine learning video screening trained using self-supervised training can be added, certain elements can be combined, and / or certain elements can be removed.

[0143] 10 is a diagram of an example of a graphical representation of predicate screening data 1000 relative to a reference video and an input or probe video. The example graphical representation of predicate screening data 1000 shows a horizontal axis 1010 corresponding to the reference video, with increasing time sequence shown from left to right. The example graphical representation of predicate screening data 1000 shows a vertical axis 1020 corresponding to the probe video, with increasing time sequence shown from bottom to top.

[0144] The squares are shown as a matrix of rows and columns with the horizontal axis 1010 representing the reference video and the vertical axis 1020 representing the probe video.

[0145] The columns of the shown matrix correspond to respective temporal positions, or frames, from the reference video, e.g., the leftmost column shown corresponds to the first, or earliest, frame in time or order, of the reference video.

[0146] The rows of the matrix shown correspond to respective temporal positions, or frames, from the probe video, e.g., the bottom row shown corresponds to the first, or earliest, frame in time or order, of the probe video.

[0147] The background of each square indicates the determined similarity between the corresponding frame of the probe video and the corresponding frame of the reference video. For brevity and clarity, five levels of similarity, L0, L1, L2, L3, and L4, are shown in Figure 10.

[0148] Each pair of a frame of the probe video and a corresponding frame of the reference video for which data indicating similarity (L0) is unavailable or otherwise indicates that sufficient similarity has not been detected, such as exceeding a predefined threshold, is shown with a white background, as in 1030.

[0149] Each pair of a frame of the probe video and a corresponding frame of the reference video, where data showing a relatively low similarity (L1) greater than a predefined threshold is shown in the predicate screening data, is shown with a wide diagonal background pointing downward and to the left, as in 1031.

[0150] Each pair of a frame of the probe video and a corresponding frame of the reference video where data showing a relatively moderately low similarity (L2) that is greater than a defined threshold and greater than a relatively low similarity is shown in the predicate screening data is shown with a narrow diagonal background pointing downward to the right, as in 1032.

[0151] Each pair of a frame of the probe video and a corresponding frame of the reference video where data showing a relatively moderately high similarity (L3), which is greater than a defined threshold and greater than a relatively moderately low similarity, is shown in the predicate screening data is shown with a dotted background, as at 1033.

[0152] Each pair of a frame of the probe video and a corresponding frame of the reference video for which data showing a relatively high similarity (L4), greater than a defined threshold and greater than a relatively moderately high similarity, is shown in the predicate screening data is shown with a black background, as in 1034.

[0153] The segments identified in or by the predicate screening data are shown as diagonal lines 1040 .

[0154] Each pair of a training frame from the probe video and a screen frame from the reference video that may be included in the extended time segment described with respect to obtaining candidate screening data as shown at 930 in FIG. 9 is shown with a thick border, such as at 1050 in FIG. 10.

[0155] 11 is a block diagram of another example method of video screening using a machine learning video screening model trained using self-supervised training 1100. Video screening using the machine learning video screening model trained using self-supervised training 1100, or a portion or portions thereof, may be implemented by a computing device such as computing device 100 shown in FIG. 1, one or more of computing and communication devices 100A, 100B, 100C shown in FIG. 2, a video hosting system such as video hosting system 400 shown in FIG. 4, or a component thereof, or a video similarity engine such as video similarity engine 500 shown in FIG.

[0156] Video screening using a machine learning video screening model trained using self-supervised learning 1100 includes an active phase 1110 and a training phase 1120. The active phase 1110 and the training phase 1120 may be performed sequentially and iteratively, with the current iteration of the active phase 1110 being performed before a subsequent iteration of the training phase 1120 and the current iteration of the training phase 1120 being performed after a previous iteration of the active phase 1110.

[0157] The active phase 1110 may be similar to video screening using a machine learning video screening model trained using self-supervised training 700 shown in Figure 7, except as described herein or apparent from the context. The active phase 1110 includes acquiring a reference video at 1111, acquiring a current trained machine learning video screening model trained using self-supervised training at 1112, acquiring a probe video at 1113, and generating screening data using the current trained machine learning video screening model at 1114. The active phase 1110 may be performed over a defined time range, such as one day (24 hours).

[0158] One or more reference videos are obtained at 1111. Obtaining the reference videos at 1111 may include obtaining a repository, database, data store, or other collection of videos defined or described as reference videos, or otherwise accessing a repository, database, data store, or other collection of videos. For example, the reference videos or protected videos may include content associated with copyright protection or other content for which access restrictions are defined. The repository of reference videos may include or be otherwise associated with previously generated fingerprint data for each reference video. Although not explicitly shown in FIG. 11 , one or more reference videos may be obtained for which previously generated fingerprint data is not available, and video screening using machine learning video screening trained using self-supervised training 1100 may include automatically generating corresponding fingerprint data.

[0159] A current trained machine learning video screening model trained using self-supervised training is obtained at 1112. The current trained machine learning video screening model is obtained from the output of a previous iteration of the training stage 1120, as indicated by the dashed arrow at 1150, and the reference video obtained at 1111 is the reference video used in the previous iteration of the training stage 1120. Obtaining a current trained machine learning video screening model trained using self-supervised training may be similar to obtaining a trained machine learning video screening model trained using self-supervised training, as shown at 710 in FIG. 7, except as described herein or otherwise clear from the context. The current trained machine learning video screening model is implemented using a predefined high-confidence confidence threshold.

[0160] One or more current probe videos, input videos, or unscreened videos are obtained at 1113. For example, the unscreened videos may be uploaded or otherwise made available to a computing device, or a system including a computing device, as shown at 1130, and the current video may be obtained from the unscreened videos. Obtaining the current probe video, input video, or unscreened video at 1112 may be similar to obtaining the current video as shown at 720 in FIG. 7, except as described herein or otherwise apparent from the context. The current probe video, input video, or unscreened video is a video obtained or otherwise made available according to a defined time range of the current iteration of the active phase 1110. In some implementations, other videos, such as other unscreened videos, may be used.

[0161] A screening video is generated at 1114 using the reference video acquired at 1111, the current trained machine learning video screening model acquired at 1112, and the current probe video acquired at 1113. Generating screening data at 1114 includes generating or inputting a screening data log including the screening data or a portion thereof, and storing or otherwise outputting the screening data log. Except as described herein or otherwise apparent from the context, generating screening data at 1114 may be similar to obtaining screening data from a trained machine learning video screening model trained using self-supervised training shown in 730 of FIG. 7. For example, the screening data stored in the screening data log may be used as predicate screening data for subsequent iterations of the training phase 1120. An exemplary graphical representation of a screening data log is shown in FIG. 10.

[0162] The screening data indicates the similarity between spatial subframes from each time sequence from each current probe video, input video, or unscreened video acquired at 1113, identified by the current trained machine learning video screening model acquired at 1112, and corresponding frames from each reference video acquired at 1111.

[0163] Each record, row, or entry in the screening data includes an identifier for the current probe video, input video, or unscreened video. Each record, row, or entry in the screening data includes an identifier for the reference video. Each record, row, or entry in the screening data includes a segment start position, such as the temporal or sequential position of a frame in the current probe video, input video, or unscreened video. Each record, row, or entry in the screening data includes a segment end position, such as the temporal or sequential position of a frame in the current probe video, input video, or unscreened video that follows the frame in the current probe video, input video, or unscreened video that corresponds to the segment start position. Each record, row, or entry in the screening data includes a reference start position, such as the temporal or sequential position of a frame in the reference video. Each record, row, or entry in the screening data includes a reference end position, such as the temporal or sequential position of a frame in the reference video that follows the frame in the reference video that corresponds to the reference start position. The segment start position and segment end position describe an interval or segment (current time segment) in the current probe video, input video, or unscreened video that has been identified by the current trained machine learning video screening model as being similar to a corresponding segment (reference time segment) in the reference video, identified with a confidence equal to or greater than a predefined minimum high-confidence confidence threshold, and the reference start position and reference end position describe the corresponding reference time segment in the reference video.

[0164] Each record, row, or entry in the screening data includes a center value indicating the spatial center of a sub-frame in the current probe video, input video, or unscreened video corresponding to the identified current time segment, which may be an aggregate, such as an average, of the sub-frame positions for each frame from each frame corresponding to the identified current time segment. Each record, row, or entry in the screening data includes an area, such as an area for each frame, of a sub-frame in the current probe video, input video, or unscreened video corresponding to the identified current time segment, which may be an aggregate, such as an average, of the sub-frame areas for each frame from each frame corresponding to the identified current time segment.

[0165] Multiple similarities, such as fingerprint similarities, may be identified for each current probe video, input video, or unscreened video, and corresponding reference video, which may be represented as multiple rows, records, or entries in the screening data log. Respective similarities between each current probe video, input video, or unscreened video, and multiple reference videos may be identified.

[0166] The training stage 1120 may be similar to obtaining a trained machine learning video screening model trained using self-supervised training 800 shown in Figure 8, except as described herein or apparent from the context. The training stage 1120 includes obtaining a reference video at 1121, obtaining a current untrained machine learning video screening model at 1122, obtaining a training video at 1123, obtaining predicate screening data at 1124, and training the current untrained machine learning video screening model at 1125.

[0167] One or more reference videos are acquired at 1121. Acquiring a reference video at 1121 may be similar to acquiring a reference video at 1111, except as described herein or otherwise clear from the context. Acquiring a reference video at 1121 includes acquiring a reference video acquired at 1111, as indicated by the arrowed line between acquiring a reference video at 1111 and acquiring a reference video at 1121. Although not explicitly shown in FIG. 11 , one or more of the reference videos acquired at 1111 may be absent or omitted from the reference video acquired at 1121. Acquiring a reference video at 1121 may include acquiring a reference video other than the Reference A video acquired at 1111. For example, one or more reference videos other than the reference video acquired at 1111 may be uploaded to or otherwise included in a repository, database, data store, or other collection of videos defined or described as reference videos, as shown at 1160, which may include automatically generating fingerprint data for each reference video.

[0168] A current untrained machine learning video screening model is obtained at 1122. Obtaining the current untrained machine learning video screening model may include obtaining data describing the current untrained machine learning video screening model, such as data describing or defining the structure or architecture of the current untrained machine learning video screening model, data describing or identifying devices or components, such as hardware components, for training the current untrained machine learning video screening model, data defining or describing one or more untrained model weights of the current untrained machine learning video screening model, which may be random or pseudo-random values, data defining or describing one or more training hyperparameters for training the current untrained machine learning video screening model, which may be manually generated values. For example, the current untrained machine learning video screening model may be an artificial neural network model, and the data describing or defining the current untrained machine learning video screening model may indicate the number or cardinality of layers of the current untrained machine learning video screening model, the number or cardinality of nodes or artificial neurons, or both.

[0169] A training video is acquired at 1123. The training video acquired at 1123 includes the probe video acquired at 1113 after screening at 1114, or a portion thereof, as indicated by the arrow line between acquiring the probe video at 1113 and acquiring the training video at 1123. Acquiring the training video at 1123 may be similar to acquiring a training dataset as shown at 810 in FIG. 8, except as described herein or apparent from the context.

[0170] Predicate screening data is obtained at 1124. Obtaining the predicate screening data at 1124 includes obtaining screening data or a portion thereof output by a previously trained machine learning video screening model at 1114, as indicated by the arrow line between generating screening data using a current trained machine learning video screening model at 1114 and obtaining the predicate screening data at 1124.

[0171] The current untrained machine learning video screening model is trained at 1125 to obtain a current trained machine learning video screening model. Training the current untrained machine learning video screening model is similar to training an untrained machine learning video screening model using a training dataset, as shown at 830 in Figure 8, except as described herein or otherwise apparent from the context. The current trained machine learning video screening model output or generated by training the current untrained machine learning video screening model at 1125 may be used as the current trained machine learning video screening model for a subsequent iteration of the active phase 1110, as indicated by the arrowed line at 1150.

[0172] As used herein, the terms "optimal," "optimized," "optimization," or other forms thereof, are relative to the respective context and do not indicate absolute theoretical optimization unless expressly specified herein.

[0173] As used herein, except as expressly explained herein or otherwise clear from context, the term "set" refers to a distinguishable collection or group of zero or more distinct elements or members that may be represented as a one-dimensional array or vector.

[0174] The word "example" or "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "example" or "exemplary" should not necessarily be construed as preferred or advantageous over other aspects or designs. Rather, use of the word "example" or "exemplary" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from context, "X includes A or B" is intended to mean the natural inclusive permutation. That is, if X includes A, if X includes B, or if X includes both A and B, then "X includes A or B" is satisfied under all of the foregoing cases. Furthermore, the articles "a" and "an," as used in this application and the appended claims, should generally be construed to mean "one or more" unless otherwise specified or clear from context to refer to the singular form. Additionally, use of the terms "an embodiment" or "one embodiment" or "an implementation" or "one implementation" throughout does not refer to the same embodiment or implementation unless so described. As used herein, the terms "determining" and "identifying," or variations thereof, include selecting, ascertaining, calculating, retrieving, receiving, determining, establishing, obtaining, or otherwise identifying or determining using one or more of the devices shown in FIG.

[0175] Additionally, while for simplicity of explanation, the figures and descriptions herein may include a series of steps or acts, elements of the methods disclosed herein may occur in various orders and / or simultaneously. Moreover, elements of the methods disclosed herein may occur with other elements not explicitly shown and described herein. Furthermore, one or more elements of the methods described herein may be omitted from an implementation of a method in accordance with the disclosed subject matter.

[0176] The sending computing and communication device 100A and / or the receiving computing and communication device 100B (and the algorithms, methods, instructions, etc. stored therein and / or executed by them) can be implemented in hardware, software, or any combination thereof. Hardware can include, for example, computers, intellectual property (IP) cores, application-specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, microcontrollers, servers, microprocessors, digital signal processors, or any other suitable circuitry. In the claims, the term "processor" should be understood to encompass any of the above hardware, either alone or in combination. The terms "signal" and "data" are used interchangeably. Furthermore, portions of the sending computing and communication device 100A and the receiving computing and communication device 100B need not necessarily be implemented in the same way.

[0177] Additionally, in some implementations, for example, the sending computing and communication device 100A or the receiving computing and communication device 100B may be implemented using a computer program that, when executed, performs any of the respective methods, algorithms, and / or instructions described herein. Additionally or alternatively, a special-purpose computer / processor may be utilized, which may include dedicated hardware for executing any of the methods, algorithms, or instructions described herein.

[0178] Furthermore, all or part of the implementation may be in the form of a computer program accessible from, for example, a tangible computer-usable or computer-readable medium. The computer-usable or computer-readable medium may be, for example, any device that can tangibly contain, store, communicate, or transport a program for use by or in connection with any processor. The medium may be, for example, an electronic, optical, electromagnetic, or semiconductor device. Other suitable media may also be available.

[0179] It will be appreciated that aspects can be implemented in any convenient form. For example, aspects can be implemented by a suitable computer program that can be carried on a suitable carrier medium, which can be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communication signal). Aspects can also be implemented using a suitable apparatus, which can be in the form of a programmable computer that executes a computer program configured to implement the methods and / or techniques disclosed herein. Aspects can be combined such that features described in connection with one aspect can be implemented in another aspect.

[0180] The above implementations are set forth to facilitate understanding of the present application and are not intended to be limiting. To the contrary, the present application covers various modifications and equivalent arrangements included within the scope of the appended claims, which scope is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures permitted under law.

Claims

1. 1. A computer-implemented method for video content screening using a video screening model trained using self-supervised training, comprising: screening a current video in response to automatically identified screening data obtained from a trained video screening model trained using self-supervised training, the screening data indicating a degree of similarity between the current video and a reference video, the self-supervised training comprising: and obtaining the trained video screening model by training an untrained video screening model using an automatically generated training dataset, wherein the automatically generated training dataset comprises: obtaining automatically generated predicate screening data indicative of predicate time segments in a training video and corresponding reference time segments in the reference video; obtaining candidate screening data for an extended time segment from the training video, the extended time segment including the predicate time segment and at least one frame from the training video adjacent to the predicate time segment, the candidate screening data indicating similarity between a screening frame from the reference video and a candidate sub-frame, the candidate sub-frame being a spatial portion of a candidate frame from the extended time segment; in response to determining that the determined similarity between the candidate subframe and the screening frame is equal to or greater than a predefined similarity threshold, including in the automatically generated training data set training example data indicative of the similarity between the candidate subframe and the screening frame; Automatically generated by screening 11. A computer-implemented method comprising:

2. The self-supervised training obtaining the training video from a plurality of training videos corresponding to the automatically generated predicate screening data; obtaining the reference video from a plurality of reference videos; obtaining the predicate screening data, the predicate screening data having been previously generated by screening the plurality of training videos with respect to the plurality of reference videos using a previously trained video screening model using a first predefined confidence threshold; The method of claim 1 , comprising:

3. 3. The method of claim 2, wherein the automatically identified screening data is obtained from the trained video screening model, which was trained using self-supervised training using the first predefined confidence threshold.

4. including the training example data in the automatically generated training data set; obtaining filtered screening data by screening the training video with respect to the reference video using the previously trained video screening model using the first predefined confidence threshold; including the training example data in the automatically generated training data set in response to determining that no data indicative of the similarity between the candidate subframe and the screening frame is present in the filtered screening data; and omitting the training example data from the automatically generated training data set in response to determining that the filtered screening data includes data indicative of the similarity between the candidate subframe and the screening frame; and The method of claim 2 , comprising:

5. 3. The method of claim 2, wherein obtaining the candidate screening data comprises obtaining the candidate screening data from the previously trained video screening model using a second predefined confidence threshold that is lower than the first predefined confidence threshold.

6. including the training example data in the automatically generated training data set; acquiring a portion of the screening frame; obtaining a fingerprint of the portion of the screening frame; obtaining a fingerprint of the screening frame; determining a similarity value indicative of a measure of similarity between the fingerprint of the portion of the screening frame and the fingerprint of the screening frame; including the training example data in the automatically generated training data set in response to determining that the similarity value is less than a predefined similarity threshold; omitting the training example data from the automatically generated training data set in response to determining that the similarity value is equal to or greater than the predefined similarity threshold. and The method of claim 1 , comprising:

7. including the training example data in the automatically generated training data set; omitting the training example data from the automatically generated training data set in response to determining that the size of the candidate subframe is less than a defined minimum size. The method of claim 1 , comprising:

8. The method of any of claims 1 to 7, wherein the self-supervised training comprises removing bias from the automatically generated training dataset.

9. 1. A computer-implemented method for video content screening using a video screening model trained using self-supervised training, comprising: obtaining an input video; obtaining screening data indicative of automatically identified associations between the input video and a reference video from a trained video screening model trained using self-supervised training, wherein the self-supervised training includes: Obtaining an automatically generated training dataset, Obtaining training videos and obtaining the reference video; obtaining predicate screening data generated using a first previously trained video screening model with respect to the training video and the reference video, the predicate screening data indicating predicate time segments in the training video and corresponding reference time segments in the reference video; obtaining, from a second previously trained video screening model, candidate screening data for an extended time segment from the training video, the extended time segment including the predicate time segment and at least one of a frame from the training video preceding the predicate time segment or a frame from the training video following the predicate time segment, the candidate screening data indicating an association between a screening frame from the reference video and a candidate sub-frame, the candidate sub-frame being a spatial portion of a candidate frame from the extended time segment; the determined similarity value for the candidate subframe and the screening frame is greater than or equal to a first predefined similarity threshold; data indicative of the association between the candidate subframe and the screening frame is not present in filtered screening data obtained from the first previously trained video screening model with respect to the training video and the reference video; a similarity value for the screening frame and a spatial portion of the screening frame is less than a second predefined similarity threshold; In response to the determination that including in the automatically generated training data set training example data indicative of the association between the candidate subframes and the screening frames; obtaining the automatically generated training data set, removing bias from the automatically generated training dataset; and Obtaining an untrained video screening model; and obtaining the trained video screening model by training the untrained screening model using the automatically generated training dataset; obtaining said screening data, identifying the input video as a screened video in response to obtaining the screening data; 11. A computer-implemented method comprising:

10. the input video is one of a plurality of input videos; the reference video is one of a plurality of reference videos; obtaining the screening data includes obtaining screening data for the plurality of input videos with respect to the plurality of reference videos; 10. The computer-implemented method of claim 9.

11. the similarity between the input video and the reference video is a similarity between an automatically generated fingerprint for the input video and an automatically generated fingerprint for the reference video; the similarity between the screening frame and the candidate subframe is a similarity between an automatically generated fingerprint for the screening frame and an automatically generated fingerprint for the candidate subframe.

10. The computer-implemented method of claim 9.

12. the predicate screening data is generated using the first previously trained video screening model with a first predefined confidence threshold; obtaining the candidate screening data from the second previously trained video screening model includes obtaining the candidate screening data from the first previously trained video screening model using a second predefined confidence threshold, the first predefined confidence threshold being greater than the second predefined confidence threshold; 10. The computer-implemented method of claim 9.

13. 13. The computer-implemented method of claim 9, wherein identifying the input video as a screened video comprises generating fingerprint data of a portion of the input video indicated by the screening data.

14. 1. A system for training a video screening model using self-supervised training, comprising: a non-transitory memory for storing instructions; Processor and wherein the processor executes the instructions to obtain a trained video screening model; To obtain the trained video screening model, the processor executes the instructions to train an untrained video screening model using a training dataset, and to automatically generate the training dataset, the processor: obtaining automatically generated predicate screening data indicative of predicate time segments in the training video and corresponding reference time segments in the reference video; obtaining candidate screening data for an extended time segment from the training video, the extended time segment including the predicate time segment and at least one frame from the training video adjacent to the predicate time segment, the candidate screening data indicating similarity between a screening frame from the reference video and a candidate sub-frame, the candidate sub-frame being a spatial portion of a candidate frame from the extended time segment; in response to determining that the determined similarity between the candidate subframe and the screening frame is equal to or greater than a predefined similarity threshold, including in the automatically generated training data set training example data indicative of the similarity between the candidate subframe and the screening frame; and executing the instructions to perform the steps.

15. To automatically generate the training data set, the processor: obtaining the training video from a plurality of training videos corresponding to the automatically generated predicate screening data; obtaining the reference video from a plurality of reference videos; obtaining the predicate screening data, the predicate screening data having been previously generated by screening the plurality of training videos with respect to the plurality of reference videos using a previously trained video screening model using a first predefined confidence threshold; 15. The system of claim 14, wherein the system executes instructions to:

16. screening a current video in response to automatically identified screening data obtained from the trained video screening model, the screening data indicating similarity between the current video and a reference video, the screening data obtained from the trained video screening model using the first predefined confidence threshold. The system of claim 15 further comprising:

17. screening the current video, identifying the current video as a screening video; generating fingerprint data of the portion of the current video indicated by the screening data; comparing the fingerprint data of the portion of the current video with fingerprint data of the reference video to determine whether the portion of the current video is similar to a respective portion of the respective reference video from the reference video; 17. The system of claim 16, comprising:

18. To include the training example data in the automatically generated training data set, the processor: obtaining filtered screening data by screening the training video with respect to the reference video using the previously trained video screening model using the first predefined confidence threshold; including the training example data in the automatically generated training data set in response to determining that no data indicative of the similarity between the candidate subframe and the screening frame is present in the filtered screening data; and omitting the training example data from the automatically generated training data set in response to determining that the filtered screening data includes data indicative of the similarity between the candidate subframe and the screening frame; and 17. The system of claim 15 or 16, wherein the system executes the instructions to:

19. 17. The system of claim 15 or 16, wherein the processor executes the instructions to obtain the candidate screening data from the previously trained video screening model using a second predefined confidence threshold that is lower than the first predefined confidence threshold.

20. To include the training example data in the automatically generated training data set, the processor: acquiring a portion of the screening frame; obtaining a fingerprint of the portion of the screening frame; obtaining a fingerprint of the screening frame; determining a similarity value indicative of a measure of similarity between the fingerprint of the portion of the screening frame and the fingerprint of the screening frame; including the training example data in the automatically generated training data set in response to determining that the similarity value is less than a predefined similarity threshold; omitting the training example data from the automatically generated training data set in response to determining that the similarity value is equal to or greater than the predefined similarity threshold.

16. The system of claim 14 or 15, wherein the system executes the instructions to:

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and program

    JP2017033175A

  • Event evaluation support system, event evaluation support device, and event evaluation support program

    JP2018206085A

  • Autonomous task performance based on visual embeddings

    WO2021015883A1