Image processing system and method therefor
By utilizing an artificial neural network deep learning model for encoding and decoding images, the system effectively hides and restores secret messages in real-time, addressing bandwidth and performance issues in existing systems.
Patent Information
- Application Number
- PCT/KR2023/019432
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-05
AI Technical Summary
Existing image processing systems face challenges in efficiently hiding and restoring secret messages in images without significantly increasing network bandwidth or compromising real-time performance.
The system employs an artificial neural network deep learning model for both encoding and decoding images, allowing for real-time hiding and restoration of secret messages within images, using techniques such as dividing images into sub-images and selecting appropriate channels for message insertion.
This approach enables efficient hiding and restoration of secret messages without increasing image size, thus reducing network bandwidth requirements and maintaining real-time capabilities.
Smart Images

Figure KR2023019432_05062025_PF_FP_ABST
Abstract
Description
Image processing system and method thereof
[0001] The present disclosure relates to an image processing system and method thereof, and more particularly, to a device and method for encoding and decoding an image using a steganography technique capable of hiding and restoring information in an image.
[0002] Steganography refers to the technique or process of hiding (or concealing) information (e.g., a secret message) in media data. For example, by hiding a secret message within the pixel values of an image, the image's appearance can be minimized, or by hiding the secret message within a music file or audio stream, the secret message can be hidden in a way that makes it nearly undetectable. In other words, steganography is a method of concealing information so that third parties (or users) cannot detect the presence of a secret message in the media using their five senses.
[0003] Conventionally, when hiding secret messages in media data, the Random LSB (Least Significant Bit) method was used, inserting the secret message into a random bit. Then, after inserting the secret message, lossless compression techniques were used to convert the data back into media data. This increased the size of the media data containing the secret message, requiring significant network bandwidth for transmission and reception. Furthermore, the increased size of the media data containing the secret message makes it unsuitable for real-time environments.
[0004] The present disclosure is proposed to solve the above-mentioned problems and various problems related thereto, and the purpose of the present disclosure is to provide a device and method for hiding a secret message in an image in an encoder and restoring the secret message from an image captured by a camera in a decoder.
[0005] Another object of the present invention is to provide a device and method for hiding a secret message in an image using an encoder that applies an artificial neural network deep learning model, and for restoring the secret message hidden in the image using a decoder that applies an artificial neural network deep learning model.
[0006] Another object of the present disclosure is to provide a device and method for hiding a secret message in an image by applying an artificial neural network deep learning model in real time and non-real time, and for restoring the hidden secret message from the image by applying the artificial neural network deep learning model.
[0007] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.
[0008] In order to achieve the above-described purpose and other advantages, the image processing system may include an encoder that generates a second image by hiding a secret message in a first image in a first video, a data hiding device that transmits a second video including the second image, an image output device that receives and plays the second video, a decoder for restoring the secret message, a hidden data restoration device that photographs the image output device while the second video is being played with at least one camera, and the secret message is restored by the decoder from the photographed image, and provides a user experience based on the restored secret message, and a server that generates the encoder and the decoder by applying an artificial neural network deep learning model, provides the encoder to the data hiding device, and provides the decoder to the hidden data restoration device.
[0009] According to embodiments, an encoder included in the data hiding device may divide the first image into a plurality of sub-images, select one or more first sub-images from among the plurality of sub-images into which the secret message is to be inserted, select a channel from among channels constituting pixels of the one or more selected first sub-images into which the secret message is to be inserted, copy the secret message to a number and a size identical to that of the one or more first sub-images to generate one or more sub-secret messages, apply an encoding model learned from data of the selected channels of the one or more first sub-images and data of the one or more sub-secret messages to generate one or more second sub-images into which the secret message is inserted, and replace one or more first sub-images from among the plurality of sub-images of the first image with the one or more second sub-images to generate the second image into which the secret message is inserted.
[0010] According to embodiments, the encoder may select a B channel as a channel into which to insert the secret message, if the channels constituting the pixels of the selected one or more first sub-images are RGB channels.
[0011] According to embodiments, the encoder may select a Y channel as a channel into which to insert the secret message, if the channels constituting the pixels of the selected one or more first sub-images are YUV channels.
[0012] According to embodiments, the secret message includes content identification information, and the content identification information can be used to identify content for providing the user experience.
[0013] According to embodiments, the secret message further includes time information, and the time information may include at least one of information for identifying a playback time of the content identified by the content identification information, information for identifying a playback time, or information for identifying a playback end time.
[0014] According to embodiments, the hidden data restoration device may include a storage unit that stores one or more contents to provide the user experience, an image input unit that photographs an image output device on which the second video is played, an analysis unit that detects an edge of the image output device from an image photographed by the image input unit, the decoder that restores the secret message from an image corresponding to the edge of the image output device detected by the analysis unit, and an output unit that extracts content from the storage unit based on content identification information included in the restored secret message to provide a user experience.
[0015] According to embodiments, the analysis unit extracts an outline of the captured image, extracts closed polygons based on the extracted outline, selects one or more rectangles including one or more closed polygons by merging closed polygons close to the center of the image, compares the one or more rectangles with one or more threshold values, determines at least one of the one or more rectangles as an edge candidate of the image output device, and when there are multiple edge candidates of the image output device, the rectangles are sorted in order of their area to be larger and then decoding is applied to each of them to finally determine the edge of the image output device.
[0016] According to embodiments, the decoder of the hidden data restoration device may divide an image corresponding to an edge of the image output device into a plurality of sub-images, select one or more third sub-images in which the secret message is inserted from among the plurality of sub-images, select a channel in which the secret message is inserted from among channels constituting pixels of the one or more selected third sub-images, and obtain the secret message by applying a decoding model learned from data of the selected channel of the one or more third sub-images.
[0017] According to embodiments, the number and size of one or more first sub-images and the selected channel selected for insertion of the secret message in the encoder of the data hiding device may be the same as the number and size of one or more third sub-images and the selected channel in which the secret message is inserted in the decoder of the hidden data restoration device.
[0018] According to embodiments, the server may include an encoder, a decoder, and a loss function calculation unit.
[0019] According to embodiments, the encoder included in the server divides a third image in a randomly input video into a plurality of sub-images, selects one or more fourth sub-images from among the plurality of sub-images into which a randomly input secret message is to be inserted, and selects a channel from among channels constituting pixels of the one or more selected fourth sub-images into which the input secret message is to be inserted.
[0020] According to embodiments, the input secret message may be copied to generate one or more sub-secret messages in the same number and size as the one or more fourth sub-images, an encoding model learned up to now from data of a selected channel of the one or more fourth sub-images and data of the one or more sub-secret messages may be applied to generate one or more fifth sub-images in which the input secret message is inserted, and one or more fourth sub-images among a plurality of sub-images of the third image may be replaced with the one or more fifth sub-images to generate the fourth image in which the input secret message is inserted.
[0021] According to embodiments, the decoder included in the server can generate a fifth image by changing data of the fourth image within a certain range, and obtain an output secret message by applying a decoding model learned so far from data of a selected channel of one or more sub-images in which the input secret message is inserted among a plurality of sub-images of the fifth image.
[0022] According to embodiments, the loss function calculation unit may calculate a loss function that evaluates the difference between the input secret message and the output secret message, evaluates the difference between the third image and the fourth image, and continuously updates the encoder and the decoder in a direction in which the loss function is minimized, in order to determine an encoder to be provided to the data hiding device and a decoder to be provided to the hidden data restoration device.
[0023] According to embodiments, the hidden data recovery device may be smart glasses.
[0024] According to embodiments, an encoding method in an image processing system may include the steps of dividing a first image in an original video into a plurality of sub-images, selecting one or more first sub-images from among the plurality of sub-images into which a secret message is to be inserted, selecting a channel from among channels constituting pixels of the one or more selected first sub-images into which the secret message is to be inserted, generating one or more sub-secret messages by copying the secret message in a number and size identical to that of the one or more first sub-images, generating one or more second sub-images into which the secret message is inserted by applying an encoding model learned from data of a selected channel of the one or more first sub-images and data of the one or more sub-secret messages, and generating a second image into which the secret message is inserted by replacing one or more first sub-images from among the plurality of sub-images of the original video with the one or more second sub-images.
[0025] According to embodiments, the channel selection step may select a B channel as a channel into which the secret message is to be inserted, if the channels constituting the pixels of the selected one or more first sub-images are RGB channels.
[0026] According to embodiments, the channel selection step may select a Y channel as a channel into which the secret message is to be inserted, if the channels constituting the pixels of the selected one or more first sub-images are YUV channels.
[0027] According to embodiments, the secret message hiding and transmission may be performed in real time or non-real time.
[0028] According to embodiments, the secret message includes content identification information, and the content identification information can be used to identify content for providing the user experience.
[0029] According to embodiments, the secret message further includes time information, and the time information may include at least one of information for identifying a playback time of the content identified by the content identification information, information for identifying a playback time, or information for identifying a playback end time.
[0030] According to embodiments, a decoding method in an image processing system may include the steps of photographing an image output device on which a video with a secret message inserted is played by at least one camera, detecting an edge of the image output device from the photographed image, restoring the secret message by applying a decoding model learned from an image corresponding to the detected edge of the image output device, and extracting content identified by content identification information included in the restored secret message from a storage unit in which one or more contents for providing a user experience are stored, thereby providing the user experience.
[0031] According to embodiments, the edge detection step of the image output device may include the steps of extracting an outline of the captured image, extracting closed polygons based on the extracted outline, selecting one or more rectangles including one or more closed polygons by merging closed polygons close to the center of the image, comparing the one or more rectangles with one or more threshold values to determine at least one of the one or more rectangles as an edge candidate of the image output device, and, if there are a plurality of edge candidates of the image output device, sorting the rectangles in order of their areas and applying a learned decoding model to each of them to finally determine the edge of the image output device.
[0032] According to embodiments, the secret message restoration step may include the steps of dividing an image corresponding to an edge of the image output device into a plurality of sub-images, selecting one or more sub-images in which the secret message is inserted from among the plurality of sub-images, selecting a channel in which the secret message is inserted from among channels constituting pixels of the one or more selected sub-images, and applying a decoding model learned from data of the selected channel of the one or more sub-images to obtain the secret message.
[0033] The device and method according to the embodiments enable a user to obtain useful information simply by watching a video being played through a device equipped with a camera without any separate operation.
[0034] The device and method according to the embodiments enable a user to obtain useful information simply by watching a video being played while wearing glasses equipped with a camera without any separate operation.
[0035] The device and method according to the embodiments allows one to watch a movie in an outdoor movie theater such as a drive-in theater and simultaneously see and hear movie subtitles and sounds through glasses equipped with cameras.
[0036] The device and method according to the embodiments can be used to view exhibited works in an exhibition such as a museum or art gallery, and simultaneously see and hear additional explanations or supplementary information about the works through glasses equipped with cameras.
[0037] In addition to the technical effects explicitly mentioned above, effects that are obvious to those skilled in the art can also be inferred from the entire description of this specification.
[0038] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments.
[0039] FIG. 1 is a block diagram showing an example of an image processing system according to embodiments.
[0040] FIG. 2 is a diagram showing an example of an image processing system for hiding a secret message in a cover video in non-real time according to embodiments.
[0041] FIG. 3 is a diagram showing an example of an image processing system for hiding a secret message in a cover video in real time according to embodiments.
[0042] FIG. 4 is a diagram showing an example of a method for generating a stego video by hiding a secret message in a cover video in an encoder according to embodiments.
[0043] FIG. 5 is a diagram showing an example of a method for recovering a secret message hidden in a cover video in a decoder according to embodiments.
[0044] FIG. 6 is a diagram showing an example of a learning process for creating an encoder and a decoder by applying an artificial neural network deep learning model in a server according to embodiments.
[0045] FIG. 7 is a drawing showing an example of a method for extracting the boundary of an image output device in an analysis unit according to embodiments.
[0046] FIGS. 8(a) to 8(f) are drawings showing examples of each process of extracting the boundary of an image output device in an analysis unit according to embodiments.
[0047] Figures 9 to 14 are diagrams showing examples of user experiences (UX) provided based on restored secret messages according to embodiments.
[0048] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be assigned the same reference numbers, and redundant descriptions thereof will be omitted. It should be noted that the following embodiments are intended only to concretize the present disclosure and do not limit or restrict the scope of the present disclosure. Anything that a specialist in the technical field to which the present disclosure pertains can easily infer from the detailed description and embodiments of the present disclosure is interpreted as falling within the scope of the present document.
[0049] The detailed description in this document should not be construed in any way as restrictive, but rather as illustrative. The scope of this document should be determined by a reasonable interpretation of the appended claims, and all changes within the scope of equivalents herein are intended to be included within the scope of this document.
[0050] Meanwhile, the blocks of the attached block diagram or the steps of the flowchart may be implemented directly in hardware, implemented as software modules executed by hardware, or implemented by a combination of these. In addition, the blocks of the attached block diagram or the steps of the flowchart may be interpreted as computer program instructions that are loaded into the processor or memory of a data processing device such as a general-purpose computer, a special-purpose computer, a portable notebook computer, or a network computer to perform designated functions. Since these computer program instructions may be stored in a memory provided in a computer device or a computer-readable memory, the functions described in the blocks of the block diagram or the steps of the flowchart may be produced as a product that includes a command means for performing the same. In addition, each block or each step may represent a module, segment, or part of code that includes one or more executable instructions for performing a specific logical function(s). Furthermore, in some alternative embodiments, the functions mentioned in the blocks or steps may be executed out of the specified order. For example, two blocks or steps shown in succession may be performed substantially simultaneously, or may be performed in reverse order, or in some cases, some blocks or steps may be performed with some blocks or steps omitted.
[0051] The present disclosure provides an encoding device that conceals an invisible secret message in a cover image using an artificial neural network deep learning model, and a decoding device that captures the secret message hidden in a transmitted image through a camera provided in the decoding device and restores the hidden message using an artificial neural network deep learning model. In the present disclosure, the artificial neural network deep learning model is also referred to as an artificial neural network steganography deep learning model.
[0052] In the present disclosure, a secret message hidden in an image may be referred to as a message, a secret message, information, additional information data, or secret data.
[0053] In this disclosure, the term "image" may be either a still image or a moving image. "Video" includes locally stored video, streaming video, live broadcasts, etc., while "still image" includes photographs, drawings, etc. In this disclosure, the term "moving image" is used interchangeably with "video."
[0054] Additionally, in the present disclosure, a video may be in the form of a file, and a cover video refers to a video in which a secret message is not hidden, and may also be referred to as an original video. Furthermore, a stego video refers to a video in which a secret message is hidden. In other words, a cover video refers to a video for inserting a secret message, and a stego video refers to a video created by inserting a secret message into a cover video (i.e., an original video).
[0055] In the present disclosure, the hidden data recovery device may be any portable device equipped with at least one camera and a display function. In the present disclosure, the hidden data recovery device is also referred to as a decoding device.
[0056] For example, in the present disclosure, the hidden data recovery device may be a smart phone (or mobile phone) or one of XR (AR / MR / VR) glasses. In the present disclosure, XR glasses will be referred to as smart glasses.
[0057] The present disclosure describes a smart glass as an embodiment of a hidden data recovery device. In the present disclosure, the smart glass may be referred to as a glass.
[0058] The present disclosure is directed to hiding secret messages in not only locally stored videos but also streaming videos and live broadcasts, and to enabling the restoration of hidden secret messages through smart glasses.
[0059] In this disclosure, non-real-time service means that the data hiding process and the process of restoring the hidden data are not performed continuously in real time.
[0060] In the present disclosure, a real-time service means that a data hiding process and a process of restoring hidden data are performed continuously, and an example of an embodiment is a real-time broadcasting service.
[0061] FIG. 1 is a block diagram showing an example of an image processing system according to embodiments.
[0062] An image processing system according to embodiments may include a data hiding device (110), an image output device (120), a hidden data restoration device (130), and a server (140).
[0063] According to embodiments, the data hiding device (110) may include an encoder that generates a stego video by hiding a secret message in a cover video. In the present disclosure, the data hiding device (110) may be referred to as an encoding device.
[0064] According to embodiments, the data hiding device (110) may include at least one of a content provider, a content distribution provider, an advertisement provider, and a cloud server, and the encoder may be included in at least one of the content provider, the content distribution provider, the advertisement provider, or the cloud server. The cloud server may be a server that provides cloud services and may be a service provider.
[0065] According to embodiments, the data hiding device (110) can transmit a secret message hidden in a cover video in real time or non-real time.
[0066] According to embodiments, when hiding a secret message in a locally stored video or content that does not require real-time transmission, the secret message hiding is performed in non-real time, and when hiding a secret message in a streaming video or live broadcast, the secret message hiding is performed in real time.
[0067] The present disclosure may be referred to as a non-real-time service when an encoder is included in a content provision unit, the content provision unit hides a secret message in the content, and then provides the content with the secret message hidden to a video output device through a content distribution unit, and may be referred to as a real-time service when an encoder is included in a content distribution unit or a cloud server, and then the content distribution unit or the cloud server hides a secret message in a streaming video or live broadcast, and then provides the content with the secret message hidden directly to a video output device.
[0068] In this disclosure, the term "image" may be either a still image or a moving image. "Video" includes locally stored video, streaming video, live broadcasts, etc., while "still image" includes photographs, drawings, etc. In this disclosure, the term "moving image" is used interchangeably with "video."
[0069] According to embodiments, the content provider may be a broadcasting station, etc., the content distribution provider may be a home shopping network, etc., and the advertising provider may be an advertising production company or an advertiser, etc.
[0070] According to embodiments, the video output device (120) is a device that processes and displays a cover video or stego video provided from the data hiding device (100).
[0071] In the present disclosure, the image output device (120) may be any device capable of processing and displaying images. For example, the image output device (120) may be a television (TV), a smart phone, a laptop, a tablet PC (Personal Computer), a desktop PC, a laptop PC, etc. In the present disclosure, the image output device (120) may also be a digital signage that displays advertising images or general images in places such as subway stations, department stores, bus stops, airports, inside subways, shopping malls, building rooftops or exterior walls, and outdoors.
[0072] In the present disclosure, the encoder may be included in the video output device (120). In this case, the video output device (120) may insert a secret message into the cover video. For example, the video output device (120) may insert a secret message into the video being played in real time.
[0073] According to embodiments, the hidden data recovery device (130) may include a decoder that recovers a secret message hidden in a cover video.
[0074] In the present disclosure, the hidden data recovery device (130) may be any device equipped with a camera capable of capturing images and a function capable of displaying images, text, information, etc., or outputting them as audio. In the present disclosure, the hidden data recovery device (130) is described as a smart glass, such as AR glasses, as an example.
[0075] According to embodiments, the server (140) generates an encoder and a decoder through a deep learning learning process. Then, the encoder generated through the deep learning learning process is provided to a data hiding device (110) and / or an image output device (120), and the decoder is provided to a hidden data restoration device (130).
[0076] For this purpose, the server (140), the data hiding device (110), the image output device (120), and / or the hidden data restoration device (130) may be connected via a wired / wireless network.
[0077] Additionally, a connection may be made between a data hiding device (110) and an image output device (120) or between an image output device (120) and a hidden data restoration device (130) via a wired / wireless network.
[0078] According to embodiments, the wired network may include various wired communication modules such as a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, as well as various cable communication modules such as a Universal Serial Bus (USB), a High Definition Multimedia Interface (HDMI), a Digital Visual Interface (DVI), RS-232 (recommended standard232), power line communication, or plain old telephone service (POTS).
[0079] According to embodiments, the wireless network may include a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), 4G, 5G, and 6G, in addition to a WiFi module and a Wireless broadband module.
[0080] Additionally, the wireless network may include a short-range communication module, and the short-range communication module may support short-range communication using at least one of Bluetooth, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi (Wireless-Fidelity), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies.
[0081] Figure 2 is a diagram illustrating an example of an image processing system for hiding a secret message in a cover video in non-real time according to embodiments. In the present disclosure, a non-real-time service means that the data hiding process and data restoration process do not occur continuously in real time.
[0082] According to embodiments, the data hiding device (110) in the image processing system of FIG. 2 may include a content providing unit (211) that provides content, an encoder (213) that hides a secret message in a cover video, and a content distribution unit (215) that distributes the cover video or the stego video. Here, the content providing and content distribution may be performed by the content providing unit (211) and the content distribution unit (215), respectively, as in FIG. 2, or the content providing unit (211) may perform both the content providing and content distribution. That is, the content providing unit (211) and the content distribution unit (215) may be the same.
[0083] According to embodiments, the content provider (211) may be referred to as a content provider, and the content distribution unit (213) may be referred to as a content distributor.
[0084] According to embodiments, the content provision unit (211) creates or produces various types of content, and the content distribution unit (215) provides or distributes the content to users (e.g., video output devices (120)) through various routes or platforms.
[0085] In the present disclosure, the encoder (213) may be provided in at least one of the content provision unit (211) and the content distribution unit (215), or may be located in a cloud server via a wired / wireless network. The wired / wireless network has been described in detail in FIG. 1, and thus will be omitted here to avoid redundant description.
[0086] In the present disclosure, the encoder (213) is positioned in the content providing unit (211) as an example.
[0087] According to embodiments, the content provider (211) of the data hiding device (110) provides a secret message and a cover video to be hidden to the encoder (213). Alternatively, the content provider (211) may provide the secret message and the cover video to the encoder (213) provided therein.
[0088] In the present disclosure, the encoder (213) is, in one embodiment, a video encoder of SaaS (Service as a Service) or SDK (Software Development Kit).
[0089] According to embodiments, SaaS represents a business model that provides software as a service. In this model, the software is typically accessed through a web browser, and users can utilize the service via the web without having to purchase or install the software. In other words, SaaS software can be accessed from anywhere with an internet connection, and the provider automatically manages and updates the software without requiring users to update it separately.
[0090] According to embodiments, an SDK represents a set of tools for software development, typically including the tools, libraries, documentation, and example code necessary for developers to create and run software. SDKs are typically specific to a specific programming language or environment and can be used with specific platforms or frameworks to simplify and accelerate the development process.
[0091] In the present disclosure, the encoder (213) creates a stego video by hiding a secret message (Secret Message) in a cover video (Video Source) based on the encoding of an artificial neural network steganography deep learning model.
[0092] More specifically, the encoder (213) generates a stego video by hiding a secret message in each frame of the cover video (Video Source) using an artificial neural network steganography deep learning model. In this case, the stego video is in the form of a file, as an example.
[0093] The stego video generated by the above encoder (213) is output to the content distribution unit (215). That is, the content provision unit (211) provides the stego video to the content distribution unit (215).
[0094] According to embodiments, the content distribution unit (215) transmits the stego video to the video output device (120). At this time, the stego video may be transmitted in non-real time.
[0095] The above video output device (120) receives and processes a stego video and then displays (or plays) it. At this time, a secret message is hidden in the stego video displayed (or played) on the video output device (120), but the user viewing the video output device is unaware of the existence of the secret message.
[0096] That is, when a hidden data recovery device (130) equipped with a camera captures an image output device (120), a secret message hidden in the image displayed (or played) by the image output device (120) or content corresponding to the secret message is displayed on the screen of the hidden data recovery device (130) or output as sound. Here, the capture may be performed automatically or manually at the user's instruction.
[0097] For example, if the hidden data recovery device (130) is a smart glass, when the user wears the smart glass and looks at the image output device (120), a secret message hidden in the image displayed (or played) on the image output device (120) or content corresponding to the secret message can be displayed on the screen of the smart glass or output as sound. This assumes that the location the user looks at while wearing the smart glass is automatically captured by the camera provided in the smart glass. If the capture is manual, the user can also directly instruct the capture.
[0098] As another example, if the hidden data recovery device (130) is a smart phone, when a user takes a picture of the image output device (120) with the smart phone, a secret message hidden in the image displayed (or played) on the image output device (120) or content corresponding to the secret message may be displayed on the screen of the smart phone or output as sound.
[0099] To this end, the hidden data restoration device (130) may include an image input unit (231) capable of capturing an image being displayed on an image output device (120), an analysis unit (233) for analyzing an image boundary from an image captured by the image input unit (231), a decoder (235) for restoring a secret message included in the captured image based on the analysis result of the analysis unit (233), a storage unit (237) for storing content corresponding to the restored secret message, and an output unit (239) for displaying the content provided from the storage unit (237) in the form of a video and / or sound. According to embodiments, the video input unit (231) may be one or more cameras, and the output unit (239) may include a display unit and / or an audio output unit. In addition, the content corresponding to the restored secret message may be in the form of a video, an audio, or a form including both. The above storage unit (237) may be referred to as a database, and the output unit (239) may be referred to as an out-of-screen.
[0100] That is, if the hidden data recovery device (130) is a smart glass, the hidden data recovery device (130) worn by the user can capture an image being played on the image output device (120). Here, the capture is performed through an image input unit (231) composed of one or more cameras. At this time, the capture may be performed automatically or manually at the user's instruction.
[0101] At this time, due to the position or angle of view of the camera, the image captured by the image input unit (231) may include not only the image being played back on the image output device (120) but also the surrounding background.
[0102] In this case, the analysis unit (233) analyzes the image boundary in real time, and extracts only the image corresponding to the image being played on the image output device (120) from the captured image based on the analysis result, and provides the extracted image to the decoder (235). The decoder (235) applies the decoding of the artificial neural network steganography deep learning model to the image provided by the analysis unit (233) to restore (decode) the secret message hidden in the image being played on the image output device (120). That is, the hidden data restoration device (130) worn by the user, for example, smart glasses equipped with a camera, photographs the image output device (120) and the surrounding background through the camera. Then, based on the captured image, the image being played on the screen of the image output device (120) is recognized, the recognized image is analyzed in real time, and the secret message hidden in the image being played on the screen of the image output device (120) is restored (decoded) through the artificial neural network steganography deep learning model.
[0103] In the present disclosure, the analysis unit (233) detects the exact edge of the image output device (120) from an image including the image output device (120) to provide only the image corresponding to the image being played back from among the captured images to the decoder (235), extracts the image corresponding to the image being played back from the captured image, and then provides the extracted image to the decoder (235).
[0104] According to embodiments, when a secret message is restored in the decoder (235), content corresponding to the restored secret message may be extracted from the storage unit (237) and provided to the output unit (239). The output unit (239) may display the content provided from the storage unit (237) on a screen and / or output it as audio.
[0105] Alternatively, depending on the restored secret message, the output unit (239) may display a UI (User Interface) on the screen or output a sound.
[0106] If the content corresponding to the restored secret message is stored in the storage unit (237), the role of extracting the content corresponding to the restored secret message from the storage unit (237) and providing it to the output unit (239) may be performed by the decoder (235), or may be performed by the storage unit (237), or may be performed by a separate module (e.g., at least one of hardware and software).
[0107] According to embodiments, when inserting a secret message into a cover video in a data hiding device (110), the data hiding device (110) may store information included in the secret message (e.g., content identifier and / or time information) and content identified by the content identifier in the storage unit (237) of the hidden data restoration device (130) via a wired / wireless network.
[0108] The present disclosure may refer to the UI, content, sound, etc. displayed on the output unit (239) according to the secret message as the user experience (UX).
[0109] That is, in the present disclosure, the smart glasses obtain relevant content from the storage unit (237, database) based on a secret message restored from a photographed image when necessary and then provide a user experience through the output unit (239).
[0110] In this disclosure, a secret message may include a content identifier (Contents ID). The content identifier may be used to distinguish and output the corresponding content from the storage unit (237). The content may include advertisements, logos, additional information related to the video being played on the video output device (120), or additional information unrelated to the video being played on the video output device (120).
[0111] In the present disclosure, a secret message may include time information. The time information may include time information (e.g., playback time) related to a video being played on the video output device (120) and / or time information related to content identified by a content identifier included in the secret message. According to embodiments, the time information related to the content identified by the content identifier may include the playback time, playback duration, etc. of the content. For example, the time information may include information for identifying when and for how long the content identified by the content identifier will be displayed on smart glasses. As another example, the playback time of the currently playing video can be determined through the time information.
[0112] In the present disclosure, the secret message may include information that can be displayed directly through the output unit (239) without going through the storage unit (237). For example, it may include a company logo, a UI (User Interface), a URL (Uniform Resource Locator) for visiting a website, a special discount code for a product or location related to a video being played on the video output device (120), etc. However, such information may be stored in the storage unit (237) and provided to the output unit (239) by being identified by a content ID included in the restored secret message.
[0113] As described above, the secret message in the present disclosure may include one or more of content ID, time information, and directly printable information.
[0114] FIG. 3 is a diagram illustrating an example of an image processing system for hiding secret messages in a cover video in real time according to embodiments. In the present disclosure, "real-time service" means a process in which data hiding and the process of restoring hidden data are performed continuously, and an example of an embodiment is a real-time broadcasting service.
[0115] According to embodiments, the data hiding device (110) in the image processing system of FIG. 3 may include a content providing unit (311) for providing content (e.g., a cover video), an advertisement providing unit (312) for providing advertisements, a content distribution unit (313) for distributing real-time video streaming or stego video streaming, an encoder (314) for hiding a secret message in the real-time video streaming (or cover video), and a storage unit (315) for storing the secret message. Here, the content providing and the content distributing may be performed by the content providing unit (311) and the content distributing unit (313) respectively as in FIG. 3, or the content providing unit (311) may perform both the content providing and the content distributing. That is, the content providing unit (311) and the content distributing unit (313) may be the same. In addition, the advertisement providing unit (312) may provide advertisements and / or provide secret messages to be inserted into the cover video or real-time video streaming.
[0116] According to embodiments, the content provider (311) may be referred to as a content provider, the advertisement provider (312) may be referred to as a content advertiser, and the content distribution unit (313) may be referred to as a content distributor.
[0117] According to embodiments, the content provider (311) generates or produces various types of content, and the content distribution unit (313) provides or distributes the content to users (e.g., video output devices (120)) through various routes or platforms. That is, the content provider (311) provides a cover video to the content distribution unit (313), and the content distribution unit (313) can distribute the cover video in the form of VoD streaming or real-time broadcasting. In addition, the advertisement provider (312) can provide an advertisement as a single content to the content provider (311) and / or the content distribution unit (313), or store it in the storage unit (315). The storage unit (315) can be referred to as a database.
[0118] In the present disclosure, the storage unit (315) may be provided within at least one of the content provider unit (311), the advertisement provider unit (312), or the content distribution unit (313), or may be located in a cloud server via a wired / wireless network. In addition, in the present disclosure, the encoder (314) may be provided within at least one of the content provider unit (311), the advertisement provider unit (312), or the content distribution unit (313), or may be located in a cloud server via a wired / wireless network. In the present disclosure, the encoder (314) and the storage unit (315) are located in a cloud server as an embodiment. Since the wired / wireless network has been described in detail in FIG. 1, a detailed description thereof will be omitted here to avoid redundant description.
[0119] In one embodiment of the present disclosure, the content provider (311) and / or the advertisement provider (312) requests the insertion of a secret message into a real-time video streaming (or cover video) and stores the secret message to be inserted and the content identified by the content identifier included in the secret message in the storage unit (315). In one embodiment, the request for the insertion of the secret message is made to a cloud server.
[0120] According to embodiments, the secret message stored in the storage unit (315) may include a content identifier (contents ID) and / or time information related to the content identified by the content identifier. For example, the time information may include the time at which the secret message is to be inserted (i.e., relative to the start of the video). Furthermore, one or more pieces of time information may be stored in relation to the secret message and / or content.
[0121] In addition, even if the advertisement provider (312) did not create the content, it can request the insertion of a secret message it wants to hide in a desired video in the same manner as the content provider (311) and store the secret message in the storage unit (315). In this case, the storage unit (315) can store the secret message desired by the advertisement provider (312) and the content identified by the content identifier included in the secret message.
[0122] And, when the content provider (311) and / or the advertisement provider (312) stores a secret message and related content in the storage unit (315) of the cloud server, the cloud server can store information (e.g., content identifier and / or time information) included in the secret message and content identified by the content identifier in the storage unit (237) of the hidden data restoration device (130) via a wired / wireless network.
[0123] In the present disclosure, the encoder (314) is, as an example, a video encoder of SaaS (Service as a Service) or SDK (Software Development Kit). Since the description of SaaS and SDK is provided in FIG. 2, a detailed description thereof will be omitted here.
[0124] According to embodiments, when a request for insertion of a secret message is received from a content provider (311) or an advertisement provider (312), the encoder (314) generates a stego video stream by hiding the secret message in real-time video streaming provided from a content distributor (313) based on encoding of an artificial neural network steganography deep learning model. Here, the real-time video streaming may be VOD (Video On Demand) streaming or real-time broadcasting (or real-time broadcasting streaming).
[0125] More specifically, the encoder (314) receives a video streaming, such as a VOD streaming or a real-time broadcast being distributed by the content distribution unit (313), and searches (or queries) the storage unit (315) for a secret message to be hidden using information of the video (e.g., a content identifier) and time information (e.g., playback time information). If it is determined that there is a secret message to be hidden, the secret message is hidden in a frame of the received video streaming, i.e., a cover image, through an artificial neural network steganography deep learning model, thereby generating a stego video streaming, i.e., a stego image. The stego video streaming thus generated is provided to the video output device (120). If it is determined that there is no secret message to be hidden, the received video streaming is provided to the video output device (120) without modification. At this time, the video streaming (i.e., stego video streaming or original video streaming) may be packetized (e.g., RTP packets) using a video streaming protocol such as RTP (Real-time Transport Protocol) and then provided to the video output device (120) via a wired / wireless network. If the encoder (314) is located in close proximity to the video output device (120), the video streaming (i.e., stego video streaming or original video streaming) may be provided to the video output device (120) via HDMI (High Definition Multimedia Interface).
[0126] In this disclosure, RTP is a protocol for transmitting data in real time over a network, and provides timestamps, sequence numbers, sender and receiver information, etc. for real-time communication. In this disclosure, HDMI is an interface for transmitting digital video and audio signals.
[0127] The above video output device (120) receives and processes video streaming (i.e., stego video streaming or original video streaming) and then displays (or plays) it. At this time, even if the video streaming received by the video output device (120) is stego video streaming, the user viewing the video output device (120) is unaware of the existence of the secret message.
[0128] At this time, when the hidden data restoration device (130) equipped with a camera captures the image output device (120), the secret message hidden in the image displayed (or played) on the image output device (120) is restored, and the content corresponding to the content identifier included in the restored secret message is displayed on the screen of the hidden data restoration device (130) or output as sound. Here, the shooting may be performed automatically or manually by the user's instruction. At this time, the content corresponding to the secret message may be an advertisement provided by the advertisement provider (312). In this case, the time information may include information for identifying the start time of the advertisement, the display time of the advertisement, etc. Alternatively, the time information may include information such as how many seconds after the user starts viewing the image output device (120) through the glasses, the advertisement identified by the content identifier will be played. The time information may be a timestamp.
[0129] For example, if the hidden data restoration device (130) is smart glasses, when the user wears the smart glasses and looks at the image output device (120), the secret message hidden in the image displayed (or played) on the image output device (120) is restored, and the content identified by the content identifier included in the restored secret message can be displayed on the screen of the smart glasses or output as sound. This assumes that the location the user looks at while wearing the smart glasses is automatically captured by the camera provided in the smart glasses. If the capture is manual, the user can also directly instruct the capture.
[0130] As another example, if the hidden data restoration device (130) is a smart phone, when a user takes a picture of the image output device (120) with the smart phone, the hidden secret message in the image displayed (or played) on the image output device (120) is restored, and the content corresponding to the content identifier included in the restored secret message can be displayed on the screen of the smart phone or output as sound.
[0131] To this end, the hidden data restoration device (130) may include an image input unit (231) capable of capturing an image being displayed on an image output device (120), an analysis unit (233) for analyzing an image boundary from an image captured by the image input unit (231), a decoder (235) for restoring a secret message included in the captured image based on the analysis result of the analysis unit (233), a storage unit (237) for storing content corresponding to the restored secret message, and an output unit (239) for displaying the content provided from the storage unit (237) in the form of a video and / or sound. According to embodiments, the image input unit (231) may be one or more cameras, and the output unit (239) may include a display unit and / or an audio output unit. In addition, the content corresponding to the restored secret message may be in the form of a video, an audio, or a form including both. The storage unit (237) may be referred to as a database.
[0132] In the present disclosure, the storage unit (315) of the data hiding device (110) and the storage unit (237) of the hidden data restoration device (130) are synchronized, as an example. For example, the same information (e.g., content ID, content, time information) is stored in the two storage units (315, 237) with respect to the secret message being hidden and restored.
[0133] That is, if the hidden data recovery device (130) is a smart glass, the hidden data recovery device (130) worn by the user can capture an image being played on the image output device (120). Here, the capture is performed through an image input unit (231) composed of one or more cameras. In the present disclosure, the capture may be performed automatically or manually at the user's instruction.
[0134] At this time, due to the position or angle of view of the camera, the image captured through the image input unit (231) may include not only the image being played back on the image output device (120) but also the surrounding background.
[0135] In this case, the analysis unit (233) analyzes the captured image in real time, and based on the analysis result, extracts only the image corresponding to the image being played on the image output device (120) from the captured image and provides it to the decoder (235). The decoder (235) applies the decoding of an artificial neural network steganography deep learning model to the image provided by the analysis unit (233) to restore (decode) the secret message hidden in the image being played on the image output device (120).
[0136] That is, a hidden data restoration device (130) worn by a user, for example, smart glasses, captures an image output device (120) and the surrounding background through a camera. Then, based on the captured image, the image being played on the screen of the image output device (120) is recognized, the recognized image is analyzed in real time, and a secret message hidden in the image being played on the screen of the image output device (120) is restored (decoded) through an artificial neural network steganography deep learning model.
[0137] In the present disclosure, the analysis unit (233) detects the exact edge of the image output device (120) from an image including the image output device (120) to provide only the image corresponding to the image being played back from among the captured images to the decoder (235), extracts the image corresponding to the image being played back from the captured image, and then provides the extracted image to the decoder (235).
[0138] According to embodiments, when a secret message is restored in the decoder (235), content corresponding to the restored secret message may be extracted from the storage unit (237) and provided to the output unit (239). The output unit (239) may display the content provided from the storage unit (237) on a screen and / or output it as audio.
[0139] Alternatively, depending on the restored secret message, the output unit (239) may display a UI (User Interface) on the screen or output a sound.
[0140] If the content corresponding to the restored secret message is stored in the storage unit (237), the role of extracting the content corresponding to the restored secret message from the storage unit (237) and providing it to the output unit (239) may be performed by the decoder (235), or may be performed by the storage unit (237), or may be performed by a separate module (e.g., at least one of hardware and software).
[0141] The present disclosure may refer to the UI, content, sound, etc. displayed on the output unit (239) according to the secret message as the user experience (UX).
[0142] That is, in the present disclosure, the smart glasses obtain relevant content from the storage unit (237, database) when necessary based on a secret message restored from a captured image, thereby providing a user experience.
[0143] In this disclosure, a secret message may include a content identifier (Contents ID). The content identifier may be used to distinguish and extract the corresponding content from the storage unit (237). The content may include an advertisement, a logo, additional information related to the video being played on the video output device (120), or additional information unrelated to the video being played on the video output device (120).
[0144] In the present disclosure, a secret message may include time information. The time information may include time information (e.g., playback time) related to a video being played on the video output device (120) and / or time information related to content identified by a content identifier included in the secret message. According to embodiments, the time information related to content identified by the content identifier may include the playback time, playback time, etc. of the content. For example, the time information may be information instructing the hidden data restoration device (130) to read and play content (e.g., an advertisement) identified by a content identifier (contents ID) from the storage unit (237) for a few seconds after the time at which the secret message is restored. In addition, the playback time of the currently playing video can be determined through the time information.
[0145] In the present disclosure, the secret message may include information that can be displayed directly through the output unit (239) without going through the storage unit (237). For example, it may include a company logo, a UI (User Interface), a URL (Uniform Resource Locator) for visiting a website, a special discount code for a product or location related to a video being played on the video output device (120), etc. However, such information may be stored in the storage unit (237) and provided to the output unit (239) by being identified by a content ID included in the restored secret message.
[0146] As described above, the secret message in the present disclosure may include one or more of content ID, time information, and directly printable information.
[0147] FIG. 4 is a diagram illustrating an example of a method for generating a stego video by hiding a secret message in a cover video using an encoder according to embodiments. In FIG. 4, the encoder may be the encoder (212) of FIG. 2 or the encoder (314) of FIG. 3 included in a data hiding device (110). In one embodiment, the data hiding device (110) receives an encoder generated by applying an artificial neural network deep learning model from a server (140).
[0148] In Fig. 4, the encoder divides the input cover video into multiple cover images (S411). That is, if the cover video is a real-time or non-real-time video, the video is composed of multiple still images (also called frames or images), and thus the cover video can be divided into multiple cover images. Here, each cover image includes color information and brightness information.
[0149] That is, step S411 extracts all cover images that constitute a cover video having a horizontal size of W and a vertical size of H. According to embodiments, the pixels of each cover image may be composed of three color channels (or RGB channels) in the width W * height H. According to embodiments, the pixels of each cover image may be composed of three brightness channels (or YCbCr / YUV channels) in the width W * height H.
[0150] And, each cover image can be divided into multiple sub-images (S412). For example, in step S412, each cover image is divided into multiple sub-images having a horizontal size of M and a vertical size of N. Here, the number of sub-images may vary depending on the sizes of the sub-images (e.g., M, N). For example, the sizes of M and N may be selected to be close to 200*200. This is one embodiment, and other values may be selected as the sizes of M and N. Fig. 4 shows an example in which a cover image is divided into 3 horizontally and 4 vertically to include 12 sub-images to help those skilled in the art understand.
[0151] In the present disclosure, a secret message can be inserted into one or more sub-images. In other words, inserting a secret message into an image means corrupting the image. Furthermore, when an image is corrupted, depending on the characteristics of the image, some areas may be easily recognized as corrupted by the user, while others may not be. Therefore, the present disclosure divides each cover image into multiple sub-images to detect areas within the image where the secret message is not easily recognized by the user.
[0152] As described above, when the cover image is divided into multiple sub-images in step S412, one or more sub-images are selected from among the multiple sub-images into which a secret message is to be inserted (S413). In the present disclosure, one embodiment selects up to K sub-images.
[0153] And, in step S413, the sub-images selected for inserting the secret message are indicated by T, and the unselected sub-images are indicated by F. If K sub-images are selected for inserting the secret message, T is indicated on K sub-images, and F is indicated on the remaining sub-images.
[0154] That is, step S413 determines whether or not to insert a secret message for each sub-image, and marks T in the sub-images that are decided to be inserted, and marks F in the sub-images that are decided not to be inserted.
[0155] In other words, inserting a secret message into an image alters the image's data, resulting in image distortion. Therefore, to enhance the opacity of the secret message insertion, a sub-image that is difficult for humans to perceive is selected.
[0156] In one embodiment, it may be decided not to embed secret messages in sub-images composed of a single color, as such sub-images are easily recognizable to humans when parts of the image are changed.
[0157] In another embodiment, since images with many different objects are difficult for humans to recognize even if some parts of the image are changed, it may be decided to embed secret messages in sub-images with many different objects.
[0158] And, the sub-image to which encoding is applied to insert a secret message is denoted as T, and the sub-image to which encoding is not applied is denoted as F. At this time, the number of sub-images with inserted secret messages can be up to K (K is 1 or more), and K is variable.
[0159] As described above, once one or more sub-images into which a secret message is to be inserted are determined in step S413, a channel into which the actual secret message is to be inserted is selected (S414). That is, step S414 determines which data the secret message is to be applied to for each sub-image into which the secret message is to be inserted.
[0160] In this disclosure, one embodiment selects Y data of a YUV channel or B data of an RGB channel, taking into account the characteristics of a sub-image. Fig. 4 illustrates an example of selecting Y data of a YUV channel. This is merely an embodiment, and other color systems and data may be selected. That is, a secret message is inserted by modifying the Y data of one or more pixels of a sub-image indicated by T.
[0161] For example, if an image contains a lot of blue colors, even if some of the B data in the RGB channels is changed, it is highly unperceivable. Or, if an image contains a lot of lines, even if the Y data in the YUV channel is changed, it is highly unperceivable.
[0162] Once the selection of sub-images and channels into which secret messages are to be inserted is completed through steps S413 and S414, a data matrix of width W / M * height H / N * K is generated based on the data information of the K sub-images into which secret messages are to be inserted and the channels selected from the K sub-images (S415). That is, the data of the channels selected from the K sub-images having width W / M * height H / N are changed due to the insertion of the secret message during the encoding process.
[0163] In addition, the encoder copies the secret message in the form of an input bit array multiple times to set it to the same size as the sub-image (S416). That is, step S416 converts the input secret message into a binary array and then copies it to the same size as the data generated in step S415 to generate K sub-secret messages having a width of W / M * height of H / N. In other words, a data matrix of width W / M * height H / N * K is created for the secret message by using data copying, etc. This is because the sizes of the sub-image and the secret message are different. For example, if the data size of the sub-image is 500 bytes and the data size of the secret message is 10 bytes, the secret message can be copied 50 times to make it 500 bytes. If the sizes of the sub-image and the secret message are the same, step S416 can be omitted.
[0164] The learned encoding model is applied to the data of the K sub-images generated in step S415 and the data of the K secret messages (or sub-secret messages) generated in step S416 to generate a stego sub-image with a secret message inserted (S417). That is, K stego sub-images having a size of width W / M * height H / N are generated.
[0165] Then, by replacing the K sub-images marked as T in step S413 among the cover images with the K stego sub-images generated in step S417, a stego image is generated (S418). That is, a secret message is hidden in the sub-images marked as T among the cover images, and a completed stego image without a hidden secret message is generated in the sub-images marked as F.
[0166] When the stego images generated in this way are collected, a stego video (or stego video streaming) is generated, and this stego video (or stego video streaming) is provided to and played by a video output device (120). At this time, the secret message may be inserted into all cover images of the cover video, or may be inserted into some of the cover images. Therefore, some of the images included in the stego video may be stego images, and some may be original images without the secret message inserted.
[0167] FIG. 5 is a diagram illustrating an example of a method for recovering a secret message hidden in a cover video in a decoder according to embodiments. In FIG. 5, the decoder may be the decoder (235) of FIG. 2 or the decoder (235) of FIG. 3 included in a hidden data recovery device (130). In one embodiment, the hidden data recovery device (130) receives a decoder generated by applying an artificial neural network deep learning model from a server (140). In FIG. 5, the hidden data recovery device (130) is described as a smart glass such as AR glasses, as an example. However, in the present disclosure, the hidden data recovery device (130) is not limited to smart glasses. That is, any device having one or more cameras and display functions can be a hidden data recovery device.
[0168] In the present disclosure, the decoding process for recovering a secret message from a stego image is performed in the reverse order of the encoding process for hiding the secret message in the cover video (or video streaming) of FIG. 4.
[0169] That is, the image input unit (231) composed of one or more cameras of the hidden data recovery device (130) captures the image input device (120) (e.g., TV) that is playing the stego video (S511). At this time, the capture may be performed automatically or manually by a user's instruction. For example, assuming that the hidden data recovery device (130) is a smart glass, the area that the user looks at while wearing the smart glass, i.e., the image output device (120) and its surroundings, may be automatically captured.
[0170] In this way, the image captured by the image input unit (231) may include not only the image being played back on the image output device (120), but also the image output device (120) and its surroundings. That is, the image input unit (231) may capture the surroundings including the image output device (120) that plays the stego video and scenes of the image being played back on the image output device (120). At this time, the image input unit (231) may capture several tens of images per second (S512). In addition, the image acquired by capturing may include stego images constituting the stego video and the image output device (120) and its surrounding background. That is, as shown in FIG. 5, each image acquired by capturing is composed of a background and a stego image.
[0171] In one embodiment, steps after step S512 are performed by an application installed in the hidden data recovery device (130).
[0172] That is, the analysis unit (233) recognizes the boundary between the background and the image output device (120) from the image captured in step S512 and separates the stego image within the image output device (120) (S513). For example, if the image output device (120) is a TV, the image analysis unit (233) recognizes the boundary between the background and the TV from the image captured by the image input unit (231) and then separates the stego image being played on the TV. In one embodiment, each stego image separated in step S513 is composed of three color channels (RGB) in the width W * height H. According to embodiments, each stego image may be composed of three brightness channels (or YCbCr / YUV channels) in the width W * height H.
[0173] The details of extracting the TV boundary in the above analysis section (233) will be explained later.
[0174] In step S513, when the stego image is separated, it is divided into multiple sub-images (S514). In this case, the number of sub-images to be separated and the size of each sub-image are, in one embodiment, the same as the size of each sub-image and the number of sub-images to be divided in the encoding process of FIG. 4. That is, in step S514, each stego image is divided into 12 sub-images, each having a horizontal size of M and a vertical size of N.
[0175] In step S514, when the stego image is divided into multiple sub-images, the sub-images with embedded secret messages are selected (S515) using the same algorithm as that used in step S413 of FIG. 4. That is, as in FIG. 4, up to K sub-images with embedded secret messages can be selected. At this time, the sub-image with embedded secret messages is indicated as T, and the sub-image without embedded secret messages is indicated as F. Even without separate information, the decoder can determine whether the corresponding sub-image is T or F by analyzing the image.
[0176] When sub-images with inserted secret messages are selected in step S515, for each sub-image, a channel with inserted secret messages is selected using the same algorithm as used in step S415 of FIG. 4 (S516). Since Y data of the YUV channel was selected in FIG. 4, as an example, Y data of the YUV channel is also selected in FIG. 5. If B data of the RGB channel was selected in FIG. 4, B data of the RGB channel is also selected in FIG. 5.
[0177] When a decoding channel is selected in step S516, a data matrix of width W / M * height H / N * K is generated based on K sub-images with secret messages inserted and data information of the selected channel from the K sub-images (S517).
[0178] Then, the learned decoding model is applied to the data matrix of width W / M * height H / N * K generated in step S517 to obtain a secret message from K sub-images (S518). That is, the secret message can be obtained from the sub-images into which the secret message is inserted.
[0179] The secret message obtained in this manner may include a content identifier and time information. Then, the decoder (235) extracts the content corresponding to the content identifier from the storage unit (237) and provides it to the output unit (239), and the output unit (239) displays and / or expresses the input content as sound based on the time information.
[0180] For example, if a travel video is being played on the video output device (120), the smart glasses can display additional information related to the travel video being played, i.e., detailed information (or specific information) of a building or place shown in the travel video.
[0181] As another example, if a sports video is being played on the video output device (120), the smart glasses can display additional information related to the sports video being played, i.e., detailed information (or specific information) of a specific number of players appearing in the sports video.
[0182] In FIGS. 4 and 5, the size information of the cover image (or stego image), the size information of each sub-image divided from the cover image (or stego image), the number information of the sub-images, and the channel information for inserting and restoring a secret message are determined in a learning process of creating an encoder and a decoder by applying an artificial neural network deep learning model in a server (140), as an example.
[0183] FIG. 6 is a diagram showing an example of a learning process for creating an encoder and a decoder by applying an artificial neural network deep learning model in a server according to embodiments. In FIG. 6, the server may be the server (140) of FIG. 1, and after creating an encoder and a decoder by applying an artificial neural network deep learning model, the encoder is provided to a data hiding device (110), and the decoder is provided to a hidden data restoration device (130). In one embodiment, the encoder may be provided as an image output device (120). In this case, the process of hiding a secret message in an image is performed in real time or non-real time in the image output device (120). That is, the encoder and decoder created in the server (140) are a pair (i.e., a set). In other words, the encoder and decoder are created as a pair, and the variables (or parameters) used in the encoder and the variables (or parameters) used in the decoder are the same. At this time, the encoder and decoder are each stored in the server (140) in the form of a file, and the encoder in the form of a file is provided to a data hiding device (110) and / or an image output device (120), and the decoder in the form of a file is provided to a hidden data restoration device (130).
[0184] In one embodiment, steps S611 to S615 of FIG. 6 operate in the same manner as steps S411 to S415 of FIG. 4.
[0185] That is, in FIG. 6, the server (140) divides the input cover video into a plurality of cover images (S611). Here, the cover video is composed of random images that have no relation to the video to be played back on the video output device (120). For example, the cover video may be a collection of images randomly collected from the Internet, etc. In other words, the current step is before the content provider creates the content, and is a learning step for creating an encoder and decoder, so the cover images have no relation to the video to be played back on the video output device (120).
[0186] In this way, step S611 extracts cover images having a width of W and a height of H from a randomly input cover video. According to embodiments, each cover image may be composed of three color channels (or RGB channels) in width W * height H. According to embodiments, each cover image may be composed of three brightness channels (or YCbCr / YUV channels) in width W * height H.
[0187] Then, each cover image is divided into multiple sub-images (S612). As in Fig. 4, in step S612, each cover image is divided into multiple sub-images with a horizontal size of M and a vertical size of N.
[0188] In step S612, when the cover image is divided into multiple sub-images, at most K sub-images are selected from among the multiple sub-images into which a secret message is to be inserted (S613). Then, as in Fig. 4, the sub-image(s) selected for inserting the secret message are indicated by T, and the unselected sub-image(s) are indicated by F.
[0189] When one or more sub-images into which a secret message is to be inserted are determined in step S613, a channel into which an actual secret message is to be inserted is selected (S614).
[0190] In Fig. 6, as an example, Y data of a YUV channel is selected as in Fig. 4. Alternatively, B data of an RGB channel may be selected.
[0191] Once the selection of sub-images and channels into which secret messages are to be inserted is completed through steps S613 and S614, a data matrix of width W / M * height H / N * K is generated based on the data information of the K sub-images into which secret messages are to be inserted and the channels selected from the K sub-images (S615). That is, the data of the selected channels of the K sub-images having width W / M * height H / N are changed due to the insertion of the secret message during the encoding process.
[0192] At the same time, the server (140) generates a random secret message, converts the generated secret message into a binary format (or a bit array format), and then copies the converted secret message multiple times to make it the same size as the sub-image generated in step S615 (S616). That is, step S616 converts the input secret message into a binary array and then copies it to the same size as the data of the sub-image generated in step S615 to generate K sub-secret messages having a width of W / M * a height of H / N. In other words, for a randomly input random secret message, a data matrix of a width of W / M * a height of H / N * K is created by using data copying, etc.
[0193] Then, the encoding model learned up to the current step (i.e., the present) is applied to the data of the K (e.g., 6) sub-images generated in step S615 and the data of the K secret messages (or sub-secret messages) generated in step S616 to generate a stego sub-image with a secret message inserted (S617). That is, K (e.g., 6) stego sub-images having a size of width W / M * height H / N are generated.
[0194] Then, by replacing the K sub-images marked with T in step S613 among the cover images with the K stego sub-images generated in step S617, a stego image in frame units is generated (S618). That is, a secret message is hidden in the sub-images marked with T among the cover images, and a completed stego image without a hidden secret message is generated in the sub-images marked with F.
[0195] The stego image generated in this way is provided to a discriminator (651), a loss function calculation unit (653), and a start step (S619) for decoding. That is, the stego image generated in step S618 is trained to minimize the difference in the image hash of the cover image and the stego image, and to minimize the discriminator.
[0196] Step S619 is the starting step of decoding, and generates a corrupted stego image by corrupting a part of the stego image generated in step S618.
[0197] That is, in order to reflect the distortion of the image that occurs in the process of playing the stego image in the image output device (120) in the data hiding device (110) and capturing it with a camera in the hidden data restoration device (130) in the decoding process, step S619 arbitrarily changes the stego image. For example, the stego image is forcibly damaged within a certain range by applying blur, random noise addition, random brightness, contrast, luminance change, rotation, image decompression, etc. to the stego image. This is because even if the image captured by the camera and the image displayed on the image output device (120) are the same, there may be differences in the recognized colors, etc. For example, even if the color is the same green, the camera may recognize it as a darker green, and the image output device (120) may recognize it as a lighter green.
[0198] If the stego image is damaged in step S619, step S620 generates (or configures) a data matrix of width W / M * height H / N * K based on K sub-images with secret messages inserted from the damaged stego image and data information of channels selected from the K sub-images.
[0199] Then, the decoding model learned up to the current step (or present) is applied to the data matrix of width W / M * height H / N * K generated in step S620 to obtain a secret message from K sub-images (S621).
[0200] According to embodiments, the loss function calculation unit (653) of the present disclosure may perform a step of determining the level of the current learning stage and defining a loss function for proceeding to the next learning stage. According to embodiments, the step of defining the loss function consists of the following three functions.
[0201] 1) A function that evaluates the loss of intelligibility of an image due to the insertion of a secret message into a stego image generated from a cover image (i.e., a function that evaluates the quality difference between the cover image and the stego image). Specifically, the encoder is updated to minimize the difference between the cover image and the stego image by comparing them pixel by pixel for each RGB pixel. In other words, continuous iterative learning is performed in the direction of decreasing the difference between RGB pixels.
[0202] 2) An evaluation function that determines whether a real image is fake from a stego image as a discriminator for Generative Adversarial Networks (i.e., a function that evaluates whether the cover image is a fake image created by AI). In other words, it is a function for excluding fake images created by AI.
[0203] 3) A function that evaluates the difference between the input secret message and the output secret message. That is, the decoder is updated to minimize the difference between the input secret message and the output secret message (i.e., the secret message obtained from the forcibly corrupted stego image). In other words, the difference between the input and output secret messages is compared in binary form, and continuous iterative learning is performed to minimize the difference.
[0204] That is, the loss function calculation unit (653) applies normalization to the result values of the three functions through a scale factor, converts them into a single floating-point number, and calculates the value of the loss function. Then, to minimize this, the encoder, decoder, and discriminator (651) are updated and the learning proceeds to the next stage.
[0205] To this end, the loss function calculation unit (653) receives a cover image extracted from a randomly input cover video and a stego image generated in step S618 (S623). In addition, the result of the discriminator of the generative adversarial network is received from the discriminator (651) (S624). In addition, the unit receives a randomly input secret message and the output secret message obtained in step S621 (S625). In other words, the loss function calculation unit (653) continuously updates the encoder and decoder in a direction that minimizes the loss function.
[0206] As described above, by repeating the update process (i.e., learning process), when the quality difference between the cover image and the stego image becomes minimal (e.g., below the first threshold value) and the difference between the input secret message and the output secret message becomes minimal (e.g., below the second threshold value), the encoder and decoder at this time are saved in file form. Then, the encoder in file form is provided to the data hiding device (110) and / or the image output device (120), and the decoder in file form is provided to the hidden data restoration device (130).
[0207] As described above, in order to insert a secret message into a cover image using an artificial neural network, the present disclosure divides each frame (i.e., a cover image) into sub-images of size M x N, classifies the sub-images into textured sub-images and near-monochrome sub-images in consideration of the characteristics of the sub-images, and inserts the secret message into the textured sub-images. To this end, the textured sub-images are denoted as T, and the near-monochrome sub-images are denoted as F. That is, since the textured sub-images are relatively complex images, even if a secret message is hidden in the sub-images, it is difficult to recognize the difference between the stego data and the cover data, thereby improving unrecognizability. In addition, since only up to K sub-images among the cover images corresponding to the frames are encoded, distortion does not occur throughout the cover image, improving unrecognizability, and the encoding speed is also improved, thereby enabling real-time encoding.
[0208] Additionally, the decoding process has the same logic as the encoding process, and since only sub-images with embedded secret messages are decoded, the decoding process is faster and simpler. Therefore, the present disclosure is suitable for lightweight AR devices.
[0209] FIG. 7 is a drawing showing an example of a method for extracting the boundary of an image output device in an analysis unit according to embodiments.
[0210] Figures 8(a) to 8(f) illustrate the process of extracting the boundary of an image output device in an analysis unit according to embodiments.
[0211] In FIGS. 7 and 8(a) to 8(f), the analysis unit may be an analysis unit (233) included in a hidden data restoration device (130). In addition, in FIGS. 7 and 8(a) to 8(f), the image output device (120) is a TV, as an example. That is, the analysis unit (233) extracts a TV boundary from an image captured by the image input unit (231). This is to provide only the image corresponding to the image being played back in the image output device (120) among the captured images to the decoder (235).
[0212] In the present disclosure, an image captured by the image input unit (231) may include one or more image output devices, or may include not only an image being played on the image output device, but also the image output device and its surroundings.
[0213] In the present disclosure, the analysis unit (233) uses a deep learning-based object detection technology for an input image acquired by the image input unit (231) as shown in FIG. 8(a) to find the boundary of an image output device (S711). In the present disclosure, the image output device may be a device such as a TV, a monitor, a tablet, or a signage, and the analysis unit (233) finds the boundary of an image output device such as a TV, a monitor, a tablet, or a signage from an input image.
[0214] At this time, one or more image output devices may be included in the image acquired by the image input unit (231). For example, a TV and a monitor may be included in one image.
[0215] In the present disclosure, when two or more image output devices are detected in an input image, one of them is selected (S712). In one embodiment of the present disclosure, the analysis unit (233) selects the image output device closest to the center of the input image by considering the user's line of sight. For example, if the input image includes both a monitor and a TV, and the TV is closer to the center of the input image than the monitor, the analysis unit (233) selects the TV. In other words, the analysis unit (233) can select one image output device with the highest confidence among multiple image input devices included in the captured image.
[0216] At this time, since the object detection technology cannot accurately detect the boundary of the object on a pixel-by-pixel basis, the analysis unit (233) may find a boundary that includes more of the outside of the object or only a portion of the object. In other words, the object detection technology has the advantage of quickly finding the boundary of the object, but has the disadvantage that the boundary of the recognized object is not accurate. Therefore, the analysis unit (233) applies upper, lower, left, and right margins to the boundary of the detected object to create a wider boundary and extracts an image within the wider boundary (S713). At this time, the left and right margins and the upper and lower margins may be based on 10% of the width and height of the detected object, respectively. However, these margins may be changed.
[0217] In the present disclosure, object detection technology may be omitted. That is, in one embodiment, the application of the object detection method in the present disclosure is optional.
[0218] When an image is extracted within a boundary created by adding a margin in step S713, the outline of the extracted image must be extracted. The present disclosure, as an example, extracts the outline of the image extracted in step S713 using a gray scale and Canny algorithm, as shown in FIG. 8(b) (S714).
[0219] At this time, if object detection technology is applied as in Fig. 8(a), gray scale and Canny algorithm are applied to the image extracted within the boundary created by adding upper, lower, left, and right margins to the boundary of the detected object, and if object detection technology is omitted, gray scale and Canny algorithm are applied to the entire image to extract the outline of the image.
[0220] According to the embodiments, the grayscale and Canny algorithms applied to extract the outline of an image are image processing functions, where grayscale is a process of converting an image to black and white, and the Canny algorithm is an edge detection algorithm that finds the outline (or contour) of all objects in an image converted to black and white.
[0221] In step S714, when the outlines of all objects in the image are extracted, closed polygons are acquired (or extracted) from the extracted outlines as in Fig. 8(c) (S715). In the present disclosure, a closed polygon has a closed boundary and is a completely closed shape with a start point and an end point connected, such as a triangle, a square, a pentagon, or a circle.
[0222] At this time, in order to quickly extract the boundary of the image output device, the following closed polygons among the multiple closed polygons extracted in step S715 can be excluded.
[0223] For example, if a closed polygon has an area smaller than a preset threshold (e.g., Area_Threshold), the closed polygon can be excluded. That is, polygons that are too small are discarded as they are unlikely to be TV edges.
[0224] In the present disclosure, the threshold value (Area_Threshold) is set to 0.0005 times the width of the boundary of the image input device selected in step S712, as an example. This is just one example, and the threshold value (Area_Threshold) can be changed.
[0225] Then, even if some of the closed polygons are excluded according to the threshold value (Area_Threshold), multiple closed polygons may exist as in Fig. 8(d).
[0226] Therefore, in this case, closed polygons are merged in the following order (S716). Step S716 is a filtering and merging step of closed polygons.
[0227] That is, starting from the center point of the boundary of the video output device selected in step S712, the distances from the center of gravity of each closed polygon are accumulated in order of proximity to obtain a rectangle with the minimum area that includes the set of closed polygons. In other words, merging is performed starting from the closed polygon closest to the center point of the TV. At this time, the location of the closed polygon is the center of gravity, and shapes smaller than a threshold value in terms of area are discarded. Then, the bounding rectangle is obtained after merging.
[0228] To do this, we find one or more closed polygons closest to the center point. If object detection technology is not applied, we find the closed polygon closest to the center point in the entire image. If object detection technology is applied, we find the closed polygon closest to the center point in the TV in the image.
[0229] In this way, closed polygons are searched in order of proximity to the center point, and a rectangle that can contain the found closed polygons (i.e., a rectangle among the closed polygons) is found. This process is repeated until a rectangle with the minimum area that contains all the closed polygons is found. For example, if there is a rectangle that contains two closed polygons found in order of proximity to the center point, the two closed polygons are merged into that rectangle. Then, if there is a wider rectangle that contains this rectangle and other closed polygons, they are merged into the wider rectangle. By repeating this process, the area of the rectangle gradually increases. In other words, through polygon merging, only multiple rectangles (i.e., a set of rectangles) remain among the closed polygons.
[0230] When polygon merging is performed in step S716, a TV acceptance test is performed (S717) to determine whether the rectangles resulting from the polygon merging are judged as video output devices (e.g., TV edges).
[0231] Step S717 is a step for determining whether to determine the set of rectangles obtained in step S716 as a video output device (e.g., TV edge) as in Fig. 8(e), and only rectangles satisfying the following conditions are selected to be determined as video output devices.
[0232] That is, if the difference between the center point of the boundary of the video output device selected in step S712 and the center point of the rectangle is less than or equal to a center threshold (Acceptance_Center_Threshold) (e.g., when object detection technology is applied) and / or the area of the rectangle compared to the area of the boundary of the video output device selected in step S712 is greater than or equal to an area threshold (Acceptance_Area_Threshold) (e.g., when object detection technology is not applied), the rectangle is selected to be determined as an video output device, and rectangles that do not satisfy this condition are discarded. For example, if object detection technology is applied, if the difference between the center of the boundary and the center point of the rectangle satisfies the threshold (Acceptance_Center_Threshold), that is, if the difference is less than the threshold (Acceptance_Center_Threshold), the rectangle is selected to be determined as a TV edge. Alternatively, if the ratio of the TV width to the area of the object detection is greater than or equal to the threshold (Acceptance_Area_Threshold), the rectangle is selected to be determined as a TV edge. Accordingly, in the present disclosure, one or more squares may be selected to be judged as an image output device.
[0233] In the present disclosure, the center threshold (Acceptance_Center_Threshold) is set to 0.333 as an example. This is an example, and the center threshold (Acceptance_Center_Threshold) is changeable. In addition, the area threshold (Acceptance_Area_Threshold) is set to 0.7 as an example. This is an example, and the area threshold (Acceptance_Area_Threshold) is changeable.
[0234] In step S717, one or more rectangles that passed the TV pass test, i.e., one or more rectangles selected to be determined as a video output device (e.g., a TV edge), are one or more candidates for a video output device (e.g., a TV edge).
[0235] That is, step S717 excludes, among one or more rectangles resulting from the polygon merge, if the difference between the center point of the entire image or the center point of the boundary of the image input device when object detection technology is applied and the center point of the rectangle is too large (i.e., greater than or equal to the threshold (Acceptance_Center_Threshold)), the rectangle is excluded from the candidate for the image output device (e.g., TV edge). This means that, among one or more rectangles, rectangles that are not at or near the center point of the entire image or the center point of the boundary of the image input device will not be selected as a candidate for the image output device (e.g., TV edge) and will be discarded. In addition, step S717 excludes, among one or more rectangles resulting from the polygon merge, rectangles that have an area that is too small (i.e., less than or equal to the threshold (Acceptance_Area_Threshold)) from the candidate for the image output device (e.g., TV edge). This means that, among one or more rectangles, rectangles that have an area that is too small will not be selected as a candidate for the image output device (e.g., TV edge) and will be discarded.
[0236] If there is only one square that passes the TV pass test in S717, that square is determined to be a video output device (i.e., TV edge), and the image corresponding to that square is provided to the decoder (235) for secret message restoration.
[0237] If there are two or more rectangles that passed the TV pass test in S717, the two or more rectangles are sorted in descending order of area, and then the images of each rectangle are provided to the decoder (235) in the sorted order to apply steganographic decoding. As a result, the rectangle from which the secret message is extracted among the two or more rectangles is selected as the final video output device (e.g., TV edge). Then, the decoder (235) reads the corresponding content from the storage unit (237) based on the content identifier and time information included in the secret message restored from the image of the selected rectangle and provides the user experience to the output unit (239).
[0238] That is, the set of rectangles that passed through step S717 is a candidate for the TV boundary, and in the case where multiple rectangles are extracted, the most suitable rectangle can be selected by applying steganographic decoding to all candidates, as shown in Fig. 8(f).
[0239] Additionally, the present disclosure can perform encoding by applying an Error Correction Code (ECC) in an encoder to insert a secret message into a cover video, and perform error detection and error correction using the ECC in a decoder, and decoding to restore the secret message. When this is applied to two or more squares in step S718, only one of the squares can restore the secret message without error. Therefore, the present disclosure can use this method to select one of two or more video output device candidates.
[0240] As described above, in the present disclosure, the encoder encodes a secret message input by an administrator (information provider, content provider, or content advertiser) for each frame of the cover video (Video Source) through an artificial neural network steganography deep learning model, regenerates a new stego image, and outputs it to the screen of a video output device (e.g., TV, laptop, signage, etc.). Then, the AR glasses recognize the screen of the video output device through a camera, analyze the video being played on the screen in real time, and decode the hidden secret message from the video being played through an artificial neural network steganography deep learning model. Then, the AR glasses compare the decoded secret message with the data stored in the database, and if they match, display a corresponding UI on the AR glasses screen or play a sound.
[0241] In the present disclosure, the secret message includes a content identifier (Contents ID) and time information (Timestamp), and the AR glasses can distinguish content stored in the database through the content identifier and determine the playback time of the content stored in the database through the time information.
[0242] Therefore, the present disclosure allows users to obtain useful information simply by watching a video while wearing the glasses, without any additional manipulation. For example, in an outdoor movie theater, such as a drive-in theater, users can see and hear movie subtitles and audio through the glasses. Furthermore, in exhibition halls, such as museums and art galleries, users can see and hear additional explanations or supplementary information about the artwork through the glasses.
[0243] Figures 9 to 14 are diagrams showing examples of user experiences (UX) provided according to secret messages restored in the present disclosure.
[0244] FIG. 9(a) shows an example of a travel video being played on an image output device (120) according to embodiments, and FIG. 9(b) is a diagram showing an example of additional information related to the travel video being played on the image output device (120) of FIG. 9(a) being displayed on a hidden data restoration device (130), for example, smart glasses. In the present disclosure, additional information related to the travel video may include geographic information, buildings, transportation, lodging, and restaurant information regarding the travel destination.
[0245] For example, assuming that a travel video is being played on the video output device (120) as in FIG. 9(a), the user can view additional information related to the travel video being played on the video output device (120), for example, detailed information (or specific information) of a building or place (e.g., a tourist attraction) shown in the travel video, through the smart glasses simply by looking at the video output device (120) through the smart glasses. That is, when the user wears the smart glasses and watches the travel video being played on the video output device (120), the smart glasses restore a secret message from the travel video being played on the video output device (120) and recognize a content identifier and / or time information included in the restored secret message. Then, the smart glasses retrieve and display content (e.g., detailed information on a tourist attraction) corresponding to the content identifier from the database according to the timestamp included in the time information through the neighbor-of-screen of the smart glasses. Out-of-screen refers to a screen outside the user's field of view in the smart glasses.
[0246] To this end, the storage unit (237) of the hidden data restoration device (130) stores additional information related to the travel video being played, and the secret message restored by the hidden data restoration device (130) includes a content identifier to extract the additional information related to the travel video being played from the storage unit (237). In addition, the restored secret message may further include time information (e.g., a timestamp) related to the additional information. The time information may include the start time of the playback, the playback time, or the end time of the playback of the additional information. For example, the smart glasses may display the additional information for several seconds (e.g., 5 seconds) from the moment the secret message is restored based on the time information. The additional information may be a still image such as a picture, a photograph, or text, or may be a moving video. In addition, if there is a plurality of additional information related to the travel video being played, there may also be a plurality of content identifiers and time information, or the content identifiers and time information of the remaining additional information may be inferred based on one content identifier and time information. Alternatively, a secret message may be hidden in the travel video being played, as many times as there are additional information.
[0247] That is, the encoder of the data hiding device (110) inserts a secret message including a content identifier and time information into the cover video in real time or non-real time to generate a stego video, and the image output device (120) processes and plays (or displays) the stego video. Then, the hidden data restoration device (130) captures the stego video being played in the image output device (120), restores the secret message from the captured image using a decoder, and then retrieves the corresponding content from the storage unit (237) based on the content identifier and time information included in the restored secret message and displays it on the output unit (239).
[0248] The method of creating an encoder and decoder by applying an artificial neural network steganography deep learning model in a server (140) not described in Fig. 9, the method of analyzing the edge of an image output device in a hidden data restoration device (130), etc. have been described in detail above, so they will be omitted here to avoid redundant explanation.
[0249] FIG. 10(a) shows an example of a sports game video being played on an image output device (120) according to embodiments, and FIG. 10(b) is a diagram showing an example of additional information related to a sports game video being played on the image output device (120) of FIG. 10(a) being displayed on a hidden data restoration device (130), for example, smart glasses. In the present disclosure, additional information related to a sports game video may include information on a sports team's club, league rankings, a list of starting players, key scenes from the game, and detailed information on a specific player.
[0250] For example, assuming that a sports game video is being played on the video output device (120) as in FIG. 10(a), the user can view additional information related to the sports game video being played on the video output device (120), such as key scenes of the game or detailed information (or specific information) of a specific player, through the smart glasses simply by looking at the video output device (120) through the smart glasses. That is, when the user wears the smart glasses and watches the sports game (or sports broadcast) being played on the video output device (120), the smart glasses restore a secret message from the sports game video being played on the video output device (120) and recognize a content identifier and / or time information included in the restored secret message. Then, the content (e.g., detailed information of a specific player) identified by the content identifier is retrieved from the database and displayed according to the timestamp included in the time information through the neighbor-of-screen.
[0251] To this end, the storage unit (237) of the hidden data restoration device (130) stores additional information related to the sports game video being played, and the secret message restored by the hidden data restoration device (130) includes a content identifier to extract the additional information related to the sports game video being played from the storage unit (237). In addition, the restored secret message may further include time information (e.g., a timestamp) related to the additional information. The time information may include the playback time, playback time, etc. of the additional information. For example, smart glasses may display the additional information for several seconds (e.g., 5 seconds) according to the time information. The additional information may be a still image such as a picture, a photograph, or text, or may be a moving video. In addition, if there is a plurality of additional information related to the played video, there may also be a plurality of content identifiers and time information, or the content identifiers and time information of the remaining additional information may be inferred based on one content identifier and time information. Alternatively, secret messages may be hidden in the sports game footage being played, as many as the number of additional information.
[0252] That is, the encoder of the data hiding device (110) inserts a secret message including a content identifier and time information into the cover video in real time or non-real time to generate a stego video, and the image output device (120) processes and plays (or displays) the stego video. Then, the hidden data restoration device (130) captures the stego video being played in the image output device (120), restores the secret message from the captured image using a decoder, and then retrieves the corresponding content from the storage unit (237) based on the content identifier and time information included in the restored secret message and displays it on the output unit (239).
[0253] The method of creating an encoder and decoder by applying an artificial neural network steganography deep learning model in a server (140) not described in Fig. 10, the method of analyzing the edge of an image output device in a hidden data restoration device (130), etc. have been described in detail above, so they will be omitted here to avoid redundant explanation.
[0254] FIG. 11(a) shows an example of a music video being played on an image output device (120) according to embodiments, and FIG. 11(b) is a diagram showing an example of additional information related to a music video being played on the image output device (120) of FIG. 11(a) being displayed on a hidden data restoration device (130), for example, smart glasses. In the present disclosure, additional information related to a music video may include agency information, group member introductions, performance news, song lyric subtitle information, and the like.
[0255] For example, assuming that a music video of an idol girl group is being played on an image output device (120) as in FIG. 11(a), the user can view additional information related to the music video being played on the image output device (120), such as detailed information (or specific information) of a specific member, song lyric subtitle information of the music video, etc., through the smart glasses simply by looking at the image output device (120) through the smart glasses. That is, when the user wears the smart glasses and watches the music video being played on the image output device (120), the smart glasses restore a secret message from the music video being played on the image output device (120) and recognize a content identifier and / or time information included in the restored secret message. Then, the smart glasses display content (e.g., song lyric subtitle information of the music video) corresponding to the content identifier according to a timestamp included in the time information through the neighbor-of-screen. If song lyric subtitle information is displayed, the song lyrics can be displayed in synchronization with the song of the music video being played on the image output device (120).
[0256] To this end, the storage unit (237) of the hidden data restoration device (130) stores additional information related to the music video being played, and the secret message restored by the hidden data restoration device (130) includes a content identifier to extract the additional information related to the music video being played from the storage unit (237). In addition, the restored secret message may further include time information (e.g., a timestamp) related to the additional information. The time information may include the playback time, playback time, etc. of the additional information. The additional information may be a still image such as a picture, a photograph, or text, or may be a moving image. In addition, if there is a plurality of additional information related to the played image, the content identifier and time information may also be a plurality, or the content identifier and time information of the remaining additional information may be inferred based on one content identifier and time information. Alternatively, as many secret messages as the number of additional information may be hidden in the music video being played.
[0257] That is, the encoder of the data hiding device (110) inserts a secret message including a content identifier and time information into the cover video in real time or non-real time to generate a stego video, and the image output device (120) processes and plays (or displays) the stego video. Then, the hidden data restoration device (130) captures the stego video being played in the image output device (120), restores the secret message from the captured image using a decoder, and then retrieves the corresponding content from the storage unit (237) based on the content identifier and time information included in the restored secret message and displays it on the output unit (239).
[0258] The method of creating an encoder and decoder by applying an artificial neural network steganography deep learning model in a server (140) not described in Fig. 11, the method of analyzing the edge of an image output device in a hidden data restoration device (130), etc. have been described in detail above, so they will be omitted here to avoid redundant explanation.
[0259] FIG. 12(a) shows an example of a drama being played on an image output device (120) according to embodiments, and FIG. 12(b) is a diagram showing an example of additional information related to a drama being played on the image output device (120) of FIG. 12(a) being displayed on a hidden data restoration device (130), for example, smart glasses. In the present disclosure, additional information related to a drama may include detailed information on a specific actor, product advertisement information such as clothes, accessories, shoes, etc. worn by a specific actor, etc.
[0260] For example, assuming that a drama is being played on an image output device (120) as in FIG. 12(a), the user can view additional information related to the drama being played on the image output device (120), such as purchase information for sunglasses worn by a specific actor, detailed information (or specific information), etc. of a specific actor, through the smart glasses simply by looking at the image output device (120) through the smart glasses. That is, when the user wears the smart glasses and watches the drama being played on the image output device (120), the smart glasses restore a secret message from the drama video being played on the image output device (120) and recognize a content identifier and / or time information included in the restored secret message. Then, the smart glasses display content (e.g., purchase information for a product worn by a specific actor) corresponding to the content identifier according to a timestamp included in the time information through neighbor-of-screen. For example, the purchase information for a product may include the price of the product, an image of the product, link information for purchasing the product, etc., and when the user selects the displayed additional information, the user can be moved to the shopping mall for the corresponding product.
[0261] To this end, the storage unit (237) of the hidden data restoration device (130) stores additional information related to the drama video being played, and the secret message restored by the hidden data restoration device (130) includes a content identifier to extract the additional information related to the drama video being played from the storage unit (237). In addition, the restored secret message may further include time information (e.g., a timestamp) related to the additional information. The time information may include the playback time and playback time of the additional information. For example, smart glasses may display the additional information for several seconds (e.g., 5 seconds) according to the time information. The additional information may be a still image such as a picture, a photograph, or text, or may be a moving video. In addition, if there is a plurality of additional information related to the played video, there may also be a plurality of content identifiers and time information, or the content identifiers and time information of the remaining additional information may be inferred based on one content identifier and time information. Alternatively, as many secret messages as the number of additional information may be hidden in the drama video being played.
[0262] That is, the encoder of the data hiding device (110) inserts a secret message including a content identifier and time information into the cover video in real time or non-real time to generate a stego video, and the image output device (120) processes and plays (or displays) the stego video. Then, the hidden data restoration device (130) captures the stego video being played in the image output device (120), restores the secret message from the captured image using a decoder, and then retrieves the corresponding content from the storage unit (237) based on the content identifier and time information included in the restored secret message and displays it on the output unit (239).
[0263] The method of creating an encoder and decoder by applying an artificial neural network steganography deep learning model in a server (140) not described in Fig. 12, the method of analyzing the edge of an image output device in a hidden data restoration device (130), etc. have been described in detail above, so they will be omitted here to avoid redundant explanation.
[0264] FIG. 13(a) shows an example of home shopping or an advertisement being played on an image output device (120) according to embodiments, and FIG. 13(b) is a diagram showing an example of additional information related to home shopping or an advertisement being played on the image output device (120) of FIG. 13(a) being displayed on a hidden data restoration device (130), for example, smart glasses. In the present disclosure, the image output device (120) can play home shopping or an advertisement (for example, a terrestrial advertisement played during an entertainment / sports break time), and the additional information related to home shopping or an advertisement may include detailed information on a home shopping product or an advertisement product being played, a discount coupon (or discount code) for the product, purchase link information, etc.
[0265] For example, assuming that an advertisement (e.g., LG's object dryer) is being played on the video output device (120) as in FIG. 13(a), the user can view additional information related to the advertisement being played on the video output device (120), such as a discount coupon (or discount code) for the advertised object dryer, purchase link information, etc., through the smart glasses simply by looking at the video output device (120) through the smart glasses. That is, when the user wears the smart glasses and views the advertisement being played on the video output device (120), the smart glasses restore a secret message from the advertisement video being played on the video output device (120) and recognize a content identifier and / or time information included in the restored secret message. Then, the smart glasses display content (e.g., a discount coupon or discount code for the advertised product, purchase link information) corresponding to the content identifier according to the timestamp included in the time information through the neighbor-of-screen. For example, if a user selects a purchase link via smart glasses, they can be taken to the corresponding product's online shopping mall and use the discount coupon (or code) provided as additional information when purchasing the product. In other words, discount coupons (or codes) can only be provided via smart glasses. In this way, the present disclosure can run promotions with special discount codes during home shopping or commercial broadcasts. Therefore, for companies selling smart glasses, this can increase sales demand for smart glasses to take advantage of the promotion, and can serve as a temporary marketing point. Furthermore, since viewers never know when a discount code will be exposed during home shopping or sports broadcasts, this can have the effect of keeping viewers glued to the channel. This can lead to increased viewership.
[0266] To this end, the storage unit (237) of the hidden data restoration device (130) stores additional information related to the home shopping / advertisement video being played, and the secret message restored by the hidden data restoration device (130) includes a content identifier to extract the additional information from the storage unit (237). In addition, the restored secret message may further include time information (e.g., a timestamp) related to the additional information. The time information may include the playback time, playback time, etc. of the additional information related to the played video. The additional information may be a still image such as a picture, a photograph, or text, or may be a moving video. In addition, if there is a plurality of additional information related to the played video, the content identifier and time information may also be a plurality, or the content identifier and time information of the remaining additional information may be inferred based on one content identifier and time information. Alternatively, as many secret messages as the number of additional information may be hidden in the home shopping / advertisement video being played.
[0267] That is, the encoder of the data hiding device (110) inserts a secret message including a content identifier and time information into the cover video in real time or non-real time to generate a stego video, and the image output device (120) processes and plays (or displays) the stego video. Then, the hidden data restoration device (130) captures the stego video being played in the image output device (120), restores the secret message from the captured image using a decoder, and then retrieves the corresponding content from the storage unit (237) based on the content identifier and time information included in the restored secret message and displays it on the output unit (239).
[0268] The method of creating an encoder and decoder by applying an artificial neural network steganography deep learning model in a server (140) not described in Fig. 13, the method of analyzing the edge of an image output device in a hidden data restoration device (130), etc. have been described in detail above, so they will be omitted here to avoid redundant explanation.
[0269] The image output device (120) in FIGS. 14(a) to 14(d) may be a digital signage installed in a subway, outdoors of a building, a subway station, a bus stop, an airport, a terminal, a shopping mall, etc. Digital signage is a device that displays information using digital display technology, and mainly plays still images or moving images for the purposes of advertising, information transmission, commercial purposes, brand promotion, etc.
[0270] When an advertisement is played on digital signage, as shown in FIGS. 14(a) to 14(d), if a user wearing smart glasses sees the advertisement, additional information related to the advertisement may be displayed through the smart glasses. For example, when an advertisement video is played on a digital signage screen installed at a bus stop, subway, terminal, airport, etc., if a user wearing smart glasses passes by and sees the advertisement video, additional information related to the advertisement may be displayed on the smart glasses screen and / or the advertisement sound may be heard through the smart glasses.
[0271] That is, when a user wears smart glasses and watches an advertisement being played on an image output device (120), the smart glasses restore a secret message from the advertisement video being played on the image output device (120) and recognize a content identifier and / or time information included in the restored secret message. Then, the smart glasses display content corresponding to the content identifier (e.g., purchase link information related to the advertised product) according to a timestamp included in the time information through neighbor-of-screen. At this time, if the user selects the purchase link information through the smart glasses, the user can be moved to the shopping mall for the corresponding product.
[0272] For example, smart glasses can display content corresponding to a content identifier included in a secret message a certain amount of time after the user starts watching an advertisement video playing on a digital signage by referencing the time information included in the secret message.
[0273] To this end, the storage unit (237) of the hidden data restoration device (130) stores additional information related to the reproduced video, and the secret message restored by the hidden data restoration device (130) includes a content identifier to extract the additional information related to the reproduced video from the storage unit (237). In addition, the restored secret message may further include time information (e.g., a timestamp) related to the additional information related to the reproduced video. The time information may include the reproduction time, reproduction time, etc. of the additional information related to the reproduced video. For example, the smart glasses may display the additional information related to the reproduced video about 10 seconds after the time from when the reproduction video is started to be viewed according to the time information. The additional information may be a still image such as a picture, a photograph, or text, or may be a moving video. In addition, if there is a plurality of additional information related to the reproduced video, there may also be a plurality of content identifiers and time information, or the content identifiers and time information of the remaining additional information may be inferred based on one content identifier and time information.
[0274] That is, the encoder of the data hiding device (110) inserts a secret message including a content identifier and time information into the cover video in real time or non-real time to generate a stego video, and the image output device (120) processes and plays (or displays) the stego video. Then, the hidden data restoration device (130) captures the stego video being played in the image output device (120), restores the secret message from the captured image using a decoder, and then retrieves the corresponding content from the storage unit (237) based on the content identifier and time information included in the restored secret message and displays it on the output unit (239).
[0275] The method of creating an encoder and decoder by applying an artificial neural network steganography deep learning model in a server (140) not described in Fig. 14, the method of analyzing the edge of an image output device in a hidden data restoration device (130), etc. have been described in detail above, so they will be omitted here to avoid redundant explanation.
[0276] As described above, the present disclosure captures a video being played (or displayed) on a video output device through a camera of various devices, and
[0277] Because it can extract secret messages (i.e. steganographic data) from images, it is suitable for AR devices that require decoding to restore secret messages immediately after viewing a video.
[0278] Furthermore, the present disclosure allows for the insertion of different secret messages into each frame, and since each frame is decoded independently, it can provide a customized user experience (UX) for each frame being viewed. In particular, since the insertion and restoration of secret messages are performed based on an artificial neural network deep learning model, lossy compression using various compression ratios can be performed during the learning process, enabling the extraction of secret messages even when the image is compressed.
[0279] And, in order to recognize a stego image from an image captured through a camera of a hidden data recovery device, for example, a glass, by recognizing an image output device (e.g., TV), extracting a stego image area from the device, and performing learning that takes into account the distortion of the camera device and the surrounding environment, it is possible to extract a secret message even if the stego image is distorted.
[0280] Furthermore, since the method of embedding a secret message into a cover video utilizes a deep learning model based on artificial neural network learning without using visible markers, it is possible to embed an invisible secret message into the video itself. Therefore, compared to markers, the method is far more unrecognizable when it comes to recognizing the difference between a cover video and a stego video with an embedded secret message. Furthermore, since the secret message is embedded into the video itself, the camera-equipped device does not need to communicate with the server to obtain the secret message from the video being viewed. Furthermore, the method can be applied to not only broadcast videos but also stored videos, streaming videos, and live broadcasts.
[0281] In this way, the present disclosure encodes an invisible secret message in a cover image using an artificial neural network deep learning model, and then decodes the secret message from a stego image transmitted through a camera in smart glasses, such as AR glasses, using the same artificial neural network deep learning model. Therefore, the present disclosure can hide secret messages not only in locally stored videos but also in streaming videos and live broadcasts, and then decode them through smart glasses, thereby providing diverse user experiences across metaverse devices.
[0282] Each of the parts, modules, or units described above may be software, processors, or hardware parts that execute sequential execution processes stored in memory (or storage units). Each of the steps described in the embodiments described above may be performed by processors, software, or hardware parts. Each of the modules / blocks / units described in the embodiments described above may operate as a processor, software, or hardware. In addition, the methods presented in the embodiments may be implemented as code. This code may be written on a processor-readable storage medium and thus may be read by a processor provided by an apparatus.
[0283] For convenience of explanation, this specification has been described separately in each drawing. However, it is also possible to design new embodiments by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by those skilled in the art, is also within the scope of the present disclosure.
[0284] The device and method according to the present disclosure are not limited to the configuration and method of the embodiments described above, but the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.
[0285] In this disclosure, terms such as "first" and "second" may be used to describe various components of the disclosure. However, the various components according to the disclosure should not be interpreted in a limited manner by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not necessarily mean the same user input signals unless the context clearly indicates otherwise.
[0286] The terminology used to describe the present disclosure is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. In addition, the term “and / or” is used to mean all possible combinations of terms. The term “comprises” or “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the present disclosure are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[0287] The best mode for carrying out the invention has been specifically described.
[0288] It will be apparent to those skilled in the art that various modifications and variations can be made to the present embodiments without departing from the spirit or scope of the present embodiments. Accordingly, the present embodiments are intended to include modifications and variations of the present embodiments provided they come within the scope of the appended claims and their equivalents.
Claims
1. A data hiding device comprising an encoder for generating a second image by hiding a secret message in a first image within a first video, and transmitting a second video including the second image; A video output device that receives and plays the second video; A hidden data restoration device comprising a decoder for restoring the secret message, wherein the device captures an image output device on which the second video is being played by at least one camera, and the secret message is restored from the captured image by the decoder, and provides a user experience based on the restored secret message; and An image processing system including a server that generates the encoder and the decoder by applying an artificial neural network deep learning model, provides the encoder to the data hiding device, and provides the decoder to the hidden data restoration device.
2. In the first paragraph, the encoder included in the data hiding device Divide the above first image into multiple sub-images, Select one or more first sub-images among the above multiple sub-images into which the secret message is to be inserted, Selecting a channel in which to insert the secret message among the channels constituting pixels of one or more of the selected first sub-images, Generating one or more sub secret messages by copying the secret message in the same number and size as the one or more first sub images, Generating one or more second sub-images with the secret message inserted by applying an encoding model learned from data of a selected channel of one or more of the first sub-images and data of one or more of the sub-secret messages, An image processing system that generates the second image having the secret message inserted therein by replacing one or more first sub-images among the plurality of sub-images of the first image with one or more second sub-images.
3. In the second paragraph, the encoder An image processing system for selecting a B channel as a channel into which to insert the secret message, when the channels constituting the pixels of the one or more selected first sub-images are RGB channels.
4. In the second paragraph, the encoder An image processing system for selecting a Y channel as a channel into which to insert the secret message, when the channels constituting the pixels of the one or more selected first sub-images are YUV channels.
5. In paragraph 2, The above secret message contains content identification information, The above content identification information is an image processing system used to identify content for providing the above user experience.
6. In paragraph 5, The above secret message further includes time information, An image processing system in which the time information includes at least one of information for identifying a playback time of content identified by the content identification information, information for identifying a playback time, or information for identifying a playback end time.
7. In the 6th paragraph, the hidden data restoration device, A storage unit storing one or more contents to provide the above user experience; A video input unit for shooting a video output device on which the second video is played; An analysis unit that detects the edge of the image output device from an image captured by the image input unit; The decoder that restores the secret message from the image corresponding to the edge of the image output device detected in the above analysis unit; and An image processing system including an output unit that extracts content from the storage unit based on content identification information included in the restored secret message and provides a user experience.
8. In paragraph 7, the analysis unit Extract the outline of the captured image above, Extract closed polygons based on the extracted outlines above, Select one or more rectangles containing one or more closed polygons by performing a merge from the closed polygons near the center of the image above, comparing said one or more rectangles with one or more threshold values to determine at least one of said one or more rectangles as an edge candidate of said image output device; An image processing system that, when there are multiple edge candidates for the image output device, arranges them in order of increasing area of the rectangle and then applies decoding to each of them to finally determine the edge of the image output device.
9. In the 8th paragraph, the decoder of the hidden data restoration device Dividing an image corresponding to an edge of the above image output device into multiple sub-images, Select one or more third sub-images from among the above multiple sub-images, in which the secret message is inserted, Selecting a channel in which the secret message is inserted among the channels constituting pixels of one or more of the selected third sub-images, An image processing system for obtaining the secret message by applying a decoding model learned from data of selected channels of one or more of the third sub-images.
10. In paragraph 9, An image processing system, wherein the number and size of one or more first sub-images and the selected channel selected for insertion of the secret message in the encoder of the data hiding device are the same as the number and size and the selected channel of one or more third sub-images into which the secret message is inserted in the decoder of the hidden data restoration device.
11. In paragraph 1, the server It includes an encoder, decoder and loss function calculation unit, The above encoder, Splitting a third image in a randomly input video into multiple sub-images, Select one or more fourth sub-images from among the above multiple sub-images into which to insert a randomly input secret message, Select a channel in which to insert the input secret message among the channels constituting the pixels of one or more of the fourth sub-images selected above, Generate one or more sub-secret messages by copying the input secret message to the same number and size as the one or more fourth sub-images, Generating one or more fifth sub-images with the input secret message inserted by applying an encoding model learned so far from data of a selected channel of one or more of the fourth sub-images and data of one or more of the sub-secret messages, Generating the fourth image with the input secret message inserted by replacing one or more fourth sub-images among the plurality of sub-images of the third image with the one or more fifth sub-images, The above decoder, The data of the fourth image is changed within a certain range to generate a fifth image, An output secret message is obtained by applying a decoding model learned so far from data of a selected channel of one or more sub-images in which the input secret message is inserted among a plurality of sub-images of the fifth image, The above loss function calculation section is, An image processing system that determines an encoder to be provided as the data hiding device and a decoder to be provided as the hidden data restoration device, evaluates the difference between the input secret message and the output secret message, calculates a loss function that evaluates the difference between the third image and the fourth image, and continuously updates the encoder and the decoder in a direction in which the loss function is minimized.
12. In paragraph 1, The above hidden data restoration device is an image processing system that is a smart glass.
13. A step of filming a video output device in which a video with a secret message inserted therein is played by at least one camera; A step of detecting an edge of the image output device from the captured image; A step of restoring the secret message by applying a decoding model learned from an image corresponding to the edge of the detected image output device; and A decoding method in an image processing system, comprising the step of extracting content identified by content identification information included in the restored secret message from a storage unit storing one or more contents for providing a user experience, thereby providing the user experience.
14. In the 13th paragraph, the edge detection step of the image output device, A step of extracting the outline of the above captured image; A step of extracting closed polygons based on the extracted outlines; A step of selecting one or more rectangles containing one or more closed polygons by performing merging from closed polygons near the center of the image; A step of comparing said one or more squares with one or more threshold values to determine at least one of said one or more squares as an edge candidate of said image output device; and A decoding method in an image processing system, comprising the step of, when there are multiple edge candidates of the image output device, arranging them in order of increasing area of the rectangle and then applying a learned decoding model to each of them to finally determine the edge of the image output device.
15. In the 14th paragraph, the secret message restoration step A step of dividing an image corresponding to an edge of the image output device into a plurality of sub-images; A step of selecting one or more sub-images having the secret message inserted therein among the plurality of sub-images; A step of selecting a channel in which the secret message is inserted among the channels constituting pixels of one or more of the selected sub-images; and A decoding method in an image processing system for obtaining the secret message by applying a decoding model learned from data of selected channels of one or more of the sub-images.
16. A storage unit that stores one or more contents that provide a user experience; A video input unit for filming a video output device in which a video with a secret message inserted therein is played by at least one camera; An analysis unit that detects the edge of the image output device from the captured image; A decoder that restores the secret message by applying a decoding model learned from an image corresponding to an edge of the detected image output device; and A hidden data restoration device including an output unit for extracting content from the storage unit and providing a user experience based on content identification information included in the restored secret message.
17. In paragraph 16, the analysis unit Extract the outline of the captured image above, Extract closed polygons based on the extracted outlines above, Select one or more rectangles containing one or more closed polygons by performing a merge from the closed polygons near the center of the image above, comparing said one or more rectangles with one or more threshold values to determine at least one of said one or more rectangles as an edge candidate of said image output device; A hidden data restoration device that, when there are multiple edge candidates of the image output device, sorts them in order of increasing area of the rectangle and then applies decoding to each of them to finally determine the edge of the image output device.
18. In the 17th paragraph, the decoder Dividing an image corresponding to an edge of the above image output device into multiple sub-images, Select one or more sub-images from among the above multiple sub-images in which the secret message is inserted, Selecting a channel in which the secret message is inserted among the channels constituting the pixels of one or more of the selected sub-images, A hidden data restoration device that obtains the secret message by applying a decoding model learned from data of selected channels of one or more of the sub-images.
Citation Information
Patent Citations
Method and apparatus for transmitting information on recommended accessories to a user terminal using a neural network
KR1020230110451A
Information processing device, information processing method, and recording medium
US20200364517A1
Imagery and annotations
US20230005052A1