System
The system facilitates easy creation of original stamps by guiding users through video shooting, scene selection, illustration generation, and style conversion, addressing privacy and copyright concerns.
Patent Information
- Application Number
- JP2024119016
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Users face difficulties in creating original stamps due to the requirement of specialized knowledge and design skills, and using personal photos as stamps raises privacy and copyright issues.
A system that provides instructions for shooting videos, selects specific scenes, generates original illustrations, converts styles, and allows editing, utilizing AI for easy stamp creation while addressing privacy and copyright concerns.
Enables users to easily create personalized stamps without specialized knowledge, overcoming privacy and copyright issues.
Smart Images

Figure 2026017955000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's digital communications, users want to express themselves using original stamps that reflect their individuality. However, creating stamps requires specialized knowledge and design skills, making it a hurdle for many users. Furthermore, using personal photos as stamps can raise privacy and copyright issues. Therefore, there is a need to provide a method that allows users to easily create their own original stamps while also overcoming privacy and copyright issues. [Means for solving the problem]
[0005] The present invention provides a system that provides a means for presenting instructions to a user for generating an image, a means for receiving a video shot by the user based on the instructions, a means for selecting a specific scene from the video, a means for generating an original illustration based on the specific scene, a means for converting the original illustration into a different style, and a means for the user to edit the illustration. This system allows users to easily shoot a video and, with the assistance of AI, create their own original stamps. Furthermore, style conversion also addresses privacy and copyright issues.
[0006] An "image" is digital data that represents visual information, and includes still images and moving images.
[0007] "Instructions" are guidelines or explanations that encourage specific actions or behaviors, and refer to information that users can use as a reference when shooting videos.
[0008] "User" refers to an individual or entity that uses the system to create original stamps.
[0009] "Video" is a multi-frame digital visual medium that is a sequence of images that change over time.
[0010] "Receiving" refers to the act of acquiring or receiving specific data, in this case the process by which the server receives video footage taken by the user.
[0011] A "scene" refers to a specific frame or series of frames in a video that contains a specific facial expression or action.
[0012] The "original illustration" is an initial image generated based on a scene extracted from a video, and is the material that will ultimately be used as a stamp.
[0013] "Style" refers to a method or design concept for transforming the appearance or expression of the original illustration into a specific format.
[0014] "Style conversion" refers to the process of changing the visual style of the original illustration, and specific examples include converting from a photo-realistic style to a cartoon-style or painterly style.
[0015] "Editing" refers to the act of adding text and decorations to the original illustration or stamp, and is a process in which the user customizes it.
[0016] "Interface" refers to the means, devices, and software operation screens that users use to interact with a system.
[0017] "System" refers to a comprehensive setup that has the functions of image generation, video reception, scene selection, illustration generation, style conversion, and editing. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention aims to provide a system that allows users to easily create their own original stamps. This system includes the steps of presenting instructions for image generation, allowing the user to shoot a video, selecting a specific scene from the video, generating an original illustration from the selected scene, converting the original illustration into a different style, and allowing the user to edit the illustration.
[0040] Program processing explanation
[0041] Initialization and User Authentication
[0042] When a user launches an application, the server prompts for authentication information. The user enters a username and password and submits them to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[0043] Video shooting scene suggestions
[0044] The server uses an AI module to send instructions to the device about scenes suitable for the stamps, such as specific facial expressions or specific poses, for the user to capture in the video.
[0045] Video recording
[0046] The user uses the camera function of the device to shoot a video based on instructions from the server. The video is set to be a few minutes long and ensures that the specified scenes are included.
[0047] Video upload and scene selection
[0048] Once the video is complete, the device automatically uploads the video data to the server, which then analyzes it and selects specific scenes. This process uses an AI algorithm to scan each frame in the video and find the best frame that corresponds to the recommended scene.
[0049] Image processing and style transfer
[0050] Once the selected scene is determined, the server generates an original illustration from that scene. The server then converts the generated original illustration into a different style. This conversion process uses machine learning algorithms such as generative adversarial networks (GANs) to convert the illustration into a style such as photorealistic, cartoonish, or pictorial.
[0051] Preview and edit stamps
[0052] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, such as emojis, stickers, and custom text.
[0053] Creating and saving stamp packs
[0054] Once the user has finished editing, the device will send the final stamp pack to the server, which will save it and issue a download link, through which the user can download the stamp pack and import it into LINE.
[0055] Specific examples
[0056] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to the stamp to create their own original stamp.
[0057] In this way, the present invention provides a method and system that allows easy creation of original stamps tailored to the needs of individual users.
[0058] The processing flow will be explained below.
[0059] Step 1:
[0060] The user launches the application, and the login screen appears.
[0061] Step 2:
[0062] The server prompts the user for authentication information: a login screen prompts for username and password.
[0063] Step 3:
[0064] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[0065] Step 4:
[0066] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[0067] Step 5:
[0068] The server uses an AI module to instruct the user on the scenes (e.g., emotions) that are suitable for the stamp. The instructions are displayed on the device screen.
[0069] Step 6:
[0070] The user uses the device's camera function to shoot a video based on the specified scene, and the user manually starts and stops video shooting.
[0071] Step 7:
[0072] Once the video recording is complete, the device will automatically upload the video data to the server, and the upload progress will be displayed on the screen.
[0073] Step 8:
[0074] The server receives the uploaded video data and uses AI algorithms to analyze each frame of the video, automatically selecting the best frame that corresponds to the recommended scene.
[0075] Step 9:
[0076] The server generates an original illustration based on the selected frame, and this original illustration is saved in the system.
[0077] Step 10:
[0078] The server converts the original illustration into a style selected by the user (e.g., photo-realistic, cartoon-like, or painted style). Style conversion is performed using machine learning algorithms such as generative adversarial networks (GANs).
[0079] Step 11:
[0080] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[0081] Step 12:
[0082] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[0083] Step 13:
[0084] The device sends the edited stamp pack to the server, which stores the received stamp pack.
[0085] Step 14:
[0086] Once the server has finished saving the stamp pack, it will generate a download link and notify the user of this link.
[0087] Step 15:
[0088] Users can download the stamp pack via the provided link and import it into the LINE app, allowing them to use their custom stamps on LINE.
[0089] Example 1
[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0091] In conventional stamp creation systems, the process for users to create their own unique images is complicated and requires many operations, making it difficult to easily generate stamps. Furthermore, manual image editing and style conversion require specialized knowledge, making them difficult for general users to use. This has led to a demand for a new method and system that allows users to easily create original stamps that are optimal for each individual user.
[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0093] In this invention, the server is a system that allows users to easily create their own original stamps, and includes: means for requesting authentication information from the user and authenticating them by referencing a database; means for generating instructions regarding scenes suitable for stamps using an AI model and transmitting the instructions to the terminal; means for receiving from the terminal a video shot by the user based on the instructions; means for analyzing the uploaded video and selecting a specific scene; means for generating an original illustration based on the specific scene and converting the generated original illustration into a different style; means for transmitting the converted illustration to the terminal and allowing the user to edit the illustration; and means for saving the edited stamp pack and issuing a download link. This allows users to easily create original stamps and use them in messaging apps without specialized knowledge.
[0094] "User authentication" is the process of verifying the identity of a user using an application and granting appropriate access permissions.
[0095] "AI Module" means a technological component based on artificial intelligence, used to perform a specific task or function.
[0096] A "prompt sentence" is a sentence that includes a message or instruction that intentionally prompts the user to take a specific action.
[0097] "Video data" refers to information saved in the file format of a video shot by a user.
[0098] "Scene selection" is the process of identifying specific frames or moments within a video and extracting important parts.
[0099] An "original illustration" is an initial image generated based on a scene extracted from a video.
[0100] "Style conversion" is an image processing process used to change the appearance and texture of the original illustration.
[0101] A "generative adversarial network (GAN)" is an algorithm consisting of two neural networks for the purpose of data generation. One creates fakes, and the other identifies them, improving accuracy.
[0102] "Preview" is an intermediate display function that allows the user to check the final output.
[0103] A "stamp pack" is a data format that compiles multiple stamp images into one set.
[0104] A "download link" is a URL that allows you to access and download data stored in online storage.
[0105] A "messaging app" is a software application that allows users to exchange text, images, videos, etc.
[0106] A "database" is a system for systematically storing and managing data.
[0107] A "session ID" is a unique identifier used to identify a communication session between a server and a client.
[0108] An "AI model for scene detection" is an artificial intelligence algorithm designed to identify and extract specific scenes within videos.
[0109] A "UI framework" is a set of software tools to assist in the design and implementation of user interfaces.
[0110] "Storage" means hardware or cloud services for permanent or temporary storage of data.
[0111] The present invention relates to a system that allows users to easily create their own original stamps. Specific embodiments of the system will be described below.
[0112] When a user launches an application, the server requests user authentication information. The server checks a database (e.g., MySQL) to verify the validity of the provided authentication information. If authentication is successful, the server generates a session ID and sends it to the device.
[0113] The server then uses an AI module (e.g., OpenAI's GPT-3) to send prompts to the device with instructions about scenes suitable for the stamp, such as "Please smile for the camera and hold the pose for three seconds."
[0114] The user activates the device's camera function and captures video based on instructions provided by the server. The device records the video using the camera module (e.g., AVFoundation in iOS). The captured video is several minutes long and ensures that the specified scenes are included.
[0115] Once the video recording is complete, the device automatically uploads the recorded video data to the server. The server saves the video data in storage (e.g., Amazon S3) and begins scene analysis. This analysis uses an AI model for scene detection (e.g., TensorFlow) to detect specific frames within the video.
[0116] The identified scene is generated as a source illustration by the server, and the source illustration is then transformed into a different style, such as photorealistic, cartoonish, or painted, using a generative adversarial network (GAN) algorithm (e.g., Pix2Pix).
[0117] The generated illustration is sent from the server to the device, where it is displayed on the device's preview screen, where the user can add text and decorations (e.g., emojis, stickers, custom text) to the stamp.
[0118] Once the user has finished editing, the device sends the final sticker pack data to the server, which saves the sticker pack in a database (e.g., PostgreSQL) and generates a download link. This link is then provided to the user, who can then download the sticker pack and import it into a messaging app (e.g., LINE).
[0119] As a concrete example, consider a scenario where a user captures a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). The user can then add appropriate text and decorations to create their own original stamps.
[0120] Examples of prompts for a generative AI model include:
[0121] "Give specific instructions for capturing a smiling scene. For example, smile for the camera and hold the pose for three seconds."
[0122] Provide instructions for capturing a scene showing a surprised expression. For example, hold your hand in front of your face as if something surprising has happened and hold the surprised expression for 3 seconds.
[0123] As described above, the present invention provides an innovative system that allows users to easily create original stamps and use them in messaging apps, even if they do not have specialized knowledge.
[0124] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0125] Step 1:
[0126] When a user launches an application, the server requests user authentication information. The terminal prompts the user to enter a username and password and sends that information to the server. The server then references a database to verify the authentication information. If authentication is successful, the server generates a session ID and sends it to the terminal.
[0127] Input: Username, Password
[0128] Data Processing: Credential Verification
[0129] Output: Session ID
[0130] Step 2:
[0131] The server uses an AI module to generate a prompt message for the scene that matches the stamp and send it to the device, specifically, a prompt message containing the instruction "Smile for the camera and hold the pose for three seconds."
[0132] Input: Session ID
[0133] Data processing: Prompt sentence generation by AI module
[0134] Output: prompt statement
[0135] Step 3:
[0136] The user activates the device's camera function and takes a video based on instructions provided by the server. The device uses the camera module to record the video and takes a few minutes of video based on specific instructions from the server.
[0137] Input: prompt statement
[0138] Data processing: Video shooting by users
[0139] Output: Video data
[0140] Step 4:
[0141] Once the video recording is complete, the device uploads the recorded video data to the server. The server receives the video data and stores it in storage. Scene analysis then begins.
[0142] Input: Video data
[0143] Data processing: Uploading video data and saving it to storage
[0144] Output: Saved video data
[0145] Step 5:
[0146] The server analyzes the stored video data using an AI model for scene analysis to detect specific frames within the video. For example, it analyzes scenes where people are smiling and selects the most suitable frame. As a result, the specific scene is extracted.
[0147] Input: Saved video data
[0148] Data processing: Frame selection using AI model for scene analysis
[0149] Output: Specific Scene
[0150] Step 6:
[0151] The server generates a source illustration using a generative adversarial network (GAN) based on the identified scene, and then proceeds to convert the generated source illustration into a different style.
[0152] Input: A specific scene
[0153] Data processing: Generating original illustrations using GAN
[0154] Output: Original illustration
[0155] Step 7:
[0156] The server converts the generated original illustration into the specified style (e.g., photo-like, cartoon-like, or painting-like), and the converted illustration is sent to the device.
[0157] Input: Original illustration
[0158] Data processing: Transformation using style transfer algorithms
[0159] Output: Converted illustration
[0160] Step 8:
[0161] The converted illustration sent to the device is displayed as a preview to the user. The user can add text and decorations to the stamp on the preview screen. This editing is done using the UI framework.
[0162] Input: Converted illustration
[0163] Data processing: Editing by the user
[0164] Output: Edited stamp
[0165] Step 9:
[0166] Once the user has finished editing, the device sends the final stamp pack data to the server, which stores the stamp pack in a database and generates a download link that can be provided to the user, who can then download the stamp pack and import it into their messaging app.
[0167] Input:Edited stamp
[0168] Data processing: Saving stamp pack data and creating links
[0169] Output: Download link
[0170] (Application example 1)
[0171] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0172] In conventional quality control systems, determining product anomalies detected on the production line mainly relies on human visual inspection and manual recording, which is inefficient and has a high risk of false positives. Furthermore, it is difficult to record the detection results and immediately link them to the quality control system, resulting in a waste of resources. To solve these problems, automated stamp generation and rapid linkage with the quality control system are needed.
[0173] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0174] In this invention, the server includes means for presenting instructions to a user for generating an image, means for receiving a video shot by the user based on the instructions, means for selecting a specific scene from the video, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style, means for the user to edit the illustration, means for automatically generating a stamp based on the condition of parts or products detected on the production line, and means for linking the stamp to a quality control system. This allows a quality check stamp to be automatically generated upon detection of a product anomaly, enabling rapid linkage with the quality control system.
[0175] The "means for presenting instructions to the user for generating an image" is a process for displaying specific instructions for capturing an appropriate scene when the user shoots a video.
[0176] The "means for receiving the video shot by the user based on the instruction" is a process for transferring the video shot by the user to the server.
[0177] The "means for selecting a specific scene from the video" is a process for identifying and extracting a scene suitable for quality check from the received video data.
[0178] The "means for generating an original illustration based on the specified scene" is a process for generating a base image from the specified scene.
[0179] The "means for converting the original illustration into a different style" refers to a process for converting the generated original illustration into a different style such as a cartoon style, a live-action style, or a painting style.
[0180] The "means for the user to edit the illustration" is a process by which the user can add text and decorations to the generated illustration.
[0181] "Means for automatically generating stamps based on the condition of parts or products detected on the production line" refers to a process for automatically generating stamps indicating quality conditions based on the detection results on the production line.
[0182] The "means for linking the stamp to the quality control system" is a process for linking the created stamp to the quality control system.
[0183] This invention is implemented in a system for automating quality control in factories. The main hardware of the system includes a production line robot and a camera, and the software uses OpenCV and TensorFlow. Specifically, the invention is implemented in the following stages:
[0184] First, the server presents the user with instructions for taking a video, including specific examples of scenes that are easy to detect (e.g., scratches on the surface, dirt, abnormal shapes, etc.). Once the user takes a video based on the instructions, the video is uploaded to the server.
[0185] The server analyzes the uploaded video and selects a specific scene. Using OpenCV and TensorFlow, it scans each frame in the video and finds the best frame suitable for quality check. From this identified scene, the original illustration is generated.
[0186] The server then converts the generated original illustration into a different style, such as cartoon, photorealistic, or painterly, using a generative adversarial network (GAN) algorithm.
[0187] Users can preview the converted illustration and edit it by adding text and decorations on it, including specific labels and comments to indicate its quality status.
[0188] Finally, the server generates the edited stamp and automatically connects it to the quality control system, allowing the product quality status to be systematically recorded and quickly addressed.
[0189] For example, if abnormal soldering is detected in a factory that manufactures electronic circuit boards, the server analyzes the video and identifies the abnormal area. Next, a stamp is generated based on the scene, and text such as "defective" or "re-inspection" is added, which is then linked to the quality control system. This allows information about the abnormal area to be shared with engineers in real time.
[0190] An example of a prompt is:
[0191] Prompt statement:
[0192] Design a system to detect product anomalies to automate quality checks on electronic circuit board manufacturing lines. The system analyzes captured video, detects anomalies, and generates original stamps highlighting the anomalies. Specifically, the system detects abnormal solder joints and selects the best frame from the scene. Based on the selected frame, a generative adversarial network (GAN) is used to create a stamp, which can then be integrated into the quality control system by an operator adding text.
[0193] This allows for effective and efficient quality control.
[0194] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0195] Step 1:
[0196] The server presents the user with instructions for generating a video. Specifically, guidelines are displayed for scenes where product anomalies are likely to be detected (e.g., scratches on the surface, dirt, abnormal shapes, etc.). The user then shoots a video based on these guidelines. The input includes the scene guidelines, and the output is a state in which the user is ready to shoot.
[0197] Step 2:
[0198] The user shoots a video based on instructions from the server. The user uses a camera to shoot the specified scene and generate a video file. This video file becomes the input of the system. The output is the shot video file.
[0199] Step 3:
[0200] The user's device uploads the video they have taken to the server. The device receives the video file and sends it to the server over the network. In this process, the video file is the input and the video data stored on the server is the output.
[0201] Step 4:
[0202] The server analyzes the received video data and selects specific scenes. Specifically, it uses AI modules for video analysis, such as OpenCV and TensorFlow, to scan each frame in the video and find the best frame suitable for quality check. The input is the video data uploaded to the server, and the output is the frame corresponding to the selected specific scene.
[0203] Step 5:
[0204] The server generates a source illustration based on a specific scene. Here, the selected frame image is processed and generated as the source illustration. The generated illustration is the basis for quality checks. The input is a frame image of a specific scene, and the output is the source illustration.
[0205] Step 6:
[0206] The server converts the original illustration into a different style using a generative adversarial network (GAN). The input is the original illustration, and the output is the style-converted illustration.
[0207] Step 7:
[0208] The user edits the converted illustration. The server presents an editing interface to the user, who uses this interface to add text and decorations to the illustration. The input is the style-converted illustration, and the output is the final illustration after user editing.
[0209] Step 8:
[0210] The server generates a stamp based on the edited illustration and links it to the quality control system. The generated stamp is automatically sent to the quality control system, where the quality status of the product is recorded. The input is the final illustration edited by the user, and the output is a stamp linked to the quality control system.
[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0212] The present invention is a system that allows users to easily create their own original stamps. In particular, it combines an emotion engine to recognize the user's emotions and generate stamps based on those emotions. This system consists of the following main steps:
[0213] Program processing explanation
[0214] Initialization and User Authentication
[0215] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[0216] Video shooting scene suggestions
[0217] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[0218] Video recording
[0219] The user uses the device's camera function to shoot video based on the specified scene, and the user manually starts and stops shooting.
[0220] emotion recognition
[0221] Once the video is uploaded, the server uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[0222] Video upload and scene selection
[0223] Video data is uploaded from the device to the server, which then selects specific scenes based on the results of the emotion engine, selecting scenes that most clearly express specific emotions (e.g., joy, surprise, sadness).
[0224] Image processing and style transfer
[0225] Once the selected scene is determined, the server generates a source illustration from that scene, which is then converted into the user's chosen style (e.g., photorealistic, cartoonish, or painted). The conversion process uses machine learning algorithms such as generative adversarial networks (GANs).
[0226] Preview and edit stamps
[0227] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, including emojis, stickers, and custom text.
[0228] Creating and saving stamp packs
[0229] Once the user has completed the editing process, the device sends the final stamp pack to the server, which saves it and generates a download link, which is then sent to the user.
[0230] Import to LINE
[0231] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[0232] Specific examples
[0233] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[0234] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[0235] The processing flow will be explained below.
[0236] Step 1:
[0237] The user launches the application, and the login screen appears.
[0238] Step 2:
[0239] The server prompts the user for authentication information: a login screen prompts for username and password.
[0240] Step 3:
[0241] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[0242] Step 4:
[0243] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[0244] Step 5:
[0245] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[0246] Step 6:
[0247] The user uses the device's camera to shoot a video based on the specified scene. The user presses the recording button to start recording. The video is set to be a few minutes long and includes facial expressions of joy, anger, sadness, and happiness, as well as other recommended scenes.
[0248] Step 7:
[0249] When the user presses the end button, the video recording is complete. The device automatically saves the video data and prepares it for uploading to the server.
[0250] Step 8:
[0251] The device uploads the video data to the server, where the video file is compressed and converted into the appropriate format.
[0252] Step 9:
[0253] The server receives the uploaded video data and uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions, voice, and movements to identify each emotion: joy, anger, sadness, and happiness.
[0254] Step 10:
[0255] The server selects specific scenes based on the analysis results of the emotion engine, for example, selecting frames that show strong emotions, such as when the user is laughing the most or looking the most surprised.
[0256] Step 11:
[0257] The server generates an original illustration from the selected scene, which is temporarily stored in the system.
[0258] Step 12:
[0259] The server converts the original illustration into the style selected by the user. For example, the process of converting from a realistic style to a cartoon style or a painting style uses machine learning algorithms such as generative adversarial networks (GANs).
[0260] Step 13:
[0261] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[0262] Step 14:
[0263] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[0264] Step 15:
[0265] The device sends the edited stamp pack to the server, which stores the received stamp pack and generates a download link.
[0266] Step 16:
[0267] The server then sends the generated download link to the user, who then downloads the stamp pack and imports it into the LINE app.
[0268] Step 17:
[0269] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[0270] Example 2
[0271] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0272] In conventional technologies, the process for users to create their own original stamps is complicated, and creating stamps that reflect emotions is particularly difficult. Another issue is that the processes for style conversion and emotion recognition are complex, requiring a high level of technical knowledge for the average user.
[0273] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server provides a means for analyzing a user's emotions using an emotion recognition engine, a means for suggesting video scenes to present to the user, a means for receiving a video shot by the user based on the suggestions, a means for selecting a specific scene from the video based on the emotion, a means for generating an original illustration based on the specific scene, a means for converting the original illustration into a different style using a generative adversarial network, and a means for the user to edit the illustration. This allows the user to easily create original stamps that reflect their emotions.
[0274] An "emotion recognition engine" is a software or hardware system that analyzes a user's facial expressions, tone of voice, etc., and identifies emotions.
[0275] The "means for suggesting video shooting scenes" is a function that presents scenes containing specific emotions to the user and gives instructions for the user to shoot a video based on those scenes.
[0276] "Means for receiving video" refers to a mechanism for transmitting and receiving video data shot by a user to a server or system.
[0277] "Means for selecting specific scenes" refers to an algorithm or process that automatically selects scenes from a user's video that most clearly express a specific emotion.
[0278] The "means for generating the original illustration" is the technique or process used to create an illustration from the selected scene.
[0279] A "generative adversarial network" is a type of machine learning algorithm that generates data by pitting two neural networks against each other.
[0280] "Methods of converting into a different style" refers to techniques or processes used to convert the original illustration into a specific style (e.g., cartoon style, live-action style, painterly style, etc.).
[0281] "Means for the user to edit the illustration" refers to an interface or tool that allows the user to add text or decorations to the illustration.
[0282] The present invention is a system that allows users to easily create their own original stamps, and in particular provides a function that recognizes the user's emotions by combining an emotion engine and generates stamps based on those emotions. This system consists of the following main steps:
[0283] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[0284] The server uses an AI module to instruct the user on scenes (e.g., emotions) that are suitable for the stamps. These instructions are displayed on the device screen. The user then uses the device's camera function to shoot a video based on the instructed scenes. The user manually starts and stops recording.
[0285] Once the video has been filmed and uploaded, the server uses an emotion engine to analyze the user's emotions in the video. Existing emotion engines, such as Google Cloud AI, AWS Rekognition, and Microsoft Azure Emotion API, can be used. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[0286] Video data is uploaded from the device to the server. The server selects specific scenes based on the results of the emotion engine. It selects scenes that most clearly express specific emotions (e.g., joy, surprise, sadness). The server analyzes the emotion recognition results and determines the most suitable scene. It then extracts specific frames from the selected scene.
[0287] Once the selected scene is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs) to generate the original illustration, which is then converted into the user's chosen style (e.g., cartoon, live action, painterly).
[0288] The converted stamp image is sent to the device and a preview is displayed to the user. The user can then add text and decorations (e.g., emojis, stickers, custom text, etc.) to the stamp. Once the user has finished editing, the device sends the final stamp pack to the server. The server saves the stamp pack and generates a download link for the user.
[0289] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[0290] Specific examples
[0291] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[0292] Example input to a generative AI model
[0293] An example of a prompt is:
[0294] "Take a video that expresses the emotion of joy."
[0295] "Please choose the scene in the video where you can see your smile most clearly."
[0296] "Take frames extracted from that scene and convert them into a cartoon-like style."
[0297] "Add custom text and emojis to the generated stamps."
[0298] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[0299] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0300] Step 1: Initialization and User Authentication
[0301] When a user launches an application, the server requests the user's authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information against a database. The input is the username and password, and the output is the authentication result. If authentication is successful, the user proceeds to the next step.
[0302] Specific behavior:
[0303] The user launches the app and enters their username and password on the authentication screen.
[0304] The authentication information is sent to the server, which then performs authentication by referencing a database.
[0305] The user is notified whether the authentication was successful or not.
[0306] Step 2: Propose a video shooting scene
[0307] The server uses an AI module to suggest scenes suitable for stamps (e.g., emotions). The suggested scenes are displayed on the device screen. The input is a request for a scene to be captured, and the output is a list of scenes generated by the AI.
[0308] Specific behavior:
[0309] The server sends a request for scene suggestions to the AI module.
[0310] The AI module generates multiple scenes, and a list of these is provided to the device via the server.
[0311] The proposed scene is displayed to the user.
[0312] Step 3: Record a video
[0313] The user shoots a video using the device's camera function based on the selected scene. The input is the user's scene selection, and the output is the shot video. The user manually starts and stops shooting.
[0314] Specific behavior:
[0315] The user selects one of the suggested scenes.
[0316] Launch the camera app and tap the record button to start recording.
[0317] When you're done recording, tap the stop button.
[0318] Step 4: Upload video and recognize emotions
[0319] The video is uploaded from the device to the server, where the server uses an emotion engine to analyze the user's emotions in the video. The input is the video, and the output is the result of the emotion analysis.
[0320] Specific behavior:
[0321] The device uploads the captured video file to the server.
[0322] The server passes the video to the emotion engine and requests analysis.
[0323] The emotion recognition results are sent back to the server.
[0324] Step 5: Scene selection and frame extraction
[0325] The server selects the scene that most clearly expresses emotion based on the results of the emotion engine, and then extracts the optimal frame from that scene. The input is the emotion analysis result, and the output is the selected frame.
[0326] Specific behavior:
[0327] The server analyzes the emotion recognition results and determines the most appropriate scene.
[0328] Extract a specific frame from the selected scene.
[0329] Step 6: Image processing and style transfer
[0330] Once the selected frame is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs). The input is the extracted frame, and the output is the illustration with the converted style.
[0331] Specific behavior:
[0332] The server generates the original illustration from the extracted frames.
[0333] The illustration is passed through a style conversion algorithm to convert it to the specified style.
[0334] Step 7: Preview and edit your stamp
[0335] The converted stamp image is sent to the device and a preview is displayed to the user. The user can add text and decorations to the stamp on this preview screen. The input is the style-converted illustration, and the output is the final stamp image edited by the user.
[0336] Specific behavior:
[0337] The server transmits the converted stamp image to the terminal.
[0338] The user uses the editing tools to add text and decorations.
[0339] Step 8: Create and save your sticker pack
[0340] When the user finishes editing, the device sends the final stamp pack to the server, which saves the stamp pack, generates a download link, and notifies the user. The input is the final stamp image, and the output is the download link.
[0341] Specific behavior:
[0342] The user taps the "Done" button to send the stamp pack to the server.
[0343] The server stores the stamp pack and generates a download link.
[0344] The download link will be sent to the user.
[0345] Step 9: Import to LINE
[0346] Users download the stamp pack via the provided link. The stamp pack is imported into the LINE app, allowing users to use their custom stamps on LINE. The input is the download link, and the output is the stamps in the LINE app.
[0347] Specific behavior:
[0348] The user clicks on the download link to get the stamp pack.
[0349] Import the stamp pack into the LINE app.
[0350] (Application example 2)
[0351] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0352] Conventional stamp generation systems have had the problem of making it difficult for users to easily create original stamps that reflect their own emotions. Furthermore, the technology to easily change the style of stamps and make them uploadable to virtual stores was insufficient. Furthermore, it was difficult to utilize an emotion engine to automatically select scenes that best express emotions.
[0353] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for requesting user authentication information and verifying the authentication information, means for presenting instructions for generating images to the user, means for receiving a video captured by the user based on the instructions, means for selecting a specific scene from the video and recognizing the user's emotion using an emotion engine, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style and using a generative adversarial network, means for the user to edit the illustration, and means for saving the user-edited stamp pack on the server and generating a download link. This allows for easy and effective generation of original stamps that reflect the user's emotion and uploading the style-converted stamps to a virtual store, enabling rich communication.
[0354] "Authentication information" refers to information such as a username and password that a user uses when accessing a system.
[0355] A "server" is a computer system that receives requests from users and performs the processing in response to them.
[0356] "Instructions" are information that prompts the system to perform specific operations or actions on the user.
[0357] A "video" is a media file that contains a series of image frames captured by a user.
[0358] The "emotion engine" is a software module that analyzes and recognizes user emotions from video and audio.
[0359] A "specific scene" is a selected part of a video that most clearly expresses an emotion or a specific event.
[0360] An "original illustration" is an image generated based on a specific scene extracted from a video.
[0361] "Converting to a style" refers to the process of changing the look and feel of the original illustration into a different style (e.g., cartoon, live-action, painterly, etc.).
[0362] A "generative adversarial network (GAN)" is a machine learning algorithm in which one network generates images and another network evaluates the generated images.
[0363] "Editing" refers to customization work that the user performs on the generated illustration (e.g., adding text, inserting decorations, etc.).
[0364] A "stamp pack" is a package containing multiple stamp images.
[0365] A "download link" is a URL that allows a user to download a file over the Internet.
[0366] A "virtual store" is a digital shopping platform that exists on the Internet.
[0367] System Configuration
[0368] This invention is a system that allows users to easily create original stamps that reflect their emotions, convert the stamps into different styles, and make them available in virtual stores. This system is composed of a server, terminals, and related software modules.
[0369] Processing flow
[0370] User Authentication
[0371] When a user launches an application, the server requests the user's authentication information. The user enters their username and password into the terminal and sends them to the server. The server verifies the authentication information, and if authentication is successful, the user can proceed to the next step.
[0372] Directing and filming video scenes
[0373] The server provides instructions to the device for generating images. Specifically, it works in conjunction with the emotion engine to present scenes that the user should capture (e.g., smile, surprise, sadness) to the device. The user then uses the device's camera function to capture these scenes as videos and upload them to the server.
[0374] Emotion Recognition and Scene Selection
[0375] The server receives videos uploaded by users and analyzes the user's emotions in the videos using an emotion engine. Based on this emotion analysis, a specific scene is selected from multiple scenes. The selection criterion is the scene that most clearly expresses the emotion.
[0376] Style conversion and stamp generation
[0377] The server generates a source illustration based on the specific scene selected, then converts the generated source illustration into different styles (e.g., cartoon, live action, painterly) using machine learning algorithms such as generative adversarial networks (GANs).
[0378] Editing and saving stamps
[0379] The converted stamp images are sent to the device, where the user can customize them by adding text, emojis, stickers, etc. Once editing is complete, the device sends the final stamp pack to the server, which stores it and provides a download link to the user.
[0380] Hardware and software used
[0381] Hardware: Cameras on smartphones and head-mounted displays (HMDs)
[0382] software:
[0383] EmotionRecognizer (emotion recognition engine)
[0384] StyleTransfer (style transfer algorithm, GAN)
[0385] OpenCV (camera operation and image processing library)
[0386] Requests (HTTP client library for sending API requests)
[0387] Specific examples
[0388] For example, suppose a user makes a surprised expression while virtual shopping. This moment is captured by the camera, and EmotionRecognizer recognizes it as "surprise." StyleTransfer then converts the expression into a cartoon-style image based on this emotion. Finally, the generated sticker is saved in the user's account via the virtual store's API.
[0389] Prompt Sentence Examples
[0390] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[0391] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0392] Step 1: Enter and verify your credentials
[0393] Subject: User and Server
[0394] Specific operation: The user launches an application and enters authentication information (username and password) into the device. The device sends this information to the server. The server verifies the received authentication information, and if authentication is successful, proceeds to the next step.
[0395] Input: Username and Password
[0396] Output: Authentication success flag
[0397] Step 2: Video Shooting Scene Instructions
[0398] Subject: Server and Terminal
[0399] Specific operation: The server uses the emotion engine to present the scene to be captured (e.g., smile, surprise, sadness) on the device screen. The user acts out the scene according to the instructions.
[0400] Input: Authentication success flag
[0401] Output: Instructions for the shooting scene
[0402] Step 3: Record your video
[0403] Subject: User and Device
[0404] Specific operation: The user uses the device's camera function to shoot a video based on the specified scene. The user starts and stops shooting.
[0405] Input: Instructions for the shooting scene
[0406] Output: Recorded video file
[0407] Step 4: Upload your video
[0408] Subject: Terminal and Server
[0409] Specific operation: The device uploads the captured video file to the server, which receives and stores the video file.
[0410] Input: Recorded video file
[0411] Output: Video data stored on the server
[0412] Step 5: Emotion recognition and specific scene selection
[0413] Subject: Server
[0414] Specific operation: The server analyzes the received video using the emotion engine to recognize the user's emotion in the video, and then selects the specific scene that most clearly expresses the emotion from multiple scenes.
[0415] Input: Video data stored on the server
[0416] Output: Specific scene information
[0417] Step 6: Generate the original illustration
[0418] Subject: Server
[0419] Specific operation: The server generates a source illustration based on the selected specific scene. The source illustration is an image extracted from a video frame.
[0420] Input: Specific scene information
[0421] Output: Original illustration image
[0422] Step 7: Style Transformation
[0423] Subject: Server
[0424] Specific operation: The server converts the generated original illustration into different styles (e.g., cartoon, live-action, painting) using a generative adversarial network (GAN).
[0425] Input: Original illustration image
[0426] Output: Style-converted illustration
[0427] Step 8: Editing the stamp
[0428] Subject: User and Device
[0429] Specific operation: The converted stamp image is sent to the device, and the user can edit it. Users can add text, emojis, stickers, etc.
[0430] Input: Style-converted illustration
[0431] Output: Edited stamp image
[0432] Step 9: Save and share your sticker pack
[0433] Subject: Terminal and Server
[0434] Specific operation: The device sends the edited stamp pack to the server, which stores it and provides a download link to the user.
[0435] Input: Edited stamp image
[0436] Output: Saved sticker pack and download link
[0437] Prompt Sentence Examples
[0438] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[0439] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0440] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0441] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0442] [Second embodiment]
[0443] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0444] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0445] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0446] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0447] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0448] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0449] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0450] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0451] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0452] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0453] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0454] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0455] The present invention aims to provide a system that allows users to easily create their own original stamps. This system includes the steps of presenting instructions for image generation, allowing the user to shoot a video, selecting a specific scene from the video, generating an original illustration from the selected scene, converting the original illustration into a different style, and allowing the user to edit the illustration.
[0456] Program processing explanation
[0457] Initialization and User Authentication
[0458] When a user launches an application, the server prompts for authentication information. The user enters a username and password and submits them to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[0459] Video shooting scene suggestions
[0460] The server uses an AI module to send instructions to the device about scenes suitable for the stamps, such as specific facial expressions or specific poses, for the user to capture in the video.
[0461] Video recording
[0462] The user uses the camera function of the device to shoot a video based on instructions from the server. The video is set to be a few minutes long and ensures that the specified scenes are included.
[0463] Video upload and scene selection
[0464] Once the video is complete, the device automatically uploads the video data to the server, which then analyzes it and selects specific scenes. This process uses an AI algorithm to scan each frame in the video and find the best frame that corresponds to the recommended scene.
[0465] Image processing and style transfer
[0466] Once the selected scene is determined, the server generates an original illustration from that scene. The server then converts the generated original illustration into a different style. This conversion process uses machine learning algorithms such as generative adversarial networks (GANs) to convert the illustration into a style such as photorealistic, cartoonish, or pictorial.
[0467] Preview and edit stamps
[0468] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, such as emojis, stickers, and custom text.
[0469] Creating and saving stamp packs
[0470] Once the user has finished editing, the device will send the final stamp pack to the server, which will save it and issue a download link, through which the user can download the stamp pack and import it into LINE.
[0471] Specific examples
[0472] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to the stamp to create their own original stamp.
[0473] In this way, the present invention provides a method and system that allows easy creation of original stamps tailored to the needs of individual users.
[0474] The processing flow will be explained below.
[0475] Step 1:
[0476] The user launches the application, and the login screen appears.
[0477] Step 2:
[0478] The server prompts the user for authentication information: a login screen prompts for username and password.
[0479] Step 3:
[0480] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[0481] Step 4:
[0482] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[0483] Step 5:
[0484] The server uses an AI module to instruct the user on the scenes (e.g., emotions) that are suitable for the stamp. The instructions are displayed on the device screen.
[0485] Step 6:
[0486] The user uses the device's camera function to shoot a video based on the specified scene, and the user manually starts and stops video shooting.
[0487] Step 7:
[0488] Once the video recording is complete, the device will automatically upload the video data to the server, and the upload progress will be displayed on the screen.
[0489] Step 8:
[0490] The server receives the uploaded video data and uses AI algorithms to analyze each frame of the video, automatically selecting the best frame that corresponds to the recommended scene.
[0491] Step 9:
[0492] The server generates an original illustration based on the selected frame, and this original illustration is saved in the system.
[0493] Step 10:
[0494] The server converts the original illustration into a style selected by the user (e.g., photo-realistic, cartoon-like, or painted style). Style conversion is performed using machine learning algorithms such as generative adversarial networks (GANs).
[0495] Step 11:
[0496] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[0497] Step 12:
[0498] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[0499] Step 13:
[0500] The device sends the edited stamp pack to the server, which stores the received stamp pack.
[0501] Step 14:
[0502] Once the server has finished saving the stamp pack, it will generate a download link and notify the user of this link.
[0503] Step 15:
[0504] Users can download the stamp pack via the provided link and import it into the LINE app, allowing them to use their custom stamps on LINE.
[0505] Example 1
[0506] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0507] In conventional stamp creation systems, the process for users to create their own unique images is complicated and requires many operations, making it difficult to easily generate stamps. Furthermore, manual image editing and style conversion require specialized knowledge, making them difficult for general users to use. This has led to a demand for a new method and system that allows users to easily create original stamps that are optimal for each individual user.
[0508] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0509] In this invention, the server is a system that allows users to easily create their own original stamps, and includes: means for requesting authentication information from the user and authenticating them by referencing a database; means for generating instructions regarding scenes suitable for stamps using an AI model and transmitting the instructions to the terminal; means for receiving from the terminal a video shot by the user based on the instructions; means for analyzing the uploaded video and selecting a specific scene; means for generating an original illustration based on the specific scene and converting the generated original illustration into a different style; means for transmitting the converted illustration to the terminal and allowing the user to edit the illustration; and means for saving the edited stamp pack and issuing a download link. This allows users to easily create original stamps and use them in messaging apps without specialized knowledge.
[0510] "User authentication" is the process of verifying the identity of a user using an application and granting appropriate access permissions.
[0511] "AI Module" means a technological component based on artificial intelligence, used to perform a specific task or function.
[0512] A "prompt sentence" is a sentence that includes a message or instruction that intentionally prompts the user to take a specific action.
[0513] "Video data" refers to information saved in the file format of a video shot by a user.
[0514] "Scene selection" is the process of identifying specific frames or moments within a video and extracting important parts.
[0515] An "original illustration" is an initial image generated based on a scene extracted from a video.
[0516] "Style conversion" is an image processing process used to change the appearance and texture of the original illustration.
[0517] A "generative adversarial network (GAN)" is an algorithm consisting of two neural networks for the purpose of data generation. One creates fakes, and the other identifies them, improving accuracy.
[0518] "Preview" is an intermediate display function that allows the user to check the final output.
[0519] A "stamp pack" is a data format that compiles multiple stamp images into one set.
[0520] A "download link" is a URL that allows you to access and download data stored in online storage.
[0521] A "messaging app" is a software application that allows users to exchange text, images, videos, etc.
[0522] A "database" is a system for systematically storing and managing data.
[0523] A "session ID" is a unique identifier used to identify a communication session between a server and a client.
[0524] An "AI model for scene detection" is an artificial intelligence algorithm designed to identify and extract specific scenes within videos.
[0525] A "UI framework" is a set of software tools to assist in the design and implementation of user interfaces.
[0526] "Storage" means hardware or cloud services for permanent or temporary storage of data.
[0527] The present invention relates to a system that allows users to easily create their own original stamps. Specific embodiments of the system will be described below.
[0528] When a user launches an application, the server requests user authentication information. The server checks a database (e.g., MySQL) to verify the validity of the provided authentication information. If authentication is successful, the server generates a session ID and sends it to the device.
[0529] The server then uses an AI module (e.g., OpenAI's GPT-3) to send prompts to the device with instructions about scenes suitable for the stamp, such as "Please smile for the camera and hold the pose for three seconds."
[0530] The user activates the device's camera function and captures video based on instructions provided by the server. The device records the video using the camera module (e.g., AVFoundation in iOS). The captured video is several minutes long and ensures that the specified scenes are included.
[0531] Once the video recording is complete, the device automatically uploads the recorded video data to the server. The server saves the video data in storage (e.g., Amazon S3) and begins scene analysis. This analysis uses an AI model for scene detection (e.g., TensorFlow) to detect specific frames within the video.
[0532] The identified scene is generated as a source illustration by the server, and the source illustration is then transformed into a different style, such as photorealistic, cartoonish, or painted, using a generative adversarial network (GAN) algorithm (e.g., Pix2Pix).
[0533] The generated illustration is sent from the server to the device, where it is displayed on the device's preview screen, where the user can add text and decorations (e.g., emojis, stickers, custom text) to the stamp.
[0534] Once the user has finished editing, the device sends the final sticker pack data to the server, which saves the sticker pack in a database (e.g., PostgreSQL) and generates a download link. This link is then provided to the user, who can then download the sticker pack and import it into a messaging app (e.g., LINE).
[0535] As a concrete example, consider a scenario where a user captures a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). The user can then add appropriate text and decorations to create their own original stamps.
[0536] Examples of prompts for a generative AI model include:
[0537] "Give specific instructions for capturing a smiling scene. For example, smile for the camera and hold the pose for three seconds."
[0538] Provide instructions for capturing a scene showing a surprised expression. For example, hold your hand in front of your face as if something surprising has happened and hold the surprised expression for 3 seconds.
[0539] As described above, the present invention provides an innovative system that allows users to easily create original stamps and use them in messaging apps, even if they do not have specialized knowledge.
[0540] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0541] Step 1:
[0542] When a user launches an application, the server requests user authentication information. The terminal prompts the user to enter a username and password and sends that information to the server. The server then references a database to verify the authentication information. If authentication is successful, the server generates a session ID and sends it to the terminal.
[0543] Input: Username, Password
[0544] Data Processing: Credential Verification
[0545] Output: Session ID
[0546] Step 2:
[0547] The server uses an AI module to generate a prompt message for the scene that matches the stamp and send it to the device, specifically, a prompt message containing the instruction "Smile for the camera and hold the pose for three seconds."
[0548] Input: Session ID
[0549] Data processing: Prompt sentence generation by AI module
[0550] Output: prompt statement
[0551] Step 3:
[0552] The user activates the device's camera function and takes a video based on instructions provided by the server. The device uses the camera module to record the video and takes a few minutes of video based on specific instructions from the server.
[0553] Input: prompt statement
[0554] Data processing: Video shooting by users
[0555] Output: Video data
[0556] Step 4:
[0557] Once the video recording is complete, the device uploads the recorded video data to the server. The server receives the video data and stores it in storage. Scene analysis then begins.
[0558] Input: Video data
[0559] Data processing: Uploading video data and saving it to storage
[0560] Output: Saved video data
[0561] Step 5:
[0562] The server analyzes the stored video data using an AI model for scene analysis to detect specific frames within the video. For example, it analyzes scenes where people are smiling and selects the most suitable frame. As a result, the specific scene is extracted.
[0563] Input: Saved video data
[0564] Data processing: Frame selection using AI model for scene analysis
[0565] Output: Specific Scene
[0566] Step 6:
[0567] The server generates a source illustration using a generative adversarial network (GAN) based on the identified scene, and then proceeds to convert the generated source illustration into a different style.
[0568] Input: A specific scene
[0569] Data processing: Generating original illustrations using GAN
[0570] Output: Original illustration
[0571] Step 7:
[0572] The server converts the generated original illustration into the specified style (e.g., photo-like, cartoon-like, or painting-like), and the converted illustration is sent to the device.
[0573] Input: Original illustration
[0574] Data processing: Transformation using style transfer algorithms
[0575] Output: Converted illustration
[0576] Step 8:
[0577] The converted illustration sent to the device is displayed as a preview to the user. The user can add text and decorations to the stamp on the preview screen. This editing is done using the UI framework.
[0578] Input: Converted illustration
[0579] Data processing: Editing by the user
[0580] Output: Edited stamp
[0581] Step 9:
[0582] Once the user has finished editing, the device sends the final stamp pack data to the server, which stores the stamp pack in a database and generates a download link that can be provided to the user, who can then download the stamp pack and import it into their messaging app.
[0583] Input:Edited stamp
[0584] Data processing: Saving stamp pack data and creating links
[0585] Output: Download link
[0586] (Application example 1)
[0587] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0588] In conventional quality control systems, determining product anomalies detected on the production line mainly relies on human visual inspection and manual recording, which is inefficient and has a high risk of false positives. Furthermore, it is difficult to record the detection results and immediately link them to the quality control system, resulting in a waste of resources. To solve these problems, automated stamp generation and rapid linkage with the quality control system are needed.
[0589] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0590] In this invention, the server includes means for presenting instructions to a user for generating an image, means for receiving a video shot by the user based on the instructions, means for selecting a specific scene from the video, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style, means for the user to edit the illustration, means for automatically generating a stamp based on the condition of parts or products detected on the production line, and means for linking the stamp to a quality control system. This allows a quality check stamp to be automatically generated upon detection of a product anomaly, enabling rapid linkage with the quality control system.
[0591] The "means for presenting instructions to the user for generating an image" is a process for displaying specific instructions for capturing an appropriate scene when the user shoots a video.
[0592] The "means for receiving the video shot by the user based on the instruction" is a process for transferring the video shot by the user to the server.
[0593] The "means for selecting a specific scene from the video" is a process for identifying and extracting a scene suitable for quality check from the received video data.
[0594] The "means for generating an original illustration based on the specified scene" is a process for generating a base image from the specified scene.
[0595] The "means for converting the original illustration into a different style" refers to a process for converting the generated original illustration into a different style such as a cartoon style, a live-action style, or a painting style.
[0596] The "means for the user to edit the illustration" is a process by which the user can add text and decorations to the generated illustration.
[0597] "Means for automatically generating stamps based on the condition of parts or products detected on the production line" refers to a process for automatically generating stamps indicating quality conditions based on the detection results on the production line.
[0598] The "means for linking the stamp to the quality control system" is a process for linking the created stamp to the quality control system.
[0599] This invention is implemented in a system for automating quality control in factories. The main hardware of the system includes a production line robot and a camera, and the software uses OpenCV and TensorFlow. Specifically, the invention is implemented in the following stages:
[0600] First, the server presents the user with instructions for taking a video, including specific examples of scenes that are easy to detect (e.g., scratches on the surface, dirt, abnormal shapes, etc.). Once the user takes a video based on the instructions, the video is uploaded to the server.
[0601] The server analyzes the uploaded video and selects a specific scene. Using OpenCV and TensorFlow, it scans each frame in the video and finds the best frame suitable for quality check. From this identified scene, the original illustration is generated.
[0602] The server then converts the generated original illustration into a different style, such as cartoon, photorealistic, or painterly, using a generative adversarial network (GAN) algorithm.
[0603] Users can preview the converted illustration and edit it by adding text and decorations on it, including specific labels and comments to indicate its quality status.
[0604] Finally, the server generates the edited stamp and automatically connects it to the quality control system, allowing the product quality status to be systematically recorded and quickly addressed.
[0605] For example, if abnormal soldering is detected in a factory that manufactures electronic circuit boards, the server analyzes the video and identifies the abnormal area. Next, a stamp is generated based on the scene, and text such as "defective" or "re-inspection" is added, which is then linked to the quality control system. This allows information about the abnormal area to be shared with engineers in real time.
[0606] An example of a prompt is:
[0607] Prompt statement:
[0608] Design a system to detect product anomalies to automate quality checks on electronic circuit board manufacturing lines. The system analyzes captured video, detects anomalies, and generates original stamps highlighting the anomalies. Specifically, the system detects abnormal solder joints and selects the best frame from the scene. Based on the selected frame, a generative adversarial network (GAN) is used to create a stamp, which can then be integrated into the quality control system by an operator adding text.
[0609] This allows for effective and efficient quality control.
[0610] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0611] Step 1:
[0612] The server presents the user with instructions for generating a video. Specifically, guidelines are displayed for scenes where product anomalies are likely to be detected (e.g., scratches on the surface, dirt, abnormal shapes, etc.). The user then shoots a video based on these guidelines. The input includes the scene guidelines, and the output is a state in which the user is ready to shoot.
[0613] Step 2:
[0614] The user shoots a video based on instructions from the server. The user uses a camera to shoot the specified scene and generate a video file. This video file becomes the input of the system. The output is the shot video file.
[0615] Step 3:
[0616] The user's device uploads the video they have taken to the server. The device receives the video file and sends it to the server over the network. In this process, the video file is the input and the video data stored on the server is the output.
[0617] Step 4:
[0618] The server analyzes the received video data and selects specific scenes. Specifically, it uses AI modules for video analysis, such as OpenCV and TensorFlow, to scan each frame in the video and find the best frame suitable for quality check. The input is the video data uploaded to the server, and the output is the frame corresponding to the selected specific scene.
[0619] Step 5:
[0620] The server generates a source illustration based on a specific scene. Here, the selected frame image is processed and generated as the source illustration. The generated illustration is the basis for quality checks. The input is a frame image of a specific scene, and the output is the source illustration.
[0621] Step 6:
[0622] The server converts the original illustration into a different style using a generative adversarial network (GAN). The input is the original illustration, and the output is the style-converted illustration.
[0623] Step 7:
[0624] The user edits the converted illustration. The server presents an editing interface to the user, who uses this interface to add text and decorations to the illustration. The input is the style-converted illustration, and the output is the final illustration after user editing.
[0625] Step 8:
[0626] The server generates a stamp based on the edited illustration and links it to the quality control system. The generated stamp is automatically sent to the quality control system, where the quality status of the product is recorded. The input is the final illustration edited by the user, and the output is a stamp linked to the quality control system.
[0627] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0628] The present invention is a system that allows users to easily create their own original stamps. In particular, it combines an emotion engine to recognize the user's emotions and generate stamps based on those emotions. This system consists of the following main steps:
[0629] Program processing explanation
[0630] Initialization and User Authentication
[0631] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[0632] Video shooting scene suggestions
[0633] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[0634] Video recording
[0635] The user uses the device's camera function to shoot video based on the specified scene, and the user manually starts and stops shooting.
[0636] emotion recognition
[0637] Once the video is uploaded, the server uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[0638] Video upload and scene selection
[0639] Video data is uploaded from the device to the server, which then selects specific scenes based on the results of the emotion engine, selecting scenes that most clearly express specific emotions (e.g., joy, surprise, sadness).
[0640] Image processing and style transfer
[0641] Once the selected scene is determined, the server generates a source illustration from that scene, which is then converted into the user's chosen style (e.g., photorealistic, cartoonish, or painted). The conversion process uses machine learning algorithms such as generative adversarial networks (GANs).
[0642] Preview and edit stamps
[0643] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, including emojis, stickers, and custom text.
[0644] Creating and saving stamp packs
[0645] Once the user has completed the editing process, the device sends the final stamp pack to the server, which saves it and generates a download link, which is then sent to the user.
[0646] Import to LINE
[0647] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[0648] Specific examples
[0649] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[0650] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[0651] The processing flow will be explained below.
[0652] Step 1:
[0653] The user launches the application, and the login screen appears.
[0654] Step 2:
[0655] The server prompts the user for authentication information: a login screen prompts for username and password.
[0656] Step 3:
[0657] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[0658] Step 4:
[0659] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[0660] Step 5:
[0661] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[0662] Step 6:
[0663] The user uses the device's camera to shoot a video based on the specified scene. The user presses the recording button to start recording. The video is set to be a few minutes long and includes facial expressions of joy, anger, sadness, and happiness, as well as other recommended scenes.
[0664] Step 7:
[0665] When the user presses the end button, the video recording is complete. The device automatically saves the video data and prepares it for uploading to the server.
[0666] Step 8:
[0667] The device uploads the video data to the server, where the video file is compressed and converted into the appropriate format.
[0668] Step 9:
[0669] The server receives the uploaded video data and uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions, voice, and movements to identify each emotion: joy, anger, sadness, and happiness.
[0670] Step 10:
[0671] The server selects specific scenes based on the analysis results of the emotion engine, for example, selecting frames that show strong emotions, such as when the user is laughing the most or looking the most surprised.
[0672] Step 11:
[0673] The server generates an original illustration from the selected scene, which is temporarily stored in the system.
[0674] Step 12:
[0675] The server converts the original illustration into the style selected by the user. For example, the process of converting from a realistic style to a cartoon style or a painting style uses machine learning algorithms such as generative adversarial networks (GANs).
[0676] Step 13:
[0677] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[0678] Step 14:
[0679] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[0680] Step 15:
[0681] The device sends the edited stamp pack to the server, which stores the received stamp pack and generates a download link.
[0682] Step 16:
[0683] The server then sends the generated download link to the user, who then downloads the stamp pack and imports it into the LINE app.
[0684] Step 17:
[0685] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[0686] Example 2
[0687] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0688] In conventional technologies, the process for users to create their own original stamps is complicated, and creating stamps that reflect emotions is particularly difficult. Another issue is that the processes for style conversion and emotion recognition are complex, requiring a high level of technical knowledge for the average user.
[0689] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server provides a means for analyzing a user's emotions using an emotion recognition engine, a means for suggesting video scenes to present to the user, a means for receiving a video shot by the user based on the suggestions, a means for selecting a specific scene from the video based on the emotion, a means for generating an original illustration based on the specific scene, a means for converting the original illustration into a different style using a generative adversarial network, and a means for the user to edit the illustration. This allows the user to easily create original stamps that reflect their emotions.
[0690] An "emotion recognition engine" is a software or hardware system that analyzes a user's facial expressions, tone of voice, etc., and identifies emotions.
[0691] The "means for suggesting video shooting scenes" is a function that presents scenes containing specific emotions to the user and gives instructions for the user to shoot a video based on those scenes.
[0692] "Means for receiving video" refers to a mechanism for transmitting and receiving video data shot by a user to a server or system.
[0693] "Means for selecting specific scenes" refers to an algorithm or process that automatically selects scenes from a user's video that most clearly express a specific emotion.
[0694] The "means for generating the original illustration" is the technique or process used to create an illustration from the selected scene.
[0695] A "generative adversarial network" is a type of machine learning algorithm that generates data by pitting two neural networks against each other.
[0696] "Methods of converting into a different style" refers to techniques or processes used to convert the original illustration into a specific style (e.g., cartoon style, live-action style, painterly style, etc.).
[0697] "Means for the user to edit the illustration" refers to an interface or tool that allows the user to add text or decorations to the illustration.
[0698] The present invention is a system that allows users to easily create their own original stamps, and in particular provides a function that recognizes the user's emotions by combining an emotion engine and generates stamps based on those emotions. This system consists of the following main steps:
[0699] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[0700] The server uses an AI module to instruct the user on scenes (e.g., emotions) that are suitable for the stamps. These instructions are displayed on the device screen. The user then uses the device's camera function to shoot a video based on the instructed scenes. The user manually starts and stops recording.
[0701] Once the video has been filmed and uploaded, the server uses an emotion engine to analyze the user's emotions in the video. Existing emotion engines, such as Google Cloud AI, AWS Rekognition, and Microsoft Azure Emotion API, can be used. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[0702] Video data is uploaded from the device to the server. The server selects specific scenes based on the results of the emotion engine. It selects scenes that most clearly express specific emotions (e.g., joy, surprise, sadness). The server analyzes the emotion recognition results and determines the most suitable scene. It then extracts specific frames from the selected scene.
[0703] Once the selected scene is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs) to generate the original illustration, which is then converted into the user's chosen style (e.g., cartoon, live action, painterly).
[0704] The converted stamp image is sent to the device and a preview is displayed to the user. The user can then add text and decorations (e.g., emojis, stickers, custom text, etc.) to the stamp. Once the user has finished editing, the device sends the final stamp pack to the server. The server saves the stamp pack and generates a download link for the user.
[0705] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[0706] Specific examples
[0707] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[0708] Example input to a generative AI model
[0709] An example of a prompt is:
[0710] "Take a video that expresses the emotion of joy."
[0711] "Please choose the scene in the video where you can see your smile most clearly."
[0712] "Take frames extracted from that scene and convert them into a cartoon-like style."
[0713] "Add custom text and emojis to the generated stamps."
[0714] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[0715] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0716] Step 1: Initialization and User Authentication
[0717] When a user launches an application, the server requests the user's authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information against a database. The input is the username and password, and the output is the authentication result. If authentication is successful, the user proceeds to the next step.
[0718] Specific behavior:
[0719] The user launches the app and enters their username and password on the authentication screen.
[0720] The authentication information is sent to the server, which then performs authentication by referencing a database.
[0721] The user is notified whether the authentication was successful or not.
[0722] Step 2: Propose a video shooting scene
[0723] The server uses an AI module to suggest scenes suitable for stamps (e.g., emotions). The suggested scenes are displayed on the device screen. The input is a request for a scene to be captured, and the output is a list of scenes generated by the AI.
[0724] Specific behavior:
[0725] The server sends a request for scene suggestions to the AI module.
[0726] The AI module generates multiple scenes, and a list of these is provided to the device via the server.
[0727] The proposed scene is displayed to the user.
[0728] Step 3: Record a video
[0729] The user shoots a video using the device's camera function based on the selected scene. The input is the user's scene selection, and the output is the shot video. The user manually starts and stops shooting.
[0730] Specific behavior:
[0731] The user selects one of the suggested scenes.
[0732] Launch the camera app and tap the record button to start recording.
[0733] When you're done recording, tap the stop button.
[0734] Step 4: Upload video and recognize emotions
[0735] The video is uploaded from the device to the server, where the server uses an emotion engine to analyze the user's emotions in the video. The input is the video, and the output is the result of the emotion analysis.
[0736] Specific behavior:
[0737] The device uploads the captured video file to the server.
[0738] The server passes the video to the emotion engine and requests analysis.
[0739] The emotion recognition results are sent back to the server.
[0740] Step 5: Scene selection and frame extraction
[0741] The server selects the scene that most clearly expresses emotion based on the results of the emotion engine, and then extracts the optimal frame from that scene. The input is the emotion analysis result, and the output is the selected frame.
[0742] Specific behavior:
[0743] The server analyzes the emotion recognition results and determines the most appropriate scene.
[0744] Extract a specific frame from the selected scene.
[0745] Step 6: Image processing and style transfer
[0746] Once the selected frame is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs). The input is the extracted frame, and the output is the illustration with the converted style.
[0747] Specific behavior:
[0748] The server generates the original illustration from the extracted frames.
[0749] The illustration is passed through a style conversion algorithm to convert it to the specified style.
[0750] Step 7: Preview and edit your stamp
[0751] The converted stamp image is sent to the device and a preview is displayed to the user. The user can add text and decorations to the stamp on this preview screen. The input is the style-converted illustration, and the output is the final stamp image edited by the user.
[0752] Specific behavior:
[0753] The server transmits the converted stamp image to the terminal.
[0754] The user uses the editing tools to add text and decorations.
[0755] Step 8: Create and save your sticker pack
[0756] When the user finishes editing, the device sends the final stamp pack to the server, which saves the stamp pack, generates a download link, and notifies the user. The input is the final stamp image, and the output is the download link.
[0757] Specific behavior:
[0758] The user taps the "Done" button to send the stamp pack to the server.
[0759] The server stores the stamp pack and generates a download link.
[0760] The download link will be sent to the user.
[0761] Step 9: Import to LINE
[0762] Users download the stamp pack via the provided link. The stamp pack is imported into the LINE app, allowing users to use their custom stamps on LINE. The input is the download link, and the output is the stamps in the LINE app.
[0763] Specific behavior:
[0764] The user clicks on the download link to get the stamp pack.
[0765] Import the stamp pack into the LINE app.
[0766] (Application example 2)
[0767] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0768] Conventional stamp generation systems have had the problem of making it difficult for users to easily create original stamps that reflect their own emotions. Furthermore, the technology to easily change the style of stamps and make them uploadable to virtual stores was insufficient. Furthermore, it was difficult to utilize an emotion engine to automatically select scenes that best express emotions.
[0769] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for requesting user authentication information and verifying the authentication information, means for presenting instructions for generating images to the user, means for receiving a video captured by the user based on the instructions, means for selecting a specific scene from the video and recognizing the user's emotion using an emotion engine, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style and using a generative adversarial network, means for the user to edit the illustration, and means for saving the user-edited stamp pack on the server and generating a download link. This allows for easy and effective generation of original stamps that reflect the user's emotion and uploading the style-converted stamps to a virtual store, enabling rich communication.
[0770] "Authentication information" refers to information such as a username and password that a user uses when accessing a system.
[0771] A "server" is a computer system that receives requests from users and performs the processing in response to them.
[0772] "Instructions" are information that prompts the system to perform specific operations or actions on the user.
[0773] A "video" is a media file that contains a series of image frames captured by a user.
[0774] The "emotion engine" is a software module that analyzes and recognizes user emotions from video and audio.
[0775] A "specific scene" is a selected part of a video that most clearly expresses an emotion or a specific event.
[0776] An "original illustration" is an image generated based on a specific scene extracted from a video.
[0777] "Converting to a style" refers to the process of changing the look and feel of the original illustration into a different style (e.g., cartoon, live-action, painterly, etc.).
[0778] A "generative adversarial network (GAN)" is a machine learning algorithm in which one network generates images and another network evaluates the generated images.
[0779] "Editing" refers to customization work that the user performs on the generated illustration (e.g., adding text, inserting decorations, etc.).
[0780] A "stamp pack" is a package containing multiple stamp images.
[0781] A "download link" is a URL that allows a user to download a file over the Internet.
[0782] A "virtual store" is a digital shopping platform that exists on the Internet.
[0783] System Configuration
[0784] This invention is a system that allows users to easily create original stamps that reflect their emotions, convert the stamps into different styles, and make them available in virtual stores. This system is composed of a server, terminals, and related software modules.
[0785] Processing flow
[0786] User Authentication
[0787] When a user launches an application, the server requests the user's authentication information. The user enters their username and password into the terminal and sends them to the server. The server verifies the authentication information, and if authentication is successful, the user can proceed to the next step.
[0788] Directing and filming video scenes
[0789] The server provides instructions to the device for generating images. Specifically, it works in conjunction with the emotion engine to present scenes that the user should capture (e.g., smile, surprise, sadness) to the device. The user then uses the device's camera function to capture these scenes as videos and upload them to the server.
[0790] Emotion Recognition and Scene Selection
[0791] The server receives videos uploaded by users and analyzes the user's emotions in the videos using an emotion engine. Based on this emotion analysis, a specific scene is selected from multiple scenes. The selection criterion is the scene that most clearly expresses the emotion.
[0792] Style conversion and stamp generation
[0793] The server generates a source illustration based on the specific scene selected, then converts the generated source illustration into different styles (e.g., cartoon, live action, painterly) using machine learning algorithms such as generative adversarial networks (GANs).
[0794] Editing and saving stamps
[0795] The converted stamp images are sent to the device, where the user can customize them by adding text, emojis, stickers, etc. Once editing is complete, the device sends the final stamp pack to the server, which stores it and provides a download link to the user.
[0796] Hardware and software used
[0797] Hardware: Cameras on smartphones and head-mounted displays (HMDs)
[0798] software:
[0799] EmotionRecognizer (emotion recognition engine)
[0800] StyleTransfer (style transfer algorithm, GAN)
[0801] OpenCV (camera operation and image processing library)
[0802] Requests (HTTP client library for sending API requests)
[0803] Specific examples
[0804] For example, suppose a user makes a surprised expression while virtual shopping. This moment is captured by the camera, and EmotionRecognizer recognizes it as "surprise." StyleTransfer then converts the expression into a cartoon-style image based on this emotion. Finally, the generated sticker is saved in the user's account via the virtual store's API.
[0805] Prompt Sentence Examples
[0806] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[0807] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0808] Step 1: Enter and verify your credentials
[0809] Subject: User and Server
[0810] Specific operation: The user launches an application and enters authentication information (username and password) into the device. The device sends this information to the server. The server verifies the received authentication information, and if authentication is successful, proceeds to the next step.
[0811] Input: Username and Password
[0812] Output: Authentication success flag
[0813] Step 2: Video Shooting Scene Instructions
[0814] Subject: Server and Terminal
[0815] Specific operation: The server uses the emotion engine to present the scene to be captured (e.g., smile, surprise, sadness) on the device screen. The user acts out the scene according to the instructions.
[0816] Input: Authentication success flag
[0817] Output: Instructions for the shooting scene
[0818] Step 3: Record your video
[0819] Subject: User and Device
[0820] Specific operation: The user uses the device's camera function to shoot a video based on the specified scene. The user starts and stops shooting.
[0821] Input: Instructions for the shooting scene
[0822] Output: Recorded video file
[0823] Step 4: Upload your video
[0824] Subject: Terminal and Server
[0825] Specific operation: The device uploads the captured video file to the server, which receives and stores the video file.
[0826] Input: Recorded video file
[0827] Output: Video data stored on the server
[0828] Step 5: Emotion recognition and specific scene selection
[0829] Subject: Server
[0830] Specific operation: The server analyzes the received video using the emotion engine to recognize the user's emotion in the video, and then selects the specific scene that most clearly expresses the emotion from multiple scenes.
[0831] Input: Video data stored on the server
[0832] Output: Specific scene information
[0833] Step 6: Generate the original illustration
[0834] Subject: Server
[0835] Specific operation: The server generates a source illustration based on the selected specific scene. The source illustration is an image extracted from a video frame.
[0836] Input: Specific scene information
[0837] Output: Original illustration image
[0838] Step 7: Style Transformation
[0839] Subject: Server
[0840] Specific operation: The server converts the generated original illustration into different styles (e.g., cartoon, live-action, painting) using a generative adversarial network (GAN).
[0841] Input: Original illustration image
[0842] Output: Style-converted illustration
[0843] Step 8: Editing the stamp
[0844] Subject: User and Device
[0845] Specific operation: The converted stamp image is sent to the device, and the user can edit it. Users can add text, emojis, stickers, etc.
[0846] Input: Style-converted illustration
[0847] Output: Edited stamp image
[0848] Step 9: Save and share your sticker pack
[0849] Subject: Terminal and Server
[0850] Specific operation: The device sends the edited stamp pack to the server, which stores it and provides a download link to the user.
[0851] Input: Edited stamp image
[0852] Output: Saved sticker pack and download link
[0853] Prompt Sentence Examples
[0854] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[0855] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0856] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0857] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0858] [Third embodiment]
[0859] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0860] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0861] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0862] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0863] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0864] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0865] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0866] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0867] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0868] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0869] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0870] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0871] The present invention aims to provide a system that allows users to easily create their own original stamps. This system includes the steps of presenting instructions for image generation, allowing the user to shoot a video, selecting a specific scene from the video, generating an original illustration from the selected scene, converting the original illustration into a different style, and allowing the user to edit the illustration.
[0872] Program processing explanation
[0873] Initialization and User Authentication
[0874] When a user launches an application, the server prompts for authentication information. The user enters a username and password and submits them to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[0875] Video shooting scene suggestions
[0876] The server uses an AI module to send instructions to the device about scenes suitable for the stamps, such as specific facial expressions or specific poses, for the user to capture in the video.
[0877] Video recording
[0878] The user uses the camera function of the device to shoot a video based on instructions from the server. The video is set to be a few minutes long and ensures that the specified scenes are included.
[0879] Video upload and scene selection
[0880] Once the video is complete, the device automatically uploads the video data to the server, which then analyzes it and selects specific scenes. This process uses an AI algorithm to scan each frame in the video and find the best frame that corresponds to the recommended scene.
[0881] Image processing and style transfer
[0882] Once the selected scene is determined, the server generates an original illustration from that scene. The server then converts the generated original illustration into a different style. This conversion process uses machine learning algorithms such as generative adversarial networks (GANs) to convert the illustration into a style such as photorealistic, cartoonish, or pictorial.
[0883] Preview and edit stamps
[0884] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, such as emojis, stickers, and custom text.
[0885] Creating and saving stamp packs
[0886] Once the user has finished editing, the device will send the final stamp pack to the server, which will save it and issue a download link, through which the user can download the stamp pack and import it into LINE.
[0887] Specific examples
[0888] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to the stamp to create their own original stamp.
[0889] In this way, the present invention provides a method and system that allows easy creation of original stamps tailored to the needs of individual users.
[0890] The processing flow will be explained below.
[0891] Step 1:
[0892] The user launches the application, and the login screen appears.
[0893] Step 2:
[0894] The server prompts the user for authentication information: a login screen prompts for username and password.
[0895] Step 3:
[0896] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[0897] Step 4:
[0898] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[0899] Step 5:
[0900] The server uses an AI module to instruct the user on the scenes (e.g., emotions) that are suitable for the stamp. The instructions are displayed on the device screen.
[0901] Step 6:
[0902] The user uses the device's camera function to shoot a video based on the specified scene, and the user manually starts and stops video shooting.
[0903] Step 7:
[0904] Once the video recording is complete, the device will automatically upload the video data to the server, and the upload progress will be displayed on the screen.
[0905] Step 8:
[0906] The server receives the uploaded video data and uses AI algorithms to analyze each frame of the video, automatically selecting the best frame that corresponds to the recommended scene.
[0907] Step 9:
[0908] The server generates an original illustration based on the selected frame, and this original illustration is saved in the system.
[0909] Step 10:
[0910] The server converts the original illustration into a style selected by the user (e.g., photo-realistic, cartoon-like, or painted style). Style conversion is performed using machine learning algorithms such as generative adversarial networks (GANs).
[0911] Step 11:
[0912] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[0913] Step 12:
[0914] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[0915] Step 13:
[0916] The device sends the edited stamp pack to the server, which stores the received stamp pack.
[0917] Step 14:
[0918] Once the server has finished saving the stamp pack, it will generate a download link and notify the user of this link.
[0919] Step 15:
[0920] Users can download the stamp pack via the provided link and import it into the LINE app, allowing them to use their custom stamps on LINE.
[0921] Example 1
[0922] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0923] In conventional stamp creation systems, the process for users to create their own unique images is complicated and requires many operations, making it difficult to easily generate stamps. Furthermore, manual image editing and style conversion require specialized knowledge, making them difficult for general users to use. This has led to a demand for a new method and system that allows users to easily create original stamps that are optimal for each individual user.
[0924] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0925] In this invention, the server is a system that allows users to easily create their own original stamps, and includes: means for requesting authentication information from the user and authenticating them by referencing a database; means for generating instructions regarding scenes suitable for stamps using an AI model and transmitting the instructions to the terminal; means for receiving from the terminal a video shot by the user based on the instructions; means for analyzing the uploaded video and selecting a specific scene; means for generating an original illustration based on the specific scene and converting the generated original illustration into a different style; means for transmitting the converted illustration to the terminal and allowing the user to edit the illustration; and means for saving the edited stamp pack and issuing a download link. This allows users to easily create original stamps and use them in messaging apps without specialized knowledge.
[0926] "User authentication" is the process of verifying the identity of a user using an application and granting appropriate access permissions.
[0927] "AI Module" means a technological component based on artificial intelligence, used to perform a specific task or function.
[0928] A "prompt sentence" is a sentence that includes a message or instruction that intentionally prompts the user to take a specific action.
[0929] "Video data" refers to information saved in the file format of a video shot by a user.
[0930] "Scene selection" is the process of identifying specific frames or moments within a video and extracting important parts.
[0931] An "original illustration" is an initial image generated based on a scene extracted from a video.
[0932] "Style conversion" is an image processing process used to change the appearance and texture of the original illustration.
[0933] A "generative adversarial network (GAN)" is an algorithm consisting of two neural networks for the purpose of data generation. One creates fakes, and the other identifies them, improving accuracy.
[0934] "Preview" is an intermediate display function that allows the user to check the final output.
[0935] A "stamp pack" is a data format that compiles multiple stamp images into one set.
[0936] A "download link" is a URL that allows you to access and download data stored in online storage.
[0937] A "messaging app" is a software application that allows users to exchange text, images, videos, etc.
[0938] A "database" is a system for systematically storing and managing data.
[0939] A "session ID" is a unique identifier used to identify a communication session between a server and a client.
[0940] An "AI model for scene detection" is an artificial intelligence algorithm designed to identify and extract specific scenes within videos.
[0941] A "UI framework" is a set of software tools to assist in the design and implementation of user interfaces.
[0942] "Storage" means hardware or cloud services for permanent or temporary storage of data.
[0943] The present invention relates to a system that allows users to easily create their own original stamps. Specific embodiments of the system will be described below.
[0944] When a user launches an application, the server requests user authentication information. The server checks a database (e.g., MySQL) to verify the validity of the provided authentication information. If authentication is successful, the server generates a session ID and sends it to the device.
[0945] The server then uses an AI module (e.g., OpenAI's GPT-3) to send prompts to the device with instructions about scenes suitable for the stamp, such as "Please smile for the camera and hold the pose for three seconds."
[0946] The user activates the device's camera function and captures video based on instructions provided by the server. The device records the video using the camera module (e.g., AVFoundation in iOS). The captured video is several minutes long and ensures that the specified scenes are included.
[0947] Once the video recording is complete, the device automatically uploads the recorded video data to the server. The server saves the video data in storage (e.g., Amazon S3) and begins scene analysis. This analysis uses an AI model for scene detection (e.g., TensorFlow) to detect specific frames within the video.
[0948] The identified scene is generated as a source illustration by the server, and the source illustration is then transformed into a different style, such as photorealistic, cartoonish, or painted, using a generative adversarial network (GAN) algorithm (e.g., Pix2Pix).
[0949] The generated illustration is sent from the server to the device, where it is displayed on the device's preview screen, where the user can add text and decorations (e.g., emojis, stickers, custom text) to the stamp.
[0950] Once the user has finished editing, the device sends the final sticker pack data to the server, which saves the sticker pack in a database (e.g., PostgreSQL) and generates a download link. This link is then provided to the user, who can then download the sticker pack and import it into a messaging app (e.g., LINE).
[0951] As a concrete example, consider a scenario where a user captures a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). The user can then add appropriate text and decorations to create their own original stamps.
[0952] Examples of prompts for a generative AI model include:
[0953] "Give specific instructions for capturing a smiling scene. For example, smile for the camera and hold the pose for three seconds."
[0954] Provide instructions for capturing a scene showing a surprised expression. For example, hold your hand in front of your face as if something surprising has happened and hold the surprised expression for 3 seconds.
[0955] As described above, the present invention provides an innovative system that allows users to easily create original stamps and use them in messaging apps, even if they do not have specialized knowledge.
[0956] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0957] Step 1:
[0958] When a user launches an application, the server requests user authentication information. The terminal prompts the user to enter a username and password and sends that information to the server. The server then references a database to verify the authentication information. If authentication is successful, the server generates a session ID and sends it to the terminal.
[0959] Input: Username, Password
[0960] Data Processing: Credential Verification
[0961] Output: Session ID
[0962] Step 2:
[0963] The server uses an AI module to generate a prompt message for the scene that matches the stamp and send it to the device, specifically, a prompt message containing the instruction "Smile for the camera and hold the pose for three seconds."
[0964] Input: Session ID
[0965] Data processing: Prompt sentence generation by AI module
[0966] Output: prompt statement
[0967] Step 3:
[0968] The user activates the device's camera function and takes a video based on instructions provided by the server. The device uses the camera module to record the video and takes a few minutes of video based on specific instructions from the server.
[0969] Input: prompt statement
[0970] Data processing: Video shooting by users
[0971] Output: Video data
[0972] Step 4:
[0973] Once the video recording is complete, the device uploads the recorded video data to the server. The server receives the video data and stores it in storage. Scene analysis then begins.
[0974] Input: Video data
[0975] Data processing: Uploading video data and saving it to storage
[0976] Output: Saved video data
[0977] Step 5:
[0978] The server analyzes the stored video data using an AI model for scene analysis to detect specific frames within the video. For example, it analyzes scenes where people are smiling and selects the most suitable frame. As a result, the specific scene is extracted.
[0979] Input: Saved video data
[0980] Data processing: Frame selection using AI model for scene analysis
[0981] Output: Specific Scene
[0982] Step 6:
[0983] The server generates a source illustration using a generative adversarial network (GAN) based on the identified scene, and then proceeds to convert the generated source illustration into a different style.
[0984] Input: A specific scene
[0985] Data processing: Generating original illustrations using GAN
[0986] Output: Original illustration
[0987] Step 7:
[0988] The server converts the generated original illustration into the specified style (e.g., photo-like, cartoon-like, or painting-like), and the converted illustration is sent to the device.
[0989] Input: Original illustration
[0990] Data processing: Transformation using style transfer algorithms
[0991] Output: Converted illustration
[0992] Step 8:
[0993] The converted illustration sent to the device is displayed as a preview to the user. The user can add text and decorations to the stamp on the preview screen. This editing is done using the UI framework.
[0994] Input: Converted illustration
[0995] Data processing: Editing by the user
[0996] Output: Edited stamp
[0997] Step 9:
[0998] Once the user has finished editing, the device sends the final stamp pack data to the server, which stores the stamp pack in a database and generates a download link that can be provided to the user, who can then download the stamp pack and import it into their messaging app.
[0999] Input:Edited stamp
[1000] Data processing: Saving stamp pack data and creating links
[1001] Output: Download link
[1002] (Application example 1)
[1003] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1004] In conventional quality control systems, determining product anomalies detected on the production line mainly relies on human visual inspection and manual recording, which is inefficient and has a high risk of false positives. Furthermore, it is difficult to record the detection results and immediately link them to the quality control system, resulting in a waste of resources. To solve these problems, automated stamp generation and rapid linkage with the quality control system are needed.
[1005] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1006] In this invention, the server includes means for presenting instructions to a user for generating an image, means for receiving a video shot by the user based on the instructions, means for selecting a specific scene from the video, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style, means for the user to edit the illustration, means for automatically generating a stamp based on the condition of parts or products detected on the production line, and means for linking the stamp to a quality control system. This allows a quality check stamp to be automatically generated upon detection of a product anomaly, enabling rapid linkage with the quality control system.
[1007] The "means for presenting instructions to the user for generating an image" is a process for displaying specific instructions for capturing an appropriate scene when the user shoots a video.
[1008] The "means for receiving the video shot by the user based on the instruction" is a process for transferring the video shot by the user to the server.
[1009] The "means for selecting a specific scene from the video" is a process for identifying and extracting a scene suitable for quality check from the received video data.
[1010] The "means for generating an original illustration based on the specified scene" is a process for generating a base image from the specified scene.
[1011] The "means for converting the original illustration into a different style" refers to a process for converting the generated original illustration into a different style such as a cartoon style, a live-action style, or a painting style.
[1012] The "means for the user to edit the illustration" is a process by which the user can add text and decorations to the generated illustration.
[1013] "Means for automatically generating stamps based on the condition of parts or products detected on the production line" refers to a process for automatically generating stamps indicating quality conditions based on the detection results on the production line.
[1014] The "means for linking the stamp to the quality control system" is a process for linking the created stamp to the quality control system.
[1015] This invention is implemented in a system for automating quality control in factories. The main hardware of the system includes a production line robot and a camera, and the software uses OpenCV and TensorFlow. Specifically, the invention is implemented in the following stages:
[1016] First, the server presents the user with instructions for taking a video, including specific examples of scenes that are easy to detect (e.g., scratches on the surface, dirt, abnormal shapes, etc.). Once the user takes a video based on the instructions, the video is uploaded to the server.
[1017] The server analyzes the uploaded video and selects a specific scene. Using OpenCV and TensorFlow, it scans each frame in the video and finds the best frame suitable for quality check. From this identified scene, the original illustration is generated.
[1018] The server then converts the generated original illustration into a different style, such as cartoon, photorealistic, or painterly, using a generative adversarial network (GAN) algorithm.
[1019] Users can preview the converted illustration and edit it by adding text and decorations on it, including specific labels and comments to indicate its quality status.
[1020] Finally, the server generates the edited stamp and automatically connects it to the quality control system, allowing the product quality status to be systematically recorded and quickly addressed.
[1021] For example, if abnormal soldering is detected in a factory that manufactures electronic circuit boards, the server analyzes the video and identifies the abnormal area. Next, a stamp is generated based on the scene, and text such as "defective" or "re-inspection" is added, which is then linked to the quality control system. This allows information about the abnormal area to be shared with engineers in real time.
[1022] An example of a prompt is:
[1023] Prompt statement:
[1024] Design a system to detect product anomalies to automate quality checks on electronic circuit board manufacturing lines. The system analyzes captured video, detects anomalies, and generates original stamps highlighting the anomalies. Specifically, the system detects abnormal solder joints and selects the best frame from the scene. Based on the selected frame, a generative adversarial network (GAN) is used to create a stamp, which can then be integrated into the quality control system by an operator adding text.
[1025] This allows for effective and efficient quality control.
[1026] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1027] Step 1:
[1028] The server presents the user with instructions for generating a video. Specifically, guidelines are displayed for scenes where product anomalies are likely to be detected (e.g., scratches on the surface, dirt, abnormal shapes, etc.). The user then shoots a video based on these guidelines. The input includes the scene guidelines, and the output is a state in which the user is ready to shoot.
[1029] Step 2:
[1030] The user shoots a video based on instructions from the server. The user uses a camera to shoot the specified scene and generate a video file. This video file becomes the input of the system. The output is the shot video file.
[1031] Step 3:
[1032] The user's device uploads the video they have taken to the server. The device receives the video file and sends it to the server over the network. In this process, the video file is the input and the video data stored on the server is the output.
[1033] Step 4:
[1034] The server analyzes the received video data and selects specific scenes. Specifically, it uses AI modules for video analysis, such as OpenCV and TensorFlow, to scan each frame in the video and find the best frame suitable for quality check. The input is the video data uploaded to the server, and the output is the frame corresponding to the selected specific scene.
[1035] Step 5:
[1036] The server generates a source illustration based on a specific scene. Here, the selected frame image is processed and generated as the source illustration. The generated illustration is the basis for quality checks. The input is a frame image of a specific scene, and the output is the source illustration.
[1037] Step 6:
[1038] The server converts the original illustration into a different style using a generative adversarial network (GAN). The input is the original illustration, and the output is the style-converted illustration.
[1039] Step 7:
[1040] The user edits the converted illustration. The server presents an editing interface to the user, who uses this interface to add text and decorations to the illustration. The input is the style-converted illustration, and the output is the final illustration after user editing.
[1041] Step 8:
[1042] The server generates a stamp based on the edited illustration and links it to the quality control system. The generated stamp is automatically sent to the quality control system, where the quality status of the product is recorded. The input is the final illustration edited by the user, and the output is a stamp linked to the quality control system.
[1043] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1044] The present invention is a system that allows users to easily create their own original stamps. In particular, it combines an emotion engine to recognize the user's emotions and generate stamps based on those emotions. This system consists of the following main steps:
[1045] Program processing explanation
[1046] Initialization and User Authentication
[1047] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[1048] Video shooting scene suggestions
[1049] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[1050] Video recording
[1051] The user uses the device's camera function to shoot video based on the specified scene, and the user manually starts and stops shooting.
[1052] emotion recognition
[1053] Once the video is uploaded, the server uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[1054] Video upload and scene selection
[1055] Video data is uploaded from the device to the server, which then selects specific scenes based on the results of the emotion engine, selecting scenes that most clearly express specific emotions (e.g., joy, surprise, sadness).
[1056] Image processing and style transfer
[1057] Once the selected scene is determined, the server generates a source illustration from that scene, which is then converted into the user's chosen style (e.g., photorealistic, cartoonish, or painted). The conversion process uses machine learning algorithms such as generative adversarial networks (GANs).
[1058] Preview and edit stamps
[1059] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, including emojis, stickers, and custom text.
[1060] Creating and saving stamp packs
[1061] Once the user has completed the editing process, the device sends the final stamp pack to the server, which saves it and generates a download link, which is then sent to the user.
[1062] Import to LINE
[1063] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[1064] Specific examples
[1065] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[1066] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[1067] The processing flow will be explained below.
[1068] Step 1:
[1069] The user launches the application, and the login screen appears.
[1070] Step 2:
[1071] The server prompts the user for authentication information: a login screen prompts for username and password.
[1072] Step 3:
[1073] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[1074] Step 4:
[1075] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[1076] Step 5:
[1077] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[1078] Step 6:
[1079] The user uses the device's camera to shoot a video based on the specified scene. The user presses the recording button to start recording. The video is set to be a few minutes long and includes facial expressions of joy, anger, sadness, and happiness, as well as other recommended scenes.
[1080] Step 7:
[1081] When the user presses the end button, the video recording is complete. The device automatically saves the video data and prepares it for uploading to the server.
[1082] Step 8:
[1083] The device uploads the video data to the server, where the video file is compressed and converted into the appropriate format.
[1084] Step 9:
[1085] The server receives the uploaded video data and uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions, voice, and movements to identify each emotion: joy, anger, sadness, and happiness.
[1086] Step 10:
[1087] The server selects specific scenes based on the analysis results of the emotion engine, for example, selecting frames that show strong emotions, such as when the user is laughing the most or looking the most surprised.
[1088] Step 11:
[1089] The server generates an original illustration from the selected scene, which is temporarily stored in the system.
[1090] Step 12:
[1091] The server converts the original illustration into the style selected by the user. For example, the process of converting from a realistic style to a cartoon style or a painting style uses machine learning algorithms such as generative adversarial networks (GANs).
[1092] Step 13:
[1093] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[1094] Step 14:
[1095] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[1096] Step 15:
[1097] The device sends the edited stamp pack to the server, which stores the received stamp pack and generates a download link.
[1098] Step 16:
[1099] The server then sends the generated download link to the user, who then downloads the stamp pack and imports it into the LINE app.
[1100] Step 17:
[1101] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[1102] Example 2
[1103] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1104] In conventional technologies, the process for users to create their own original stamps is complicated, and creating stamps that reflect emotions is particularly difficult. Another issue is that the processes for style conversion and emotion recognition are complex, requiring a high level of technical knowledge for the average user.
[1105] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server provides a means for analyzing a user's emotions using an emotion recognition engine, a means for suggesting video scenes to present to the user, a means for receiving a video shot by the user based on the suggestions, a means for selecting a specific scene from the video based on the emotion, a means for generating an original illustration based on the specific scene, a means for converting the original illustration into a different style using a generative adversarial network, and a means for the user to edit the illustration. This allows the user to easily create original stamps that reflect their emotions.
[1106] An "emotion recognition engine" is a software or hardware system that analyzes a user's facial expressions, tone of voice, etc., and identifies emotions.
[1107] The "means for suggesting video shooting scenes" is a function that presents scenes containing specific emotions to the user and gives instructions for the user to shoot a video based on those scenes.
[1108] "Means for receiving video" refers to a mechanism for transmitting and receiving video data shot by a user to a server or system.
[1109] "Means for selecting specific scenes" refers to an algorithm or process that automatically selects scenes from a user's video that most clearly express a specific emotion.
[1110] The "means for generating the original illustration" is the technique or process used to create an illustration from the selected scene.
[1111] A "generative adversarial network" is a type of machine learning algorithm that generates data by pitting two neural networks against each other.
[1112] "Methods of converting into a different style" refers to techniques or processes used to convert the original illustration into a specific style (e.g., cartoon style, live-action style, painterly style, etc.).
[1113] "Means for the user to edit the illustration" refers to an interface or tool that allows the user to add text or decorations to the illustration.
[1114] The present invention is a system that allows users to easily create their own original stamps, and in particular provides a function that recognizes the user's emotions by combining an emotion engine and generates stamps based on those emotions. This system consists of the following main steps:
[1115] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[1116] The server uses an AI module to instruct the user on scenes (e.g., emotions) that are suitable for the stamps. These instructions are displayed on the device screen. The user then uses the device's camera function to shoot a video based on the instructed scenes. The user manually starts and stops recording.
[1117] Once the video has been filmed and uploaded, the server uses an emotion engine to analyze the user's emotions in the video. Existing emotion engines, such as Google Cloud AI, AWS Rekognition, and Microsoft Azure Emotion API, can be used. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[1118] Video data is uploaded from the device to the server. The server selects specific scenes based on the results of the emotion engine. It selects scenes that most clearly express specific emotions (e.g., joy, surprise, sadness). The server analyzes the emotion recognition results and determines the most suitable scene. It then extracts specific frames from the selected scene.
[1119] Once the selected scene is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs) to generate the original illustration, which is then converted into the user's chosen style (e.g., cartoon, live action, painterly).
[1120] The converted stamp image is sent to the device and a preview is displayed to the user. The user can then add text and decorations (e.g., emojis, stickers, custom text, etc.) to the stamp. Once the user has finished editing, the device sends the final stamp pack to the server. The server saves the stamp pack and generates a download link for the user.
[1121] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[1122] Specific examples
[1123] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[1124] Example input to a generative AI model
[1125] An example of a prompt is:
[1126] "Take a video that expresses the emotion of joy."
[1127] "Please choose the scene in the video where you can see your smile most clearly."
[1128] "Take frames extracted from that scene and convert them into a cartoon-like style."
[1129] "Add custom text and emojis to the generated stamps."
[1130] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[1131] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1132] Step 1: Initialization and User Authentication
[1133] When a user launches an application, the server requests the user's authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information against a database. The input is the username and password, and the output is the authentication result. If authentication is successful, the user proceeds to the next step.
[1134] Specific behavior:
[1135] The user launches the app and enters their username and password on the authentication screen.
[1136] The authentication information is sent to the server, which then performs authentication by referencing a database.
[1137] The user is notified whether the authentication was successful or not.
[1138] Step 2: Propose a video shooting scene
[1139] The server uses an AI module to suggest scenes suitable for stamps (e.g., emotions). The suggested scenes are displayed on the device screen. The input is a request for a scene to be captured, and the output is a list of scenes generated by the AI.
[1140] Specific behavior:
[1141] The server sends a request for scene suggestions to the AI module.
[1142] The AI module generates multiple scenes, and a list of these is provided to the device via the server.
[1143] The proposed scene is displayed to the user.
[1144] Step 3: Record a video
[1145] The user shoots a video using the device's camera function based on the selected scene. The input is the user's scene selection, and the output is the shot video. The user manually starts and stops shooting.
[1146] Specific behavior:
[1147] The user selects one of the suggested scenes.
[1148] Launch the camera app and tap the record button to start recording.
[1149] When you're done recording, tap the stop button.
[1150] Step 4: Upload video and recognize emotions
[1151] The video is uploaded from the device to the server, where the server uses an emotion engine to analyze the user's emotions in the video. The input is the video, and the output is the result of the emotion analysis.
[1152] Specific behavior:
[1153] The device uploads the captured video file to the server.
[1154] The server passes the video to the emotion engine and requests analysis.
[1155] The emotion recognition results are sent back to the server.
[1156] Step 5: Scene selection and frame extraction
[1157] The server selects the scene that most clearly expresses emotion based on the results of the emotion engine, and then extracts the optimal frame from that scene. The input is the emotion analysis result, and the output is the selected frame.
[1158] Specific behavior:
[1159] The server analyzes the emotion recognition results and determines the most appropriate scene.
[1160] Extract a specific frame from the selected scene.
[1161] Step 6: Image processing and style transfer
[1162] Once the selected frame is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs). The input is the extracted frame, and the output is the illustration with the converted style.
[1163] Specific behavior:
[1164] The server generates the original illustration from the extracted frames.
[1165] The illustration is passed through a style conversion algorithm to convert it to the specified style.
[1166] Step 7: Preview and edit your stamp
[1167] The converted stamp image is sent to the device and a preview is displayed to the user. The user can add text and decorations to the stamp on this preview screen. The input is the style-converted illustration, and the output is the final stamp image edited by the user.
[1168] Specific behavior:
[1169] The server transmits the converted stamp image to the terminal.
[1170] The user uses the editing tools to add text and decorations.
[1171] Step 8: Create and save your sticker pack
[1172] When the user finishes editing, the device sends the final stamp pack to the server, which saves the stamp pack, generates a download link, and notifies the user. The input is the final stamp image, and the output is the download link.
[1173] Specific behavior:
[1174] The user taps the "Done" button to send the stamp pack to the server.
[1175] The server stores the stamp pack and generates a download link.
[1176] The download link will be sent to the user.
[1177] Step 9: Import to LINE
[1178] Users download the stamp pack via the provided link. The stamp pack is imported into the LINE app, allowing users to use their custom stamps on LINE. The input is the download link, and the output is the stamps in the LINE app.
[1179] Specific behavior:
[1180] The user clicks on the download link to get the stamp pack.
[1181] Import the stamp pack into the LINE app.
[1182] (Application example 2)
[1183] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1184] Conventional stamp generation systems have had the problem of making it difficult for users to easily create original stamps that reflect their own emotions. Furthermore, the technology to easily change the style of stamps and make them uploadable to virtual stores was insufficient. Furthermore, it was difficult to utilize an emotion engine to automatically select scenes that best express emotions.
[1185] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for requesting user authentication information and verifying the authentication information, means for presenting instructions for generating images to the user, means for receiving a video captured by the user based on the instructions, means for selecting a specific scene from the video and recognizing the user's emotion using an emotion engine, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style and using a generative adversarial network, means for the user to edit the illustration, and means for saving the user-edited stamp pack on the server and generating a download link. This allows for easy and effective generation of original stamps that reflect the user's emotion and uploading the style-converted stamps to a virtual store, enabling rich communication.
[1186] "Authentication information" refers to information such as a username and password that a user uses when accessing a system.
[1187] A "server" is a computer system that receives requests from users and performs the processing in response to them.
[1188] "Instructions" are information that prompts the system to perform specific operations or actions on the user.
[1189] A "video" is a media file that contains a series of image frames captured by a user.
[1190] The "emotion engine" is a software module that analyzes and recognizes user emotions from video and audio.
[1191] A "specific scene" is a selected part of a video that most clearly expresses an emotion or a specific event.
[1192] An "original illustration" is an image generated based on a specific scene extracted from a video.
[1193] "Converting to a style" refers to the process of changing the look and feel of the original illustration into a different style (e.g., cartoon, live-action, painterly, etc.).
[1194] A "generative adversarial network (GAN)" is a machine learning algorithm in which one network generates images and another network evaluates the generated images.
[1195] "Editing" refers to customization work that the user performs on the generated illustration (e.g., adding text, inserting decorations, etc.).
[1196] A "stamp pack" is a package containing multiple stamp images.
[1197] A "download link" is a URL that allows a user to download a file over the Internet.
[1198] A "virtual store" is a digital shopping platform that exists on the Internet.
[1199] System Configuration
[1200] This invention is a system that allows users to easily create original stamps that reflect their emotions, convert the stamps into different styles, and make them available in virtual stores. This system is composed of a server, terminals, and related software modules.
[1201] Processing flow
[1202] User Authentication
[1203] When a user launches an application, the server requests the user's authentication information. The user enters their username and password into the terminal and sends them to the server. The server verifies the authentication information, and if authentication is successful, the user can proceed to the next step.
[1204] Directing and filming video scenes
[1205] The server provides instructions to the device for generating images. Specifically, it works in conjunction with the emotion engine to present scenes that the user should capture (e.g., smile, surprise, sadness) to the device. The user then uses the device's camera function to capture these scenes as videos and upload them to the server.
[1206] Emotion Recognition and Scene Selection
[1207] The server receives videos uploaded by users and analyzes the user's emotions in the videos using an emotion engine. Based on this emotion analysis, a specific scene is selected from multiple scenes. The selection criterion is the scene that most clearly expresses the emotion.
[1208] Style conversion and stamp generation
[1209] The server generates a source illustration based on the specific scene selected, then converts the generated source illustration into different styles (e.g., cartoon, live action, painterly) using machine learning algorithms such as generative adversarial networks (GANs).
[1210] Editing and saving stamps
[1211] The converted stamp images are sent to the device, where the user can customize them by adding text, emojis, stickers, etc. Once editing is complete, the device sends the final stamp pack to the server, which stores it and provides a download link to the user.
[1212] Hardware and software used
[1213] Hardware: Cameras on smartphones and head-mounted displays (HMDs)
[1214] software:
[1215] EmotionRecognizer (emotion recognition engine)
[1216] StyleTransfer (style transfer algorithm, GAN)
[1217] OpenCV (camera operation and image processing library)
[1218] Requests (HTTP client library for sending API requests)
[1219] Specific examples
[1220] For example, suppose a user makes a surprised expression while virtual shopping. This moment is captured by the camera, and EmotionRecognizer recognizes it as "surprise." StyleTransfer then converts the expression into a cartoon-style image based on this emotion. Finally, the generated sticker is saved in the user's account via the virtual store's API.
[1221] Prompt Sentence Examples
[1222] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[1223] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1224] Step 1: Enter and verify your credentials
[1225] Subject: User and Server
[1226] Specific operation: The user launches an application and enters authentication information (username and password) into the device. The device sends this information to the server. The server verifies the received authentication information, and if authentication is successful, proceeds to the next step.
[1227] Input: Username and Password
[1228] Output: Authentication success flag
[1229] Step 2: Video Shooting Scene Instructions
[1230] Subject: Server and Terminal
[1231] Specific operation: The server uses the emotion engine to present the scene to be captured (e.g., smile, surprise, sadness) on the device screen. The user acts out the scene according to the instructions.
[1232] Input: Authentication success flag
[1233] Output: Instructions for the shooting scene
[1234] Step 3: Record your video
[1235] Subject: User and Device
[1236] Specific operation: The user uses the device's camera function to shoot a video based on the specified scene. The user starts and stops shooting.
[1237] Input: Instructions for the shooting scene
[1238] Output: Recorded video file
[1239] Step 4: Upload your video
[1240] Subject: Terminal and Server
[1241] Specific operation: The device uploads the captured video file to the server, which receives and stores the video file.
[1242] Input: Recorded video file
[1243] Output: Video data stored on the server
[1244] Step 5: Emotion recognition and specific scene selection
[1245] Subject: Server
[1246] Specific operation: The server analyzes the received video using the emotion engine to recognize the user's emotion in the video, and then selects the specific scene that most clearly expresses the emotion from multiple scenes.
[1247] Input: Video data stored on the server
[1248] Output: Specific scene information
[1249] Step 6: Generate the original illustration
[1250] Subject: Server
[1251] Specific operation: The server generates a source illustration based on the selected specific scene. The source illustration is an image extracted from a video frame.
[1252] Input: Specific scene information
[1253] Output: Original illustration image
[1254] Step 7: Style Transformation
[1255] Subject: Server
[1256] Specific operation: The server converts the generated original illustration into different styles (e.g., cartoon, live-action, painting) using a generative adversarial network (GAN).
[1257] Input: Original illustration image
[1258] Output: Style-converted illustration
[1259] Step 8: Editing the stamp
[1260] Subject: User and Device
[1261] Specific operation: The converted stamp image is sent to the device, and the user can edit it. Users can add text, emojis, stickers, etc.
[1262] Input: Style-converted illustration
[1263] Output: Edited stamp image
[1264] Step 9: Save and share your sticker pack
[1265] Subject: Terminal and Server
[1266] Specific operation: The device sends the edited stamp pack to the server, which stores it and provides a download link to the user.
[1267] Input: Edited stamp image
[1268] Output: Saved sticker pack and download link
[1269] Prompt Sentence Examples
[1270] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[1271] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1272] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1273] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1274] [Fourth embodiment]
[1275] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1276] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1277] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1278] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1279] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1280] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1281] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1282] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1283] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1284] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1285] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1286] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1287] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1288] The present invention aims to provide a system that allows users to easily create their own original stamps. This system includes the steps of presenting instructions for image generation, allowing the user to shoot a video, selecting a specific scene from the video, generating an original illustration from the selected scene, converting the original illustration into a different style, and allowing the user to edit the illustration.
[1289] Program processing explanation
[1290] Initialization and User Authentication
[1291] When a user launches an application, the server prompts for authentication information. The user enters a username and password and submits them to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[1292] Video shooting scene suggestions
[1293] The server uses an AI module to send instructions to the device about scenes suitable for the stamps, such as specific facial expressions or specific poses, for the user to capture in the video.
[1294] Video recording
[1295] The user uses the camera function of the device to shoot a video based on instructions from the server. The video is set to be a few minutes long and ensures that the specified scenes are included.
[1296] Video upload and scene selection
[1297] Once the video is complete, the device automatically uploads the video data to the server, which then analyzes it and selects specific scenes. This process uses an AI algorithm to scan each frame in the video and find the best frame that corresponds to the recommended scene.
[1298] Image processing and style transfer
[1299] Once the selected scene is determined, the server generates an original illustration from that scene. The server then converts the generated original illustration into a different style. This conversion process uses machine learning algorithms such as generative adversarial networks (GANs) to convert the illustration into a style such as photorealistic, cartoonish, or pictorial.
[1300] Preview and edit stamps
[1301] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, such as emojis, stickers, and custom text.
[1302] Creating and saving stamp packs
[1303] Once the user has finished editing, the device will send the final stamp pack to the server, which will save it and issue a download link, through which the user can download the stamp pack and import it into LINE.
[1304] Specific examples
[1305] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to the stamp to create their own original stamp.
[1306] In this way, the present invention provides a method and system that allows easy creation of original stamps tailored to the needs of individual users.
[1307] The processing flow will be explained below.
[1308] Step 1:
[1309] The user launches the application, and the login screen appears.
[1310] Step 2:
[1311] The server prompts the user for authentication information: a login screen prompts for username and password.
[1312] Step 3:
[1313] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[1314] Step 4:
[1315] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[1316] Step 5:
[1317] The server uses an AI module to instruct the user on the scenes (e.g., emotions) that are suitable for the stamp. The instructions are displayed on the device screen.
[1318] Step 6:
[1319] The user uses the device's camera function to shoot a video based on the specified scene, and the user manually starts and stops video shooting.
[1320] Step 7:
[1321] Once the video recording is complete, the device will automatically upload the video data to the server, and the upload progress will be displayed on the screen.
[1322] Step 8:
[1323] The server receives the uploaded video data and uses AI algorithms to analyze each frame of the video, automatically selecting the best frame that corresponds to the recommended scene.
[1324] Step 9:
[1325] The server generates an original illustration based on the selected frame, and this original illustration is saved in the system.
[1326] Step 10:
[1327] The server converts the original illustration into a style selected by the user (e.g., photo-realistic, cartoon-like, or painted style). Style conversion is performed using machine learning algorithms such as generative adversarial networks (GANs).
[1328] Step 11:
[1329] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[1330] Step 12:
[1331] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[1332] Step 13:
[1333] The device sends the edited stamp pack to the server, which stores the received stamp pack.
[1334] Step 14:
[1335] Once the server has finished saving the stamp pack, it will generate a download link and notify the user of this link.
[1336] Step 15:
[1337] Users can download the stamp pack via the provided link and import it into the LINE app, allowing them to use their custom stamps on LINE.
[1338] Example 1
[1339] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1340] In conventional stamp creation systems, the process for users to create their own unique images is complicated and requires many operations, making it difficult to easily generate stamps. Furthermore, manual image editing and style conversion require specialized knowledge, making them difficult for general users to use. This has led to a demand for a new method and system that allows users to easily create original stamps that are optimal for each individual user.
[1341] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1342] In this invention, the server is a system that allows users to easily create their own original stamps, and includes: means for requesting authentication information from the user and authenticating them by referencing a database; means for generating instructions regarding scenes suitable for stamps using an AI model and transmitting the instructions to the terminal; means for receiving from the terminal a video shot by the user based on the instructions; means for analyzing the uploaded video and selecting a specific scene; means for generating an original illustration based on the specific scene and converting the generated original illustration into a different style; means for transmitting the converted illustration to the terminal and allowing the user to edit the illustration; and means for saving the edited stamp pack and issuing a download link. This allows users to easily create original stamps and use them in messaging apps without specialized knowledge.
[1343] "User authentication" is the process of verifying the identity of a user using an application and granting appropriate access permissions.
[1344] "AI Module" means a technological component based on artificial intelligence, used to perform a specific task or function.
[1345] A "prompt sentence" is a sentence that includes a message or instruction that intentionally prompts the user to take a specific action.
[1346] "Video data" refers to information saved in the file format of a video shot by a user.
[1347] "Scene selection" is the process of identifying specific frames or moments within a video and extracting important parts.
[1348] An "original illustration" is an initial image generated based on a scene extracted from a video.
[1349] "Style conversion" is an image processing process used to change the appearance and texture of the original illustration.
[1350] A "generative adversarial network (GAN)" is an algorithm consisting of two neural networks for the purpose of data generation. One creates fakes, and the other identifies them, improving accuracy.
[1351] "Preview" is an intermediate display function that allows the user to check the final output.
[1352] A "stamp pack" is a data format that compiles multiple stamp images into one set.
[1353] A "download link" is a URL that allows you to access and download data stored in online storage.
[1354] A "messaging app" is a software application that allows users to exchange text, images, videos, etc.
[1355] A "database" is a system for systematically storing and managing data.
[1356] A "session ID" is a unique identifier used to identify a communication session between a server and a client.
[1357] An "AI model for scene detection" is an artificial intelligence algorithm designed to identify and extract specific scenes within videos.
[1358] A "UI framework" is a set of software tools to assist in the design and implementation of user interfaces.
[1359] "Storage" means hardware or cloud services for permanent or temporary storage of data.
[1360] The present invention relates to a system that allows users to easily create their own original stamps. Specific embodiments of the system will be described below.
[1361] When a user launches an application, the server requests user authentication information. The server checks a database (e.g., MySQL) to verify the validity of the provided authentication information. If authentication is successful, the server generates a session ID and sends it to the device.
[1362] The server then uses an AI module (e.g., OpenAI's GPT-3) to send prompts to the device with instructions about scenes suitable for the stamp, such as "Please smile for the camera and hold the pose for three seconds."
[1363] The user activates the device's camera function and captures video based on instructions provided by the server. The device records the video using the camera module (e.g., AVFoundation in iOS). The captured video is several minutes long and ensures that the specified scenes are included.
[1364] Once the video recording is complete, the device automatically uploads the recorded video data to the server. The server saves the video data in storage (e.g., Amazon S3) and begins scene analysis. This analysis uses an AI model for scene detection (e.g., TensorFlow) to detect specific frames within the video.
[1365] The identified scene is generated as a source illustration by the server, and the source illustration is then transformed into a different style, such as photorealistic, cartoonish, or painted, using a generative adversarial network (GAN) algorithm (e.g., Pix2Pix).
[1366] The generated illustration is sent from the server to the device, where it is displayed on the device's preview screen, where the user can add text and decorations (e.g., emojis, stickers, custom text) to the stamp.
[1367] Once the user has finished editing, the device sends the final sticker pack data to the server, which saves the sticker pack in a database (e.g., PostgreSQL) and generates a download link. This link is then provided to the user, who can then download the sticker pack and import it into a messaging app (e.g., LINE).
[1368] As a concrete example, consider a scenario where a user captures a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server automatically selects scenes containing these expressions and extracts the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). The user can then add appropriate text and decorations to create their own original stamps.
[1369] Examples of prompts for a generative AI model include:
[1370] "Give specific instructions for capturing a smiling scene. For example, smile for the camera and hold the pose for three seconds."
[1371] Provide instructions for capturing a scene showing a surprised expression. For example, hold your hand in front of your face as if something surprising has happened and hold the surprised expression for 3 seconds.
[1372] As described above, the present invention provides an innovative system that allows users to easily create original stamps and use them in messaging apps, even if they do not have specialized knowledge.
[1373] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1374] Step 1:
[1375] When a user launches an application, the server requests user authentication information. The terminal prompts the user to enter a username and password and sends that information to the server. The server then references a database to verify the authentication information. If authentication is successful, the server generates a session ID and sends it to the terminal.
[1376] Input: Username, Password
[1377] Data Processing: Credential Verification
[1378] Output: Session ID
[1379] Step 2:
[1380] The server uses an AI module to generate a prompt message for the scene that matches the stamp and send it to the device, specifically, a prompt message containing the instruction "Smile for the camera and hold the pose for three seconds."
[1381] Input: Session ID
[1382] Data processing: Prompt sentence generation by AI module
[1383] Output: prompt statement
[1384] Step 3:
[1385] The user activates the device's camera function and takes a video based on instructions provided by the server. The device uses the camera module to record the video and takes a few minutes of video based on specific instructions from the server.
[1386] Input: prompt statement
[1387] Data processing: Video shooting by users
[1388] Output: Video data
[1389] Step 4:
[1390] Once the video recording is complete, the device uploads the recorded video data to the server. The server receives the video data and stores it in storage. Scene analysis then begins.
[1391] Input: Video data
[1392] Data processing: Uploading video data and saving it to storage
[1393] Output: Saved video data
[1394] Step 5:
[1395] The server analyzes the stored video data using an AI model for scene analysis to detect specific frames within the video. For example, it analyzes scenes where people are smiling and selects the most suitable frame. As a result, the specific scene is extracted.
[1396] Input: Saved video data
[1397] Data processing: Frame selection using AI model for scene analysis
[1398] Output: Specific Scene
[1399] Step 6:
[1400] The server generates a source illustration using a generative adversarial network (GAN) based on the identified scene, and then proceeds to convert the generated source illustration into a different style.
[1401] Input: A specific scene
[1402] Data processing: Generating original illustrations using GAN
[1403] Output: Original illustration
[1404] Step 7:
[1405] The server converts the generated original illustration into the specified style (e.g., photo-like, cartoon-like, or painting-like), and the converted illustration is sent to the device.
[1406] Input: Original illustration
[1407] Data processing: Transformation using style transfer algorithms
[1408] Output: Converted illustration
[1409] Step 8:
[1410] The converted illustration sent to the device is displayed as a preview to the user. The user can add text and decorations to the stamp on the preview screen. This editing is done using the UI framework.
[1411] Input: Converted illustration
[1412] Data processing: Editing by the user
[1413] Output: Edited stamp
[1414] Step 9:
[1415] Once the user has finished editing, the device sends the final stamp pack data to the server, which stores the stamp pack in a database and generates a download link that can be provided to the user, who can then download the stamp pack and import it into their messaging app.
[1416] Input:Edited stamp
[1417] Data processing: Saving stamp pack data and creating links
[1418] Output: Download link
[1419] (Application example 1)
[1420] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1421] In conventional quality control systems, determining product anomalies detected on the production line mainly relies on human visual inspection and manual recording, which is inefficient and has a high risk of false positives. Furthermore, it is difficult to record the detection results and immediately link them to the quality control system, resulting in a waste of resources. To solve these problems, automated stamp generation and rapid linkage with the quality control system are needed.
[1422] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1423] In this invention, the server includes means for presenting instructions to a user for generating an image, means for receiving a video shot by the user based on the instructions, means for selecting a specific scene from the video, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style, means for the user to edit the illustration, means for automatically generating a stamp based on the condition of parts or products detected on the production line, and means for linking the stamp to a quality control system. This allows a quality check stamp to be automatically generated upon detection of a product anomaly, enabling rapid linkage with the quality control system.
[1424] The "means for presenting instructions to the user for generating an image" is a process for displaying specific instructions for capturing an appropriate scene when the user shoots a video.
[1425] The "means for receiving the video shot by the user based on the instruction" is a process for transferring the video shot by the user to the server.
[1426] The "means for selecting a specific scene from the video" is a process for identifying and extracting a scene suitable for quality check from the received video data.
[1427] The "means for generating an original illustration based on the specified scene" is a process for generating a base image from the specified scene.
[1428] The "means for converting the original illustration into a different style" refers to a process for converting the generated original illustration into a different style such as a cartoon style, a live-action style, or a painting style.
[1429] The "means for the user to edit the illustration" is a process by which the user can add text and decorations to the generated illustration.
[1430] "Means for automatically generating stamps based on the condition of parts or products detected on the production line" refers to a process for automatically generating stamps indicating quality conditions based on the detection results on the production line.
[1431] The "means for linking the stamp to the quality control system" is a process for linking the created stamp to the quality control system.
[1432] This invention is implemented in a system for automating quality control in factories. The main hardware of the system includes a production line robot and a camera, and the software uses OpenCV and TensorFlow. Specifically, the invention is implemented in the following stages:
[1433] First, the server presents the user with instructions for taking a video, including specific examples of scenes that are easy to detect (e.g., scratches on the surface, dirt, abnormal shapes, etc.). Once the user takes a video based on the instructions, the video is uploaded to the server.
[1434] The server analyzes the uploaded video and selects a specific scene. Using OpenCV and TensorFlow, it scans each frame in the video and finds the best frame suitable for quality check. From this identified scene, the original illustration is generated.
[1435] The server then converts the generated original illustration into a different style, such as cartoon, photorealistic, or painterly, using a generative adversarial network (GAN) algorithm.
[1436] Users can preview the converted illustration and edit it by adding text and decorations on it, including specific labels and comments to indicate its quality status.
[1437] Finally, the server generates the edited stamp and automatically connects it to the quality control system, allowing the product quality status to be systematically recorded and quickly addressed.
[1438] For example, if abnormal soldering is detected in a factory that manufactures electronic circuit boards, the server analyzes the video and identifies the abnormal area. Next, a stamp is generated based on the scene, and text such as "defective" or "re-inspection" is added, which is then linked to the quality control system. This allows information about the abnormal area to be shared with engineers in real time.
[1439] An example of a prompt is:
[1440] Prompt statement:
[1441] Design a system to detect product anomalies to automate quality checks on electronic circuit board manufacturing lines. The system analyzes captured video, detects anomalies, and generates original stamps highlighting the anomalies. Specifically, the system detects abnormal solder joints and selects the best frame from the scene. Based on the selected frame, a generative adversarial network (GAN) is used to create a stamp, which can then be integrated into the quality control system by an operator adding text.
[1442] This allows for effective and efficient quality control.
[1443] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1444] Step 1:
[1445] The server presents the user with instructions for generating a video. Specifically, guidelines are displayed for scenes where product anomalies are likely to be detected (e.g., scratches on the surface, dirt, abnormal shapes, etc.). The user then shoots a video based on these guidelines. The input includes the scene guidelines, and the output is a state in which the user is ready to shoot.
[1446] Step 2:
[1447] The user shoots a video based on instructions from the server. The user uses a camera to shoot the specified scene and generate a video file. This video file becomes the input of the system. The output is the shot video file.
[1448] Step 3:
[1449] The user's device uploads the video they have taken to the server. The device receives the video file and sends it to the server over the network. In this process, the video file is the input and the video data stored on the server is the output.
[1450] Step 4:
[1451] The server analyzes the received video data and selects specific scenes. Specifically, it uses AI modules for video analysis, such as OpenCV and TensorFlow, to scan each frame in the video and find the best frame suitable for quality check. The input is the video data uploaded to the server, and the output is the frame corresponding to the selected specific scene.
[1452] Step 5:
[1453] The server generates a source illustration based on a specific scene. Here, the selected frame image is processed and generated as the source illustration. The generated illustration is the basis for quality checks. The input is a frame image of a specific scene, and the output is the source illustration.
[1454] Step 6:
[1455] The server converts the original illustration into a different style using a generative adversarial network (GAN). The input is the original illustration, and the output is the style-converted illustration.
[1456] Step 7:
[1457] The user edits the converted illustration. The server presents an editing interface to the user, who uses this interface to add text and decorations to the illustration. The input is the style-converted illustration, and the output is the final illustration after user editing.
[1458] Step 8:
[1459] The server generates a stamp based on the edited illustration and links it to the quality control system. The generated stamp is automatically sent to the quality control system, where the quality status of the product is recorded. The input is the final illustration edited by the user, and the output is a stamp linked to the quality control system.
[1460] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1461] The present invention is a system that allows users to easily create their own original stamps. In particular, it combines an emotion engine to recognize the user's emotions and generate stamps based on those emotions. This system consists of the following main steps:
[1462] Program processing explanation
[1463] Initialization and User Authentication
[1464] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[1465] Video shooting scene suggestions
[1466] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[1467] Video recording
[1468] The user uses the device's camera function to shoot video based on the specified scene, and the user manually starts and stops shooting.
[1469] emotion recognition
[1470] Once the video is uploaded, the server uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[1471] Video upload and scene selection
[1472] Video data is uploaded from the device to the server, which then selects specific scenes based on the results of the emotion engine, selecting scenes that most clearly express specific emotions (e.g., joy, surprise, sadness).
[1473] Image processing and style transfer
[1474] Once the selected scene is determined, the server generates a source illustration from that scene, which is then converted into the user's chosen style (e.g., photorealistic, cartoonish, or painted). The conversion process uses machine learning algorithms such as generative adversarial networks (GANs).
[1475] Preview and edit stamps
[1476] The converted stamp image is sent to the device and a preview is displayed to the user, where the user can add text and decorations to the stamp, including emojis, stickers, and custom text.
[1477] Creating and saving stamp packs
[1478] Once the user has completed the editing process, the device sends the final stamp pack to the server, which saves it and generates a download link, which is then sent to the user.
[1479] Import to LINE
[1480] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[1481] Specific examples
[1482] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[1483] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[1484] The processing flow will be explained below.
[1485] Step 1:
[1486] The user launches the application, and the login screen appears.
[1487] Step 2:
[1488] The server prompts the user for authentication information: a login screen prompts for username and password.
[1489] Step 3:
[1490] The user enters the authentication information and presses the send button, which sends the entered authentication information to the server.
[1491] Step 4:
[1492] The server checks the received authentication information against a database and performs authentication. If authentication is successful, the user proceeds to the next step. If authentication fails, an error message is displayed.
[1493] Step 5:
[1494] The server uses an AI module to instruct the user on the appropriate scenes (e.g., emotions) for the stamps. These instructions are displayed on the device screen.
[1495] Step 6:
[1496] The user uses the device's camera to shoot a video based on the specified scene. The user presses the recording button to start recording. The video is set to be a few minutes long and includes facial expressions of joy, anger, sadness, and happiness, as well as other recommended scenes.
[1497] Step 7:
[1498] When the user presses the end button, the video recording is complete. The device automatically saves the video data and prepares it for uploading to the server.
[1499] Step 8:
[1500] The device uploads the video data to the server, where the video file is compressed and converted into the appropriate format.
[1501] Step 9:
[1502] The server receives the uploaded video data and uses an emotion engine to analyze the user's emotions in the video. The emotion engine analyzes the user's facial expressions, voice, and movements to identify each emotion: joy, anger, sadness, and happiness.
[1503] Step 10:
[1504] The server selects specific scenes based on the analysis results of the emotion engine, for example, selecting frames that show strong emotions, such as when the user is laughing the most or looking the most surprised.
[1505] Step 11:
[1506] The server generates an original illustration from the selected scene, which is temporarily stored in the system.
[1507] Step 12:
[1508] The server converts the original illustration into the style selected by the user. For example, the process of converting from a realistic style to a cartoon style or a painting style uses machine learning algorithms such as generative adversarial networks (GANs).
[1509] Step 13:
[1510] The server sends the converted stamp image to the device, which displays the stamp image as a preview on the device screen.
[1511] Step 14:
[1512] Users can preview, add text and decorations (e.g. emojis, stickers, custom text) to the stamp, and then click the save button when they're done editing.
[1513] Step 15:
[1514] The device sends the edited stamp pack to the server, which stores the received stamp pack and generates a download link.
[1515] Step 16:
[1516] The server then sends the generated download link to the user, who then downloads the stamp pack and imports it into the LINE app.
[1517] Step 17:
[1518] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[1519] Example 2
[1520] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1521] In conventional technologies, the process for users to create their own original stamps is complicated, and creating stamps that reflect emotions is particularly difficult. Another issue is that the processes for style conversion and emotion recognition are complex, requiring a high level of technical knowledge for the average user.
[1522] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server provides a means for analyzing a user's emotions using an emotion recognition engine, a means for suggesting video scenes to present to the user, a means for receiving a video shot by the user based on the suggestions, a means for selecting a specific scene from the video based on the emotion, a means for generating an original illustration based on the specific scene, a means for converting the original illustration into a different style using a generative adversarial network, and a means for the user to edit the illustration. This allows the user to easily create original stamps that reflect their emotions.
[1523] An "emotion recognition engine" is a software or hardware system that analyzes a user's facial expressions, tone of voice, etc., and identifies emotions.
[1524] The "means for suggesting video shooting scenes" is a function that presents scenes containing specific emotions to the user and gives instructions for the user to shoot a video based on those scenes.
[1525] "Means for receiving video" refers to a mechanism for transmitting and receiving video data shot by a user to a server or system.
[1526] "Means for selecting specific scenes" refers to an algorithm or process that automatically selects scenes from a user's video that most clearly express a specific emotion.
[1527] The "means for generating the original illustration" is the technique or process used to create an illustration from the selected scene.
[1528] A "generative adversarial network" is a type of machine learning algorithm that generates data by pitting two neural networks against each other.
[1529] "Methods of converting into a different style" refers to techniques or processes used to convert the original illustration into a specific style (e.g., cartoon style, live-action style, painterly style, etc.).
[1530] "Means for the user to edit the illustration" refers to an interface or tool that allows the user to add text or decorations to the illustration.
[1531] The present invention is a system that allows users to easily create their own original stamps, and in particular provides a function that recognizes the user's emotions by combining an emotion engine and generates stamps based on those emotions. This system consists of the following main steps:
[1532] When a user launches an application, the server prompts the user for authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information and, if successful, proceeds to the next step.
[1533] The server uses an AI module to instruct the user on scenes (e.g., emotions) that are suitable for the stamps. These instructions are displayed on the device screen. The user then uses the device's camera function to shoot a video based on the instructed scenes. The user manually starts and stops recording.
[1534] Once the video has been filmed and uploaded, the server uses an emotion engine to analyze the user's emotions in the video. Existing emotion engines, such as Google Cloud AI, AWS Rekognition, and Microsoft Azure Emotion API, can be used. The emotion engine analyzes the user's facial expressions and tone of voice to recognize emotions.
[1535] Video data is uploaded from the device to the server. The server selects specific scenes based on the results of the emotion engine. It selects scenes that most clearly express specific emotions (e.g., joy, surprise, sadness). The server analyzes the emotion recognition results and determines the most suitable scene. It then extracts specific frames from the selected scene.
[1536] Once the selected scene is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs) to generate the original illustration, which is then converted into the user's chosen style (e.g., cartoon, live action, painterly).
[1537] The converted stamp image is sent to the device and a preview is displayed to the user. The user can then add text and decorations (e.g., emojis, stickers, custom text, etc.) to the stamp. Once the user has finished editing, the device sends the final stamp pack to the server. The server saves the stamp pack and generates a download link for the user.
[1538] Users download the stamp pack via the provided link, which is then imported into the LINE app, allowing users to use their custom stamps on LINE.
[1539] Specific examples
[1540] For example, a user can capture a scene containing a specific facial expression, such as "smile," "surprise," or "sadness." The server uses an emotion engine to automatically select scenes containing these expressions and extract the best frames from each scene. These frames are then generated as original illustrations and converted into the style selected by the user (e.g., manga style). Finally, the user can add appropriate text and decorations to create their own original stamp.
[1541] Example input to a generative AI model
[1542] An example of a prompt is:
[1543] "Take a video that expresses the emotion of joy."
[1544] "Please choose the scene in the video where you can see your smile most clearly."
[1545] "Take frames extracted from that scene and convert them into a cartoon-like style."
[1546] "Add custom text and emojis to the generated stamps."
[1547] In this way, the present invention provides a system that, by combining an emotion engine, enables the user to easily create original stamps that more faithfully reflect the emotions expressed by the user.
[1548] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1549] Step 1: Initialization and User Authentication
[1550] When a user launches an application, the server requests the user's authentication information. The user enters a username and password and sends the authentication information to the server. The server validates the authentication information against a database. The input is the username and password, and the output is the authentication result. If authentication is successful, the user proceeds to the next step.
[1551] Specific behavior:
[1552] The user launches the app and enters their username and password on the authentication screen.
[1553] The authentication information is sent to the server, which then performs authentication by referencing a database.
[1554] The user is notified whether the authentication was successful or not.
[1555] Step 2: Propose a video shooting scene
[1556] The server uses an AI module to suggest scenes suitable for stamps (e.g., emotions). The suggested scenes are displayed on the device screen. The input is a request for a scene to be captured, and the output is a list of scenes generated by the AI.
[1557] Specific behavior:
[1558] The server sends a request for scene suggestions to the AI module.
[1559] The AI module generates multiple scenes, and a list of these is provided to the device via the server.
[1560] The proposed scene is displayed to the user.
[1561] Step 3: Record a video
[1562] The user shoots a video using the device's camera function based on the selected scene. The input is the user's scene selection, and the output is the shot video. The user manually starts and stops shooting.
[1563] Specific behavior:
[1564] The user selects one of the suggested scenes.
[1565] Launch the camera app and tap the record button to start recording.
[1566] When you're done recording, tap the stop button.
[1567] Step 4: Upload video and recognize emotions
[1568] The video is uploaded from the device to the server, where the server uses an emotion engine to analyze the user's emotions in the video. The input is the video, and the output is the result of the emotion analysis.
[1569] Specific behavior:
[1570] The device uploads the captured video file to the server.
[1571] The server passes the video to the emotion engine and requests analysis.
[1572] The emotion recognition results are sent back to the server.
[1573] Step 5: Scene selection and frame extraction
[1574] The server selects the scene that most clearly expresses emotion based on the results of the emotion engine, and then extracts the optimal frame from that scene. The input is the emotion analysis result, and the output is the selected frame.
[1575] Specific behavior:
[1576] The server analyzes the emotion recognition results and determines the most appropriate scene.
[1577] Extract a specific frame from the selected scene.
[1578] Step 6: Image processing and style transfer
[1579] Once the selected frame is determined, the server generates an illustration based on that frame. This process uses machine learning algorithms such as generative adversarial networks (GANs). The input is the extracted frame, and the output is the illustration with the converted style.
[1580] Specific behavior:
[1581] The server generates the original illustration from the extracted frames.
[1582] The illustration is passed through a style conversion algorithm to convert it to the specified style.
[1583] Step 7: Preview and edit your stamp
[1584] The converted stamp image is sent to the device and a preview is displayed to the user. The user can add text and decorations to the stamp on this preview screen. The input is the style-converted illustration, and the output is the final stamp image edited by the user.
[1585] Specific behavior:
[1586] The server transmits the converted stamp image to the terminal.
[1587] The user uses the editing tools to add text and decorations.
[1588] Step 8: Create and save your sticker pack
[1589] When the user finishes editing, the device sends the final stamp pack to the server, which saves the stamp pack, generates a download link, and notifies the user. The input is the final stamp image, and the output is the download link.
[1590] Specific behavior:
[1591] The user taps the "Done" button to send the stamp pack to the server.
[1592] The server stores the stamp pack and generates a download link.
[1593] The download link will be sent to the user.
[1594] Step 9: Import to LINE
[1595] Users download the stamp pack via the provided link. The stamp pack is imported into the LINE app, allowing users to use their custom stamps on LINE. The input is the download link, and the output is the stamps in the LINE app.
[1596] Specific behavior:
[1597] The user clicks on the download link to get the stamp pack.
[1598] Import the stamp pack into the LINE app.
[1599] (Application example 2)
[1600] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1601] Conventional stamp generation systems have had the problem of making it difficult for users to easily create original stamps that reflect their own emotions. Furthermore, the technology to easily change the style of stamps and make them uploadable to virtual stores was insufficient. Furthermore, it was difficult to utilize an emotion engine to automatically select scenes that best express emotions.
[1602] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for requesting user authentication information and verifying the authentication information, means for presenting instructions for generating images to the user, means for receiving a video captured by the user based on the instructions, means for selecting a specific scene from the video and recognizing the user's emotion using an emotion engine, means for generating an original illustration based on the specific scene, means for converting the original illustration into a different style and using a generative adversarial network, means for the user to edit the illustration, and means for saving the user-edited stamp pack on the server and generating a download link. This allows for easy and effective generation of original stamps that reflect the user's emotion and uploading the style-converted stamps to a virtual store, enabling rich communication.
[1603] "Authentication information" refers to information such as a username and password that a user uses when accessing a system.
[1604] A "server" is a computer system that receives requests from users and performs the processing in response to them.
[1605] "Instructions" are information that prompts the system to perform specific operations or actions on the user.
[1606] A "video" is a media file that contains a series of image frames captured by a user.
[1607] The "emotion engine" is a software module that analyzes and recognizes user emotions from video and audio.
[1608] A "specific scene" is a selected part of a video that most clearly expresses an emotion or a specific event.
[1609] An "original illustration" is an image generated based on a specific scene extracted from a video.
[1610] "Converting to a style" refers to the process of changing the look and feel of the original illustration into a different style (e.g., cartoon, live-action, painterly, etc.).
[1611] A "generative adversarial network (GAN)" is a machine learning algorithm in which one network generates images and another network evaluates the generated images.
[1612] "Editing" refers to customization work that the user performs on the generated illustration (e.g., adding text, inserting decorations, etc.).
[1613] A "stamp pack" is a package containing multiple stamp images.
[1614] A "download link" is a URL that allows a user to download a file over the Internet.
[1615] A "virtual store" is a digital shopping platform that exists on the Internet.
[1616] System Configuration
[1617] This invention is a system that allows users to easily create original stamps that reflect their emotions, convert the stamps into different styles, and make them available in virtual stores. This system is composed of a server, terminals, and related software modules.
[1618] Processing flow
[1619] User Authentication
[1620] When a user launches an application, the server requests the user's authentication information. The user enters their username and password into the terminal and sends them to the server. The server verifies the authentication information, and if authentication is successful, the user can proceed to the next step.
[1621] Directing and filming video scenes
[1622] The server provides instructions to the device for generating images. Specifically, it works in conjunction with the emotion engine to present scenes that the user should capture (e.g., smile, surprise, sadness) to the device. The user then uses the device's camera function to capture these scenes as videos and upload them to the server.
[1623] Emotion Recognition and Scene Selection
[1624] The server receives videos uploaded by users and analyzes the user's emotions in the videos using an emotion engine. Based on this emotion analysis, a specific scene is selected from multiple scenes. The selection criterion is the scene that most clearly expresses the emotion.
[1625] Style conversion and stamp generation
[1626] The server generates a source illustration based on the specific scene selected, then converts the generated source illustration into different styles (e.g., cartoon, live action, painterly) using machine learning algorithms such as generative adversarial networks (GANs).
[1627] Editing and saving stamps
[1628] The converted stamp images are sent to the device, where the user can customize them by adding text, emojis, stickers, etc. Once editing is complete, the device sends the final stamp pack to the server, which stores it and provides a download link to the user.
[1629] Hardware and software used
[1630] Hardware: Cameras on smartphones and head-mounted displays (HMDs)
[1631] software:
[1632] EmotionRecognizer (emotion recognition engine)
[1633] StyleTransfer (style transfer algorithm, GAN)
[1634] OpenCV (camera operation and image processing library)
[1635] Requests (HTTP client library for sending API requests)
[1636] Specific examples
[1637] For example, suppose a user makes a surprised expression while virtual shopping. This moment is captured by the camera, and EmotionRecognizer recognizes it as "surprise." StyleTransfer then converts the expression into a cartoon-style image based on this emotion. Finally, the generated sticker is saved in the user's account via the virtual store's API.
[1638] Prompt Sentence Examples
[1639] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[1640] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1641] Step 1: Enter and verify your credentials
[1642] Subject: User and Server
[1643] Specific operation: The user launches an application and enters authentication information (username and password) into the device. The device sends this information to the server. The server verifies the received authentication information, and if authentication is successful, proceeds to the next step.
[1644] Input: Username and Password
[1645] Output: Authentication success flag
[1646] Step 2: Video Shooting Scene Instructions
[1647] Subject: Server and Terminal
[1648] Specific operation: The server uses the emotion engine to present the scene to be captured (e.g., smile, surprise, sadness) on the device screen. The user acts out the scene according to the instructions.
[1649] Input: Authentication success flag
[1650] Output: Instructions for the shooting scene
[1651] Step 3: Record your video
[1652] Subject: User and Device
[1653] Specific operation: The user uses the device's camera function to shoot a video based on the specified scene. The user starts and stops shooting.
[1654] Input: Instructions for the shooting scene
[1655] Output: Recorded video file
[1656] Step 4: Upload your video
[1657] Subject: Terminal and Server
[1658] Specific operation: The device uploads the captured video file to the server, which receives and stores the video file.
[1659] Input: Recorded video file
[1660] Output: Video data stored on the server
[1661] Step 5: Emotion recognition and specific scene selection
[1662] Subject: Server
[1663] Specific operation: The server analyzes the received video using the emotion engine to recognize the user's emotion in the video, and then selects the specific scene that most clearly expresses the emotion from multiple scenes.
[1664] Input: Video data stored on the server
[1665] Output: Specific scene information
[1666] Step 6: Generate the original illustration
[1667] Subject: Server
[1668] Specific operation: The server generates a source illustration based on the selected specific scene. The source illustration is an image extracted from a video frame.
[1669] Input: Specific scene information
[1670] Output: Original illustration image
[1671] Step 7: Style Transformation
[1672] Subject: Server
[1673] Specific operation: The server converts the generated original illustration into different styles (e.g., cartoon, live-action, painting) using a generative adversarial network (GAN).
[1674] Input: Original illustration image
[1675] Output: Style-converted illustration
[1676] Step 8: Editing the stamp
[1677] Subject: User and Device
[1678] Specific operation: The converted stamp image is sent to the device, and the user can edit it. Users can add text, emojis, stickers, etc.
[1679] Input: Style-converted illustration
[1680] Output: Edited stamp image
[1681] Step 9: Save and share your sticker pack
[1682] Subject: Terminal and Server
[1683] Specific operation: The device sends the edited stamp pack to the server, which stores it and provides a download link to the user.
[1684] Input: Edited stamp image
[1685] Output: Saved sticker pack and download link
[1686] Prompt Sentence Examples
[1687] Take a photo of your facial expressions, such as "smile," "surprise," or "sadness," and generate a sticker based on that expression. The sticker will then be converted into a cartoon style and finally uploaded to your virtual store.
[1688] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1689] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1690] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1691] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1692] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1693] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1694] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1695] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1696] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1697] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1698] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1699] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1700] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1701] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1702] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1703] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1704] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1705] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1706] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1707] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1708] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1709] The following is further disclosed regarding the above embodiment.
[1710] (Claim 1)
[1711] means for presenting instructions to a user for generating an image;
[1712] means for receiving the video captured by the user based on the instruction;
[1713] means for selecting a specific scene from the video;
[1714] a means for generating an original illustration based on the specific scene;
[1715] means for converting the original illustration into a different style;
[1716] The system provides a means for the user to edit the illustration.
[1717] (Claim 2)
[1718] The system according to claim 1, characterized in that the original illustration is converted into a style such as a cartoon style, a live-action style, or a painting style.
[1719] (Claim 3)
[1720] 2. The system of claim 1, wherein the selection of the particular scene includes an algorithm for selecting the best frame from a plurality of scenes.
[1721] (Claim 4)
[1722] 2. The system according to claim 1, further comprising means for providing an interface for a user to add text and decorations to the original illustration.
[1723] (Claim 5)
[1724] The system of claim 1, characterized in that an AI module is used to generate and convert the original illustration into a different style.
[1725] "Example 1"
[1726] (Claim 1)
[1727] It is a system that allows users to easily create their own original stamps.
[1728] a means for requesting authentication information from a user and authenticating the user by referencing a database;
[1729] means for generating and transmitting instructions for a scene suitable for the stamp to a device using an AI model;
[1730] means for receiving from a terminal the video captured by the user based on the instruction;
[1731] A means to analyze uploaded videos and select specific scenes,
[1732] A method for generating original illustrations based on specific scenes and converting the generated original illustrations into different styles.
[1733] A means for transmitting the converted illustration to a terminal and allowing a user to edit the illustration;
[1734] A way to save edited stamp packs and issue download links
[1735] Including system.
[1736] (Claim 2)
[1737] It features technology that converts original illustrations into manga, live-action, painting, and other styles.
[1738] 10. The system of claim 1.
[1739] (Claim 3)
[1740] It features an AI algorithm that selects the best frame from multiple scenes when selecting a specific scene.
[1741] 10. The system of claim 1.
[1742] "Application Example 1"
[1743] (Claim 1)
[1744] means for presenting instructions to a user for generating an image;
[1745] means for receiving the video captured by the user based on the instruction;
[1746] means for selecting a specific scene from the video;
[1747] a means for generating an original illustration based on the specific scene;
[1748] means for converting the original illustration into a different style;
[1749] A means for a user to edit the illustration;
[1750] means for automatically generating stamps based on the condition of parts or products detected on the manufacturing line;
[1751] A system that provides a means for linking said stamps to a quality control system.
[1752] (Claim 2)
[1753] The system according to claim 1, characterized in that the original illustration is converted into a style such as a cartoon style, a live-action style, or a painting style.
[1754] (Claim 3)
[1755] 2. The system of claim 1, wherein the selection of the particular scene includes an algorithm for selecting the best frame from a plurality of scenes.
[1756] "Example 2: Combining Emotion Engines"
[1757] (Claim 1)
[1758] means for analyzing a user's emotions using an emotion recognition engine;
[1759] A means for proposing video shooting scenes to be presented to a user;
[1760] A means for receiving a video that the user has shot based on the suggestion;
[1761] means for selecting a specific scene from the video based on emotion;
[1762] a means for generating an original illustration based on the specific scene;
[1763] A means for converting the original illustration into a different style using a generative adversarial network;
[1764] The system provides a means for the user to edit the illustration.
[1765] (Claim 2)
[1766] The system according to claim 1, characterized in that the original illustration is converted into a style such as a cartoon style, a live-action style, or a painting style.
[1767] (Claim 3)
[1768] 2. The system of claim 1, wherein the selection of the particular scene includes an algorithm for selecting the best frame from a plurality of scenes.
[1769] "Application example 2 when combining emotion engines"
[1770] (Claim 1)
[1771] means for requesting user authentication information and verifying the authentication information;
[1772] means for presenting instructions to a user for generating an image;
[1773] means for receiving the video captured by the user based on the instruction;
[1774] a means for selecting a specific scene from the video and recognizing a user's emotion using an emotion engine;
[1775] a means for generating an original illustration based on the specific scene;
[1776] A means for transforming the original illustration into a different style and utilizing a generative adversarial network;
[1777] A means for a user to edit the illustration;
[1778] A system that provides a means for users to save edited stamp packs on a server and generate download links.
[1779] (Claim 2)
[1780] The system according to claim 1, characterized in that the original illustration can be converted into a style such as a cartoon style, a live-action style, or a painting style, and can be uploaded as a sticker pack for a virtual store.
[1781] (Claim 3)
[1782] 10. The system of claim 1, further comprising an algorithm for selecting the best frames from multiple scenes and using an emotion engine to select the scene that most clearly reflects the user's emotion. [Explanation of symbols]
[1783] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for presenting instructions to a user for generating an image; means for receiving the video captured by the user based on the instruction; means for selecting a specific scene from the video; a means for generating an original illustration based on the specific scene; means for converting the original illustration into a different style; The system provides a means for the user to edit the illustration.
2. 2. The system according to claim 1, wherein the original illustration is converted into a style such as a cartoon style, a live-action style, or a painting style.
3. 2. The system of claim 1, wherein said selection of said particular scene includes an algorithm for selecting the best frame from a plurality of scenes.
4. 2. The system according to claim 1, further comprising means for providing an interface for a user to add text and decorations to the original illustration.
5. The system of claim 1, wherein an AI module is used to generate and convert the original illustration into different styles.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A